# ESC-EXP-27：Shared Semantic Residual Factorization

**系列：** Extensional Structural Convergence — Experimental Phase  
**文件編號：** ESC-EXP-27  
**版本：** v0.1  
**日期：** 2026-09-23  
**前置：** ESC-00 ～ ESC-06、ESC-EXP-00 ～ ESC-EXP-26  
**狀態：** Shared Latent Factor / Overlapping Task Provenance Experiment

**作者：** Neo.K  
**機構：** EveMissLab／一言諾科技有限公司  

---

## 摘要

ESC-EXP-26 建立了：

$$
\boxed{
R_J(H)
=
H(J\mid H)
}
$$

並證明：

> future task若只需要 underlying state 的一部分 distinction，就不必每次解碼完整 universal provenance。

但 EXP-26 deliberately 使用兩個完全互補、互不重疊的 selective task layers：

$$
MID,
\quad
OUTER.
$$

它們的 provenance bits沒有重複，因此：

$$
R_{\mathrm{MID}}
+
R_{\mathrm{OUTER}}
=
R_{\mathrm{FULL}}.
$$

這是一個乾淨 baseline，但真實 task families通常不是這樣。

多個 tasks 很可能：

- 各自需要一些 private distinctions；
- 同時重複依賴一批 shared latent distinctions。

如果我們直接：

> 每個 task各存一份 task-specific provenance code，

那 shared latent bits 就會被重複保存。

本輪因此故意建立兩個 overlap tasks：

$$
J_A
$$

與：

$$
J_B.
$$

Hot representation仍是 H4。

Hot 已知：

$$
(T,D),
$$

隱藏：

$$
G.
$$

Task A 需要 $G$ 的 Hot regions：

$$
\boxed{
\{00,01,10\}
}
$$

Task B 需要：

$$
\boxed{
\{01,10,11\}.
}
$$

所以：

### A-only region

$$
\{00\}.
$$

### Shared region

$$
\boxed{
\{01,10\}.
}
$$

### B-only region

$$
\{11\}.
$$

兩個 tasks 的 union：

$$
\boxed{
A\cup B
=
\{00,01,10,11\}
}
$$

剛好覆蓋完整 hidden residual。

因此本輪可以比較兩種 coding architecture。

## Naive Task-Wise Coding

分別保存：

$$
Z_A
$$

與：

$$
Z_B.
$$

storage：

$$
\boxed{
H(J_A\mid H)
+
H(J_B\mid H).
}
$$

Shared distinctions被存兩次。

## Shared-Factor Coding

改存：

$$
\boxed{
Z_{A\text{-only}},
\quad
Z_{\mathrm{shared}},
\quad
Z_{B\text{-only}}.
}
$$

Task A讀：

$$
Z_{A\text{-only}}
+
Z_{\mathrm{shared}}.
$$

Task B讀：

$$
Z_{\mathrm{shared}}
+
Z_{B\text{-only}}.
$$

Joint A+B讀三層。

因此理想 storage：

$$
\boxed{
H(
Z_{A\text{-only}},
Z_{\mathrm{shared}},
Z_{B\text{-only}}
\mid H
).
}
$$

本輪主要結果：

### Uniform Prior

Task A：

$$
0.75
\text{ bit}.
$$

Task B：

$$
0.75
\text{ bit}.
$$

Joint：

$$
1
\text{ bit}.
$$

所以 naive task-wise storage：

$$
1.5.
$$

factorized storage：

$$
1.
$$

省：

$$
\boxed{
0.5\text{ bit/state}
}
$$

即：

$$
\boxed{
33.33\%.
}
$$

這 $0.5$ bit正好等於 shared factor本身：

$$
\boxed{
H(Z_{\mathrm{shared}}\mid H)
=
0.5.
}
$$

### Skewed Prior

Task A：

$$
0.648390484031.
$$

Task B：

$$
0.441549035869.
$$

Joint：

$$
0.811278124459.
$$

所以 naive storage：

$$
1.089939519900.
$$

factorized storage：

$$
0.811278124459.
$$

省：

$$
\boxed{
0.278661395441
\text{ bits/state}.
}
$$

即：

$$
\boxed{
25.57\%.
}
$$

這個 savings 又精確等於 shared factor：

$$
\boxed{
0.278661395441.
}
$$

因此：

$$
\boxed{
\text{conditional task redundancy}
=
\text{shared latent factor information}
}
$$

在本 finite overlap geometry 中完全成立。

---

# 1. Runtime

EXP-27 regression：

```text
6 passed
```

測試內容：

- uniform overlap geometry 精確為：
  $$
  R_A=R_B=0.75,
  \quad
  R_S=0.5,
  \quad
  R_{AB}=1;
  $$
- shared factor等於 conditional task redundancy；
- A+B joint tasks精確恢復 full residual；
- factorized storage小於 naive per-task storage；
- factorized retrieval同時優於 monolithic與 naive task-wise read；
- uniform / skewed shared factor都非零。

---

# 2. Hot Geometry

Hot H4：

$$
\{
\{\varnothing,G\},
\{D,DG\},
\{T,TG\},
\{TD,TDG\}
\}.
$$

四個 Hot regions可用：

$$
(T,D)
$$

標記：

$$
00,\quad01,\quad10,\quad11.
$$

每個 region內只剩：

$$
G
$$

尚未決定。

---

# 3. Overlap Tasks

Task A：

$$
\boxed{
J_A
=
G
\quad
\text{on }
\{00,01,10\}.
}
$$

Task B：

$$
\boxed{
J_B
=
G
\quad
\text{on }
\{01,10,11\}.
}
$$

所以：

$$
A\cap B
=
\boxed{
\{01,10\}
}.
$$

---

# 4. Factor Regions

定義三個 semantic factors：

$$
Z_A
$$

只服務：

$$
00.
$$

$$
Z_S
$$

服務 shared：

$$
01,10.
$$

$$
Z_B
$$

只服務：

$$
11.
$$

所以：

$$
\boxed{
J_A
\Longleftrightarrow
(Z_A,Z_S)
}
$$

在給定 Hot 後成立。

同理：

$$
\boxed{
J_B
\Longleftrightarrow
(Z_S,Z_B).
}
$$

而：

$$
\boxed{
(J_A,J_B)
\Longleftrightarrow
(Z_A,Z_S,Z_B).
}
$$

---

# 5. Joint Tasks Recover Full Hidden State

因：

$$
A\cup B
=
\{00,01,10,11\},
$$

joint task family知道每個 Hot region中的：

$$
G.
$$

所以：

$$
\boxed{
H(
J_A,J_B
\mid H
)
=
H(
G
\mid H
).
}
$$

也就是：

$$
\boxed{
R_{AB}
=
R_{\mathrm{FULL}}.
}
$$

Uniform / skewed都精確成立。

---

# 6. Conditional Redundancy

對兩 tasks：

$$
J_A,J_B,
$$

定義：

$$
\boxed{
\mathcal R_{AB\mid H}
=
H(J_A\mid H)
+
H(J_B\mid H)
-
H(J_A,J_B\mid H).
}
$$

如果：

$$
\mathcal R>0,
$$

代表兩個 task-specific code中有 shared information。

---

# 7. Uniform Conditional Redundancy

Uniform：

$$
H(J_A\mid H)
=
0.75,
$$

$$
H(J_B\mid H)
=
0.75,
$$

$$
H(J_A,J_B\mid H)
=
1.
$$

所以：

$$
\boxed{
\mathcal R_{AB\mid H}
=
0.5.
}
$$

---

# 8. Shared Factor Information

Shared regions：

$$
01,
10.
$$

每個 Hot region probability：

$$
0.25,
$$

而其 hidden $G$ entropy為：

$$
1.
$$

所以 shared residual：

$$
\boxed{
R_S
=
0.25+0.25
=
0.5.
}
$$

因此：

$$
\boxed{
\mathcal R_{AB\mid H}
=
R_S.
}
$$

---

# 9. Skewed Conditional Redundancy

Skewed：

$$
R_A
=
0.648390484031,
$$

$$
R_B
=
0.441549035869,
$$

$$
R_{AB}
=
0.811278124459.
$$

所以：

$$
\boxed{
\mathcal R_{AB\mid H}
=
0.278661395441.
}
$$

---

# 10. Skewed Shared Factor

直接計算 shared regions：

$$
01,
10
$$

所對應的 hidden conditional entropy contribution：

$$
\boxed{
R_S
=
0.278661395441.
}
$$

所以再次：

$$
\boxed{
\mathcal R_{AB\mid H}
=
R_S.
}
$$

---

# 11. Task Boundary vs Latent-Factor Boundary

Naive task-wise coding以：

$$
\boxed{
\text{task}
}
$$

當 coding boundary。

所以：

$$
Z_A
$$

與：

$$
Z_B
$$

各自包含 shared residual。

Factorized coding改以：

$$
\boxed{
\text{latent distinction region}
}
$$

當 boundary。

因此 shared部分只存一次。

這是本輪最核心的 coding architecture差異。

---

# 12. Uniform Storage Comparison

Naive：

$$
R_A+R_B
=
0.75+0.75
=
\boxed{
1.5.
}
$$

Factorized：

$$
R_{A\text{-only}}
+
R_S
+
R_{B\text{-only}}
$$

$$
=
0.25+0.5+0.25
=
\boxed{
1.
}
$$

所以：

$$
\boxed{
33.33\%
}
$$

storage reduction。

---

# 13. Skewed Storage Comparison

A-only：

$$
\boxed{
0.369729088590.
}
$$

Shared：

$$
\boxed{
0.278661395441.
}
$$

B-only：

$$
\boxed{
0.162887640428.
}
$$

總：

$$
\boxed{
0.811278124459.
}
$$

naive：

$$
0.648390484031
+
0.441549035869
=
1.089939519900.
$$

省：

$$
\boxed{
0.278661395441.
}
$$

---

# 14. Storage Savings 正好是 Shared Factor

一般在本 separable overlap construction：

$$
R_A
=
R_{A\text{-only}}
+
R_S,
$$

$$
R_B
=
R_S
+
R_{B\text{-only}},
$$

而：

$$
R_{AB}
=
R_{A\text{-only}}
+
R_S
+
R_{B\text{-only}}.
$$

所以：

$$
R_A+R_B-R_{AB}
=
\boxed{
R_S.
}
$$

即：

$$
\boxed{
\text{duplicated storage}
=
\text{shared latent information}.
}
$$

---

# 15. Huffman Storage

Uniform one-shot Huffman：

naive：

$$
1.5.
$$

factorized：

$$
1.
$$

仍省：

$$
0.5.
$$

Skewed：

naive：

$$
\boxed{
1.343484419263.
}
$$

factorized：

$$
\boxed{
1.
}
$$

省：

$$
\boxed{
0.343484419263.
}
$$

---

# 16. 為什麼 Huffman Savings 比 Shannon Shared Bit 更大？

Skewed Shannon shared factor：

$$
0.278661395441.
$$

Huffman storage savings：

$$
0.343484419263.
$$

因 task-wise one-shot coding不只重複 shared information，

還各自支付 prefix quantization overhead。

factorization同時消掉：

- semantic duplication；
- 部分 coding overhead duplication。

---

# 17. Retrieval Workload

本輪 workload：

$$
P(A)=0.45,
$$

$$
P(B)=0.45,
$$

$$
P(AB)=0.10.
$$

所有 queries都需要某種 hidden residual。

---

# 18. Universal Monolithic Read

如果每次都讀完整 universal residual：

Uniform：

$$
\boxed{
1\text{ bit/query}.
}
$$

Skewed：

$$
\boxed{
0.811278124459.
}
$$

---

# 19. Naive Task-Wise Read

A query讀：

$$
R_A.
$$

B query讀：

$$
R_B.
$$

AB query因兩份 task code彼此獨立，要讀：

$$
R_A+R_B.
$$

Uniform：

$$
\boxed{
0.825.
}
$$

Skewed：

$$
\boxed{
0.599466735945.
}
$$

---

# 20. Factorized Read

A query讀：

$$
Z_A+Z_S.
$$

B query讀：

$$
Z_S+Z_B.
$$

AB query讀：

$$
Z_A+Z_S+Z_B.
$$

Uniform：

$$
\boxed{
0.775.
}
$$

Skewed：

$$
\boxed{
0.571600596401.
}
$$

---

# 21. Factorized vs Monolithic

Uniform：

$$
1
\rightarrow
0.775.
$$

下降：

$$
\boxed{
22.5\%.
}
$$

Skewed：

$$
0.811278
\rightarrow
0.571601.
$$

下降：

$$
\boxed{
29.54\%.
}
$$

---

# 22. Factorized vs Naive Task-Wise

Uniform：

$$
0.825
\rightarrow
0.775.
$$

下降：

$$
\boxed{
6.06\%.
}
$$

Skewed：

$$
0.599467
\rightarrow
0.571601.
$$

下降：

$$
\boxed{
4.65\%.
}
$$

這個差距主要來自：

$$
P(AB)=0.10
$$

的 joint queries不再重複讀 shared factor。

若 joint-task probability更高，這個收益會更大。

---

# 23. Storage Savings 與 Retrieval Savings 是不同來源

Storage savings：

> Shared factor不用存兩次。

Retrieval savings：

> Joint query不用讀 Shared factor兩次。

所以：

$$
\boxed{
\text{shared factorization improves both write/storage and read/retrieval sides}.
}
$$

---

# 24. Shared Factor Fraction

Uniform full residual：

$$
1.
$$

shared：

$$
0.5.
$$

所以：

$$
\boxed{
50\%
}
$$

的 universal residual同時被 A 與 B 使用。

Skewed：

$$
\frac{
0.278661395441
}{
0.811278124459
}
=
\boxed{
34.35\%.
}
$$

所以相同 geometric overlap，在不同 prior下有不同 information overlap。

---

# 25. Geometry Overlap ≠ Information Overlap

A/B overlap regions占：

$$
2/4
=
50\%.
$$

但 skewed information overlap只有：

$$
34.35\%.
$$

因此：

$$
\boxed{
\text{overlap in state-space geometry}
\neq
\text{overlap in information mass}.
}
$$

這和 EXP-25 的 block-count vs bits差異完全一致。

---

# 26. Shared-Latent Factor Value

可以定義：

$$
\boxed{
V_S
=
H(J_A\mid H)
+
H(J_B\mid H)
-
H(J_A,J_B\mid H).
}
$$

它就是：

> task-wise provenance storage中最少可被共享／去重的 conditional information。

本輪：

Uniform：

$$
V_S=0.5.
$$

Skewed：

$$
V_S=0.278661395441.
$$

---

# 27. Shared Semantic Factor

因此可以把：

$$
Z_S
$$

定義為：

$$
\boxed{
\text{能同時服務多個 future tasks 的最小共同 residual component}.
}
$$

本 finite construction中它具有非常直接的 latent-region realization。

---

# 28. Task Factorization vs Common Information

一般情況下，shared factor未必能像本輪這麼乾淨地對應 disjoint state regions。

更一般可能需要尋找：

$$
Z_S
$$

使：

$$
H(J_A\mid H,Z_A,Z_S)=0,
$$

$$
H(J_B\mid H,Z_B,Z_S)=0,
$$

並最小化：

$$
H(Z_A,Z_S,Z_B\mid H).
$$

這開始接近 conditional common-information / multiterminal coding 類型的問題。

本輪只是最簡單 finite baseline。

---

# 29. Shared Factor 不能只靠 Task 名稱決定

如果兩個 tasks：

- 名稱不同；
- output格式不同；

但 underlying residual需求高度重疊，

仍應共用同一 latent layer。

反過來，同一 task family中的不同 contexts也可能需要完全不同 residual factors。

所以：

$$
\boxed{
\text{code factorization should follow latent dependency,
not API/task labels}.
}
$$

---

# 30. Semantic Residual Factor Graph

可以把 future tasks與 latent factors建成 bipartite graph：

$$
\boxed{
G_{\mathrm{task-factor}}
=
(
\mathcal J,
\mathcal Z,
E
).
}
$$

本輪：

$$
J_A
\leftrightarrow
\{
Z_A,Z_S
\},
$$

$$
J_B
\leftrightarrow
\{
Z_S,Z_B
\}.
$$

這已經是一個最小 shared-factor graph。

---

# 31. Factor Degree

Shared factor：

$$
Z_S
$$

degree：

$$
2.
$$

private factors：

$$
Z_A,Z_B
$$

degree：

$$
1.
$$

未來如果有大量 tasks，

高-degree latent factors可能值得：

- Hot-cache；
- 優先 replicate；
- 強保護；
- 高可用性 storage。

因此 factor degree會成為新的 tier-placement signal。

---

# 32. Frequency × Sharedness

一個 residual factor的 operational value不只取決於：

$$
H(Z\mid H).
$$

還取決於：

$$
\boxed{
\sum_{J:
Z\in I(J)}
P(J).
}
$$

也就是它被多少／多常 tasks使用。

所以 Hot promotion policy可以由：

$$
\boxed{
\text{information bits}
\times
\text{task reuse frequency}
}
$$

推動。

---

# 33. Shared Factor 優先 Hot 的可能性

如果：

$$
Z_S
$$

被 A / B / C / D 多個 tasks共同依賴，

即使它本身 information size不大，

把它放 Hot可能非常划算。

private factors則留 Cold。

所以 multi-tier coding的 natural architecture可能是：

$$
\boxed{
\text{Hot shared core}
+
\text{Cold private residuals}.
}
$$

---

# 34. 這與 Representation Core 很接近

如果跨大量 tasks都反覆使用同一 latent residual，

那這個 residual開始具有：

$$
\boxed{
\text{representation core}
}
$$

性質。

所以 ESC可以從：

> shared provenance factor

進一步發展成：

> minimal common latent substrate across task families。

---

# 35. Factorized Task Lattice

本輪 factors：

$$
Z_A,
Z_S,
Z_B.
$$

可形成 layer subsets。

重要 nodes：

$$
E_0,
$$

$$
E_A,
$$

$$
E_S,
$$

$$
E_B,
$$

$$
E_{AS},
$$

$$
E_{SB},
$$

$$
E_{ASB}.
$$

其中：

$$
E_{AS}
$$

足以回答 Task A，

$$
E_{SB}
$$

足以回答 Task B，

$$
E_{ASB}
$$

足以回答 joint/full task。

---

# 36. Task Lattice 不再只是 $B_2$

EXP-26 是兩個完全互補 task layers，

所以 recoverability requirement lattice是：

$$
B_2.
$$

EXP-27 加入 shared factor後，

底層 factor subset lattice更接近：

$$
\boxed{
B_3
}
$$

但只有部分 nodes對應完整 task contracts。

所以：

$$
\boxed{
\text{factor lattice}
\neq
\text{task lattice}.
}
$$

這是很重要的新區分。

---

# 37. Task Quotient of Factor Lattice

多個 factor subsets可能對目前 task set具有相同 operational utility。

所以可以再做 quotient：

$$
\boxed{
\mathcal P(\mathcal Z)
\rightarrow
\text{Task Capability Classes}.
}
$$

這其實又回到 ESC effect quotient / contextual equivalence。

---

# 38. Representation Factorization 又回到 ESC 主題

我們現在有：

- underlying state；
- Hot projection；
- task outputs；
- residual factors；
- factor subsets；
- task capability equivalence。

不同 representations再次需要比較：

$$
\boxed{
\text{what distinctions they preserve,
what tasks they support,
what future composition they enable}.
}
$$

這幾乎正是 ESC 最早的共同底空間命題。

---

# 39. Shared Factor Storage Lower Bound

本輪 separable overlap case：

$$
\boxed{
R_S
=
\mathcal R_{AB\mid H}.
}
$$

因此任何不重複保存 task-specific codes的 factorization，至少需要保存這一份 shared information。

---

# 40. Naive Duplication Tax

定義：

$$
\boxed{
T_{\mathrm{dup}}
=
R_A+R_B-R_{AB}.
}
$$

本輪：

$$
\boxed{
T_{\mathrm{dup}}
=
R_S.
}
$$

它可以被解讀成：

> 使用 task boundaries 而不是 latent-factor boundaries 所支付的 storage tax。

---

# 41. Uniform Duplication Tax

$$
\boxed{
T_{\mathrm{dup}}
=
0.5\text{ bit/state}.
}
$$

相對 naive：

$$
1.5,
$$

比例：

$$
\boxed{
33.33\%.
}
$$

---

# 42. Skewed Duplication Tax

$$
\boxed{
0.278661395441
\text{ bits/state}.
}
$$

相對：

$$
1.089939519900,
$$

比例：

$$
\boxed{
25.57\%.
}
$$

---

# 43. One-Shot Coding Duplication Tax

Skewed Huffman：

naive：

$$
1.343484419263.
$$

factorized：

$$
1.
$$

所以：

$$
\boxed{
T_{\mathrm{dup,Huff}}
=
0.343484419263.
}
$$

這比 Shannon semantic redundancy更高，因還包含 duplicated coding overhead。

---

# 44. Shared Factorization Principle

因此可以提出：

$$
\boxed{
\text{Shared Semantic Residual Factorization Principle}.
}
$$

若多個 future tasks對 Hot representation存在重疊 residual dependency，

則 provenance code不應簡單按 task各自獨立儲存。

應優先尋找：

$$
\boxed{
\text{private factors}
+
\text{shared factors}
}
$$

使 joint task code避免重複保存 shared information。

---

# 45. 本輪錨點

$$
\boxed{
\textbf{ESC-EXP-27.A}
\quad
\mathcal R_{AB\mid H}
=
H(J_A\mid H)
+
H(J_B\mid H)
-
H(J_A,J_B\mid H).
}
$$

$$
\boxed{
\textbf{ESC-EXP-27.B}
\quad
\mathcal R_{AB\mid H}
=
R_S
}
$$

在本 finite overlap construction精確成立。

$$
\boxed{
\textbf{ESC-EXP-27.C}
\quad
\text{factorized storage相對 naive task-wise storage降低 }33.33\%\text{（uniform）與 }25.57\%\text{（skewed）。}
}
$$

$$
\boxed{
\textbf{ESC-EXP-27.D}
\quad
\text{factorized expected retrieval同時低於 universal monolithic 與 naive task-wise retrieval。}
}
$$

$$
\boxed{
\textbf{ESC-EXP-27.E}
\quad
\text{code boundary應追隨 shared latent dependency，而不應只追隨 task/API boundary。}
}
$$

---

# 46. 從 EXP-26 到 EXP-27

EXP-26：

$$
\boxed{
\text{Which provenance layer does each task need?}
}
$$

EXP-27：

$$
\boxed{
\text{Which provenance layers are actually shared by multiple tasks?}
}
$$

所以：

$$
\boxed{
\text{Task-Adaptive Coding}
\rightarrow
\text{Shared Latent Residual Factorization}.
}
$$

---

# 47. 下一輪：ESC-EXP-28

EXP-27 的 shared factor還是人工設計：

$$
A\text{-only},
\quad
\text{Shared},
\quad
B\text{-only}.
$$

真正下一步應該讓 runtime自己找 factorization。

也就是：

$$
\boxed{
\text{Automatic Semantic Residual Factor Discovery}.
}
$$

給定：

- Hot representation $H$ ；
- 多個 future tasks：
  $$
  J_1,\dots,J_m;
  $$
- task demand distribution；
- storage / read cost；

自動搜尋 latent factors：

$$
Z_1,\dots,Z_k
$$

以及每個 task要讀的 factor subset：

$$
I(J_i).
$$

目標：

$$
\boxed{
\min
\left[
\text{total stored provenance}
+
\lambda
\text{expected task read cost}
\right]
}
$$

subject to：

$$
\boxed{
H(
J_i
\mid
H,
Z_{I(J_i)}
)
=
0
\quad
\forall i.
}
$$

真正問題變成：

> **AI 能不能自己從 task family 中發現「哪些 hidden distinctions 應該共用一層 residual factor」？**

這會把人工 semantic decomposition推進到真正可自動編譯的：

$$
\boxed{
\text{provenance factor compiler}.
}
$$
