# ESC-EXP-17：Approximate Contextual Minimization

**系列：** Extensional Structural Convergence — Experimental Phase  
**文件編號：** ESC-EXP-17  
**版本：** v0.1  
**日期：** 2026-09-23  
**前置：** ESC-00 ～ ESC-06、ESC-EXP-00 ～ ESC-EXP-16  
**狀態：** Error-Bounded Contextual Compression Experiment / Canonical UTF-8 Source

**作者：** Neo.K  
**機構：** EveMissLab／一言諾科技有限公司  

---

## 摘要

ESC-EXP-16 已解出 exact contextual minimization：

$$
\boxed{
\text{在不丟失任何 admissible future epistemic behavior 的前提下，
求最大的安全 quotient。}
}
$$

exact case使用：

$$
\operatorname{Eff}_1
=
\operatorname{Eff}_2.
$$

但真實 AI、world model、agent memory 與 approximate representation 很少要求精確相等。

更自然的是：

$$
\boxed{
d(
\operatorname{Eff}_1,
\operatorname{Eff}_2
)
\le
\varepsilon.
}
$$

本輪因此建立：

$$
\boxed{
\varepsilon\text{-Contextual Minimization}.
}
$$

但一進入：

$$
\varepsilon>0,
$$

exact quotient 的一個重要性質立即消失：

> **距離不超過 $\varepsilon$ 通常不是等價關係。**

因此 approximate compression 一般不再是：

$$
S/\sim
$$

形式的唯一 quotient，

而是：

$$
\boxed{
\text{在 error budget 下尋找 minimum-cardinality safe partition。}
}
$$

本輪對 8 個 mechanism states 的全部：

$$
\boxed{
B_8=4140
}
$$

個 set partitions 做完整枚舉。

最重要的 finite 結果：

1. approximate safe compression 隨 $\varepsilon$ 單調增加；
2. future horizon 增加時 safe compression 單調下降；
3. $\varepsilon$ -closeness 確實出現非傳遞反例；
4. approximate optimum 可以不唯一；
5. skewed prior 下，小到約 $0.00597$ bit 的 current error，在 future composition 下可放大到約 $0.86877$ bit；
6. 最大 finite amplification ratio 約：
   $$
   \boxed{
   145.48\times.
   }
   $$
7. uniform synergy regime 中有：
   $$
   \boxed{
   15
   }
   $$
   組 mechanism pairs 具有：
   $$
   d_0=0
   $$
   但：
   $$
   d_{\mathrm{future}}>0.
   $$

所以：

$$
\boxed{
\text{small current approximation error}
\not\Rightarrow
\text{small future compositional error}.
}
$$

---

# 1. Runtime

EXP-17 新增 regression：

```text
5 passed
```

測試內容：

- 8 mechanisms 的 set partition 總數為：
  $$
  4140;
  $$
- minimum class count 對 $\varepsilon$ 單調不增；
- required class count 對 future horizon 單調不減；
- uniform prior 存在 zero-now / positive-future activation；
- skewed approximate minimization 存在多個 optimal partitions。

---

# 2. Approximate Contextual Distance

對 mechanism states：

$$
s,t,
$$

target family：

$$
\mathcal J,
$$

future action alphabet：

$$
\Sigma,
$$

future horizon：

$$
h,
$$

定義 contexts：

$$
\mathcal C_{\Sigma,h}
=
\{
c:
|c|\le h
\}.
$$

在目前 monotone-union mechanism 中，context 可以用 primitive-effect subset 表示。

定義：

$$
\boxed{
d_h(s,t)
=
\max_{
c\in\mathcal C_{\Sigma,h}
}
\left\|
\operatorname{Eff}(s\vee c)
-
\operatorname{Eff}(t\vee c)
\right\|_\infty.
}
$$

這是把所有允許 future contexts 下的最壞 epistemic difference 聚合成一個 contextual pseudometric。

---

# 3. Horizon Monotonicity

因為：

$$
\mathcal C_{\Sigma,h}
\subseteq
\mathcal C_{\Sigma,h+1},
$$

所以：

$$
\boxed{
d_{h+1}(s,t)
\ge
d_h(s,t).
}
$$

也就是：

> 看得越遠，兩個 mechanisms 只能變得更容易區分，不會更難區分。

這是 approximate contextual compression 的基本 monotonicity。

---

# 4. $\varepsilon$ -Safe Block

對 mechanism block：

$$
B\subseteq S,
$$

定義 contextual diameter：

$$
\boxed{
\operatorname{diam}_{h}(B)
=
\max_{s,t\in B}
d_h(s,t).
}
$$

若：

$$
\boxed{
\operatorname{diam}_h(B)
\le
\varepsilon,
}
$$

則稱：

$$
B
$$

為：

$$
(\varepsilon,h)\text{-safe block}.
$$

---

# 5. Approximate Safe Partition

partition：

$$
\mathcal P
=
\{B_1,\dots,B_k\}
$$

合法若：

$$
\forall B_i\in\mathcal P,
\quad
\operatorname{diam}_h(B_i)
\le
\varepsilon.
$$

最佳化：

$$
\boxed{
k^\ast(\varepsilon,h)
=
\min_{\mathcal P}
|\mathcal P|
}
$$

subject to上述 diameter constraint。

定義 compression：

$$
\boxed{
C^\ast(\varepsilon,h)
=
1-
\frac{k^\ast(\varepsilon,h)}{|S|}.
}
$$

---

# 6. Exact Case 是 Approximate Case 的特例

若：

$$
\varepsilon=0,
$$

則同 block 中任意兩個 states 必須：

$$
d_h(s,t)=0.
$$

zero-distance relation 是 equivalence relation。

因此 exact case重新得到：

$$
\boxed{
\text{唯一的 coarsest behavioral quotient}.
}
$$

但：

$$
\varepsilon>0
$$

後一般不再成立。

---

# 7. Pairwise $\varepsilon$ -Closeness 不傳遞

skewed prior、full future horizon 下，取：

$$
\boxed{
\varepsilon
=
0.868767195327.
}
$$

找到反例：

$$
x=\varnothing,
$$

$$
y=T,
$$

$$
z=TG.
$$

距離：

$$
d(x,y)
\approx
0.868767
\le
\varepsilon,
$$

$$
d(y,z)
\approx
0.811278
\le
\varepsilon,
$$

但：

$$
\boxed{
d(x,z)
\approx
0.968520
>
\varepsilon.
}
$$

因此：

$$
x\sim_\varepsilon y,
$$

$$
y\sim_\varepsilon z,
$$

但：

$$
x\not\sim_\varepsilon z.
$$

所以：

$$
\boxed{
\sim_\varepsilon
\text{ 一般不是 equivalence relation}.
}
$$

---

# 8. Approximate Compression 不是普通 Quotient

因為 threshold relation 不傳遞，

不能直接寫：

$$
S/\sim_\varepsilon.
$$

真正需要的是：

$$
\boxed{
\text{partition optimization}.
}
$$

等價地，可以建立 compatibility graph：

$$
G_\varepsilon,
$$

其中：

$$
(s,t)\in E
\iff
d_h(s,t)\le\varepsilon.
$$

每個合法 compression block 必須是：

$$
G_\varepsilon
$$

中的 clique。

因此 minimum-class approximate compression 是：

$$
\boxed{
\text{minimum clique cover}
}
$$

問題，

亦等價於 complement graph 的 coloring problem。

所以 finite approximate minimization 的計算複雜度已經和 exact partition refinement 明顯不同。

---

# 9. 為什麼本輪可以完整枚舉？

mechanism states只有：

$$
|S|=8.
$$

set partition總數：

$$
\boxed{
B_8=4140.
}
$$

因此本輪直接完整枚舉所有 partitions，

而不是使用 heuristic clustering。

這保證每一個報告的：

$$
k^\ast(\varepsilon,h)
$$

都是 exact finite optimum。

---

# 10. Uniform Prior

uniform prior 保留 EXP-15 最乾淨的 pure synergy。

full future horizon 的 contextual distances 只產生兩級：

$$
0,
\quad
1.
$$

所以：

$$
\varepsilon=0
$$

時：

$$
k^\ast=8,
$$

$$
C^\ast=0.
$$

一旦：

$$
\varepsilon=1,
$$

全部 mechanisms 都可以放進同一 block：

$$
k^\ast=1,
$$

$$
\boxed{
C^\ast=87.5\%.
}
$$

uniform prior 因此不是 graded approximation 的好例子，但非常適合研究 latent activation。

---

# 11. Uniform Zero-Now Activation

本輪找到：

$$
\boxed{
15
}
$$

組 pairs：

$$
(s,t)
$$

滿足：

$$
d_0(s,t)=0
$$

但：

$$
d_3(s,t)>0.
$$

因此它們：

> 現在完全無法區分，但 future composition 可以把差異放大成可觀察 effect。

這是 EXP-15：

$$
\text{current zero gain}
\neq
\text{future zero capability}
$$

的 approximate-error 版本。

---

# 12. Infinite Amplification 的自然情況

若：

$$
d_0=0,
$$

而：

$$
d_h>0,
$$

則普通 ratio：

$$
\frac{d_h}{d_0}
$$

沒有有限值。

因此本輪不把它硬寫成數值。

而標記成：

$$
\boxed{
\text{latent activation}.
}
$$

這比宣稱「無限倍」更精確。

---

# 13. Skewed Prior

為了研究 graded error budget，本輪加入 nonuniform prior。

它對：

- $T$ ；
- $D$ ；
- $G$ ；

設定不同 base rates，

並對：

$$
T=D
$$

加入 mild correlation。

因此 reveal 單一 primitive 後，

對 XOR target也可能產生部分資訊，

不再只有：

$$
0/1
$$

gain。

---

# 14. Skewed $\varepsilon$ –Compression Curve

full contextual horizon 下：

| $\varepsilon$ | $k^\ast$ | compression | optimal partitions |
|---:|---:|---:|---:|
| 0 | 8 | 0% | 1 |
| 0.811278 | 4 | 50% | 1 |
| 0.868767 | 4 | 50% | 16 |
| 0.922142 | 2 | 75% | 1 |
| 0.928114 | 2 | 75% | 4 |
| 0.968520 | 2 | 75% | 16 |
| 0.980864 | 2 | 75% | 64 |
| 0.982256 | 1 | 87.5% | 1 |

這是本系列第一條真正 graded：

$$
\boxed{
\text{error budget}
\longrightarrow
\text{safe compression ratio}
}
$$

曲線。

---

# 15. Compression Phase Transitions

曲線不是平滑的。

class count：

$$
8
\rightarrow
4
\rightarrow
2
\rightarrow
1.
$$

因此 approximate abstraction 具有：

$$
\boxed{
\text{discrete compression phase transitions}.
}
$$

增加很小的 error budget，

可能完全不改 optimum class count；

跨過某些 critical distance 後，

class count會突然下降。

---

# 16. Critical Error Thresholds

在本 skewed finite domain：

第一個 compression threshold：

$$
\boxed{
\varepsilon_1
\approx
0.811278.
}
$$

從：

$$
8
\rightarrow4.
$$

第二個主要 threshold：

$$
\boxed{
\varepsilon_2
\approx
0.922142.
}
$$

從：

$$
4
\rightarrow2.
$$

最後：

$$
\boxed{
\varepsilon_3
\approx
0.982256.
}
$$

從：

$$
2
\rightarrow1.
$$

---

# 17. Approximate Optimum 可以不唯一

exact：

$$
\varepsilon=0
$$

只有：

$$
1
$$

個 optimal partition。

但 approximate regime 中：

$$
\varepsilon
\approx
0.868767
$$

有：

$$
\boxed{
16
}
$$

個不同 4-class optimal partitions。

$$
\varepsilon
\approx
0.928114
$$

有：

$$
4
$$

個 2-class optima。

$$
\varepsilon
\approx
0.968520
$$

有：

$$
16.
$$

而：

$$
\varepsilon
\approx
0.980864
$$

甚至有：

$$
\boxed{
64
}
$$

個不同的 2-class optimal partitions。

---

# 18. Approximate Quotient 的 Canonicality Gap

因此可以定義：

$$
\boxed{
N_{\mathrm{opt}}(\varepsilon,h)
}
$$

為 minimum-class safe partitions 的數量。

若：

$$
N_{\mathrm{opt}}=1,
$$

approximate abstraction具有唯一 optimum。

若：

$$
N_{\mathrm{opt}}>1,
$$

則：

$$
\boxed{
\text{class count 唯一，
grouping structure 不唯一}.
}
$$

這是 exact quotient 完全沒有的問題。

---

# 19. Approximate Abstraction 需要 Secondary Criterion

當：

$$
N_{\mathrm{opt}}>1,
$$

只指定：

$$
(\varepsilon,h)
$$

不足以決定唯一 abstraction。

還需要 secondary criterion，例如：

- nestedness；
- stability；
- provenance preservation；
- minimal representation churn；
- weighted semantic preference；
- canonical tie-breaking。

所以：

$$
\boxed{
\text{approximate minimization}
=
\text{compression optimization}
+
\text{canonicalization policy}.
}
$$

---

# 20. Future Horizon 對 Approximate Compression 的影響

選：

$$
\boxed{
\varepsilon
=
0.868767195327.
}
$$

這個 threshold 同時具有：

- 非傳遞反例；
- horizon-sensitive class count；
- multiple optima。

skewed prior：

### Horizon 0

$$
k^\ast=2,
$$

compression：

$$
\boxed{
75\%.
}
$$

optimal partitions：

$$
2.
$$

### Horizon 1

$$
k^\ast=4,
$$

compression：

$$
\boxed{
50\%.
}
$$

optimal partitions：

$$
16.
$$

### Horizon 2

仍：

$$
k^\ast=4.
$$

### Horizon 3

仍：

$$
k^\ast=4.
$$

所以：

$$
\boxed{
\text{只要求再看一步未來，
safe compression 就從 }75\%\text{ 降到 }50\%.
}
$$

---

# 21. Approximate Future-Obligation Price

對固定：

$$
\varepsilon,
$$

可定義：

$$
\boxed{
P_{\mathrm{future}}^\varepsilon
=
C^\ast(\varepsilon,0)
-
C^\ast(\varepsilon,h).
}
$$

在上述例子：

$$
\boxed{
P_{\mathrm{future}}^\varepsilon
=
0.75-0.50
=
0.25.
}
$$

也就是一層 future compositional guarantee 要付出：

$$
\boxed{
25\text{ percentage points}
}
$$

的 compression budget。

---

# 22. Error Amplification

定義：

$$
\boxed{
A_h(s,t)
=
\frac{
d_h(s,t)
}{
d_0(s,t)
}
}
$$

當：

$$
d_0>0.
$$

若：

$$
d_0=0<d_h,
$$

改記為 latent activation。

---

# 23. Skewed 最大有限 Amplification

最強 finite pair：

$$
s=\varnothing,
$$

$$
t=T.
$$

current difference：

$$
d_0
\approx
0.00597183.
$$

full contextual difference：

$$
d_3
\approx
0.86876720.
$$

因此：

$$
\boxed{
A_3
\approx
145.48.
}
$$

也就是：

> 現在只有千分位到百分位等級的小 effect difference，未來 composition 後可以被放大兩個數量級以上。

---

# 24. 另一個對稱 Amplification

$$
G
$$

與：

$$
TG
$$

同樣：

$$
d_0
\approx
0.00597183,
$$

而：

$$
d_3
\approx
0.86876720.
$$

所以同樣：

$$
A_3
\approx145.48.
$$

這不是單一 pair 偶然。

---

# 25. 為什麼會放大？

原因仍是：

$$
\boxed{
\text{synergy activation}.
}
$$

current output只看到：

$$
\operatorname{Eff}(s),
$$

而 future context 可以補齊 latent dependency set。

一旦 interaction term 被啟動，

原本很小的 marginal difference可能變成接近完整 target entropy 的差異。

---

# 26. Current Error 不是 Future Error Bound

因此不能只用：

$$
d_0(s,t)\le\varepsilon
$$

就宣稱：

> 這兩個 representations 可以安全合併。

真正需要的是：

$$
\boxed{
d_h(s,t)\le\varepsilon
}
$$

或更強：

$$
\boxed{
d_\infty(s,t)\le\varepsilon.
}
$$

這正是 approximate contextual criterion。

---

# 27. Horizon-Bounded Approximation

如果 runtime只承諾：

$$
h
$$

步 future composition，

則：

$$
d_h
$$

就足夠。

不需要為永遠不會發生的 contexts 支付 representation cost。

所以：

$$
\boxed{
\text{approximate safety}
}
$$

仍然應該是 obligation-conditioned。

---

# 28. Error Budget × Horizon Geometry

本輪現在得到一個二維 compression surface：

$$
\boxed{
C^\ast(\varepsilon,h).
}
$$

其基本單調性：

$$
\boxed{
\varepsilon_1\le\varepsilon_2
\Rightarrow
C^\ast(\varepsilon_1,h)
\le
C^\ast(\varepsilon_2,h),
}
$$

以及：

$$
\boxed{
h_1\le h_2
\Rightarrow
C^\ast(\varepsilon,h_1)
\ge
C^\ast(\varepsilon,h_2).
}
$$

所以：

- 放寬 error budget → 可以壓更多；
- 拉長 future guarantee → 必須保留更多。

---

# 29. Exact ESC 是 Surface 的一條邊

當：

$$
\varepsilon=0,
$$

$$
C^\ast(0,h)
$$

就退回 EXP-16 的 exact contextual compression。

因此：

$$
\boxed{
\text{Exact Contextual Minimization}
}
$$

不是另一套理論，

而是 approximate compression surface 的：

$$
\boxed{
\varepsilon=0
}
$$

boundary。

---

# 30. Approximate CER

ESC 最早的 CER 要求：

$$
\Pi(x_1)=\Pi(x_2)
\Rightarrow
x_1=x_2.
$$

現在可定義 approximate behavioral version：

$$
\boxed{
q(s_1)=q(s_2)
\Rightarrow
d_h(s_1,s_2)\le\varepsilon.
}
$$

即 abstraction fiber 的 contextual diameter受：

$$
\varepsilon
$$

控制。

這可以視為：

$$
\boxed{
\varepsilon\text{-Contextual Recoverability}.
}
$$

---

# 31. Approximate Quotient 不應濫用「等價類」

因為：

$$
d\le\varepsilon
$$

不傳遞，

所以稱為：

$$
\boxed{
\text{safe block}
}
$$

比：

$$
\text{equivalence class}
$$

更準確。

只有：

$$
\varepsilon=0
$$

或 distance具有特殊 ultrametric / transitivity property 時，

才自然回到 quotient language。

---

# 32. Canonicality 成為新問題

exact minimization：

$$
\boxed{
\text{coarsest quotient unique}.
}
$$

approximate minimization：

$$
\boxed{
\text{minimum class count may unique，
optimal partition may non-unique}.
}
$$

所以 approximate representation system必須多回答一個問題：

> **若有 64 種一樣小、都符合 error budget 的 abstraction，選哪一種？**

這是 EXP-17 新增的結構問題。

---

# 33. Compression Stability

可以定義 abstraction：

$$
\mathcal P^\ast(\varepsilon,h)
$$

對：

- $\varepsilon$ perturbation；
- prior shift；
- horizon change；

的穩定性。

一個 class count 很好、但 grouping structure稍微改參數就完全重排的 abstraction，

可能不適合真實 runtime。

因此：

$$
\boxed{
\text{minimum size}
\neq
\text{maximum operational quality}.
}
$$

---

# 34. 對 Agent Memory Compression 的直接意義

假設兩個 skill states現在只有：

$$
0.006
$$

bit effect difference。

current-only compressor 可能幾乎必然把它們合併。

但若未來 tool composition 可能使：

$$
0.006
\rightarrow0.869,
$$

則這個 merge 會毀掉大量 future capability。

所以 agent compression需要的不是：

$$
\boxed{
\text{current embedding similarity}
}
$$

而更接近：

$$
\boxed{
\text{future contextual distortion}.
}
$$

---

# 35. 對模型量化／Representation Compression 的意義

同理，

若某個 latent channel現在只有極小 observable effect，

不能只因：

$$
\Delta_{\mathrm{now}}
\ll1
$$

就刪掉。

需要先檢查：

$$
\boxed{
\sup_c
\Delta_{\mathrm{future}}(c).
}
$$

尤其含有：

- gating；
- routing；
- multiplicative interaction；
- parity-like coupling；
- rare-condition activation；

的 representation。

---

# 36. Synergy-Sensitive Compression

因此可以提出：

$$
\boxed{
\text{Synergy-Sensitive Compression Principle}
}
$$

任何 approximate merge：

$$
s_1\sim_\varepsilon s_2
$$

在接受前，至少要測：

$$
\boxed{
\max_{c\in\mathcal C_{\mathrm{adm}}}
d(
\operatorname{Eff}(s_1\circ c),
\operatorname{Eff}(s_2\circ c)
).
}
$$

current marginal distance只是一個 lower bound。

---

# 37. 本輪錨點

$$
\boxed{
\textbf{ESC-EXP-17.A}
\quad
\varepsilon>0
\text{ 時 contextual closeness 一般不再是 equivalence relation。}
}
$$

$$
\boxed{
\textbf{ESC-EXP-17.B}
\quad
\text{approximate safe compression 是 minimum-diameter partition problem，而不是普通 quotient。}
}
$$

$$
\boxed{
\textbf{ESC-EXP-17.C}
\quad
C^\ast(\varepsilon,h)
\text{ 對 }\varepsilon\text{ 單調不減，對 }h\text{ 單調不增。}
}
$$

$$
\boxed{
\textbf{ESC-EXP-17.D}
\quad
\text{skewed finite benchmark 中 current error 可被 future composition 放大約 }145.48\times.
}
$$

$$
\boxed{
\textbf{ESC-EXP-17.E}
\quad
\text{approximate optimum 可以多解，因此 canonicalization policy 是 error-bounded compression 的必要額外層。}
}
$$

---

# 38. 下一輪：ESC-EXP-18

EXP-17 現在留下的最大問題已經不是：

> 可以壓多少？

而是：

> **如果有很多同樣 optimal 的安全 partition，哪一個應該成為 canonical approximation？**

所以 EXP-18 最自然的是：

$$
\boxed{
\text{Canonical Approximate Abstraction and Stability}.
}
$$

可以研究：

- optimal partition 的 nestedness；
- $\varepsilon$ 變動時 representation churn；
- horizon 變動時 class split stability；
- canonical tie-breaking；
- prior shift robustness；
- 是否存在 hierarchical abstraction tree；
- 若 contextual distance 不是 ultrametric，如何找最接近的 stable hierarchy。

也就是把：

$$
\boxed{
\text{minimum safe compression}
}
$$

提升成：

$$
\boxed{
\text{stable, canonical, operationally reusable safe compression}.
}
$$
