# ESC-EXP-18：Canonical Approximate Abstraction and Stability

**系列：** Extensional Structural Convergence — Experimental Phase  
**文件編號：** ESC-EXP-18  
**版本：** v0.1  
**日期：** 2026-09-23  
**前置：** ESC-00 ～ ESC-06、ESC-EXP-00 ～ ESC-EXP-17  
**狀態：** Canonicalization / Stability / Hierarchical Approximation Experiment

**作者：** Neo.K  
**機構：** EveMissLab／一言諾科技有限公司  

---

## 摘要

ESC-EXP-17 已經證明 approximate contextual minimization 與 exact quotient 有一個根本差異：

$$
\boxed{
\text{minimum class count 可以唯一，
但 optimal partition 本身可以有多個。}
}
$$

在 skewed finite benchmark 中，某些 $\varepsilon$ 甚至有：

$$
16,
\quad
64
$$

個 equally optimal partitions。

因此本輪研究的問題不再是：

> 可以壓到多小？

而是：

> **在所有一樣小、都滿足 error budget 的 abstractions 中，哪一個應成為 canonical representation？**

本輪研究三種穩定性要求：

1. **low churn**：相鄰 operating points 間 representation 少重排；
2. **nestedness**： $\varepsilon$ 放寬時只 merge、不 split；
3. **prior robustness**：prior 輕微改變時 abstraction structure 盡量不變。

主要結果：

$$
\boxed{
\text{pointwise optimal}
\not\Rightarrow
\text{globally stable}.
}
$$

具體而言：

- 每個 $\varepsilon$ 各自選 lexicographic optimum，總 churn：
  $$
  1.5;
  $$
- 若在所有 pointwise-optimal paths 中全域最小化 churn，總 churn：
  $$
  \boxed{
  1.142857
  }
  $$
  ，下降約：
  $$
  \boxed{
  23.81\%.
  }
  $$

更重要的是：

$$
\boxed{
\text{不存在一條全程 nested、且每一點都保持 class-count optimal 的 }\varepsilon\text{-path。}
}
$$

第一個不可兼容點出現在：

$$
\boxed{
\varepsilon
=
0.922142264459.
}
$$

若強制 hierarchy，必須暫時保留：

$$
4\text{ classes}
$$

而 pointwise optimum 已經只需：

$$
2.
$$

所以 canonical hierarchical stability 必須付出：

$$
\boxed{
\text{compression overhead}.
}
$$

另一方面，固定：

$$
\varepsilon
=
0.868767195327
$$

而增加 future horizon 時，反而存在：

$$
\boxed{
8
}
$$

條完全 nested 的 optimal paths。

因此：

$$
\boxed{
\text{epsilon-axis canonicalization}
}
$$

與：

$$
\boxed{
\text{horizon-axis canonicalization}
}
$$

具有不同幾何。

最後，prior temperature perturbation：

$$
\tau
=
0.85,
1.0,
1.15
$$

會明顯移動 critical $\varepsilon$，

但 4-class 與 2-class 的 canonical partition structure 在本測試中完全不變：

$$
\boxed{
\text{churn}=0.
}
$$

也就是：

> **threshold 不穩定，不代表 abstraction topology 不穩定。**

---

# 1. Runtime

EXP-18 新增 regression：

```text
5 passed
```

測試包括：

- 不存在全程 nested pointwise-optimal epsilon path；
- failure point 精確落在：
  $$
  \varepsilon=0.922142264459;
  $$
- fixed- $\varepsilon$ horizon optima 存在 nested path；
- global minimum-churn path 不劣於 independent lexicographic selection；
- prior robustness 同時輸出 4-class / 2-class compression levels。

---

# 2. Partition Churn

對同一 mechanism set 上的兩個 partitions：

$$
\mathcal P,
\mathcal Q,
$$

考慮所有 unordered state pairs：

$$
\{x,y\}.
$$

若：

- 在 $\mathcal P$ 中同 block；
- 在 $\mathcal Q$ 中不同 block；

或反之，

則該 pair 發生 coassignment change。

定義：

$$
\boxed{
\operatorname{Churn}
(\mathcal P,\mathcal Q)
=
\frac{
\#\text{pairwise coassignment changes}
}{
\binom{|S|}{2}
}.
}
$$

本輪：

$$
|S|=8,
$$

所以分母：

$$
\binom82
=
28.
$$

---

# 3. Independent Canonicalization 的問題

最簡單 policy：

> 每一個 $\varepsilon$ 都獨立選 lexicographically first optimal partition。

這保證：

- deterministic；
- 每點 class-count optimal。

但它完全不考慮前一個 representation。

結果總 churn：

$$
\boxed{
1.5.
}
$$

這意味沿整條 epsilon sweep：

28 個 state pairs 的 coassignment 關係累積發生大量重排。

---

# 4. Global Minimum-Churn Optimal Path

本輪把每個 epsilon threshold 的所有 pointwise-optimal partitions 當作 layer。

相鄰 partitions 的 edge cost：

$$
\operatorname{Churn}.
$$

然後做 shortest-path dynamic programming：

$$
\boxed{
\min
\sum_i
\operatorname{Churn}
(
\mathcal P_i,
\mathcal P_{i+1}
)
}
$$

subject to：

$$
\mathcal P_i
$$

在每個 operating point 都是 minimum-class safe partition。

結果：

$$
\boxed{
\operatorname{Churn}_{\min}
=
1.142857.
}
$$

相較 independent lex：

$$
1.5
$$

下降：

$$
\boxed{
23.81\%.
}
$$

---

# 5. Minimum-Churn Path

最小 churn path：

### $\varepsilon=0$

$$
8\text{ singleton classes}.
$$

### $\varepsilon=0.811278$

$$
4\text{ classes}:
$$

$$
\{\varnothing,G\},
$$

$$
\{D,DG\},
$$

$$
\{T,TG\},
$$

$$
\{TD,TDG\}.
$$

### $\varepsilon=0.868767$

仍然：

$$
4\text{ classes},
$$

但為降低下一階段 churn，會選另一個 equally optimal grouping：

$$
\{\varnothing,G\},
$$

$$
\{D\},
$$

$$
\{T,TG\},
$$

$$
\{DG,TD,TDG\}.
$$

這是一個非常重要的結果：

> **canonical abstraction 可能故意不選局部看起來最「自然」的 grouping，而選一個更適合未來 trajectory 的 equally optimal grouping。**

---

# 6. Path-Dependent Canonicality

所以 canonicalization 不能只寫：

$$
\boxed{
\operatorname{Canon}(\varepsilon)
}
$$

更完整應是：

$$
\boxed{
\operatorname{Canon}
(
\varepsilon_i
\mid
\mathcal P_{i-1},
\Theta_{\mathrm{future}}
)
}
$$

其中：

$$
\Theta_{\mathrm{future}}
$$

可以代表預期的未來 operating-point trajectory。

這表示 approximate canonicalization 本質上是 path-dependent planning problem。

---

# 7. Nestedness

若：

$$
\varepsilon_1
<
\varepsilon_2,
$$

一個理想 hierarchical abstraction 應滿足：

$$
\boxed{
\mathcal P_{\varepsilon_1}
\preceq
\mathcal P_{\varepsilon_2}
}
$$

也就是：

> error budget 放寬時，只允許 classes merge，不允許先 merge 後又重新 split/re-group。

這會產生真正的 abstraction hierarchy。

---

# 8. Pointwise Optimal Hierarchy 不存在

本輪 exhaustively 搜尋所有 optimal partitions。

結果：

$$
\boxed{
\text{不存在全程 nested optimal path。}
}
$$

第一個 failure：

$$
\boxed{
\varepsilon
=
0.922142264459.
}
$$

---

# 9. 為什麼會失敗？

在：

$$
\varepsilon
=
0.811278124459
$$

唯一 4-class optimum 是：

$$
\{\varnothing,G\},
$$

$$
\{D,DG\},
$$

$$
\{T,TG\},
$$

$$
\{TD,TDG\}.
$$

但在：

$$
\varepsilon
=
0.922142264459
$$

唯一 2-class pointwise optimum 是：

$$
\{\varnothing,D,G,T\},
$$

與：

$$
\{DG,TD,TG,TDG\}.
$$

後者**不是**前者的 coarsening。

因此不能靠單純 merge 從第一個 optimum 到第二個 optimum。

必須重新分組。

---

# 10. Stability–Optimality Conflict

所以 approximate abstraction 中出現新的基本張力：

$$
\boxed{
\text{pointwise compression optimality}
\quad\text{vs.}\quad
\text{hierarchical stability}.
}
$$

如果 insist 每一個 epsilon 都用最少 classes，

必須接受 representation reorganization。

如果 insist hierarchy，

必須接受某些 epsilon 下不是 pointwise minimal。

---

# 11. Canonical Nested Hierarchy

本輪建立一個 deterministic hierarchy-preserving policy：

每一個新的 epsilon：

1. 只考慮前一 partition 的 coarsenings；
2. 保留 epsilon-safe；
3. 選 class count 最少者；
4. lexicographic tie-break。

這保證：

$$
\boxed{
\text{nested}
}
$$

與：

$$
\boxed{
\text{deterministic}.
}
$$

但不宣稱是所有 nested paths 中的 global optimum。

---

# 12. Hierarchy Overhead

在：

$$
\varepsilon
=
0.922142264459
$$

pointwise optimum：

$$
2\text{ classes}.
$$

nested canonical hierarchy：

$$
4\text{ classes}.
$$

overhead：

$$
\boxed{
+2.
}
$$

下一個 threshold：

$$
\varepsilon
=
0.928114090327
$$

仍是：

$$
+2.
$$

之後到：

$$
\varepsilon
=
0.968519847998
$$

hierarchy 才可以安全 merge 到：

$$
2\text{ classes}.
$$

整條 sweep 的 cumulative class overhead：

$$
\boxed{
4.
}
$$

---

# 13. Hierarchical Delay

這可以解讀成：

> hierarchy 要等待更寬鬆的 error budget，才能做 pointwise optimizer 已經提前做的 compression。

因此可以定義：

$$
\boxed{
\text{Hierarchical Compression Delay}.
}
$$

對某個 merge event：

$$
\varepsilon_{\mathrm{hier}}
-
\varepsilon_{\mathrm{opt}}.
$$

本例從：

$$
0.922142
$$

延遲到：

$$
0.968520.
$$

約：

$$
\boxed{
0.046378.
}
$$

---

# 14. Horizon Axis 的情況不同

固定：

$$
\varepsilon
=
0.868767195327,
$$

從：

$$
h=0
\rightarrow1
\rightarrow2
\rightarrow3.
$$

pointwise optimal class counts：

$$
2
\rightarrow4
\rightarrow4
\rightarrow4.
$$

因 future obligation增加，partition應該做 refinement：

$$
\text{coarse}
\rightarrow
\text{fine}.
$$

本輪找到：

$$
\boxed{
8
}
$$

條完全 nested 且 pointwise-optimal 的 horizon paths。

所以 horizon axis 上：

$$
\boxed{
\text{stability}
}
$$

與：

$$
\boxed{
\text{pointwise optimality}
}
$$

並沒有衝突。

---

# 15. Epsilon Axis vs Horizon Axis

因此兩個 operating parameters 不可混為同一種 hierarchy。

## $\varepsilon\uparrow$

error budget 放寬：

$$
\text{理想方向：merge}.
$$

但 pointwise optimum 可能要求 regrouping。

## $h\uparrow$

future obligation增加：

$$
\text{理想方向：split/refine}.
$$

在本 benchmark 中存在 optimal nested chain。

所以：

$$
\boxed{
\text{canonical abstraction geometry is parameter-axis dependent}.
}
$$

---

# 16. Prior Perturbation

本輪對 skewed prior 做 temperature transform：

$$
w_\tau(x)
\propto
w(x)^\tau,
$$

測：

$$
\tau
=
0.85,
1.0,
1.15.
$$

 $\tau<1$ 使 prior較平，

 $\tau>1$ 使 prior較尖。

---

# 17. 4-Class Critical Threshold

要第一次安全壓到：

$$
4\text{ classes},
$$

critical epsilon：

$$
\tau=0.85:
\quad
0.858364,
$$

$$
\tau=1:
\quad
0.811278,
$$

$$
\tau=1.15:
\quad
0.760876.
$$

threshold 明顯移動。

但是三個 prior 的唯一 optimal 4-class partition完全相同。

所以：

$$
\boxed{
\text{partition churn}=0.
}
$$

---

# 18. 2-Class Critical Threshold

同樣：

$$
\tau=0.85:
\quad
0.944097,
$$

$$
\tau=1:
\quad
0.922142,
$$

$$
\tau=1.15:
\quad
0.896802.
$$

critical epsilon仍然明顯改變。

但 canonical 2-class partition也完全不變：

$$
\boxed{
\text{churn}=0.
}
$$

---

# 19. Threshold Stability 與 Topology Stability

因此需要分開：

## Threshold robustness

$$
\varepsilon_c(\mu)
$$

對 prior shift 是否穩定。

## Topology robustness

最優 grouping：

$$
\mathcal P^\ast(\mu)
$$

是否穩定。

本輪得到一個很清楚的例子：

$$
\boxed{
\text{threshold sensitive}
}
$$

但：

$$
\boxed{
\text{topology robust}.
}
$$

---

# 20. 這對 Runtime 很重要

若 prior drift 只改 critical threshold，

但不改 abstraction topology，

runtime 不需要重建 representation graph。

只需要更新：

$$
\boxed{
\text{何時切換 abstraction level}.
}
$$

這比每次 prior shift 都重新分群成本低很多。

---

# 21. Canonicality 的三個不同目標

到目前可以區分：

### Static canonicality

同一 operating point：

$$
\theta
$$

選唯一 partition。

### Path canonicality

沿：

$$
\theta_0,\theta_1,\dots
$$

最小 representation churn。

### Hierarchical canonicality

要求不同 abstraction levels 形成 nested partitions。

三者不一定能同時達到 pointwise optimum。

---

# 22. Operational Canonicality

因此真正 runtime 需要的 canonicality 可以定義成 multi-objective：

$$
\boxed{
\min
\left(
\text{class count},
\text{churn},
\text{hierarchy violations},
\text{prior sensitivity}
\right).
}
$$

而不是簡單：

$$
\min |\mathcal P|.
$$

---

# 23. Representation Churn 是真成本

representation regrouping 可能造成：

- cache invalidation；
- memory relocation；
- embedding remap；
- agent-state migration；
- tool routing changes；
- provenance remapping。

所以即使兩個 partitions 在數學上同樣 optimal，

churn：

$$
\boxed{
\neq0
}
$$

可能代表真實系統成本。

EXP-18 第一次把這一項正式放進 ESC abstraction selection。

---

# 24. Canonical Path Optimization

因此可以定義：

$$
\boxed{
\min_{\mathcal P_0,\dots,\mathcal P_n}
\sum_i
\operatorname{Churn}
(
\mathcal P_i,\mathcal P_{i+1}
)
}
$$

subject to：

$$
\mathcal P_i
\in
\operatorname{Opt}(\theta_i).
$$

本輪的 global min-churn path 就是第一個有限實例。

---

# 25. Pointwise Optimality 不是 Dynamic Optimality

這也是一個一般性的結論：

$$
\boxed{
\arg\min_{\mathcal P_i}
C(\mathcal P_i)
}
$$

逐點解，

不一定等於：

$$
\boxed{
\arg\min_{\mathcal P_{0:n}}
\left[
\sum_i C(\mathcal P_i)
+
\lambda
\sum_i
\operatorname{TransitionCost}
\right].
}
$$

EXP-18 中 lexicographic independent policy 就輸給 global min-churn policy。

---

# 26. Approximate Abstraction 開始變成 Control Problem

EXP-17 是：

$$
\text{optimization over partitions}.
$$

EXP-18 進一步變成：

$$
\boxed{
\text{optimization over partition trajectories}.
}
$$

也就是 state：

$$
\mathcal P_t,
$$

control：

$$
\mathcal P_{t+1},
$$

environment：

$$
\theta_t
=
(\varepsilon,h,\mu).
$$

這已經開始接近 dynamic abstraction control。

---

# 27. Stable Hierarchy 不必完全 Pointwise Optimal

如果 application 需要：

- progressive loading；
- hierarchical cache；
- multi-resolution memory；
- stable parent-child class structure；

那麼多保留：

$$
2
$$

個 classes 一段 epsilon 區間，

可能比立即從 4 壓成 2、之後重組 representation 更划算。

所以：

$$
\boxed{
\text{smallest model}
\neq
\text{best operational model}.
}
$$

---

# 28. Approximate Abstraction Tree

如果強制 nestedness，

所有 abstraction levels可以形成：

$$
\boxed{
\text{hierarchical abstraction tree}.
}
$$

每一個更高 error budget level只做 node merge。

這種 structure 可以：

- incremental update；
- stable IDs；
- provenance inheritance；
- coarse-to-fine query；
- progressive decoding。

所以即使有 class-count overhead，也可能具有很高工程價值。

---

# 29. 不是所有 Optimal Partitions 都能放進同一 Tree

EXP-18 的 no-nested-optimal-chain 結果非常重要：

$$
\boxed{
\text{set of pointwise optima}
}
$$

一般不能直接組成一棵 hierarchy。

所以若系統要求 tree representation，

必須在：

$$
\boxed{
\text{optimality}
}
$$

與：

$$
\boxed{
\text{hierarchical consistency}
}
$$

之間做明確 trade-off。

---

# 30. Canonicality Gap

可定義：

$$
\boxed{
\Delta_{\mathrm{canon}}
=
C(
\mathcal P_{\mathrm{canonical}}
)
-
C(
\mathcal P_{\mathrm{pointwise\ optimum}}
).
}
$$

如果成本用 class count：

本輪 nested hierarchy 在兩個 thresholds：

$$
\Delta_{\mathrm{canon}}=2.
$$

這就是 hierarchy 所支付的 canonicality gap。

---

# 31. Stability Gap

同樣可定義：

$$
\boxed{
\Delta_{\mathrm{churn}}
=
\operatorname{Churn}_{\mathrm{naive}}
-
\operatorname{Churn}_{\mathrm{stable}}.
}
$$

本輪：

$$
1.5
-
1.142857
=
\boxed{
0.357143.
}
$$

相對下降：

$$
\boxed{
23.81\%.
}
$$

---

# 32. Prior Robustness 目前是好消息

在有限 temperature perturbation：

$$
0.85
\le
\tau
\le
1.15
$$

中，

4-class / 2-class canonical topology 都沒有改變。

這表示至少在這個 finite regime：

$$
\boxed{
\text{abstraction shape 比 error threshold 更 robust}.
}
$$

未來可擴張到：

- stronger prior shifts；
- support perturbations；
- target reweighting；
- domain drift。

---

# 33. Exact Canonicality vs Approximate Canonicality

Exact：

$$
\varepsilon=0
$$

通常由 equivalence classes 自然提供 canonical quotient。

Approximate：

$$
\varepsilon>0
$$

則需要額外 policy。

因此：

$$
\boxed{
\text{canonicality is free in exact quotient,
but becomes a design resource in approximate abstraction}.
}
$$

---

# 34. ESC Approximate Abstraction Stack

到目前可以形成一個完整 stack：

### Level 1

$$
d_h
$$

contextual distortion。

### Level 2

$$
\varepsilon\text{-safe blocks}.
$$

### Level 3

minimum-class safe partitions。

### Level 4

canonical partition selection。

### Level 5

stable partition trajectory。

### Level 6

hierarchical / prior-robust abstraction system。

這已經遠超單純「壓縮」。

---

# 35. 本輪錨點

$$
\boxed{
\textbf{ESC-EXP-18.A}
\quad
\text{pointwise optimal approximate partitions 不一定存在全程 nested chain。}
}
$$

$$
\boxed{
\textbf{ESC-EXP-18.B}
\quad
\text{global min-churn path 可在保持 pointwise class-count optimal 的前提下顯著降低 representation churn。}
}
$$

$$
\boxed{
\textbf{ESC-EXP-18.C}
\quad
\text{hierarchical canonicality 可能要求 class-count overhead。}
}
$$

$$
\boxed{
\textbf{ESC-EXP-18.D}
\quad
\text{threshold robustness 與 abstraction-topology robustness 必須分開測量。}
}
$$

$$
\boxed{
\textbf{ESC-EXP-18.E}
\quad
\text{canonical approximate abstraction 是 trajectory-level optimization，而非只是一個 pointwise clustering 問題。}
}
$$

---

# 36. 下一輪：ESC-EXP-19

現在下一個真正自然的方向是：

$$
\boxed{
\text{Dynamic Abstraction Control}.
}
$$

讓 operating point：

$$
\theta_t
=
(\varepsilon_t,h_t,\mu_t)
$$

隨時間變化，

並把 abstraction：

$$
\mathcal P_t
$$

當成 controlled state。

目標可以寫成：

$$
\boxed{
\min
\sum_t
\left[
\alpha |\mathcal P_t|
+
\beta \operatorname{Distortion}_t
+
\gamma \operatorname{Churn}
(
\mathcal P_{t-1},
\mathcal P_t
)
\right].
}
$$

同時滿足：

$$
\boxed{
\operatorname{diam}_{d_{h_t}}
(B)
\le
\varepsilon_t.
}
$$

這就會把 ESC approximate abstraction 真正推進到：

> **在世界、任務與 future obligation 持續變動時，representation 應該何時 merge、何時 split、何時維持不動？**

也就是從 static representation theory 進入真正的：

$$
\boxed{
\text{representation dynamics}.
}
$$
