# ESC-EXP-07：Bridge-Cost Pareto Frontier 與 Pairwise ESC Distance

**系列：** Extensional Structural Convergence — Experimental Phase  
**文件編號：** ESC-EXP-07  
**版本：** v0.1  
**日期：** 2026-09-22  
**前置：** ESC-00 ～ ESC-06、ESC-EXP-00 ～ ESC-EXP-06  
**狀態：** Executable Finite-Model Evidence / Canonical UTF-8 Source

**作者：** Neo.K  
**機構：** EveMissLab／一言諾科技有限公司  

---

## 摘要

ESC-EXP-06 已把共同 kernel 對齊寫成最小 observation-cost optimization。

本輪將單一六維 target kernel 擴張為所有非空子集：

$$
\varnothing
\neq
\mathcal J'
\subseteq
\mathcal J.
$$

六個 target 產生：

$$
2^6-1=63
$$

個非空 kernel subsets。

對每個：

$$
\mathcal J'
$$

本輪計算：

1. F/C/I 各自最小 adaptation cost；
2. 三路共同最低 bridge cost；
3. target-count coverage；
4. joint-information coverage；
5. pairwise target-conditioned ESC distance；
6. bridge-cost / information-coverage Pareto frontier。

因此 EXP-07 第一次回答：

> 若不要求一次保留全部共同核心，哪些 kernel 最「便宜」？多得到一點共同結構，需要支付多少額外 observation cost？

本輪 controlled domain 仍為：

$$
624
$$

個 canonical $n=3$ substrate classes。

完整六維 kernel entropy：

$$
\boxed{
H(\mathcal J)
=
6.079214\text{ bits}.
}
$$

---

# 1. Runtime

EXP-07 tests：

```text
5 passed
```

完整 experimental regression：

```text
37 passed
```

並成功枚舉：

$$
\boxed{
63
}
$$

個非空 target-kernel subsets。

---

# 2. Adaptation Cost

對 projection：

$$
P\in\{F,C,I\},
$$

定義：

$$
\boxed{
A_P(\mathcal J')
=
\min_{S_P}
H(S_P)
}
$$

subject to：

$$
\mathcal J'
$$

可由：

$$
(\Pi_P,S_P)
$$

恢復。

所以：

$$
A_P
$$

是單一 representation 補到指定 kernel 所需的最低資訊成本。

---

# 3. Pairwise ESC Bridge Distance

本輪定義 target/domain-conditioned symmetric bridge distance：

$$
\boxed{
d_{\mathrm{ESC}}
(P,Q\mid\mathcal J',\mathcal D)
=
A_P(\mathcal J')
+
A_Q(\mathcal J').
}
$$

並定義：

$$
d_{\mathrm{ESC}}(P,P)=0.
$$

這是一個以共同 kernel 為中心的 star-type bridge pseudometric 候選。

它具有：

- 非負；
- 對稱；
- triangle inequality；

但：

$$
d(P,Q)=0
$$

只表示兩邊都不需額外 bridge 才能恢復該 target kernel，

**不表示：**

$$
P=Q.
$$

所以不能把它誤寫成 representation identity metric。

---

# 4. Full-Kernel Distance

對完整六維：

$$
\mathcal J
$$

各 representation adaptation cost：

$$
A_F=1\text{ bit},
$$

$$
A_C=2\text{ bits},
$$

$$
A_I\approx3.713545\text{ bits}.
$$

因此：

$$
\boxed{
d_{\mathrm{ESC}}(F,C)=3.000000\text{ bits}
}
$$

$$
\boxed{
d_{\mathrm{ESC}}(F,I)\approx4.713545\text{ bits}
}
$$

$$
\boxed{
d_{\mathrm{ESC}}(C,I)\approx5.713545\text{ bits}.
}
$$

所以在這個 target kernel / finite domain 下：

$$
\boxed{
F\text{ 與 }C\text{ 最近},
}
$$

$$
\boxed{
C\text{ 與 }I\text{ 最遠}.
}
$$

這不是抽象本體距離，而是**對齊到同一共同 kernel 的最小橋接成本距離**。

---

# 5. Normalized Full-Kernel Distance

因 pair distance 是兩邊 adaptation cost 相加，可用：

$$
2H(\mathcal J)
$$

正規化。

得到約：

$$
\widetilde d(F,C)
\approx
0.246742,
$$

$$
\widetilde d(F,I)
\approx
0.387677,
$$

$$
\widetilde d(C,I)
\approx
0.469925.
$$

這只是一個 finite normalized score，不宣稱具有跨 domain 絕對尺度。

---

# 6. Native Specialization

不加任何 bridge 時：

F 原生可恢復：

$$
5/6
$$

個 targets：

- partition；
- type；
- direction；
- context；
- history。

只缺：

$$
\text{probe}.
$$

C 原生可恢復：

$$
4/6
$$

個 targets：

- partition；
- type；
- direction；
- history。

缺：

$$
\text{context},
\quad
\text{probe}.
$$

I 原生可恢復：

$$
2/6
$$

個 targets：

- partition；
- probe。

因此：

$$
\boxed{
F
}
$$

在本 kernel 上是最接近完整共同 view 的 representation。

---

# 7. 單一 Target Pairwise Distance

## Partition

三條原生都可恢復：

$$
d_{FC}=d_{FI}=d_{CI}=0.
$$

這不表示三條表示相同，只表示 partition target 在 full-support controlled domain 不需要 bridge。

## Type

F/C 原生可恢復；I 需：

$$
0.941829\text{ bits}.
$$

因此：

$$
d_{FC}=0,
$$

$$
d_{FI}=d_{CI}\approx0.941829.
$$

## Direction

同理：

$$
d_{FI}=d_{CI}\approx0.830973.
$$

## Context

F 原生可恢復，C/I 各需 $1$ bit：

$$
d_{FC}=1,
$$

$$
d_{FI}=1,
$$

$$
d_{CI}=2.
$$

## Probe

I 原生可恢復，F/C 各需 $1$ bit：

$$
d_{FC}=2,
$$

$$
d_{FI}=1,
$$

$$
d_{CI}=1.
$$

## History

F/C 原生可恢復；I 需：

$$
0.941829\text{ bits}.
$$

---

# 8. Distance Is Target-Dependent

因此 pair ordering 不是絕對固定。

例如：

對 type：

$$
d(F,C)=0.
$$

對 probe：

$$
d(F,C)=2.
$$

所以：

$$
\boxed{
d_{\mathrm{ESC}}
=
d_{\mathrm{ESC}}
(P,Q\mid\mathcal J',\mathcal D).
}
$$

沒有 target kernel 與 domain，就沒有完整定義的 ESC distance。

---

# 9. Minimum Cost by Kernel Size

每個 target-count 下的最低 bridge cost：

- 1 targets: `0.000000` bits; best-info set: `partition_block_count`; info fraction `22.81%`
- 2 targets: `0.830973` bits; best-info set: `partition_block_count + directional_source_entropy`; info fraction `36.47%`
- 3 targets: `1.771717` bits; best-info set: `partition_block_count + type_entropy + directional_source_entropy`; info fraction `51.91%`
- 4 targets: `2.713545` bits; best-info set: `partition_block_count + type_entropy + directional_source_entropy + history_length`; info fraction `67.10%`
- 5 targets: `4.713545` bits; best-info set: `partition_block_count + type_entropy + directional_source_entropy + context_mode + history_length / partition_block_count + type_entropy + directional_source_entropy + probe_mode + history_length`; info fraction `83.55%`
- 6 targets: `6.713545` bits; best-info set: `partition_block_count + type_entropy + directional_source_entropy + context_mode + probe_mode + history_length`; info fraction `100.00%`

可以看到：

$$
1\text{ target}
$$

時甚至存在：

$$
0\text{ cost}
$$

的 partition kernel。

而隨共同 kernel 擴大，最低成本單調不減。

---

# 10. Cheap Common Core

最便宜的逐步擴張序列非常清楚。

### 1 target

$$
\{\text{partition}\}
$$

成本：

$$
0.
$$

資訊覆蓋：

$$
22.81\%.
$$

### 2 targets

加入 direction：

$$
\{\text{partition},\text{direction}\}
$$

成本：

$$
0.830973.
$$

資訊覆蓋：

$$
36.47\%.
$$

### 3 targets

再加入 type：

$$
\{\text{partition},\text{type},\text{direction}\}
$$

成本：

$$
1.771717.
$$

資訊覆蓋：

$$
51.91\%.
$$

### 4 targets

再加入 history：

$$
\{\text{partition},\text{type},\text{direction},\text{history}\}
$$

成本：

$$
2.713545.
$$

資訊覆蓋：

$$
67.10\%.
$$

這形成一個相當便宜的：

$$
\boxed{
\text{structural + relational + provenance-lite kernel}.
}
$$

---

# 11. Semantic Coordinates 比較昂貴

第五個 target 若加入：

$$
\text{context}
$$

或：

$$
\text{probe},
$$

最低成本都跳到：

$$
4.713545\text{ bits}.
$$

資訊覆蓋：

$$
83.55\%.
$$

所以本 domain 中，真正昂貴的是：

$$
\boxed{
\text{semantic / observer-frame alignment}.
}
$$

而不是 partition 本身。

---

# 12. Full Kernel

完整六維：

$$
\mathcal J
$$

成本：

$$
6.713545\text{ bits},
$$

資訊覆蓋：

$$
100\%.
$$

所以從約：

$$
67.1\%
$$

資訊 kernel 推到 full semantic kernel，需要額外約：

$$
4.0\text{ bits}.
$$

這再次顯示：

$$
\boxed{
\text{結構共同性便宜，語義／觀察框架共同性較昂貴。}
}
$$

至少在本 finite domain 成立。

---

# 13. Entropy-Cost Pareto Frontier

本輪共有：

$$
16
$$

個 nondominated Pareto points。

完整 frontier：

- cost `0.000000` bits → info `22.81%` → `partition_block_count`
- cost `0.830973` bits → info `36.47%` → `partition_block_count, directional_source_entropy`
- cost `0.941829` bits → info `38.30%` → `partition_block_count, history_length`
- cost `0.941829` bits → info `38.30%` → `partition_block_count, type_entropy`
- cost `1.771717` bits → info `51.91%` → `partition_block_count, type_entropy, directional_source_entropy`
- cost `1.883657` bits → info `53.60%` → `partition_block_count, type_entropy, history_length`
- cost `2.713545` bits → info `67.10%` → `partition_block_count, type_entropy, directional_source_entropy, history_length`
- cost `3.771717` bits → info `68.36%` → `partition_block_count, type_entropy, directional_source_entropy, context_mode`
- cost `3.771717` bits → info `68.36%` → `partition_block_count, type_entropy, directional_source_entropy, probe_mode`
- cost `3.883657` bits → info `70.05%` → `partition_block_count, type_entropy, context_mode, history_length`
- cost `3.883657` bits → info `70.05%` → `partition_block_count, type_entropy, probe_mode, history_length`
- cost `4.713545` bits → info `83.55%` → `partition_block_count, type_entropy, directional_source_entropy, context_mode, history_length`
- cost `4.713545` bits → info `83.55%` → `partition_block_count, type_entropy, directional_source_entropy, probe_mode, history_length`
- cost `5.771717` bits → info `84.81%` → `partition_block_count, type_entropy, directional_source_entropy, context_mode, probe_mode`
- cost `5.883657` bits → info `86.50%` → `partition_block_count, type_entropy, context_mode, probe_mode, history_length`
- cost `6.713545` bits → info `100.00%` → `partition_block_count, type_entropy, directional_source_entropy, context_mode, probe_mode, history_length`

這裡的 Pareto 意義是：

> 不存在另一個 kernel subset 同時成本更低或相等，且 joint-information coverage 更高或相等，並至少一項嚴格更好。

---

# 14. Pareto Frontier 的形狀

frontier 顯示三個明顯區段。

## A. Zero / Low Cost Structural Core

$$
0
\rightarrow
2.713545\text{ bits}
$$

可以從：

$$
22.8\%
$$

推到：

$$
67.1\%
$$

資訊覆蓋。

## B. Semantic Expansion

加入 context / probe 後，成本跳升至：

$$
3.77
\sim
5.88\text{ bits},
$$

資訊覆蓋約：

$$
68\%
\sim
86.5\%.
$$

## C. Full Closure

最後：

$$
6.713545\text{ bits}
$$

才到：

$$
100\%.
$$

因此 Pareto 曲線並非線性。

---

# 15. Kernel Coverage 不能只用 Target Count

例如兩個不同 5-target subsets 可能 target count 相同，但 joint entropy 不同。

所以本輪同時保存：

$$
\boxed{
\text{target fraction}
}
$$

與：

$$
\boxed{
\text{information fraction}.
}
$$

真正 Pareto optimization 以：

$$
H(\mathcal J')
$$

作為主要 coverage 軸，而不是只數 target 數量。

---

# 16. ESC Distance 與 Representation Specialization

full-kernel adaptation：

$$
A_F<A_C<A_I.
$$

這不表示：

$$
F
$$

一般比 I 「更強」。

只表示對本六維 target kernel：

F 原本保存較多 target-relevant observables。

如果 target 改成純 probe geometry，排序會改變。

因此：

$$
\boxed{
\text{specialization}
\neq
\text{global superiority}.
}
$$

---

# 17. 一個更成熟的「理論距離」概念

ESC 現在可以把「兩套表示多像」拆成三種不同問題：

## 17.1 Native overlap

兩邊原生共同 recoverable 的 targets 有多少？

## 17.2 Bridge distance

兩邊補到共同 kernel 要多少成本？

## 17.3 Pareto profile

當 kernel coverage 從小到大時，距離如何變化？

所以單一 scalar：

$$
d_{\mathrm{ESC}}
$$

只是其中一個截面。

完整關係其實更接近：

$$
\boxed{
\mathcal P_{PQ}
:
\text{coverage}
\mapsto
\text{minimum bridge cost}.
}
$$

---

# 18. EXP-07 對「三套理論到底有多遠」的答案

在目前 full kernel：

$$
\boxed{
F-C
<
F-I
<
C-I.
}
$$

但 target-wise：

- relation semantics 上 F/C 幾乎重合；
- probe frame 上 I 更接近 F/C 的補充端；
- context 上 F 與 C/I 分離；
- history 上 I 與 F/C 分離。

因此 representation difference 具有：

$$
\boxed{
\text{anisotropy}.
}
$$

也就是在不同 structural axis 上距離不同。

---

# 19. Anisotropic ESC Geometry

這表示未來不應只建一個 scalar distance。

更完整可以寫成 distance vector：

$$
\boxed{
\mathbf d_{\mathrm{ESC}}(P,Q)
=
(
d_{\pi},
d_T,
d_D,
d_\Gamma,
d_P,
d_H,
\dots
).
}
$$

而 full-kernel scalar distance只是對這些軸在指定成本函數下的聚合。

---

# 20. 本輪錨點

$$
\boxed{
\textbf{ESC-EXP-07.A}
\quad
d_{\mathrm{ESC}}(P,Q\mid\mathcal J,\mathcal D)
\text{ 必須 target/domain-conditioned。}
}
$$

$$
\boxed{
\textbf{ESC-EXP-07.B}
\quad
\text{本 finite full kernel 中 }F-C<F-I<C-I.
}
$$

$$
\boxed{
\textbf{ESC-EXP-07.C}
\quad
\text{共同 kernel 的 bridge-cost / information-coverage frontier 非線性。}
}
$$

$$
\boxed{
\textbf{ESC-EXP-07.D}
\quad
\text{結構／關係 kernel 相對便宜，而 semantic / observer-frame alignment 成本較高。}
}
$$

$$
\boxed{
\textbf{ESC-EXP-07.E}
\quad
\text{representation distance 具有 structural-axis anisotropy。}
}
$$

---

# 21. 下一輪：ESC-EXP-08

下一輪最自然的是把：

$$
d_{\mathrm{ESC}}
$$

從單一 controlled domain 推到：

$$
\boxed{
\text{cross-domain stability}.
}
$$

也就是比較：

- 不同 $n$ ；
- 不同 support coverage；
- 不同 target priorities；
- 不同 channel-cost model；

之下：

$$
F-C<F-I<C-I
$$

是否仍成立。

如果排序會翻轉，就建立真正的：

$$
\boxed{
\text{ESC distance field}
}
$$

而不是單一 distance matrix。
