← Archive
lm-002724 · 2026-08

多尺度概率幾何:微觀、中觀、宏觀與全域判定空間

下載 MD 檔 ⬇
📎 附件 · Companion files — 隨文交付的程式 / 證明 / 資料,可獨立下載重驗

title: "多尺度概率幾何:微觀、中觀、宏觀與全域判定空間" english_title: "Multi-Scale Probability Geometry: Micro, Meso, Macro, and Global Judgment Spaces" series: "判定域概率論與超概率研究" series_id: "JDPSP" paper_id: "JDPSP-04" author: "Neo.K" organization: "EveMissLab" version: "0.1.0" status: "研究初稿 / geometric framework proposal" date: "2026-08-13" language: "zh-TW"

多尺度概率幾何:微觀、中觀、宏觀與全域判定空間

Multi-Scale Probability Geometry: Micro, Meso, Macro, and Global Judgment Spaces

作者: Neo.K
機構: EveMissLab
系列: 判定域概率論與超概率研究,Paper 04
版本: v0.1.0
日期: 2026-08-13

摘要

前兩篇工作分別建立判定域

D=(X,Σ;r,s,c)\mathfrak D=(X,\Sigma;r,s,c)

以及跨判定域的提升見證

(D,P)W(E,Q).(\mathfrak D,P) \xRightarrow{\mathcal W} (\mathfrak E,Q).

本文把單一判定域與單一 lifting arrow 擴張成一個多尺度空間,研究微觀、中觀、宏觀與明示全域概率如何被排列、比較、聚合與沿路徑搬運。

本文首先區分兩種不同的「概率幾何」。第一種是既有的 probability-distribution geometry,例如 Fisher information geometry、divergence geometry 與 Wasserstein / optimal-transport geometry。這些方向已具有成熟理論,本文不宣稱重新發明。第二種是本文提出的 judgment-domain geometry:把判定域作為節點,把 admissible transport 作為箭頭,把尺度索引作為基底,研究 probability objects 如何在不同判定域纖維之間移動。

令尺度系統為偏序或更一般 category:

S,\mathbf S,

判定域 category 為:

JDom,\mathbf{JDom},

並設尺度投影:

π:JDomS.\pi: \mathbf{JDom} \rightarrow \mathbf S.

尺度 ss 上的判定域纖維記為:

JDoms=π1(s).\mathbf{JDom}_s = \pi^{-1}(s).

對每一判定域 D\mathfrak D,再附上其概率對象集合:

Prob(D).\mathbf{Prob}(\mathfrak D).

由此形成一個 bundle-like / indexed probability structure,而非預先宣稱為微分幾何意義的 fiber bundle。

本文引入三類幾何量。第一,尺度距離 dSd_S 描述判定域在尺度圖或尺度偏序中的結構距離;第二,纖維內距離使用既有 Fisher、divergence 或 Wasserstein 結構比較同一判定域中的概率分布;第三,跨域比較成本先把兩個概率 transport 到共同比較域,再計算分布差異與 transport cost,因此一般只稱 comparison cost,而不無條件宣稱為 metric。

本文進一步定義路徑 transport:

Tγ=TenTe1,T_{\gamma} = T_{e_n}\circ\cdots\circ T_{e_1},

路徑依賴:

Δγ1,γ2(P)=D(Tγ1P,Tγ2P),\Delta_{\gamma_1,\gamma_2}(P) = D( T_{\gamma_1}P, T_{\gamma_2}P ),

以及閉路 defect:

Hγ(P)=D(P,TγP).H_{\gamma}(P) = D( P,T_{\gamma}P ).

若尺度系統是偏序 category,且 transport assignment 構成 functor,則任意兩條由同一尺度起點到終點的 refinement chains 必須給出同一 composite transport;本文稱之為 Path-Independence Proposition。反之,非零 path defect 代表尺度路徑、語境選擇或 aggregation mechanism 具有不可忽略的結構作用。

本文亦利用 data-processing principle 定義 stochastic coarse-graining 的 distinguishability loss:

Lf(K;P,Q)=Df(PQ)Df(KPKQ)0,L_f(K;P,Q) = D_f(P\Vert Q) - D_f(K_{\star}P\Vert K_{\star}Q) \ge0,

並證明串接 stochastic channels 時此 loss 具有望遠鏡式可加性。這提供一種可審計的跨尺度資訊損失帳本。

最後,本文重新定義「全域」:全域不是沒有尺度下標的默認狀態,而應是研究尺度 category 中明示的 terminal / designated maximal judgment domain。若尺度結構不存在唯一 terminal object,則不應假設存在唯一 canonical global probability。

本文因此將多尺度概率表示為:

Scale Base+Judgment-Domain Fibers+Probability Fibers+Transport Geometry+Path / Loss Structure.\boxed{ \text{Scale Base} + \text{Judgment-Domain Fibers} + \text{Probability Fibers} + \text{Transport Geometry} + \text{Path / Loss Structure}. }

這為下一篇「遞歸概率場」提供幾何底座:當每一點不只承載 PP,而承載

P,P(P),P2(P),P,\mathsf P(P),\mathsf P^2(P),\ldots

時,概率階數將成為此幾何的第二個獨立方向。

關鍵詞: 多尺度概率、判定域、概率幾何、information geometry、Wasserstein geometry、coarse-graining、data processing、fiber、path dependence、holonomy defect、Markov kernel、全域概率、尺度空間


1. 問題起點:概率不只存在於一個空間,而存在於一族尺度空間

Paper 02 建立:

D=(X,Σ;r,s,c).\mathfrak D = (X,\Sigma;r,s,c).

Paper 03 建立:

(D,P)W(E,Q).(\mathfrak D,P) \xRightarrow{\mathcal W} (\mathfrak E,Q).

這已足以描述一次跨域概率轉換。

然而實際多尺度系統通常不是:

DE\mathfrak D \rightarrow \mathfrak E

的一條箭頭,而更接近:

Dμ,1,Dμ,2,Dm,1,Dm,2,DM,1,DG.\mathfrak D_{\mu,1}, \mathfrak D_{\mu,2}, \ldots \rightarrow \mathfrak D_{m,1}, \mathfrak D_{m,2}, \ldots \rightarrow \mathfrak D_{M,1}, \ldots \rightarrow \mathfrak D_G.

其中:

  • 多個 micro domains 可以共同形成一個 meso domain;
  • 多個 meso domains 可能聚合成不同 macro descriptions;
  • 不同 aggregation paths 可能導向不同宏觀概率;
  • 某些尺度不存在自然 total order;
  • 某些研究甚至不存在唯一 global object。

因此本文的核心問題變成:

一族判定域如何構成概率空間的幾何?\boxed{ \text{一族判定域如何構成概率空間的幾何?} }

2. 「概率幾何」已有成熟理論:本文不能重新命名

2.1 Information geometry

Information geometry 已把 statistical models 視為幾何空間,研究 Fisher metric、Amari--Chentsov tensor、divergences、dual connections 與 sufficient statistics。

對參數模型:

{pθ}θΘ,\{ p_{\theta} \}_{\theta\in\Theta},

Fisher information metric 可寫為:

gij(θ)=Eθ[ilogpθ(X)jlogpθ(X)].g_{ij}(\theta) = \mathbb E_{\theta} \left[ \partial_i\log p_{\theta}(X) \, \partial_j\log p_{\theta}(X) \right].

Markov morphisms 與 sufficient statistics 下的單調性/不變性亦是 information geometry 的核心結果。

因此本文不能把:

概率分布可以形成幾何空間\boxed{ \text{概率分布可以形成幾何空間} }

宣稱為新發現。

2.2 Wasserstein / optimal-transport geometry

若:

(X,d)(X,d)

為 metric space,則具有有限 pp 階矩的 probability measures 可形成:

Pp(X).\mathcal P_p(X).

Wasserstein distance 定義為:

Wp(P,Q)=(infγΠ(P,Q)X×Xd(x,y)pdγ(x,y))1/p.W_p(P,Q) = \left( \inf_{\gamma\in\Pi(P,Q)} \int_{X\times X} d(x,y)^p \,d\gamma(x,y) \right)^{1/p}.

因此 probability measures 本身已具有非常成熟的 transport geometry。

近年的研究甚至直接對 Wasserstein spaces 中的一系列 probability measures 建立 multiscale transform、refinement operator 與跨尺度幾何分析。

所以本文不能把:

multiscale geometry of probability measures\boxed{ \text{multiscale geometry of probability measures} }

本身宣稱為首次提出。

2.3 本文真正研究的另一層幾何

本文新增的問題是:

概率分布所在的判定域本身如何排列成尺度空間?\boxed{ \text{概率分布所在的判定域本身如何排列成尺度空間?} }

也就是把幾何分成:

geometry inside probability spaces\boxed{ \text{geometry inside probability spaces} }

與:

geometry between judgment domains.\boxed{ \text{geometry between judgment domains}. }

本文研究後者,並把前者作為可以嵌入每一個判定域纖維中的成熟工具。


3. 尺度基底

3.1 尺度不必是一條直線

最簡模型可以取:

sμsmsMsG.s_{\mu} \preceq s_m \preceq s_M \preceq s_G.

但真實多尺度系統未必是 total order。

例如:

sindividual,s_{\mathrm{individual}}, steam,s_{\mathrm{team}}, sdepartment,s_{\mathrm{department}}, scompanys_{\mathrm{company}}

可能是一條鏈。

但同一研究同時可能有:

sgeographic,sorganizational,stemporal,s_{\mathrm{geographic}}, \qquad s_{\mathrm{organizational}}, \qquad s_{\mathrm{temporal}},

形成不可線性排序的尺度族。

因此本文令:

S\mathbf S

為尺度 category。

最簡情況下:

S\mathbf S

可以是偏序 category。

更一般情況下,可以具有多種 typed morphisms。

3.2 尺度節點

每個:

sOb(S)s\in\operatorname{Ob}(\mathbf S)

表示一個被明示的 observational / organizational scale。

例如:

S={μ,m,M,G}.\mathcal S = \{ \mu,m,M,G \}.

但這只是一個實作方便的四層模型,不是宇宙中唯一自然分層。

3.3 尺度箭頭

若:

sts\rightarrow t

是尺度箭頭,表示:

允許從尺度 s 轉譯到尺度 t.\text{允許從尺度 }s\text{ 轉譯到尺度 }t.

不同箭頭可以有不同型別,例如:

refine,\mathrm{refine}, aggregate,\mathrm{aggregate}, forget,\mathrm{forget}, project,\mathrm{project}, compare.\mathrm{compare}.

所以:

sts\rightarrow t

不必只代表「變大」。


4. 判定域在尺度基底上的排列

Paper 02 的判定域為:

D=(X,Σ;r,s,c).\mathfrak D = (X,\Sigma;r,s,c).

現在把所有判定域組成:

JDom.\mathbf{JDom}.

定義尺度投影:

π:JDomS.\boxed{ \pi: \mathbf{JDom} \rightarrow \mathbf S. }

滿足:

π(D)=s.\pi(\mathfrak D)=s.

4.1 尺度纖維

對每個尺度:

s,s,

定義:

JDoms=π1(s).\boxed{ \mathbf{JDom}_s = \pi^{-1}(s). }

同一尺度上可以有很多不同判定域。

例如 micro scale 中可以同時有:

Dμ,1=單一 token,\mathfrak D_{\mu,1} = \text{單一 token}, Dμ,2=單一 agent action,\mathfrak D_{\mu,2} = \text{單一 agent action}, Dμ,3=單一個體事件.\mathfrak D_{\mu,3} = \text{單一個體事件}.

因此:

scalesingle domain.\boxed{ \text{scale} \neq \text{single domain}. }

4.2 不是先宣稱為 fiber bundle

嚴格微分幾何的 fiber bundle 需要 topology、local triviality 等結構。

本文目前只建立:

bundle-like indexed family.\boxed{ \text{bundle-like indexed family}. }

只有在未來證明:

  • topology;
  • local trivializations;
  • transition maps;
  • compatibility;

之後,才能升格為更嚴格的 fiber-bundle terminology。


5. 每個判定域再承載一個概率纖維

對每個:

D=(X,Σ;κ),\mathfrak D = (X,\Sigma;\kappa),

定義:

Prob(D)\mathbf{Prob}(\mathfrak D)

為在該判定域形式 carrier 上允許的 probability objects。

在標準 Kolmogorov 情況:

Prob(D)={P:P is a probability measure on (X,Σ)}.\mathbf{Prob}(\mathfrak D) = \{ P: P\text{ is a probability measure on }(X,\Sigma) \}.

於是整體結構至少具有兩級索引:

sDP.s \mapsto \mathfrak D \mapsto P.

可以寫成:

SJDomProbObj.\boxed{ \mathbf S \leftarrow \mathbf{JDom} \leftarrow \mathbf{ProbObj}. }

其中:

ProbObj\mathbf{ProbObj}

暫時表示所有帶判定域型別的概率對象。


6. 三種不同的幾何距離

本文不使用一個「萬能距離」把所有東西硬塞在一起,而分成三層。

6.1 尺度距離

若尺度圖具有 edge weights:

w(e)0,w(e)\ge0,

可以定義結構距離:

dS(s,t)=infγ:steγw(e).d_S(s,t) = \inf_{\gamma:s\leadsto t} \sum_{e\in\gamma}w(e).

這只是尺度圖上的 graph / path distance。

它不等於 probability distance。

6.2 纖維內概率距離

若:

P,QProb(D),P,Q\in\mathbf{Prob}(\mathfrak D),

且該判定域具有適當結構,可以使用既有:

  • Fisher--Rao distance;
  • Kullback--Leibler divergence;
  • Jensen--Shannon divergence;
  • total variation;
  • Hellinger distance;
  • Wasserstein distance。

記為:

dD(P,Q).d_{\mathfrak D}(P,Q).

本文不規定所有判定域必須使用同一種距離。

6.3 跨域 comparison cost

若:

PProb(D),P\in\mathbf{Prob}(\mathfrak D), QProb(E),Q\in\mathbf{Prob}(\mathfrak E),

PPQQ 可能連 sample space 都不同,不能直接計算:

D(P,Q).D(P,Q).

因此先選共同比較域:

C,\mathfrak C,

以及 transports:

TDC,T_{\mathfrak D\rightarrow\mathfrak C}, TEC.T_{\mathfrak E\rightarrow\mathfrak C}.

定義:

CC(P,Q)=DC(TDCP,TECQ)+λ[c(TDC)+c(TEC)].C_{\mathfrak C}(P,Q) = D_{\mathfrak C} \left( T_{\mathfrak D\rightarrow\mathfrak C}P, T_{\mathfrak E\rightarrow\mathfrak C}Q \right) + \lambda \left[ c(T_{\mathfrak D\rightarrow\mathfrak C}) + c(T_{\mathfrak E\rightarrow\mathfrak C}) \right].

再定義:

CJ(P,Q)=infC,T,UCC(P,Q).\boxed{ C_J(P,Q) = \inf_{\mathfrak C,T,U} C_{\mathfrak C}(P,Q). }

目前本文只稱其為:

cross-domain comparison cost.\boxed{ \text{cross-domain comparison cost}. }

不無條件宣稱它滿足 metric axioms。


7. 為什麼跨域距離不能直接硬定義?

假設:

PtokenP_{\mathrm{token}}

是 token sequence space 上的 distribution,

而:

PstrategyP_{\mathrm{strategy}}

是 strategy labels 上的 distribution。

裸寫:

W2(Ptoken,Pstrategy)W_2( P_{\mathrm{token}}, P_{\mathrm{strategy}} )

通常沒有定義,因為兩者不在同一 metric carrier。

因此任何跨尺度概率距離,都必須先回答:

比較在哪裡發生?\boxed{ \text{比較在哪裡發生?} }

這就是共同比較域:

C.\mathfrak C.

換句話說:

cross-scale distance requires a comparison witness.\boxed{ \text{cross-scale distance requires a comparison witness}. }

這是 Paper 03 的 lifting witness 在幾何層的延伸。


8. Coarse-graining 與 information loss

若:

K:XYK:X\rightsquigarrow Y

為 stochastic channel,則大量 divergence 滿足 data-processing inequality。

對適當 ff -divergence:

Df(KPKQ)Df(PQ).D_f( K_{\star}P \Vert K_{\star}Q ) \le D_f( P\Vert Q ).

因此可定義 coarse-graining distinguishability loss:

Lf(K;P,Q)=Df(PQ)Df(KPKQ).\boxed{ L_f( K;P,Q ) = D_f(P\Vert Q) - D_f( K_{\star}P \Vert K_{\star}Q ). }

由 data processing:

Lf(K;P,Q)0.L_f( K;P,Q ) \ge0.

這個非負性不是本文新定理,而是既有 data-processing principle 的直接結果。

本文的新用途是把它放入 scale-lifting ledger。


9. 命題一:串接 stochastic channels 的 loss 可望遠鏡分解

設:

K1:XY,K_1:X\rightsquigarrow Y, K2:YZ.K_2:Y\rightsquigarrow Z.

定義:

P1=(K1)P,P_1=(K_1)_{\star}P, Q1=(K1)Q.Q_1=(K_1)_{\star}Q.

則:

Lf(K2K1;P,Q)=Df(PQ)Df((K2)P1(K2)Q1).L_f( K_2\circ K_1;P,Q ) = D_f(P\Vert Q) - D_f( (K_2)_{\star}P_1 \Vert (K_2)_{\star}Q_1 ).

加入並減去:

Df(P1Q1),D_f(P_1\Vert Q_1),

得到:

Lf(K2K1;P,Q)=[Df(PQ)Df(P1Q1)]+[Df(P1Q1)Df((K2)P1(K2)Q1)].L_f( K_2\circ K_1;P,Q ) = \left[ D_f(P\Vert Q) - D_f(P_1\Vert Q_1) \right] + \left[ D_f(P_1\Vert Q_1) - D_f( (K_2)_{\star}P_1 \Vert (K_2)_{\star}Q_1 ) \right].

因此:

Lf(K2K1;P,Q)=Lf(K1;P,Q)+Lf(K2;P1,Q1).\boxed{ L_f( K_2\circ K_1;P,Q ) = L_f(K_1;P,Q) + L_f(K_2;P_1,Q_1). } \boxed{\square}

對:

KnK1,K_n\circ\cdots\circ K_1,

可遞推得到:

Lf(KnK1;P,Q)=j=1nLf(Kj;Pj1,Qj1).\boxed{ L_f( K_n\circ\cdots\circ K_1;P,Q ) = \sum_{j=1}^{n} L_f( K_j;P_{j-1},Q_{j-1} ). }

其中:

Pj=(Kj)Pj1,P_j=(K_j)_{\star}P_{j-1}, Qj=(Kj)Qj1.Q_j=(K_j)_{\star}Q_{j-1}.

這提供:

cumulative scale-information-loss ledger.\boxed{ \text{cumulative scale-information-loss ledger}. }

10. Scale Distortion

information loss 只量測某種 distinguishability 被 channel 壓縮多少。

但尺度提升還可能有另一種失真:

不同提升路徑得到不同 target probability.\text{不同提升路徑得到不同 target probability}.

因此本文區分:

loss\boxed{ \text{loss} }

與:

path distortion.\boxed{ \text{path distortion}. }

前者可以即使所有 transport 都完全合法仍發生。

後者則表示:

尺度路徑本身影響結果.\boxed{ \text{尺度路徑本身影響結果}. }

11. 路徑 transport

令尺度/判定域圖中存在 path:

γ:D0D1Dn.\gamma : \mathfrak D_0 \rightarrow \mathfrak D_1 \rightarrow \cdots \rightarrow \mathfrak D_n.

每條 edge:

ei:Di1Die_i: \mathfrak D_{i-1} \rightarrow \mathfrak D_i

具有 transport:

Tei.T_{e_i}.

定義 path transport:

Tγ=TenTe1.\boxed{ T_{\gamma} = T_{e_n} \circ \cdots \circ T_{e_1}. }

因此:

Pn=Tγ(P0).P_n = T_{\gamma}(P_0).

這讓多尺度概率不再只是一組分布,而是一個:

probability-flow network.\boxed{ \text{probability-flow network}. }

12. 路徑依賴

若存在兩條不同 path:

γ1,γ2:DE,\gamma_1, \gamma_2: \mathfrak D \leadsto \mathfrak E,

則可比較:

Tγ1(P)T_{\gamma_1}(P)

與:

Tγ2(P).T_{\gamma_2}(P).

選擇 target-domain divergence:

DE,D_{\mathfrak E},

定義:

Δγ1,γ2(P)=DE(Tγ1(P),Tγ2(P)).\boxed{ \Delta_{\gamma_1,\gamma_2}(P) = D_{\mathfrak E} \left( T_{\gamma_1}(P), T_{\gamma_2}(P) \right). }

若:

Δγ1,γ2(P)=0\Delta_{\gamma_1,\gamma_2}(P)=0

對所有 admissible PP 成立,則稱兩條 path 在概率 transport 上等價。

若:

Δγ1,γ2(P)>0,\Delta_{\gamma_1,\gamma_2}(P)>0,

則存在:

path dependence.\boxed{ \text{path dependence}. }

13. 命題二:Functorial transport 導致尺度鏈路徑獨立

假設:

S\mathbf S

是一個偏序 category。

因此若:

st,s\preceq t,

則 category 中至多只有一個 morphism:

st.s\rightarrow t.

又假設 probability transport assignment 構成 functor:

F:SProbTrans.F: \mathbf S \rightarrow \mathbf{ProbTrans}.

取任意兩條 refinement chains:

s=s0s1sn=t,s=s_0\preceq s_1\preceq\cdots\preceq s_n=t,

以及:

s=t0t1tm=t.s=t_0\preceq t_1\preceq\cdots\preceq t_m=t.

在偏序 category 中,兩條 composite morphisms 都等於唯一:

st.s\rightarrow t.

由 functoriality:

F(st)=F(sn1sn)F(s0s1),F(s\rightarrow t) = F(s_{n-1}\rightarrow s_n) \circ \cdots \circ F(s_0\rightarrow s_1),

也等於:

F(tm1tm)F(t0t1).F(t_{m-1}\rightarrow t_m) \circ \cdots \circ F(t_0\rightarrow t_1).

故:

Tγ1=Tγ2.\boxed{ T_{\gamma_1} = T_{\gamma_2}. }

因此:

Δγ1,γ2(P)=0\boxed{ \Delta_{\gamma_1,\gamma_2}(P)=0 }

對所有 PP 成立。

\boxed{\square}

13.1 解讀

這個命題不是說真實世界必然 path-independent。

它告訴我們:

如果我們希望尺度 aggregation 真的是一個單純的 functorial coarse-graining,則不同 refinement chains 必須 commute。

若實驗發現:

Δγ1,γ2(P)>0,\Delta_{\gamma_1,\gamma_2}(P)>0,

則至少有一個假設不成立:

  • scale base 不是單純 poset;
  • arrows 具有不同語義;
  • transport assignment 不是 functor;
  • intermediate context 改變了模型;
  • aggregation 本身具有 path dependence。

14. 閉路 defect 與概率 holonomy 類比

如果存在 loop:

γ:DD,\gamma: \mathfrak D \leadsto \mathfrak D,

定義:

Hγ(P)=DD(P,Tγ(P)).\boxed{ H_{\gamma}(P) = D_{\mathfrak D} \left( P,T_{\gamma}(P) \right). }

若:

Hγ(P)=0,H_{\gamma}(P)=0,

則這個閉路在概率對象 PP 上沒有產生可觀測 defect。

若:

Hγ(P)>0,H_{\gamma}(P)>0,

則沿一圈尺度/語境 transport 後,概率沒有回到原點。

本文暫稱:

probability holonomy defect.\boxed{ \text{probability holonomy defect}. }

必須強調:這是幾何類比與新定義,不是在宣稱已建立 differential-geometric holonomy theorem。

真正使用「holonomy」作為嚴格幾何術語之前,未來還需建立 connection、parallel transport 與適當 bundle structure。


15. 微觀、中觀、宏觀不應只是一條硬編碼階梯

先前研究常用:

μmM.\mu \rightarrow m \rightarrow M.

這在工程上很有用,但本文把它一般化。

一個 macro domain 可能同時接收:

Dμ,1,Dμ,2,,Dμ,n\mathfrak D_{\mu,1}, \mathfrak D_{\mu,2}, \ldots, \mathfrak D_{\mu,n}

的輸入。

因此更接近:

{Dμ,i}i=1nDmDM.\{ \mathfrak D_{\mu,i} \}_{i=1}^{n} \rightarrow \mathfrak D_m \rightarrow \mathfrak D_M.

甚至不同 meso clustering:

C1,C2C_1, C_2

可能得到不同:

Dm(1),Dm(2).\mathfrak D_{m}^{(1)}, \qquad \mathfrak D_{m}^{(2)}.

最後造成:

PM(1)PM(2).P_M^{(1)} \neq P_M^{(2)}.

因此:

meso layer is often a modeling choice, not a uniquely given natural layer.\boxed{ \text{meso layer is often a modeling choice, not a uniquely given natural layer}. }

這正是 path-dependence 分析的重要來源。


16. 多尺度 aggregation 的兩步結構

從 micro 到 macro 不應直接寫:

PμPM.P_{\mu} \rightarrow P_M.

一般至少需要:

第一步:形成 joint / relational structure

{Pμ,i}+dependence modelQμ(n).\{ P_{\mu,i} \} + \text{dependence model} \rightarrow Q_{\mu}^{(n)}.

第二步:形成 macro observable

g:XμnXM,g: X_{\mu}^{n} \rightarrow X_M,

因此:

PM=g#Qμ(n).P_M = g_{\#}Q_{\mu}^{(n)}.

所以:

micro probabilities+dependency geometry+aggregation operator=macro probability.\boxed{ \text{micro probabilities} + \text{dependency geometry} + \text{aggregation operator} = \text{macro probability}. }

這也是 Paper 03 的 Aggregation Witness 在幾何層的重新表達。


17. Scale Chart

為了逐步接近真正的 atlas,本文提出弱版 chart。

定義 17.1:Scale Chart

一個 scale chart 為:

U=(U,ZU,{TDZU}),\mathcal U = (U,Z_U,\{T_{\mathfrak D\rightarrow Z_U}\}),

其中:

  • UOb(JDom)U\subseteq\operatorname{Ob}(\mathbf{JDom}) 是一組鄰近或可共同比較的判定域;
  • ZUZ_U 是共同 representation / comparison space;
  • 每個 DU\mathfrak D\in U 都有 admissible transport:
TDZU.T_{\mathfrak D\rightarrow Z_U}.

chart 的目的,是讓:

{PD}DU\{ P_{\mathfrak D} \}_{\mathfrak D\in U}

可以被 transport 到共同空間:

ZUZ_U

中比較。

17.2 Chart overlap

若:

UV,U\cap V\neq\varnothing,

則需要 transition relation:

ΦUV:ZUZV.\Phi_{UV}: Z_U \rightsquigarrow Z_V.

未來若:

  • transitions 可逆;
  • cocycle conditions 成立;
  • topology 已建立;

才有資格討論更嚴格的 atlas / bundle geometry。

因此 v0.1.0 只把:

chart\boxed{ \text{chart} }

當成「共同比較坐標」概念。


18. 全域不是默認,而是結構

Paper 02 已提出 Explicit Globality Principle。

本文給出更強版本。

定義 18.1:Study-Global Domain

對特定研究 scope,若:

DG\mathfrak D_G

JDom\mathbf{JDom} 中一個 designated terminal object,使所有 relevant domains:

Di\mathfrak D_i

都存在:

DiDG,\mathfrak D_i \rightarrow \mathfrak D_G,

則稱:

DG\mathfrak D_G

為該研究的 study-global domain。

18.1 若沒有 terminal object?

可能存在兩個 incomparable maximal domains:

DM,1,DM,2.\mathfrak D_{M,1}, \qquad \mathfrak D_{M,2}.

且沒有 canonical:

DG.\mathfrak D_G.

那麼:

the study has no canonical global probability.\boxed{ \text{the study has no canonical global probability}. }

這不是概率論失敗,而是 domain geometry 本身不存在唯一 global object。

18.2 相對全域

因此「global」永遠是相對於:

declared study universe.\boxed{ \text{declared study universe}. }

一個研究中的 global,在更大的研究中可能只是 meso 或 local。

這與微、中、宏尺度的遞歸性完全相容。


19. Globality Certificate

如果一篇研究宣稱:

PG(A)=p,P_G(A)=p,

本文建議同步提供:

G=(DG,U,{Ti},ΓG),\mathcal G = ( \mathfrak D_G, \mathcal U, \{T_i\}, \Gamma_G ),

其中:

  • DG\mathfrak D_G:global domain;
  • U\mathcal U:被宣稱覆蓋的 local / meso / macro domains;
  • {Ti}\{T_i\}:各 domain 到 global 的 lifts;
  • ΓG\Gamma_G:一致性與 coverage assumptions。

因此 global claim 不只是:

P(A)=p.\boxed{ P(A)=p. }

而是:

probability+coverage certificate.\boxed{ \text{probability} + \text{coverage certificate}. }

20. Information Geometry 在本框架中的位置

Information geometry 可以嵌入:

Prob(D)\mathbf{Prob}(\mathfrak D)

的單一 fiber。

例如對 parametric family:

MD={PθD}θΘ,\mathcal M_{\mathfrak D} = \{ P_{\theta}^{\mathfrak D} \}_{\theta\in\Theta},

可以使用 Fisher metric:

gijD.g^{\mathfrak D}_{ij}.

若存在 stochastic coarse-graining:

K:DE,K: \mathfrak D \rightarrow \mathfrak E,

則 Fisher information 在適當 Markov morphism 下具有單調性。

因此:

information geometry\boxed{ \text{information geometry} }

可以測量「一個 fiber 內部」與「transport 後信息如何縮減」。

但:

gDg^{\mathfrak D}

不自動定義:

d(D,E).d( \mathfrak D, \mathfrak E ).

這正是本文把 in-fiber geometry 與 domain geometry 分離的原因。


21. Wasserstein Geometry 在本框架中的位置

如果兩個 probability objects 已被 transport 到同一 metric carrier:

(Z,dZ),(Z,d_Z),

則:

WpW_p

可以作為非常自然的 comparison geometry。

因此 chart:

U=(U,ZU,{Ti})\mathcal U = (U,Z_U,\{T_i\})

中可以定義:

dU(Pi,Pj)=Wp(TiPi,TjPj).d_U( P_i,P_j ) = W_p( T_iP_i, T_jP_j ).

這提供:

local Wasserstein chart.\boxed{ \text{local Wasserstein chart}. }

但不同 chart 之間是否相容,仍然需要 transition maps。

這也是為什麼:

Wasserstein geometryjudgment-domain geometry.\boxed{ \text{Wasserstein geometry} \neq \text{judgment-domain geometry}. }

前者可以成為後者的 local metric ingredient。


22. Scale Curvature:暫不定義數值,只定義研究問題

既然有 path dependence:

Δγ1,γ2(P),\Delta_{\gamma_1,\gamma_2}(P),

很自然會想把它叫成「曲率」。

但 v0.1.0 暫時不直接定義:

Rscale.R_{\mathrm{scale}}.

原因是嚴格 curvature 需要更明確的:

  • connection;
  • tangent structure;
  • infinitesimal transport;
  • loop limit;
  • smoothness 或離散 curvature framework。

因此本文只提出研究問題:

判定域 probability transport 的 non-commutativity 是否可以在離散幾何、category-theoretic curvature 或 information-geometric curvature 中得到自然形式化?

目前只保留:

Hγ(P)H_{\gamma}(P)

作為 loop defect。

這能避免過早把漂亮類比宣稱成已完成 differential geometry。


23. AI 概率場:第一個具體多尺度幾何

考慮:

s0=token,s_0 = \text{token}, s1=phrase / local semantic,s_1 = \text{phrase / local semantic}, s2=response semantic,s_2 = \text{response semantic}, s3=strategy,s_3 = \text{strategy}, s4=task behavior,s_4 = \text{task behavior}, s5=model-level experimental profile.s_5 = \text{model-level experimental profile}.

這不是宣稱 AI 的唯一自然尺度。

它只是某個 experiment atlas。

每個尺度都有:

Dsi,\mathfrak D_{s_i},

以及:

Psi.P_{s_i}.

跨尺度可以有:

T01,T12,T23,T34,T45.T_{01}, T_{12}, T_{23}, T_{34}, T_{45}.

因此:

Ps5=T45T34T23T12T01(Ps0)P_{s_5} = T_{45} \circ T_{34} \circ T_{23} \circ T_{12} \circ T_{01} (P_{s_0})

只在所有中間 transports 都被明確定義時才有意義。


24. AI 中的 path dependence

例如 token sequence 可以經由兩條不同 route 形成 strategy distribution。

第一條:

tokensemantic classstrategy.\text{token} \rightarrow \text{semantic class} \rightarrow \text{strategy}.

第二條:

tokenlatent embeddingstrategy.\text{token} \rightarrow \text{latent embedding} \rightarrow \text{strategy}.

得到:

Pstrategy(A)P_{\mathrm{strategy}}^{(A)}

與:

Pstrategy(B).P_{\mathrm{strategy}}^{(B)}.

則:

ΔA,B=Dstrategy(Pstrategy(A),Pstrategy(B)).\Delta_{A,B} = D_{\mathrm{strategy}} \left( P_{\mathrm{strategy}}^{(A)}, P_{\mathrm{strategy}}^{(B)} \right).

如果:

ΔA,B0,\Delta_{A,B} \gg0,

就表示:

strategy probability is representation-path dependent.\boxed{ \text{strategy probability is representation-path dependent}. }

這是一個真正可以實驗的命題。


25. AI Probability Fingerprint 的幾何化

先前 AI probability-field 研究可使用:

Ptoken,Psemantic,Pstrategy,Ptask,Psuccess.P_{\mathrm{token}}, P_{\mathrm{semantic}}, P_{\mathrm{strategy}}, P_{\mathrm{task}}, P_{\mathrm{success}}.

現在可以把一個 model 的 fingerprint 寫成:

F(M)=(Ps0,Ps1,,Psn,{Tij},{Lij},{Δγa,γb}).\mathcal F(M) = \left( P_{s_0}, P_{s_1}, \ldots, P_{s_n}, \{T_{ij}\}, \{L_{ij}\}, \{\Delta_{\gamma_a,\gamma_b}\} \right).

所以 model fingerprint 不再只是:

一串 entropy values.\boxed{ \text{一串 entropy values}. }

而是:

a multi-scale probability geometry.\boxed{ \text{a multi-scale probability geometry}. }

其中包含:

  • 各尺度 distribution;
  • scale transports;
  • information losses;
  • path dependence;
  • approximation error;
  • context dependence。

26. 微—中—宏的遞歸性

某一層的 macro object 可以在更大系統中成為 micro object。

例如:

individual tokenresponseconversationagentmulti-agent system.\text{individual token} \rightarrow \text{response} \rightarrow \text{conversation} \rightarrow \text{agent} \rightarrow \text{multi-agent system}.

對 conversation 而言,response 是 micro / meso。

對 multi-agent system 而言,整個 agent 仍可能只是 micro component。

因此尺度不是絕對 label。

更精確應表示:

s=s(U),s = s(\mathcal U),

即:

scale is relative to a declared universe of analysis.\boxed{ \text{scale is relative to a declared universe of analysis}. }

這與本文「全域也是相對 study universe」的結論一致。


27. 多尺度概率矩陣

工程上,可以把一個研究表示成:

P=[Pμ,1Pμ,2Pm,1Pm,2PM,1PM,2PG,1PG,2].\mathbf P = \begin{bmatrix} P_{\mu,1} & P_{\mu,2} & \cdots \\ P_{m,1} & P_{m,2} & \cdots \\ P_{M,1} & P_{M,2} & \cdots \\ P_{G,1} & P_{G,2} & \cdots \end{bmatrix}.

但這個矩陣只是一個索引表。

真正結構還包括 transport tensor / edge family:

T={Tij},\mathbf T = \{ T_{ij} \},

以及 loss:

L={Lij},\mathbf L = \{ L_{ij} \},

path defects:

Δ={Δγa,γb}.\mathbf \Delta = \{ \Delta_{\gamma_a,\gamma_b} \}.

因此完整對象更接近:

GP=(JDom,P,T,L,Δ).\boxed{ \mathfrak G_P = ( \mathbf{JDom}, \mathbf P, \mathbf T, \mathbf L, \mathbf\Delta ). }

本文稱其為:

Judgment-Domain Probability Geometry.\boxed{ \text{Judgment-Domain Probability Geometry}. }

28. Geometry Ledger

每個 domain node 至少保存:

domain_id
reference_scope
scale
context
formal_space
event_structure
probability_object
intrinsic_metric_or_divergence

每個 edge 至少保存:

source_domain
target_domain
transport_type
transport_witness
transport_cost
information_loss
approximation_error

每個 path 可以保存:

path_id
nodes
edges
composite_transport
cumulative_loss
target_probability

若存在 alternative path:

path_pair
path_defect
comparison_divergence

這讓多尺度概率幾何具有可審計工程形式。


29. 幾何合法性的四個層次

G0:Indexed

只有 domain / scale indexing。

G1:Connected

已有 admissible transports。

G2:Comparable

部分 fibers 已具有 comparison charts、metrics 或 divergences。

G3:Geometric

已建立足夠:

  • topology;
  • metric / divergence structure;
  • path composition;
  • transition consistency。

G4:Bundle / Connection Level

若進一步建立:

  • local trivialization;
  • connection;
  • parallel transport;
  • curvature / holonomy;

才進入真正更強的 fiber-bundle geometry。

因此本文目前主要位於:

G1G2,\boxed{ G1\text{--}G2, }

並提出通往 G3G3G4G4 的研究綱領。

這樣可以避免把「畫一張尺度圖」過度聲稱為完整幾何理論。


30. 新穎性邊界

本文不宣稱首次提出:

  • Fisher information geometry;
  • probability manifolds;
  • Wasserstein geometry;
  • optimal transport;
  • multiscale Wasserstein analysis;
  • data-processing inequality;
  • Markov morphism monotonicity;
  • stochastic coarse-graining;
  • projective systems;
  • fiber bundles;
  • categorical probability。

本文提出的是:

把判定域、尺度、概率纖維與 transport witness 統一成一個幾何化研究對象.\boxed{ \text{把判定域、尺度、概率纖維與 transport witness 統一成一個幾何化研究對象}. }

核心新增概念候選包括:

  1. judgment-domain scale base;
  2. judgment-domain fibers;
  3. probability fibers over judgment domains;
  4. cross-domain comparison witness;
  5. path transport;
  6. path-dependence defect;
  7. probability holonomy defect;
  8. cumulative scale-information-loss ledger;
  9. study-global domain / globality certificate;
  10. geometry-legality levels。

這些目前是 framework proposal,而不是宣稱已形成標準幾何分支。


31. 待證猜想與後續定理

31.1 Minimal Globality Conjecture

若尺度 category 不存在 terminal object,則任何「唯一 canonical global probability」都必須額外依賴某個 completion 或 selection principle。

31.2 Scale Path Equivalence Problem

給定兩條:

γ1,γ2:DE,\gamma_1,\gamma_2: \mathfrak D\leadsto\mathfrak E,

判定:

Tγ1=Tγ2T_{\gamma_1} = T_{\gamma_2}

的最小充分條件是什麼?

31.3 Distortion Accumulation Theorem

若每條 edge 只有 approximate transport:

d(TeP,Qe)εe,d( T_eP, Q_e ) \le \varepsilon_e,

則一條長 path 的總幾何失真可以得到何種 sharp bound?

31.4 Chart Compatibility Problem

何時兩個 comparison charts:

(U,ZU),(U,Z_U), (V,ZV)(V,Z_V)

能在 overlap 上形成一致 probability geometry?

31.5 Holonomy Classification Problem

非零:

Hγ(P)H_{\gamma}(P)

究竟來自:

  • representation change;
  • information loss;
  • noncommuting aggregation;
  • context shift;
  • stochasticity;

中的哪一種結構?


32. 通往 Paper 05:遞歸概率場

到目前為止,一個判定域:

D\mathfrak D

承載一階概率:

PD.P_{\mathfrak D}.

但最初的研究問題還包含:

概率場的概率場的概率場\boxed{ \text{概率場的概率場的概率場}\ldots }

因此下一篇將加入 probability order:

r=0,1,2,r = 0,1,2,\ldots

並定義:

P0(X)=X,\mathsf P^0(X) = X, Pr+1(X)=P(Pr(X)).\mathsf P^{r+1}(X) = \mathsf P( \mathsf P^r(X) ).

到那時本文的尺度軸:

ss

與概率階軸:

rr

會真正交叉成:

(s,r).\boxed{ (s,r). }

如果再加入時間:

t,t,

便得到:

Ps,t(r).\boxed{ \mathfrak P_{s,t}^{(r)}. }

因此 Paper 04 的真正功能,是先建立:

概率階數進入之前的尺度幾何底座.\boxed{ \text{概率階數進入之前的尺度幾何底座}. }

33. 結論

本文把判定域概率論從:

單域\text{單域}

推向:

域網路.\text{域網路}.

從:

單箭頭 lifting\text{單箭頭 lifting}

推向:

path transport.\text{path transport}.

從:

local/global 二分\text{local/global 二分}

推向:

multi-scale geometry.\text{multi-scale geometry}.

整體結構可概括為:

SJDomProbObj.\boxed{ \mathbf S \leftarrow \mathbf{JDom} \leftarrow \mathbf{ProbObj}. }

其中尺度基底 S\mathbf S 組織 micro、meso、macro 與 study-global 關係;每個尺度上可以有多個判定域;每個判定域再承載 probability objects。

概率分布內部已有成熟的 Fisher / information geometry 與 Wasserstein / optimal-transport geometry。

本文真正新增的研究方向則是:

geometry between the domains that host those probabilities.\boxed{ \text{geometry between the domains that host those probabilities}. }

跨尺度比較必須先找到共同 comparison domain。

跨尺度提升必須提供 transport witness。

coarse-graining 的 information loss 可以由 data-processing divergence loss 記錄。

多步 stochastic transport 的 loss 可以望遠鏡式分解。

不同 aggregation paths 若產生不同結果,則可由:

Δγ1,γ2(P)\Delta_{\gamma_1,\gamma_2}(P)

測量 path dependence。

閉路後無法回復原分布,則由:

Hγ(P)H_{\gamma}(P)

記錄 probability holonomy defect。

最重要的是,全域不再被當成:

沒有下標的概率.\text{沒有下標的概率}.

而是:

一個被明示、被覆蓋、被證明可由研究域接入的 global judgment domain.\boxed{ \text{一個被明示、被覆蓋、被證明可由研究域接入的 global judgment domain}. }

若不存在唯一 terminal / designated maximal domain,就不應假定存在唯一 canonical global probability。

因此截至 Paper 04,本系列已形成:

HistoryJudgment DomainLifting CalculusMulti-Scale Probability Geometry.\boxed{ \text{History} \rightarrow \text{Judgment Domain} \rightarrow \text{Lifting Calculus} \rightarrow \text{Multi-Scale Probability Geometry}. }

下一步將沿著第二個獨立方向展開:

Probability Order.\boxed{ \text{Probability Order}. }

也就是:

Recursive Probability Fields.\boxed{ \text{Recursive Probability Fields}. }

參考文獻

[1] Neo.K. (2026). 《判定域概率論:概率之前的容器、空間與量詞》. JDPSP-02, EveMissLab.

[2] Neo.K. (2026). 《局部概率與全域概率:尺度提升的合法性與判定域提升演算》. JDPSP-03, EveMissLab.

[3] Ay, N., Jost, J., Lê, H. V., & Schwachhöfer, L. (2017). Information Geometry. Springer. Earlier foundational paper: Information Geometry and Sufficient Statistics, arXiv:1207.6736.

[4] Lê, H. V. (2017). The Uniqueness of the Fisher Metric as Information Metric. Annals of the Institute of Statistical Mathematics, 69, 879--896. Preprint: arXiv:1306.1465.

[5] Panaretos, V. M., & Zemel, Y. (2019). Statistical Aspects of Wasserstein Distances. Annual Review of Statistics and Its Application, 6, 405--431. Preprint: arXiv:1806.05500.

[6] Seguy, V., & Cuturi, M. (2015). Principal Geodesic Analysis for Probability Measures under the Optimal Transport Metric. arXiv:1506.07944.

[7] Itoh, M., & Satoh, H. (2022). Information Geometry of the Space of Probability Measures and Barycenter Maps. arXiv:2208.11861.

[8] Perrone, P. (2022). Markov Categories and Entropy. arXiv:2212.11719.

[9] Hilder, B., & Sharma, U. (2022). Quantitative Coarse-Graining of Markov Chains. arXiv:2201.10256.

[10] Fritz, T., Klingler, A., McNeely, D., Shah-Mohammed, A., & Wang, Y. (2024). Hidden Markov Models and the Bayes Filter in Categorical Probability. arXiv:2401.14669.

[11] Mattar, W., & Sharon, N. (2025). Multiscaling in Wasserstein Spaces. arXiv:2509.10415.

[12] Maity, R., Mahapatra, S., & Sarkar, T. (2015). Information Geometry and the Renormalization Group. arXiv:1503.03978.

[13] Berman, D. S., & Klinger, M. S. (2022). The Inverse of Exact Renormalization Group Flows as Statistical Inference. arXiv:2212.11379.


Appendix A. 最小符號表

符號 意義
S\mathbf S 尺度 category / base
JDom\mathbf{JDom} 判定域 category / indexed domain family
π:JDomS\pi:\mathbf{JDom}\to\mathbf S 尺度投影
JDoms\mathbf{JDom}_s 尺度 ss 上的判定域纖維
Prob(D)\mathbf{Prob}(\mathfrak D) 判定域 D\mathfrak D 上的概率對象集合
dSd_S 尺度結構距離
dDd_{\mathfrak D} 判定域 fiber 內的 probability distance/divergence
CJC_J 跨判定域 comparison cost
TeT_e edge transport
TγT_{\gamma} path transport
LfL_f stochastic coarse-graining distinguishability loss
Δγ1,γ2\Delta_{\gamma_1,\gamma_2} path-dependence defect
HγH_{\gamma} loop / probability-holonomy defect
DG\mathfrak D_G study-global judgment domain
G\mathcal G globality certificate

Appendix B. 幾何合法性分級

G0G1G2G3G4.G0 \rightarrow G1 \rightarrow G2 \rightarrow G3 \rightarrow G4.
  • G0G0:Indexed;
  • G1G1:Connected by admissible transports;
  • G2G2:Comparable through charts / metrics / divergences;
  • G3G3:具有穩定 topology、path composition 與 transition consistency;
  • G4G4:具有 bundle / connection / parallel transport / curvature 類結構。

Paper 04 v0.1.0 主要聲稱建立:

G1G2G1\text{--}G2

的形式化接口,而非宣稱已完成 G4G4 微分幾何。

Appendix C. Canonical Geometry Record

node:
  domain_id:
  reference_scope:
  scale:
  context:
  formal_space:
  event_structure:
  probability_object:
  intrinsic_metric_or_divergence:

edge:
  source_domain:
  target_domain:
  transport_type:
  transport_witness:
  transport_cost:
  information_loss:
  approximation_error:

path:
  path_id:
  nodes:
  edges:
  composite_transport:
  cumulative_loss:
  target_probability:

path_comparison:
  path_a:
  path_b:
  target_domain:
  comparison_divergence:
  path_defect:

globality:
  global_domain:
  covered_domains:
  lifting_witnesses:
  coverage_conditions:
  terminal_or_designated_maximal: