GIRA-A03|資訊海不是世界模型:去重、版本、時態、語義與 X 次結構化
The Information Ocean Is Not a World Model: Deduplication, Versioning, Temporality, Semantics, and X-Order Structuring
系列: Global Intelligence: Existence, Recognition, and Operational Reach(GIRA)系列中文名: 全域智能:存在、識別與操作域系列篇次: Paper 03 / 09作者: Neo.K研究協作: Aletheia(GPT-5.6 Sol)機構: EveMissLab/一言諾科技有限公司版本: v0.1日期: 2026-09-04狀態: Canonical Source / UTF-8 Markdown文件性質: Global AI 資訊架構/動態世界狀態/語義演化/知識工程
摘要
GIRA-A01 已區分 ASI 與 Global AI,指出超級智能能力不自動等於全域操作架構;GIRA-A02 則進一步提出 Global Cognitive Atlas,主張全域認知不應被理解為單一巨大表示,而應維持多觀察者、多方法、多表示的有效域、轉換、邊界與 coherence。本文處理第三個基礎問題:即使一個 AI 已能存取極大規模的全球網路、資料庫、新聞、論文、企業資訊、感測資料與歷史檔案,它仍然不因此擁有世界模型。
本文提出:
Information Ocean ≠ World Model . \boxed{
\text{Information Ocean}
\neq
\text{World Model}.
} Information Ocean = World Model .
全球資訊具有高度重複、轉述、版本分歧、時間錯位、語義異名、來源衝突、推論混入事實、過期狀態與不同觀察尺度。若 AI 只是持續收集與召回,它得到的可能是更大的資訊海,而不是更高品質的世界認知。真正的 Global AI 必須把原始資訊經過多次、不同目的的結構轉換,使之成為可追蹤、可更新、可比較、可查詢、可驗證且可用於決策的世界狀態。
本文將原始資訊量記為:
D r a w ( t ) , D_{\mathrm{raw}}(t), D raw ( t ) ,
而將有效世界資訊記為:
D e f f ( t ) . D_{\mathrm{eff}}(t). D eff ( t ) .
一般而言:
D e f f ( t ) ≪ D r a w ( t ) . \boxed{
D_{\mathrm{eff}}(t)
\ll
D_{\mathrm{raw}}(t).
} D eff ( t ) ≪ D raw ( t ) .
這個差距並不表示大部分資料「沒有價值」,而是表示大部分資料不能直接作為世界狀態使用。它們必須先經過 identity resolution、deduplication、source binding、temporal normalization、claim separation、provenance、semantic alignment、conflict preservation、versioning 與 state extraction。
本文將這個過程概念化為 X-Order Structuring :
S ( 0 ) → S ( 1 ) → S ( 2 ) → ⋯ → S ( x ) , S^{(0)}
\rightarrow
S^{(1)}
\rightarrow
S^{(2)}
\rightarrow
\cdots
\rightarrow
S^{(x)}, S ( 0 ) → S ( 1 ) → S ( 2 ) → ⋯ → S ( x ) ,
其中 x x x 不是固定常數,而是依任務、領域、時間尺度與所需精度動態決定。最簡化地:
S ( 0 ) = raw signals , S^{(0)}
=
\text{raw signals}, S ( 0 ) = raw signals ,
S ( 1 ) = canonicalized entities and sources , S^{(1)}
=
\text{canonicalized entities and sources}, S ( 1 ) = canonicalized entities and sources ,
S ( 2 ) = claims, events, relations, and versions , S^{(2)}
=
\text{claims, events, relations, and versions}, S ( 2 ) = claims, events, relations, and versions ,
S ( 3 ) = temporal state and dependency structure , S^{(3)}
=
\text{temporal state and dependency structure}, S ( 3 ) = temporal state and dependency structure ,
S ( 4 ) = causal, strategic, and systemic abstractions , S^{(4)}
=
\text{causal, strategic, and systemic abstractions}, S ( 4 ) = causal, strategic, and systemic abstractions ,
S ( 5 ) = query-conditioned active world projection . S^{(5)}
=
\text{query-conditioned active world projection}. S ( 5 ) = query-conditioned active world projection .
本文強調,真正的高階結構化不是「壓縮越多越好」。每一層都可能丟失資訊,因此必須保留 canonical source、provenance、representation identity 與 rebuildability。這與既有 Multi-Representation Memory Fabric 的原則一致:
R i ( m ) ≠ m . \boxed{
R_i(m)
\neq
m.
} R i ( m ) = m .
任何向量、圖、矩陣、摘要、狀態表、embedding 或 active context,都只是受治理記憶的 representation,而不是 memory identity 或 epistemic authority 本身。
本文亦承接 SEDB 的事件溯源與主張帳本模型,採用:
K t = ( V t , E t , C t , P t , T t , Λ t , Δ t ) \boxed{
\mathcal K_t
=
(
V_t,
E_t,
C_t,
P_t,
T_t,
\Lambda_t,
\Delta_t
)
} K t = ( V t , E t , C t , P t , T t , Λ t , Δ t )
作為動態知識狀態的參考骨架,並將狀態更新寫為:
K t + 1 = U ( K t , Δ t ) . \boxed{
\mathcal K_{t+1}
=
\mathcal U(
\mathcal K_t,
\Delta_t
).
} K t + 1 = U ( K t , Δ t ) .
Global AI 因而不應在每次新資料出現時重新「理解整個 Internet」,而應持續估計:
Δ W t = W t + 1 − W t , \boxed{
\Delta W_t
=
W_{t+1}
-
W_t,
} Δ W t = W t + 1 − W t ,
即世界究竟改變了什麼。
本文最後提出 Global State Refinement Pipeline :
Observe → Preserve → Resolve → Deduplicate → Temporalize → Semanticize → Relate → Abstract → Project → Verify . \boxed{
\text{Observe}
\rightarrow
\text{Preserve}
\rightarrow
\text{Resolve}
\rightarrow
\text{Deduplicate}
\rightarrow
\text{Temporalize}
\rightarrow
\text{Semanticize}
\rightarrow
\text{Relate}
\rightarrow
\text{Abstract}
\rightarrow
\text{Project}
\rightarrow
\text{Verify}.
} Observe → Preserve → Resolve → Deduplicate → Temporalize → Semanticize → Relate → Abstract → Project → Verify .
其中 Preserve 必須先於高階壓縮,以確保任何 derived world state 都可以追溯、重算與撤回。真正的 Global AI 不應擁有一個「永遠正確的真理資料庫」,而應維持一套可以知道何者是觀察、推論、假說、衝突、撤回、過期、未知與可驗證狀態的動態資訊基礎設施。
關鍵詞: Global AI、資訊海、世界模型、X 次結構化、去重、版本控制、時態資料、SEDB、語義演化、事件溯源、主張帳本、Provenance、Dynamic World State、Delta State、Multi-Representation Memory
1. 問題:如果 AI 能讀完整個網路,它就理解世界了嗎?
令網路與外部資料形成原始資訊域:
I r a w . \mathcal I_{\mathrm{raw}}. I raw .
AI 可以具有非常大的可達範圍:
A d a t a ≈ I r a w . A_{\mathrm{data}}
\approx
\mathcal I_{\mathrm{raw}}. A data ≈ I raw .
這仍然不能推出:
A d a t a = W t . \boxed{
A_{\mathrm{data}}
=
W_t.
} A data = W t .
其中 W t W_t W t 表示時間 t t t 的有效世界狀態。網路保存的是大量表述,不是天然正規化的世界狀態。
2. 一個世界事件可以生成大量資訊物件
假設世界中發生事件:
e . e. e .
它可能生成官方公告、新聞、公司聲明、社群貼文、分析文章、二次轉述、多語言翻譯、影片與後續更正。因此一個事件可以對應:
{ d 1 , d 2 , … , d n } . \{d_1,d_2,\ldots,d_n\}. { d 1 , d 2 , … , d n } .
若 AI 把每個 d i d_i d i 都當作獨立世界事實,則:
representation multiplicity → false world multiplicity . \boxed{
\text{representation multiplicity}
\rightarrow
\text{false world multiplicity}.
} representation multiplicity → false world multiplicity .
3. Entity Resolution
不同來源可能使用全名、縮寫、舊名、品牌名、子公司名、翻譯名或拼寫差異表示同一實體。因此:
n s u r f a c e > n e n t i t y . n_{\mathrm{surface}}
>
n_{\mathrm{entity}}. n surface > n entity .
Global AI 需要 canonical entity identity:
E ∗ . E^\ast. E ∗ .
並將:
e 1 , e 2 , … , e k e_1,e_2,\ldots,e_k e 1 , e 2 , … , e k
映射到:
E ∗ . E^\ast. E ∗ .
這個映射不能只靠字串相似度,還需要 time、jurisdiction、ownership、identifier、address、legal status 與 provenance。
所以:
Name Similarity ≠ Entity Identity . \boxed{
\text{Name Similarity}
\neq
\text{Entity Identity}.
} Name Similarity = Entity Identity .
4. 去重不是刪除重複文字
Global AI 所需的去重至少包括:
exact duplicate;
semantic duplicate;
event duplicate;
derivative duplicate;
version duplicate;
partial duplicate。
因此:
Deduplication ≠ Deletion . \boxed{
\text{Deduplication}
\neq
\text{Deletion}.
} Deduplication = Deletion .
更合理的是建立 equivalence、derivation 與 version relation。
5. 去重後仍應保留來源差異
假設:
d 1 ∼ d 2 d_1
\sim
d_2 d 1 ∼ d 2
表示兩份資料指向同一事件。這不代表:
d 1 = d 2 . d_1=d_2. d 1 = d 2 .
因為兩者可能有不同時間、作者、證據、wording、authority 與 correction history。
因此 canonical state 應保留:
E ∗ + { d i } + { P i } , E^\ast
+
\{d_i\}
+
\{P_i\}, E ∗ + { d i } + { P i } ,
而不是把所有來源抹平成一份摘要。
6. Claim 不等於 Fact
SEDB 已提出:
Claim ≠ Fact . \boxed{
\text{Claim}
\neq
\text{Fact}.
} Claim = Fact .
對來源 s s s 提出的命題 c c c ,應表示為:
C = ( s , p , o , scope , status , time , provenance ) . C
=
(
s,
p,
o,
\text{scope},
\text{status},
\text{time},
\text{provenance}
). C = ( s , p , o , scope , status , time , provenance ) .
而不是直接寫入 WorldFact。
7. 認識論狀態必須是一等資料
Global AI 至少應區分:
O B S ≠ I N F ≠ H Y P ≠ C O N ≠ R E T ≠ U N K . \boxed{
\mathrm{OBS}
\neq
\mathrm{INF}
\neq
\mathrm{HYP}
\neq
\mathrm{CON}
\neq
\mathrm{RET}
\neq
\mathrm{UNK}.
} OBS = INF = HYP = CON = RET = UNK .
若沒有這層,AI 很容易形成:
Inference t → Fact t + 1 , \text{Inference}_t
\rightarrow
\text{Fact}_{t+1}, Inference t → Fact t + 1 ,
經過多輪後造成自我污染。
8. 時間不是單一 timestamp
至少應區分:
t v a l i d , t r e c o r d , t p u b l i s h , t m e n t i o n , t i n f e r e n c e . t_{\mathrm{valid}},
\quad
t_{\mathrm{record}},
\quad
t_{\mathrm{publish}},
\quad
t_{\mathrm{mention}},
\quad
t_{\mathrm{inference}}. t valid , t record , t publish , t mention , t inference .
因此:
Timestamp ≠ Temporal Semantics . \boxed{
\text{Timestamp}
\neq
\text{Temporal Semantics}.
} Timestamp = Temporal Semantics .
9. 多時間系統為什麼重要?
假設政策 P P P 在 t 1 t_1 t 1 生效,但新聞在 t 0 < t 1 t_0<t_1 t 0 < t 1 提前報導。如果 AI 只保留 publish time,它可能在歷史重建時把政策提前生效。
因此:
publication history ≠ world-state history . \boxed{
\text{publication history}
\neq
\text{world-state history}.
} publication history = world-state history .
10. 版本不等於覆寫
傳統 CRUD 很容易:
v 1 → v 2 v_1
\rightarrow
v_2 v 1 → v 2
後只保留 v 2 v_2 v 2 。但 Global AI 需要知道 v 1 v_1 v 1 曾經存在、何時被取代、哪些結論基於 v 1 v_1 v 1 、 v 2 v_2 v 2 改了什麼,以及是否需要重算 downstream state。
所以:
Update ≠ Overwrite . \boxed{
\text{Update}
\neq
\text{Overwrite}.
} Update = Overwrite .
更適合:
v 1 → Δ 1 v 2 → Δ 2 v 3 . v_1
\xrightarrow{\Delta_1}
v_2
\xrightarrow{\Delta_2}
v_3. v 1 Δ 1 v 2 Δ 2 v 3 .
11. Event Sourcing 的世界歷史啟發
若每次狀態改變都保存事件:
Δ t , \Delta_t, Δ t ,
則:
W t + 1 = U ( W t , Δ t ) . W_{t+1}
=
\mathcal U(
W_t,
\Delta_t
). W t + 1 = U ( W t , Δ t ) .
本文不要求特定 Event Sourcing 實作,只要求:
State change should remain reconstructable . \boxed{
\text{State change should remain reconstructable}.
} State change should remain reconstructable .
12. 網路資訊具有高度派生性
假設原始資料:
d 0 d_0 d 0
被 d 1 d_1 d 1 引用,再被 d 2 d_2 d 2 轉述,再被 d 3 d_3 d 3 摘要。
如果 AI 只看到:
d 1 , d 2 , d 3 , d_1,d_2,d_3, d 1 , d 2 , d 3 ,
可能誤以為有三個獨立來源。
因此需要 derivation graph:
d 0 → d 1 → d 2 → d 3 . d_0
\rightarrow
d_1
\rightarrow
d_2
\rightarrow
d_3. d 0 → d 1 → d 2 → d 3 .
13. Provenance 不是附加 metadata
W3C PROV 將 provenance 表示為 Entity、Activity、Agent 與 derivation 等關係。對 Global AI 而言,provenance 應視為:
epistemic structure . \boxed{
\text{epistemic structure}.
} epistemic structure .
因為兩個內容相同的 claim,如果來源鏈不同,其可信度與獨立性可能完全不同。
14. 多來源不等於多證據
假設:
s 2 ← s 1 , s_2
\leftarrow
s_1, s 2 ← s 1 ,
s 3 ← s 1 . s_3
\leftarrow
s_1. s 3 ← s 1 .
那麼:
N s o u r c e s = 3 N_{\mathrm{sources}}=3 N sources = 3
卻可能只有:
N i n d e p e n d e n t e v i d e n c e = 1. N_{\mathrm{independent\ evidence}}=1. N independent evidence = 1.
因此:
Source Count ≠ Evidence Independence . \boxed{
\text{Source Count}
\neq
\text{Evidence Independence}.
} Source Count = Evidence Independence .
15. 歷史網路需要時間旅行
若今天 t 2 t_2 t 2 查詢過去 t 1 t_1 t 1 ,Global AI 不應只用今天頁面的最新版本回答。它應區分:
What is known now about t 1 \text{What is known now about }t_1 What is known now about t 1
與:
What was knowable at t 1 . \text{What was knowable at }t_1. What was knowable at t 1 .
HTTP Memento 以 datetime negotiation、Memento 與 TimeMap 等機制提供 Web 資源時間版本存取的標準化思路。
因此:
Historical Reconstruction ≠ Present-Day Retelling . \boxed{
\text{Historical Reconstruction}
\neq
\text{Present-Day Retelling}.
} Historical Reconstruction = Present-Day Retelling .
16. 文明動態記憶的核心是 State Difference
既有「網路資訊海作為文明動態記憶」提出:
N d ( t ) = Top3 ( Δ K d ( t ) ) . N_d(t)
=
\operatorname{Top3}(
\Delta K_d(t)
). N d ( t ) = Top3 ( Δ K d ( t )) .
其中最重要的不是 Top3,而是:
Δ K . \boxed{
\Delta K.
} Δ K .
世界持續變化時,真正稀缺的認知資源不是再讀一次全部資料,而是知道什麼變了。
17. 世界狀態與世界差分
令世界狀態:
W t . W_t. W t .
Global AI 的更新更接近:
W t + 1 = Ψ ( W t , D t + 1 , Δ D t , P t ) . \boxed{
W_{t+1}
=
\Psi(
W_t,
D_{t+1},
\Delta D_t,
P_t
).
} W t + 1 = Ψ ( W t , D t + 1 , Δ D t , P t ) .
並顯式保存:
Δ W t = W t + 1 − W t . \Delta W_t
=
W_{t+1}
-
W_t. Δ W t = W t + 1 − W t .
18. Delta-First 世界模型
本文提出:
Delta-First World Maintenance . \boxed{
\text{Delta-First World Maintenance}.
} Delta-First World Maintenance .
系統平常優先計算:
Δ W t \Delta W_t Δ W t
而不是反覆全量重建 W t W_t W t 。只有在 schema 變更、ontology 失效、大規模資料污染、新方法出現或 consistency check 失敗時,才做較大範圍 rebuild。
19. 什麼叫 X 次結構化?
正式定義:
S ( 0 ) = I r a w . S^{(0)}
=
\mathcal I_{\mathrm{raw}}. S ( 0 ) = I raw .
第 k k k 層結構化為:
S ( k + 1 ) = Φ k ( S ( k ) , Q t , M t , B t ) . \boxed{
S^{(k+1)}
=
\Phi_k(
S^{(k)},
Q_t,
M_t,
B_t
).
} S ( k + 1 ) = Φ k ( S ( k ) , Q t , M t , B t ) .
其中 Q t Q_t Q t 是當前問題, M t M_t M t 是方法集合, B t B_t B t 是資源預算。
因此:
x = x ( Q t , Ω , ϵ , B t ) . x
=
x(
Q_t,
\Omega,
\epsilon,
B_t
). x = x ( Q t , Ω , ϵ , B t ) .
20. 第一層:Canonicalization
S ( 0 ) → S ( 1 ) . S^{(0)}
\rightarrow
S^{(1)}. S ( 0 ) → S ( 1 ) .
處理 source identity、document identity、entity identity、format normalization、language mapping、exact duplicate 與 version linkage。
目標是:
知道哪些東西其實是同一個東西。 \boxed{
\text{知道哪些東西其實是同一個東西。}
} 知道哪些東西其實是同一個東西。
21. 第二層:Claim and Event Structuring
S ( 1 ) → S ( 2 ) . S^{(1)}
\rightarrow
S^{(2)}. S ( 1 ) → S ( 2 ) .
將文件拆成 event、claim、observation、actor、relation、evidence、version 與 scope。
此時 Document 不再是核心最小單位。
22. 第三層:Temporal and Relational State
S ( 2 ) → S ( 3 ) . S^{(2)}
\rightarrow
S^{(3)}. S ( 2 ) → S ( 3 ) .
建立:
G t = ( V t , E t ) G_t
=
(V_t,E_t) G t = ( V t , E t )
與多時間系統 T t T_t T t 。
系統開始能回答:誰和誰有關、何時開始、何時失效、哪條關係是推論、哪條關係有直接來源。
23. 第四層:Dependency and System Structure
S ( 3 ) → S ( 4 ) . S^{(3)}
\rightarrow
S^{(4)}. S ( 3 ) → S ( 4 ) .
這一層處理 dependency、bottleneck、substitution、flow、hierarchy、cluster、feedback 與 propagation。
它是從 knowledge graph 走向 system model 的關鍵。
24. 第五層:Causal and Strategic Abstraction
S ( 4 ) → S ( 5 ) . S^{(4)}
\rightarrow
S^{(5)}. S ( 4 ) → S ( 5 ) .
在有足夠證據與方法時,建立 causal hypothesis、intervention model、game structure、strategic option、risk propagation 與 counterfactual dependency。
但這些必須保留:
I N F \mathrm{INF} INF
或:
H Y P \mathrm{HYP} HYP
狀態。
25. 第六層:Query-Conditioned Active Projection
對查詢:
q t q_t q t
生成:
P t ( q ) = Project ( S ( ≤ x ) , q t , B t ) . \boxed{
P_t(q)
=
\operatorname{Project}(
S^{(\leq x)},
q_t,
B_t
).
} P t ( q ) = Project ( S ( ≤ x ) , q t , B t ) .
這是 working context。
因此:
World State ≠ Working Context . \boxed{
\text{World State}
\neq
\text{Working Context}.
} World State = Working Context .
26. X 次結構化不是單向 ETL
實際系統會回流。
例如 S ( 5 ) S^{(5)} S ( 5 ) 發現因果模型矛盾,可能回到 S ( 2 ) S^{(2)} S ( 2 ) 重新檢查 claim identity。
所以:
S ( k ) ↔ S ( j ) . \boxed{
S^{(k)}
\leftrightarrow
S^{(j)}.
} S ( k ) ↔ S ( j ) .
X 次結構化更接近多層可逆 refinement graph。
27. 結構化越高,資訊損失風險越高
假設:
R k : S ( 0 ) → S ( k ) . R_k:
S^{(0)}
\to
S^{(k)}. R k : S ( 0 ) → S ( k ) .
通常:
Info ( S ( k ) ) < Info ( S ( 0 ) ) , \operatorname{Info}(S^{(k)})
<
\operatorname{Info}(S^{(0)}), Info ( S ( k ) ) < Info ( S ( 0 ) ) ,
但:
Actionability ( S ( k ) ) > Actionability ( S ( 0 ) ) \operatorname{Actionability}(S^{(k)})
>
\operatorname{Actionability}(S^{(0)}) Actionability ( S ( k ) ) > Actionability ( S ( 0 ) )
可能成立。
因此真正的工程張力是:
Compression ↔ Reconstructibility . \boxed{
\text{Compression}
\leftrightarrow
\text{Reconstructibility}.
} Compression ↔ Reconstructibility .
28. Canonical Source 與 Derived Representation 必須分離
Multi-Representation Memory Fabric 提出:
R i ( m ) ≠ m . \boxed{
R_i(m)
\neq
m.
} R i ( m ) = m .
對 canonical object m m m ,可以產生 vector、graph、matrix、summary 與 state representation。
任何 derived representation 都不能單獨取得:
epistemic authority . \boxed{
\text{epistemic authority}.
} epistemic authority .
29. 摘要資料庫的風險
若:
s = Summarize ( d ) , s
=
\operatorname{Summarize}(d), s = Summarize ( d ) ,
然後刪除 d d d ,則:
summary error → irreversible memory corruption . \boxed{
\text{summary error}
\rightarrow
\text{irreversible memory corruption}.
} summary error → irreversible memory corruption .
更安全的是保留:
s → source d . s
\xrightarrow{\operatorname{source}}
d. s source d .
30. Embedding 全部內容也不夠
向量表示擅長 semantic retrieval,但通常不原生表示 valid time、version lineage、provenance、epistemic status、exact identity、contradiction type 與 authority。
因此:
Vector Search ≠ World-State Maintenance . \boxed{
\text{Vector Search}
\neq
\text{World-State Maintenance}.
} Vector Search = World-State Maintenance .
31. 圖資料庫也不是世界模型本身
圖:
G = ( V , E ) G=(V,E) G = ( V , E )
非常適合表示關係。
但若邊沒有 time、scope、claim identity、source、status 與 version,則:
A → supports B A\xrightarrow{\operatorname{supports}}B A supports B
仍然過度粗糙。
所以:
Graph ≠ Epistemically Governed Graph . \boxed{
\text{Graph}
\neq
\text{Epistemically Governed Graph}.
} Graph = Epistemically Governed Graph .
32. 不同資料表示不需要互相取代
本文採取:
One governed state + many representations . \boxed{
\text{One governed state}
+
\text{many representations}.
} One governed state + many representations .
不同 representation 可服務 exact retrieval、fuzzy discovery、graph traversal、matrix computation、temporal comparison、symbolic reasoning 與 human audit。
33. 資料新不等於世界新
網路每天有大量:
Δ D t . \Delta D_t. Δ D t .
但其中可能大部分只是 repost、commentary、translation、repackaging 與 repeated statistics。
因此:
Δ D t ≠ Δ W t . \boxed{
\Delta D_t
\neq
\Delta W_t.
} Δ D t = Δ W t .
34. World Novelty
對新資料 d d d ,定義概念量:
N W ( d ) = Distance ( W t , W t ⊕ d ) . \boxed{
N_W(d)
=
\operatorname{Distance}
(
W_t,
W_t\oplus d
).
} N W ( d ) = Distance ( W t , W t ⊕ d ) .
如果加入 d d d 後 W t W_t W t 幾乎不變,則:
N W ( d ) ≈ 0. N_W(d)\approx0. N W ( d ) ≈ 0.
這表示內容新,但世界狀態不新。
35. State-Changing Information
定義:
I Δ W = { d : Δ W ( d ) ≠ 0 } . \boxed{
\mathcal I_{\Delta W}
=
\{d:
\Delta W(d)
\neq0
\}.
} I Δ W = { d : Δ W ( d ) = 0 } .
Global AI 的 attention 應優先處理 I Δ W \mathcal I_{\Delta W} I Δ W ,而不是只追求 latest。
因此:
latest ≠ important . \boxed{
\text{latest}
\neq
\text{important}.
} latest = important .
36. 過期資訊污染
若 claim c t c_t c t 在時間 t t t 正確,後來被 c t + 1 c_{t+1} c t + 1 取代,而 retrieval 仍以語義相似度召回 c t c_t c t ,就可能造成 stale truth contamination。
因此每個 claim 需要 validity state。
37. Stale 不等於 False
歷史 claim:
c t c_t c t
在今天可能不是 current,但仍可能是描述當時世界的正確歷史資料。
所以:
stale ≠ false . \boxed{
\text{stale}
\neq
\text{false}.
} stale = false .
38. 同一事件存在不同觀察尺度
事件 e e e 可以在 micro、meso、macro 尺度被描述。
因此:
same raw event → multiple scale states . \boxed{
\text{same raw event}
\rightarrow
\text{multiple scale states}.
} same raw event → multiple scale states .
這也是 X 次結構化的另一個維度。
39. 高階趨勢不能直接寫回底層事實
若 AI 從大量 micro events 推導:
H m a c r o , H_{\mathrm{macro}}, H macro ,
應標記為 I N F \mathrm{INF} INF 或 H Y P \mathrm{HYP} HYP ,不能再把它當作獨立觀察支持自己,否則形成:
epistemic feedback loop . \boxed{
\text{epistemic feedback loop}.
} epistemic feedback loop .
40. 自我引用污染
若 AI 推論 I t I_t I t 被公開後,又被下一輪系統抓回並誤認為獨立來源:
I t → I t w e b → I t + 1 , I_t
\rightarrow
I_t^{\mathrm{web}}
\rightarrow
I_{t+1}, I t → I t web → I t + 1 ,
就會形成虛假增強。
因此 Global AI 必須標記:
derivation lineage . \boxed{
\text{derivation lineage}.
} derivation lineage .
41. Dynamic Database 與 Static Corpus 的差別
Static Corpus 假設資料相對固定。
Dynamic World Database 則要求:
D t ≠ D t + 1 . D_t
\neq
D_{t+1}. D t = D t + 1 .
更重要的是 schema 甚至也可能變:
S c h e m a t ≠ S c h e m a t + 1 . Schema_t
\neq
Schema_{t+1}. S c h e m a t = S c h e m a t + 1 .
概念會 split、merge、rename、deprecate、revive、generalize 與 specialize。
42. 世界本體也可能演化
令 ontology:
O t . \mathcal O_t. O t .
Global AI 不應假設:
O t = O 0 ∀ t . \mathcal O_t
=
\mathcal O_0
\quad
\forall t. O t = O 0 ∀ t .
新科技、新制度與新概念可能造成:
O t + 1 = Γ ( O t , Δ W t ) . \mathcal O_{t+1}
=
\Gamma(
\mathcal O_t,
\Delta W_t
). O t + 1 = Γ ( O t , Δ W t ) .
43. Schema Migration 本身是認知事件
若概念 C C C 分裂成 C 1 , C 2 C_1,C_2 C 1 , C 2 ,舊資料不能只留在舊分類中,而需要重分類並保留歷史分類狀態。
因此:
schema evolution ⊂ world-model history . \boxed{
\text{schema evolution}
\subset
\text{world-model history}.
} schema evolution ⊂ world-model history .
44. 世界模型不是最新真相表
更合理的世界模型包含:
current state + history + provenance + uncertainty + conflict + unknown . \boxed{
\text{current state}
+
\text{history}
+
\text{provenance}
+
\text{uncertainty}
+
\text{conflict}
+
\text{unknown}.
} current state + history + provenance + uncertainty + conflict + unknown .
可寫成:
W t = ( S t , H t , P t , U t , C t , N t ) . W_t
=
(
S_t,
H_t,
P_t,
U_t,
C_t,
N_t
). W t = ( S t , H t , P t , U t , C t , N t ) .
45. Conflict Preservation
若兩個高品質來源衝突,Global AI 不應立即 SelectOne。
它可以保留:
C t = { c 1 , c 2 } \boxed{
C_t
=
\{c_1,c_2\}
} C t = { c 1 , c 2 }
直到取得足夠 evidence。
Conflict 也是 world state,不只是 database error。
46. Unknown Preservation
同樣:
U N K \mathrm{UNK} UNK
是一種:
explicit epistemic state . \boxed{
\text{explicit epistemic state}.
} explicit epistemic state .
若系統不能保存 unknown,就容易用 plausible generation 把 unknown 填滿。
47. Global State Refinement Pipeline
本文提出:
Observe → Preserve → Resolve → Deduplicate → Temporalize → Semanticize → Relate → Abstract → Project → Verify . \boxed{
\text{Observe}
\rightarrow
\text{Preserve}
\rightarrow
\text{Resolve}
\rightarrow
\text{Deduplicate}
\rightarrow
\text{Temporalize}
\rightarrow
\text{Semanticize}
\rightarrow
\text{Relate}
\rightarrow
\text{Abstract}
\rightarrow
\text{Project}
\rightarrow
\text{Verify}.
} Observe → Preserve → Resolve → Deduplicate → Temporalize → Semanticize → Relate → Abstract → Project → Verify .
48. Preserve 為什麼必須靠前?
任何後續步驟都可能錯。
若結構化錯誤,但原始 d d d 仍存在,可以 rebuild。若原始資料被不可逆壓縮掉,就無法復原。
因此:
Preserve before irreversible abstraction . \boxed{
\text{Preserve before irreversible abstraction}.
} Preserve before irreversible abstraction .
49. Resolve
Resolve 處理 identity、source、duplicate、version、language、format 與 entity binding。
這是資料從 documents 進入 world objects 前的第一道門。
50. Temporalize
Temporalize 不只是加日期,而是建立:
what was true when \boxed{
\text{what was true when}
} what was true when
以及:
what was known when . \boxed{
\text{what was known when}.
} what was known when .
51. Semanticize
Semanticize 將資料轉成 claim、event、relation、concept、state、observation 與 hypothesis,但每個轉換都保留 SourceBinding。
52. Relate
Relate 建立:
G t , G_t, G t ,
包含 support、contradiction、dependency、derivation、ownership、flow、sequence 與 causality candidate。
53. Abstract
Abstract 產生 trend、bottleneck、regime、cluster、system risk 與 strategic variable。
所有高階 abstraction 都應明示 derived status。
54. Project
Project 依當前問題 q t q_t q t 生成 active context:
C t ( q ) . C_t(q). C t ( q ) .
因此:
Canonical World State ≠ Active Context . \boxed{
\text{Canonical World State}
\neq
\text{Active Context}.
} Canonical World State = Active Context .
55. Verify
Verify 可以包括 source check、contradiction search、temporal consistency、schema validation、formal proof、simulation、independent Agent review 與 human review。
驗證後才允許 state promotion。
56. State Promotion
令新 claim c c c 初始為:
I N F . \mathrm{INF}. INF .
驗證後可能:
I N F → O B S , \mathrm{INF}
\rightarrow
\mathrm{OBS}, INF → OBS ,
或:
I N F → C O N , \mathrm{INF}
\rightarrow
\mathrm{CON}, INF → CON ,
甚至:
I N F → R E T . \mathrm{INF}
\rightarrow
\mathrm{RET}. INF → RET .
因此世界資料庫是:
epistemic state machine . \boxed{
\text{epistemic state machine}.
} epistemic state machine .
57. 壓縮率不是越高越好
定義:
ρ k = ∣ S ( k ) ∣ ∣ S ( 0 ) ∣ . \rho_k
=
\frac{
|S^{(k)}|
}{
|S^{(0)}|
}. ρ k = ∣ S ( 0 ) ∣ ∣ S ( k ) ∣ .
真正關心的是:
η k = decision-relevant structure retained storage or context cost . \boxed{
\eta_k
=
\frac{
\text{decision-relevant structure retained}
}{
\text{storage or context cost}
}.
} η k = storage or context cost decision-relevant structure retained .
不是單純追求:
ρ k → 0. \rho_k\to0. ρ k → 0.
58. 錯誤壓縮可能比不壓縮更危險
若壓縮遺失一個低頻但高影響信號:
x ∗ , x^\ast, x ∗ ,
可能有:
Importance ( x ∗ ) ≫ Frequency ( x ∗ ) . \operatorname{Importance}(x^\ast)
\gg
\operatorname{Frequency}(x^\ast). Importance ( x ∗ ) ≫ Frequency ( x ∗ ) .
因此 frequency-based compression 不適合作為唯一策略。
59. Importance-Preserving Compression
本文提出概念量:
IPC ( C ) = P ( critical structure survives compression ) . \boxed{
\operatorname{IPC}(C)
=
P(
\text{critical structure survives compression}
).
} IPC ( C ) = P ( critical structure survives compression ) .
壓縮器不只要保持平均語義,還要保持 chokepoint、anomaly、rare risk、irreversible event 與 strategic dependency。
60. 資料品質不是單一分數
一個資料物件可以有:
Q ( d ) = ( q s o u r c e , q t i m e , q i d e n t i t y , q i n d e p e n d e n c e , q p r e c i s i o n , q s c o p e , q f r e s h n e s s ) . Q(d)
=
(
q_{\mathrm{source}},
q_{\mathrm{time}},
q_{\mathrm{identity}},
q_{\mathrm{independence}},
q_{\mathrm{precision}},
q_{\mathrm{scope}},
q_{\mathrm{freshness}}
). Q ( d ) = ( q source , q time , q identity , q independence , q precision , q scope , q freshness ) .
因此:
High Quality ≠ High Relevance . \boxed{
\text{High Quality}
\neq
\text{High Relevance}.
} High Quality = High Relevance .
61. 新穎性與品質必須分離
定義資料品質 Q ( d ) Q(d) Q ( d ) 、世界新穎度 N W ( d ) N_W(d) N W ( d ) 與當前查詢相關度 R q ( d ) R_q(d) R q ( d ) 。
則:
Value ( d ) = F ( Q ( d ) , N W ( d ) , R q ( d ) ) . \boxed{
\operatorname{Value}(d)
=
F(
Q(d),
N_W(d),
R_q(d)
).
} Value ( d ) = F ( Q ( d ) , N W ( d ) , R q ( d )) .
62. Canonicality
Canonical 不表示唯一真理,而表示系統目前正式承認、可追溯、可版本化的受治理狀態入口。
因此:
Canonical ≠ infallible . \boxed{
\text{Canonical}
\neq
\text{infallible}.
} Canonical = infallible .
canonical state 可以 revise、retract、supersede、branch 與 merge。
63. 世界狀態也可能分支
若兩個高可信模型無法調和:
W t ( a ) W_t^{(a)} W t ( a )
與:
W t ( b ) , W_t^{(b)}, W t ( b ) ,
系統可暫時維持:
{ W t ( a ) , W t ( b ) } . \boxed{
\{W_t^{(a)},W_t^{(b)}\}.
} { W t ( a ) , W t ( b ) } .
而不是強迫統一。
64. Branch 不等於失敗
如果世界證據本身未決,分支可能比單一答案更正確。
Global AI 的成熟度可以部分由:
ability to preserve unresolved branches \boxed{
\text{ability to preserve unresolved branches}
} ability to preserve unresolved branches
衡量。
65. Dynamic World State 需要 Wake Conditions
若某狀態 s s s 目前穩定,系統不需要一直高成本重算。
可以設定:
Wake ( s ) = { 1 , if triggering delta appears 0 , otherwise . \operatorname{Wake}(s)
=
\begin{cases}
1,&\text{if triggering delta appears}\\
0,&\text{otherwise}.
\end{cases} Wake ( s ) = { 1 , 0 , if triggering delta appears otherwise .
因此:
Persistent World State ≠ Continuous Full Inference . \boxed{
\text{Persistent World State}
\neq
\text{Continuous Full Inference}.
} Persistent World State = Continuous Full Inference .
66. Hot / Warm / Cold State
概念上:
W = W h o t ∪ W w a r m ∪ W c o l d . \mathcal W
=
W_{\mathrm{hot}}
\cup
W_{\mathrm{warm}}
\cup
W_{\mathrm{cold}}. W = W hot ∪ W warm ∪ W cold .
Hot 對應高變動、高決策相關狀態;Warm 對應中等變化依賴;Cold 對應歷史與低變動狀態。
這不只是 storage optimization,也是 attention architecture。
67. 世界狀態與注意力耦合
若:
Δ W t ( v ) \Delta W_t(v) Δ W t ( v )
很大,則:
a t + 1 ( v ) ↑ . a_{t+1}(v)
\uparrow. a t + 1 ( v ) ↑ .
若:
Δ W t ( v ) ≈ 0 \Delta W_t(v)\approx0 Δ W t ( v ) ≈ 0
且無高風險依賴,則:
a t + 1 ( v ) ↓ . a_{t+1}(v)
\downarrow. a t + 1 ( v ) ↓ .
這自然接到 GIRA-A04。
68. Information-to-State Conversion Efficiency
定義:
η I S = verified state-changing information raw information processed . \boxed{
\eta_{IS}
=
\frac{
\text{verified state-changing information}
}{
\text{raw information processed}
}.
} η I S = raw information processed verified state-changing information .
若:
η I S ≪ 1 , \eta_{IS}\ll1, η I S ≪ 1 ,
表示系統花大量算力處理重複與低狀態價值資訊。
Global AI 的競爭力可能很大部分來自:
η I S ↑ . \eta_{IS}\uparrow. η I S ↑ .
69. State Reconstruction Cost
令:
C R ( W t ) C_R(W_t) C R ( W t )
表示從保存的 source、events、versions 與 representations 重建 W t W_t W t 的成本。
系統需要在高壓縮與低重建成本之間取得平衡。
70. Epistemic Debt
若系統持續產生 derived state,卻不保留 source、lineage、version、status 與 uncertainty,則形成:
D E = Epistemic Debt . \boxed{
D_E
=
\text{Epistemic Debt}.
} D E = Epistemic Debt .
隨時間累積後,世界模型可能變得不可審計。
71. 全域資料優勢不只是資料更多
真正的優勢可能是:
better state conversion . \boxed{
\text{better state conversion}.
} better state conversion .
包括更少重複、更準 entity identity、更完整 temporal state、更強 provenance、更低 stale contamination、更好 unknown preservation 與更有效 active projection。
因此:
Global Data Advantage ≠ Data Volume Advantage . \boxed{
\text{Global Data Advantage}
\neq
\text{Data Volume Advantage}.
} Global Data Advantage = Data Volume Advantage .
72. 對 ASI 的限制
即使:
I A S I ≫ I h u m a n , I_{\mathrm{ASI}}
\gg
I_{\mathrm{human}}, I ASI ≫ I human ,
若它被餵入:
D r a w D_{\mathrm{raw}} D raw
而缺乏 structuring、temporality、provenance 與 verification,仍可能受到 information architecture bottleneck。
因此:
ASI + Internet ⇏ Global World Model . \boxed{
\text{ASI}
+
\text{Internet}
\not\Rightarrow
\text{Global World Model}.
} ASI + Internet ⇒ Global World Model .
73. ASI 可以自行發明這些方法嗎?
可能。
如果 ASI 具有 meta-cognition、architecture invention、autonomous experimentation 與 persistent tooling,它可能自行建立:
Φ 1 , Φ 2 , … , Φ x . \Phi_1,\Phi_2,\ldots,\Phi_x. Φ 1 , Φ 2 , … , Φ x .
但這仍然證明需要這些轉換,而不是證明它們不需要存在。
因此:
method discoverability ≠ method dispensability . \boxed{
\text{method discoverability}
\neq
\text{method dispensability}.
} method discoverability = method dispensability .
74. Global AI 資訊層的七條不變量
Source ≠ Claim \boxed{
\text{Source}
\neq
\text{Claim}
} Source = Claim
Claim ≠ Fact \boxed{
\text{Claim}
\neq
\text{Fact}
} Claim = Fact
Current ≠ Historical \boxed{
\text{Current}
\neq
\text{Historical}
} Current = Historical
State ≠ Projection \boxed{
\text{State}
\neq
\text{Projection}
} State = Projection
New Content ≠ New World State \boxed{
\text{New Content}
\neq
\text{New World State}
} New Content = New World State
Representation ≠ Authority \boxed{
\text{Representation}
\neq
\text{Authority}
} Representation = Authority
Compression ≠ Deletion . \boxed{
\text{Compression}
\neq
\text{Deletion}.
} Compression = Deletion .
75. Global Information Architecture
綜合本文:
I G = ( S , E , C , P , T , V , R , W , Δ ) . \boxed{
\mathfrak I_G
=
(
\mathcal S,
\mathcal E,
\mathcal C,
\mathcal P,
\mathcal T,
\mathcal V,
\mathcal R,
\mathcal W,
\Delta
).
} I G = ( S , E , C , P , T , V , R , W , Δ ) .
其中:
S \mathcal S S :sources;
E \mathcal E E :entities/events;
C \mathcal C C :claims;
P \mathcal P P :provenance;
T \mathcal T T :temporal semantics;
V \mathcal V V :versions;
R \mathcal R R :representations;
W \mathcal W W :world states;
Δ \Delta Δ :state-change events。
76. 與 Global Cognitive Atlas 的關係
A02 的:
A G \mathfrak A_G A G
處理不同 observer、method、representation 如何形成全域認知。
A03 的:
I G \mathfrak I_G I G
處理原始資訊如何被轉成可供這些認知 charts 使用的受治理狀態。
因此兩者形成:
I G ↔ A G . \boxed{
\mathfrak I_G
\leftrightarrow
\mathfrak A_G.
} I G ↔ A G .
77. 可觀測預測
本文提出六個預測:
高階 Agent 的競爭差異會越來越多來自 state architecture,而不只是 model benchmark。
大型企業 AI 將逐步從 RAG 文件庫走向 claim、event、provenance、temporal state 型架構。
AI memory 將由單一 vector store 走向 multi-representation governed fabric。
世界監控型 AI 的核心輸出將從 summary 轉向 Δ W t \Delta W_t Δ W t 。
重要資訊排序將越來越重視 state-changing novelty,而不是 publication recency。
高自治 Agent 若缺乏 provenance 與 epistemic status,會累積自我引用污染與 stale truth。
78. 與既有 EveMissLab 研究的關係
78.1 網路資訊海作為文明動態記憶
既有研究提出:
Daily Delta → Temporal Record → Domain History → Dynamic Memory . \text{Daily Delta}
\rightarrow
\text{Temporal Record}
\rightarrow
\text{Domain History}
\rightarrow
\text{Dynamic Memory}. Daily Delta → Temporal Record → Domain History → Dynamic Memory .
本文將它提升為 Global AI 的世界狀態維護層。
78.2 SEDB
SEDB 已提出:
K t = ( V t , E t , C t , P t , T t , Λ t , Δ t ) . \mathcal K_t
=
(
V_t,
E_t,
C_t,
P_t,
T_t,
\Lambda_t,
\Delta_t
). K t = ( V t , E t , C t , P t , T t , Λ t , Δ t ) .
本文採用其 claim-first、event sourcing、provenance、multi-time、epistemic status 與 semantic evolution 作為 X 次結構化的重要工程候選。
78.3 Multi-Representation Memory Fabric
MRMF 已提出:
R i ( m ) ≠ m . R_i(m)\neq m. R i ( m ) = m .
以及 One Governed Memory State + Many Replaceable Representations。
本文將其擴張到 Global AI:世界狀態可以有多表示,但任何表示不能自行取得 world authority。
78.4 GIRA-A02
A02 提出 Global AI 需要 atlas of world models。
本文補上:atlas 所依賴的資料本身必須先經受治理的動態結構化。
79. 外部標準支點
W3C PROV 提供 Entity、Activity、Agent 與 derivation 等 provenance 表示框架,說明 provenance 可以成為可互操作的一等資料結構,而不只是純文字註記。
RFC 7089 HTTP Memento 提供 Web 資源 datetime negotiation、Memento 與 TimeMap 機制,說明時間版本與某個時間點的資源狀態存取可以被正式協議化。
本文不主張直接以 PROV-O 或 Memento 實作全部 Global AI world state;它們只是提供成熟的外部支點。
80. 結論
本文的核心命題是:
Information Ocean ≠ World Model . \boxed{
\text{Information Ocean}
\neq
\text{World Model}.
} Information Ocean = World Model .
以及:
More Information ⇏ More Effective Global Cognition . \boxed{
\text{More Information}
\not\Rightarrow
\text{More Effective Global Cognition}.
} More Information ⇒ More Effective Global Cognition .
真正的 Global AI 需要把:
D r a w D_{\mathrm{raw}} D raw
經過:
S ( 0 ) → S ( 1 ) → ⋯ → S ( x ) S^{(0)}
\rightarrow
S^{(1)}
\rightarrow
\cdots
\rightarrow
S^{(x)} S ( 0 ) → S ( 1 ) → ⋯ → S ( x )
轉換成:
W t . W_t. W t .
而世界狀態又必須持續更新:
W t → Δ W t W t + 1 . W_t
\xrightarrow{\Delta W_t}
W_{t+1}. W t Δ W t W t + 1 .
因此未來真正稀缺的能力可能不是誰能收集最多資料,而是:
誰能以最低認知與計算成本,把最大規模的異質資訊轉換成最可靠、可追溯、可重建、可更新的世界狀態。
可以概括為:
Global Intelligence requires information-to-state conversion intelligence. \boxed{
\text{Global Intelligence}
\text{ requires information-to-state conversion intelligence.}
} Global Intelligence requires information-to-state conversion intelligence.
最後提出 Global AI 的資料原則:
Preserve the source, govern the state, derive the representation, and update by delta. \boxed{
\text{Preserve the source, govern the state, derive the representation, and update by delta.}
} Preserve the source, govern the state, derive the representation, and update by delta.
只有當 AI 能持續知道哪些資料是同一件事、哪些只是轉述、哪些已過期、哪些是觀察、哪些是推論、哪些仍有爭議、哪些真的改變了世界狀態,以及哪些高階表示可以被重建,它才真正從「讀取資訊海」走向「維持世界模型」。
下一篇 GIRA-A04 將研究:
當世界模型已經建立,AI 如何知道現在真正重要的是哪裡? \boxed{
\text{當世界模型已經建立,AI 如何知道現在真正重要的是哪裡?}
} 當世界模型已經建立, AI 如何知道現在真正重要的是哪裡?
也就是動態關鍵節點、瓶頸、注意力與全域認知資源配置。
參考文獻與前置研究
EveMissLab / Neo.K 既有研究
Neo.K with Aletheia, GIRA-A01|ASI 不等於 Global AI:智能能力類別與全域操作架構類別的分離 , 2026.
Neo.K with Aletheia, GIRA-A02|局部全域與真正全域認知:觀察者、方法論座標與認知域 , 2026.
Neo.K, 網路資訊海作為文明動態記憶:總論與未來命題 , EML-IIODO-TH-10, 2026.
Neo.K, AI 原生語義演化資料庫:從事件溯源、主張帳本到動態語義圖查詢語言的系統架構 , SEDB / SEQL Technical Whitepaper v1.0, 2026.
Neo.K, 多表示記憶 Fabric:向量、符號、矩陣、圖、原文與狀態庫的共存 , AI 自主上下文記憶與認知編譯系列 Paper 05, 2026.
Neo.K with Aletheia, ACWC-01|母模型+超算為何仍不等於 AI 原生計算世界 , 2026.
外部參考
W3C Provenance Working Group, PROV-O: The PROV Ontology , W3C Recommendation, 2013.
W3C Provenance Working Group, PROV-DM: The PROV Data Model , W3C Recommendation, 2013.
H. Van de Sompel et al., RFC 7089 — HTTP Framework for Time-Based Access to Resource States (Memento) , 2013.
Canonical Source Note
本文件的正式原稿為此 UTF-8 Markdown source。聊天介面的渲染版本不應被視為 canonical source。
數學公式 canonical delimiter 僅使用:
inline math:$...$
display math:$$...$$
不得以 Unicode 數學字元替換 LaTeX source,不進行 unicode_escape 類 round-trip,不自行改寫反斜線、delimiter 或公式原始碼。