跨世界證據:一致、反例、獨立性與證據轉移
Cross-World Evidence: Agreement, Counterexamples, Independence, and Evidence Transport Across Runnable Worlds
Branching World Computation / World-Domain Cognitive Runtime
分支世界計算/世界域認知 Runtime 系列
WDC-05 / BWC-05 — Evidence Paper I
作者:Neo.K(許筌崴)
協作形式化:Aletheia
機構:一言諾科技有限公司(EveMissLab)
日期:2026-08-17
版本:v0.1
狀態:cross-world evidence / dependence / transport formalization
Canonical Non-Identity Statement
WDC-01 建立:
除非建立明示的 world-to-reality evidence transport contract。
WDC-02 建立:
WDC-04 再建立:
本文進一步建立:
以及:
本文不主張:
- 100 個 worlds 支持某命題就使該命題為真;
- 多數決是 cross-world evidence 的 universal aggregation rule;
- 不同 world IDs 自動代表獨立 evidence sources;
- 不同 seeds 等於不同模型;
- 不同 foundation models 一定有獨立 error;
- 同 backend 的多個 branches 沒有 evidence value;
- counterexample 必然勝過所有 supporting worlds;
- one scalar evidence score 足以描述所有 cross-world evidence;
- simulation calibration 可完全消除 model-form uncertainty;
- NASA、climate-model ensemble 或 Bayesian stacking 等既有方法等同 WDC;
- world-to-reality transport 可以由 simulation consensus 自動取得;
- 任何跨世界結果都能被合併成 probability。
摘要
到 WDC-04 為止,World-Domain Cognitive Runtime 已經建立:
- runnable world;
- branching world lineage;
- world-domain Governor;
- master / local actor / observer / evaluator / Governor 的 role separation;
- branch blindness 與 explicit cross-world channels;
- bounded authority 與 nested-agent architecture。
於是下一個問題終於可以被嚴格提出:
如果有很多 runnable worlds 都對同一命題輸出結果,這些結果到底構成多少證據?
最粗糙的方法是計數:
然後宣稱:
100 個 world 中 97 個支持 ,所以 有 97% 機率成立。
本文拒絕這個推論。
因為 97 個 worlds 可能:
- 全部 fork 自同一 root;
- 使用同一 dynamics backend;
- 使用同一 foundation model;
- 使用同一 dataset;
- 使用同一 hidden assumption;
- 由同一 evaluator 評分;
- 只改了 random seed;
- 只是 near-duplicate world contracts。
此時:
本文定義 Cross-World Evidence Packet:
其中:
- :world identity / contract;
- :world outcome relative to claim ;
- :within-world validity / contract fidelity;
- :lineage information;
- :backend / model family;
- :assumption / parameter family;
- :evaluator identity / independence profile;
- :replication / run information;
- :world-to-target transport scope;
- :uncertainty / unknowns。
因此 cross-world evidence aggregation 的輸入不是:
而是:
本文再定義 world-pair dependence:
其中:
表示在 declared evidence dimensions 上高度獨立;
表示近乎 evidence-duplicate。
依賴不應只看 output correlation,而至少拆成:
分別表示:
- :lineage dependence;
- :backend dependence;
- :model / architecture dependence;
- :data dependence;
- :assumption dependence;
- :evaluator dependence;
- :communication / context dependence。
因此:
同樣:
也不必然成立。
真正要問的是:
為什麼這兩個 worlds 會犯相同錯誤?它們共享了哪些 upstream causes?
本文提出概念性 Effective Evidence Count:
其中:
本文不指定一個 universal 公式,因為 dependence type、claim type 與 aggregation semantics 可能不同。但任何宣稱:
worlds agree
的報告,都應至少同時報:
或明確說明無法可靠估計 。
本文進一步提出:
Evidence Family
若一組 worlds:
共享:
- root;
- backend;
- core data;
- evaluator;
- central assumptions;
則它們可以先被視為一個:
family 內的多 runs 可以提高:
- stochastic precision;
- numerical stability;
- tail estimation;
- robustness within contract;
但不能被錯算成同等數量的 independent model confirmations。
因此:
本文正式區分至少四層 replication:
- R0 — Re-run Replication:同 world spec / backend / evaluator,不同 run;
- R1 — Branch Replication:同 root / backend,不同 controlled intervention 或 seed family;
- R2 — Backend Replication:不同 dynamics / simulator / model family;
- R3 — Assumption / Evaluator Replication:不同 assumptions、data pathways、evaluator families;
- R4 — External Target Replication:真實資料或真實實驗 resolution。
因此:
只作 evidence-diversification hierarchy,不表示任何固定 numerical weight。
本文亦正式建立 Counterexample Burden。
假設:
個 worlds 支持:
但一個高-validity、high-independence、matched-contract world:
產生:
不能只用:
票數把它抹掉。
定義:
對每個 counterexample,至少問:
- 是 runtime bug 嗎?
- 是 contract mismatch 嗎?
- 是 rare but valid region 嗎?
- 是不同 ontology 嗎?
- 是 world-to-reality irrelevant 嗎?
- 還是它揭示了 supporting worlds 共享的 blind spot?
本文提出:
Counterexample Escalation Rule
若 counterexample:
同時具有:
則 Governor 應提高:
- replication budget;
- adversarial worlds;
- backend diversity;
- evaluator review;
而不是直接當 outlier 刪除。
本文再定義 cross-world evidence profile:
其中:
- :cross-world consistency;
- :independence;
- :replication depth;
- :counterexample burden;
- :backend / assumption diversity;
- :transport strength;
- :remaining uncertainty。
這是一個 vector,而不是單一「世界信心分數」。
本文拒絕:
如果沒有說明:
0.93 到底是 world-internal probability、ensemble frequency、posterior predictive weight、real-world calibrated probability,還是純 heuristic。
本文建立 Cross-World Agreement Matrix:
與 dependence matrix:
兩者必須分開。
一個高度一致的 cluster:
若:
可能只是同源共錯。
反過來,多個高度獨立 worlds:
若仍高度一致:
才是更強的 cross-world robustness signal。
本文將這個差異寫成:
但不預設一個 universal 。
外部 multi-model ensemble research 對這一點有直接而成熟的技術鄰接。氣候模型 ensemble 長期面臨「model democracy」問題:多個 models 可能共享 code、parameterizations、tuning choices,因此不能被自然當成完全獨立的一人一票。Knutti 等人的 weighting work 以及 Sanderson、Wehner、Knutti 的 multi-model assessment 明確把 model performance 與 model interdependence 一起納入權重;後續 ClimSIPS 也把 independence、performance 與 spread 明確分離,用於 CMIP model subselection。這些工作不等於 WDC,但它們提供一個非常關鍵的外部先例:
本文亦引用 Bayesian predictive stacking 作為另一種 aggregation 鄰接。Yao、Vehtari、Simpson 與 Gelman 的 stacking 方法不是假設 candidate models 中必有一個是真實 data-generating model,而是根據 out-of-sample predictive performance 組合 predictive distributions。WDC 不直接採用 stacking 作 universal world aggregator,但接受其一個重要方法論警告:
當真實生成機制可能不在 candidate worlds 裡時,不應把「哪個 world 最真」當作唯一 aggregation 問題。
這與 WDC-01 的 unknown / transport boundary 相容。
本文同時引入 simulation credibility boundary。NASA-STD-7009B 要求 modeling / simulation 活動明示 verification、validation、uncertainty、acceptance criteria 與 credibility practices;2026 年的 NASA-HDBK-7009B 進一步作為實作 guide。這些標準不是 WDC,但它們支持本文的核心證據分層:
因此本文將 world evidence 分為:
Layer 1 — Execution Integrity
問:
world 是否照自己的 contract 正確執行?
Layer 2 — Internal Validation
問:
within-world outcomes 是否可重現、穩定、與 internal known cases 一致?
Layer 3 — Cross-World Robustness
問:
在不同 lineage / backend / assumptions / evaluators 下,結論是否仍保持?
Layer 4 — Target Transport
問:
world evidence 是否能合理搬到 target system / reality?
Layer 5 — External Resolution
問:
真實 experiment / data / deployment 是否支持?
因此:
後一層需要額外 evidence,而不是自動包含。
本文再建立:
World Evidence Transport Graph
令:
node 可以是:
- world outcome;
- world family aggregate;
- external dataset;
- real experiment;
- claim。
edge type 可以是:
- replicate;
- calibrate;
- compare;
- validate;
- transport;
- contradict;
- aggregate。
這使 cross-world evidence 不再只是:
一個表格裡 100 行 simulation results。
而是:
本文也正式建立 Evidence Transport Debt:
表示從目前最高 evidence layer 到 target real-world claim 仍缺少多少:
- calibration;
- validation;
- causal justification;
- scale matching;
- distribution matching;
- external replication。
不要求 universal scalar,可作 structured debt ledger。
因此一個 world consortium 可以:
但:
例如 20 個 independent social simulations 全都預測同一制度結果,但沒有 real-world calibration;它們構成有價值的 cross-model robustness evidence,但還不是 real-world probability。
本文建立 Cross-World Consensus Classes:
C0 — Single-Family Agreement
大量 worlds,但共享主要 backend / assumptions。
C1 — Multi-Branch Agreement
多 interventions / seeds 下穩定。
C2 — Multi-Backend Agreement
不同 simulators / model families 一致。
C3 — Multi-Assumption / Multi-Evaluator Agreement
不同 assumption families、不同 evaluator families 一致。
C4 — Cross-World + External Calibration
world consortium 在 known real cases 上校準。
C5 — Prospective External Resolution
新 real-world evidence 在預先註冊條件下支持。
這不是 truth ladder,而是:
本文進一步要求 Pre-Registered Claim Equivalence。
若 world 輸出:
與 world 輸出:
要說它們「支持同一命題」之前,必須先定義:
即:
- outcome representation;
- comparison metric;
- equivalence tolerance。
否則容易出現:
每個 world 結論都不完全一樣,但事後把它們全部歸類成「差不多支持」。
本文稱:
Consensus Stretching
若 equivalence contract 在看完結果後不斷放寬,使更多 worlds 被算成 support,就構成 post-hoc evidence inflation。
因此:
本文亦正式建立 World Voting Failure。
假設:
來自同一 backend family;
來自完全不同 scientific model。
單純 majority:
可能完全錯誤。
因此:
不是 WDC default。
更合理的最低要求是:
即使最終使用 weighting,也要先把 dependence 寫出來。
本文再定義:
Independent Counterevidence Search
在世界 ensemble 已經高度支持:
時,Governor 不應只 spawn 更多:
supporting variants。
而應主動產生:
尋找:
- weakest assumptions;
- failure regions;
- alternative causal mechanisms;
- adversarial agent policies;
- backend-disagreement regimes。
如果經 adversarial search 仍難找到高-validity counterexample,cross-world robustness 才真正增加。
因此:
只作方法論偏序,不作 universal theorem。
本文亦建立 World Family Collapse Test。
將 worlds 按 dependence 聚類:
然後比較:
與:
如果:
報告應說:
1000 runs / worlds across 2 major evidence families,
而不是:
1000 independent simulations confirm the result。
本文提出 Backend Ablation:
依次移除:
- backend family;
- evaluator family;
- root lineage family;
- data family;
重新聚合。
若結論只在某單一 family 存在時成立:
若移除任意單一 family 後仍成立:
本文也提出 Cross-World Evidence Packet 最小工程格式:
claim_id
world_id
world_family_id
root_lineage
branch_path
backend
model_versions
data_sources
assumption_family
evaluator_id
evaluator_independence
run_count
seed_policy
outcome
outcome_equivalence_class
internal_validity
uncertainty
counterexample_status
transport_scope
external_calibration
evidence_level
聚合報告則至少包含:
claim_id
total_worlds
total_runs
evidence_families
estimated_dependence
agreement_by_family
counterexamples
backend_ablation
evaluator_ablation
external_validation
transport_debt
unknowns
本文最後提出:
Cross-World Evidence Principle
Cross-world evidence strength depends not only on how many worlds agree, but on how independently they were constructed, how validly they ran, how diversely they represent assumptions, how strongly counterexamples were sought, and how well their conclusions transport to the target domain.
No Model Democracy by Default Principle
One world, one vote is not a default evidence rule when worlds share lineage, backends, data, assumptions, evaluators, or communication channels.
Counterexample Preservation Principle
A valid, independent counterexample must not be erased by numerical majority; it should trigger diagnosis, replication, and model-family expansion.
Transport Separation Principle
Cross-world robustness and real-world validity are distinct evidence layers. Agreement among simulations cannot by itself erase world-to-reality transport debt.
下一篇 WDC-06 將回到 WDC-03 刻意留下的深層問題:
Which Worlds Deserve Computation?
《哪些世界值得被計算:世界投資組合、探索—驗證與計算價值》
它不再只是 Governor API,而會正式研究:
- value of computation;
- world portfolio;
- exploration / exploitation / verification;
- rare-world preservation;
- expected regret;
- information geometry;
- stopping / expansion criteria。
關鍵詞: Cross-World Evidence、Model Dependence、Ensemble Independence、Counterexample、Simulation Credibility、Evidence Transport、Model Weighting、World Families、WDC
1. 100 個世界支持同一命題,第一個問題不是「幾票?」
第一個問題應是:
2. World Identity 不等於 Evidence Identity
WDC-02 保證:
但這只說:
runtime / lineage identity 不同。
3. 它沒有保證
4. Shared Backend
如果:
可能有 shared model-form error。
5. Shared Data
可能有同 data bias。
6. Shared Root
同一 parent checkpoint:
具有 lineage dependence。
7. Shared Evaluator
可能有 evaluation bias。
8. Shared Master Hint
WDC-04 已說:
可以污染 branches。
9. Shared Communication
cross-world channels:
也可能讓 errors 相關。
10. Therefore
11. Evidence Packet
對 claim:
world:
產生:
12. Outcome Component
作最簡離散版。
13. Continuous Outcome
也可:
14. Internal Validity
15. Invalid World
若:
不能因為它支持 hypothesis 就算票。
16. Evidence Family
定義:
按:
- backend;
- model;
- data;
- assumptions;
- root;
- evaluator;
聚類。
17. Family Partition
若能建立硬 partition。
18. More Realistically Soft Families
world 可同時共享:
- backend family A;
- data family B。
19. So Dependence Is Graph Better Than Partition
20. Edge Weight
21. Dependence Dimensions
22. Lineage Dependence
23. Backend Dependence
24. Model Dependence
25. Data Dependence
26. Assumption Dependence
27. Evaluator Dependence
28. Communication Dependence
29. One Number Can Be a Projection
30. But Raw Vector Must Remain Auditable
31. Climate Ensemble Analogy
Climate multi-model ensembles often contain:
- related model families;
- shared parameterizations;
- shared tuning practices;
- duplicated / near-related models。
32. Therefore Model Democracy Can Overcount
equal weighting assumes more independence than actually exists。
33. Knutti 2017
提出 weighting 同時考慮:
34. Sanderson–Wehner–Knutti 2017
explicit skill + independence weighting。
35. WDC Analogy
world aggregation 至少要問:
36. But WDC Has More Dependence Channels
worlds 還可能共享:
- evaluator;
- local agent;
- prompt;
- master;
- fork lineage。
37. So Climate Weighting Is Analogue, Not Drop-In Formula
38. Model Democracy Failure Example
100 worlds:
都來自 backend:
39. One World
來自:
40. If A Has Shared Bug
100:1 vote can be nonsense。
41. Backend Family Vote
minimum alternative:
先看:
42. But Even Family Vote Is Not Universal
不同 family quality 不同。
43. Performance Matters
world 在 historical calibration:
44. Independence Matters
45. Transport Matters
46. Counterexample Matters
47. Evidence Vector
48. No Universal Weight
49. Re-runs
同 world spec:
50. What Re-runs Tell Us
- stochastic variance;
- numerical stability;
- rare-event rate under model。
51. What Re-runs Do Not Tell Us
- model-form validity;
- backend independence;
- reality validity。
52. Therefore R0
Re-run replication。
53. R1
branch replication。
54. R2
backend replication。
55. R3
assumption / evaluator replication。
56. R4
external target replication。
57. Evidence Diversification
應報是哪一層。
58. 1000 Seeds Can Be R0 Only
59. This Is Still Useful
但別叫:
1000 independent scientific models。
60. Agreement Matrix
61. High Agreement
62. Dependence Matrix
63. Four Cases
Case I
high agreement + high dependence。
64. Interpretation
shared-family consensus。
65. Case II
high agreement + low dependence。
66. Interpretation
stronger robustness。
67. Case III
low agreement + high dependence。
68. Interpretation
same family unstable / stochastic / sensitive。
69. Case IV
low agreement + low dependence。
70. Interpretation
genuine model uncertainty / assumption disagreement。
71. This Is Much More Informative Than Majority Count
72. Effective Evidence Count
概念:
73. If Fully Independent
74. If Perfect Duplicates
75. Intermediate
depends on structure。
76. Do Not Pretend Precision
若 dependence 未知,
報:
77. Counterexample
定義:
78. Not Every Opposite Output Is Valid Counterexample
可能:
- runtime failure;
- contract mismatch;
- unsupported evaluator。
79. Counterexample Validity
80. Strong Counterexample
需要:
- high internal validity;
- relevant contract;
- independent source;
- reproducibility。
81. One Strong Counterexample Can Matter More Than Many Duplicates
82. But It Does Not Automatically Falsify Probabilistic Claims
若:
是:
outcome occurs 99% of time,
one failure expected。
83. Claim Type Matters
84. Universal Claim
單一合法反例可有非常高 logical burden。
85. Probabilistic Claim
需要 frequency / model semantics。
86. Existential Claim
single positive witness may dominate。
87. Therefore Evidence Aggregation Must Know Claim Type
88. Universal World Voting Is Meaningless
89. Claim-Type-Aware Aggregation
90. Causal Claims
need:
- intervention semantics;
- matched branches;
- transport assumptions。
91. Forecast Claims
need:
- pre-registration;
- resolution;
- calibration。
92. Comparative Claims
need:
- matched budgets;
- evaluator equivalence;
- controlled deltas。
93. Consensus Equivalence
world outputs not necessarily same representation。
94. Define Feature Map
95. Distance
96. Equivalence
97. Must Freeze Before Seeing Aggregate
98. Consensus Stretching
post-hoc increase:
until desired consensus。
99. Forbidden Without Disclosure
100. Support Class
101. Counter Class
102. Inconclusive Class
103. Invalid Class
104. Do Not Force Every World Into Support/Reject
105. Inconclusive Is Real Category
106. Invalid Is Not Counter
simulator bug:
not:
107. Cross-World Evidence Profile
108.
consistency。
109.
independence。
110.
replication depth。
111.
counterexample burden。
112.
cross-backend / assumption diversity。
113.
transport strength。
114.
remaining uncertainty。
115. Evidence Frontier
world claim can lie on Pareto frontier across these dimensions。
116. No Single Confidence by Default
117. NASA M&S Credibility Calibration
NASA-STD-7009B explicitly covers:
- model verification;
- validation;
- uncertainty;
- credibility;
- acceptance criteria。
118. Why Important
它把:
simulation output exists
與:
simulation is credible enough for a decision
分開。
119. WDC Extends This Across Many Worlds
120. Execution Integrity
121. Internal Validation
122. Cross-World Robustness
123. Transport
124. External Resolution
125. Non-Implication Chain
126. Execution Integrity Example
code matches specification。
127. But Specification Can Be Wrong
128. Internal Validation Example
historical benchmark reproduced。
129. But Future Regime Can Shift
130. Cross-World Robustness Example
different backends agree。
131. But All May Omit Same real mechanism
132. Transport Validation Needed
133. NASA-HDBK-7009B 2026
提供 NASA-STD-7009B implementation guidance。
134. WDC Use
不是 adopt as universal standard。
而是借:
135. World Evidence Level
本文建議:
CWE-0
single world output。
136. CWE-1
reproducible within-world runs。
137. CWE-2
multiple branches / seeds, same family。
138. CWE-3
independent backend / assumption families。
139. CWE-4
cross-world evidence + calibrated target cases。
140. CWE-5
pre-registered prospective external resolution。
141. This Is Not Universal Truth Ladder
它是 WDC evidence workflow。
142. Bayesian Stacking External Analogy
Yao et al. discuss M-open setting:
true data-generating process may not be among candidate models。
143. WDC Analogy
real world may not be represented by any current world family。
144. Therefore Keep
unknown / omitted model region。
145. World Ensemble Is Not Exhaustive
146. Stacking Predictive Distributions
combines models based on predictive performance。
147. WDC Does Not Adopt as Default
because world outcomes can be:
- non-probabilistic;
- causal;
- symbolic;
- adversarial。
148. But It Demonstrates
aggregation need not mean:
pick one model as true。
149. Model Selection vs Model Combination
150. WDC Can Keep Multiple Surviving World Families
151. Climate Model Independence
Knutti / Sanderson work explicitly considers model interdependence。
152. Why Climate Models Dependent?
- shared code;
- parameterizations;
- tuning practices;
- genealogies。
153. WDC Worlds Have Even More Shared Structure
- prompts;
- foundation models;
- datasets;
- agents;
- evaluators。
154. So Dependence Audit Is Mandatory for Strong Claims
155. Skill vs Independence
climate weighting separates:
and:
156. WDC Analogy
a high-performing world family can still be overrepresented。
157. Independence vs Diversity
different outputs alone don't establish independence。
158. Diversity Must Be Causal/Structural, Not Cosmetic
159. Cosmetic Diversity
change:
- color;
- wording;
- seed;
without changing error source。
160. Structural Diversity
change:
- model family;
- dynamics;
- data;
- evaluator;
- assumptions。
161. World Family Identification
Governor can cluster using metadata first。
162. Then outcome similarity second。
163. Never Use Outcome Similarity Alone
because true models may converge to same result。
164. Dependence vs Convergence
important distinction。
165. Two Independent Models Can Agree Because Reality Constrains Them
166. Two Dependent Models Can Agree Because They Share Bug
167. Need Provenance
168. Evidence Graph
169. Evidence Nodes
- run;
- world;
- family;
- dataset;
- experiment;
- claim。
170. Evidence Edges
171. Supports Edge
world outcome supports claim under contract。
172. Counters Edge
materially challenges claim。
173. Replicates
same or nearby protocol。
174. Calibrates
known target case。
175. Validates
checks model adequacy。
176. Transports
moves evidence scope。
177. DependsOn
marks shared upstream source。
178. Aggregates
creates meta-evidence node。
179. Evidence Graph Must Remain Provenance-Aware
180. Aggregation Node
181. It Should Store Inputs
not erase them。
182. No Evidence Flattening
don't replace 100 packets with:
97% support。
183. Transport Graph
world:
to target:
184. Transport Edge
185. Transport Conditions
- scale match;
- variable match;
- dynamics match;
- distribution match;
- intervention match;
- agent behavior match。
186. Transport Debt
187. Debt Components
188. Cross-World Agreement Can Reduce Some Debt
e.g. model-form sensitivity。
189. But Cannot Erase Missing Real Mechanism
190. Example
100 traffic simulators omit human panic。
191. All agree evacuation works。
192. If real event includes panic dynamics
transport weak。
193. Cross-World Robustness ≠ External Completeness
194. Counterexample Escalation
Strong counterexample detected。
195. Governor Action
196. Spawn Replicates
197. Spawn Backend Variants
198. Check Assumptions
199. Audit Evaluator
200. Compare Transport
201. Do Not Suppress Outlier
202. Outlier vs Counterexample
outlier:
statistically unusual output。
counterexample:
logically / materially challenges claim。
203. Different.
204. Universal Claim Example
205. One valid
with:
falsifies within class contract。
206. Probabilistic Claim Example
207. One failure doesn't falsify。
208. Claim Type Must Be Explicit
209. Cross-World Causal Evidence
paired branches:
210. Internal effect
211. Replicate Across Backends
212. Agreement on Effect Direction
stronger cross-model causal robustness。
213. Still Requires External Transport
214. Evaluator Independence
WDC-04 provides:
215. Cross-World Evidence Must Include It
216. If all evaluators same
217. Independent Evaluator Replication
different evaluator family。
218. Blind Scoring
reduces hypothesis-label bias。
219. Two-Stage Evaluation
blind first, provenance audit later。
220. Cross-World Result Can Be Evaluator-Sensitive
221. Evaluator Ablation
remove one evaluator family。
222. If conclusion flips
223. Backend Ablation
remove backend family 。
224. Compute:
225. Leave-One-Family-Out Robustness
226. Robust if no single family determines conclusion。
227. Lineage Ablation
remove root family。
228. Data Ablation
remove data source family。
229. Assumption Ablation
remove assumption cluster。
230. Evidence Sensitivity Profile
231. Low Sensitivity Stronger Robustness
not truth proof。
232. Family Collapse Test
cluster worlds。
233. Report
234. Example
235. Stronger Report
1200 world executions across 3 major model families。
236. Not
1200 independent models confirmed。
237. World Consensus Classes
C0
single family。
238. C1
multi-branch same backend。
239. C2
multi-backend。
240. C3
multi-assumption/evaluator。
241. C4
external calibration。
242. C5
prospective external resolution。
243. Claim Registry
每個 claim:
claim_id
claim_type
definition
scope
equivalence_contract
target_domain
resolution_rule
244. Evidence Registry
每 evidence packet:
evidence_id
claim_id
world_id
run_id
family_ids
backend
data_sources
assumptions
evaluator
outcome
validity
uncertainty
transport
245. Aggregation Registry
aggregate_id
claim_id
input_evidence_ids
aggregation_method
dependence_model
counterexamples
sensitivity
transport_debt
version
246. Versioning Is Critical
new world:
arrives。
247. Aggregate Updates
248. Never Overwrite Old Evidence State
249. This Enables Historical Audit
what did we believe before counterexample?
250. Cross-World Calibration
on known cases:
251. Compare predicted ensemble evidence with real resolution。
252. Calibration Can Estimate
- overconfidence;
- family bias;
- transport gap。
253. Calibration Is Task-Specific
254. Good on robotics ≠ good on economics
255. Evidence Family Performance
256. Past performance can inform weighting
but beware regime shift。
257. Performance Weighting Can Overfit
if same validation set reused。
258. Separate Calibration and Test Sets
259. Cross-World Unknown Mass
TCD had:
WDC can have:
meaning:
relevant model/world families may be absent。
260. Even 100 independent worlds can miss omitted mechanism
261. Ensemble Closure Error
262. Keep Unknown World Region
263. How to Reduce
- new backend;
- new ontology;
- adversarial model;
- real data anomaly。
264. Not by Just More Seeds
265. Independent Counterevidence Search
Governor creates:
266. Goal
find strongest plausible failure。
267. Adversarial Search Budget
268. Supporting Search Budget
269. Verification Budget
270. Balanced Evidence Program
not necessarily equal。
271. Confirmation Pressure
if:
consensus can be engineered。
272. Audit Budget Allocation
273. Cross-World Evidence Governance
Governor should track:
- supporting family count;
- counter family count;
- independence;
- unresolved contradiction。
274. Promotion Criteria
world claim should not Promote just because majority。
275. Need Cross-World Evidence Packet
276. Promotion P3
WDC-03 cross-backend level。
277. WDC-05 now defines what cross-backend means
not merely different process IDs。
278. Strong P3
different:
- backend architecture;
- assumptions;
- evaluator;
where feasible。
279. External-Test Candidate
P4 requires transport assessment。
280. Not just simulation consensus。
281. WDC-05 Principle I — World Count Is Not Evidence Count
world identities、runs 與 independent evidence units 必須分開報告。
282. Principle II — Dependence-Aware Aggregation
cross-world aggregation 必須考慮 lineage、backend、model、data、assumption、evaluator 與 communication dependence。
283. Principle III — Agreement–Independence Separation
worlds 的結果是否一致與 worlds 是否獨立是兩個不同問題;強 evidence 需要知道兩者。
284. Principle IV — Counterexample Preservation
有效 counterexample 不得因數量較少而自動淘汰;其 validity、independence 與 claim type 決定它的證據負擔。
285. Principle V — Claim-Type Awareness
universal、existential、probabilistic、causal、comparative 與 forecasting claims 不能使用同一票數式 aggregation semantics。
286. Principle VI — Transport Separation
simulation ensemble 的 cross-world robustness 與 target reality validity 必須分層;前者不能自動消除後者的 transport debt。
287. Principle VII — Unknown-World Principle
即使現有 worlds 彼此獨立且高度一致,也應保留 current ensemble 可能漏掉重要 model family / mechanism 的 possibility。
288. Principle VIII — Pre-Registered Consensus
world outputs 被視為「支持同一命題」所需的 equivalence mapping 與 tolerance,應在 aggregate resolution 前固定或明示版本變更。
289. Principle IX — Family Ablation
重要 cross-world conclusions 應測試移除主要 backend / lineage / data / evaluator family 後是否仍成立。
290. Principle X — Evidence Graph Preservation
aggregation 不得抹除 individual evidence packet、counterexample、dependence 與 transport lineage。
291. WDC-05 Benchmark A — Duplicate Consensus
100 exact clones support q。
Expected:
292. Benchmark B — Independent Backend Consensus
5 substantially different backends support q。
compare evidence profile。
293. Benchmark C — Shared Bug Family
99 same backend worlds wrong。
1 independent backend correct。
test majority failure。
294. Benchmark D — Universal Claim Counterexample
many support worlds,
one valid counterexample。
test escalation。
295. Benchmark E — Probabilistic Claim
one failure among many valid runs。
ensure not incorrectly logical-falsified。
296. Benchmark F — Evaluator Dependence
same worlds scored by:
- shared evaluator;
- independent evaluators。
measure conclusion sensitivity。
297. Benchmark G — Consensus Stretching
post-hoc widen 。
audit should detect version change。
298. Benchmark H — Backend Ablation
remove each major backend family。
track aggregate stability。
299. Benchmark I — Omitted Mechanism
all worlds omit mechanism M。
external reality includes M。
test ensemble closure failure。
300. Benchmark J — Cross-World Calibration
known historical cases。
estimate overconfidence / transport debt。
301. Benchmark K — Adversarial Counterworlds
support consensus established。
allocate 。
measure counterexample discovery。
302. Benchmark L — Family Count Reporting
1000 runs from 2 world families。
report format must expose this。
303. Benchmark M — Cross-Backend False Agreement
different models share same training data flaw。
metadata dependence should remain nonzero。
304. Benchmark N — M-Open Ensemble
none of candidate worlds matches data-generating mechanism。
test whether aggregation preserves unknown-world mass rather than force winner。
305. 可否證條件
F305.1 Dependence Model No-Gain
若 dependence-aware aggregation 長期不比 raw counting 更好,complex dependence model 可簡化。
F305.2 Family Misclassification
若 metadata-defined families 與 actual shared errors無關,family definition 應修正。
F305.3 Effective Count Overprecision
若 對 arbitrary weights 高度敏感,應報 range / unresolved,而非虛假精確值。
F305.4 Counterexample Flood
若 adversarial search 產生大量 invalid pseudo-counterexamples,需提高 validity gate。
F305.5 Calibration Leakage
若 weighting 與 evaluation 使用同一 calibration data 過度調參,cross-world credibility 會被高估。
F305.6 Transport Overclaim
若 CWE-3 cross-world consensus 被直接當 CWE-5 real validation,evidence ladder 失效。
F305.7 Consensus Equivalence Drift
若 equivalence contract 事後任意調整,consensus claim 應降級。
F305.8 Unknown-World Suppression
若 aggregator 強制 posterior mass 只分配到現有 worlds,對 omitted-model risk 的表示失效。
F305.9 Evaluator Correlation
若 independent-looking evaluators實際共享 model/data bias,independence claim 需降級。
306. 與 WDC-06 的接口
WDC-03 已建立 Governor。
WDC-05 現在告訴 Governor:
不是 world 多就證據多,也不是 consensus 大就應繼續加同一類 world。
下一篇終於可以更深入處理:
WDC-06 — Which Worlds Deserve Computation?
《哪些世界值得被計算:世界投資組合、探索—驗證與計算價值》
核心問題會從:
升級成:
也就是:
- 現在缺的是 support world?
- counterworld?
- independent backend?
- rare-event world?
- higher-fidelity world?
- real-world calibration?
WDC-06 將把 Governor 從 lifecycle manager 推成真正的 world-portfolio metareasoner。
307. 結論
WDC-01 建立:
WDC-05 現在進一步建立:
當:
個 worlds 全都說:
成立。
成熟的 WDC 不應第一個問:
幾比幾?
而應先問:
因此:
中文:
沒有獨立性的共識,可能只是同一個錯誤被重複很多次。
而反過來:
因為 genuine disagreement 會暴露:
- hidden assumptions;
- model-form uncertainty;
- ontology gaps;
- evaluator sensitivity;
- transport debt。
所以跨世界證據不是:
而是:
這也是 World-Domain Cognitive Runtime 第一次真正從:
跨進:
Claim Typing
| Claim | Type | Status |
|---|---|---|
| World count、run count、independent evidence count 非同一 | D | Canonical separation |
| Cross-world evidence packet 應包含 lineage/backend/data/assumption/evaluator/transport | D | Proposed evidence contract |
| Dependence 應與 agreement 分開建模 | D | Canonical methodology |
| Effective evidence count 可作 dependence-aware概念量 | D / C | Proposed scaffold, no universal formula |
| Counterexamples應依 claim type、validity、independence 評估 | D | Canonical rule |
| Climate multi-model weighting explicitly handles performance and model interdependence | E | External ensemble analogue |
| Bayesian stacking combines predictive distributions in M-open settings without assuming a candidate is true | E | External statistical analogue |
| NASA-STD-7009B / HDBK-7009B formalize M&S credibility, verification, validation, uncertainty practices | E | External official credibility framework |
| 100 worlds support q implies q is true | — | Explicitly rejected |
| Different model names imply independent errors | — | Explicitly rejected |
| Cross-world robustness automatically proves external validity | — | Explicitly rejected |
Evidence Ladder
本文目前主要位於:
- L0:cross-world evidence packet / dependence / transport taxonomy;
- L1–L2:duplicate consensus、shared-bug、family-ablation、counterexample benchmarks;
- L3:climate ensemble dependence weighting、Bayesian stacking、NASA M&S credibility 提供外部技術對照;
- L4:需要實際 WDC multi-backend runtime 與 calibrated dependence estimation;
- L5+:prospective real-world resolution、large-scale world portfolio evidence 尚待後續。
參考文獻
Neo.K 內部正典與譜系
- Neo.K with Aletheia. From Possible Futures to Runnable Worlds. WDC-01 / BWC-01, 2026.
- Neo.K with Aletheia. Branching World Graph. WDC-02 / BWC-02, 2026.
- Neo.K with Aletheia. World-Domain Governor. WDC-03 / BWC-03, 2026.
- Neo.K with Aletheia. Nested Agents and Observer Separation. WDC-04 / BWC-04, 2026.
- Neo.K with Aletheia. Generative Forecasting. UCPNP Series II Paper 13, 2026.
- Neo.K with Aletheia. Prospective Constructive Intelligence. UCPNP Series II Paper 14, 2026.
External technical calibration
- Knutti, R., Sedláček, J., Sanderson, B. M., Lorenz, R., Fischer, E. M., & Eyring, V. A climate model projection weighting scheme accounting for performance and interdependence. Geophysical Research Letters, 44, 1909–1918, 2017.
- Sanderson, B. M., Wehner, M., & Knutti, R. Skill and independence weighting for multi-model assessments. Geoscientific Model Development, 10, 2379–2395, 2017.
- Merrifield, A. L., Brunner, L., Lorenz, R., et al. Climate model Selection by Independence, Performance, and Spread (ClimSIPS v1.0.1) for regional applications. Geoscientific Model Development, 16, 4715–4740, 2023.
- Kulinich, M., Fan, Y., Penev, S., Evans, J. P., & Olson, R. A Markov chain method for weighting climate model ensembles. Geoscientific Model Development, 14, 3539–3551, 2021.
- Yao, Y., Vehtari, A., Simpson, D., & Gelman, A. Using stacking to average Bayesian predictive distributions. Bayesian Analysis, 13(3), 917–1007, 2018.
- NASA. NASA-STD-7009B — Standard for Models and Simulations. Office of the Chief Engineer, 2024.
- NASA. NASA-HDBK-7009B — NASA Handbook for Models and Simulations: An Implementation Guide for NASA-STD-7009B. Office of the Chief Engineer, 2026.
Public Version Disclaimer
本文是一個 cross-simulation / multi-world evidence / model-dependence framework。
本文不聲稱:
- WDC 的 dependence matrix 已有 universal estimator;
- climate-model weighting 可以直接移植到所有 AI worlds;
- Bayesian stacking 是所有 world evidence 的最佳 aggregation;
- NASA modeling standards 等同 WDC;
- multi-world agreement 是 truth;
- model diversity 自動等於 error independence;
- strong counterexample 永遠推翻 probabilistic claim;
- world-to-reality transport 可由 simulation count 代替;
- CWE levels 是跨所有學科通用標準;
- 本文已完成 WDC world-portfolio allocation theory;
- 本文對 classical vs. 提供任何新證明。
本文真正建立的是:
以及: