哪些世界值得被計算:世界投資組合、探索—驗證與計算價值
Which Worlds Deserve Computation? World Portfolios, Exploration–Verification Allocation, and the Value of Computation
Branching World Computation / World-Domain Cognitive Runtime
分支世界計算/世界域認知 Runtime 系列
WDC-06 / BWC-06 — Metareasoning Paper I
作者:Neo.K(許筌崴)
協作形式化:Aletheia
機構:一言諾科技有限公司(EveMissLab)
日期:2026-08-17
版本:v0.1
狀態:world-portfolio / value-of-computation / metareasoning formalization
Canonical Non-Identity Statement
WDC-03 已建立:
並將 Governor 定義為:
的有限資源管理器。
WDC-05 再建立:
本文進一步建立:
以及:
本文不主張:
- 某個 world 本體上「值得存在」或「不值得存在」;
- highest-posterior world 一定應取得最多算力;
- highest-information-gain computation 一定改善實際決策;
- expected information gain 等於 value of computation;
- value of computation 存在單一 universal scalar;
- cheap low-fidelity world 一定值得先算;
- high-fidelity world 一定比 symbolic / coarse world 更有價值;
- counterworld 永遠比 support world 重要;
- exploration、verification、transport、tail-risk 有固定最佳比例;
- world portfolio optimization 已等同通用科學方法;
- 本文的 world portfolio 可直接取代 Bayesian experimental design、bandits、MCTS 或 multi-fidelity optimization;
- 本文已解決所有長期 world-governance normative questions。
摘要
WDC-03 建立 World-Domain Governor,提供:
- Admit;
- Schedule;
- Allocate;
- Pause;
- Resume;
- Preempt;
- Kill;
- Archive;
- Promote。
但 WDC-03 刻意沒有封頂一個更深的問題:
下一單位 computation 究竟應該投到哪個 world?
WDC-05 又進一步指出,多跑一個與現有 worlds 高度相依的 world,可能幾乎沒有增加 independent evidence;相反地,一個低成功率、但能測試 hidden assumption 的 counterworld,可能具有更高證據價值。
因此本文主張:
真正應被 Governor 比較的,不是「哪個 world 比較好」,而是「下一個 world-computation action 的邊際價值」。
本文定義 world-computation action:
其中:
- :world、world family、claim、unknown region 或 external target;
- :要執行的 computation operation;
- :追加 budget;
- :fidelity / resolution level;
- :run horizon / deadline;
- :此 computation 服務的 claim / decision;
- :computation contract。
至少可以包括:
因此:
不只等於:
再讓某個 world 多跑幾步。
它也可以是:
建立一個完全不同 backend 的獨立 world。
或:
專門建立最強 counterexample branch。
或:
不再增加 simulation,而把 budget 投向 real-data calibration。
本文定義當前 evidence / decision state:
其中:
- :目前 cross-world evidence graph;
- :world dependence structure;
- :cross-world evidence profiles;
- :unknown / omitted world-family region;
- :world-to-target transport debt;
- :待做的 real decision set;
- :decision deadline。
本文再建立 computation-value vector:
其中:
- :decision-improvement value;
- :information gain;
- :verification value;
- :counterexample / falsification value;
- :independence / diversity gain;
- :transport-debt reduction;
- :tail-risk / failure-region value;
- :option value / future-computation enabling value;
- :compute / human / verification cost;
- :latency;
- :safety / containment burden。
本文拒絕預設:
但在明示 task utility:
後,可以局部使用:
權重必須:
- task-relative;
- versioned;
- auditable;
- 不得事後偷偷改到支持已選 computation。
本文接著把 Decision Value of Computation 寫成:
若 computation cost 可與 decision utility 同尺度比較,則:
但本文同時強調:
Expected Information Gain 可以寫成:
一個 computation 可以帶來大量 information,卻完全不改變當前 decision;反過來,一個只消除 decision boundary 附近少量 uncertainty 的 world,information gain 不大,卻可能直接改變 action。
因此:
外部 Bayesian optimal experimental design 正是有限實驗資源下以 expected information gain 選擇實驗的成熟框架;Foster 等人的 variational BOED 工作專門處理 EIG estimation 的昂貴性,而 Zheng 等人的 sequential BOED 進一步直接研究 variable experiment / computation costs 下如何自適應分配額外 computation。這些工作與 WDC 結構相鄰,但不表示 WDC 必須以 mutual information 作唯一 utility。
本文亦建立 Transport Value:
若目前最大的問題不是:
worlds 不夠多,
而是:
simulation 與 reality 的 bridge 太弱,
那麼新的 simulation world 可能:
此時最值得的 computation 可能是:
- historical backtest;
- real-data calibration;
- hardware experiment;
- shadow deployment;
- external measurement。
因此:
本文再建立 Independence Value:
一個新的 independent backend 即使 prediction 與現有 worlds 一樣,也可能具有高 ,因為它降低:
所有 agreement 都來自同一 shared bug
的可能性。
相反地,第 1000 個同 backend seed:
但可能仍有 stochastic-precision value。
本文建立 Counterexample Value:
對 universal / safety claims, 可以非常高;對普通 probabilistic frequency estimate,單一 counterexample 的邏輯作用則不同。
因此 computation value 必須知道:
本文進一步提出 Epistemic Deficit Vector:
其中:
- :within-world stochastic / numerical uncertainty;
- :independent evidence family deficit;
- :counterexample-search deficit;
- :world-to-target transport deficit;
- :resolution / model-detail deficit;
- :rare-event / risk coverage deficit;
- :omitted-world-family / ontology deficit;
- :decision-critical unresolved uncertainty。
這讓 Governor 不再只是問:
哪個 world 分數最高?
而可以先問:
我們現在真正缺哪一種證據?
若:
優先:
若:
優先:
若:
優先:
若:
優先:
若:
優先:
若:
優先:
這是本文最重要的策略轉換:
本文進一步建立:
World Computation Portfolio
世界計算投資組合
定義:
其中:
是 allocation。
要求:
Portfolio 可以至少包含:
- support / refinement computations;
- counterworld computations;
- independent-backend computations;
- stochastic replication;
- transport / calibration;
- rare-event stress worlds;
- unknown-region exploration;
- fidelity escalation。
本文拒絕:
因為這容易造成 WDC-03 的:
World-Mode Collapse
本文提出:
應同時考慮:
- expected utility;
- evidence diversity;
- counterexample coverage;
- dependence;
- transport;
- latency;
- safety;
- compute cost。
本文進一步引入 Myopic vs Dynamic Computation Value。
Myopic value:
只看 computation 完成後立即增加多少 value。
但一個 cheap computation 可能本身價值不大,卻能決定:
是否值得啟動一個極昂貴 high-fidelity world。
因此 dynamic value 需要考慮:
本文將概念性 dynamic value 寫成:
這與 Sezener 與 Dayan 在 MCTS 中區分 static / dynamic values of computation 的問題高度相鄰:一次 simulation 的價值不只可能影響現在 action,也可能影響之後還會做哪些 computations。
因此:
可以具有高:
即 option / computation-enabling value。
本文將這種 dependency 明確表示為 Computation Graph:
node:
是 possible computations;
edge:
表示:
做完 後,才知道是否值得/是否允許執行 。
例如:
因此 WDC portfolio 不是一個 flat ranking list。
它可以是:
外部 Hyperband 正是 adaptive resource allocation / early stopping 的重要工程類比:先對大量 configurations 配少量資源,再把更多資源投入較有希望者。本文接受其 coarse-to-fine resource idea,但不接受把「performance winner」直接當 WDC promotion criterion。
WDC 的 cheap world 即使顯示:
hypothesis 失敗,
也可能因為它是一個高價值 counterexample 而被 Promote。
所以:
本文將 WDC 的多 fidelity 升級機制命名為:
Successive Evidence Escalation
最小層級:
其中:
- :cheap abstract / symbolic screening;
- :low-fidelity runnable world;
- :medium-fidelity replicated world;
- :high-fidelity / cross-backend world;
- :external calibration / real test candidate。
升級條件不是:
結果好看。
而可以是:
- decision sensitivity high;
- counterexample importance high;
- unresolved disagreement high;
- transport value high;
- rare risk high;
- information gain high。
本文同時加入 Low-Fidelity Reliability Gate。
外部 multi-fidelity Bayesian optimization 已清楚顯示:便宜 approximation 可以降低昂貴 evaluation 的需求,但若 low-fidelity source 與 target 關係很差,它甚至可能讓總 optimization cost 更高。Mikkola 等人的 robust MFBO 正是針對 unreliable information sources 提出保護機制。
因此 WDC 定義:
表示 fidelity level 對 target evidence 的 calibrated reliability。
若:
則 cheap world:
因此:
本文提出:
Fidelity Escalation Rule
低 fidelity 可以:
- screen;
- prioritize;
- identify gross failures;
但若其 reliability 未被校準,不應單獨:
- falsify real-world claim;
- block critical high-fidelity validation;
- authorize deployment。
本文再建立 World Computation Opportunity Cost。
當 Global Budget:
固定,
給:
更多 compute,
就代表:
少一些。
因此:
即使:
也可能:
此時不該選它。
本文再加入 Deadline-Adjusted Value。
若 decision deadline:
而 computation result time:
則結果來得太晚:
可能對當前 decision value近乎:
定義概念:
或使用 domain-specific soft decay。
因此:
可能不如:
本文亦建立 Meta-Reasoning Cost。
估計:
本身也需要:
- compute;
- modeling;
- bookkeeping;
- uncertainty estimation。
所以:
如果:
接近或超過 world computation cost,
過度精細的 Governor 反而浪費資源。
因此:
Meta-Boundedness Principle
A metareasoner should not spend more resources deciding what to compute than the decision is worth.
可以概念寫成:
才值得進一步精算;否則可使用 simpler heuristics。
本文進一步定義 Stopping Condition:
對 active computation:
若:
且不存在:
- minimum replication obligation;
- unresolved safety obligation;
- rare-event obligation;
- counterexample obligation;
- deadline-independent archival need;
則可以停止。
但:
可能只是:
在目前 budget 下,不值得繼續計算。
本文再定義 Expansion Condition:
如果:
即 WDC-05 的 backend / lineage / evaluator sensitivity 高,
則不是繼續跑同 family,
而應:
如果:
則:
如果:
則:
因此:
本文提出 World Portfolio Regret。
在 finite benchmark 中,如果 exhaustive computation 可以知道最佳 allocation:
而 Governor 實際使用:
則:
真實世界通常不知道:
但有限 simulator benchmark 可用此測 Governor。
本文進一步建立 Computation-Value Calibration。
對每個 computation:
執行前預測:
執行後測:
保存:
長期可檢查:
- information-gain overestimate;
- decision-value overestimate;
- counterexample underestimation;
- transport-value calibration;
- latency errors。
因此 Governor 不只 calibration worlds,
也 calibration:
本文稱:
Meta-Calibration
若 Governor 永遠認為:
support worlds 很有價值,
但實際 realized gain 長期低,
它必須調整 allocation model。
本文再建立 World Portfolio Modes。
Mode S — Support Refinement
精化目前 leading family。
Mode C — Counterexample Search
找 strongest plausible failure。
Mode I — Independence Expansion
不同 backend / data / evaluator。
Mode R — Replication
降低 stochastic / numerical uncertainty。
Mode T — Transport
calibration / external mapping。
Mode H — High-Fidelity Escalation
增加 dynamics / actor / temporal fidelity。
Mode U — Unknown-Region Exploration
建立新 ontology / new world family。
Mode X — Tail / Stress Exploration
低機率、高 impact 或安全失敗區域。
因此:
這只是 budget-coordinate representation,不要求:
永遠以固定比例分配。
本文提出 Adaptive Portfolio Rebalancing:
如果新 counterexample 出現:
可能合理。
如果跨 backend 已非常穩,但 transport debt 高:
如果 deadline 接近:
本文再建立 World Portfolio Frontier:
WDC 不要求 portfolio 被壓成:
本文同時處理 Rare-World Preservation。
假設某 world:
被 current model 認為:
但:
若它涉及:
- catastrophic safety failure;
- irreversible option loss;
- systemic collapse;
- security breach;
則其 computation value 可能高。
因此:
這與 WDC-03 的 diversity reserve 及 WDC-05 的 counterexample preservation 接軌。
本文提出 Tail Computation Reserve:
作為可選 governance component,用於:
- stress test;
- rare-event search;
- worst-case branch;
- adversarial agent policy。
它不是要求所有系統固定保留某百分比,而是禁止:
因為 leading posterior 很高,就把所有 tail worlds 永久砍光。
本文再建立 Discriminative World Value。
有時兩個 hypotheses:
都能解釋現有 evidence。
最有價值 world 不是:
最可能支持 的 world。
而是:
最能讓 與 產生不同 predictions 的 world。
定義:
或其他 task-specific discriminability metric。
這與 optimal experimental design 的核心精神相鄰:
選擇最能減少 relevant uncertainty 的 experiment。
因此:
本文再建立 Calibration World。
如果 target reality 有 historical cases:
可以專門建立:
用於:
- fit;
- validate;
- transport-debt estimation。
其價值不一定來自產生新 future,而是:
因此:
可以比:
更值得算。
本文再建立 World-Portfolio Dependency Penalty。
若 proposed computation:
高度依賴 active portfolio 中同一 family:
則新增 evidence 的 marginal independence value下降。
概念:
這直接延續 WDC-05。
本文因此提出:
Correlation-Aware World Allocation
Allocate compute according to the marginal contribution of a world computation to the evidence portfolio, not merely to the standalone quality of its world.
一個 standalone 很強的 world:
如果 portfolio 已有:
個幾乎相同 worlds,
其 marginal value 可能低。
反之,一個 standalone fidelity 稍低但真正獨立的 world:
可能有更高 portfolio value。
本文接著建立 Computation Admission Record:
computation_id
target_world_or_claim
operation
requested_budget
fidelity
deadline
expected_decision_value
expected_information_gain
expected_verification_value
expected_counterexample_value
expected_independence_gain
expected_transport_gain
tail_risk_value
estimated_cost
estimated_latency
meta_cost
portfolio_dependence
decision
執行後建立:
computation_id
realized_cost
realized_latency
realized_information_gain
decision_changed
counterexample_found
new_family_created
transport_debt_reduced
world_promoted
world_invalidated
followup_actions
這使 WDC-06 可以真正被實驗。
本文最後提出:
World Computation Value Principle
The object of allocation is not a world's abstract worth but the expected marginal value of a specific next computation on the current evidence and decision state.
Deficit-Directed Computation Principle
World computation should target the dominant epistemic or decision deficit—replication, independence, counterexample, transport, fidelity, tail risk, or unknown-world coverage—rather than reflexively adding more samples to the current leading family.
Portfolio, Not Winner Principle
A mature WDC Governor should maintain a portfolio of complementary computation modes rather than only expanding the current highest-ranked world.
Decision–Information Separation Principle
Expected information gain and expected decision improvement are distinct; a computation may be scientifically informative yet decision-irrelevant, or decision-critical with modest total information gain.
Dynamic Computation Principle
The value of a computation can include the future computations it enables or prevents; therefore world allocation should not be assumed purely myopic.
Fidelity Reliability Principle
Low-fidelity worlds may be used for screening and prioritization only to the degree that their relationship to the target is calibrated; cheap but unreliable worlds can increase rather than decrease total cost.
Meta-Boundedness Principle
Estimating the value of computation is itself a computation and must be bounded by its expected benefit.
下一篇 WDC-07 將研究:
World Ensemble Learning
《世界集合學習:如何讓跨世界結果反過來更新生成器、模型、Governor 與未來空間》
也就是:
我們已經會挑 world、跑 world、比較 world。那這些結果如何真正讓整個 WDC system 變聰明,而不是每輪重新從頭開始?
關鍵詞: Value of Computation、World Portfolio、Metareasoning、Bayesian Experimental Design、Multi-Fidelity Computation、Counterworlds、Transport Debt、World-Domain Governor、WDC
1. WDC-06 的真正 allocation unit 不是 World
Governor 若只對:
打永久分數,
容易產生錯誤。
2. Because the same World may need different next computations
例如:
可以:
- run more;
- fork;
- replicate;
- upgrade fidelity;
- change evaluator。
3. Therefore
4. Computation Action
5. Target Can Be Existing World
6. Or Missing World Family
7. Or Claim
8. Or Transport Gap
9. Operation Can Create a New World
10. Or No New World
11. This Is Important
Sometimes best next action is:
12. And Measure Reality
13. Current Evidence State
14. It Includes Cross-World Evidence Graph
15. Dependence
16. Unknown World Mass
17. Transport Debt
18. Decision Set
19. Deadline
20. Computation Changes Evidence State
21. WDC-06 Is a Meta-Level MDP-Like Problem
State:
Action:
Outcome:
22. But We Do Not Claim Standard MDP Assumptions Always Hold
23. Metareasoning
The system reasons about:
24. Rational Metareasoning External Calibration
Lieder et al. formulate algorithm selection as rational metareasoning。
25. Core Neighboring Idea
use expected consequences of computations to choose among cognitive strategies。
26. WDC Extension
strategy is not just algorithm。
It can be:
27. Decision VOC
28. Baseline Decision
29. After Computation
30. Improvement
31. Cost
Net:
32. This Requires Utility Model
If unavailable:
report unprojected vector。
33. Information Value
when probabilistic latent target exists。
34. Expected Information Gain
35. But EIG Can Be High for Irrelevant Parameter
36. Example
Learn precise parameter that doesn't affect action。
37. Decision Value Near Zero
38. Conversely
one binary test near action threshold。
39. Low entropy reduction
but flips decision。
40. Therefore
41. BOED External Calibration
Foster et al. call EIG central but computationally difficult。
42. This Meta-Point Matters
Even estimating experiment value can be expensive。
43. Zheng et al. 2020
explicitly allocate computation to refine MI estimates with variable costs。
44. WDC Analogy
We may need to allocate compute:
to estimate which world computation is worth doing。
45. Meta-Compute
46. Infinite Meta-Regress Risk
Should we compute the value of computing the value of computing?
47. Stop by bounded metareasoning
48. Or simple heuristic gate
49. Meta-Heuristic
if one candidate obviously dominates in cost/value bounds, act。
50. Only expensive VOC estimation near ambiguity
51. Value Bounds
52. If Upper Bound Below Cost
reject。
53. If Lower Bound Clearly Highest
admit。
54. This Mirrors Sequential Bound Refinement Idea
55. Dynamic VOC
Sezener & Dayan distinguish values accounting for future computations。
56. WDC Example
cheap world says:
possible failure。
57. That unlocks
58. So cheap world's total value includes follow-up opportunities
59. Option Value
60. Negative Option Value Also Possible
Computation may make deadline miss。
61. Computation Dependency Graph
62. Node = computation
63. Edge = unlock / condition / block
64. Example
65. If positive branch
66. If failure branch
67. If transport promising
68. This Is Sequential Program
69. Hyperband External Calibration
Hyperband treats configurations as arms and allocates finite resources adaptively。
70. Early Stop Poor Configurations
71. WDC Difference
Poor performance world can still be valuable counterevidence。
72. Therefore Elimination Criterion Is Evidence Value, Not Performance Alone
73. Successive Evidence Escalation
74. L0 Abstract Screening
cheap symbolic / analytical。
75. L1 Low-Fidelity World
76. L2 Replicated Medium Fidelity
77. L3 High Fidelity / Cross Backend
78. L4 External Test
79. Promotion Trigger
can be:
- support;
- counterexample;
- disagreement;
- tail risk;
- transport value。
80. Not Positive Result Only
81. Multi-Fidelity External Calibration
Song et al. study different mutually dependent information sources with different costs。
82. Key Neighboring Idea
cost-sensitive information allocation across fidelities。
83. WDC Extension
world fidelities may differ in:
- dynamics;
- actor realism;
- resolution;
- backend;
- observation detail。
84. Low Fidelity Reliability
85. Mikkola et al. Warning
unreliable lower-fidelity sources can worsen total optimization cost。
86. WDC Consequence
A cheap world family with bad transport can misallocate Governor budget。
87. Reliability-Aware Promotion
88. No Auto-Prune from Uncalibrated Low Fidelity
89. Epistemic Deficit
90. Why Deficit First?
Because same evidence count can hide different missingness。
91. Example A
1000 stochastic runs same backend。
92. Stochastic deficit small
93. Independence deficit high
94. Correct action
CrossBackend。
95. Wrong action
Run 1000 more seeds。
96. Example B
5 independent backends agree。
97. But no real calibration。
98. Transport deficit high。
99. Correct action
Calibrate。
100. Example C
Strong consensus, no counter search。
101. Counter deficit high。
102. Correct action
ForkCounter。
103. Example D
Model family known bad in rare tail。
104. Tail deficit high。
105. Correct action
StressTail。
106. Deficit-Directed Routing Table
107. This Can Be Learned
but initially rule-based。
108. Portfolio
109. Why Portfolio?
Because unknowns are plural。
110. One computation mode cannot cover all。
111. Support Budget
112. Counter Budget
113. Independence Budget
114. Replication Budget
115. Transport Budget
116. Fidelity Budget
117. Unknown Exploration
118. Tail Stress
119. Total
120. Not Fixed Percentages
121. Adaptive Rebalance
122. Portfolio Dependence
Even different modes can share same backend。
123. Need dependence penalty
124. Marginal Independence Gain
125. If Proposed World Is Near Duplicate
126. But Replication Value May Still Be High
127. Separate Them
128. Portfolio Diversity Is Functional
not visual diversity。
129. Cosmetic World Diversity
different prompts / colors。
130. Structural World Diversity
different error sources。
131. World-Mode Collapse
if all budget to one family。
132. But Anti-Mode-Collapse Can Overexplore
133. Need opportunity cost
134. Opportunity Cost
135. Compare with next-best computation
136. A Positive Value Computation Can Still Be Suboptimal
137. Resource Knapsack View
finite resources:
138. Each computation consumes vector cost。
139. Portfolio is constrained multi-resource selection
140. But outcomes adaptive
so static knapsack insufficient。
141. Sequential Policy
142. Computation Policy
This is Governor's meta-policy。
143. Myopic Policy
choose:
144. Dynamic Policy
accounts downstream。
145. Exact Dynamic Planning Often Intractable
146. Bounded Approximation Needed
147. WDC Doesn't Promise Optimal Governor
148. Deadline
149. Slow high-value computation
may arrive after decision。
150. Value Decays
151. Deadline-Adjusted VOC
152. Result Latency Is First-Class
153. Parallelism
Can run several worlds simultaneously。
154. But parallel computations may become redundant after one finishes
155. Batch Value
not necessarily additive。
156. Redundant Batch
two worlds answer same question。
157. Complementary Batch
one support, one counter, one transport。
158. Batch Scheduling
needs joint value estimate。
159. Diversity Reserve Helps
160. But Reserve Itself Costs compute
161. Tail-Risk Worlds
low probability/high impact。
162. Why Majority Scheduler Fails
posterior-low branches get no budget。
163. But safety may demand coverage。
164. Tail Utility
165. Domain-specific risk measure
Could be:
- failure probability;
- worst-case;
- CVaR;
- reachability risk。
166. WDC Does Not Mandate One
167. Safety-Critical Override
Some worlds computed due obligation
even if expected decision gain low。
168. Mandatory Compute
169. Example
minimum adversarial test before deployment。
170. Thus Portfolio Has Hard Constraints
not only utility optimization。
171. Decision Boundary Worlds
If two actions nearly tied,
compute worlds that discriminate。
172. Discriminative Value
173. If outcome distributions same under H1/H2
low discriminative value。
174. If strongly different
high value。
175. This Is Experimental-Design Logic
176. Scientific Hypothesis Testing
world can be chosen to maximize expected separation。
177. Not to maximize preferred hypothesis success。
178. Counterworld Design
Generate world where leading claim most likely fails
while remaining plausible。
179. Strongest Counterworld
conceptual only。
180. Avoid Unrealistic Strawman Counterworld
181. Counterworld Plausibility Contract
must remain inside admissible target assumptions。
182. Calibration World
world designed to test known real cases。
183. Its outcome is not new future prediction
but model credibility measurement。
184. Calibration Value
185. Transport Value
related but not identical。
186. Calibration can reveal
- systematic bias;
- scale mismatch;
- dynamics error。
187. Transport Debt Decomposition
from WDC-05:
188. Choose computation targeting largest debt component。
189. Example
high。
190. Need human behavior calibration
not higher physics fidelity。
191. Example
high。
192. Need better dynamics backend。
193. Therefore High Fidelity Must Be Typed
194. Fidelity Vector
195. Upgrade Relevant Coordinate Only
196. Don't pay for photorealism if claim about symbolic dynamics
197. Fidelity Waste
198. World Portfolio Frontier
multi-objective frontier。
199. Governor can present frontier to master/human
rather than hidden scalar ranking。
200. Explain Allocation
For each selected computation:
selected because independence deficit high。
201. Explain Rejection
rejected as redundant with family F3。
202. Explain Pause
marginal stochastic precision below threshold。
203. Explain Promotion
strong counterexample needs high-fidelity replication。
204. Explain External Test
simulation robustness high; transport debt now dominant。
205. This Is More Auditable Than One Score
206. Meta-Calibration
Governor predicts:
207. After outcome
measure:
208. Store pair
209. Calibration Curve
predicted vs realized gain。
210. Governor Overconfidence
if predicted high, realized low。
211. Governor Underexploration
if killed computations later found valuable。
212. Governance Miss
from WDC-03。
213. WDC-06 Reuses Misses to Learn Computation Value
214. Archived World as Training Data
Governance history can train meta-policy。
215. But Avoid Self-Confirmation
if only executed computations have labels。
216. Selection Bias
unexecuted worlds have unknown realized value。
217. Need exploration audit
randomly sample some low-priority computations。
218. Off-policy Evaluation Difficult
219. Finite Benchmark Helps
where exhaustive tree known。
220. Benchmark A — Exhaustive Small Portfolio
all computations run。
221. Learn optimal allocation under budget。
222. Compare Governor regret。
223. Benchmark B — Information vs Decision
one high EIG irrelevant world。
one low EIG decision-flipping world。
224. Governor should distinguish。
225. Benchmark C — Independence Deficit
many same-family runs。
one independent backend。
226. Correct next compute = independent backend。
227. Benchmark D — Counterexample Deficit
strong consensus but no adversarial search。
228. Correct = counterworld。
229. Benchmark E — Transport Deficit
many simulations agree。
known real calibration absent。
230. Correct = calibration / real-test proposal。
231. Benchmark F — Unreliable Low Fidelity
cheap world misranks candidates。
232. Governor should learn low 。
233. Benchmark G — Hyperband-Like Screening
many candidates, small cheap budgets。
promote informative subset。
234. Compare compute savings。
235. Benchmark H — Negative Result Promotion
cheap world finds critical failure。
should promote despite poor success metric。
236. Benchmark I — Tail Risk
rare catastrophic world with low probability。
ensure risk reserve。
237. Benchmark J — Deadline
accurate slow world vs coarser fast world。
238. Benchmark K — Meta-Overhead
VOC estimation expensive。
simple heuristic should win。
239. Benchmark L — Dynamic VOC
cheap screening unlocks valuable expensive experiment。
myopic policy misses it。
240. Benchmark M — Portfolio Redundancy
several computations look individually high-value but overlap heavily。
241. Benchmark N — World Family Expansion
all existing families share same assumption。
unknown-world exploration creates new family。
242. Benchmark O — Calibration World
world intentionally reproduces historical cases。
measure transport debt reduction。
243. Benchmark P — Governance Selection Bias
meta-policy trained only on executed worlds。
test random audit worlds reveal bias。
244. WDC-06 Principle I — Computation Unit
Governor allocation should target specific next computations, not attach permanent worth labels to worlds.
245. Principle II — Deficit-Directed Allocation
Choose computations according to the dominant unresolved evidence or decision deficit rather than automatically adding more runs to the leading world family.
246. Principle III — Information–Decision Separation
Expected information gain and expected decision improvement are distinct quantities and should be reported separately where both matter.
247. Principle IV — Portfolio Complementarity
The value of a world computation depends on what the current portfolio already contains; near-duplicate evidence has lower marginal independence value than genuinely complementary computation.
248. Principle V — Dynamic Computation
A computation can be valuable because it changes which later computations become worthwhile; next-computation value should not be assumed purely myopic.
249. Principle VI — Fidelity Reliability
Low-fidelity worlds should only be allowed to screen or terminate higher-fidelity computation to the extent that their relation to the target has been calibrated.
250. Principle VII — Tail Preservation
Low-probability worlds can retain high computation value when they probe catastrophic, irreversible, safety-critical, or highly discriminative regions.
251. Principle VIII — Transport First When Transport Is the Bottleneck
When cross-world robustness is already high but world-to-target transport debt dominates, additional simulation may be inferior to calibration or external testing.
252. Principle IX — Meta-Boundedness
Computing the value of computation has its own cost; WDC metareasoning must itself stop when additional allocation analysis is not worth its cost.
253. Principle X — Meta-Calibration
A Governor should compare predicted computation value with realized information, verification, transport, and decision gain, and update its allocation model over time.
254. 可否證條件
F254.1 Portfolio No-Gain
若 portfolio allocation 長期不優於簡單 equal / random / FIFO policies,WDC-06 應簡化。
F254.2 VOC Miscalibration
若 predicted VOC 與 realized decision improvement無關,decision-value model 需要重建。
F254.3 EIG Misuse
若高 EIG computations 持續不影響 target decisions,不能把 information gain 當 decision value proxy。
F254.4 Counterworld Waste
若 counterworld budget 只產生 unrealistic invalid failures,counterworld generator 需提高 admissibility。
F254.5 Independence Mispricing
若所謂 independent worlds 仍共享主要 error source, 被高估。
F254.6 Low-Fidelity Misrouting
若 low-fidelity worlds frequently prune high-value high-fidelity worlds,reliability gate 失效。
F254.7 Dynamic-VOC Overhead
若 non-myopic computation planning成本大於其收益,應退回 myopic / heuristic policy。
F254.8 Tail Overallocation
若 tail reserve 消耗大量 budget、卻對安全或 decision 完全無增量,應調整。
F254.9 Transport Neglect
若 simulation consensus 不斷增加但 external calibration 始終不進行,portfolio 失衡。
F254.10 Meta-Selection Bias
若 Governor 只從已執行 computations 學習而忽略 rejected-world outcomes,meta-calibration 可能自我封閉。
255. 與 WDC-07 的接口
WDC-06 現在回答:
下一個 world computation 應該做什麼?
但還有下一個問題:
這些 world computations 做完後,系統如何真正學習?
如果每輪只是:
- spawn;
- run;
- aggregate;
- archive;
然後下一輪重新從零開始,
WDC 仍只是昂貴的 simulation factory。
下一篇:
WDC-07 — World Ensemble Learning
《世界集合學習:跨世界結果如何更新生成器、模型、Governor 與未來空間》
將研究:
以及更危險的問題:
如果 Generator 只學自己以前生成的 worlds,會不會形成自我封閉的 world ontology?
256. 結論
WDC-03 問:
worlds 太多怎麼管理?
WDC-06 現在把問題再往前推:
管理的真正目標是什麼?
答案不是:
也不是:
真正 allocation unit 是:
而它的價值取決於:
與:
當 stochastic uncertainty 最大時:
多跑幾次。
當 independent evidence 不足時:
換 backend。
當 consensus 太舒服時:
找 counterworld。
當 tail risk 未知時:
stress-test。
當 transport debt 最大時:
停止增加模擬,去 calibration。
當兩個 hypotheses 都說得通時:
找最能區分它們的 world。
因此:
中文:
下一個最值得計算的世界,是對目前決策與證據投資組合具有最高邊際增量的世界,而不一定是現在最被看好的世界。
所以 WDC-06 真正把:
從一個 lifecycle / scheduler manager,
推成了:
World-Portfolio Metareasoner
它不再只問:
哪個 process 還有 GPU?
而問:
這也是 Branching World Computation 第一次真正碰到:
Claim Typing
| Claim | Type | Status |
|---|---|---|
| World worth 與 next-computation value 非同一 | D | Canonical separation |
| WDC allocation unit 可定義為 typed world-computation action | D | Proposed core formalization |
| Decision VOC 與 EIG 非同一 | D / E | Core distinction + external metareasoning/BOED analogue |
| Epistemic deficit 可引導不同 computation modes | D | Proposed routing framework |
| World portfolio 應包含 support/counter/independence/replication/transport/tail 等互補模式 | D | Proposed portfolio framework |
| Dynamic VOC 可包含 downstream computation value | D / E | Proposed WDC extension + MCTS analogue |
| Hyperband demonstrates adaptive finite-resource allocation and early stopping | E | External resource-allocation analogue |
| Multi-fidelity BO trades cheap dependent information sources against expensive evaluations | E | External multi-fidelity analogue |
| Unreliable low-fidelity sources can increase total optimization cost | E | External warning / calibration analogue |
| Highest posterior world should always receive most compute | — | Explicitly rejected |
| High information gain always means high decision value | — | Explicitly rejected |
| Low-probability world has low computation value | — | Explicitly rejected |
Evidence Ladder
本文目前主要位於:
- L0:world-computation action / value vector / deficit-directed portfolio;
- L1–L2:finite exhaustive portfolio benchmarks、VOC calibration、multi-fidelity routing;
- L3:rational metareasoning、MCTS VOC、BOED、Hyperband、multi-fidelity optimization 提供外部技術對照;
- L4:需要實際 WDC Governor 執行 adaptive world portfolio experiments;
- L5+:long-horizon real-world meta-calibration、world ensemble learning、autonomous portfolio adaptation 尚待後續。
參考文獻
Neo.K 內部正典與譜系
- Neo.K with Aletheia. World-Domain Governor. WDC-03 / BWC-03, 2026.
- Neo.K with Aletheia. Cross-World Evidence. WDC-05 / BWC-05, 2026.
- Neo.K with Aletheia. Branching World Graph. WDC-02 / BWC-02, 2026.
- Neo.K with Aletheia. Nested Agents and Observer Separation. WDC-04 / BWC-04, 2026.
- Neo.K with Aletheia. Six-Way Temporal Coupling. TCD-07, 2026.
- Neo.K with Aletheia. Prospective Constructive Intelligence. UCPNP Series II Paper 14, 2026.
External technical calibration
- Lieder, F., Plunkett, D., Hamrick, J. B., Russell, S. J., Hay, N. J., & Griffiths, T. L. Algorithm selection by rational metareasoning as a model of human strategy selection. NeurIPS 27, 2014.
- Sezener, E., & Dayan, P. Static and Dynamic Values of Computation in MCTS. UAI, PMLR 124:31–40, 2020.
- Foster, A., Jankowiak, M., Bingham, E., Horsfall, P., Teh, Y. W., Rainforth, T., & Goodman, N. Variational Bayesian Optimal Experimental Design. NeurIPS 32, 2019.
- Zheng, S., Hayden, D., Pacheco, J., & Fisher III, J. W. Sequential Bayesian Experimental Design with Variable Cost Structure. NeurIPS 33, 2020.
- Li, L., Jamieson, K., DeSalvo, G., Rostamizadeh, A., & Talwalkar, A. Hyperband: A Novel Bandit-Based Approach to Hyperparameter Optimization. JMLR 18(185):1–52, 2018.
- Song, J., Chen, Y., & Yue, Y. A General Framework for Multi-fidelity Bayesian Optimization with Gaussian Processes. AISTATS, PMLR 89:3158–3167, 2019.
- Mikkola, P., Martinelli, J., Filstroff, L., & Kaski, S. Multi-Fidelity Bayesian Optimization with Unreliable Information Sources. AISTATS, PMLR 206:7425–7454, 2023.
- Fan, M., Yoon, B.-J., Dougherty, E., Urban, N., Alexander, F., Arróyave, R., & Qian, X. Multi-fidelity Bayesian Optimization with Multiple Information Sources of Input-dependent Fidelity. UAI, PMLR 244, 2024.
- Foster, A., et al. Reverse-Annealed Sequential Monte Carlo for Efficient Bayesian Optimal Experiment Design. NeurIPS, 2025.
Public Version Disclaimer
本文是一個 world-computation allocation / metareasoning / simulation-portfolio framework。
本文不聲稱:
- world 的 intrinsic worth 可以由 computation value 衡量;
- WDC 的 VOC vector 有 universal estimator;
- EIG 是所有 science / engineering decisions 的最佳 utility;
- Hyperband、BOED、MCTS 或 multi-fidelity BO 等同 WDC;
- low-fidelity worlds 必然能省成本;
- counterworlds 應獲固定比例算力;
- tail-world budget 有跨 domain 通用最優值;
- Governor 可以只靠自動 scoring 取代人類/制度 authority;
- world portfolio metareasoning 已解決所有 normative research-priority 問題;
- 本文對 classical vs. 提供任何新證明。
本文真正建立的是:
以及: