title: "TCFT-04|高階推理不等於高階策略:認知成本、注意力分配與停止治理"
title_en: "High-Order Reasoning Is Not High-Order Strategy: Cognitive Cost, Attention Allocation, and Stopping Governance"
series: "Temporal Cognitive Frontier Theory (TCFT) / 時代認知前沿理論"
paper_no: "04"
version: "v0.1"
date: "2026-08-31"
author: "Neo.K"
affiliation: "EveMissLab / 一言諾科技有限公司"
document_type: "理論論文 / Reasoning Allocation Governance 篇"
language: "zh-Hant"
status: "正式系列初稿"
previous_paper: "TCFT-03|反身性推理:當推理本身進入被推理世界"
next_paper: "TCFT-05|時間正規化新穎度:Frozen-Time Prior Art 與同期認知基準"
TCFT-04|高階推理不等於高階策略:認知成本、注意力分配與停止治理
High-Order Reasoning Is Not High-Order Strategy: Cognitive Cost, Attention Allocation, and Stopping Governance
系列: Temporal Cognitive Frontier Theory, TCFT / 時代認知前沿理論篇次: 04版本: v0.1日期: 2026-08-31作者: Neo.K機構脈絡: EveMissLab / 一言諾科技有限公司
摘要
TCFT-01 至 TCFT-03 依序擴張了一個認知主體可處理的未來結構:未來底空間使更多候選世界進入可思考範圍;反事實視界使未發生世界成為可生成、比較與翻轉決策的候選;反身性推理則把推理者、模型、預測、公開、行動與他者反應重新納入下一輪世界與認知狀態。若只沿能力軸前進,一個自然但錯誤的結論是:想得越多、越深、越高階,策略就越好。
本文拒絕此命題。
High-Order Reasoning ≠ High-Order Strategy . \boxed{
\text{High-Order Reasoning}
\neq
\text{High-Order Strategy}.
} High-Order Reasoning = High-Order Strategy .
原因並不只在於人類有疲勞、時間有限或 AI 有 token budget。更一般地,任何有限智慧體都面對一個 Reasoning Allocation Problem(推理資源配置問題) :在多個對象、多個域、多個時間尺度、多種推理操作與多條候選分支之間,應把有限計算、注意力、搜尋、驗證與反身性深度配置在哪裡?
本文提出 Reasoning Allocation Governance(RAG,推理配置治理) 作為 TCFT 的策略層候選。對時間 t t t 的認知預算:
B t c o g = B t a t t + B t s e a r c h + B t s i m + B t c f + B t r r + B t v e r i f y + B t s w i t c h + B t m e t a . \boxed{
B_t^{cog}
=
B_t^{att}
+
B_t^{search}
+
B_t^{sim}
+
B_t^{cf}
+
B_t^{rr}
+
B_t^{verify}
+
B_t^{switch}
+
B_t^{meta}.
} B t co g = B t a tt + B t se a r c h + B t s im + B t c f + B t r r + B t v er i f y + B t s w i t c h + B t m e t a .
其中各項分別表示注意力、搜尋、模擬、反事實、反身推理、驗證、切換與元推理成本。此分解不是神經或計算架構的唯一真實分區,而是一個治理帳本。
本文將 Russell 與 Wefald 的 utility of computation / value of computation 、bounded rationality、anytime algorithm stopping、resource-rationality、rational inattention 與 expected value of control 視為重要前置文獻。它們共同指出:計算、注意與控制本身具有成本,元層決策必須回答「下一個計算步驟值不值得做」。TCFT-04 在此基礎上加入三個專門面向:第一,未來底空間、反事實與反身性會使搜索空間快速擴張;第二,推理資源配置本身具有 domain / target opportunity cost ;第三,局部模型解析度提高可能同時降低整體世界與其他主體的覆蓋。
本文提出:
NVOC t ( r ) = E [ Δ U t ( r ) ] − C t d i r e c t ( r ) − C t o p p ( r ) − C t r i s k ( r ) − C t d e l a y ( r ) \boxed{
\operatorname{NVOC}_t(r)
=
\mathbb E[
\Delta U_t(r)
]
-
C_t^{direct}(r)
-
C_t^{opp}(r)
-
C_t^{risk}(r)
-
C_t^{delay}(r)
} NVOC t ( r ) = E [ Δ U t ( r )] − C t d i r ec t ( r ) − C t o pp ( r ) − C t r i s k ( r ) − C t d e l a y ( r )
其中 NVOC \operatorname{NVOC} NVOC 為 Net Value of Computation , r r r 表示下一個 reasoning action。若:
NVOC t ( r ) ≤ 0 , \boxed{
\operatorname{NVOC}_t(r)\le 0,
} NVOC t ( r ) ≤ 0 ,
則在目前 task、budget、time horizon 與 uncertainty 下,繼續該 reasoning action 不再有正邊際價值。這提供一個比「想累了才停」更一般的停止條件。
本文進一步區分:
Capability ≠ Activation ≠ Allocation ≠ Termination . \boxed{
\text{Capability}
\neq
\text{Activation}
\neq
\text{Allocation}
\neq
\text{Termination}.
} Capability = Activation = Allocation = Termination .
一個 agent 可以具有很強的反身性推理能力,但在某一 task 選擇不用;也可以只進行低解析度人物模型,因為該模型已足以支撐行動。這不是能力不足,而可能是更高階的策略配置。
本文提出 Target Saturation、Domain Neglect、Local-Model Overinvestment、Reasoning Lock-In、Meta-Reasoning Regress、Switching Cost、Strategic Coarsening 與 Sufficient Resolution 等概念。特別地:
High-Resolution Local Model ⇏ High-Quality Global Strategy . \boxed{
\text{High-Resolution Local Model}
\not\Rightarrow
\text{High-Quality Global Strategy}.
} High-Resolution Local Model ⇒ High-Quality Global Strategy .
若把大量資源投入單一他者:
B t a r g e t ↑ , B_{target}\uparrow, B t a r g e t ↑ ,
則:
B s e l f + B o t h e r s + B w o r l d + B f u t u r e ↓ B_{self}
+
B_{others}
+
B_{world}
+
B_{future}
\downarrow B se l f + B o t h er s + B w or l d + B f u t u r e ↓
可能同時成立。此時即使對該他者的 inference accuracy 提升,global utility 仍可能下降。
本文最後把 RAG 接回 Cognitive Operator-Domain Theory(CODT)與既有「內外總作用量原理」。CODT 已把高階 method 寫成:
M e t h o d = P r o g r a m ( O p e r a t o r s , T o p o l o g y , C o n t e x t , P o l i c y , B u d g e t ) , Method
=
Program(
Operators,
Topology,
Context,
Policy,
Budget
), M e t h o d = P r o g r am ( O p er a t or s , T o p o l o g y , C o n t e x t , P o l i cy , B u d g e t ) ,
而既有總作用量理論已提出:
E [ Δ U n e x t ] ≤ Δ S n e x t \mathbb E[
\Delta U_{next}
]
\le
\Delta \mathcal S_{next} E [ Δ U n e x t ] ≤ Δ S n e x t
作為外部展開停止候選。TCFT-04 將其泛化到內部推理:推理、反事實、反身性與驗證也都是有成本的作用,策略優劣不在於是否把可用能力全部啟動,而在於是否把正確能力配置到正確地方、正確深度與正確時間。
本文中心命題為:
Strategic intelligence is not merely the ability to reason deeply, but the ability to allocate, switch, coarse-grain, and stop reasoning well. \boxed{
\text{Strategic intelligence is not merely the ability to reason deeply,
but the ability to allocate, switch, coarse-grain, and stop reasoning well.}
} Strategic intelligence is not merely the ability to reason deeply, but the ability to allocate, switch, coarse-grain, and stop reasoning well.
關鍵詞: 高階推理、高階策略、推理資源配置、Value of Computation、元推理、注意力成本、停止條件、bounded rationality、resource rationality、rational inattention、CODT、TCFT
Abstract
TCFT-01 through TCFT-03 progressively expanded the cognitive space available to an agent: future base-spaces introduced candidate futures; counterfactual horizons introduced unrealized alternatives; reflexive reasoning reintroduced the reasoner, its models, predictions, disclosures, actions, and induced reactions into subsequent world and reasoning states. A tempting but incorrect conclusion is that greater reasoning depth, breadth, or reflexivity automatically yields better strategy.
This paper rejects that conclusion:
High-Order Reasoning ≠ High-Order Strategy . \boxed{
\text{High-Order Reasoning}
\neq
\text{High-Order Strategy}.
} High-Order Reasoning = High-Order Strategy .
The reason is more general than human fatigue or AI token limits. Any finite cognitive system faces a Reasoning Allocation Problem : how should limited attention, computation, search, simulation, counterfactual expansion, reflexive depth, verification, switching, and metareasoning be allocated across targets, domains, timescales, and possible reasoning actions?
We propose Reasoning Allocation Governance (RAG) as a strategy-layer candidate within TCFT. A cognitive budget is decomposed as a governance ledger rather than a claim about literal neural modules:
B t c o g = B t a t t + B t s e a r c h + B t s i m + B t c f + B t r r + B t v e r i f y + B t s w i t c h + B t m e t a . \boxed{
B_t^{cog}
=
B_t^{att}
+
B_t^{search}
+
B_t^{sim}
+
B_t^{cf}
+
B_t^{rr}
+
B_t^{verify}
+
B_t^{switch}
+
B_t^{meta}.
} B t co g = B t a tt + B t se a r c h + B t s im + B t c f + B t r r + B t v er i f y + B t s w i t c h + B t m e t a .
Building on utility-of-computation metareasoning, bounded rationality, anytime-algorithm stopping, resource-rational analysis, rational inattention, and expected-value-of-control frameworks, we define a candidate Net Value of Computation:
NVOC t ( r ) = E [ Δ U t ( r ) ] − C t d i r e c t ( r ) − C t o p p ( r ) − C t r i s k ( r ) − C t d e l a y ( r ) . \boxed{
\operatorname{NVOC}_t(r)
=
\mathbb E[
\Delta U_t(r)
]
-
C_t^{direct}(r)
-
C_t^{opp}(r)
-
C_t^{risk}(r)
-
C_t^{delay}(r).
} NVOC t ( r ) = E [ Δ U t ( r )] − C t d i r ec t ( r ) − C t o pp ( r ) − C t r i s k ( r ) − C t d e l a y ( r ) .
When NVOC t ( r ) ≤ 0 \operatorname{NVOC}_t(r)\le 0 NVOC t ( r ) ≤ 0 , continuing that reasoning action no longer has positive marginal value under the current task, budget, uncertainty, and horizon.
The paper distinguishes capability, activation, allocation, and termination. It introduces Target Saturation, Domain Neglect, Local-Model Overinvestment, Reasoning Lock-In, Meta-Reasoning Regress, Switching Cost, Strategic Coarsening, and Sufficient Resolution. A particularly important claim is that increasing the resolution of a local model can reduce global strategic quality by consuming resources that could have been allocated to self-modeling, third parties, environment, future contingencies, or validation.
Finally, the framework is connected to Cognitive Operator-Domain Theory (CODT) and prior work on total internal-external action costs. The central proposition is:
Strategic intelligence is not merely the ability to reason deeply, but the ability to allocate, switch, coarse-grain, and stop reasoning well. \boxed{
\text{Strategic intelligence is not merely the ability to reason deeply,
but the ability to allocate, switch, coarse-grain, and stop reasoning well.}
} Strategic intelligence is not merely the ability to reason deeply, but the ability to allocate, switch, coarse-grain, and stop reasoning well.
Keywords: high-order reasoning; strategy; value of computation; metareasoning; attention allocation; stopping; bounded rationality; resource rationality; rational inattention; TCFT
1. 導論:能力擴張之後,為什麼需要治理?
前面三篇逐步回答:
我能想到哪些未來? \boxed{
\text{我能想到哪些未來?}
} 我能想到哪些未來?
我能生成哪些未發生世界? \boxed{
\text{我能生成哪些未發生世界?}
} 我能生成哪些未發生世界?
我能否把自己與他人的反應放回推理? \boxed{
\text{我能否把自己與他人的反應放回推理?}
} 我能否把自己與他人的反應放回推理?
但一個真正有限的智慧體還必須回答:
我現在應該把計算資源花在哪裡? \boxed{
\text{我現在應該把計算資源花在哪裡?}
} 我現在應該把計算資源花在哪裡?
以及:
我什麼時候應該停止? \boxed{
\text{我什麼時候應該停止?}
} 我什麼時候應該停止?
2. 高階推理不等於高階策略
假設 agent A 可以做:
R R 10 , RR_{10}, R R 10 ,
agent B 只做:
R R 2 . RR_2. R R 2 .
不能推出:
S t r a t e g y ( A ) > S t r a t e g y ( B ) . Strategy(A)>Strategy(B). S t r a t e g y ( A ) > S t r a t e g y ( B ) .
因為:
R R 10 RR_{10} R R 10
可能提供幾乎零新增決策價值,
卻消耗大量:
time;
attention;
compute;
opportunity;
verification;
delay budget。
因此:
Depth ⇏ Value . \boxed{
\text{Depth}
\not\Rightarrow
\text{Value}.
} Depth ⇒ Value .
3. Strategy 是 Reasoning over Reasoning Allocation
普通 reasoning:
R : S t a t e → C o n c l u s i o n . R:
State
\rightarrow
Conclusion. R : S t a t e → C o n c l u s i o n .
strategy-level metareasoning 則問:
M R : S t a t e → Which reasoning action should be executed next? \boxed{
M_R:
State
\rightarrow
\text{Which reasoning action should be executed next?}
} M R : S t a t e → Which reasoning action should be executed next?
所以 reasoning action 本身成為 decision variable。
4. Computational Action
令:
r t r_t r t
表示一個 reasoning action。
它可以是:
搜尋一個來源;
展開一個 counterfactual;
增加一階 recursive ToM;
驗證一個 assumption;
模擬一個 scenario;
改變 representation;
切換 domain;
查證一個 prior;
停止。
因此:
r t ∈ R t m e t a . \boxed{
r_t
\in
\mathcal R_t^{meta}.
} r t ∈ R t m e t a .
5. Value of Computation
Russell 與 Wefald 的 metareasoning 傳統核心問題之一正是:
一個計算步驟因為可能改變外部行動,其價值是多少?
TCFT 沿用此精神。
令:
U t ∗ U_t^* U t ∗
為目前最佳可行決策的 expected utility。
執行 reasoning action:
r r r
之後,可能更新為:
U t + 1 ∗ . U_{t+1}^*. U t + 1 ∗ .
則粗略 expected improvement:
VOC ( r ) = E [ U t + 1 ∗ − U t ∗ ] . \boxed{
\operatorname{VOC}(r)
=
\mathbb E[
U_{t+1}^*-U_t^*
].
} VOC ( r ) = E [ U t + 1 ∗ − U t ∗ ] .
6. Net Value of Computation
TCFT 進一步扣除多種成本:
NVOC t ( r ) = E [ Δ U t ( r ) ] − C t d i r e c t ( r ) − C t o p p ( r ) − C t r i s k ( r ) − C t d e l a y ( r ) . \boxed{
\operatorname{NVOC}_t(r)
=
\mathbb E[
\Delta U_t(r)
]
-
C_t^{direct}(r)
-
C_t^{opp}(r)
-
C_t^{risk}(r)
-
C_t^{delay}(r).
} NVOC t ( r ) = E [ Δ U t ( r )] − C t d i r ec t ( r ) − C t o pp ( r ) − C t r i s k ( r ) − C t d e l a y ( r ) .
其中:
C d i r e c t C^{direct} C d i r ec t
為直接計算/注意/工具成本。
C o p p C^{opp} C o pp
為因執行 r r r 而不能做其他事的機會成本。
C r i s k C^{risk} C r i s k
為錯誤展開、過度自信、資訊污染或不可逆行動相關風險。
C d e l a y C^{delay} C d e l a y
為延遲決策造成的價值損失。
7. 停止條件
若:
NVOC t ( r ) ≤ 0 , \boxed{
\operatorname{NVOC}_t(r)\le 0,
} NVOC t ( r ) ≤ 0 ,
則繼續這個 reasoning action 沒有正邊際價值。
因此:
Stop ⇏ No More Reasoning Is Possible . \boxed{
\text{Stop}
\not\Rightarrow
\text{No More Reasoning Is Possible}.
} Stop ⇒ No More Reasoning Is Possible .
而是:
No Currently Available Reasoning Action Has Positive Net Value . \boxed{
\text{No Currently Available Reasoning Action Has Positive Net Value}.
} No Currently Available Reasoning Action Has Positive Net Value .
8. Global Stopping
若:
max r ∈ R t m e t a NVOC t ( r ) ≤ 0 , \max_{r\in\mathcal R_t^{meta}}
\operatorname{NVOC}_t(r)
\le0, r ∈ R t m e t a max NVOC t ( r ) ≤ 0 ,
可候選性地停止整個 deliberation:
Act / Commit / Defer / Idle . \boxed{
\text{Act / Commit / Defer / Idle}.
} Act / Commit / Defer / Idle .
實際 outcome 仍依任務與權限而定。
9. Bounded Rationality
bounded rationality 的核心不是:
有限智慧體比較笨。
而是:
optimization itself has costs and limits . \boxed{
\text{optimization itself has costs and limits}.
} optimization itself has costs and limits .
即使存在理論上的最佳答案:
a ∗ , a^*, a ∗ ,
找到:
a ∗ a^* a ∗
可能比使用一個足夠好的:
a ~ \tilde a a ~
成本高得多。
10. Bounded Optimality 與 TCFT
因此真正策略目標可以從:
arg max a U ( a ) \arg\max_a U(a) arg a max U ( a )
轉向:
arg max π E [ U ( O u t c o m e ( π ) ) − C ( π ) ] . \boxed{
\arg\max_{\pi}
\mathbb E[
U(
Outcome(\pi)
)
-
C(\pi)
].
} arg π max E [ U ( O u t co m e ( π )) − C ( π )] .
其中:
π \pi π
包含整個 reasoning policy,
而不是單一外部 action。
11. Resource Rationality
resource-rational analysis 把認知策略視為在有限 computational resources 下的適應性/近似最適配置。
TCFT 接受這個重要方向。
但 TCFT-04 特別強調:
the resource budget is distributed across domains, targets, reasoning modes, and reflexive depths . \boxed{
\text{the resource budget is distributed across domains,
targets, reasoning modes, and reflexive depths}.
} the resource budget is distributed across domains, targets, reasoning modes, and reflexive depths .
不是只有同一問題內的計算深度。
12. Rational Inattention
rational inattention 提醒:
information processing itself is costly . \boxed{
\text{information processing itself is costly}.
} information processing itself is costly .
因此不是所有可取得資訊都值得處理。
13. Expected Value of Control
Expected Value of Control 模型把 expected payoff、control intensity 與 effort cost 放進同一控制配置問題。
TCFT 將此精神泛化為:
How much cognitive control should be allocated, to what, and for how long? \boxed{
\text{How much cognitive control should be allocated,
to what, and for how long?}
} How much cognitive control should be allocated, to what, and for how long?
14. Cognitive Budget Ledger
定義:
B t c o g = B t a t t + B t s e a r c h + B t s i m + B t c f + B t r r + B t v e r i f y + B t s w i t c h + B t m e t a . \boxed{
B_t^{cog}
=
B_t^{att}
+
B_t^{search}
+
B_t^{sim}
+
B_t^{cf}
+
B_t^{rr}
+
B_t^{verify}
+
B_t^{switch}
+
B_t^{meta}.
} B t co g = B t a tt + B t se a r c h + B t s im + B t c f + B t r r + B t v er i f y + B t s w i t c h + B t m e t a .
它不是聲稱認知真的必須物理切成八個池,而是一個治理帳本。
15. Capability、Activation、Allocation、Termination
TCFT-04 的核心 separation:
Capability ≠ Activation ≠ Allocation ≠ Termination . \boxed{
\text{Capability}
\neq
\text{Activation}
\neq
\text{Allocation}
\neq
\text{Termination}.
} Capability = Activation = Allocation = Termination .
一個 agent 有能力:
R R 10 RR_{10} R R 10
不代表每次都該啟動:
R R 10 . RR_{10}. R R 10 .
16. Strategic Restraint
若 task 只需:
R R 2 , RR_2, R R 2 ,
但:
C a p ( R R 10 ) = 1 , Cap(RR_{10})=1, C a p ( R R 10 ) = 1 ,
合理策略仍可:
A c t ( R R 10 ) = 0. \boxed{
Act(RR_{10})=0.
} A c t ( R R 10 ) = 0.
這不是能力不足,而可能是:
Strategic Restraint . \boxed{
\text{Strategic Restraint}.
} Strategic Restraint .
17. Sufficient Resolution
對 target X X X ,令模型解析度為:
ρ X . \rho_X. ρ X .
若:
D e c i s i o n ( M X ρ ) = D e c i s i o n ( M X ρ + Δ ) \boxed{
Decision(
M_X^{\rho}
)
=
Decision(
M_X^{\rho+\Delta}
)
} D ec i s i o n ( M X ρ ) = D ec i s i o n ( M X ρ + Δ )
在合理 uncertainty 與 perturbation 下持續穩定,則可候選地稱為:
Sufficient Resolution . \boxed{
\text{Sufficient Resolution}.
} Sufficient Resolution .
18. Strategic Coarsening
如果:
M X c o a r s e M_X^{coarse} M X co a r se
已足以支持 action,
則:
Strategic Coarsening \boxed{
\text{Strategic Coarsening}
} Strategic Coarsening
可能比最大解析度更好。
粗略不必等於錯誤。
19. Local-Model Overinvestment
定義:
Local-Model Overinvestment \boxed{
\text{Local-Model Overinvestment}
} Local-Model Overinvestment
當:
Δ A c c u r a c y t a r g e t > 0 \Delta Accuracy_{target}>0 Δ A cc u r a c y t a r g e t > 0
但:
Δ U g l o b a l < 0. \Delta U_{global}<0. Δ U g l o ba l < 0.
也就是:
局部模型變準,整體策略反而變差。
20. 為什麼局部變準仍可能全域變差?
因為:
B t o t a l = B t a r g e t + B s e l f + B o t h e r s + B w o r l d + B f u t u r e + B v e r i f y . \boxed{
B_{total}
=
B_{target}
+
B_{self}
+
B_{others}
+
B_{world}
+
B_{future}
+
B_{verify}.
} B t o t a l = B t a r g e t + B se l f + B o t h er s + B w or l d + B f u t u r e + B v er i f y .
若:
B t a r g e t ↑ , B_{target}\uparrow, B t a r g e t ↑ ,
其他項可能下降。
21. Target Saturation
若對 target j j j 的邊際增益:
Δ U j ( b ) \Delta U_j(b) Δ U j ( b )
隨 budget 下降,
且:
∂ E [ U ] ∂ b j ≤ 0 , \boxed{
\frac{\partial \mathbb E[U]}{\partial b_j}
\le0,
} ∂ b j ∂ E [ U ] ≤ 0 ,
則可候選稱為:
Target Saturation . \boxed{
\text{Target Saturation}.
} Target Saturation .
22. Domain Neglect
若 agent 長期大量投入:
D 1 D_1 D 1
而忽略:
D 2 , … , D n , D_2,\ldots,D_n, D 2 , … , D n ,
即使:
P e r f o r m a n c e ( D 1 ) Performance(D_1) P er f or man ce ( D 1 )
非常高,
仍可能形成:
Domain Neglect . \boxed{
\text{Domain Neglect}.
} Domain Neglect .
23. Domain Allocation Problem
令:
D = { D 1 , … , D n } . \mathcal D
=
\{D_1,\ldots,D_n\}. D = { D 1 , … , D n } .
分配:
b 1 , … , b n b_1,\ldots,b_n b 1 , … , b n
subject to:
∑ k = 1 n b k ≤ B . \sum_{k=1}^{n}b_k\le B. k = 1 ∑ n b k ≤ B .
策略問題:
max b E [ U ( b 1 , … , b n ) ] . \boxed{
\max_{\mathbf b}
\mathbb E[
U(
b_1,\ldots,b_n
)
].
} b max E [ U ( b 1 , … , b n )] .
24. 平均分配不等於合理
最佳配置一般不要求:
b 1 = b 2 = ⋯ = b n . b_1=b_2=\cdots=b_n. b 1 = b 2 = ⋯ = b n .
因為不同 domain 具有不同:
urgency;
uncertainty;
expected impact;
risk;
decision sensitivity。
25. Attention Opportunity Cost
把注意力放在 X X X ,意味此刻無法完整放在 Y Y Y 。
因此:
C o p p ( X ) \boxed{
C^{opp}(X)
} C o pp ( X )
是推理配置的實質成本,而不只是心理疲勞。
26. 自己也是被配置的 target
分析:
M A ( B ) M_A(B) M A ( B )
時,還要問:
M A ( A ∣ M A ( B ) ) . \boxed{
M_A(
A\mid M_A(B)
).
} M A ( A ∣ M A ( B )) .
也就是:
我正在因為研究 B 而做什麼?
27. Self-Game
現在的自己:
A t A_t A t
可以選擇繼續深挖。
未來的自己:
A t + 1 A_{t+1} A t + 1
承擔延遲、時間與 lost alternatives。
所以:
A t ↔ A t + 1 \boxed{
A_t
\leftrightarrow
A_{t+1}
} A t ↔ A t + 1
形成一種 intertemporal self-game。
28. Value of Information 不等於 Value of Computation
某資訊如果拿到可能極有價值:
V O I ( I ) ≫ 0. VOI(I)\gg0. V O I ( I ) ≫ 0.
但取得它可能極昂貴。
因此:
V O I ≠ V O C . \boxed{
VOI
\neq
VOC.
} V O I = V O C .
29. Value of Search
搜尋 s s s 的預期價值可候選寫成:
V O S ( s ) = E [ V O I ( R e s u l t ( s ) ) ] − C ( s ) . \boxed{
VOS(s)
=
\mathbb E[
VOI(
Result(s)
)
]
-
C(s).
} V O S ( s ) = E [ V O I ( R es u l t ( s ))] − C ( s ) .
30. Value of Verification
驗證 v v v :
V O V ( v ) = E x p e c t e d L o s s A v o i d e d ( v ) − C ( v ) . \boxed{
VOV(v)
=
ExpectedLossAvoided(v)
-
C(v).
} V O V ( v ) = E x p ec t e d L oss A v o i d e d ( v ) − C ( v ) .
因此不需要所有 claim 都用相同驗證強度。
31. Verification Allocation
高風險、不可逆、高外部性的 claim:
B v e r i f y ↑ . B_{verify}\uparrow. B v er i f y ↑ .
探索性、低風險假說:
B v e r i f y B_{verify} B v er i f y
可以較低。
32. Delay Cost
若 deadline:
T d , T_d, T d ,
reasoning 耗時:
τ r , \tau_r, τ r ,
則可能產生:
C d e l a y ( r ) . \boxed{
C^{delay}(r).
} C d e l a y ( r ) .
因此:
better answer later \text{better answer later} better answer later
可能輸給:
good-enough answer now . \text{good-enough answer now}. good-enough answer now .
33. Anytime Reasoning
anytime algorithm 的精神是:
多算通常可以改善答案,但任何時點都可返回目前解。
TCFT 將其抽象到 reasoning。
真正需要決定:
t s t o p . \boxed{
t_{stop}.
} t s t o p .
34. Stop 是一個 Meta-Action
停止不是沒有 action。
S t o p ∈ R m e t a . \boxed{
Stop
\in
\mathcal R^{meta}.
} S t o p ∈ R m e t a .
與:
Search;
Expand;
Verify;
Simulate;
同樣是策略選項。
35. Optimal Stopping Candidate
如果:
max r NVOC ( r ) ≤ 0 , \max_r
\operatorname{NVOC}(r)
\le0, r max NVOC ( r ) ≤ 0 ,
則:
S t o p \boxed{
Stop
} S t o p
可能是最合理選擇。
所以:
Knowing when not to think further is part of strategic intelligence . \boxed{
\text{Knowing when not to think further
is part of strategic intelligence}.
} Knowing when not to think further is part of strategic intelligence .
36. Premature Stopping
若存在:
r ∗ r^* r ∗
使:
NVOC ( r ∗ ) ≫ 0 \operatorname{NVOC}(r^*)\gg0 NVOC ( r ∗ ) ≫ 0
但 agent 已停止,
則:
Premature Stopping . \boxed{
\text{Premature Stopping}.
} Premature Stopping .
37. Overthinking
若持續執行:
r 1 , r 2 , … r_1,r_2,\ldots r 1 , r 2 , …
而:
NVOC ( r k ) < 0 \operatorname{NVOC}(r_k)<0 NVOC ( r k ) < 0
持續成立,
可候選性地稱:
Overthinking . \boxed{
\text{Overthinking}.
} Overthinking .
這是一個功能定義,不是心理診斷。
38. Reasoning Lock-In
若 agent 因已投入大量:
C p a s t C_{past} C p a s t
而繼續同一路徑,即使未來邊際價值低,
形成:
Reasoning Lock-In . \boxed{
\text{Reasoning Lock-In}.
} Reasoning Lock-In .
已付出的 reasoning cost 不應自動增加未來 reasoning value。
39. Switching Cost
從:
D i D_i D i
切換到:
D j D_j D j
具有:
C s w i t c h ( i , j ) . C_{switch}(i,j). C s w i t c h ( i , j ) .
可能包括:
context rebuild;
memory load;
representation shift;
tool shift;
synchronization。
40. Contextual Inertia
如果:
C s w i t c h C_{switch} C s w i t c h
很高,
agent 可能即使知道別的 domain 更重要,也不切。
形成:
Contextual Inertia . \boxed{
\text{Contextual Inertia}.
} Contextual Inertia .
41. Thrashing
反過來,頻繁切換會使:
C s w i t c h ↑ C_{switch}\uparrow C s w i t c h ↑
形成:
Thrashing . \boxed{
\text{Thrashing}.
} Thrashing .
策略必須在 Lock-In 與 Thrashing 之間取得平衡。
42. Domain Switching Policy
候選:
π s w i t c h ( S t ) → { S t a y , S w i t c h ( D j ) , P a u s e , M e r g e , D e l e g a t e } . \boxed{
\pi_{switch}(S_t)
\rightarrow
\{
Stay,
Switch(D_j),
Pause,
Merge,
Delegate
\}.
} π s w i t c h ( S t ) → { S t a y , S w i t c h ( D j ) , P a u se , M er g e , D e l e g a t e } .
43. Delegation
如果 reasoning action:
r r r
可交給另一 agent 或 tool,
策略集合就不是:
{ D o , N o t D o } \{Do,NotDo\} { D o , N o t D o }
而是:
{ D o , D e l e g a t e , D e f e r , S t o p } . \boxed{
\{Do,Delegate,Defer,Stop\}.
} { D o , D e l e g a t e , D e f er , S t o p } .
44. Delegation Value
V O D ( r , j ) = E x p e c t e d G a i n ( r , j ) − C o o r d i n a t i o n C o s t − V e r i f i c a t i o n C o s t − T r u s t R i s k . \boxed{
VOD(r,j)
=
ExpectedGain(r,j)
-
CoordinationCost
-
VerificationCost
-
TrustRisk.
} V O D ( r , j ) = E x p ec t e d G ain ( r , j ) − C oor d ina t i o n C os t − V er i f i c a t i o n C os t − T r u s tR i s k .
45. Multi-Agent Allocation
對 agents:
A 1 , … , A n , A_1,\ldots,A_n, A 1 , … , A n ,
目標不是讓所有 agent 重複相同 computation,
而是:
maximize complementary coverage per total cost . \boxed{
\text{maximize complementary coverage per total cost}.
} maximize complementary coverage per total cost .
46. Redundancy 不是永遠浪費
如果獨立 replication 能顯著降低 model error,
則:
Redundancy ≠ Waste . \boxed{
\text{Redundancy}
\neq
\text{Waste}.
} Redundancy = Waste .
47. Correlated Redundancy
若所有 agents 共用:
same model;
same data;
same prompt;
same blindspot;
重複十次可能只產生:
correlated confidence . \boxed{
\text{correlated confidence}.
} correlated confidence .
48. Attention to Other Others
在 social reasoning 中,如果只模型:
B , B, B ,
容易忽略:
C , D , E , … C,D,E,\ldots C , D , E , …
但 relationship outcome 可能主要受到第三方支配。
所以:
Target-Centric Reasoning ⇏ System-Centric Strategy . \boxed{
\text{Target-Centric Reasoning}
\not\Rightarrow
\text{System-Centric Strategy}.
} Target-Centric Reasoning ⇒ System-Centric Strategy .
49. World Outside the Social Graph
甚至不能把所有注意力都留給 agents。
還有:
environment;
technology;
institution;
deadline;
random shock;
resource state。
因此:
B s o c i a l ≠ B g l o b a l . \boxed{
B_{social}
\neq
B_{global}.
} B soc ia l = B g l o ba l .
50. Strategic Scope
令:
S t \mathcal S_t S t
為目前納入策略的 scope。
若:
S t \mathcal S_t S t
過窄,
則高深度 reasoning 可能是:
locally sophisticated, globally blind . \boxed{
\text{locally sophisticated, globally blind}.
} locally sophisticated, globally blind .
51. Scope-Depth Tradeoff
固定 budget 下常出現:
D e p t h ↑ ⇒ B r e a d t h ↓ \boxed{
Depth\uparrow
\Rightarrow
Breadth\downarrow
} D e pt h ↑⇒ B r e a d t h ↓
的資源 tradeoff。
這不是邏輯必然,而是有限資源條件。
52. Strategic Intelligence Vector
TCFT-04 可保存:
S ⃗ i = ( A l l o c a t i o n , S w i t c h i n g , S t o p p i n g , C o a r s e n i n g , V e r i f i c a t i o n , D e l e g a t i o n , S c o p e C o n t r o l , M e t a E f f i c i e n c y ) . \boxed{
\vec S_i
=
(
Allocation,
Switching,
Stopping,
Coarsening,
Verification,
Delegation,
ScopeControl,
MetaEfficiency
).
} S i = ( A l l oc a t i o n , S w i t c hin g , S t o pp in g , C o a r se nin g , V er i f i c a t i o n , D e l e g a t i o n , S co p e C o n t r o l , M e t a E f f i c i e n cy ) .
而不是壓成一個「策略 IQ」。
53. Reasoning Efficiency
定義候選:
η R = Δ U d e c i s i o n C c o g . \boxed{
\eta_R
=
\frac{
\Delta U_{decision}
}{
C_{cog}
}.
} η R = C co g Δ U d ec i s i o n .
但:
E f f i c i e n c y ≠ O b j e c t i v e V a l i d i t y . \boxed{
Efficiency
\neq
ObjectiveValidity.
} E f f i c i e n cy = O bj ec t i v e V a l i d i t y .
更有效率地完成錯誤目標仍然是錯。
54. Goal Review
因此治理中要保留:
B g o a l − r e v i e w . \boxed{
B_{goal-review}.
} B g o a l − r e v i e w .
在:
anomaly;
repeated failure;
regime shift;
時提升。
55. Adaptive Governance
若:
R e g i m e t ≠ R e g i m e t + 1 , Regime_t
\neq
Regime_{t+1}, R e g im e t = R e g im e t + 1 ,
原配置策略可以:
π R ∗ ( t ) ≠ π R ∗ ( t + 1 ) . \boxed{
\pi_R^*(t)
\neq
\pi_R^*(t+1).
} π R ∗ ( t ) = π R ∗ ( t + 1 ) .
RAG 必須可版本化與情境化。
56. Cognitive Cost 不只有時間
至少可以拆:
C c o g = C t i m e + C c o m p u t e + C a t t e n t i o n + C m e m o r y + C s w i t c h + C v e r i f y + C d e l a y + C r i s k . \boxed{
C_{cog}
=
C_{time}
+
C_{compute}
+
C_{attention}
+
C_{memory}
+
C_{switch}
+
C_{verify}
+
C_{delay}
+
C_{risk}.
} C co g = C t im e + C co m p u t e + C a tt e n t i o n + C m e m or y + C s w i t c h + C v er i f y + C d e l a y + C r i s k .
57. Token Cost 只是 AI 的一部分
對 LLM:
C t o k e n C_{token} C t o k e n
只是:
C t o k e n ⊂ C c o g − t o t a l . \boxed{
C_{token}
\subset
C_{cog-total}.
} C t o k e n ⊂ C co g − t o t a l .
工具、搜尋、驗證、狀態同步與 human review 可能更昂貴。
58. Low Subjective Effort 不等於 Low Strategic Cost
即使人類主觀上「不累」,
推理仍消耗:
time;
opportunity;
attention channel;
delayed action。
所以:
Low Subjective Effort ⇏ Low Strategic Cost . \boxed{
\text{Low Subjective Effort}
\not\Rightarrow
\text{Low Strategic Cost}.
} Low Subjective Effort ⇒ Low Strategic Cost .
59. More Capability Can Increase Allocation Complexity
如果:
C a p ↑ , Cap\uparrow, C a p ↑ ,
可選 reasoning actions:
∣ R m e t a ∣ ↑ . |\mathcal R^{meta}|\uparrow. ∣ R m e t a ∣ ↑ .
因此:
More Capability ⇏ Simpler Strategy . \boxed{
\text{More Capability}
\not\Rightarrow
\text{Simpler Strategy}.
} More Capability ⇒ Simpler Strategy .
60. Reasoning Inflation
AI 讓單次 reasoning 變便宜:
C ( r ) ↓ . C(r)\downarrow. C ( r ) ↓ .
但如果使用量:
N ( r ) ↑ ↑ , N(r)\uparrow\uparrow, N ( r ) ↑↑ ,
總成本:
N ( r ) C ( r ) N(r)C(r) N ( r ) C ( r )
不一定下降。
本文暫稱:
Reasoning Inflation . \boxed{
\text{Reasoning Inflation}.
} Reasoning Inflation .
61. Verification Bottleneck
AI 可以快速生成:
10 4 10^4 1 0 4
候選,
但 human / verifier 只能檢查:
20. 20. 20.
則真正 bottleneck 是:
B v e r i f y . \boxed{
B_{verify}.
} B v er i f y .
這時再增加 generation 可能沒有正邊際價值。
62. TCFT-01 的停止問題
未來底空間:
Ω t \Omega_t Ω t
何時不再擴張?
當:
E [ Δ U n e w − f u t u r e ] ≤ C e x p a n d + C e v a l u a t e + C d e l a y . \boxed{
\mathbb E[
\Delta U_{new-future}
]
\le
C_{expand}
+
C_{evaluate}
+
C_{delay}.
} E [ Δ U n e w − f u t u r e ] ≤ C e x p an d + C e v a l u a t e + C d e l a y .
63. TCFT-02 的停止問題
反事實視界:
C t \mathcal C_t C t
何時停止?
當:
max c n e w E x p e c t e d R e v e r s a l V a l u e ( c n e w ) ≤ C o s t ( c n e w ) . \boxed{
\max_{c_{new}}
ExpectedReversalValue(c_{new})
\le
Cost(c_{new}).
} c n e w max E x p ec t e d R e v er s a l V a l u e ( c n e w ) ≤ C os t ( c n e w ) .
64. TCFT-03 的停止問題
反身深度:
d R d_R d R
何時停止?
當:
Δ D e c i s i o n V a l u e ( d R + 1 ) ≤ Δ C o s t ( d R + 1 ) . \boxed{
\Delta DecisionValue(d_R+1)
\le
\Delta Cost(d_R+1).
} Δ D ec i s i o nV a l u e ( d R + 1 ) ≤ Δ C os t ( d R + 1 ) .
65. Unified Marginal Stopping Principle
因此可統一成:
E [ Δ U n e x t ] ≤ Δ C n e x t + Δ C o p p , n e x t + Δ C r i s k , n e x t . \boxed{
\mathbb E[
\Delta U_{next}
]
\le
\Delta C_{next}
+
\Delta C_{opp,next}
+
\Delta C_{risk,next}.
} E [ Δ U n e x t ] ≤ Δ C n e x t + Δ C o pp , n e x t + Δ C r i s k , n e x t .
則考慮:
S t o p / S w i t c h / D e l e g a t e / A c t . \boxed{
Stop / Switch / Delegate / Act.
} S t o p / S w i t c h / D e l e g a t e / A c t .
66. 與內外總作用量原理的接口
既有 DIEEC 已提出:
E [ Δ U e x p a n d ] > Δ S e x p a n d \boxed{
\mathbb E[
\Delta U_{expand}
]
>
\Delta \mathcal S_{expand}
} E [ Δ U e x p an d ] > Δ S e x p an d
才值得繼續展開,
而:
E [ Δ U n e x t ] ≤ Δ S n e x t \boxed{
\mathbb E[
\Delta U_{next}
]
\le
\Delta \mathcal S_{next}
} E [ Δ U n e x t ] ≤ Δ S n e x t
作為停止候選。
TCFT-04 把此原則泛化到內部 reasoning action。
67. Cognitive Action Cost
對 reasoning action:
r t , r_t, r t ,
可定義:
L t c o g ( r t ) = C a t t + C c o m p u t e + C m e m o r y + C s w i t c h + C v e r i f y + C d e l a y + C r i s k . \boxed{
\mathcal L_t^{cog}(r_t)
=
C_{att}
+
C_{compute}
+
C_{memory}
+
C_{switch}
+
C_{verify}
+
C_{delay}
+
C_{risk}.
} L t co g ( r t ) = C a tt + C co m p u t e + C m e m or y + C s w i t c h + C v er i f y + C d e l a y + C r i s k .
68. Cognitive Action Path
整條 reasoning trajectory:
Γ R = { S 0 , r 0 , S 1 , r 1 , … , S T } . \Gamma_R
=
\{
S_0,r_0,S_1,r_1,\ldots,S_T
\}. Γ R = { S 0 , r 0 , S 1 , r 1 , … , S T } .
總成本:
S R [ Γ R ] = ∑ t = 0 T − 1 L t c o g ( r t ) + C t e r m i n a l . \boxed{
\mathcal S_R[\Gamma_R]
=
\sum_{t=0}^{T-1}
\mathcal L_t^{cog}(r_t)
+
C_{terminal}.
} S R [ Γ R ] = t = 0 ∑ T − 1 L t co g ( r t ) + C t er mina l .
69. Strategic Objective
候選:
π R ∗ = arg max π R [ E [ U t a s k ] − S R [ π R ] ] . \boxed{
\pi_R^*
=
\arg\max_{\pi_R}
\left[
\mathbb E[
U_{task}
]
-
\mathcal S_R[\pi_R]
\right].
} π R ∗ = arg π R max [ E [ U t a s k ] − S R [ π R ] ] .
subject to:
safety;
legality;
epistemic minimum;
deadline;
minimum verification。
70. CODT 接口
CODT 已有:
M e t h o d = P r o g r a m ( O p e r a t o r s , T o p o l o g y , C o n t e x t , P o l i c y , B u d g e t ) . \boxed{
Method
=
Program(
Operators,
Topology,
Context,
Policy,
Budget
).
} M e t h o d = P r o g r am ( O p er a t or s , T o p o l o g y , C o n t e x t , P o l i cy , B u d g e t ) .
TCFT-04 的 RAG 主要對應:
P o l i c y + B u d g e t + T e r m i n a t i o n + R o u t i n g . \boxed{
Policy
+
Budget
+
Termination
+
Routing.
} P o l i cy + B u d g e t + T er mina t i o n + R o u t in g .
71. RAG 不是 Domain 宣告
依 CODT:
R A G ⇏ P r o m o t e d D o m a i n . \boxed{
RAG
\not\Rightarrow
PromotedDomain.
} R A G ⇒ P r o m o t e d D o main .
目前最保守定位是:
meta-policy / scheduler / governance-layer candidate . \boxed{
\text{meta-policy / scheduler / governance-layer candidate}.
} meta-policy / scheduler / governance-layer candidate .
72. Think-Act Boundary
繼續思考與立即行動本身也是競爭選項。
如果:
V t h i n k = max r N V O C ( r ) V_{think}
=
\max_r NVOC(r) V t hink = r max N V O C ( r )
低於立即行動的淨價值,
就應:
A c t N o w . \boxed{
ActNow.
} A c tN o w .
實際形式需要 task-specific 定義。
73. Defer 與 Idle
如果目前資訊不足,但未來自然會有高價值 evidence,
則:
D e f e r \boxed{
Defer
} D e f er
可能優於:
T h i n k F o r e v e r . ThinkForever. T hink F or e v er .
若沒有必要行動或繼續推理:
I d l e \boxed{
Idle
} I d l e
也可以是合法策略。
74. Reasoning Allocation Governance
RAG 最小輸出候選:
G R ( S t ) → { R e a s o n , S e a r c h , V e r i f y , E x p a n d C F , E x p a n d R R , S w i t c h , D e l e g a t e , A c t , D e f e r , S t o p , I d l e } . \boxed{
G_R(S_t)
\rightarrow
\{
Reason,
Search,
Verify,
ExpandCF,
ExpandRR,
Switch,
Delegate,
Act,
Defer,
Stop,
Idle
\}.
} G R ( S t ) → { R e a so n , S e a r c h , V er i f y , E x p an d C F , E x p an d R R , S w i t c h , D e l e g a t e , A c t , D e f er , S t o p , I d l e } .
75. Confidence 不是停止準則
高 confidence 仍可能有:
support failure;
blindspot;
self-confirmation。
低 confidence 也不一定值得繼續,
如果下一步成本過高。
所以:
C o n f i d e n c e ≠ S t o p p i n g C r i t e r i o n . \boxed{
Confidence
\neq
StoppingCriterion.
} C o n f i d e n ce = S t o pp in g C r i t er i o n .
76. Decision Sensitivity
真正重要的常常不是 uncertainty 單獨大小,
而是:
Uncertainty × Decision Sensitivity . \boxed{
\text{Uncertainty}
\times
\text{Decision Sensitivity}.
} Uncertainty × Decision Sensitivity .
如果 uncertainty 很高但 action 幾乎不變,
繼續推理價值可能很低。
77. Expected Regret Reduction
可以候選定義:
E R R ( r ) = E x p e c t e d R e g r e t b e f o r e − E x p e c t e d R e g r e t a f t e r r . \boxed{
ERR(r)
=
ExpectedRegret_{before}
-
ExpectedRegret_{after\ r}.
} E R R ( r ) = E x p ec t e d R e g r e t b e f or e − E x p ec t e d R e g r e t a f t er r .
如果:
E R R ( r ) ERR(r) E R R ( r )
很小,
則 r r r 可能沒有策略價值。
78. Risk-Sensitive Allocation
若有 low-probability high-impact branch:
P ( c ∗ ) ≪ 1 , P(c^*)\ll1, P ( c ∗ ) ≪ 1 ,
但:
L o s s ( c ∗ ) ≫ 1 , Loss(c^*)\gg1, L oss ( c ∗ ) ≫ 1 ,
它仍可能:
N V O C ( c ∗ ) > 0. NVOC(c^*)>0. N V O C ( c ∗ ) > 0.
所以不能只用 probability pruning。
79. Myopic 與 Non-Myopic Metareasoning
只看:
N V O C ( r t ) NVOC(r_t) N V O C ( r t )
可能漏掉:
這一步本身沒直接收益,但會開啟高價值 reasoning region。
因此需要:
non-myopic metareasoning . \boxed{
\text{non-myopic metareasoning}.
} non-myopic metareasoning .
80. Gateway Computation
某 reasoning action:
r g r_g r g
直接價值:
Δ U ≈ 0 , \Delta U\approx0, Δ U ≈ 0 ,
但會解鎖:
r h i g h . r_{high}. r hi g h .
因此:
V ( r g ) = V d i r e c t + V o p t i o n . \boxed{
V(r_g)
=
V_{direct}
+
V_{option}.
} V ( r g ) = V d i r ec t + V o pt i o n .
81. Option Value of Computation
定義候選:
O V C ( r ) \boxed{
OVC(r)
} O V C ( r )
表示:
這個 reasoning action 是否打開未來高價值計算選項?
82. Amortized Cognitive Value
某些昂貴結構:
ontology;
index;
reusable model;
toolchain;
可以跨任務攤銷。
所以:
C n o w > B e n e f i t n o w C_{now}
>
Benefit_{now} C n o w > B e n e f i t n o w
不代表長期不值得。
可候選寫:
A C V ( r ) = ∑ k E x p e c t e d B e n e f i t k ( r ) − L i f e c y c l e C o s t ( r ) . \boxed{
ACV(r)
=
\sum_k ExpectedBenefit_k(r)
-
LifecycleCost(r).
} A C V ( r ) = k ∑ E x p ec t e d B e n e f i t k ( r ) − L i f ecy c l e C os t ( r ) .
83. Meta-Reasoning Regress
如果用:
M 2 M_2 M 2
決定:
M 1 , M_1, M 1 ,
再用:
M 3 M_3 M 3
決定:
M 2 , M_2, M 2 ,
形成:
M 1 ← M 2 ← M 3 ← ⋯ M_1
\leftarrow
M_2
\leftarrow
M_3
\leftarrow
\cdots M 1 ← M 2 ← M 3 ← ⋯
則產生:
Meta-Reasoning Regress . \boxed{
\text{Meta-Reasoning Regress}.
} Meta-Reasoning Regress .
meta 層也必須接受:
N V O C m e t a ≤ 0 \boxed{
NVOC^{meta}\le0
} N V O C m e t a ≤ 0
時停止。
84. Allocation Audit
對 AI runtime 特別重要的是保留:
R e a s o n i n g L e d g e r t = ( A c t i o n , T a r g e t , D o m a i n , B u d g e t , E x p e c t e d G a i n , A c t u a l G a i n , C o s t , S t o p R e a s o n ) . \boxed{
ReasoningLedger_t
=
(
Action,
Target,
Domain,
Budget,
ExpectedGain,
ActualGain,
Cost,
StopReason
).
} R e a so nin g L e d g e r t = ( A c t i o n , T a r g e t , D o main , B u d g e t , E x p ec t e d G ain , A c t u a l G ain , C os t , S t o pR e a so n ) .
如此才能回放:
為什麼當時選擇不繼續想?
85. 可反證命題
H1:Reasoning Depth 與 Strategic Utility 可分離
若在控制 task difficulty 後,reasoning depth 總是單調提高 cost-adjusted strategic outcome,TCFT-04 的核心 separation 被削弱。
H2:NVOC-Based Stopping 優於固定深度
若 fixed-depth reasoning 在 heterogeneous tasks 中不劣於 cost-aware stopping,NVOC 的工程必要性降低。
H3:Strategic Coarsening 可提高 Global Utility
若 coarse local model 永遠無法在 cost-adjusted global performance 上超越 fine model,Strategic Coarsening 應降級。
H4:Domain Switching 改善 Multi-Domain Tasks
若 explicit switch policy 沒有帶來收益,domain-level routing 可能不必要。
H5:Verification Allocation 優於 Uniform Verification
若所有 claim 同樣驗證比 risk-sensitive allocation 更好,selective verification 沒有必要。
H6:Meta-Reasoning Cost 不可忽略
若 meta-control cost 始終 negligible, B m e t a B^{meta} B m e t a 不必單獨建模。
86. 實驗設計
Experiment A:Fixed Depth vs Cost-Aware Stopping
比較:
D e p t h = k Depth=k D e pt h = k
與:
N V O C -stop . NVOC\text{-stop}. N V O C -stop .
量測:
utility;
accuracy;
latency;
compute;
regret。
Experiment B:Single-Target Overinvestment
建立:
A , B , C , E n v i r o n m e n t . A,B,C,Environment. A , B , C , E n v i r o nm e n t .
agent 可大量分析 B B B ,但 outcome 同時由 C C C 與 environment 影響。
檢查 high-resolution M ( B ) M(B) M ( B ) 是否造成 global performance 下降。
Experiment C:Strategic Coarsening
提供 coarse、medium、fine models,比較:
U t i l i t y − C o s t F r o n t i e r . \boxed{
Utility-Cost Frontier.
} U t i l i t y − C os tF r o n t i er .
Experiment D:Domain Switching
比較 no-switch、periodic、value-guided、random switching。
Experiment E:Verification Allocation
比較 uniform、risk-weighted、uncertainty-weighted 與 NVOC-weighted verification。
Experiment F:Human-AI Attention Allocation
提供大量 AI-generated branches,限制 human review budget,測不同 attention policy。
Experiment G:Meta-Reasoning Regress
允許 agent 花資源「思考如何思考」,測 meta-overhead 是否爆炸。
Experiment H:Deadline Shift
縮短 deadline,測合理策略是否自動調低不必要深度與增加 action urgency。
Experiment I:Risk Shift
提高 tail loss,測是否對 low-probability high-impact branches 增加 counterfactual / verification budget。
Experiment J:Amortized Reasoning
比較一次性 reasoning 與建立 reusable cognitive infrastructure 的跨任務成本。
87. Temporal Strategic Advancement
TCFT 的時代前沿不只測:
reasoning capability . \text{reasoning capability}. reasoning capability .
還可以測:
reasoning governance capability . \boxed{
\text{reasoning governance capability}.
} reasoning governance capability .
某個歷史人物未必比所有人「想得最深」,
但可能更早展現:
selective attention;
correct stopping;
domain switching;
coarse but sufficient models;
delegation;
uncertainty-sensitive verification。
這也可以是 temporal cognitive advancement。
88. Temporal Baseline for Strategy
因此:
B t s t r a t e g y \boxed{
B_t^{strategy}
} B t s t r a t e g y
必須包含同期可取得的:
tools;
advisers;
institutions;
communication;
computational resources。
不能只比較裸人腦。
89. AI-Augmented Baseline
進入 AI 時代後:
B t h u m a n + A I B_t^{human+AI} B t h u man + A I
應成為現代 benchmark 的一部分。
如果 AI 已經可以廉價完成某種 reasoning,
那麼一個現代人的「超前」必須扣除這個工具條件。
90. Strategy Is Task-Relative
策略優劣必須 index by:
T a s k , G o a l , R i s k , T i m e , R e s o u r c e s , C o n s t r a i n t s . \boxed{
Task,
Goal,
Risk,
Time,
Resources,
Constraints.
} T a s k , G o a l , R i s k , T im e , R eso u r ces , C o n s t r ain t s .
不存在一個無條件固定的:
globally optimal reasoning depth . \text{globally optimal reasoning depth}. globally optimal reasoning depth .
91. Dynamic Strategic Invariant
高階穩定不一定是:
π R ( t + 1 ) = π R ( t ) . \pi_R(t+1)=\pi_R(t). π R ( t + 1 ) = π R ( t ) .
而可能是:
I ( R A G t ) ≈ I ( R A G t + 1 ) \boxed{
I(RAG_t)
\approx
I(RAG_{t+1})
} I ( R A G t ) ≈ I ( R A G t + 1 )
即治理原則穩定,但資源配置動態改變。
例如:
高風險需要更多驗證;
低邊際價值停止;
persistent failure 觸發切換。
92. 與 TCFT-05 的接口
TCFT-04 最後還沒有回答:
某個人很早就懂得這種資源治理,是真的超前,還是我們把今天的語言投射回去?
這需要:
F r o z e n − T i m e P r i o r − A r t A u d i t . \boxed{
Frozen-Time Prior-Art Audit.
} F r oz e n − T im e P r i or − A r t A u d i t .
以及:
T i m e − N o r m a l i z e d N o v e l t y . \boxed{
Time-Normalized Novelty.
} T im e − N or ma l i z e d N o v e l t y .
這就是 TCFT-05。
93. 侷限
第一,NVOC 中的效用與成本通常只能估計。
第二,機會成本高度 task-dependent。
第三,人類 cognitive effort 與 AI compute cost 不可直接等量。
第四,reasoning quality 不一定隨 compute 單調增加。
第五,meta-reasoning 本身會消耗資源。
第六,過度成本最小化可能造成 premature stopping。
第七,高風險 task 可能需要非期望值式風險準則。
第八,Strategic Coarsening 若錯用可能遮蔽必要細節。
第九,domain switching 需要維持 state consistency。
第十,本文不宣稱已找到 globally optimal reasoning allocation。
94. 結論
TCFT-01 至 TCFT-03 主要在問:
How much more can an agent think? \boxed{
\text{How much more can an agent think?}
} How much more can an agent think?
TCFT-04 第一次反過來問:
How much of that capability should actually be used? \boxed{
\text{How much of that capability should actually be used?}
} How much of that capability should actually be used?
因此:
Capability ≠ Activation ≠ Allocation ≠ Termination . \boxed{
\text{Capability}
\neq
\text{Activation}
\neq
\text{Allocation}
\neq
\text{Termination}.
} Capability = Activation = Allocation = Termination .
高階推理不是高階策略。
因為策略不只要求:
Think Deeply . \text{Think Deeply}. Think Deeply .
還要求:
Think at the right depth, about the right target, in the right domain, for the right duration, at the right cost. \boxed{
\text{Think at the right depth,
about the right target,
in the right domain,
for the right duration,
at the right cost.
}
} Think at the right depth, about the right target, in the right domain, for the right duration, at the right cost.
所以:
High-Resolution Local Model ⇏ High-Quality Global Strategy . \boxed{
\text{High-Resolution Local Model}
\not\Rightarrow
\text{High-Quality Global Strategy}.
} High-Resolution Local Model ⇒ High-Quality Global Strategy .
一個 agent 對單一他者的模型可以越來越精細,
但如果它把:
B s e l f , B o t h e r s , B w o r l d , B f u t u r e B_{self},
B_{others},
B_{world},
B_{future} B se l f , B o t h er s , B w or l d , B f u t u r e
全部擠出去,
整體策略可能反而惡化。
同樣地,反事實可以繼續生成、反身性可以繼續升階、資訊可以繼續搜尋;但當:
NVOC ≤ 0 , \boxed{
\operatorname{NVOC}\le0,
} NVOC ≤ 0 ,
繼續推理不再自動構成智慧。
真正成熟的策略能力包含:
Allocate + Switch + Coarsen + Verify + Delegate + Stop . \boxed{
\text{Allocate}
+
\text{Switch}
+
\text{Coarsen}
+
\text{Verify}
+
\text{Delegate}
+
\text{Stop}.
} Allocate + Switch + Coarsen + Verify + Delegate + Stop .
因此 TCFT-04 的中心命題是:
Strategic intelligence is not merely the ability to reason deeply, but the ability to allocate, switch, coarse-grain, and stop reasoning well. \boxed{
\text{Strategic intelligence is not merely the ability to reason deeply,
but the ability to allocate, switch, coarse-grain, and stop reasoning well.}
} Strategic intelligence is not merely the ability to reason deeply, but the ability to allocate, switch, coarse-grain, and stop reasoning well.
一個存在真正高階的地方,有時不是:
他能想到第十層。
而是:
他知道第二層已經夠了。
甚至:
他知道這次根本不值得把主要注意力放在這個人、這個問題或這個域。
下一篇 TCFT-05 將處理另一個必要問題:
即使某種思維看起來超前, 我們怎麼知道它在當時真的超前? \boxed{
\text{即使某種思維看起來超前,
我們怎麼知道它在當時真的超前?}
} 即使某種思維看起來超前, 我們怎麼知道它在當時真的超前?
也就是:
Frozen-Time Prior Art + Temporal Baseline + Time-Normalized Novelty . \boxed{
\text{Frozen-Time Prior Art}
+
\text{Temporal Baseline}
+
\text{Time-Normalized Novelty}.
} Frozen-Time Prior Art + Temporal Baseline + Time-Normalized Novelty .
References
Russell, S. J., & Wefald, E. H. (1991). Principles of metareasoning. Artificial Intelligence , 49(1–3), 361–395. DOI: 10.1016/0004-3702(91)90015-C.
Russell, S. J., & Wefald, E. H. (1991). Do the Right Thing: Studies in Limited Rationality . MIT Press.
Hansen, E. A., & Zilberstein, S. (2001). Monitoring and control of anytime algorithms: A dynamic programming approach. Artificial Intelligence , 126(1–2), 139–157. DOI: 10.1016/S0004-3702(00)00068-0.
Zilberstein, S. (2011). Metareasoning and bounded rationality. In M. T. Cox & A. Raja (Eds.), Metareasoning: Thinking about Thinking . MIT Press.
Lieder, F., & Griffiths, T. L. (2020). Resource-rational analysis: Understanding human cognition as the optimal use of limited computational resources. Behavioral and Brain Sciences , 43, e1. DOI: 10.1017/S0140525X1900061X.
Sims, C. A. (2003). Implications of rational inattention. Journal of Monetary Economics , 50(3), 665–690.
Sims, C. A. (2006). Rational Inattention: Beyond the Linear-Quadratic Case. American Economic Review , 96(2), 158–163. DOI: 10.1257/000282806777212431.
Sims, C. A. (2010). Rational inattention and monetary economics. Handbook of Monetary Economics , 3, 155–181. DOI: 10.1016/B978-0-444-53238-1.00004-1.
Shenhav, A., Botvinick, M. M., & Cohen, J. D. (2013). The Expected Value of Control: An Integrative Theory of Anterior Cingulate Cortex Function. Neuron , 79(2), 217–240. DOI: 10.1016/j.neuron.2013.07.007.
Botvinick, M., & Braver, T. (2015). Motivation and Cognitive Control: From Behavior to Neural Mechanism. Annual Review of Psychology , 66, 83–113. DOI: 10.1146/annurev-psych-010814-015044.
Fulton, C. (2022). Choosing what to pay attention to. Theoretical Economics .
Neo.K. (2026). 內外總作用量原理:從 TOKEN 機率到世界展開成本 . EveMissLab.
Neo.K. (2026). Cognitive Operator-Domain Theory (CODT) Series 01–10 . EveMissLab.
Neo.K. (2026). TCFT-00|超前認知不是預言:時代認知前沿的問題設定 . EveMissLab.
Neo.K. (2026). TCFT-01|未來底空間選擇:想像力、候選世界與問題先行 . EveMissLab.
Neo.K. (2026). TCFT-02|反事實視界:未發生世界、問題生成與認知覆蓋 . EveMissLab.
Neo.K. (2026). TCFT-03|反身性推理:當推理本身進入被推理世界 . EveMissLab.
Canonical Note
本文件正式原始碼使用 UTF-8。
數學原始碼只使用 canonical delimiters:
inline math:$...$
display math:$$...$$
不進行 unicode_escape 類 round-trip;不把 LaTeX 轉為 Unicode 數學字元後再作 canonical source;不以聊天 rendering view 作為正式原稿。