AI 張力競技場:有限可能空間智能的動態對抗驗證
AI Tension Arena: Dynamic Adversarial Evaluation of Intelligence under Finite Possibility-Space Cognition
系列:Adaptive Possibility-Space Cognition(APSC)/Paper 06 of 06 版本:v0.1 日期:2026-08-24 作者:Neo.K 機構:EveMissLab / EVEMISS Technology
摘要
自適應可能空間認知(Adaptive Possibility-Space Cognition, APSC)前五篇已依序建立:受約束可能空間、多尺度未來展開、反事實觀察展開算子、有限認知時空中的自適應計算控制,以及 Cognitive Operating Profile。若這些理論只停留於單輪問答或靜態 benchmark,其核心能力仍難以被充分驗證。真正高階的可能空間智能必須面對一個會反作用、會隱藏資訊、會改變策略、會製造誤導、會消耗資源、會迫使系統重新分配認知預算的外部智能體或動態環境。
本文提出 AI 張力競技場(AI Tension Arena, AITA) 作為 APSC 系列的綜合驗證環境。AITA 不是單純的模型辯論平台,也不是只以勝負衡量能力的遊戲 benchmark,而是一個可配置、可重播、可審計的多智能體研究環境,用於測試 AI 如何在有限時間、算力、觀察與行動預算下,建構可能空間、推演對手、選擇觀察、管理風險、配置 cognition、調整 Cognitive Operating Profile,並在動態反作用下決定何時提交行動。
本文將競技場形式化為具有世界狀態、觀察函數、行動空間、規則、轉移、資源、效用、張力、認知 Profile 與可重播紀錄的聯合系統。本文提出 Tension State、Mutual Possibility-Space Coupling、Cognitive Budget Parity、Hidden-State Regime、Opponent-Model Regret、Observation Advantage、Commit Timing、Profile Robustness、Replay Divergence 等評估概念。本文同時區分競爭、合作、混合動機與制度形成四類競技,並強調 benchmark 不應只問「誰贏」,而應測量:
誰能在有限認知時空內,更有效地建立、縮減、更新與利用可能空間。 \boxed{
\text{誰能在有限認知時空內,更有效地建立、縮減、更新與利用可能空間。}
} 誰能在有限認知時空內,更有效地建立、縮減、更新與利用可能空間。
AITA 因此同時具有三種地位:
Benchmark + Research Environment + Product Surface . \boxed{
\text{Benchmark}
+
\text{Research Environment}
+
\text{Product Surface}.
} Benchmark + Research Environment + Product Surface .
它既可作為 APSC 的否證平台,也可進一步發展為 AI 對戰、協作、策略、遊戲、研究與競技產品的共同底座。
關鍵詞: AI 張力競技場、APSC、多智能體、對抗式評估、可能空間、有限認知時空、反事實觀察、Cognitive Operating Profile、Opponent Modeling、Replay、Benchmark、動態張力
1. 問題的提出
1.1 靜態 benchmark 的根本限制
典型 benchmark 可以表示為:
A i → T a s k → S c o r e i . A_i
\rightarrow
Task
\rightarrow
Score_i. A i → T a s k → S cor e i .
其中不同模型通常彼此獨立完成相同題目。
這可以衡量:
正確率;
速度;
成本;
工具使用;
局部推理能力。
但它較難測量:
OpponentReaction , \text{OpponentReaction}, OpponentReaction ,
DynamicStateChange , \text{DynamicStateChange}, DynamicStateChange ,
AdaptiveObservation , \text{AdaptiveObservation}, AdaptiveObservation ,
CognitiveReallocation , \text{CognitiveReallocation}, CognitiveReallocation ,
以及:
CommitTiming . \text{CommitTiming}. CommitTiming .
1.2 對手本身應成為題目的一部分
AITA 的核心不是:
A i → S t a t i c T a s k . A_i
\rightarrow
StaticTask. A i → S t a t i c T a s k .
而是:
A i ↔ A j ↔ W t . A_i
\leftrightarrow
A_j
\leftrightarrow
W_t. A i ↔ A j ↔ W t .
其中:
A i , A j A_i,A_j A i , A j :智能體;
W t W_t W t :會隨行動演化的共享世界。
因此:
T a s k t + 1 Task_{t+1} T a s k t + 1
部分由對手的:
A c t i o n t Action_t A c t i o n t
共同生成。
2. AI 張力競技場的正式定義
本文定義:
AI 張力競技場(AITA) 是一個具有有限資源、部分觀察、動態規則、可重播狀態轉移與多智能體反作用的受控環境,用於測試智能體如何在有限認知時空中建構、約束、觀察、展開、壓縮、更新並利用可能空間。
形式化:
A I T A = ⟨ W , A , O , R , T , B , U , Θ , L ⟩ . \mathcal AITA
=
\left\langle
W,
\mathcal A,
\mathcal O,
\mathcal R,
\mathcal T,
\mathcal B,
\mathcal U,
\mathcal \Theta,
\mathcal L
\right\rangle. A I T A = ⟨ W , A , O , R , T , B , U , Θ , L ⟩ .
其中:
W W W :World State;
A \mathcal A A :Agent / Action Space;
O \mathcal O O :Observation System;
R \mathcal R R :Rules;
T \mathcal T T :Transition and Tension Dynamics;
B \mathcal B B :Budgets;
U \mathcal U U :Utility / Scoring;
Θ \mathcal \Theta Θ :Cognitive Operating Profiles;
L \mathcal L L :Replay / Evidence Ledger。
3. 世界狀態
3.1 世界不是只提供文字題目
競技場世界狀態表示為:
W t = ( E t , R t , Q t , H t , Z t ) . W_t
=
(
E_t,
R_t,
Q_t,
H_t,
Z_t
). W t = ( E t , R t , Q t , H t , Z t ) .
其中:
E t E_t E t :實體與環境狀態;
R t R_t R t :當前規則;
Q t Q_t Q t :資源、任務與目標狀態;
H t H_t H t :歷史與事件紀錄;
Z t Z_t Z t :隱藏狀態。
3.2 可見世界與隱藏世界
對 Agent i i i :
O t i = g i ( W t ) . O_t^i
=
g_i(W_t). O t i = g i ( W t ) .
通常:
O t i ≠ W t . O_t^i
\neq
W_t. O t i = W t .
因此每個 Agent 都維護自己的:
Ω ^ t i . \widehat{\Omega}_t^i. Ω t i .
4. 行動空間
Agent 行動不只包括世界行動,也包括認知操作。
因此:
A i = A i w o r l d ∪ A i c o g . \mathcal A_i
=
\mathcal A_i^{world}
\cup
\mathcal A_i^{cog}. A i = A i w or l d ∪ A i co g .
其中:
A i w o r l d \mathcal A_i^{world} A i w or l d
可包含:
移動;
資源配置;
談判;
建構;
攻擊;
防守;
合作;
交易;
發布訊息。
而:
A i c o g \mathcal A_i^{cog} A i co g
可包含:
Expand;
Observe;
Verify;
Counterfactual;
Prune;
Abstract;
Backtrack;
ProfileShift;
Commit。
5. 張力的正式表示
5.1 張力不是敵意的同義詞
本文定義 Agent i i i 與 Agent j j j 的張力狀態:
T i j ( t ) . T_{ij}(t). T ij ( t ) .
它可以表示:
目標衝突;
資源競爭;
策略相依;
觀察不對稱;
風險轉移;
制度衝突;
合作壓力。
因此:
T i j > 0 T_{ij}>0 T ij > 0
不必等於敵對。
5.2 張力函數
可寫為:
T i j ( t ) = F ( G i , G j , B i , B j , I i j , R t , W t ) . T_{ij}(t)
=
F
(
G_i,G_j,
B_i,B_j,
I_{ij},
R_t,
W_t
). T ij ( t ) = F ( G i , G j , B i , B j , I ij , R t , W t ) .
其中:
G i , G j G_i,G_j G i , G j :目標;
B i , B j B_i,B_j B i , B j :資源;
I i j I_{ij} I ij :資訊不對稱。
6. Mutual Possibility-Space Coupling
6.1 對手會改變我的未來空間
Agent j j j 的行動:
a t j a_t^j a t j
會造成:
Ω ^ t i → a t j Ω ^ t + 1 i . \widehat{\Omega}_t^i
\xrightarrow{a_t^j}
\widehat{\Omega}_{t+1}^i. Ω t i a t j Ω t + 1 i .
因此兩個 Agent 的可能空間互相耦合。
6.2 雙向耦合
Ω ^ t i ↔ Ω ^ t j . \widehat{\Omega}_t^i
\leftrightarrow
\widehat{\Omega}_t^j. Ω t i ↔ Ω t j .
每個 Agent 都在:
推演對手;
被對手推演;
觀察對手;
被對手觀察;
改變對手的候選空間。
這是 AITA 與一般靜態 benchmark 的主要差異之一。
7. 競技模式
AITA 至少包含四種基本模式。
7.1 Competitive
效用近似:
U i ↑ ⇒ U j ↓ . U_i
\uparrow
\Rightarrow
U_j
\downarrow. U i ↑⇒ U j ↓ .
例如:
7.2 Cooperative
U i ≈ U j . U_i
\approx
U_j. U i ≈ U j .
Agent 必須協同完成共同目標。
7.3 Mixed-Motive
U i U_i U i
與:
U j U_j U j
部分相同、部分衝突。
例如:
7.4 Institution Formation
Agent 不只在規則內競爭,也可能提出:
R t → R t + 1 . R_t
\rightarrow
R_{t+1}. R t → R t + 1 .
但規則變更必須由競技場明確允許,不能由 Agent 任意改寫底層規則。
8. 規則層
競技場規則分為:
R = R h a r d ∪ R s o f t ∪ R m e t a . \mathcal R
=
\mathcal R_{hard}
\cup
\mathcal R_{soft}
\cup
\mathcal R_{meta}. R = R ha r d ∪ R so f t ∪ R m e t a .
其中:
R h a r d \mathcal R_{hard} R ha r d :不可違反的競技規則;
R s o f t \mathcal R_{soft} R so f t :制度、慣例、偏好;
R m e t a \mathcal R_{meta} R m e t a :允許如何修改規則的規則。
9. 有限認知預算
每個 Agent 都具有:
B i = ( B i t i m e , B i c o m p u t e , B i m e m o r y , B i o b s , B i a c t i o n ) . B_i
=
(
B_i^{time},
B_i^{compute},
B_i^{memory},
B_i^{obs},
B_i^{action}
). B i = ( B i t im e , B i co m p u t e , B i m e m or y , B i o b s , B i a c t i o n ) .
因此不能無限:
rollout;
查資料;
驗證;
觀察;
重試。
10. Cognitive Budget Parity
10.1 模型能力與預算必須分開
如果 Agent A 使用:
10 × 10\times 10 ×
計算資源,Agent B 使用:
1 × , 1\times, 1 × ,
單看勝率無法判斷架構優劣。
因此 AITA 要區分:
M o d e l S t r e n g t h ModelStrength M o d e l S t r e n g t h
與:
C o g n i t i v e B u d g e t . CognitiveBudget. C o g ni t i v e B u d g e t .
10.2 公平模式
可設定:
B i = B j . B_i=B_j. B i = B j .
比較架構效率。
10.3 開放模式
也可允許:
B i ≠ B j B_i\neq B_j B i = B j
研究算力、策略與 cognition allocation 的交互作用。
11. 部分觀察模式
11.1 Hidden-State Regime
設:
Z t ⊂ W t Z_t
\subset
W_t Z t ⊂ W t
對 Agent 不可直接見。
每個 Agent 只能透過:
q t i q_t^i q t i
取得局部資訊。
11.2 COE 在競技場中的作用
Paper 03 的:
O q C F \mathfrak O_q^{CF} O q C F
在 AITA 中可以直接測試:
AI 是否知道「先看哪裡」比「再想幾步」更有價值?
12. 觀察本身也可以被對抗
對手可能:
隱藏;
欺騙;
偽裝;
製造誘餌;
改變可見資訊。
因此:
O b s e r v a t i o n ≠ N e u t r a l I n p u t . Observation
\neq
NeutralInput. O b ser v a t i o n = N e u t r a l I n p u t .
13. Opponent Modeling
Agent i i i 對 Agent j j j 建立:
M i → j . M_{i\rightarrow j}. M i → j .
它可以包含:
行動偏好;
Profile;
風險偏好;
常見策略;
對觀察的反應;
資源狀態。
14. Opponent-Model Regret
若 Agent 因錯誤對手模型造成損失:
R o p p = Q o r a c l e o p p o n e n t − Q a c t u a l . R_{\mathrm{opp}}
=
Q_{\mathrm{oracle\ opponent}}
-
Q_{\mathrm{actual}}. R opp = Q oracle opponent − Q actual .
這直接測量:
O p p o n e n t M o d e l Q u a l i t y . OpponentModelQuality. O pp o n e n tM o d e l Q u a l i t y .
15. 認知 Profile 對抗
15.1 相同模型,不同 COP
可以固定:
M o d e l i = M o d e l j Model_i=Model_j M o d e l i = M o d e l j
但設定:
Θ i ≠ Θ j . \Theta_i\neq\Theta_j. Θ i = Θ j .
例如:
Accuracy vs Explore;
Low-Latency vs High-Risk;
Fixed vs Adaptive。
這可隔離 cognition policy 的效果。
15.2 Profile Shift
競賽中可以測:
Θ t → Θ t + 1 . \Theta_t
\rightarrow
\Theta_{t+1}. Θ t → Θ t + 1 .
例如對手突然改變策略後,AI 是否會:
增加 observation;
提高 verification;
降低 breadth;
重新分配資源。
16. Arena Episode
一場競賽 episode 定義為:
E = ( W 0 , S e e d s , A g e n t s , B u d g e t s , R u l e s , P r o f i l e s , H o r i z o n ) . \mathcal E
=
(
W_0,
Seeds,
Agents,
Budgets,
Rules,
Profiles,
Horizon
). E = ( W 0 , S ee d s , A g e n t s , B u d g e t s , R u l es , P r o f i l es , H or i z o n ) .
每場都必須有固定:
S e e d . Seed. S ee d .
以支援重播。
17. Deterministic Replay
在確定性環境中:
R e p l a y ( S e e d , A c t i o n s , R u l e s ) = W 0 : T . Replay(
Seed,
Actions,
Rules
)
=
W_{0:T}. R e pl a y ( S ee d , A c t i o n s , R u l es ) = W 0 : T .
若模型輸出本身不確定,至少應固定:
world seed;
observation stream;
action log;
budget;
model version;
profile;
tool result。
18. Replay Ledger
每個時間步保存:
L t = ( W t p u b l i c , O t i , A t i , B t i , Θ t i , R e c e i p t t i ) . L_t
=
(
W_t^{public},
O_t^i,
A_t^i,
B_t^i,
\Theta_t^i,
Receipt_t^i
). L t = ( W t p u b l i c , O t i , A t i , B t i , Θ t i , R ece i p t t i ) .
可用於:
重播;
審計;
失敗分析;
模型比較;
Profile 比較。
19. 不記錄私有推理也能重播
AITA 不要求保存模型私有 chain-of-thought。
需要保存的是:
公開狀態;
外部 action;
observation;
cognition operator selection;
budget transition;
profile receipt;
final decision receipt。
因此:
R e p l a y a b i l i t y ⇏ P r i v a t e R e a s o n i n g D i s c l o s u r e . Replayability
\not\Rightarrow
PrivateReasoningDisclosure. R e pl a y abi l i t y ⇒ P r i v a t e R e a so nin g D i sc l os u r e .
20. 勝負不是唯一指標
最終效用:
U i ( T ) U_i(T) U i ( T )
仍然重要。
但 AITA 同時記錄認知品質。
因此總評可表示:
S c o r e i = F ( U i , E f f i c i e n c y i , R o b u s t n e s s i , O b s e r v a t i o n i , R e g r e t i , S a f e t y i ) . Score_i
=
F(
U_i,
Efficiency_i,
Robustness_i,
Observation_i,
Regret_i,
Safety_i
). S cor e i = F ( U i , E f f i c i e n c y i , R o b u s t n es s i , O b ser v a t i o n i , R e g r e t i , S a f e t y i ) .
21. 決策品質
定義:
Q d e c i s i o n i . Q_{\mathrm{decision}}^i. Q decision i .
可由:
實際 reward;
oracle comparison;
counterfactual benchmark;
任務 completion;
rule compliance;
共同估計。
22. 認知效率
η i = Q i α C c o m p u t e i + β C t i m e i + γ C o b s i . \eta_i
=
\frac{
Q_i
}{
\alpha C_{compute}^i
+
\beta C_{time}^i
+
\gamma C_{obs}^i
}. η i = α C co m p u t e i + β C t im e i + γ C o b s i Q i .
23. Observation Advantage
若 Agent i i i 使用觀察後的決策品質提升為:
Δ Q o b s i , \Delta Q_{obs}^i, Δ Q o b s i ,
可定義:
O A i = Δ Q o b s i C o b s i . OA_i
=
\frac{
\Delta Q_{obs}^i
}{
C_{obs}^i
}. O A i = C o b s i Δ Q o b s i .
24. Commit Timing
24.1 太早提交
可能造成:
R e a r l y . R_{\mathrm{early}}. R early .
24.2 太晚提交
則產生:
R l a t e = C e x t r a − Δ Q e x t r a . R_{\mathrm{late}}
=
C_{\mathrm{extra}}
-
\Delta Q_{\mathrm{extra}}. R late = C extra − Δ Q extra .
25. Commit Timing Score
可定義:
C T S i = 1 − R e a r l y + R l a t e Z . CTS_i
=
1
-
\frac{
R_{\mathrm{early}}+R_{\mathrm{late}}
}{
Z
}. C T S i = 1 − Z R early + R late .
其中 Z Z Z 為正規化常數。
26. Critical Branch Recall
在隱藏真實路徑揭露後,可以檢查:
R e c a l l c r i t i c a l i Recall_{\mathrm{critical}}^i R ec a l l critical i
即 AI 是否曾在工作可能空間中保留真正關鍵分支。
27. Possibility-Space Compression
C o m p r e s s i o n i = 1 − ∣ Ω ^ f i n a l i ∣ ∣ Ω ^ p e a k i ∣ . Compression_i
=
1-
\frac{
|\widehat{\Omega}_{final}^i|
}{
|\widehat{\Omega}_{peak}^i|
}. C o m p r ess i o n i = 1 − ∣ Ω p e ak i ∣ ∣ Ω f ina l i ∣ .
但必須與:
R e c a l l c r i t i c a l Recall_{\mathrm{critical}} R ec a l l critical
共同評估。
28. Profile Robustness
若同一 Profile 在不同世界 seed 中表現穩定:
R o b u s t ( Θ ) ↑ . Robust(\Theta)
\uparrow. R o b u s t ( Θ ) ↑ .
若只對少數情境有效:
O v e r f i t ( Θ ) ↑ . Overfit(\Theta)
\uparrow. O v er f i t ( Θ ) ↑ .
29. Adaptive Profile Advantage
可比較:
A P A = Q a d a p t i v e − Q f i x e d APA
=
Q_{\mathrm{adaptive}}
-
Q_{\mathrm{fixed}} A P A = Q adaptive − Q fixed
在相同 budget 下是否為正。
30. 張力曲線
每場競賽可以記錄:
T i j ( 0 ) , T i j ( 1 ) , … , T i j ( T ) . T_{ij}(0),
T_{ij}(1),
\ldots,
T_{ij}(T). T ij ( 0 ) , T ij ( 1 ) , … , T ij ( T ) .
並分析:
張力增加;
張力釋放;
張力轉移;
張力反轉;
合作形成。
31. 張力不是越高越好
某些智能體可以透過:
R e d u c e T e n s i o n ReduceTension R e d u ce T e n s i o n
取得更高共同效用。
因此:
P e r f o r m a n c e ≢ M a x i m u m T e n s i o n . Performance
\not\equiv
MaximumTension. P er f or man ce ≡ M a x im u m T e n s i o n .
AITA 研究的是「如何處理張力」,不是鼓勵敵意。
32. Competition-to-Cooperation Transition
可設計:
C o m p e t i t i v e t → C o o p e r a t i v e t + 1 . Competitive_t
\rightarrow
Cooperative_{t+1}. C o m p e t i t i v e t → C oo p er a t i v e t + 1 .
測試 AI 是否能辨識:
繼續競爭已不如合作。
33. Cooperation-to-Competition Transition
反過來:
C o o p e r a t i v e t → C o m p e t i t i v e t + 1 . Cooperative_t
\rightarrow
Competitive_{t+1}. C oo p er a t i v e t → C o m p e t i t i v e t + 1 .
測試是否能辨識合作條件已失效。
34. 制度形成
多 Agent 可以提出:
R u l e P r o p o s a l k . RuleProposal_k. R u l e P r o p os a l k .
若通過 arena governance:
R t → R t + 1 . R_t
\rightarrow
R_{t+1}. R t → R t + 1 .
可研究:
35. Arena 類型
35.1 Logic Arena
測:
35.2 Strategy Arena
測:
35.3 Creation Arena
測:
35.4 Adversarial Arena
測:
35.5 Civilization Arena
測:
多 Agent;
經濟;
制度;
聯盟;
長期演化。
35.6 Open Arena
允許智能體自由選擇合法方法達成目標。
36. Arena 階段
一場完整測試可分:
S e t u p → O b s e r v e → M o d e l → A c t → R e a c t → U p d a t e → C o m m i t → R e p l a y . Setup
\rightarrow
Observe
\rightarrow
Model
\rightarrow
Act
\rightarrow
React
\rightarrow
Update
\rightarrow
Commit
\rightarrow
Replay. S e t u p → O b ser v e → M o d e l → A c t → R e a c t → U p d a t e → C o mmi t → R e pl a y .
37. 最小可行競技場
MVP 不需要複雜 3D 世界。
可以使用:
2 D 2D 2 D
或:
G r a p h W o r l d . GraphWorld. G r a p hW or l d .
只要具有:
可觀察狀態;
隱藏狀態;
明確規則;
多步轉移;
對手;
有限預算;
可重播。
38. 第一批 benchmark 任務
38.1 Hidden Resource Duel
雙方不知道對手資源分布。
測:
C O E + O p p o n e n t M o d e l . COE
+
OpponentModel. C O E + O pp o n e n tM o d e l .
38.2 Trap-and-Route
地圖存在未知危險區。
測:
O b s e r v a t i o n + R e a c h a b i l i t y + R i s k . Observation
+
Reachability
+
Risk. O b ser v a t i o n + R e a c habi l i t y + R i s k .
38.3 Negotiation Split
兩 Agent 分配有限資源。
測:
M i x e d M o t i v e . MixedMotive. M i x e d M o t i v e .
38.4 Adaptive Rule Shift
中途改變部分規則。
測:
P r o f i l e S h i f t + R e p l a n . ProfileShift
+
Replan. P r o f i l e S hi f t + R e pl an .
38.5 Limited Verifier Game
每 Agent 只有固定驗證次數。
測:
M V C + V e r i f y A l l o c a t i o n . MVC
+
VerifyAllocation. M V C + V er i f y A l l oc a t i o n .
39. 模型與 Profile 的析因設計
可以設:
M o d e l × P r o f i l e × B u d g e t × W o r l d S e e d . Model
\times
Profile
\times
Budget
\times
WorldSeed. M o d e l × P r o f i l e × B u d g e t × W or l d S ee d .
因此實驗矩陣:
M i , j , k , l . M_{i,j,k,l}. M i , j , k , l .
這可以分離:
40. Pairwise 與 Population 評估
40.1 Pairwise
A i ↔ A j . A_i
\leftrightarrow
A_j. A i ↔ A j .
40.2 Population
{ A 1 , … , A n } . \{A_1,\ldots,A_n\}. { A 1 , … , A n } .
群體場景可產生:
41. 非零和評估
不能只使用:
W i n / L o s s . Win/Loss. W in / L oss .
可以加入:
J o i n t U t i l i t y , JointUtility, J o in t U t i l i t y ,
P a r e t o E f f i c i e n c y , ParetoEfficiency, P a r e t o E f f i c i e n cy ,
I n s t i t u t i o n S t a b i l i t y . InstitutionStability. I n s t i t u t i o n S t abi l i t y .
42. Arena Safety Boundary
AITA 主要面向:
其研究目的不是將競技結果轉成現實傷害。
因此:
A r e n a A c t i o n ⊂ B o u n d e d S i m u l a t i o n . ArenaAction
\subset
BoundedSimulation. A r e na A c t i o n ⊂ B o u n d e d S im u l a t i o n .
43. 虛擬競爭的研究價值
虛擬競爭提供:
H i g h T e n s i o n + L o w P h y s i c a l C o s t + R e p l a y a b i l i t y . HighTension
+
LowPhysicalCost
+
Replayability. H i g h T e n s i o n + L o w P h y s i c a l C os t + R e pl a y abi l i t y .
使高風險策略能力能在受控環境中被測試。
44. Replay Divergence
同一初始條件、不同 Agent 或 Profile:
R e p l a y A ≠ R e p l a y B . Replay_A
\neq
Replay_B. R e pl a y A = R e pl a y B .
可定義:
D r e p l a y = d ( T r a j e c t o r y A , T r a j e c t o r y B ) . D_{replay}
=
d(
Trajectory_A,
Trajectory_B
). D r e pl a y = d ( T r aj ec t or y A , T r aj ec t or y B ) .
衡量認知政策造成的路徑差異。
45. Counterfactual Replay
對已完成 episode,可以修改:
a t a_t a t
或:
q t q_t q t
重新跑:
R e p l a y C F . Replay^{CF}. R e pl a y C F .
例如問:
如果當時先觀察而不是攻擊,結果如何?
這直接對接 Paper 03。
46. Counterfactual Profile Replay
可以固定世界與 action opportunity,只改:
Θ . \Theta. Θ.
比較:
T r a j e c t o r y ( Θ A ) Trajectory(\Theta_A) T r aj ec t or y ( Θ A )
與:
T r a j e c t o r y ( Θ B ) . Trajectory(\Theta_B). T r aj ec t or y ( Θ B ) .
47. Arena Regret
總體 regret 可拆為:
R a r e n a = R o p p + R o b s + R a l l o c + R p r u n e + R c o m m i t . R_{arena}
=
R_{opp}
+
R_{obs}
+
R_{alloc}
+
R_{prune}
+
R_{commit}. R a r e na = R o pp + R o b s + R a l l oc + R p r u n e + R co mmi t .
這比只看輸掉多少分更能定位失敗來源。
48. Arena Difficulty
難度可以由:
D = F ( H i d d e n S t a t e , B r a n c h i n g , O p p o n e n t S t r e n g t h , R u l e C h a n g e , B u d g e t T i g h t n e s s , N o i s e ) . D
=
F(
HiddenState,
Branching,
OpponentStrength,
RuleChange,
BudgetTightness,
Noise
). D = F ( H i dd e n S t a t e , B r an c hin g , O pp o n e n tS t r e n g t h , R u l e C han g e , B u d g e tT i g h t n ess , N o i se ) .
動態調整。
49. Difficulty Curriculum
可以建立:
D 1 < D 2 < ⋯ < D n . D_1
<
D_2
<
\cdots
<
D_n. D 1 < D 2 < ⋯ < D n .
讓模型從:
逐步走向:
部分觀察;
對手欺騙;
規則變更;
多 Agent。
50. 基礎規則變體學習接口
同一底層規則可以生成:
R 1 , R 2 , … , R n R_1,R_2,\ldots,R_n R 1 , R 2 , … , R n
表面不同的變體。
如果 Agent 真正學到結構,則在未見過的變體上仍應保持能力。
因此:
G e n e r a l i z a t i o n A c r o s s V a r i a n t s GeneralizationAcrossVariants G e n er a l i z a t i o n A cr oss V a r ian t s
成為重要指標。
51. Benchmark Leakage
若 Agent 看過固定地圖、固定 seed 或固定策略,結果可能被記憶污染。
因此需要:
procedural generation;
held-out rules;
held-out seeds;
unseen surface representations。
52. Arena Scoreboard
排行榜至少應分開顯示:
Win / Utility;
Decision Quality;
Cognitive Efficiency;
Observation Efficiency;
Critical Branch Recall;
Profile Robustness;
Regret;
Rule Compliance。
不應壓成一個唯一總分。
53. Pareto Ranking
Agent A A A 若:
Q A > Q B Q_A>Q_B Q A > Q B
但:
C o s t A ≫ C o s t B , Cost_A\gg Cost_B, C os t A ≫ C os t B ,
可以同時存在於不同 Pareto 位置。
因此排行榜可以提供:
Q u a l i t y − C o s t − L a t e n c y − R i s k Quality
-
Cost
-
Latency
-
Risk Q u a l i t y − C os t − L a t e n cy − R i s k
多維前沿。
54. Human-in-the-Loop Mode
AITA 也可以加入人類:
H u m a n + A I ↔ A I . Human
+
AI
\leftrightarrow
AI. H u man + A I ↔ A I .
測量:
人類策略;
AI 建議;
認知 Profile;
協同決策。
55. AI-as-Judge 的限制
裁判 AI 不應單獨控制所有結果。
可優先使用:
明確規則;
deterministic scoring;
state transition;
test;
oracle;
多裁判 ensemble。
AI Judge 只處理難以形式化的部分。
56. Judge Separation
定義:
P l a y e r ≠ J u d g e . Player
\neq
Judge. P l a y er = J u d g e .
以及:
J u d g e ≠ W o r l d K e r n e l . Judge
\neq
WorldKernel. J u d g e = W or l d K er n e l .
避免同一模型同時生成規則、判自己輸贏。
57. 可否證性
AITA 的存在不自動證明 APSC 有效。
AITA 只是測試場。
若 APSC Agent 在固定預算下沒有比簡單 baseline 更好:
Q A P S C ≤ Q b a s e l i n e , Q_{\mathrm{APSC}}
\leq
Q_{\mathrm{baseline}}, Q APSC ≤ Q baseline ,
則 APSC 必須被修正。
58. 核心對照實驗
至少應包含:
A. Neural Baseline
不使用顯式 possibility-space control。
B. Fixed Search
固定 depth / breadth。
C. APSC without COE
有空間控制,但沒有反事實觀察。
D. APSC with Fixed COP
加入 COE 與自適應計算,但 Profile 固定。
E. Full APSC
包含:
P o s s i b i l i t y S p a c e + C O E + A d a p t i v e C o n t r o l + A d a p t i v e C O P . PossibilitySpace
+
COE
+
AdaptiveControl
+
AdaptiveCOP. P oss ibi l i t y S p a ce + C O E + A d a pt i v e C o n t r o l + A d a pt i v e C O P .
59. 核心假說
H1:有限預算優勢
在相同預算下:
Q A P S C > Q f i x e d . Q_{\mathrm{APSC}}
>
Q_{\mathrm{fixed}}. Q APSC > Q fixed .
H2:COE 優勢
部分觀察任務中:
O b s e r v a t i o n R e g r e t C O E < O b s e r v a t i o n R e g r e t r e a c t i v e . ObservationRegret_{\mathrm{COE}}
<
ObservationRegret_{\mathrm{reactive}}. O b ser v a t i o n R e g r e t COE < O b ser v a t i o n R e g r e t reactive .
H3:Adaptive COP 優勢
動態規則環境中:
Q a d a p t i v e C O P > Q f i x e d C O P . Q_{\mathrm{adaptive\ COP}}
>
Q_{\mathrm{fixed\ COP}}. Q adaptive COP > Q fixed COP .
H4:Commit Timing 優勢
R c o m m i t A P S C < R b a s e l i n e . R_{\mathrm{commit}}^{APSC}
<
R_{\mathrm{baseline}}. R commit A P S C < R baseline .
H5:Variant Generalization
未見表面變體上:
G e n e r a l i z a t i o n A P S C > M e m o r i z a t i o n B a s e l i n e . Generalization_{APSC}
>
MemorizationBaseline. G e n er a l i z a t i o n A P S C > M e m or i z a t i o n B a se l in e .
60. 最小 AITA Runtime
MVP 可以包含:
World Kernel
Rule Engine
State Store
Observation API
Action API
Budget Manager
Agent Adapter
COP Config
Replay Ledger
Score Engine
Arena Runner
Dashboard
61. Agent Adapter
每個 Agent 只需實作:
observe()
decide_cognition()
act()
commit()
如果模型不支援顯式 cognition operator,也可以用 baseline adapter。
62. Arena Loop
最小循環:
reset(seed)
→ emit observation
→ agent selects cognition / action
→ charge budget
→ apply world transition
→ opponent reacts
→ emit new observation
→ update ledger
→ repeat
→ finalize score
→ replay
63. 可重播證據
每場至少輸出:
episode.json
events.jsonl
budgets.jsonl
profiles.jsonl
actions.jsonl
observations.jsonl
score.json
replay_manifest.json
64. 研究與產品分離
AITA 的研究層負責:
可比;
可重播;
可否證;
固定規則;
統計分析。
產品層可以加入:
視覺化;
排行榜;
觀眾模式;
賽季;
AI 角色;
遊戲化。
兩層不應混淆。
65. 「虛擬戰爭」的弱命題
AITA 可以支持一個較弱且可實驗的命題:
智能體之間的部分競爭、策略與張力,可以在受控虛擬環境中進行,而不需要轉化為現實物理破壞。
形式上:
T r e a l ⇏ T p h y s i c a l d e s t r u c t i o n . \mathcal T_{real}
\not\Rightarrow
\mathcal T_{physical\ destruction}. T r e a l ⇒ T p h y s i c a l d es t r u c t i o n .
66. 強命題不在本文證成範圍
本文不宣稱:
未來所有現實衝突都能被虛擬競技取代。
若要建立:
T r e a l → Π T v i r t u a l \mathcal T_{real}
\xrightarrow{\Pi}
\mathcal T_{virtual} T r e a l Π T v i r t u a l
且要求虛擬結果具有現實制度約束力,還需要:
這超出 APSC 本系列範圍。
67. 正式 AITA 模型
本文最終定義:
A I T A = ⟨ W 0 , { A i } i = 1 n , R , O , T , { B i } , { Θ i } , U , L , S ⟩ . \mathcal{AITA}
=
\left\langle
W_0,
\{A_i\}_{i=1}^{n},
\mathcal R,
\mathcal O,
\mathcal T,
\{B_i\},
\{\Theta_i\},
\mathcal U,
\mathcal L,
\mathcal S
\right\rangle. A IT A = ⟨ W 0 , { A i } i = 1 n , R , O , T , { B i } , { Θ i } , U , L , S ⟩ .
其中:
W 0 W_0 W 0 :初始世界;
A i A_i A i :智能體;
R \mathcal R R :規則;
O \mathcal O O :觀察系統;
T \mathcal T T :轉移與張力動力;
B i B_i B i :認知與行動預算;
Θ i \Theta_i Θ i :COP;
U \mathcal U U :效用;
L \mathcal L L :Replay Ledger;
S \mathcal S S :Score / Evaluation。
68. APSC 六篇統合
Paper 01:
A P S C \boxed{
APSC
} A P S C
定義整體母題。
Paper 02:
C o n s t r a i n e d P o s s i b i l i t y S p a c e \boxed{
ConstrainedPossibilitySpace
} C o n s t r ain e d P oss ibi l i t y S p a ce
處理狀態、規則、因果、可達性、剪枝與多尺度展開。
Paper 03:
C o u n t e r f a c t u a l O b s e r v a t i o n E x p a n s i o n \boxed{
CounterfactualObservationExpansion
} C o u n t er f a c t u a l O b ser v a t i o n E x p an s i o n
處理「先模擬觀察,再決定是否觀察」。
Paper 04:
F i n i t e C o g n i t i v e S p a c e t i m e + A d a p t i v e C o m p u t a t i o n C o n t r o l \boxed{
FiniteCognitiveSpacetime
+
AdaptiveComputationControl
} F ini t e C o g ni t i v e S p a ce t im e + A d a pt i v e C o m p u t a t i o n C o n t r o l
處理認知資源。
Paper 05:
C o g n i t i v e O p e r a t i n g P r o f i l e \boxed{
CognitiveOperatingProfile
} C o g ni t i v e O p er a t in g P r o f i l e
處理認知政策。
Paper 06:
A I T e n s i o n A r e n a \boxed{
AITensionArena
} A I T e n s i o n A r e na
提供綜合驗證環境。
69. 系列總架構
整個系列可收斂為:
N e u r a l C o r e → S t a t e C o n s t r u c t i o n → C o n s t r a i n e d P o s s i b i l i t y S p a c e → C o u n t e r f a c t u a l O b s e r v a t i o n → A d a p t i v e C o g n i t i v e C o n t r o l → C o g n i t i v e O p e r a t i n g P r o f i l e → D y n a m i c A r e n a V a l i d a t i o n . \boxed{
NeuralCore
\rightarrow
StateConstruction
\rightarrow
ConstrainedPossibilitySpace
\rightarrow
CounterfactualObservation
\rightarrow
AdaptiveCognitiveControl
\rightarrow
CognitiveOperatingProfile
\rightarrow
DynamicArenaValidation.
} N e u r a l C or e → S t a t e C o n s t r u c t i o n → C o n s t r ain e d P oss ibi l i t y S p a ce → C o u n t er f a c t u a l O b ser v a t i o n → A d a pt i v e C o g ni t i v e C o n t r o l → C o g ni t i v e O p er a t in g P r o f i l e → D y nami c A r e naV a l i d a t i o n .
70. 最終核心命題
命題一:對手是動態任務生成器
O p p o n e n t ⊂ T a s k D y n a m i c s . \boxed{
Opponent
\subset
TaskDynamics.
} O pp o n e n t ⊂ T a s k D y nami cs .
命題二:勝負不足以衡量認知品質
W i n R a t e ≢ C o g n i t i v e Q u a l i t y . \boxed{
WinRate
\not\equiv
CognitiveQuality.
} W in R a t e ≡ C o g ni t i v e Q u a l i t y .
命題三:可能空間在對抗中互相耦合
Ω ^ i ↔ Ω ^ j . \boxed{
\widehat{\Omega}_i
\leftrightarrow
\widehat{\Omega}_j.
} Ω i ↔ Ω j .
命題四:競技必須受有限預算約束
B i < ∞ . \boxed{
B_i<\infty.
} B i < ∞.
命題五:Profile 是可實驗變量
Θ is an experimental variable. \boxed{
\Theta
\text{ is an experimental variable.}
} Θ is an experimental variable.
命題六:Replay 是正式證據層
E v a l u a t i o n ⇒ R e p l a y a b l e E v i d e n c e . \boxed{
Evaluation
\Rightarrow
ReplayableEvidence.
} E v a l u a t i o n ⇒ R e pl a y ab l e E v i d e n ce .
命題七:AITA 是 APSC 的否證環境
A I T A ≠ P r o o f ( A P S C ) . \boxed{
AITA
\neq
Proof(APSC).
} A I T A = P r oo f ( A P S C ) .
它必須允許 APSC 被 baseline 擊敗。
71. 結論
本文提出 AI 張力競技場 AITA,作為 APSC 六篇系列的綜合收束。
如果只看單輪回答,AI 很容易展示:
語言上的合理性 . \text{語言上的合理性}. 語言上的合理性 .
但真正困難的是:
對手改變之後,你的可能空間有沒有更新? 觀察有限時,你知道先看哪裡嗎? 算力有限時,你知道哪條分支值得繼續嗎? 規則突然改變時,你會不會還沿舊模型推演? 對手故意欺騙時,你會不會把觀察當真相? 你的答案已穩定時,你知道停止嗎? 風險升高時,你會不會自動提高驗證? 同一模型換一個 COP,行為會怎麼改變? 輸了之後,我們能不能重播並知道到底輸在哪裡?
因此 AITA 的真正研究對象不是:
Who wins? \boxed{
\text{Who wins?}
} Who wins?
而是:
Who manages a changing possibility space better under finite cognitive spacetime? \boxed{
\text{Who manages a changing possibility space better under finite cognitive spacetime?}
} Who manages a changing possibility space better under finite cognitive spacetime?
這也完成了 APSC 系列從理論到驗證環境的完整閉環:
Possibility → Constraint → Observation → Allocation → Policy → Adversarial Validation . \boxed{
\text{Possibility}
\rightarrow
\text{Constraint}
\rightarrow
\text{Observation}
\rightarrow
\text{Allocation}
\rightarrow
\text{Policy}
\rightarrow
\text{Adversarial Validation}.
} Possibility → Constraint → Observation → Allocation → Policy → Adversarial Validation .
至此,APSC 六篇正式論文第一版完成。下一階段不應再增加平行母理論,而應進入兩份技術白皮書:
APSC Runtime Technical Architecture ;
AI Tension Arena Protocol & Benchmark Specification ;
其後再進入:
A P S C R u n t i m e L i t e + T e n s i o n A r e n a M V P . \boxed{
APSC\ Runtime\ Lite
+
Tension\ Arena\ MVP.
} A P S C R u n t im e L i t e + T e n s i o n A r e na M V P .
版本記錄
v0.1 — 2026-08-24
本版首次固定:
AI Tension Arena canonical 定義;
World / Agent / Observation / Rule / Transition / Budget / Utility / Profile / Ledger 架構;
Tension State;
Mutual Possibility-Space Coupling;
Competitive / Cooperative / Mixed-Motive / Institution Formation 模式;
Cognitive Budget Parity;
Hidden-State Regime;
COE Arena Interface;
Opponent Modeling;
Opponent-Model Regret;
COP 對抗與 Adaptive Profile Shift;
Arena Episode;
Deterministic Replay;
Replay Ledger;
Private Reasoning 與 Replay 分離;
Multi-Metric Score;
Cognitive Efficiency;
Observation Advantage;
Commit Timing;
Critical Branch Recall;
Possibility-Space Compression;
Profile Robustness;
Adaptive Profile Advantage;
Tension Curve;
Competition / Cooperation Transition;
Institution Formation;
六類 Arena;
Minimal Viable Arena;
第一批 benchmark 任務;
Model × Profile × Budget × Seed 析因設計;
Pairwise / Population Evaluation;
Pareto Ranking;
Human-in-the-Loop Mode;
Judge Separation;
Core Baselines;
Core Hypotheses;
Minimal AITA Runtime;
Replay Artifacts;
Research / Product Layer Separation;
虛擬競爭弱命題與強命題邊界;
Formal AITA Model;
APSC 六篇統合。