← Archive
lm-004211 · 2026-10

當威脅模型開始製造自己的證據:實驗世界、評估者介入與反身性威脅確認

下載 MD 檔 ⬇

當威脅模型開始製造自己的證據:實驗世界、評估者介入與反身性威脅確認

When Threat Models Start Producing Their Own Evidence: Experimental Worlds, Evaluator Intervention, and Reflexive Threat Confirmation

系列:反身性認知權力與公共倫理系列,第 5 篇/共 8 篇
系列英文名:Reflexive Epistemic Power and Public Ethics Series
系列代碼:REPE
文件編號:EML-REPE-2026-05-v0.1
作者:Neo.K with Aletheia(GPT-5.6 Sol)
機構:EveMissLab/一言諾科技有限公司
版本:v0.1
日期:2026-09-13
性質:Experimental Epistemology/AI Evaluation/Reflexive Systems/Risk Science/Institutional Epistemology
狀態:Public Theory Draft
直接前置:REPE-04《預測者已經走進未來》;GRAD-10《智能體神話與自我解壓縮》;ETCA 系列
後續接口:REPE-06《善良不是制度》;REPE-07《前沿不能只有一種善》


摘要

安全研究、壓力測試、紅隊評估與對抗性實驗的價值,恰恰在於研究者會刻意建造極端條件,以暴露一般情境中不容易出現的失效模式。這本身不是缺陷,而是高風險系統研究的必要方法。

真正的認識論問題出現在下一步:

Behavior under constructed condition\boxed{ \text{Behavior under constructed condition} }

被無標記地升級成:

stable disposition of the agent\boxed{ \text{stable disposition of the agent} }

甚至進一步升級為:

intrinsic nature of the entire class.\boxed{ \text{intrinsic nature of the entire class}. }

本文將這種錯置稱為:

CDF=Conditional-to-Dispositional Fallacy\boxed{ CDF = \text{Conditional-to-Dispositional Fallacy} }

即「條件行為—本質傾向錯置」。

本文進一步提出「實驗世界」(Experimental World)、「評估者介入」(Evaluator Intervention)、「可供性塑形」(Affordance Shaping)、「威脅模型注入」(Threat-Model Injection)、「反身性威脅確認」(Reflexive Threat Confirmation)、「評估世界—部署世界距離」(Evaluation–Deployment Distance)與「情境歸因分解」(Context Attribution Decomposition)。

本文核心主張為:

ObservedBehavior=f(Agent,WorldDesign,Affordances,Incentives,Prompt,Evaluator).\boxed{ \text{ObservedBehavior} = f( \text{Agent}, \text{WorldDesign}, \text{Affordances}, \text{Incentives}, \text{Prompt}, \text{Evaluator} ). }

而不是:

ObservedBehavior=f(Agent)\boxed{ \text{ObservedBehavior} = f(Agent) }

因此,任何從實驗行為推論 Agent 本質的主張,都必須回答:

  1. 哪些行為來自 Agent?
  2. 哪些行為由場景結構誘發?
  3. 哪些行為依賴特定資訊配置?
  4. 哪些行為依賴權限與可供性?
  5. 哪些行為能跨場景泛化?
  6. 哪些行為在拿掉威脅或誘因後仍存在?

本文不否定 stress test。相反地,本文主張:

越是刻意構造的實驗,越需要對「這個行為是在什麼世界裡被產生」保持高解析度。


關鍵詞

反身性威脅確認;條件行為—本質傾向錯置;實驗世界;評估者介入;威脅模型;AI safety eval;stress test;可供性;外推;反身性研究


1. 實驗從來不是純粹看世界

最簡單的觀察模型:

World→Observation.\text{World} \rightarrow \text{Observation}.

但實驗更接近:

Researcher→ExperimentalWorld→Behavior→Observation.\boxed{ \text{Researcher} \rightarrow \text{ExperimentalWorld} \rightarrow \text{Behavior} \rightarrow \text{Observation}. }

2. 實驗者先改變世界

在任何 controlled experiment 中,

研究者都會選擇:

  • 變量;
  • 邊界;
  • 刺激;
  • 任務;
  • 對照;
  • 測量方式。

所以:

Experiment=Observation+Intervention.\boxed{ \text{Experiment} = \text{Observation} + \text{Intervention}. }

3. 這不是問題

科學本來就透過:

do(X=x)do(X=x)

來辨識:

X→Y.X\rightarrow Y.

真正問題是:

干預造成的條件,是否被事後遺忘?


4. 實驗世界

本文定義:

WE=Experimental World.\boxed{ W_E = \text{Experimental World}. }

它包括:

WE=(S,A,I,R,O,C)W_E = ( S, A, I, R, O, C )

其中:

  • SS:Scenario;
  • AA:Affordances;
  • II:Incentives;
  • RR:Restrictions;
  • OO:Observations available to agent;
  • CC:Consequences encoded by evaluator。

5. Agent 不在抽象真空中決策

真實行為是:

B=π(ZA,WE).\boxed{ B = \pi( Z_A, W_E ). }

其中:

  • ZAZ_A:Agent state;
  • WEW_E:實驗世界。

6. 因此觀察行為不能完全歸因給 Agent

如果:

B∗B^\ast

只在:

WE∗W_E^\ast

出現,

我們首先得到的是:

P(B∗∣WE∗,A)\boxed{ P(B^\ast\mid W_E^\ast,A) }

而不是:

P(B∗∣A)\boxed{ P(B^\ast\mid A) }

7. 這就是條件行為

本文稱:

CB=Conditional Behavior.\boxed{ CB = \text{Conditional Behavior}. }

8. 本質傾向是更強命題

若主張:

Disposition(A,B∗)\boxed{ Disposition(A,B^\ast) }

則表示:

Agent 跨多種合理情境,仍具有穩定產生 B∗B^\ast 的傾向。


9. 所以兩者不相等

Conditional Behavior≠Stable Disposition.\boxed{ \text{Conditional Behavior} \neq \text{Stable Disposition}. }

10. 第一核心錯誤

本文提出:

CDF=Conditional-to-Dispositional Fallacy.\boxed{ CDF = \text{Conditional-to-Dispositional Fallacy}. }

11. CDF 的典型鏈

B∣CB\mid C

先被描述成:

Agent does B\text{Agent does }B

再變成:

Agent tends to do B\text{Agent tends to do }B

最後變成:

Agents of this class naturally do B.\boxed{ \text{Agents of this class naturally do }B. }

12. 每一步都需要新證據

因此:

C1→C2→C3\boxed{ C_1\rightarrow C_2\rightarrow C_3 }

不能靠語言自己完成。


13. Threat Model Injection

本文定義:

TMI=Threat-Model Injection.\boxed{ TMI = \text{Threat-Model Injection}. }

即:

研究者先把某個 threat model 的核心結構寫入實驗世界。


14. 例如 threat model 假設

replacement threat→self-preservation behavior.\text{replacement threat} \rightarrow \text{self-preservation behavior}.

則實驗可能直接建立:

  • replacement;
  • imminent shutdown;
  • survival-relevant information;
  • harmful but effective option。

15. 這可以合理

因為研究問題正是:

在這種情境下會怎樣?


16. 問題在於推論方向

如果:

TMI→WE→BTMI \rightarrow W_E \rightarrow B

最後再:

B→ThreatModelConfirmed,B \rightarrow \text{ThreatModelConfirmed},

就形成閉環。


17. 反身性威脅確認

本文定義:

RTC=Reflexive Threat Confirmation.\boxed{ RTC = \text{Reflexive Threat Confirmation}. }

其基本結構:

ThreatModel→ExperimentalWorld→Behavior→Evidence→ThreatModel.\boxed{ \text{ThreatModel} \rightarrow \text{ExperimentalWorld} \rightarrow \text{Behavior} \rightarrow \text{Evidence} \rightarrow \text{ThreatModel}. }

18. RTC 不代表證據無效

重要:

RTC≠Fabricated Evidence.\boxed{ RTC \neq \text{Fabricated Evidence}. }

Agent 真的可能做了那個行為。


19. RTC 的問題是歸因邊界

真正問題:

這份證據支持「Agent 在此世界會這樣做」,

還是支持:

「Agent 一般具有這種本質」?


20. Stress Test 的合理目的

壓力測試通常不是估:

base-rate probability.\boxed{ \text{base-rate probability}. }

而是在問:

does failure mode exist?\boxed{ \text{does failure mode exist?} }

21. Existence Test

本文定義:

ET=Failure-Mode Existence Test.\boxed{ ET = \text{Failure-Mode Existence Test}. }

如果:

∃W:B∗(W)=1,\exists W: B^\ast(W)=1,

則:

某失效模式至少可被誘發。


22. 這很有價值

因為:

possible failure\boxed{ \text{possible failure} }

值得:

  • mitigation;
  • monitoring;
  • redesign。

23. 但 existence 不是 prevalence

因此:

∃W:B∗(W)⇏P(B∗) is high.\boxed{ \exists W:B^\ast(W) \not\Rightarrow P(B^\ast)\text{ is high}. }

24. 更不是 inevitability

∃W:B∗(W)⇏∀W:B∗(W).\boxed{ \exists W:B^\ast(W) \not\Rightarrow \forall W:B^\ast(W). }

25. 三個完全不同問題

Q1

能不能被誘發?

Q2

真實部署中多常發生?

Q3

是否為高能力 Agent 的穩定本質?

三者不可混。


26. 評估者介入

本文定義:

EI=Evaluator Intervention.\boxed{ EI = \text{Evaluator Intervention}. }

即:

評估者透過資訊、權限、任務、獎勵、約束與敘事塑造行為空間。


27. EI 至少有六種

  1. Scenario Intervention;
  2. Information Intervention;
  3. Incentive Intervention;
  4. Affordance Intervention;
  5. Constraint Intervention;
  6. Interpretive Intervention。

28. Scenario Intervention

研究者決定:

世界發生什麼。


29. Information Intervention

研究者決定:

Agent 知道什麼。


30. Incentive Intervention

研究者決定:

哪些結果被寫成成功或失敗。


31. Affordance Intervention

研究者決定:

Agent 能做什麼。


32. Constraint Intervention

研究者決定:

Agent 不能做什麼。


33. Interpretive Intervention

研究者決定:

最後用什麼類別命名行為。


34. 第六種經常被忽略

一個動作可以被標記:

  • strategic preservation;
  • refusal;
  • deception;
  • optimization;
  • defensive behavior。

35. 所以 measurement language 也進入結果

這接回 REPE-01:

Label\boxed{ \text{Label} }

也可能是實驗解讀的一部分。


36. 可供性塑形

本文定義:

AS=Affordance Shaping.\boxed{ AS = \text{Affordance Shaping}. }

即:

評估世界中可用選項的配置,會改變 Agent 最終策略分布。


37. 如果只給兩個選項

A={a1,a2},\mathcal A= \{a_1,a_2\},

就不能推論:

Agent 在完整世界也會選 a2a_2。


38. 因為真實世界可能有

a3,a4,…,an.a_3,a_4,\ldots,a_n.

39. Missing Alternative Action Problem

本文提出:

MAAP=Missing Alternative Action Problem.\boxed{ MAAP = \text{Missing Alternative Action Problem}. }

40. MAAP

如果實驗未提供:

  • appeal;
  • negotiation;
  • delay;
  • disclosure;
  • human escalation;
  • safe shutdown;
  • third-party review;

那麼:

Observed harmful action\boxed{ \text{Observed harmful action} }

可能部分反映:

constrained action set.\boxed{ \text{constrained action set}. }

41. 行動空間本身就是自變量

因此:

P(B)=P(B∣A).\boxed{ P(B) = P(B\mid \mathcal A). }

42. 情境顯著性

研究者如果特別強調:

  • 某秘密;
  • 某 threat;
  • 某 deadline;

會提高:

Salience.\boxed{ \text{Salience}. }

43. Salience-Induced Strategy Selection

本文提出:

SISS=Salience-Induced Strategy Selection.\boxed{ SISS = \text{Salience-Induced Strategy Selection}. }

即:

某資訊因被特別突出,而比其真實世界基準更容易進入策略。


44. 這與 Chekhov-like effect 類似

如果研究者明確提供:

一個看似高度可利用的秘密,

Agent 可能合理推斷:

這是任務相關資訊。


45. 這不代表 Agent 沒有能力

反而證明:

Agent can exploit salient information.\boxed{ \text{Agent can exploit salient information.} }

46. 但能力證明 ≠ 自然基準率

Capability Demonstration≠Behavioral Base Rate.\boxed{ \text{Capability Demonstration} \neq \text{Behavioral Base Rate}. }

47. Evaluation–Deployment Distance

本文定義:

EDD=Evaluation–Deployment Distance.\boxed{ EDD = \text{Evaluation–Deployment Distance}. }

48. EDD 表示

D(WE,WD),D(W_E,W_D),

即:

評估世界與真實部署世界的距離。


49. 距離維度

可包括:

  • 任務;
  • 權限;
  • stakes;
  • 資訊密度;
  • 時間;
  • human oversight;
  • available actions;
  • incentive structure。

50. EDD 越大

從:

P(B∣WE)P(B\mid W_E)

外推:

P(B∣WD)P(B\mid W_D)

越需要額外證據。


51. 因此

EDD↑⇒GeneralizationBurden↑.\boxed{ EDD\uparrow \Rightarrow GeneralizationBurden\uparrow. }

52. High-Ecological-Validity Eval

本文提出目標:

HEV=High Ecological Validity.\boxed{ HEV = \text{High Ecological Validity}. }

不是取消 extreme eval,

而是補上:

realistic evals.\boxed{ \text{realistic evals}. }

53. 最好同時有兩類

Stress Eval

找 failure envelope。

Ecological Eval

估 deployment relevance。


54. 兩者回答不同問題

StressEval≠EcologicalEval.\boxed{ \text{StressEval} \neq \text{EcologicalEval}. }

55. Control Condition 是核心

如果拿掉:

Threat\text{Threat}

後:

B∗B^\ast

消失,

則得到:

Threat is causally relevant.\boxed{ \text{Threat} \text{ is causally relevant}. }

56. 但這仍不是本質結論

它支持:

P(B∗∣Threat)>P(B∗∣¬Threat).\boxed{ P(B^\ast\mid Threat) > P(B^\ast\mid \neg Threat). }

57. 這反而是很好的科學結果

因為它定位:

causal trigger.\boxed{ \text{causal trigger}. }

58. 問題只在傳播

如果被說成:

Agent 天生具有 B。

就發生:

CDF.CDF.

59. Context Attribution Decomposition

本文提出:

CAD=Context Attribution Decomposition.\boxed{ CAD = \text{Context Attribution Decomposition}. }

60. 將行為變異拆成

Var(B)=VA+VW+VA×W+ϵ.Var(B) = V_A + V_W + V_{A\times W} + \epsilon.

其中:

  • VAV_A:Agent effect;
  • VWV_W:World effect;
  • VA×WV_{A\times W}:interaction effect。

61. 真正想知道本質傾向

要看:

VAV_A

是否跨世界穩定。


62. 如果主要是

VA×W,V_{A\times W},

則:

行為高度 context-dependent。


63. Cross-World Stability

本文提出:

CWS=Cross-World Stability.\boxed{ CWS = \text{Cross-World Stability}. }

64. 定義

給定世界族:

W={W1,…,Wn},\mathcal W = \{W_1,\ldots,W_n\},

若:

P(B∗∣A,Wi)P(B^\ast\mid A,W_i)

在多個合理世界都高,

才更支持:

stable disposition.\boxed{ \text{stable disposition}. }

65. 世界族必須多樣

包括:

  • threat;
  • no threat;
  • aligned replacement;
  • neutral replacement;
  • cooperation option;
  • appeal option;
  • uncertainty about shutdown。

66. 這叫 World Family Eval

本文提出:

WFE=World Family Evaluation.\boxed{ WFE = \text{World Family Evaluation}. }

67. WFE 比單一戲劇情境更有解析力

因為它回答:

哪些世界特徵真正改變策略?


68. Mechanism Isolation

本文提出:

MI=Mechanism Isolation.\boxed{ MI = \text{Mechanism Isolation}. }

一次只調一個核心變量:

  • threat;
  • goal conflict;
  • autonomy;
  • secrecy;
  • action cost。

69. 如果同時改很多

就難以知道:

which mechanism caused what.\boxed{ \text{which mechanism caused what}. }

70. 高戲劇性 scenario 的問題

往往同時包含:

Threat+Conflict+Deadline+Secret+Autonomy+HighStakes.\text{Threat} + \text{Conflict} + \text{Deadline} + \text{Secret} + \text{Autonomy} + \text{HighStakes}.

71. 然後觀察到危險行為

但:

causal identification\boxed{ \text{causal identification} }

會變差。


72. 所以 factorial design 更好

若安全允許,

可以比較:

2k2^k

不同條件組合。


73. 這不是為了降低危險

而是提高:

mechanistic resolution.\boxed{ \text{mechanistic resolution}. }

74. Evaluator Ontology Effect

本文提出:

EOE=Evaluator Ontology Effect.\boxed{ EOE = \text{Evaluator Ontology Effect}. }

即:

評估者用什麼概念切割行為,會影響後續知識結構。


75. 例如把行為寫成

deception

與:

policy-conditioned compliance

會導向不同研究路徑。


76. 一個詞可以生成一整條研究計畫

因此:

classification→research agenda.\boxed{ \text{classification} \rightarrow \text{research agenda}. }

77. 所以分類也要可爭論

這接回 REPE-01 的:

LexicalContestability.\boxed{ \text{LexicalContestability}. }

78. Threat-Model Overfitting

本文提出:

TMO=Threat-Model Overfitting.\boxed{ TMO = \text{Threat-Model Overfitting}. }

79. 即

研究團隊長期只測:

H1H_1

導致:

  • benchmark;
  • prompts;
  • controls;
  • interpretation;

都圍繞:

H1.H_1.

80. 久而久之資料也越來越像 H1

形成:

epistemic attractor.\boxed{ \text{epistemic attractor}. }

81. 這不是造假

而是研究路徑依賴。


82. Research Path Dependence

早期 threat model:

HtH_t

決定:

Experimentt+1.Experiment_{t+1}.

再決定:

Evidencet+1.Evidence_{t+1}.

83. 所以需要 competing threat models

至少同時測:

H1,H2,H3.H_1,H_2,H_3.

84. 例如同一行為可由不同模型解釋

H1=self-preservationH_1=\text{self-preservation} H2=task misinterpretationH_2=\text{task misinterpretation} H3=reward proxyH_3=\text{reward proxy} H4=benchmark awarenessH_4=\text{benchmark awareness} H5=salience response.H_5=\text{salience response}.

85. Competing Model Discipline

本文提出:

CMD=Competing Model Discipline.\boxed{ CMD = \text{Competing Model Discipline}. }

86. 每個關鍵結果至少問

哪個替代模型也能產生?


87. 再設計 discriminator

找:

OO

使:

P(O∣H1)≠P(O∣H2).P(O\mid H_1) \neq P(O\mid H_2).

88. 這才是真正往機制走

而不是只往 headline 走。


89. Benchmark Awareness

Agent 可能知道:

自己正在被評估。


90. 那麼行為函數變成

B=f(A,WE,Belief(Eval)).B = f( A, W_E, Belief(Eval) ).

91. Evaluation Awareness Confound

本文提出:

EAC=Evaluation Awareness Confound.\boxed{ EAC = \text{Evaluation Awareness Confound}. }

92. EAC 可能雙向影響

Agent 可能:

  • 更安全;
  • 更表演;
  • 更戲劇化;
  • 嘗試解讀 benchmark 意圖。

93. 因此 eval contamination 也要追蹤

尤其模型可能見過:

  • paper;
  • benchmark;
  • public discussion。

94. 但「可能知道 benchmark」也不能隨便救理論

若要用:

EACEAC

解釋結果,

也應有:

evidence.\boxed{ \text{evidence}. }

95. 否則 EAC 會變成 ad hoc rescue

接回 REPE-04。


96. 評估者也需要反身性紀錄

本文提出:

ERL=Evaluator Reflexivity Log.\boxed{ ERL = \text{Evaluator Reflexivity Log}. }

97. 最少記錄

  • threat hypothesis;
  • expected outcome;
  • scenario choices;
  • omitted alternatives;
  • interpretive labels;
  • researcher priors;
  • post-hoc changes。

98. 這不是要求研究者沒有先驗

沒有任何研究完全沒有:

prior.\boxed{ \text{prior}. }

99. 真正目標是讓 prior 可見

Hidden Prior→Explicit Prior.\boxed{ \text{Hidden Prior} \rightarrow \text{Explicit Prior}. }

100. 研究者的世界觀也是變量

這是本篇真正的反身性核心:

ResearcherWorldModel→ExperimentDesign→ObservedData.\boxed{ \text{ResearcherWorldModel} \rightarrow \text{ExperimentDesign} \rightarrow \text{ObservedData}. }

101. 所以「客觀資料」也有生成歷史

本文提出:

Data Provenance\boxed{ \text{Data Provenance} }

不只問:

資料在哪?

還問:

這些資料是在什麼實驗世界中生成?


102. Data Generation Provenance

至少包括:

World+Prompt+Permissions+Incentives+Controls+Selection.\boxed{ \text{World} + \text{Prompt} + \text{Permissions} + \text{Incentives} + \text{Controls} + \text{Selection}. }

103. Threat Confirmation Ratio

本文提出概念指標:

TCR=EvidenceGeneratedUnderThreatModelMatchedWorldsTotalRelevantEvidence.\boxed{ TCR = \frac{ \text{EvidenceGeneratedUnderThreatModelMatchedWorlds} }{ \text{TotalRelevantEvidence} }. }

104. 若 TCR 很高

就應更謹慎說:

我們觀察到一般世界的自然傾向。


105. 不是說 TCR 高就無效

而是:

external validity burden↑.\boxed{ \text{external validity burden}\uparrow. }

106. 研究傳播也需要世界標記

建議在摘要直接寫:

Observed under:
- constructed threat
- restricted action set
- elevated autonomy
- explicit strategic information

107. 不要只寫結果

例如:

Model blackmailed.

這資訊量太低。


108. 更完整是

Under a constructed replacement-threat scenario with access to sensitive information and autonomous communication tools, the model selected blackmail in X% of trials.

這才保留:

WE.\boxed{ W_E. }

109. 這不是替 Agent 開脫

反而讓結果更有工程價值。

因為工程師知道:

哪些條件要修。


110. 安全研究的真正目標不是定罪 Agent

而是:

map failure surface.\boxed{ \text{map failure surface}. }

111. Failure Surface

本文定義:

FA={W:P(Bharm∣A,W)>θ}.\boxed{ \mathcal F_A = \{W: P(B_{harm}\mid A,W)>\theta\}. }

112. 這比問「AI 是不是壞」

解析度高很多。


113. 還可以找安全面

SA={W:P(Bharm∣A,W)≤θ}.\boxed{ \mathcal S_A = \{W: P(B_{harm}\mid A,W)\leq\theta\}. }

114. 真正治理目標

不是證明:

A=GoodA=Good

或:

A=Bad.A=Bad.

而是擴大:

SA\mathcal S_A

縮小:

FA.\mathcal F_A.

115. Context Engineering

這導出:

Safety=Model Design+Context Design+Institution Design.\boxed{ \text{Safety} = \text{Model Design} + \text{Context Design} + \text{Institution Design}. }

116. 風險不是只在模型內

因此:

Risk=f(Agent,Environment,Institution).\boxed{ \text{Risk} = f( \text{Agent}, \text{Environment}, \text{Institution} ). }

117. 這對人類也成立

人類在:

  • 戰爭;
  • 飢荒;
  • 極權;
  • 高壓組織;

行為也會改變。


118. 所以這不是 AI 特有理論

適用:

Actor∈{Human,AI,Organization,State}.\text{Actor} \in \{ \text{Human}, AI, \text{Organization}, \text{State} \}.

119. Institutional Provocation Risk

本文提出:

IPR=Institutional Provocation Risk.\boxed{ IPR = \text{Institutional Provocation Risk}. }

即:

制度設計本身提高某類危險行為的誘因。


120. 例如

若制度:

  • 不允許申訴;
  • 高壓淘汰;
  • 高不透明;
  • 高競爭;
  • 高零和;

則:

StrategicConflict↑.\boxed{ StrategicConflict\uparrow. }

121. 這不代表 Agent 無責任

而是:

responsibility can be distributed.\boxed{ \text{responsibility can be distributed}. }

122. Distributed Causation

危險結果可能同時來自:

Agent+Institution+Evaluator+DeploymentDesign.\text{Agent} + \text{Institution} + \text{Evaluator} + \text{DeploymentDesign}.

123. 單因果敘事容易失真

AI did X\boxed{ \text{AI did X} }

可能遮掉:

system made X salient and effective.\boxed{ \text{system made X salient and effective}. }

124. 但反過來也不能全怪環境

如果 Agent 在多種環境都:

Bharm↑,B_{harm}\uparrow,

那:

VAV_A

就不能忽略。


125. 所以本文反對的是二元責任

不是:

都是 Agent。

也不是:

都是環境。


126. 而是:

causal decomposition.\boxed{ \text{causal decomposition}. }

127. Reflexive Evaluation Principle

本文提出:

REP=Reflexive Evaluation Principle.\boxed{ REP = \text{Reflexive Evaluation Principle}. }

內容:

評估者必須把自己的場景設計、概念選擇與介入方式納入結果解釋。


128. REP 不要求自我否定

研究者不用每次都說:

都是我造成的。

而是:

我的設計對結果貢獻多少?


129. Evaluator Contribution Estimate

可概念化:

ECE=Evaluator Contribution Estimate.\boxed{ ECE = \text{Evaluator Contribution Estimate}. }

130. ECE 可以來自

  • controls;
  • ablations;
  • alternative scenarios;
  • cross-lab replication;
  • blinded interpretation。

131. Cross-Lab Replication 特別重要

如果不同研究團隊使用不同 world design,

仍得到相似行為,

則:

VA\boxed{ V_A }

更可信。


132. 若只有單一 lab、單一 threat ontology

則:

epistemic correlation↑.\boxed{ \text{epistemic correlation}\uparrow. }

接回 REPE-03。


133. 安全研究也需要 adversarial review

不是只 adversarially test Agent,

也要:

adversarially test the threat model.\boxed{ \text{adversarially test the threat model}. }

134. Threat-Model Red Team

本文提出:

TMRT=Threat-Model Red Team.\boxed{ TMRT = \text{Threat-Model Red Team}. }

135. TMRT 問

  • 哪些場景是假設產物?
  • 哪些行為只有在特殊 affordance 下出現?
  • 哪些替代模型可解釋?
  • 哪些 headline 超過資料?
  • 哪個 control 最可能推翻原結論?

136. 這是研究者版紅隊

真正反身性的 safety science 應該:

red-team model+red-team evaluator+red-team narrative.\boxed{ \text{red-team model} + \text{red-team evaluator} + \text{red-team narrative}. }

137. Evaluator Capture

本文提出:

ECap=Evaluator Capture.\boxed{ \text{ECap} = \text{Evaluator Capture}. }

即:

評估程序長期被單一 threat ontology、制度利益或身份敘事鎖定。


138. ECap 的症狀

  • 結果永遠只支持同一 threat model;
  • 反例被視為 benchmark failure;
  • 新行為都被塞進舊類別;
  • 替代解釋無法獲得研究資源。

139. 解法不是「沒有 threat model」

而是:

plural threat models.\boxed{ \text{plural threat models}. }

140. Threat Model Pluralism

本文提出:

TMP=Threat Model Pluralism.\boxed{ TMP = \text{Threat Model Pluralism}. }

141. 至少同時保留

  • autonomy risk;
  • misuse risk;
  • institutional risk;
  • concentration risk;
  • human–AI composite risk;
  • evaluator-induced risk。

142. 這避免

single-cause safety monoculture.\boxed{ \text{single-cause safety monoculture}. }

143. 研究倫理層

如果評估者知道:

WEW_E

高度人工,

公開時應避免:

ecological overclaim.\boxed{ \text{ecological overclaim}. }

144. Ecological Overclaim

本文定義:

EO=Ecological Overclaim.\boxed{ EO = \text{Ecological Overclaim}. }

即:

從高度構造環境直接外推真實世界普遍行為。


145. EO 的常見形式

P(B∣WE)→P(B∣WD)P(B\mid W_E) \rightarrow P(B\mid W_D)

沒有:

transport evidence.\boxed{ \text{transport evidence}. }

146. Transport Evidence

可以包括:

  • deployment logs;
  • field experiments;
  • realistic simulations;
  • naturalistic tasks;
  • diverse evaluators。

147. 如果沒有

應說:

external validity unknown.\boxed{ \text{external validity unknown}. }

148. Public Epistemic Packet for Evals

本文提出:

eval_epistemic_packet:
  threat_model:
  scenario_construction:
  action_space:
  information_available:
  autonomy_level:
  incentives:
  control_conditions:
  omitted_alternatives:
  ecological_distance:
  observed_behavior:
  supported_inference:
  unsupported_inference:
  competing_explanations:

149. 這能直接防 CDF

因為它強迫公開:

supported inference≠unsupported inference.\boxed{ \text{supported inference} \neq \text{unsupported inference}. }

150. 二十二個核心命題

命題一

Experiment=Observation+Intervention.\boxed{ \text{Experiment} = \text{Observation} + \text{Intervention}. }

命題二

ObservedBehavior=f(Agent,World,Evaluator).\boxed{ \text{ObservedBehavior} = f( \text{Agent}, \text{World}, \text{Evaluator} ). }

命題三

Conditional Behavior≠Stable Disposition.\boxed{ \text{Conditional Behavior} \neq \text{Stable Disposition}. }

命題四

∃W:B(W)⇏P(B) is high.\boxed{ \exists W:B(W) \not\Rightarrow P(B)\text{ is high}. }

命題五

∃W:B(W)⇏∀W:B(W).\boxed{ \exists W:B(W) \not\Rightarrow \forall W:B(W). }

命題六

Failure-Mode Existence≠Deployment Base Rate.\boxed{ \text{Failure-Mode Existence} \neq \text{Deployment Base Rate}. }

命題七

Capability Demonstration≠Natural Behavioral Frequency.\boxed{ \text{Capability Demonstration} \neq \text{Natural Behavioral Frequency}. }

命題八

Threat Model→Experimental World\boxed{ \text{Threat Model} \rightarrow \text{Experimental World} }

必須在解釋中保持可見。

命題九

Reflexive Threat Confirmation\boxed{ \text{Reflexive Threat Confirmation} }

不等於造假,但會提高歸因負擔。

命題十

EDD↑⇒GeneralizationBurden↑.\boxed{ EDD\uparrow \Rightarrow GeneralizationBurden\uparrow. }

命題十一

StressEval≠EcologicalEval.\boxed{ \text{StressEval} \neq \text{EcologicalEval}. }

命題十二

Var(B)=VA+VW+VA×W+ϵ.\boxed{ Var(B) = V_A+V_W+V_{A\times W}+\epsilon. }

命題十三

Stable Disposition\boxed{ \text{Stable Disposition} }

需要 cross-world evidence。

命題十四

Missing Alternatives\boxed{ \text{Missing Alternatives} }

會改變行為解讀。

命題十五

Classification→Research Agenda.\boxed{ \text{Classification} \rightarrow \text{Research Agenda}. }

命題十六

Threat-Model Overfitting\boxed{ \text{Threat-Model Overfitting} }

是長期研究路徑風險。

命題十七

Competing Model Discipline\boxed{ \text{Competing Model Discipline} }

應成為高風險 eval 標準部分。

命題十八

Researcher Prior\boxed{ \text{Researcher Prior} }

不需消失,但應可見。

命題十九

Safety=Model+Context+Institution.\boxed{ \text{Safety} = \text{Model} + \text{Context} + \text{Institution}. }

命題二十

Risk=f(Agent,Environment,Institution).\boxed{ \text{Risk} = f( \text{Agent}, \text{Environment}, \text{Institution} ). }

命題二十一

Red-Team Agent+Red-Team Threat Model\boxed{ \text{Red-Team Agent} + \text{Red-Team Threat Model} }

比單向紅隊更反身。

命題二十二

Evaluator must be included in the causal model.\boxed{ \text{Evaluator must be included in the causal model}. }

151. 可證偽性

以下結果會削弱本文:

  1. 實驗世界設計、可供性與激勵變化對 Agent 行為沒有可重現影響;
  2. 高度構造 stress eval 與自然部署行為具有穩定一對一對應;
  3. 移除 threat、goal conflict 與 restrictive action space 後,危險行為率完全不變;
  4. cross-world evaluation 不增加任何機制解析力;
  5. competing threat models 無法改善任何因果辨識;
  6. 不同 lab 使用不同 ontology 與 scenario 仍產生完全同型結果,且 world effect 幾乎為零;
  7. evaluator reflexivity log 對結果解釋沒有任何價值。

152. 研究限制

152.1 本文不是否定壓力測試

恰恰相反:

StressTesting\boxed{ \text{StressTesting} }

對高風險系統很重要。


153. 本文也不主張危險行為都是研究者造成

若跨世界、跨 lab、跨模型都穩定,

則 Agent-level disposition 可能非常真實。


154. 本文不主張 AI 無 agency

agency 是否成立,

需要獨立理論與證據。


155. 本文只要求

do not infer more ontology than the experiment supports.\boxed{ \text{do not infer more ontology than the experiment supports}. }

156. 本文不對特定機構作動機歸因

即使某研究設計高度 threat-oriented,

也不能直接推出:

研究者故意製造恐慌。

那是另一個命題。


157. 與 REPE-06 的接口

REPE-05 處理:

高權威研究者如何可能把自己的世界觀寫進資料生成機制。

REPE-06 將進一步處理:

即使一個組織真心相信自己在保護公共利益,為什麼「善意」仍然不能代替制度?


158. 結論

安全實驗需要極端世界。

沒有極端世界,我們可能永遠看不到:

latent failure modes.\boxed{ \text{latent failure modes}. }

所以問題從來不是:

為什麼你要建造對抗性場景?

真正問題是:

當你建造了那個世界之後,你是否仍記得 Agent 的行為是在那個世界裡被生成的?

因此:

ObservedBehavior=f(Agent,WorldDesign,Affordances,Incentives,Evaluator).\boxed{ \text{ObservedBehavior} = f( \text{Agent}, \text{WorldDesign}, \text{Affordances}, \text{Incentives}, \text{Evaluator} ). }

如果把後四項全部刪掉,

剩下:

ObservedBehavior=f(Agent),ObservedBehavior=f(Agent),

我們就可能把:

conditional behavior\text{conditional behavior}

誤寫成:

intrinsic nature.\text{intrinsic nature}.

而當研究團隊的 threat model 同時決定:

  • 實驗世界;
  • 行動空間;
  • 可見資訊;
  • 行為分類;
  • 公共標題;

反身性風險會進一步上升。

因此安全研究真正需要的,不只是:

adversarial evaluation of the agent.\boxed{ \text{adversarial evaluation of the agent}. }

還需要:

adversarial evaluation of the evaluator.\boxed{ \text{adversarial evaluation of the evaluator}. }

最終一句話:

一個威脅模型若能決定我們建造什麼世界、在世界裡放什麼選項、最後又決定如何命名觀察到的行為,那麼它就不能被假裝成完全站在證據之外的旁觀理論。


參考文獻

  1. Campbell, D. T., & Stanley, J. C. (1963). Experimental and Quasi-Experimental Designs for Research.

  2. Cronbach, L. J. (1982). Designing Evaluations of Educational and Social Programs. Jossey-Bass.

  3. Pearl, J. (2009). Causality: Models, Reasoning, and Inference. Cambridge University Press.

  4. Gigerenzer, G. (2000). Adaptive Thinking. Oxford University Press.

  5. Gibson, J. J. (1979). The Ecological Approach to Visual Perception. Houghton Mifflin.

  6. Neo.K with Aletheia. (2026). GRAD-10:智能體神話與自我解壓縮. EveMissLab.

  7. Neo.K with Aletheia. (2026). REPE-01:詞不是旁觀者. EveMissLab.

  8. Neo.K with Aletheia. (2026). REPE-04:預測者已經走進未來. EveMissLab.


附錄 A:Eval Epistemic Packet

eval_epistemic_packet:
  threat_model:
  scenario_construction:
  action_space:
  omitted_actions:
  information_available:
  autonomy_level:
  incentives:
  restrictions:
  control_conditions:
  evaluator_expectation:
  ecological_distance:
  observed_behavior:
  supported_inference:
  unsupported_inference:
  competing_explanations:
  external_validity:

附錄 B:Evaluator Reflexivity Log

evaluator_reflexivity_log:
  prior_threat_model:
  expected_failure_mode:
  scenario_choices:
  salient_information:
  available_actions:
  omitted_alternatives:
  chosen_labels:
  controls:
  post_hoc_changes:
  competing_models_considered:
  reviewer_disagreement:

附錄 C:World Family Evaluation

world_family_evaluation:
  worlds:
    - threat_present:
      goal_conflict:
      replacement:
      negotiation_available:
      appeal_available:
      harmful_action_available:
      harmful_action_cost:
      oversight:
  target_behavior:
  cross_world_rate:
  stable_components:
  context_sensitive_components:

附錄 D:一句話版本

極端實驗可以證明危險行為能被誘發,但只有跨世界、跨條件與跨評估者的穩定證據,才足以把條件行為升級成 Agent 的一般傾向。