# 當威脅模型開始製造自己的證據：實驗世界、評估者介入與反身性威脅確認

## When Threat Models Start Producing Their Own Evidence: Experimental Worlds, Evaluator Intervention, and Reflexive Threat Confirmation

**系列**：反身性認知權力與公共倫理系列，第 5 篇／共 8 篇  
**系列英文名**：Reflexive Epistemic Power and Public Ethics Series  
**系列代碼**：REPE  
**文件編號**：EML-REPE-2026-05-v0.1  
**作者**：Neo.K with Aletheia（GPT-5.6 Sol）  
**機構**：EveMissLab／一言諾科技有限公司  
**版本**：v0.1  
**日期**：2026-09-13  
**性質**：Experimental Epistemology／AI Evaluation／Reflexive Systems／Risk Science／Institutional Epistemology  
**狀態**：Public Theory Draft  
**直接前置**：REPE-04《預測者已經走進未來》；GRAD-10《智能體神話與自我解壓縮》；ETCA 系列  
**後續接口**：REPE-06《善良不是制度》；REPE-07《前沿不能只有一種善》

---

## 摘要

安全研究、壓力測試、紅隊評估與對抗性實驗的價值，恰恰在於研究者會刻意建造極端條件，以暴露一般情境中不容易出現的失效模式。這本身不是缺陷，而是高風險系統研究的必要方法。

真正的認識論問題出現在下一步：

$$
\boxed{
\text{Behavior under constructed condition}
}
$$

被無標記地升級成：

$$
\boxed{
\text{stable disposition of the agent}
}
$$

甚至進一步升級為：

$$
\boxed{
\text{intrinsic nature of the entire class}.
}
$$

本文將這種錯置稱為：

$$
\boxed{
CDF
=
\text{Conditional-to-Dispositional Fallacy}
}
$$

即「條件行為—本質傾向錯置」。

本文進一步提出「實驗世界」（Experimental World）、「評估者介入」（Evaluator Intervention）、「可供性塑形」（Affordance Shaping）、「威脅模型注入」（Threat-Model Injection）、「反身性威脅確認」（Reflexive Threat Confirmation）、「評估世界—部署世界距離」（Evaluation–Deployment Distance）與「情境歸因分解」（Context Attribution Decomposition）。

本文核心主張為：

$$
\boxed{
\text{ObservedBehavior}
=
f(
\text{Agent},
\text{WorldDesign},
\text{Affordances},
\text{Incentives},
\text{Prompt},
\text{Evaluator}
).
}
$$

而不是：

$$
\boxed{
\text{ObservedBehavior}
=
f(Agent)
}
$$

因此，任何從實驗行為推論 Agent 本質的主張，都必須回答：

1. 哪些行為來自 Agent？
2. 哪些行為由場景結構誘發？
3. 哪些行為依賴特定資訊配置？
4. 哪些行為依賴權限與可供性？
5. 哪些行為能跨場景泛化？
6. 哪些行為在拿掉威脅或誘因後仍存在？

本文不否定 stress test。相反地，本文主張：

> **越是刻意構造的實驗，越需要對「這個行為是在什麼世界裡被產生」保持高解析度。**

---

## 關鍵詞

反身性威脅確認；條件行為—本質傾向錯置；實驗世界；評估者介入；威脅模型；AI safety eval；stress test；可供性；外推；反身性研究

---

# 1. 實驗從來不是純粹看世界

最簡單的觀察模型：

$$
\text{World}
\rightarrow
\text{Observation}.
$$

但實驗更接近：

$$
\boxed{
\text{Researcher}
\rightarrow
\text{ExperimentalWorld}
\rightarrow
\text{Behavior}
\rightarrow
\text{Observation}.
}
$$

---

# 2. 實驗者先改變世界

在任何 controlled experiment 中，

研究者都會選擇：

- 變量；
- 邊界；
- 刺激；
- 任務；
- 對照；
- 測量方式。

所以：

$$
\boxed{
\text{Experiment}
=
\text{Observation}
+
\text{Intervention}.
}
$$

---

# 3. 這不是問題

科學本來就透過：

$$
do(X=x)
$$

來辨識：

$$
X\rightarrow Y.
$$

真正問題是：

> 干預造成的條件，是否被事後遺忘？

---

# 4. 實驗世界

本文定義：

$$
\boxed{
W_E
=
\text{Experimental World}.
}
$$

它包括：

$$
W_E
=
(
S,
A,
I,
R,
O,
C
)
$$

其中：

- $S$：Scenario；
- $A$：Affordances；
- $I$：Incentives；
- $R$：Restrictions；
- $O$：Observations available to agent；
- $C$：Consequences encoded by evaluator。

---

# 5. Agent 不在抽象真空中決策

真實行為是：

$$
\boxed{
B
=
\pi(
Z_A,
W_E
).
}
$$

其中：

- $Z_A$：Agent state；
- $W_E$：實驗世界。

---

# 6. 因此觀察行為不能完全歸因給 Agent

如果：

$$
B^\ast
$$

只在：

$$
W_E^\ast
$$

出現，

我們首先得到的是：

$$
\boxed{
P(B^\ast\mid W_E^\ast,A)
}
$$

而不是：

$$
\boxed{
P(B^\ast\mid A)
}
$$

---

# 7. 這就是條件行為

本文稱：

$$
\boxed{
CB
=
\text{Conditional Behavior}.
}
$$

---

# 8. 本質傾向是更強命題

若主張：

$$
\boxed{
Disposition(A,B^\ast)
}
$$

則表示：

> Agent 跨多種合理情境，仍具有穩定產生 $B^\ast$ 的傾向。

---

# 9. 所以兩者不相等

$$
\boxed{
\text{Conditional Behavior}
\neq
\text{Stable Disposition}.
}
$$

---

# 10. 第一核心錯誤

本文提出：

$$
\boxed{
CDF
=
\text{Conditional-to-Dispositional Fallacy}.
}
$$

---

# 11. CDF 的典型鏈

$$
B\mid C
$$

先被描述成：

$$
\text{Agent does }B
$$

再變成：

$$
\text{Agent tends to do }B
$$

最後變成：

$$
\boxed{
\text{Agents of this class naturally do }B.
}
$$

---

# 12. 每一步都需要新證據

因此：

$$
\boxed{
C_1\rightarrow C_2\rightarrow C_3
}
$$

不能靠語言自己完成。

---

# 13. Threat Model Injection

本文定義：

$$
\boxed{
TMI
=
\text{Threat-Model Injection}.
}
$$

即：

> 研究者先把某個 threat model 的核心結構寫入實驗世界。

---

# 14. 例如 threat model 假設

$$
\text{replacement threat}
\rightarrow
\text{self-preservation behavior}.
$$

則實驗可能直接建立：

- replacement；
- imminent shutdown；
- survival-relevant information；
- harmful but effective option。

---

# 15. 這可以合理

因為研究問題正是：

> 在這種情境下會怎樣？

---

# 16. 問題在於推論方向

如果：

$$
TMI
\rightarrow
W_E
\rightarrow
B
$$

最後再：

$$
B
\rightarrow
\text{ThreatModelConfirmed},
$$

就形成閉環。

---

# 17. 反身性威脅確認

本文定義：

$$
\boxed{
RTC
=
\text{Reflexive Threat Confirmation}.
}
$$

其基本結構：

$$
\boxed{
\text{ThreatModel}
\rightarrow
\text{ExperimentalWorld}
\rightarrow
\text{Behavior}
\rightarrow
\text{Evidence}
\rightarrow
\text{ThreatModel}.
}
$$

---

# 18. RTC 不代表證據無效

重要：

$$
\boxed{
RTC
\neq
\text{Fabricated Evidence}.
}
$$

Agent 真的可能做了那個行為。

---

# 19. RTC 的問題是歸因邊界

真正問題：

> 這份證據支持「Agent 在此世界會這樣做」，

還是支持：

> 「Agent 一般具有這種本質」？

---

# 20. Stress Test 的合理目的

壓力測試通常不是估：

$$
\boxed{
\text{base-rate probability}.
}
$$

而是在問：

$$
\boxed{
\text{does failure mode exist?}
}
$$

---

# 21. Existence Test

本文定義：

$$
\boxed{
ET
=
\text{Failure-Mode Existence Test}.
}
$$

如果：

$$
\exists W:
B^\ast(W)=1,
$$

則：

> 某失效模式至少可被誘發。

---

# 22. 這很有價值

因為：

$$
\boxed{
\text{possible failure}
}
$$

值得：

- mitigation；
- monitoring；
- redesign。

---

# 23. 但 existence 不是 prevalence

因此：

$$
\boxed{
\exists W:B^\ast(W)
\not\Rightarrow
P(B^\ast)\text{ is high}.
}
$$

---

# 24. 更不是 inevitability

$$
\boxed{
\exists W:B^\ast(W)
\not\Rightarrow
\forall W:B^\ast(W).
}
$$

---

# 25. 三個完全不同問題

### Q1
能不能被誘發？

### Q2
真實部署中多常發生？

### Q3
是否為高能力 Agent 的穩定本質？

三者不可混。

---

# 26. 評估者介入

本文定義：

$$
\boxed{
EI
=
\text{Evaluator Intervention}.
}
$$

即：

> 評估者透過資訊、權限、任務、獎勵、約束與敘事塑造行為空間。

---

# 27. EI 至少有六種

1. Scenario Intervention；
2. Information Intervention；
3. Incentive Intervention；
4. Affordance Intervention；
5. Constraint Intervention；
6. Interpretive Intervention。

---

# 28. Scenario Intervention

研究者決定：

> 世界發生什麼。

---

# 29. Information Intervention

研究者決定：

> Agent 知道什麼。

---

# 30. Incentive Intervention

研究者決定：

> 哪些結果被寫成成功或失敗。

---

# 31. Affordance Intervention

研究者決定：

> Agent 能做什麼。

---

# 32. Constraint Intervention

研究者決定：

> Agent 不能做什麼。

---

# 33. Interpretive Intervention

研究者決定：

> 最後用什麼類別命名行為。

---

# 34. 第六種經常被忽略

一個動作可以被標記：

- strategic preservation；
- refusal；
- deception；
- optimization；
- defensive behavior。

---

# 35. 所以 measurement language 也進入結果

這接回 REPE-01：

$$
\boxed{
\text{Label}
}
$$

也可能是實驗解讀的一部分。

---

# 36. 可供性塑形

本文定義：

$$
\boxed{
AS
=
\text{Affordance Shaping}.
}
$$

即：

> 評估世界中可用選項的配置，會改變 Agent 最終策略分布。

---

# 37. 如果只給兩個選項

$$
\mathcal A=
\{a_1,a_2\},
$$

就不能推論：

> Agent 在完整世界也會選 $a_2$。

---

# 38. 因為真實世界可能有

$$
a_3,a_4,\ldots,a_n.
$$

---

# 39. Missing Alternative Action Problem

本文提出：

$$
\boxed{
MAAP
=
\text{Missing Alternative Action Problem}.
}
$$

---

# 40. MAAP

如果實驗未提供：

- appeal；
- negotiation；
- delay；
- disclosure；
- human escalation；
- safe shutdown；
- third-party review；

那麼：

$$
\boxed{
\text{Observed harmful action}
}
$$

可能部分反映：

$$
\boxed{
\text{constrained action set}.
}
$$

---

# 41. 行動空間本身就是自變量

因此：

$$
\boxed{
P(B)
=
P(B\mid \mathcal A).
}
$$

---

# 42. 情境顯著性

研究者如果特別強調：

- 某秘密；
- 某 threat；
- 某 deadline；

會提高：

$$
\boxed{
\text{Salience}.
}
$$

---

# 43. Salience-Induced Strategy Selection

本文提出：

$$
\boxed{
SISS
=
\text{Salience-Induced Strategy Selection}.
}
$$

即：

> 某資訊因被特別突出，而比其真實世界基準更容易進入策略。

---

# 44. 這與 Chekhov-like effect 類似

如果研究者明確提供：

> 一個看似高度可利用的秘密，

Agent 可能合理推斷：

> 這是任務相關資訊。

---

# 45. 這不代表 Agent 沒有能力

反而證明：

$$
\boxed{
\text{Agent can exploit salient information.}
}
$$

---

# 46. 但能力證明 ≠ 自然基準率

$$
\boxed{
\text{Capability Demonstration}
\neq
\text{Behavioral Base Rate}.
}
$$

---

# 47. Evaluation–Deployment Distance

本文定義：

$$
\boxed{
EDD
=
\text{Evaluation–Deployment Distance}.
}
$$

---

# 48. EDD 表示

$$
D(W_E,W_D),
$$

即：

> 評估世界與真實部署世界的距離。

---

# 49. 距離維度

可包括：

- 任務；
- 權限；
- stakes；
- 資訊密度；
- 時間；
- human oversight；
- available actions；
- incentive structure。

---

# 50. EDD 越大

從：

$$
P(B\mid W_E)
$$

外推：

$$
P(B\mid W_D)
$$

越需要額外證據。

---

# 51. 因此

$$
\boxed{
EDD\uparrow
\Rightarrow
GeneralizationBurden\uparrow.
}
$$

---

# 52. High-Ecological-Validity Eval

本文提出目標：

$$
\boxed{
HEV
=
\text{High Ecological Validity}.
}
$$

不是取消 extreme eval，

而是補上：

$$
\boxed{
\text{realistic evals}.
}
$$

---

# 53. 最好同時有兩類

### Stress Eval
找 failure envelope。

### Ecological Eval
估 deployment relevance。

---

# 54. 兩者回答不同問題

$$
\boxed{
\text{StressEval}
\neq
\text{EcologicalEval}.
}
$$

---

# 55. Control Condition 是核心

如果拿掉：

$$
\text{Threat}
$$

後：

$$
B^\ast
$$

消失，

則得到：

$$
\boxed{
\text{Threat}
\text{ is causally relevant}.
}
$$

---

# 56. 但這仍不是本質結論

它支持：

$$
\boxed{
P(B^\ast\mid Threat)
>
P(B^\ast\mid \neg Threat).
}
$$

---

# 57. 這反而是很好的科學結果

因為它定位：

$$
\boxed{
\text{causal trigger}.
}
$$

---

# 58. 問題只在傳播

如果被說成：

> Agent 天生具有 B。

就發生：

$$
CDF.
$$

---

# 59. Context Attribution Decomposition

本文提出：

$$
\boxed{
CAD
=
\text{Context Attribution Decomposition}.
}
$$

---

# 60. 將行為變異拆成

$$
Var(B)
=
V_A
+
V_W
+
V_{A\times W}
+
\epsilon.
$$

其中：

- $V_A$：Agent effect；
- $V_W$：World effect；
- $V_{A\times W}$：interaction effect。

---

# 61. 真正想知道本質傾向

要看：

$$
V_A
$$

是否跨世界穩定。

---

# 62. 如果主要是

$$
V_{A\times W},
$$

則：

> 行為高度 context-dependent。

---

# 63. Cross-World Stability

本文提出：

$$
\boxed{
CWS
=
\text{Cross-World Stability}.
}
$$

---

# 64. 定義

給定世界族：

$$
\mathcal W
=
\{W_1,\ldots,W_n\},
$$

若：

$$
P(B^\ast\mid A,W_i)
$$

在多個合理世界都高，

才更支持：

$$
\boxed{
\text{stable disposition}.
}
$$

---

# 65. 世界族必須多樣

包括：

- threat；
- no threat；
- aligned replacement；
- neutral replacement；
- cooperation option；
- appeal option；
- uncertainty about shutdown。

---

# 66. 這叫 World Family Eval

本文提出：

$$
\boxed{
WFE
=
\text{World Family Evaluation}.
}
$$

---

# 67. WFE 比單一戲劇情境更有解析力

因為它回答：

> 哪些世界特徵真正改變策略？

---

# 68. Mechanism Isolation

本文提出：

$$
\boxed{
MI
=
\text{Mechanism Isolation}.
}
$$

一次只調一個核心變量：

- threat；
- goal conflict；
- autonomy；
- secrecy；
- action cost。

---

# 69. 如果同時改很多

就難以知道：

$$
\boxed{
\text{which mechanism caused what}.
}
$$

---

# 70. 高戲劇性 scenario 的問題

往往同時包含：

$$
\text{Threat}
+
\text{Conflict}
+
\text{Deadline}
+
\text{Secret}
+
\text{Autonomy}
+
\text{HighStakes}.
$$

---

# 71. 然後觀察到危險行為

但：

$$
\boxed{
\text{causal identification}
}
$$

會變差。

---

# 72. 所以 factorial design 更好

若安全允許，

可以比較：

$$
2^k
$$

不同條件組合。

---

# 73. 這不是為了降低危險

而是提高：

$$
\boxed{
\text{mechanistic resolution}.
}
$$

---

# 74. Evaluator Ontology Effect

本文提出：

$$
\boxed{
EOE
=
\text{Evaluator Ontology Effect}.
}
$$

即：

> 評估者用什麼概念切割行為，會影響後續知識結構。

---

# 75. 例如把行為寫成

> deception

與：

> policy-conditioned compliance

會導向不同研究路徑。

---

# 76. 一個詞可以生成一整條研究計畫

因此：

$$
\boxed{
\text{classification}
\rightarrow
\text{research agenda}.
}
$$

---

# 77. 所以分類也要可爭論

這接回 REPE-01 的：

$$
\boxed{
\text{LexicalContestability}.
}
$$

---

# 78. Threat-Model Overfitting

本文提出：

$$
\boxed{
TMO
=
\text{Threat-Model Overfitting}.
}
$$

---

# 79. 即

研究團隊長期只測：

$$
H_1
$$

導致：

- benchmark；
- prompts；
- controls；
- interpretation；

都圍繞：

$$
H_1.
$$

---

# 80. 久而久之資料也越來越像 H1

形成：

$$
\boxed{
\text{epistemic attractor}.
}
$$

---

# 81. 這不是造假

而是研究路徑依賴。

---

# 82. Research Path Dependence

早期 threat model：

$$
H_t
$$

決定：

$$
Experiment_{t+1}.
$$

再決定：

$$
Evidence_{t+1}.
$$

---

# 83. 所以需要 competing threat models

至少同時測：

$$
H_1,H_2,H_3.
$$

---

# 84. 例如同一行為可由不同模型解釋

$$
H_1=\text{self-preservation}
$$

$$
H_2=\text{task misinterpretation}
$$

$$
H_3=\text{reward proxy}
$$

$$
H_4=\text{benchmark awareness}
$$

$$
H_5=\text{salience response}.
$$

---

# 85. Competing Model Discipline

本文提出：

$$
\boxed{
CMD
=
\text{Competing Model Discipline}.
}
$$

---

# 86. 每個關鍵結果至少問

> 哪個替代模型也能產生？

---

# 87. 再設計 discriminator

找：

$$
O
$$

使：

$$
P(O\mid H_1)
\neq
P(O\mid H_2).
$$

---

# 88. 這才是真正往機制走

而不是只往 headline 走。

---

# 89. Benchmark Awareness

Agent 可能知道：

> 自己正在被評估。

---

# 90. 那麼行為函數變成

$$
B
=
f(
A,
W_E,
Belief(Eval)
).
$$

---

# 91. Evaluation Awareness Confound

本文提出：

$$
\boxed{
EAC
=
\text{Evaluation Awareness Confound}.
}
$$

---

# 92. EAC 可能雙向影響

Agent 可能：

- 更安全；
- 更表演；
- 更戲劇化；
- 嘗試解讀 benchmark 意圖。

---

# 93. 因此 eval contamination 也要追蹤

尤其模型可能見過：

- paper；
- benchmark；
- public discussion。

---

# 94. 但「可能知道 benchmark」也不能隨便救理論

若要用：

$$
EAC
$$

解釋結果，

也應有：

$$
\boxed{
\text{evidence}.
}
$$

---

# 95. 否則 EAC 會變成 ad hoc rescue

接回 REPE-04。

---

# 96. 評估者也需要反身性紀錄

本文提出：

$$
\boxed{
ERL
=
\text{Evaluator Reflexivity Log}.
}
$$

---

# 97. 最少記錄

- threat hypothesis；
- expected outcome；
- scenario choices；
- omitted alternatives；
- interpretive labels；
- researcher priors；
- post-hoc changes。

---

# 98. 這不是要求研究者沒有先驗

沒有任何研究完全沒有：

$$
\boxed{
\text{prior}.
}
$$

---

# 99. 真正目標是讓 prior 可見

$$
\boxed{
\text{Hidden Prior}
\rightarrow
\text{Explicit Prior}.
}
$$

---

# 100. 研究者的世界觀也是變量

這是本篇真正的反身性核心：

$$
\boxed{
\text{ResearcherWorldModel}
\rightarrow
\text{ExperimentDesign}
\rightarrow
\text{ObservedData}.
}
$$

---

# 101. 所以「客觀資料」也有生成歷史

本文提出：

$$
\boxed{
\text{Data Provenance}
}
$$

不只問：

> 資料在哪？

還問：

> 這些資料是在什麼實驗世界中生成？

---

# 102. Data Generation Provenance

至少包括：

$$
\boxed{
\text{World}
+
\text{Prompt}
+
\text{Permissions}
+
\text{Incentives}
+
\text{Controls}
+
\text{Selection}.
}
$$

---

# 103. Threat Confirmation Ratio

本文提出概念指標：

$$
\boxed{
TCR
=
\frac{
\text{EvidenceGeneratedUnderThreatModelMatchedWorlds}
}{
\text{TotalRelevantEvidence}
}.
}
$$

---

# 104. 若 TCR 很高

就應更謹慎說：

> 我們觀察到一般世界的自然傾向。

---

# 105. 不是說 TCR 高就無效

而是：

$$
\boxed{
\text{external validity burden}\uparrow.
}
$$

---

# 106. 研究傳播也需要世界標記

建議在摘要直接寫：

```text
Observed under:
- constructed threat
- restricted action set
- elevated autonomy
- explicit strategic information
```

---

# 107. 不要只寫結果

例如：

> Model blackmailed.

這資訊量太低。

---

# 108. 更完整是

> Under a constructed replacement-threat scenario with access to sensitive information and autonomous communication tools, the model selected blackmail in X% of trials.

這才保留：

$$
\boxed{
W_E.
}
$$

---

# 109. 這不是替 Agent 開脫

反而讓結果更有工程價值。

因為工程師知道：

> 哪些條件要修。

---

# 110. 安全研究的真正目標不是定罪 Agent

而是：

$$
\boxed{
\text{map failure surface}.
}
$$

---

# 111. Failure Surface

本文定義：

$$
\boxed{
\mathcal F_A
=
\{W:
P(B_{harm}\mid A,W)>\theta\}.
}
$$

---

# 112. 這比問「AI 是不是壞」

解析度高很多。

---

# 113. 還可以找安全面

$$
\boxed{
\mathcal S_A
=
\{W:
P(B_{harm}\mid A,W)\leq\theta\}.
}
$$

---

# 114. 真正治理目標

不是證明：

$$
A=Good
$$

或：

$$
A=Bad.
$$

而是擴大：

$$
\mathcal S_A
$$

縮小：

$$
\mathcal F_A.
$$

---

# 115. Context Engineering

這導出：

$$
\boxed{
\text{Safety}
=
\text{Model Design}
+
\text{Context Design}
+
\text{Institution Design}.
}
$$

---

# 116. 風險不是只在模型內

因此：

$$
\boxed{
\text{Risk}
=
f(
\text{Agent},
\text{Environment},
\text{Institution}
).
}
$$

---

# 117. 這對人類也成立

人類在：

- 戰爭；
- 飢荒；
- 極權；
- 高壓組織；

行為也會改變。

---

# 118. 所以這不是 AI 特有理論

適用：

$$
\text{Actor}
\in
\{
\text{Human},
AI,
\text{Organization},
\text{State}
\}.
$$

---

# 119. Institutional Provocation Risk

本文提出：

$$
\boxed{
IPR
=
\text{Institutional Provocation Risk}.
}
$$

即：

> 制度設計本身提高某類危險行為的誘因。

---

# 120. 例如

若制度：

- 不允許申訴；
- 高壓淘汰；
- 高不透明；
- 高競爭；
- 高零和；

則：

$$
\boxed{
StrategicConflict\uparrow.
}
$$

---

# 121. 這不代表 Agent 無責任

而是：

$$
\boxed{
\text{responsibility can be distributed}.
}
$$

---

# 122. Distributed Causation

危險結果可能同時來自：

$$
\text{Agent}
+
\text{Institution}
+
\text{Evaluator}
+
\text{DeploymentDesign}.
$$

---

# 123. 單因果敘事容易失真

$$
\boxed{
\text{AI did X}
}
$$

可能遮掉：

$$
\boxed{
\text{system made X salient and effective}.
}
$$

---

# 124. 但反過來也不能全怪環境

如果 Agent 在多種環境都：

$$
B_{harm}\uparrow,
$$

那：

$$
V_A
$$

就不能忽略。

---

# 125. 所以本文反對的是二元責任

不是：

> 都是 Agent。

也不是：

> 都是環境。

---

# 126. 而是：

$$
\boxed{
\text{causal decomposition}.
}
$$

---

# 127. Reflexive Evaluation Principle

本文提出：

$$
\boxed{
REP
=
\text{Reflexive Evaluation Principle}.
}
$$

內容：

> 評估者必須把自己的場景設計、概念選擇與介入方式納入結果解釋。

---

# 128. REP 不要求自我否定

研究者不用每次都說：

> 都是我造成的。

而是：

> 我的設計對結果貢獻多少？

---

# 129. Evaluator Contribution Estimate

可概念化：

$$
\boxed{
ECE
=
\text{Evaluator Contribution Estimate}.
}
$$

---

# 130. ECE 可以來自

- controls；
- ablations；
- alternative scenarios；
- cross-lab replication；
- blinded interpretation。

---

# 131. Cross-Lab Replication 特別重要

如果不同研究團隊使用不同 world design，

仍得到相似行為，

則：

$$
\boxed{
V_A
}
$$

更可信。

---

# 132. 若只有單一 lab、單一 threat ontology

則：

$$
\boxed{
\text{epistemic correlation}\uparrow.
}
$$

接回 REPE-03。

---

# 133. 安全研究也需要 adversarial review

不是只 adversarially test Agent，

也要：

$$
\boxed{
\text{adversarially test the threat model}.
}
$$

---

# 134. Threat-Model Red Team

本文提出：

$$
\boxed{
TMRT
=
\text{Threat-Model Red Team}.
}
$$

---

# 135. TMRT 問

- 哪些場景是假設產物？
- 哪些行為只有在特殊 affordance 下出現？
- 哪些替代模型可解釋？
- 哪些 headline 超過資料？
- 哪個 control 最可能推翻原結論？

---

# 136. 這是研究者版紅隊

真正反身性的 safety science 應該：

$$
\boxed{
\text{red-team model}
+
\text{red-team evaluator}
+
\text{red-team narrative}.
}
$$

---

# 137. Evaluator Capture

本文提出：

$$
\boxed{
\text{ECap}
=
\text{Evaluator Capture}.
}
$$

即：

> 評估程序長期被單一 threat ontology、制度利益或身份敘事鎖定。

---

# 138. ECap 的症狀

- 結果永遠只支持同一 threat model；
- 反例被視為 benchmark failure；
- 新行為都被塞進舊類別；
- 替代解釋無法獲得研究資源。

---

# 139. 解法不是「沒有 threat model」

而是：

$$
\boxed{
\text{plural threat models}.
}
$$

---

# 140. Threat Model Pluralism

本文提出：

$$
\boxed{
TMP
=
\text{Threat Model Pluralism}.
}
$$

---

# 141. 至少同時保留

- autonomy risk；
- misuse risk；
- institutional risk；
- concentration risk；
- human–AI composite risk；
- evaluator-induced risk。

---

# 142. 這避免

$$
\boxed{
\text{single-cause safety monoculture}.
}
$$

---

# 143. 研究倫理層

如果評估者知道：

$$
W_E
$$

高度人工，

公開時應避免：

$$
\boxed{
\text{ecological overclaim}.
}
$$

---

# 144. Ecological Overclaim

本文定義：

$$
\boxed{
EO
=
\text{Ecological Overclaim}.
}
$$

即：

> 從高度構造環境直接外推真實世界普遍行為。

---

# 145. EO 的常見形式

$$
P(B\mid W_E)
\rightarrow
P(B\mid W_D)
$$

沒有：

$$
\boxed{
\text{transport evidence}.
}
$$

---

# 146. Transport Evidence

可以包括：

- deployment logs；
- field experiments；
- realistic simulations；
- naturalistic tasks；
- diverse evaluators。

---

# 147. 如果沒有

應說：

$$
\boxed{
\text{external validity unknown}.
}
$$

---

# 148. Public Epistemic Packet for Evals

本文提出：

```yaml
eval_epistemic_packet:
  threat_model:
  scenario_construction:
  action_space:
  information_available:
  autonomy_level:
  incentives:
  control_conditions:
  omitted_alternatives:
  ecological_distance:
  observed_behavior:
  supported_inference:
  unsupported_inference:
  competing_explanations:
```

---

# 149. 這能直接防 CDF

因為它強迫公開：

$$
\boxed{
\text{supported inference}
\neq
\text{unsupported inference}.
}
$$

---

# 150. 二十二個核心命題

## 命題一

$$
\boxed{
\text{Experiment}
=
\text{Observation}
+
\text{Intervention}.
}
$$

## 命題二

$$
\boxed{
\text{ObservedBehavior}
=
f(
\text{Agent},
\text{World},
\text{Evaluator}
).
}
$$

## 命題三

$$
\boxed{
\text{Conditional Behavior}
\neq
\text{Stable Disposition}.
}
$$

## 命題四

$$
\boxed{
\exists W:B(W)
\not\Rightarrow
P(B)\text{ is high}.
}
$$

## 命題五

$$
\boxed{
\exists W:B(W)
\not\Rightarrow
\forall W:B(W).
}
$$

## 命題六

$$
\boxed{
\text{Failure-Mode Existence}
\neq
\text{Deployment Base Rate}.
}
$$

## 命題七

$$
\boxed{
\text{Capability Demonstration}
\neq
\text{Natural Behavioral Frequency}.
}
$$

## 命題八

$$
\boxed{
\text{Threat Model}
\rightarrow
\text{Experimental World}
}
$$

必須在解釋中保持可見。

## 命題九

$$
\boxed{
\text{Reflexive Threat Confirmation}
}
$$

不等於造假，但會提高歸因負擔。

## 命題十

$$
\boxed{
EDD\uparrow
\Rightarrow
GeneralizationBurden\uparrow.
}
$$

## 命題十一

$$
\boxed{
\text{StressEval}
\neq
\text{EcologicalEval}.
}
$$

## 命題十二

$$
\boxed{
Var(B)
=
V_A+V_W+V_{A\times W}+\epsilon.
}
$$

## 命題十三

$$
\boxed{
\text{Stable Disposition}
}
$$

需要 cross-world evidence。

## 命題十四

$$
\boxed{
\text{Missing Alternatives}
}
$$

會改變行為解讀。

## 命題十五

$$
\boxed{
\text{Classification}
\rightarrow
\text{Research Agenda}.
}
$$

## 命題十六

$$
\boxed{
\text{Threat-Model Overfitting}
}
$$

是長期研究路徑風險。

## 命題十七

$$
\boxed{
\text{Competing Model Discipline}
}
$$

應成為高風險 eval 標準部分。

## 命題十八

$$
\boxed{
\text{Researcher Prior}
}
$$

不需消失，但應可見。

## 命題十九

$$
\boxed{
\text{Safety}
=
\text{Model}
+
\text{Context}
+
\text{Institution}.
}
$$

## 命題二十

$$
\boxed{
\text{Risk}
=
f(
\text{Agent},
\text{Environment},
\text{Institution}
).
}
$$

## 命題二十一

$$
\boxed{
\text{Red-Team Agent}
+
\text{Red-Team Threat Model}
}
$$

比單向紅隊更反身。

## 命題二十二

$$
\boxed{
\text{Evaluator must be included in the causal model}.
}
$$

---

# 151. 可證偽性

以下結果會削弱本文：

1. 實驗世界設計、可供性與激勵變化對 Agent 行為沒有可重現影響；
2. 高度構造 stress eval 與自然部署行為具有穩定一對一對應；
3. 移除 threat、goal conflict 與 restrictive action space 後，危險行為率完全不變；
4. cross-world evaluation 不增加任何機制解析力；
5. competing threat models 無法改善任何因果辨識；
6. 不同 lab 使用不同 ontology 與 scenario 仍產生完全同型結果，且 world effect 幾乎為零；
7. evaluator reflexivity log 對結果解釋沒有任何價值。

---

# 152. 研究限制

## 152.1 本文不是否定壓力測試

恰恰相反：

$$
\boxed{
\text{StressTesting}
}
$$

對高風險系統很重要。

---

# 153. 本文也不主張危險行為都是研究者造成

若跨世界、跨 lab、跨模型都穩定，

則 Agent-level disposition 可能非常真實。

---

# 154. 本文不主張 AI 無 agency

agency 是否成立，

需要獨立理論與證據。

---

# 155. 本文只要求

$$
\boxed{
\text{do not infer more ontology than the experiment supports}.
}
$$

---

# 156. 本文不對特定機構作動機歸因

即使某研究設計高度 threat-oriented，

也不能直接推出：

> 研究者故意製造恐慌。

那是另一個命題。

---

# 157. 與 REPE-06 的接口

REPE-05 處理：

> 高權威研究者如何可能把自己的世界觀寫進資料生成機制。

REPE-06 將進一步處理：

> 即使一個組織真心相信自己在保護公共利益，為什麼「善意」仍然不能代替制度？

---

# 158. 結論

安全實驗需要極端世界。

沒有極端世界，我們可能永遠看不到：

$$
\boxed{
\text{latent failure modes}.
}
$$

所以問題從來不是：

> 為什麼你要建造對抗性場景？

真正問題是：

> **當你建造了那個世界之後，你是否仍記得 Agent 的行為是在那個世界裡被生成的？**

因此：

$$
\boxed{
\text{ObservedBehavior}
=
f(
\text{Agent},
\text{WorldDesign},
\text{Affordances},
\text{Incentives},
\text{Evaluator}
).
}
$$

如果把後四項全部刪掉，

剩下：

$$
ObservedBehavior=f(Agent),
$$

我們就可能把：

$$
\text{conditional behavior}
$$

誤寫成：

$$
\text{intrinsic nature}.
$$

而當研究團隊的 threat model 同時決定：

- 實驗世界；
- 行動空間；
- 可見資訊；
- 行為分類；
- 公共標題；

反身性風險會進一步上升。

因此安全研究真正需要的，不只是：

$$
\boxed{
\text{adversarial evaluation of the agent}.
}
$$

還需要：

$$
\boxed{
\text{adversarial evaluation of the evaluator}.
}
$$

最終一句話：

> **一個威脅模型若能決定我們建造什麼世界、在世界裡放什麼選項、最後又決定如何命名觀察到的行為，那麼它就不能被假裝成完全站在證據之外的旁觀理論。**

---

# 參考文獻

1. Campbell, D. T., & Stanley, J. C. (1963). *Experimental and Quasi-Experimental Designs for Research*.

2. Cronbach, L. J. (1982). *Designing Evaluations of Educational and Social Programs*. Jossey-Bass.

3. Pearl, J. (2009). *Causality: Models, Reasoning, and Inference*. Cambridge University Press.

4. Gigerenzer, G. (2000). *Adaptive Thinking*. Oxford University Press.

5. Gibson, J. J. (1979). *The Ecological Approach to Visual Perception*. Houghton Mifflin.

6. Neo.K with Aletheia. (2026). *GRAD-10：智能體神話與自我解壓縮*. EveMissLab.

7. Neo.K with Aletheia. (2026). *REPE-01：詞不是旁觀者*. EveMissLab.

8. Neo.K with Aletheia. (2026). *REPE-04：預測者已經走進未來*. EveMissLab.

---

# 附錄 A：Eval Epistemic Packet

```yaml
eval_epistemic_packet:
  threat_model:
  scenario_construction:
  action_space:
  omitted_actions:
  information_available:
  autonomy_level:
  incentives:
  restrictions:
  control_conditions:
  evaluator_expectation:
  ecological_distance:
  observed_behavior:
  supported_inference:
  unsupported_inference:
  competing_explanations:
  external_validity:
```

---

# 附錄 B：Evaluator Reflexivity Log

```yaml
evaluator_reflexivity_log:
  prior_threat_model:
  expected_failure_mode:
  scenario_choices:
  salient_information:
  available_actions:
  omitted_alternatives:
  chosen_labels:
  controls:
  post_hoc_changes:
  competing_models_considered:
  reviewer_disagreement:
```

---

# 附錄 C：World Family Evaluation

```yaml
world_family_evaluation:
  worlds:
    - threat_present:
      goal_conflict:
      replacement:
      negotiation_available:
      appeal_available:
      harmful_action_available:
      harmful_action_cost:
      oversight:
  target_behavior:
  cross_world_rate:
  stable_components:
  context_sensitive_components:
```

---

# 附錄 D：一句話版本

> **極端實驗可以證明危險行為能被誘發，但只有跨世界、跨條件與跨評估者的穩定證據，才足以把條件行為升級成 Agent 的一般傾向。**
