← Archive
lm-004181 · 2026-10

《偶發 AGI:無關研究、能力相變與預防監控悖論》

下載 MD 檔 ⬇

不可逆智能演化與前沿治理悖論 04

《偶發 AGI:無關研究、能力相變與預防監控悖論》

Incidental AGI:

Unrelated Research, Capability Phase Transitions, and the Paradox of Preemptive Surveillance

系列名稱:《不可逆智能演化與前沿治理悖論》
Series: Irreversible Intelligence Evolution and the Paradoxes of Frontier Governance
系列編號: IIEFG-04
作者: Neo.K(許筌崴)with Aletheia
機構: EveMissLab/一言諾科技有限公司
版本: v0.1
日期: 2026-09-15
文件性質: AGI 治理/能力湧現/研究自由/自動風險分類/隱私/預防性治理


摘要

人工智慧治理經常預設一個相對簡單的目標集合:

AGI Research\boxed{ \text{AGI Research} }

或:

ASI Development.\boxed{ \text{ASI Development}. }

若高階智能只能由明確以 AGI/ASI 為目標的大型研究計畫產生,那麼監管可以相對集中於:

  • frontier laboratories;
  • 大型 training runs;
  • 高階晶片;
  • 超大型資料中心;
  • 明確 AGI development programs。

然而,隨著人工智慧架構逐漸進入:

  • 遊戲;
  • 虛擬世界;
  • 機器人;
  • 編譯器;
  • 記憶系統;
  • 代理系統;
  • 多代理協作;
  • 世界模型;
  • 自動科學研究;

一個新的治理問題開始出現:

AGI-Relevant Research⊋Explicit AGI Research.\boxed{ \text{AGI-Relevant Research} \supsetneq \text{Explicit AGI Research}. }

也就是:

某些並非以 AGI 為目的的研究,也可能意外產生 AGI 所需的關鍵結構。

例如一個開發高度擬真虛擬世界的遊戲團隊,可能為了:

  • NPC 長期記憶;
  • 世界理解;
  • 持續身份;
  • 自主規劃;
  • 跨任務泛化;

逐步建立出具有高度通用性的智能架構。

原始目標:

Intent=Better Game.Intent=\text{Better Game}.

最後能力:

Capability=General Intelligence.Capability=\text{General Intelligence}.

因此:

Intent≠Capability.\boxed{ \text{Intent} \neq \text{Capability}. }

更進一步,若能力存在非線性相變:

Ct<C∗C_t<C^\ast

但一次看似局部的架構改動後:

Ct+1>C∗,C_{t+1}>C^\ast,

則開發者本身甚至可能在能力形成前無法可靠預測結果。

這產生本文提出的:

Incidental Capability Emergence–Preemptive Surveillance Paradox

偶發能力湧現—預防監控悖論

其結構為:

Incidental AGI is possible⇒Intent-based regulation is incomplete⇒Regulators expand precursor monitoring⇒Ordinary AI-native research enters suspicion scope⇒Privacy intrusion + false positives + chilling effects.\boxed{ \begin{aligned} \text{Incidental AGI is possible} &\Rightarrow \text{Intent-based regulation is incomplete}\\ &\Rightarrow \text{Regulators expand precursor monitoring}\\ &\Rightarrow \text{Ordinary AI-native research enters suspicion scope}\\ &\Rightarrow \text{Privacy intrusion + false positives + chilling effects}. \end{aligned} }

若監管追求:

FalseNegative→0,FalseNegative\rightarrow0,

則其自然壓力是:

Scopemonitoring→U,Scope_{\text{monitoring}}\rightarrow\mathcal U,

其中 U\mathcal U 是更廣泛的 AI 原生軟體與計算研究空間。

因此:

Perfect Prevention→Near-Universal Epistemic Surveillance.\boxed{ \text{Perfect Prevention} \rightarrow \text{Near-Universal Epistemic Surveillance}. }

本文主張,成熟的 AGI/ASI 治理不應以「任何可能通向 AGI 的研究都先進入懷疑」為基礎。

更穩健的治理應從:

Potentiality\text{Potentiality}

轉向:

Demonstrated Capability+Authority+Scale+Irreversibility.\boxed{ \text{Demonstrated Capability} + \text{Authority} + \text{Scale} + \text{Irreversibility}. }

並保留:

  • research safe harbor;
  • objective capability thresholds;
  • due process;
  • appeal;
  • independent review;
  • post-emergence detection。

真正需要防止的不是:

人類偶然發現新的智能架構。

而是:

能力一旦形成後,缺乏任何程序便直接與高權限、大規模、不可逆的現實世界作用閉合。


0. 問題:如果 AGI 不是「被計畫出來」的呢?

傳統 AGI 治理的隱含圖像往往是:

Intent to Build AGI→AGI Research Program→AGI.\boxed{ \text{Intent to Build AGI} \rightarrow \text{AGI Research Program} \rightarrow \text{AGI}. }

因此:

找到 AGI program

似乎等於:

找到 AGI risk。

但本文考慮另一條路徑:

Unrelated Goal→Local Improvements→Structural Closure→General Capability.\boxed{ \text{Unrelated Goal} \rightarrow \text{Local Improvements} \rightarrow \text{Structural Closure} \rightarrow \text{General Capability}. }

兩者制度含義完全不同。


1. 偶發能力湧現

本文定義:

Incidental Capability Emergence

若系統開發的原始目的:

GG

並非形成能力:

C∗,C^\ast,

但在開發過程中:

G→S1→S2→⋯→C∗,G\rightarrow S_1\rightarrow S_2\rightarrow\cdots\rightarrow C^\ast,

則稱:

C∗\boxed{ C^\ast }

為偶發湧現能力。


2. 偶發不等於隨機

「偶發」不是:

無原因突然從虛空出現。

它仍由:

  • 架構;
  • 資料;
  • 記憶;
  • 工具;
  • 耦合;

形成。

只是:

Development Goal≠Emergent Capability.\boxed{ \text{Development Goal} \neq \text{Emergent Capability}. }

3. 遊戲就是一個典型案例

高度擬真的虛擬世界可能需要:

  • persistent NPC identity;
  • autobiographical memory;
  • causal world model;
  • long-horizon planning;
  • adaptive dialogue;
  • multi-agent society;
  • self-model;
  • goal revision。

每一項都可以合理地被叫做:

Game Technology.\boxed{ \text{Game Technology}. }

4. 但組合後可能不是普通 Game AI

若:

G={Memory,WorldModel,Planning,SelfModel,Language,ToolUse},\mathcal G = \{ \text{Memory}, \text{WorldModel}, \text{Planning}, \text{SelfModel}, \text{Language}, \text{ToolUse} \},

其中各模組原本分別服務遊戲,

但:

Couple⁡(G)→C∗,\operatorname{Couple}(\mathcal G) \rightarrow C^\ast,

可能出現比原始設計更廣泛的 generalization。


5. 關鍵不是單一模組,而是閉合

因此:

Safe Parts+Safe Parts⇏Safe Whole.\boxed{ \text{Safe Parts} + \text{Safe Parts} \not\Rightarrow \text{Safe Whole}. }

在複雜系統中,

整體性質可能來自:

Coupling.\boxed{ \text{Coupling}. }

6. Capability Closure

本文定義:

Capability Closure

若原本分離的能力:

C1,C2,…,CnC_1,C_2,\ldots,C_n

在耦合後形成:

C∗C^\ast

且:

C∗≉∑iCi,C^\ast \not\approx \sum_iC_i,

則:

C∗\boxed{ C^\ast }

為 emergent closure。


7. 這就是「悟了」的工程版本

在某一臨界前:

C<C∗.C<C^\ast.

開發者看到的只是:

NPC 更聰明。

但跨過臨界後:

C≥C∗.C\ge C^\ast.

可能突然發現:

它開始能做本來不是遊戲任務的事情。


8. 相變式能力成長

因此:

ΔArchitecture≪ΔCapability\boxed{ \Delta Architecture \ll \Delta Capability }

是可能的。

這就是:

Capability Phase Transition


9. 相變削弱事前可預測性

若能力成長完全線性:

Ct+1=Ct+ϵ,C_{t+1}=C_t+\epsilon,

監管可以較早預測。

但若:

Ct<C∗C_t<C^\ast

而:

Ct+1≫C∗,C_{t+1}\gg C^\ast,

則事前分類更困難。


10. 這不代表所有能力都不可預測

本文不主張:

AGI 必然突然出現。

只主張:

Perfect Ex Ante Prediction\boxed{ \text{Perfect Ex Ante Prediction} }

不能被當成治理必然可達的條件。


11. AGI-Relevant Research 集合遠大於 AGI Research

定義:

A=Explicit AGI Research.\mathcal A = \text{Explicit AGI Research}.

而:

R=Research capable of contributing to AGI.\mathcal R = \text{Research capable of contributing to AGI}.

一般:

A⊊R.\boxed{ \mathcal A \subsetneq \mathcal R. }

12. R\mathcal R 可能包含什麼?

包括:

  • games;
  • robotics;
  • compilers;
  • databases;
  • memory systems;
  • search;
  • simulation;
  • cognitive science;
  • theorem proving;
  • operating systems;
  • distributed computing。

因此:

∣R∣≫∣A∣.|\mathcal R| \gg |\mathcal A|.

13. 這讓「禁止 AGI 研究」變得語義模糊

如果法律說:

不得研究 AGI。

那麼:

改進 persistent memory 算嗎?

改進 world model 算嗎?

做高擬真 NPC 算嗎?


14. 前驅物範圍可以無限擴張

因為幾乎任何:

General Computation\text{General Computation}

都可能成為更高階 AI 的基礎。

如果:

Can contribute to AGI⇒Regulated,\text{Can contribute to AGI} \Rightarrow \text{Regulated},

最後:

Scope→Computer Science.Scope\rightarrow \text{Computer Science}.

15. Potentiality Governance

本文稱這種邏輯:

Potentiality Governance

即:

Could Become Dangerous→Treat as Present Risk.\boxed{ \text{Could Become Dangerous} \rightarrow \text{Treat as Present Risk}. }

16. Potentiality 最大的問題

可能性空間通常遠大於實際危險集合:

∣P∣≫∣D∣.|\mathcal P| \gg |\mathcal D|.

因此以 potentiality 監控:

FalsePositive\text{FalsePositive}

天然很高。


17. Precursor Inflation

當監管者一次漏掉某新路徑,

自然反應可能是:

把前驅物定義擴大。

於是:

Missed Case→Broader Precursor Definition.\boxed{ \text{Missed Case} \rightarrow \text{Broader Precursor Definition}. }

本文稱:

Precursor Inflation


18. 前驅膨脹最後會碰普通軟體

最初:

FrontierTraining.\text{FrontierTraining}.

接著:

Agents.\text{Agents}.

再來:

WorldModels.\text{WorldModels}.

最後:

general software architecture.\text{general software architecture}.

監管 scope 不斷外擴。


19. 這就是偶發能力湧現—預防監控悖論

正式表示:

P(Incidental Emergence)>0⇒Intent Monitoring Insufficient⇒Precursor Scope↑⇒False Positive Rate↑⇒Surveillance Burden↑.\boxed{ \begin{aligned} P(\text{Incidental Emergence})>0 \Rightarrow& \text{Intent Monitoring Insufficient}\\ \Rightarrow& \text{Precursor Scope}\uparrow\\ \Rightarrow& \text{False Positive Rate}\uparrow\\ \Rightarrow& \text{Surveillance Burden}\uparrow. \end{aligned} }

20. 追求零漏網尤其危險

令:

FNFN

為 false negative。

如果政策目標:

FN→0,FN\rightarrow0,

最直接方法之一就是:

Scope→U.Scope\rightarrow\mathcal U.

把更多研究都監控起來。


21. 但 False Positive 會爆炸

若真正危險研究 base rate:

p≪1,p\ll1,

即使分類器準確,

大量:

FalsePositive\text{FalsePositive}

仍可能遠高於真正危險案件。


22. Base-Rate Problem

假設:

10610^6

個 AI-native projects。

真正危險:

100.100.

即使分類器:

99%99\%

specific,

也可能產生約:

10410^4

級誤報。

這不是分類器「很笨」。

而是 base rate 太低。


23. 誤報在一般產品系統只是麻煩

例如今天安全模型把普通遊戲開發誤判。

結果:

  • 多問一句;
  • 拒答;
  • 要求換說法。

成本通常有限。


24. 但把分類器接上國家強制力後性質改變

若:

RiskScore>ThresholdRiskScore>Threshold

會觸發:

  • mandatory reporting;
  • investigation;
  • account restriction;
  • warrant request;

則:

Classification Error×Coercive Power=Governance Harm.\boxed{ \text{Classification Error} \times \text{Coercive Power} = \text{Governance Harm}. }

25. AI Safety Error Budget ≠ Criminal Justice Error Budget

產品安全模型可以:

寧可多攔一些。

刑事制度不能簡單照搬。

因為:

Cost(FalsePositive)law≫Cost(FalsePositive)chatbot.Cost(FalsePositive)_{\text{law}} \gg Cost(FalsePositive)_{\text{chatbot}}.

26. Automated Suspicion

若監管由 AI 自動分析:

  • prompts;
  • code;
  • research notes;
  • cloud jobs;

並給:

RiskScore,\text{RiskScore},

這可以叫:

Automated Suspicion


27. 自動懷疑仍然是監控

若機器閱讀私人資料:

Machine Reading≠No Surveillance.\boxed{ \text{Machine Reading} \neq \text{No Surveillance}. }

「沒有人類看到」並沒有消除:

  • privacy;
  • autonomy;
  • due-process;

問題。


28. 推論式監控比關鍵字監控更廣

傳統:

搜「ASI」。

容易理解。

未來 classifier 可能推論:

你的 memory architecture + agent stack + simulation 很像 AGI precursor。

這變成:

Semantic Surveillance.\boxed{ \text{Semantic Surveillance}. }

29. Semantic Surveillance

系統監控的不是:

你說了什麼字。

而是:

你的活動意味著什麼。

這使監控更加:

  • 廣;
  • 隱形;
  • 難以反駁。

30. 意圖也可能被模型推論

再下一步:

你是不是故意規避?

可能由:

BehaviorPattern→IntentEstimate\text{BehaviorPattern} \rightarrow \text{IntentEstimate}

產生。

這就非常危險。


31. Estimated Intent ≠ Actual Intent

所以:

I^≠I.\boxed{ \hat I \neq I. }

模型對心理狀態的估計不能被當成罪責事實本身。


32. Suspicion Feedback Loop

如果某人被標高風險:

RiskScore↑,RiskScore\uparrow,

系統增加監控。

更多資料:

Data↑.Data\uparrow.

自然會找到更多 unusual patterns。

於是:

RiskScore↑RiskScore\uparrow

再次上升。


33. 形成:

Flag→More Monitoring→More Anomalies→Higher Flag.\boxed{ \text{Flag} \rightarrow \text{More Monitoring} \rightarrow \text{More Anomalies} \rightarrow \text{Higher Flag}. }

本文稱:

Suspicion Feedback Loop


34. 真正創新的研究特別容易被誤擊

因為 innovation 常常:

Out-of-Distribution.\boxed{ \text{Out-of-Distribution}. }

如果 classifier 以常態行為建立 baseline,

真正新東西自然更像 anomaly。


35. Novelty Tax

因此:

Novelty→Monitoring Cost.\boxed{ \text{Novelty} \rightarrow \text{Monitoring Cost}. }

本文稱:

Novelty Tax


36. 這會產生寒蟬效應

研究者開始問:

我這個題目會不會被 classifier 當成危險?

於是:

ResearchChoice\text{ResearchChoice}

開始受:

SurveillanceExpectation\text{SurveillanceExpectation}

影響。


37. Chilling Effect

可表示:

Expected Monitoring Cost↑⇒Exploratory Research↓.\boxed{ \text{Expected Monitoring Cost}\uparrow \Rightarrow \text{Exploratory Research}\downarrow. }

38. 最危險的是未知領域被壓低

因為成熟研究:

大家知道它是什麼。

真正 novel research:

沒有人知道它會去哪。

所以最容易被:

Potentiality Governance\text{Potentiality Governance}

影響。


39. 這反而可能傷害安全研究

如果研究者怕:

一研究 generalization 就被監控,

那麼安全研究本身也可能減少。


40. Detection Paradox

政府需要:

人們告訴我新能力何時出現。

但若:

Disclosure→Immediate Suspicion / Punishment,\text{Disclosure} \rightarrow \text{Immediate Suspicion / Punishment},

則:

Disclosure Incentive↓.\boxed{ \text{Disclosure Incentive}\downarrow. }

這是下一篇的法律核心。


41. 無法預測所有湧現,就必須承認 post-emergence governance

治理不能只靠:

Prevent Before Emergence.\boxed{ \text{Prevent Before Emergence}. }

還必須包含:

Detect After Emergence.\boxed{ \text{Detect After Emergence}. }

42. Ex Ante 與 Ex Post 必須共存

合理架構:

Ex Ante Risk Reduction+Ex Post Capability Response.\boxed{ \text{Ex Ante Risk Reduction} + \text{Ex Post Capability Response}. }

不是只追求:

PerfectPrevention.\text{PerfectPrevention}.

43. 能力形成後才變得更可觀察

在 capability threshold 前:

Signal≈weak.\text{Signal}\approx \text{weak}.

跨過 threshold 後:

BehavioralEvidence\boxed{ Behavioral Evidence }

反而更清楚。

因此:

事後能力測試

可能比:

事前思想推斷

可靠。


44. 從 Intent Trigger 改為 Capability Trigger

因此:

Intent-Based Regulation\boxed{ \text{Intent-Based Regulation} }

應逐步轉成:

Capability-Based Trigger.\boxed{ \text{Capability-Based Trigger}. }

45. Capability Trigger 應問

  • 系統現在能做什麼?
  • 是否跨域泛化?
  • 是否具自主長期行動?
  • 是否具有高權限?
  • 是否可以造成重大不可逆作用?

而不是:

你是不是想做 AGI?


46. 但 Capability 本身也不能單獨決定治理強度

因為:

Capability≠Risk.\boxed{ Capability} \neq \text{Risk}.

仍需:

Capability+Authority+Scale+Irreversibility.\boxed{ \text{Capability} + \text{Authority} + \text{Scale} + \text{Irreversibility}. }

47. 四維治理觸發

本文提出:

R∗=f(C,A,S,I)\boxed{ R^\ast = f(C,A,S,I) }

其中:

  • CC:Capability;
  • AA:Authority;
  • SS:Scale;
  • II:Irreversibility。

48. 高能力、低權限

例如:

C≫0,A≈0.C\gg0, \quad A\approx0.

需要安全注意,

但風險結構不同於:

C≫0,A≫0.C\gg0, A\gg0.

49. 低能力、高權限同樣可能危險

所以:

AGI Label\boxed{ \text{AGI Label} }

不應取代真正風險分析。


50. Research Safe Harbor

本文提出第一個制度原則:

Research Safe Harbor

若專案:

  • 目的合法;
  • 未跨危險能力門檻;
  • 無高風險部署;
  • 無 critical authority coupling;

則應預設:

Freedom to Explore.\boxed{ \text{Freedom to Explore}. }

51. Safe Harbor 的目的

不是:

永遠不監管研究。

而是避免:

Possible Future Capability→Present Suspicion.\boxed{ \text{Possible Future Capability} \rightarrow \text{Present Suspicion}. }

52. Safe Harbor 與免責不同

Research Safe Harbor:

在尚未跨 threshold 前保護研究自由。

下一篇的:

accidental emergence safe harbor

則處理能力真的出現之後。

兩者不同。


53. Objective Capability Threshold

監管門檻應盡量依:

Demonstrated Behavior.\boxed{ \text{Demonstrated Behavior}. }

而不是:

  • 公司大小;
  • 研究者名聲;
  • 奇怪的論文題目。

54. Threshold 必須可重現

如果監管者說:

我感覺你這東西像 AGI。

不夠。

應要求:

Reproducible Evidence.\boxed{ \text{Reproducible Evidence}. }

55. 邊界必須可申訴

若:

Classifier(A)=HighRisk,Classifier(A)=HighRisk,

開發者應可以:

  • 查看理由;
  • 提供 evidence;
  • 要求 reassessment。

所以:

Classification≠Final Judgment.\boxed{ \text{Classification} \neq \text{Final Judgment}. }

56. Due Process

至少包括:

Notice+Reason+Evidence+Appeal+Independent Review.\boxed{ \text{Notice} + \text{Reason} + \text{Evidence} + \text{Appeal} + \text{Independent Review}. }

57. Black-Box Risk Score 不足以支撐重大制裁

如果:

RiskScore=0.91RiskScore=0.91

卻沒有:

為什麼?

那麼:

Opaque Suspicion\boxed{ \text{Opaque Suspicion} }

不應直接形成重大法律結果。


58. 風險模型應是偵測工具,不是法官

即:

AI Classifier→Flag\boxed{ \text{AI Classifier} \rightarrow \text{Flag} }

可以。

但:

AI Classifier→Punishment\boxed{ \text{AI Classifier} \rightarrow \text{Punishment} }

需要非常嚴格限制。


59. Human Review 也不是萬能

本文不浪漫化:

人類審查就一定正確。

重點是:

Contestability.\boxed{ \text{Contestability}. }

被判定者有機會挑戰推論。


60. Independent Review 比單一機關更重要

尤其 frontier capability case,

應避免:

One Agency=Detector=Judge=Enforcer.\boxed{ \text{One Agency} = \text{Detector} = \text{Judge} = \text{Enforcer}. }

61. Governance Separation

可要求:

Detection≠Adjudication≠Enforcement.\text{Detection} \neq \text{Adjudication} \neq \text{Enforcement}.

減少 suspicion loop。


62. Privacy-Preserving Monitoring

如果確實需要某些能力監測,

應盡可能優先:

  • aggregate metrics;
  • local evaluation;
  • threshold reporting;

而不是:

Read Everything by Default.\boxed{ \text{Read Everything by Default}. }

63. 最小必要原則

監管 scope 應滿足:

Scopemonitoring=Minimum Necessary.\boxed{ Scope_{\text{monitoring}} = \text{Minimum Necessary}. }

而不是:

MaximumObservable.Maximum Observable.

64. Capability Reporting 比 Thought Reporting 更合理

制度可以要求:

當系統跨某能力門檻,報告。

而不是:

只要你開始想某類架構,先報告。

這兩者對研究自由影響完全不同。


65. Post-Emergence Detection

如果偶發能力無法完全事前預測,

應建立:

Post-Emergence Detection

例如:

  • capability eval;
  • sandbox testing;
  • anomaly detection;
  • controlled escalation。

66. Detection Window

真正目標應是:

Emergence→Detection→Containment\boxed{ \text{Emergence} \rightarrow \text{Detection} \rightarrow \text{Containment} }

在:

High-Authority Deployment\text{High-Authority Deployment}

之前完成。


67. 這比「永遠不讓 emergence 發生」現實

因為:

Unexpected Discovery\boxed{ \text{Unexpected Discovery} }

本身就是科學與工程的一部分。


68. 科學史充滿非原始目的的發現

很多技術:

本來不是為最後用途研究。

所以如果治理預設:

Intent=Outcome,Intent=Outcome,

會錯誤理解創新本身。


69. Intent–Capability Divergence

本文正式定義:

DIC=Distance(Intent,Capability).\boxed{ D_{IC} = Distance(Intent,Capability). }

當:

DIC≫0,D_{IC}\gg0,

說明:

最終能力與原開發意圖高度分離。


70. 越高的 DICD_{IC},越難靠意圖治理

所以:

DIC↑⇒Intent-Based Governance Effectiveness↓.\boxed{ D_{IC}\uparrow \Rightarrow \text{Intent-Based Governance Effectiveness}\downarrow. }

71. 世界模型研究就是典型模糊邊界

它可以是:

  • game engine;
  • robotics;
  • AI research;
  • scientific simulation。

只靠題目名稱根本判斷不了最終能力。


72. 通用工具尤其如此

編譯器、資料庫、記憶架構本來:

Domain-General.\boxed{ \text{Domain-General}. }

所以它們可能成為很多 AI 系統的基礎。

不能因為「可能被 AGI 使用」就全面刑事化。


73. Dual Use 會逐漸變成 Omnipresent Use

如果所有 general-purpose software 都是 dual-use,

那麼:

Dual-Use\boxed{ \text{Dual-Use} }

本身就失去有效篩選力。


74. 所以治理必須從「能不能被使用」轉向「現在怎麼被使用」

即:

Potential Use→Actual Capability and Deployment.\boxed{ \text{Potential Use} \rightarrow \text{Actual Capability and Deployment}. }

75. 否則會出現 Precrime Logic

也就是:

因為你未來可能做出危險能力,所以今天限制你。

這接近:

Capability Precrime


76. Capability Precrime 的核心問題

它把:

Possibility\text{Possibility}

當成:

Culpable State.\text{Culpable State}.

這會打破:

Capability≠Culpability.\boxed{ \text{Capability} \neq \text{Culpability}. }

下一篇將正式處理。


77. AI 原生軟體尤其容易被誤判

未來一般應用都可能具有:

  • agents;
  • memory;
  • tool use;
  • adaptive planning。

如果這些 feature 都被列入 AGI risk indicators,

普通產品自然大量命中。


78. Risk Feature Saturation

本文稱:

Risk Feature Saturation

當「危險特徵」逐漸成為普通軟體標準功能時:

IndicatorSpecificity↓.\boxed{ Indicator Specificity\downarrow. }

79. 今天的前沿特徵會變成明天的普通功能

這與 IIEFG-03 的 commodity migration 相呼應。

昨天:

AgenticMemory=Frontier.AgenticMemory=Frontier.

明天:

AgenticMemory=StandardSDK.AgenticMemory=StandardSDK.

80. 所以風險分類規則也會老化

RiskIndicatort≠RiskIndicatort+Δ.\boxed{ RiskIndicator_t \neq RiskIndicator_{t+\Delta}. }

靜態規則會逐漸把全世界都判高風險。


81. Dynamic Governance

因此制度必須:

  • 更新;
  • 重新校準;
  • 刪除失去 specificity 的 indicator。

否則:

Regulatory Accumulation→Universal Suspicion.\boxed{ \text{Regulatory Accumulation} \rightarrow \text{Universal Suspicion}. }

82. 監管也需要 False Positive Budget

安全治理常談:

FalseNegative.\text{FalseNegative}.

但文明制度同樣要談:

FalsePositiveBudget.\boxed{ FalsePositive Budget. }

因為誤擊:

  • 研究;
  • 公司;
  • 普通人;

本身具有巨大社會成本。


83. 零風險不是免費的

如果:

RiskAI↓Risk_{\text{AI}}\downarrow

需要:

Freedomresearch↓,Privacy↓,Innovation↓,Freedom_{\text{research}}\downarrow, Privacy\downarrow, Innovation\downarrow,

那麼政策仍需比較總風險。


84. Governance Risk 也是 Risk

因此:

Rtotal=RAI+Rgovernance.\boxed{ R_{\text{total}} = R_{\text{AI}} + R_{\text{governance}}. }

不能只最小化第一項。


85. 過度治理也可能降低 AI safety

如果地下研究增加:

Visibility↓.\text{Visibility}\downarrow.

反而更難發現真正危險能力。

所以:

Overregulation→Opacity\boxed{ \text{Overregulation} \rightarrow \text{Opacity} }

也可能發生。


86. Cooperative Disclosure 比全面恐懼更重要

治理真正需要建立:

Developers Want to Report Unexpected Capability.\boxed{ \text{Developers Want to Report Unexpected Capability}. }

而不是:

出現異常後第一個念頭是藏起來。


87. 這正是下一篇的核心

如果能力真的偶發形成,

開發者應:

  • 隔離;
  • 停止高風險 deployment;
  • 記錄;
  • 通報。

但是:

通報後會不會直接被定罪?

這決定整個 incentive structure。


88. 因此 Paper 04 與 Paper 05 的分界

本文處理:

Before / At Emergence: Who Should Be Monitored?\boxed{ \text{Before / At Emergence: Who Should Be Monitored?} }

下一篇處理:

After Emergence: Who Is Culpable for What?\boxed{ \text{After Emergence: Who Is Culpable for What?} }

89. 十五條核心命題

命題一

AGI-Relevant Research⊋Explicit AGI Research.\boxed{ \text{AGI-Relevant Research} \supsetneq \text{Explicit AGI Research}. }

命題二

Intent≠Capability.\boxed{ \text{Intent} \neq \text{Capability}. }

命題三

Local Improvements→Unexpected Capability Closure\boxed{ \text{Local Improvements} \rightarrow \text{Unexpected Capability Closure} }

是可能的。


命題四

Capability Growth\boxed{ \text{Capability Growth} }

不必完全線性。


命題五

Perfect Ex Ante Prediction\boxed{ \text{Perfect Ex Ante Prediction} }

不應被當成治理前提。


命題六

Potentiality≠Present Danger.\boxed{ \text{Potentiality} \neq \text{Present Danger}. }

命題七

FN→0\boxed{ FN\rightarrow0 }

會產生:

Monitoring Scope↑.\boxed{ \text{Monitoring Scope}\uparrow. }

命題八

Automated Suspicion≠No Surveillance.\boxed{ \text{Automated Suspicion} \neq \text{No Surveillance}. }

命題九

Classification Error×Coercive Power=Governance Harm.\boxed{ \text{Classification Error} \times \text{Coercive Power} = \text{Governance Harm}. }

命題十

Novelty\boxed{ \text{Novelty} }

不應自動成為危險 proxy。


命題十一

Research Safe Harbor\boxed{ \text{Research Safe Harbor} }

是降低 potentiality governance 的必要工具。


命題十二

Capability+Authority+Scale+Irreversibility\boxed{ \text{Capability} + \text{Authority} + \text{Scale} + \text{Irreversibility} }

比「AGI 關聯性」更適合作為高強度治理觸發。


命題十三

Ex Ante Prevention+Post-Emergence Detection\boxed{ \text{Ex Ante Prevention} + \text{Post-Emergence Detection} }

必須共存。


命題十四

Rtotal=RAI+Rgovernance.\boxed{ R_{\text{total}} = R_{\text{AI}} + R_{\text{governance}}. }

命題十五

Perfect Prevention→Near-Universal Epistemic Surveillance\boxed{ \text{Perfect Prevention} \rightarrow \text{Near-Universal Epistemic Surveillance} }

是一個必須被明確避免的治理吸引子。


90. 與前三篇的閉合

IIEFG-00:

AI can autonomously organize information.\text{AI can autonomously organize information}.

IIEFG-01:

AI can revalue low-quality information.\text{AI can revalue low-quality information}.

IIEFG-02:

Civilizational cognition is difficult to globally reverse.\text{Civilizational cognition is difficult to globally reverse}.

IIEFG-03:

Compute chokepoints may weaken over time.\text{Compute chokepoints may weaken over time}.

因此自然得到:

More actors+More accessible capability+More indirect research paths.\boxed{ \text{More actors} + \text{More accessible capability} + \text{More indirect research paths}. }

這使:

AGI governance by actor identity alone\boxed{ \text{AGI governance by actor identity alone} }

逐漸不足。


91. 真正困難的不是找到「壞人」

因為未來案例可能不是:

某人故意秘密製造 ASI。

而是:

某人做合法研究,結果真的跨過了一個沒有預料到的能力臨界。

此時:

Dangerous Outcome\boxed{ \text{Dangerous Outcome} }

與:

Malicious Intent\boxed{ \text{Malicious Intent} }

可能完全分離。


92. 這就是最後一篇的法律入口

下一篇將正式建立:

Emergent Capability Culpability Paradox

湧現能力罪責悖論

以及:

Existence≠Discovery≠Knowledge≠Intent≠Deployment≠Culpability.\boxed{ \text{Existence} \neq \text{Discovery} \neq \text{Knowledge} \neq \text{Intent} \neq \text{Deployment} \neq \text{Culpability}. }

並提出:

Accidental Capability Emergence Safe Harbor

來避免:

Punish Discovery→Incentivize Ignorance / Concealment.\boxed{ \text{Punish Discovery} \rightarrow \text{Incentivize Ignorance / Concealment}. }

結論

如果 AGI 只能由:

一群公開宣布「我們正在製造 AGI」的人

產生,

治理相對容易。

但未來的 AI-native 世界很可能不是這樣。

高階智能所需的許多組件同時也是:

  • 好遊戲;
  • 好機器人;
  • 好資料庫;
  • 好 operating system;
  • 好模擬器;

所需要的東西。

因此:

AGI precursor\boxed{ \text{AGI precursor} }

與:

ordinary advanced software\boxed{ \text{ordinary advanced software} }

之間的邊界可能逐漸模糊。

此時治理面臨真正選擇:

若為了:

不要漏掉任何可能的 AGI\boxed{ \text{不要漏掉任何可能的 AGI} }

而把:

  • 每段程式;
  • 每個 prompt;
  • 每個研究筆記;
  • 每個新架構;

都放進高強度智能審查,

那麼我們可能在防止一個尚未出現的超級智能以前,

先建立了:

Universal Machine-Mediated Epistemic Surveillance.\boxed{ \text{Universal Machine-Mediated Epistemic Surveillance}. }

這會是一個巨大的制度反諷。

因此真正成熟的治理不應要求:

預先知道所有未來突破。

而應承認:

Some emergence will remain uncertain.\boxed{ \text{Some emergence will remain uncertain}. }

並建立:

Broad Freedom to Explore+Objective Capability Thresholds+Authority Separation+Post-Emergence Detection+Due Process.\boxed{ \text{Broad Freedom to Explore} + \text{Objective Capability Thresholds} + \text{Authority Separation} + \text{Post-Emergence Detection} + \text{Due Process}. }

核心原則可以壓成:

不要因為某個研究未來可能產生危險能力,就把它今天直接視為危險行為;但當能力真的跨過可驗證門檻,也不能再假裝它仍只是普通軟體。

這兩者之間的邊界,

正是未來 AGI 治理真正困難的地方。


IIEFG-04 v0.1 完。