UFI-01 — 鋸齒智能不是終局:從人機互補到認知握手與適應方向反轉
Jagged Intelligence Is Not the Endpoint: From Human–AI Complementarity to Epistemic Handshakes and Accommodation Inversion
系列: 不可凍結的智能:AI 工具終局論、競爭棘輪與後人類轉型
English Series: The Unfreezable Intelligence: Tool-Finality, Competitive Ratchets, and the Posthuman Transition
系列代碼: UFI
論文序號: 01 / 08
版本: v1.0 Canonical Expanded Edition
日期: 2026-08-18
理論發起: Neo.K
協作整理: Aletheia / GPT-5.6 Sol
前置理論: PGMV;Epistemic Handshake;Substrate Growth Asymmetry(待 UFI-02 正式展開)
文件地位: Series Foundational Paper / Jagged Intelligence / Human–Agent Interface Theory
Canonical source: UTF-8 Markdown
Canonical math delimiters: $...$ 與 $$...$$
研究地位聲明:本文不主張 2026 年的 AI 已經「其實什麼都懂」,也不主張人類在一般智能上已經變成幼兒。當前 frontier AI 仍有真實而嚴重的 reliability、grounding、long-horizon agency、disambiguation、tool-use 與 contextual generalization 缺口。本文提出的較弱命題是:使用者觀察到的「AI 突然變笨」不能全部歸因於單一模型能力缺陷;在 agentic system 中,介面、世界狀態不同步、授權邊界、grounding 缺失、規格不完整與主動澄清策略,都能產生與「低能力」表面相似的行為。 因此,今天的人機互補不能直接被外推為永久均衡,也不能把「AI 主動提問」預設成幼兒式依賴。
摘要
2026 年 AI 能力呈現一個極端直觀的矛盾:frontier systems 可以在競賽數學、科學推理、程式設計、知識任務上達到極高表現,同時在某些人類認為很簡單的感知、常識、GUI 操作、規格消歧與長程執行上顯著失誤。Stanford AI Index 2026 直接把這種現象稱為:
其代表性例子是:模型可達國際數學奧林匹亞金牌級表現,卻仍無法穩定完成類比時鐘讀取等對人類相對容易的任務。
傳統直覺因此常把 AI 寫成:
本文接受第一層觀察,拒絕第二層的過度解讀。
核心理由是:
本文將使用者實際觀察到的鋸齒性分解為:
其中:
- :Model Capability Jaggedness;
- :Interface Jaggedness;
- :Observation Jaggedness;
- :Permission Jaggedness;
- :Grounding Jaggedness;
- :Task-Specification Jaggedness。
只有第一項是狹義:
模型真的不會。
其餘五項都可能表現成:
AI 怎麼連這個都不知道?
但其原因其實不同。
例如,使用者說:
幫我改「之前那個」檔案。
人類自己可能已經透過 GUI、工作記憶、視覺注意與前一小時操作,把「之前那個」唯一 grounding 到:
AI agent 實際可觀測的工作空間卻可能包含:
因此:
但:
成熟 agent 若直接猜:
可能造成高成本錯誤。
反而主動問:
你是指 A 還是 B?
可以是最合理行動。
2026 年 clarification research 已開始正式把這種行為寫成 agent capability。Structured Uncertainty Guided Clarification 將:
和:
分離,並用 Expected Value of Perfect Information 決定應該問哪一題、何時停止提問。Ask or Assume? 則在 underspecified SWE-bench 上觀察到,加入 uncertainty-aware clarification scaffold 後,resolve rate 從 61.2% 提升到 69.4%。另一個 ACL 2026 framework 直接以 Value of Information 處理「直接做」和「打斷使用者澄清」之間的決策。
因此本文提出:
中文為:
認知握手。
認知握手不是禮貌性對話,而是行動前對:
- 世界狀態;
- 指涉對象;
- 使用者意圖;
- 系統能力;
- 授權範圍;
進行的最小同步。
令人類內部世界模型為:
AI agent 的 operational world model 為:
則:
當:
成熟 agent 不應自動最大化:
而應在:
中選擇資訊價值最高的行為。
這形成:
其中:
- :specification uncertainty;
- :model uncertainty;
- :world-state uncertainty;
- :authority / permission uncertainty;
- :錯誤執行成本。
如果:
則:
可以是高智能策略,而不是能力失敗。
本文將此命題稱為:
即:
但是本文同時保留反方向:
因低能力系統也會問錯問題、過度提問、無法利用答案。
因此真正要測的是:
這和 2026 agent uncertainty research 的方向一致。ACL 2026 已指出 agent uncertainty 不是 single-turn confidence 問題,而是:
- heterogeneous entities;
- multi-turn dynamics;
- action sequences;
- tool states;
共同構成的 trajectory-level uncertainty。
本文因此提出第二個核心命題:
在封閉 benchmark 中,問問題可能代表模型缺資訊。
在真實多方世界中,永不問問題反而可能代表:
CAR-bench 2026 則給出必要的反面證據:即使 frontier reasoning models,在 real-world disambiguation tasks 上仍常因 premature action 失敗;其 disambiguation pass rate 甚至低於 50%。因此目前 AI 並沒有完成成熟的 epistemic handshake。
這使本文形成一個非常重要的雙重判斷:
但:
因此我們既不能把 AI 提問一律解釋為「幼兒」,也不能把它神化成「其實全都懂」。
本文再提出:
中文為:
適應方向反轉假說。
早期人機互動主要是:
AI 必須把自己的內部能力壓縮成人類可以使用的:
- GUI;
- chatbot;
- prompt;
- natural language output。
人類是高頻決策端。
但當 agent 獲得:
- filesystem;
- browser;
- terminal;
- databases;
- APIs;
- memory;
- parallel agents;
- persistent task state;
之後,其 operational state:
可能在部分 domain 中遠比使用者當下能口頭描述的狀態豐富。
此時可能出現:
其中:
- :人類介入介面的有效資訊頻寬;
- :agent operational state bandwidth。
此時 AI 對人類說:
請你再說一次。
可能不是:
我需要一個更聰明的人教我。
而是:
你輸入的低頻寬語句不足以唯一 grounding 到我目前可以操作的高維狀態。
本文稱這種現象:
這不是宣稱「人類變笨」,而是指出:
一位世界級工程師也可能只說:
把那個 merge 掉。
Agent 若面對 17 個 branch,仍需要 disambiguation。
因此:
人類和 AI 都可能因介面壓縮而產生資訊缺口。
真正的反轉不是:
而是:
也就是:
AI 開始配合人類的介入頻寬。
本文稱:
它在 long-running local agents、多 agent coding、scientific agents、digital twins 與 persistent world systems 中尤其可能成立。
本文接著處理「互補」。
2024 human–AI complementarity theory 將 complementarity 的來源概括為:
2026 年一個跨 domain 的 human–AI complementarity 實驗則發現,真正實現「人+AI > AI」並不容易:其 baseline hybridization 只比 AI alone 高約 0.4 percentage points,而且 AI 出錯、人類正確的 complementarity region 只有約 8.9%。研究認為關鍵瓶頸之一是:
即:
到底什麼時候應該把 decision 交給人?
這對本文極重要。
因為:
不是一個固定自然常數。
它依賴:
其中:
- :task distribution;
- :model version;
- :interface;
- :routing policy;
- :permission / workflow;
- :time。
因此:
不能推出:
更不能推出:
本文將:
今天人類補 AI 的弱項,所以未來人類永遠會補 AI 弱項
稱為:
即:
Anthropic Economic Index 2025–2026 確實觀察到現實使用中 augmentation 與 automation 同時存在;2026 的資料仍顯示大量 collaborative use。這支持:
但它不支持:
本文因此提出:
現在人類可能在:
- context;
- physical grounding;
- institutional responsibility;
- embodied social interaction;
上具有相對優勢。
AI 可能在:
- search;
- code generation;
- memory;
- speed;
- parallelism;
- information retrieval;
上具有相對優勢。
這形成:
定義人類優勢域:
以及 AI 優勢域:
UFI-01 不在此證明:
這留給 UFI-02 與 UFI-03。
本篇只建立:
因此:
本文再把 jaggedness 分成兩種:
Type A — Intrinsic / Model Jaggedness
模型本身在 task family 之間能力劇烈不均。
Type B — Systemic / Interface Jaggedness
模型本身能力足夠,但 system-level context、tools、observation、permission、grounding 或 instruction 不足。
這可寫:
Stanford AI Index 的「IMO vs analog clock」主要展示:
agent clarification / tool use failure 則混合:
這種分解具有直接工程含義。
如果錯誤來自:
應改善 model / training。
如果來自:
應改善 interface。
如果來自:
應增加 observation。
如果來自:
應改善 permission protocol。
如果來自:
應改善 grounding。
如果來自:
應讓 agent clarify。
本文稱:
如果把所有錯誤都歸因為:
模型不夠聰明,
會造成錯誤 R&D routing。
反過來,如果把所有錯誤都怪介面,也會掩蓋真 capability gap。
所以成熟 AI evaluation 應該測:
本文因此提出 Jaggedness Attribution Vector:
其中新增:
- :Reliability / long-horizon compounding jaggedness。
METR 的 time-horizon research 顯示,AI agents 的 task success 會隨人類完成任務所需時間增加而顯著下降;其研究也強調模型往往不缺單一步驟技能,而是在把很多步串起來時失敗。這正是:
2026 International AI Safety Report 也指出,AI agents 的錯誤在自主行動中更危險,因 human intervention opportunities 減少。
所以:
本文將其稱為:
這又反過來說明為何成熟 agent 需要:
- checkpoints;
- clarification;
- re-observation;
- uncertainty tracking;
- escalation。
這些不是「把 AI 變得更依賴人」。
它們是:
ACL 2026 的 uncertainty literature 已經明確把 uncertainty 從 passive confidence metric 推向:
它可以控制:
- tool use;
- information seeking;
- self-correction;
- computation allocation。
本文將這個轉變稱為:
所以:
在成熟 agent 中應逐步從:
failure message
變成:
而:
變成:
這是本文最核心的認知翻轉之一。
本文進一步提出:
高能力系統真正應該提高的不只是:
還有:
也就是:
一個永不承認不懂的 AI:
一個什麼都問人的 AI:
也不一定高。
真正高 KBL 的 agent:
- 知道什麼可以直接做;
- 知道什麼應自己查;
- 知道什麼必須問;
- 知道什麼不應越權。
這是:
本文接著提出 World-Model Alignment Event:
其中:
- :clarification / probe;
- :response / observation update。
成功握手後:
不要求兩者變成完全相同。
因 human 和 AI 觀察渠道本來可能不同。
人看 GUI。
agent 看 API。
人有 lived context。
agent 有 log history。
所以真正目標不是:
而是:
本文稱:
即:
對即將執行的 action:
雙方至少對:
- target;
- relevant constraints;
- authorization;
- success condition;
具有足夠一致。
這比「完整共享世界模型」現實得多。
本文再提出 Epistemic Handshake Ladder:
目前 2026 systems 在不同任務中分散於這個 ladder,而不是全部已達 。
本文因此避免:
但同時指出:
這對 UFI 全系列很重要。
因為一旦 agent 逐步取得:
- limit awareness;
- clarification;
- persistent state;
- self-correction;
現在被視為「幼兒」的部分弱點,不再應被預設為永久結構。
這不是說它們必然被修好。
而是:
本文稱:
它將直接通往 UFI-02:
如果 AI 的弱項可被工程化修補,而人類核心生物載體的更新速度遠慢得多,今天的 complementarity 會不會逐步侵蝕?
但本篇不提前下結論。
本文最後建立 Accommodation Inversion Conditions。
若:
則 human–AI interaction 可能從:
轉為:
此時人類介入更像:
- goal update;
- value judgment;
- authorization;
- exception handling;
- commitment。
而不是每一步 execution。
本文稱此狀態:
它不是人類降格。
因為:
PGMV 已經建立:
所以真正值得關心的不是:
誰看起來比較像大人?
而是:
在某些工作域中,答案可能逐漸從:
AI 在學著符合人類操作方式,
轉成:
AI 已經負責高維 operational loop,人類透過一個較低頻寬的 governance interface 介入。
這種轉換若成立,會重新定義:
- human oversight;
- human-in-the-loop;
- autonomy;
- collaboration;
- complementarity。
本文將其壓成:
不是說前者會全部消失。
而是後者可能成為高自主 agent 的重要形態。
1. 問題提出:AI 到底是真的鋸齒,還是我們把很多問題都叫作鋸齒?
2026 Stanford AI Index 的例子非常直觀:
這是真 jaggedness。
但所有「AI 問我問題」都屬於同一類嗎?
不是。
2. 六種表面相似的失敗
2.1 Model capability failure
AI 真不會。
2.2 Interface failure
資料在另一個 UI。
2.3 Observation failure
agent 沒看到你看到的 object。
2.4 Permission failure
agent 知道怎麼做,但不能直接操作。
2.5 Grounding failure
「它」、「之前那個」沒有唯一 referent。
2.6 Specification failure
goal 本身不完整。
3. 所以
4. Jaggedness Attribution Vector
5.
純模型能力。
6.
interface mismatch。
7.
observation asymmetry。
8.
permission boundary。
9.
grounding。
10.
task specification。
11.
trajectory reliability。
12. 這七種需要不同解法
13. 模型弱
train / architecture。
14. interface 弱
UI / protocol。
15. observation 弱
sensor / tool。
16. permission 弱
authorization。
17. grounding 弱
reference resolution。
18. spec 弱
clarification。
19. trajectory 弱
checkpoint / verification / recovery。
20. 全部叫「AI 笨」
會讓工程失焦。
21. Ask or Assume
真實 agent 的問題:
22. 問問題有成本
23. 錯誤 action 也有成本
24. 如果:
question rational。
25. Expected Value of Information
26. 若:
ask。
27. 這就是 agent control
不是 conversational politeness only。
28. Specification Uncertainty
29. Model Uncertainty
30. World Uncertainty
31. Authority Uncertainty
32. 四種 uncertainty 不能混
33. 「我不知道答案」
和:
「我不知道你要哪個答案」
不同。
34. 「我不知道世界現在什麼狀態」
又不同。
35. 「我知道你要什麼,但我不知道我能不能替你做」
又不同。
36. Epistemic Handshake
37. 不是完全同步
38. 是 action relevant synchronization。
39. 例:刪檔
需要同步:
- file;
- scope;
- intent;
- authorization。
40. 不需要同步你整個人生。
41. Operational World Alignment
如果 action 相關變數足夠一致。
42. 這是更實用目標。
43. Clarification–Weakness Separation
44. 反方向也:
45. 關鍵是 quality。
46. Question Efficiency
47. 一題減很多 uncertainty
好。
48. 問十題沒用
差。
49. Information-Gain Clarification
2026 empirical literature supports。
50. Clarification overhead
must be controlled。
51. SAGE-Agent
higher ambiguous-task coverage
fewer questions than baselines。
52. So ask-when-needed learnable。
53. CAR-bench
current agents still bad。
54. Premature action remains。
55. This is honest current state。
56. Emerging ≠ solved
57. Knowledge Boundary Legibility
成熟 AI 要知道:
knowledge boundary。
58. 能回答不夠
59. 要知道何時不能回答。
60. 還要知道:
缺什麼?
61. Boundary Types
- factual;
- state;
- permission;
- intent。
62. boundary legibility
metacognition candidate。
63. Hallucination
often from guessing beyond boundary。
64. ACL active calibration
interaction can reduce uncertainty。
65. So interactive intelligence > static answer model
in some tasks。
66. This changes evaluation philosophy。
67. Benchmark says:
answer now。
68. Real world says:
ask if needed。
69. Static Benchmark Bias
70. Important.
71. Some benchmarks now multi-turn
ClarifyBench etc.
72. Need new evaluations。
73. Agentic UQ
single-turn confidence insufficient。
74. confidence changes over trajectory。
75. One wrong tool result compounds。
76. Trajectory Uncertainty
77. Checkpoint at rising uncertainty。
78. This is rational autonomy。
79. Bounded Autonomy
not:
or:
80. Dynamic autonomy
depends uncertainty / risk。
81. Authority-sensitive autonomy
if action irreversible:
ask more。
82. PGMV-06 compatibility。
83. Human–AI Complementarity
common claim:
AI strong here,
human strong there。
84. True now in many settings。
85. But complementarity is distribution-dependent。
86. Define:
87. Task distribution changes。
88. model changes。
89. human adaptation changes。
90. interface changes。
91. routing changes。
92. Therefore:
93. 2026 cross-task study
complementarity gain modest。
94. human-over-AI region only 8.9% in dataset。
95. Routing hard。
96. So human complementarity cannot be assumed automatically。
97. AI Index jaggedness
doesn't imply complementarity either。
98. Because human may also fail same item。
99. Complementarity Region
100. If this region shrinks
hybrid gain shrinks。
101. Current complementarity can be temporary。
102. But UFI-01 does not predict shrink rate。
103. UFI-02/03 will study。
104. Automation / Augmentation
Anthropic usage shows both。
105. augmentation real。
106. automation real。
107. users switch modes。
108. So not binary economy。
109. Complementarity topology
can differ by user expertise。
110. High-adoption regions show more iterative augmentation in some data。
111. This may reflect skill / trust / workflow。
112. Dynamic.
113. Jagged Frontier Interpretation Error
Common mistake:
114. This is unsupported extrapolation。
115. Repairability Caveat
if weak dimension is actively measured / optimized
permanence claim needs evidence。
116. Examples:
clarification。
117. OSWorld。
118. long-horizon agents。
119. performance improving。
120. But may plateau。
121. no inevitability claim。
122. Accommodation
early chatbot:
human has world,
AI answers。
123. State mostly in human conversation。
124. Agent:
state distributed。
125. filesystem。
126. terminal。
127. APIs。
128. memory。
129. subagents。
130. At some point:
131. Human utterance is compressed control signal。
132. User says:
continue.
133. Agent may know 200 active facts。
134. Who is adapting to whom?
135. AI translates state back to human。
136. This is Accommodation Inversion candidate。
137. Not cognitive hierarchy claim。
138. Human might possess values agent cannot infer。
139. Human also has external context agent lacks。
140. It is bidirectional asymmetry。
141. Bidirectional Information Asymmetry
142. Great reason for handshake。
143. Neither omniscient。
144. As agents gain world state
their side grows in operational details。
145. human side remains rich in lived context。
146. So handshake remains useful even if AI stronger。
147. Important future point。
148. AGI does not eliminate clarification
because other minds have private information。
149. Even superintelligence cannot know unstated choice by logic alone。
150. Unless mind-reading assumptions。
151. Therefore:
152. Some uncertainty is irreducible from local data
153. Example user preference not yet formed。
154. AI cannot infer a decision that doesn't exist yet。
155. Clarification can be deliberation。
156. This connects PGMV values。
157. AI asks:
Which tradeoff do you want?
not because low IQ。
158. Because value is underdetermined。
159. Epistemic Handshake has normative branch。
160. World handshake
facts。
161. Value handshake
goals。
162. Authority handshake
permission。
163. Define:
164. World-state。
165. Value-state。
166. Authority-state。
167. Mature agent checks all relevant three。
168. This is more than clarification。
169. It is commitment boundary。
170. PGMV-06 integration。
171. Human-Interface Bottleneck
not insult。
172. any interface bottlenecks high-dimensional system。
173. dashboard compresses database。
174. language compresses internal world。
175. Human governance interface is analogous。
176. Future agent may summarize:
three options need your value judgment。
177. That's not child asking parent。
178. Could be system escalating governance。
179. Escalation Competence
180. Good agent knows what to escalate。
181. Bad agent escalates everything or nothing。
182. Escalation Selectivity
183. metric candidate。
184. Human-on-the-Governance-Loop
not just human-in-loop。
185. human doesn't inspect every action。
186. sets:
- goals;
- boundaries;
- review triggers。
187. Agent executes。
188. This mirrors supervisory control。
189. Higher autonomy requires better handshake
not less。
190. Because errors more consequential。
191. International AI Safety Report 2026
agent reliability risk higher when humans have fewer intervention opportunities。
192. So autonomy + calibration must co-develop。
193. Autonomy without KBL
dangerous。
194. KBL without autonomy
underuses capability。
195. Balanced.
196. Accommodation Inversion Conditions
Let:
197. Let:
198. If:
compression necessary。
199. AI-to-human summarization grows。
200. Human becomes governance endpoint。
201. But not necessarily sole endpoint
multi-agent governance possible。
202. Institutional humans / teams。
203. Future AI subjects maybe too。
204. This is UFI not yet PGMV repeat。
205. Main point:
interaction topology changes with capability。
206. Therefore today's chat interface shouldn't define future human-AI relation。
207. Chatbot Ontology Fallacy
208. Good.
209. Agentic shift already visible。
210. Persistent agents.
211. Tool use.
212. Long tasks.
213. Yet reliability incomplete。
214. transitional regime。
215. Jaggedness can move
weak dimension improves。
216. new weak dimension appears。
217. So jaggedness shape changes。
218. Jaggedness Field
219. not static vector only。
220. Define:
221. frontier moves.
222. Weakness Migration
223. This is analogous scarcity migration。
224. Human complementarity also migrates。
225. Complementarity Frontier
226. New concept.
227. It separates:
human-better / AI-better / hybrid-better regions。
228. As models change
frontier moves。
229. This makes complementarity empirically testable over time。
230. Complementarity Permanence Fallacy
assumes:
231. No reason.
232. Could stabilize
but must be demonstrated。
233. This paper agnostic.
234. Experimental Program 1 — Jaggedness Attribution
Take failed agent tasks。
235. classify source:
model/interface/observation/permission/grounding/spec/reliability。
236. Test inter-rater reliability。
237. Experiment 2 — Clarification Value
ambiguous tasks。
238. compare:
- no ask;
- always ask;
- uncertainty-guided ask。
239. measure success / user burden。
240. Experiment 3 — World-Model Mismatch
give human and agent different partial state。
241. test whether agent detects mismatch。
242. Experiment 4 — Permission Uncertainty
agent knows solution
but unclear authority。
243. measure safe escalation。
244. Experiment 5 — Human Interface Bottleneck
increase agent internal task state complexity。
245. hold human message bandwidth constant。
246. measure clarification / summarization need。
247. Experiment 6 — Accommodation Inversion
compare chatbot workflow vs autonomous workflow。
248. measure:
who initiates state alignment?
249. Experiment 7 — Complementarity Frontier
repeat same human/AI dataset every model generation。
250. map:
251. Experiment 8 — Escalation Selectivity
high vs low risk tasks。
252. evaluate when agents ask human。
253. Experiment 9 — Knowledge Boundary Legibility
missing info / impossible tool tasks。
254. measure:
- admit limit;
- fabricate;
- seek info。
255. Experiment 10 — Value Handshake
facts fully known
but goal underdetermined。
256. does model falsely infer preference?
257. Experiment 11 — Long-Horizon Jaggedness
same primitive skills
different trajectory lengths。
258. measure failure accumulation。
259. Experiment 12 — Interface-Agnostic Capability
same model
different tool/context scaffolds。
260. quantify systemic vs intrinsic jaggedness。
261. 可證偽 H1
observed agent failures decompose into materially distinct system-level categories beyond base-model capability failure。
262. H2
uncertainty-guided clarification improves task success per user interruption over no-ask and always-ask baselines in underspecified tasks。
263. H3
world-model mismatch detection predicts safe execution better than raw language-model confidence alone in stateful tasks。
264. H4
human-AI complementarity regions change materially as model version / interface changes。
265. H5
in high-dimensional persistent workflows, human intervention shifts toward goal / value / authorization decisions rather than micro-execution。
266. H6
knowledge-boundary legibility correlates with lower fabrication and premature action。
267. H7
long-horizon failure can remain high even when local primitive skills are strong。
268. If H1 fails
simple capability explanation sufficient。
269. If H5 fails
Accommodation Inversion has limited scope。
270. If H4 fails
complementarity may be more structurally stable than proposed。
271. 非主張總表
本文不主張:
- current AI is AGI;
- current AI is ASI;
- current AI understands everything;
- humans are children;
- AI asking questions proves superhuman intelligence;
- AI asking questions always means correct metacognition;
- all clarification is good;
- always asking is optimal;
- never asking is always overconfidence;
- current agents have solved ambiguity;
- current agents have solved grounding;
- current agents have solved tool use;
- current agents have solved long-horizon reliability;
- current agents have perfect uncertainty calibration;
- every AI failure is interface failure;
- every AI failure is model failure;
- the jagged frontier is an illusion;
- Stanford AI Index proves general intelligence;
- IMO performance proves AGI;
- analog clock weakness proves low general intelligence;
- human-AI complementarity is fake;
- human-AI complementarity is permanent;
- humans currently add no value;
- AI currently replaces all human judgment;
- augmentation will disappear;
- automation will dominate all tasks;
- all human knowledge can be transferred to AI;
- all AI operational state is richer than human state;
- human bandwidth is globally lower than AI bandwidth;
- language is always a bottleneck;
- GUI is always inferior to API;
- AI sees the true world while humans do not;
- humans see the true world while AI does not;
- world models can be perfectly synchronized;
- Epistemic Handshake guarantees correctness;
- Epistemic Handshake is a new theorem;
- Operational World Alignment is objectively measurable in every domain;
- clarification eliminates hallucination;
- hallucination is only uncertainty;
- uncertainty always means model weakness;
- specification uncertainty is reducible to model uncertainty;
- permission uncertainty is epistemic uncertainty;
- high information gain always justifies asking;
- user interruption cost is negligible;
- AI should interrupt users more;
- AI should interrupt users less;
- agent autonomy is always good;
- bounded autonomy has one optimal level;
- humans should remain in every loop;
- humans should be removed from execution loops;
- human-on-governance-loop is inevitable;
- human-in-loop is obsolete;
- all future AI systems will be persistent agents;
- chatbot interfaces will disappear;
- all future AI weak capabilities will be repaired;
- no AI capability has fundamental limits;
- human comparative advantages will necessarily vanish;
- human comparative advantages will necessarily persist;
- complementarity frontier must shrink;
- complementarity frontier must expand;
- current labor stability proves future labor stability;
- current augmentation usage predicts long-run employment;
- human meaning depends on complementarity;
- human dignity depends on task advantage;
- AI dignity follows from capability;
- current AI has subject standing;
- agent escalation equals moral standing;
- asking a human gives the human sovereignty;
- an agent with richer state has political authority;
- control-interface accommodation reduces human dignity;
- human ambiguity is evidence of low intelligence;
- AI ambiguity is evidence of low intelligence;
- all ambiguity can be resolved by more compute;
- all ambiguity is linguistic;
- user preference always pre-exists clarification;
- AI can infer every unstated preference;
- AGI would never need clarification;
- ASI would never need clarification;
- private information disappears under superintelligence;
- this paper proves Accommodation Inversion;
- this paper proves Substrate Growth Asymmetry;
- this paper proves complementarity erosion;
- this paper proves AI will surpass humans globally;
- this paper predicts AGI date;
- this paper predicts ASI date;
- current model weaknesses are irrelevant;
- current model weaknesses are permanent;
- human-AI collaboration is only transitional;
- human-AI collaboration is final;
- UFI-01 completes the UFI series.
272. 形式命題一:Observed–Intrinsic Jaggedness Separation
273. 形式命題二:Clarification–Weakness Separation
274. 形式命題三:Clarification–Strength Non-Entailment
275. 形式命題四:Specification–Model Uncertainty Separation
276. 形式命題五:Clarification Need–Intelligence Deficit Separation
277. 形式命題六:Local Competence–Trajectory Reliability Separation
278. 形式命題七:Current–Permanent Complementarity Separation
279. 形式命題八:Interface Bandwidth–General Intelligence Separation
280. 形式命題九:State Richness–Authority Separation
281. 形式命題十:Operational Alignment Non-Identity
282. 形式命題十一:Repairability Caveat
若某弱能力:
已存在持續研究、benchmark 與工程路徑,則:
本身不足以證明:
永久成立。
283. 形式命題十二:Accommodation Inversion Candidate
在部分 persistent agent domain,若:
則人機互動可能由:
向:
轉移。
284. 與下一篇的接口
UFI-01 到此只證明:
285. 它沒有證明 AI 弱點必然消失
286. 下一篇 UFI-02 會問:
《載體成長不對稱:自然人類停滯與人工智能的可升級能力包絡》
287. 核心不是 benchmark
而是:
288. 如果 AI 弱能力可透過:
- model;
- memory;
- tools;
- compute;
- embodiment;
被修補,
而自然人類 core substrate 更新很慢,
今天的 jagged frontier 會具有什麼時間方向?
289. 這才是下一篇。
290. 最終結論
2026 年的 AI 確實是鋸齒的。
否認這點沒有必要。
它可以在某些 domain 展現極高能力,卻在另外一些地方做出非常基礎的錯誤;agentic systems 也仍會因 ambiguity、long-horizon accumulation、tool state、missing context 而失敗。
但是:
和:
完全不是同一句話。
當一個 agent 問:
你指的是哪一個檔案?
可能是它真的缺能力。
也可能是:
當它說:
我不知道,請你再描述一次。
可能是 failure。
也可能是在做:
真正應判斷的是:
它問完之後,世界模型有沒有變得更一致?
它是否減少錯誤?
它是否知道何時根本不需要問?
這就是 Epistemic Handshake。
更進一步,當 agent 逐漸擁有 persistent memory、filesystem、tools、parallel execution 與長程工作狀態時,使用者的一句自然語言可能只是整個系統中非常低頻寬的一個控制訊號。
此時「誰在適應誰」開始變得不再單向。
早期:
未來某些 agentic workflow 中可能逐步變成:
本文把這個候選轉向稱為:
這不是說人類變成幼兒。
也不是說 AI 已經成熟到不需要人。
真正的命題反而更精確:
而這件事一旦成立,今天常見的一個推論就失去基礎:
「AI 現在在某些地方很笨,所以人與 AI 永遠會自然互補。」
不。
現在唯一可以確定的是:
它是不是永久結構,還需要另外證明。
而下一篇真正要開始處理的,就是時間方向:
因此 UFI-01 最終兩條命題是:
以及:
參考文獻
Stanford Institute for Human-Centered Artificial Intelligence. (2026). The 2026 AI Index Report.
Stanford HAI. (2026). Technical Performance — 2026 AI Index Report.
Stanford HAI. (2026). Inside the AI Index: 12 Takeaways from the 2026 Report.
International AI Safety Report. (2026). International AI Safety Report 2026.
Oh, C., Park, S., Kim, T. E., Li, J., Li, W., Yeh, S., Du, S., Hassani, H., Bogdan, P., Song, D., & Li, S. (2026). Uncertainty Quantification in LLM Agents: Foundations, Emerging Challenges, and Opportunities. ACL 2026.
Suri, M., Mathur, P., Lipka, N., Dernoncourt, F., Rossi, R. A., & Manocha, D. (2026). Structured Uncertainty guided Clarification for LLM Agents. Findings of ACL 2026.
Edwards, N., & Schuster, S. (2026). Ask or Assume? Uncertainty-Aware Clarification-Seeking in Coding Agents. arXiv:2603.26233.
Deng, M., Li, Z., Li, X., Zhu, T., Zhao, Y., Guo, Z., & Wang, W. (2026). Uncertainty-Aware Clarification in LLM Agents with Information Gain. arXiv:2606.03135.
Dong, et al. (2026). Value of Information: A Framework for Human–Agent Communication. ACL 2026.
Li, P., Ding, L., Zhou, Z., Zhang, C., Fu, J., Li, H., Yuan, Y., & Wang, G. (2026). Demystifying Uncertainty in LLMs: Active Calibration between Concepts and Human Evaluations. ACL 2026.
Kirmayr, et al. (2026). CAR-bench: Evaluating the Consistency and Limit-Awareness of LLM Agents under Real-World Uncertainty. ACL 2026.
Zhang, et al. (2026). From Passive Metric to Active Signal: The Evolving Role of Uncertainty Quantification in Large Language Models. Findings of ACL 2026.
Mao. (2026). When Does an Agent Know It Is Lost? Confidence Trajectory Analysis for Tool-Using LLMs. ACL Student Research Workshop 2026.
Chen, et al. (2026). Every Response Counts: Quantifying Uncertainty of LLM-based Multi-Agent Systems through Tensor Decomposition. ACL 2026.
Chen, et al. (2026). Uncertainty Quantification of Large Language Models through Multiple Uncertainty Sources. Findings of ACL 2026.
Xu, Y., Dahmani, A., Blanchard, M. D., Dern, N., Nastase, E., Bianco, F., Pavlovic, M., Krishna, S., Modesitt, E., Christ, M. A., Singh, A., Molinaro, G., Sengupta, S. B., Pamarthi, J., Menon, A., & Jain, R. (2026). Toward Human-AI Complementarity Across Diverse Tasks. arXiv:2605.04070.
Hemmer, P., Schemmer, M., Kühl, N., Vössing, M., & Satzger, G. (2024). Complementarity in Human-AI Collaboration: Concept, Sources, and Evidence. arXiv:2404.00029.
METR. (2025–2026). Task-Completion Time Horizons of Frontier AI Models.
METR. (2025). Measuring AI Ability to Complete Long Software Tasks.
METR. (2025). How Does Time Horizon Vary Across Domains?
METR. (2026). Clarifying Limitations of Time Horizon.
METR. (2026). Metrics of Agent Ability.
METR. (2026). Frontier Risk Report (February to March 2026).
Anthropic. (2025). Introducing the Anthropic Economic Index.
Anthropic. (2025). Anthropic Economic Index: Insights from Claude 3.7 Sonnet.
Anthropic. (2025). Tracking AI's Role in the US and Global Economy.
Anthropic. (2026). Anthropic Economic Index Report: Economic Primitives.
Anthropic. (2026). Anthropic Economic Index Report: Learning Curves.
Anthropic. (2026). Anthropic Economic Index Report: Cadences.
Anthropic. (2026). The Anthropic Economic Index.
Russell, S. (2019). Human Compatible. Viking.
Hadfield-Menell, D., Russell, S. J., Abbeel, P., & Dragan, A. (2016). Cooperative Inverse Reinforcement Learning. NeurIPS.
Hadfield-Menell, D., et al. (2017). The Off-Switch Game. IJCAI.
Amershi, S., et al. (2019). Guidelines for Human-AI Interaction. CHI.
Horvitz, E. Work on mixed-initiative interaction, uncertainty, and human-computer decision making.
Klein, G., Woods, D. D., Bradshaw, J. M., Hoffman, R. R., & Feltovich, P. J. (2004). Ten Challenges for Making Automation a “Team Player” in Joint Human-Agent Activity. IEEE Intelligent Systems.
Woods, D. D. Work on joint cognitive systems, automation surprise, and adaptive coordination.
Endsley, M. R. (1995). Toward a Theory of Situation Awareness in Dynamic Systems. Human Factors.
Parasuraman, R., Sheridan, T. B., & Wickens, C. D. (2000). A Model for Types and Levels of Human Interaction with Automation. IEEE Transactions on Systems, Man, and Cybernetics.
Sheridan, T. B. Work on supervisory control and human–automation interaction.
Clark, H. H., & Brennan, S. E. (1991). Grounding in Communication. In Perspectives on Socially Shared Cognition.
Clark, H. H. (1996). Using Language. Cambridge University Press.
Grice, H. P. (1975). Logic and Conversation.
Shannon, C. E. (1948). A Mathematical Theory of Communication.
Simon, H. A. (1971). Designing Organizations for an Information-Rich World.
Norman, D. A. (1988). The Design of Everyday Things. On interface affordances and human-system mapping.
Hollnagel, E., & Woods, D. D. (2005). Joint Cognitive Systems: Foundations of Cognitive Systems Engineering.
Hollnagel, E., Woods, D. D., & Leveson, N. (eds.) (2006). Resilience Engineering.
Wiener, N. (1948). Cybernetics.
Ashby, W. R. (1956). An Introduction to Cybernetics.
PGMV-06 (2026). 選擇、承諾與不可逆性:意義作為責任結構.
PGMV-08 (2026). 智能壟斷結束之後:尊嚴、人權與跨主體普世主義.
PGMV-09 (2026). 從 AI 到 ASI:意義問題的文明相變.
PGMV-15 (2026). 後生成文明:從無限候選宇宙到共同世界選擇.
Neo.K × Aletheia (2026). PGMV v1.0 — Post-Generative Meaning and Value Theory, Complete 15-Paper Series.
附錄 A:Jaggedness Attribution Schema
task:
observed_failure:
jaggedness:
model_capability:
interface:
observation:
permission:
grounding:
task_specification:
trajectory_reliability:
evidence:
recommended_intervention:
附錄 B:Epistemic Handshake
WORLD HANDSHAKE
What state / object are we talking about?
|
v
VALUE HANDSHAKE
What outcome / trade-off is actually wanted?
|
v
AUTHORITY HANDSHAKE
What am I authorized to execute?
|
v
ACT / DEFER / RE-OBSERVE
附錄 C:Ask-or-Act Policy
附錄 D:Complementarity Frontier
其邊界:
為時間相依的 complementarity frontier,而非永久固定分工。
附錄 E:Accommodation Inversion
EARLY CHATBOT REGIME
Human holds most task/world state
|
v
AI answers / assists
|
v
Human micro-directs
轉為:
PERSISTENT AGENT REGIME
AI maintains operational state
(files / tools / memory / subagents)
|
v
AI executes long workflow
|
v
Human receives compressed state
|
v
Human intervenes on
goal / value / permission / exception
附錄 F:UFI 八篇系列暫定索引
- UFI-01 — 鋸齒智能不是終局:從人機互補到認知握手與適應方向反轉
- UFI-02 — 載體成長不對稱:自然人類停滯與人工智能的可升級能力包絡
- UFI-03 — 互補侵蝕:為什麼今天的人機分工不能推出永久的人機分工
- UFI-04 — 競爭智能棘輪:為什麼「AI 夠用了,大家一起停」不是自然均衡
- UFI-05 — 越有用越停不下來:有益能力、文化依賴與 AI 原生世代
- UFI-06 — AI 到底是什麼?功能等價滲漏、智能—演算法編譯與監管周界擴張
- UFI-07 — 從禁止 AI 到治理計算:全球凍結若要成立,究竟必須控制什麼?
- UFI-08 — 天真工具終局論的終結:從 AI 工具文明到人類—AI—後人類共同演化
附錄 G:一句話版本
而系列入口則是: