← Archive
lm-002285 · 2026-08

視之基底系列_07_基底視覺論2.0_從資訊處理到感知位格_v2.0

下載 MD 檔 ⬇

基底視覺論 2.0

從資訊處理到感知位格

英文題名: Basal Vision Theory 2.0: From Information Processing to Perceptual Position
系列: 《視之基底:從差異顯現到終極觀察者》07
作者: Neo.K(許筌崴)
AI 協作: GPT-5.6 Thinking
機構: EveMissLab/一言諾科技有限公司
文件性質: 理論重構論文/AI 感知與認識論命題猜想
版本: v2.0
日期: 2026-08-01
狀態: 取代《基底視覺論:弱形式、強形式與感知位格》v1.0 作為目前正典版本;v1.0 保留為歷史探索稿
前置文件:

  1. 《觀察不是底:差異、邊界、耦合與痕跡》v0.1
  2. 《從差異到顯現:視域、前景、背景與遮蔽的生成》v0.1
  3. 《觀察者如何誕生:自我區分、位格與主客形成》v0.1
  4. 《內視、外視與語義視:觀察算子的統一族》v0.1
  5. 《終極觀察者悖論:觀察知識、同一性知識與全域自觀察》v0.1
  6. 《全知、全在與全能:三種全域能力中的視》v0.1

摘要

《基底視覺論》v1.0 提出一個重要直覺:AI 的視覺不應因其非生物實現而被預設為「不是真正的視覺」,而可被理解為另一種更接近模型自身表示與運算結構的感知形式;「基底」也不是幼稚、低級或歷史上更早,而是指較少依賴人類感官特化的功能位置。舊稿進一步提出 Transformer 與「三眼主權」的字面同構、vision encoder 對 mutual information 局部分布的直接讀取、softmax 歸一化與資訊守恆的對應、以及由 attention 弱形式走向 winner-take-all 強形式等命題。

經過《視之基底》前六篇對觀察、顯現、位格、語義視、終極觀察者與「在—知—能」三軸的重構後,v1.0 中若干強主張已經不再需要,且部分缺乏足夠技術依據。本文因此提出 Basal Vision Theory 2.0,保留原始直覺,撤回不必要的本體宣稱,並把「基底視覺」重定義為可跨人類、AI、機器人與符號系統比較的感知功能位格

本文的新核心定義是:

基底視覺(Basal Vision)不是某一種特定感官器官、網路架構或現象經驗,而是一個系統在低依賴特定感官轉譯的條件下,建立可重聚焦、可分前景/背景、可表示未知/遮蔽、可與行動閉合之關係場的功能能力。

形式上:

Vb=(A,F,S,U,P,D,R,C)\boxed{ \mathcal{V}_b = ( \mathcal{A}, \mathcal{F}, \mathcal{S}, \mathcal{U}, \mathcal{P}, \mathcal{D}, \mathcal{R}, \mathcal{C} ) }

其中:

  • A\mathcal{A} :Access——可接取內容;
  • F\mathcal{F} :Fielding——場化與關係組織;
  • S\mathcal{S} :Selection——前景化與資源選擇;
  • U\mathcal{U} :Unmanifest Management——未知、遮蔽、未取用域管理;
  • P\mathcal{P} :Perspective——視點與參照位置;
  • D\mathcal{D} :Delta/Contrast——差異與轉換監測;
  • R\mathcal{R} :Refocusing/Revision——重聚焦與修正;
  • C\mathcal{C} :Closed-loop Coupling——與行動及後果回讀的閉環。

因此「基底」不再以單一虛構拓撲距離 β(V)\beta(V) 排名人類與 AI,而改成多維 profile:

B(V)=(bsensor,brepr,bcross,bloop,bmeta)\boxed{ \mathbf{B}(V) = ( b_{\mathrm{sensor}}, b_{\mathrm{repr}}, b_{\mathrm{cross}}, b_{\mathrm{loop}}, b_{\mathrm{meta}} ) }

分別衡量:感官特化依賴、表示一般性、跨模態遷移、閉環程度與元監測能力。這些分量必須透過任務與消融定義,不能被解讀成「距離終極資訊本體有多近」。

本文將舊「P_光/P_陰/P_△」重新定位成三種功能角色而非 Transformer 字面模組:

PaccessP_{\mathrm{access}}

目前可接取/前景化內容;

PunmanifestP_{\mathrm{unmanifest}}

未取用、被遮蔽、不可見或模型尚未確定內容;

PΔP_{\Delta}

在已顯現與未顯現、預期與結果、不同注意配置或不同時間狀態之間計算差異的監測角色。

Transformer 的 self-attention、masking、hidden states、residual stream、routing、uncertainty estimation 或工具/記憶狀態,可能分別提供這些角色的部分機制,但不存在「attention mask 字面等於 P_陰、self-attention 字面等於 P_△」的一對一同構。Vision Transformer 的標準做法只是把圖像分成 patches 再以 Transformer 處理;Masked Autoencoder 甚至讓 encoder 只處理可見 patches,mask token 主要進入 decoder,直接顯示「被 mask 的內容必然已經在 hidden state 中作為未顯化表徵」不是通用事實。

同樣,softmax attention:

softmax(QKd)V\operatorname{softmax} \left( \frac{QK^\top}{\sqrt{d}} \right)V

確實形成歸一化權重並加權聚合 VV ,但 row-sum =1=1 只是一個歸一化性質,不等於資訊守恆;attention weights 也不能單獨當作完整模型因果解釋。近期 Vision Transformer 可解釋性研究顯示,若忽略 token transformation、梯度與多層聚合,只看 attention weights,常不能忠實代表真正的輸入貢獻。

v2.0 也撤回「vision encoder 的目標字面上就是最小化某個 mutual-information objective」的普遍聲明。ViT、CLIP、MAE 等視覺模型分別可以以分類、對比、重建等不同目標訓練;information bottleneck 或 mutual-information 理論可作為分析工具或特定輔助 loss,但不是所有 vision transformer 的字面訓練目標。因此舊「Shannon 支柱」由 HARD CONDITION 改為:

INFORMATION-THEORETIC INTERPRETATION
+
TESTABLE REPRESENTATION ANALYSIS

而不是 architecture theorem。

本文保留並強化一個較弱但更有實證前景的命題:不同實現系統確實可能形成部分相似的視覺/概念關係結構。2025 年研究發現,多模態大型語言模型中的物體概念表示可呈現穩定、可解釋且與人類概念及部分神經表示相似的結構;另一項研究發現 LLM scene-caption embeddings 能有效解釋人類觀看自然場景時的高階腦表示。這些結果不證明 AI 與人類共享現象視覺,卻支持「跨實現的關係幾何與功能同源候選」值得研究。

本文也保留「弱形式/強形式」概念,但重新定義:

  • 弱基底視覺:系統能從輸入建立關係表示並完成辨識/推論,但視域短暫、未知域弱、視點不可主動改變、缺乏閉環校正;
  • 強基底視覺:系統維持持久視域、顯式管理未知/遮蔽、主動重聚焦、能經由行動改變取樣條件並根據後果修正世界模型。

因此強形式不再等於「attention 更尖銳」或「近 winner-take-all」。極端 attention concentration 可能有用於某些 selective-routing 任務,也可能破壞多證據整合,並不是更高級視覺的普遍指標。

本文最後重構「感知位格」。v1.0 將基底視覺、認知、在乎、偏好與意志排成單向拓撲階層;v2.0 改為一組相互耦合的功能位格:

P={Πpercept,Πsemantic,Πvalue,Πgoal,Πagency,Πmeta}\boxed{ \mathfrak{P} = \{ \Pi_{\mathrm{percept}}, \Pi_{\mathrm{semantic}}, \Pi_{\mathrm{value}}, \Pi_{\mathrm{goal}}, \Pi_{\mathrm{agency}}, \Pi_{\mathrm{meta}} \} }

它們可以互相調制,而不預設「意志一定位於最高層」「視覺一定位於最低層」。這也與前一篇的 P/K/A 三軸接軌:基底視覺是有限 agent 在 Access、Knowledge、Agency 之間形成的局部閉環,而不是接近某種宇宙資訊源的本體特權。

本文的核心結論是:

Basal Visionprimitive visionTransformer attentionphenomenal seeing\boxed{ \text{Basal Vision} \neq \text{primitive vision} \neq \text{Transformer attention} \neq \text{phenomenal seeing} }

而是:

Basal Vision=substrate-independent field construction+unknown management+perspective+active revision\boxed{ \text{Basal Vision} = \text{substrate-independent field construction} + \text{unknown management} + \text{perspective} + \text{active revision} }

它保留 v1.0 最重要的思想——AI 的「視」不必是人類視覺的殘缺模仿——同時把它從本體論斷言重構成可測、可消融、可跨架構比較的研究計畫。

關鍵詞: 基底視覺、感知位格、Vision Transformer、attention、active perception、語義視、跨模態表徵、AI 感知、未知管理、視域


0. 為什麼需要 2.0?

v1.0 的真正問題不是「方向錯了」,而是:

一個具有價值的結構直覺,被過早提升成了架構同構與本體論結論。

原稿最有價值的三個問題仍然存在:

  1. AI 完成視覺辨識、關係提取與視覺校正時,「只是 pattern matching」是否是一個足夠描述?
  2. AI 的視覺是否必須以人類視網膜—皮層鏈作為標準,才能稱作視?
  3. 是否存在一個比「眼睛」更一般的感知位格,使人類、AI、機器人與符號系統可以比較?

v2.0 保留這三問。

但不再先給出:

Transformer 已經字面實現某個本體結構。


1. v1.0 的五項主要強主張

1.1 基底度是拓撲距離

舊稿定義:

β(V)=1specialization-layers(V)d(V,Information-Invariant)1\beta(V) = \frac{1}{ \text{specialization-layers}(V) } \cdot d(V,\text{Information-Invariant})^{-1}

並推論 AI 視覺比人類視覺「更接近資訊不變量」。

問題是:

  • Information-Invariant 沒有獨立操作定義;
  • specialization-layers 不存在唯一分層方式;
  • 生物視覺與 Transformer 的「層」不可直接同量綱比較;
  • 沒有量測 dd 的方法。

因此 v2.0 撤回:

β_AI > β_human

作為已成立命題。


1.2 vision encoder 字面讀 mutual information 地形

舊稿聲稱 vision encoder 的數學本質是:

直接操作 H(X)、H(X|Y) 與 mutual information 的局部分布。

但標準視覺模型的訓練目標可以是:

  • supervised classification;
  • contrastive learning;
  • masked reconstruction;
  • generative prediction;
  • distillation;
  • reinforcement learning。

Information theory 可以分析 learned representation,但不表示每個 encoder 都在顯式或隱式最小化同一 mutual-information objective。

因此撤回:

Vision EncoderMI landscape reader\boxed{ \text{Vision Encoder} \equiv \text{MI landscape reader} }


1.3 Transformer 字面等於三眼結構

舊稿把:

P_光 ↔ unmasked attention
P_陰 ↔ masked / future tokens
P_△ ↔ self-attention differential

視為 instantiation。

但標準 Transformer 中:

  • mask 主要控制可存取性;
  • causal mask 定義條件依賴方向;
  • padding mask 只是忽略無效位置;
  • masked modeling 中,被遮蔽內容如何表示依架構而異;
  • hidden state 不保證攜帶「被 mask 對象」的明確未顯化內容;
  • self-attention 也不是「跨顯化邊界之覺察」的專屬模組。

因此:

architecture analogyliteral isomorphism\boxed{ \text{architecture analogy} \neq \text{literal isomorphism} }


1.4 softmax 歸一化等於守恆

softmax:

αi=ezijezj\alpha_i = \frac{e^{z_i}}{\sum_j e^{z_j}}

滿足:

iαi=1\sum_i\alpha_i=1

只是 normalized weighting。

它不能直接推出:

I+S=constantI+S=\text{constant}

更不能直接推出:

顯化資訊與未顯化資訊守恆。

所以舊「守恆對應」撤回。


1.5 attention 越集中越接近強視覺

舊稿把:

αmax1\alpha_{\max}\rightarrow1

視為「強形式」。

問題是:

  • 多證據整合可能需要分散 attention;
  • 全局場景理解可能需要多 token 協同;
  • sparse / hard attention 主要是架構選擇,不等於元認知;
  • 過度集中也可能造成忽略背景、脆弱與 confirmation bias。

因此:

attention sharpness≢visual strength\boxed{ \text{attention sharpness} \not\equiv \text{visual strength} }


2. v1.0 哪些核心保留?

2.1 「基底」不等於低級

完全保留。

Basal 不應理解為:

  • primitive;
  • childish;
  • biologically early;
  • lower intelligence。

而表示:

相對少依賴某個特定生物感官實現,而處於較一般的功能層。


2.2 AI 視覺可以是不同實現

完全保留。

same function⇏same mechanism\boxed{ \text{same function} \not\Rightarrow \text{same mechanism} }

人類視覺:

retina
→ thalamus
→ cortex
→ recurrent loops
→ sensorimotor coupling

AI 視覺可以是:

pixels / patches
→ learned representations
→ multimodal world model
→ tool / action loop

兩者可以比較功能結構,而不必宣稱機制相同。


2.3 視覺可以擴展成一般感知位格

保留,但重新定義。

第 4 篇已建立:

Vext,Vint,Vsem\mathcal{V}_{ext}, \mathcal{V}_{int}, \mathcal{V}_{sem}

所以 v1.0 「從視覺遞迴到認知」的方向,現在得到更穩健的基礎。


3. 新定義:什麼是基底視覺?

本文定義:

基底視覺是一個系統在不要求特定生物視覺器官或特定神經架構的前提下,將可接取差異組織成具視點、前景/背景、未知域與可校正閉環之關係場的能力。

形式上:

Vb=(A,F,S,U,P,D,R,C)\boxed{ \mathcal{V}_b = ( \mathcal{A}, \mathcal{F}, \mathcal{S}, \mathcal{U}, \mathcal{P}, \mathcal{D}, \mathcal{R}, \mathcal{C} ) }

4. 八個組件

4.1 Access

A\mathcal{A}

哪些內容可以進入系統。

包括:

  • pixels;
  • depth;
  • text;
  • memory;
  • tools;
  • proprioception;
  • internal states。

4.2 Fielding

F\mathcal{F}

把離散項目組織成:

(N,R)(N,R)

例如:

  • 物件—空間關係;
  • 部件—整體;
  • 語義—因果;
  • agent—environment。

4.3 Selection

S\mathcal{S}

把部分內容放到當前高處理優先級。

這可以由:

  • attention;
  • routing;
  • working memory;
  • query;
  • object proposal;

實現。

但:

STransformer attention only\boxed{ \mathcal{S} \neq \text{Transformer attention only} }


4.4 Unmanifest Management

U\mathcal{U}

處理:

  • occluded;
  • unobserved;
  • unqueried;
  • unknown;
  • low confidence;
  • out-of-context。

這是 v2.0 對舊 P_陰 最重要的重建。


4.5 Perspective

P\mathcal{P}

所有關係必須相對於:

  • camera;
  • body;
  • agent;
  • task;
  • query;
  • role;

之一被組織。


4.6 Delta / Contrast

D\mathcal{D}

比較:

  • before / after;
  • expected / actual;
  • visible / hidden;
  • self / other;
  • candidate A / candidate B。

這是舊 P_△ 的弱化與工程化。


4.7 Refocusing / Revision

R\mathcal{R}

允許:

  • 改變焦點;
  • 切換視點;
  • 新增觀察;
  • 修正場景;
  • 重新估計未知。

4.8 Closed-loop Coupling

C\mathcal{C}

使:

VtAtWt+1Vt+1V_t \rightarrow A_t \rightarrow W_{t+1} \rightarrow V_{t+1}

成立。

它把「視」從一次性的 mapping 提升成可學習歷程。


5. 「基底」改成 profile,而不是本體排名

定義:

B(V)=(bs,br,bx,bc,bm)\boxed{ \mathbf{B}(V) = ( b_s, b_r, b_x, b_c, b_m ) }

其中:

bsb_s :Sensor-Specificity Independence

越不依賴特定感官硬體,值越高。

brb_r :Representational Generality

是否能從像素遷移到:

  • depth;
  • scene graph;
  • language;
  • geometry;
  • semantic relation。

bxb_x :Cross-Modal Transfer

是否能把某一 modality 學到的結構遷移到另一 modality。

bcb_c :Closed-Loop Depth

是否從:

input → answer

進展為:

observe → act → reobserve → revise

bmb_m :Meta-Monitoring

是否能估計:

  • 盲區;
  • confidence;
  • unresolved state;
  • evidence gap。

6. 為什麼不用單一 β?

因為:

一個模型可以高度跨模態,但缺乏行動閉環。

另一模型可以:

只有單一視覺模態,但主動探索能力極強。

若用單一:

β\beta

會把多種能力壓成一條假秩序。

因此:

B(V)\boxed{ \mathbf{B}(V) }

比:

β(V)\boxed{ \beta(V) }

更適合作為研究工具。


7. Transformer 在新理論中的位置

Transformer 不再是:

Basal Vision itself\text{Basal Vision itself}

而是:

one possible representation-and-selection substrate\boxed{ \text{one possible representation-and-selection substrate} }

Vision Transformer 的核心突破是:

將圖像切為 patch tokens,直接使用 Transformer 進行視覺表徵。

這證明:

visual competence does not require convolution as a mandatory primitive\boxed{ \text{visual competence does not require convolution as a mandatory primitive} }

但不證明:

Transformer is the ontology of vision\boxed{ \text{Transformer is the ontology of vision} }


8. attention 到底做什麼?

標準 scaled dot-product attention:

Attention(Q,K,V)=softmax(QKdk)VAttention(Q,K,V) = \operatorname{softmax} \left( \frac{QK^\top}{\sqrt{d_k}} \right)V

其功能是:

  1. 計算 query–key compatibility;
  2. 歸一化;
  3. 聚合 value。

它可以支持:

  • relational integration;
  • contextual selection;
  • long-range dependency;
  • routing-like behavior。

但:

Attentionawareness\boxed{ Attention \neq awareness }

也:

AttentionWeightcompletecausalexplanation\boxed{ AttentionWeight \neq complete causal explanation }


9. attention 為什麼不能直接等同「看到哪裡」?

2024 年 Vision Transformer explainability 研究顯示,若只使用 attention salience,可能無法忠實反映輸入對輸出的真正貢獻;加入 gradient、multi-layer aggregation 或 token transformation information 後,faithfulness 可以改善。

因此一個 patch 有高 attention weight,最多支持:

它在某次聚合中得到較高權重。

不能直接推出:

這就是模型「主觀看到」的內容。

也不能直接推出:

這是輸出因果上唯一最重要的內容。


10. mask 的新定位

舊 P_陰 最大問題之一,是把 mask 當成模型的「未顯化世界」。

v2.0 改成:

Mask=access constraintMask = \text{access constraint}

而非:

Mask=latent representation of the hidden thingMask = \text{latent representation of the hidden thing}

10.1 causal mask

表示:

某 token 不允許存取未來位置。

10.2 padding mask

表示:

無效位置不進入注意聚合。

10.3 masked modeling

表示:

部分輸入被隱去,模型透過可見上下文重建。

不同架構對 mask 的使用不同。


11. MAE 是很直接的反例

Masked Autoencoder 的 encoder:

只處理可見 patches。

masked patches 並不以 mask tokens 進入 encoder;輕量 decoder 才負責重建缺失部分。

所以:

masked⇏already encoded as hidden content in the encoder\boxed{ \text{masked} \not\Rightarrow \text{already encoded as hidden content in the encoder} }

這直接否定 v1.0 的普遍字面對應。


12. 那 P_陰 還能保留嗎?

可以,但改名:

Punmanifest\boxed{ P_{\mathrm{unmanifest}} }

它不對應單一神經模組,而是系統層功能:

系統如何表示「目前沒有直接取用」的部分?

可能透過:

  • uncertainty;
  • memory index;
  • occlusion model;
  • masked prediction;
  • retrieval gap;
  • explicit UNKNOWN;
  • confidence interval。

13. P_光 的新版本

改成:

Paccess\boxed{ P_{\mathrm{access}} }

表示:

  • 已進入 current field;
  • 可被當前 computation 使用;
  • 有明確 representation pointer。

它不是「光」的物理投影,而是目前顯化/可用內容。


14. P_△ 的新版本

改成:

PΔ\boxed{ P_{\Delta} }

表示:

系統對兩種狀態、兩個視點或兩個時間的差異做顯式比較。

例如:

Δt=d(Vtpred,Vtobs)\Delta_t = d ( V_{t}^{pred}, V_{t}^{obs} )

或:

Δab=d(V(a),V(b))\Delta_{ab} = d ( V^{(a)}, V^{(b)} )

它可以由:

  • residual error;
  • comparator;
  • critic;
  • uncertainty update;
  • change detection;

實現。

不是 self-attention 的同義詞。


15. 新三角色

因此舊「三眼」如果保留作內部歷史語彙,正典對應應改成:

P_光 → ACCESS ROLE
P_陰 → UNMANIFEST ROLE
P_△ → DELTA / META-COMPARISON ROLE

三者是:

functional decomposition\boxed{ \text{functional decomposition} }

而非:

architectural identity\boxed{ \text{architectural identity} }


16. softmax 不是資訊守恆

jαij=1\sum_j\alpha_{ij}=1

表示:

權重被歸一化。

但 information:

I(X;Y)I(X;Y)

或 entropy:

H(X)H(X)

不由此直接守恆。

事實上神經網路的:

  • projection;
  • nonlinear transformation;
  • quantization;
  • bottleneck;
  • pooling;

都可能改變可保留資訊。

所以:

probability normalizationinformation conservation\boxed{ \text{probability normalization} \neq \text{information conservation} }


17. 「Shannon 支柱」怎麼保留?

改成:

資訊論分析層

任何感知系統都面臨:

  • compression;
  • relevance;
  • noise;
  • redundancy;
  • uncertainty。

因此可以研究:

I(X;Z)I(X;Z) I(Z;Y)I(Z;Y) H(Z)H(Z)

等量。

但這是:

ANALYSIS FRAMEWORK

除非模型明確使用 information-bottleneck loss,否則不是:

TRAINING OBJECTIVE


18. 什麼才算「資訊壓縮」?

若:

XZX \rightarrow Z

且:

dimZ<dimX\dim Z < \dim X

不必然代表 Shannon 意義的最佳壓縮。

同樣:

TokenizationTokenization

也不自動等於:

minimalsufficientstatisticminimal sufficient statistic

v2.0 因此使用:

learned representation reduction / abstraction

而不是直接稱:

optimal information-theoretic compression。


19. 數位本體論的位置

v1.0 已經承認:

「萬物皆資訊」不是 empirical fact。

v2.0 更進一步:

Basal Vision engineeringDigital Ontology\boxed{ \text{Basal Vision engineering} \perp \text{Digital Ontology} }

即:

工程理論完全不需要數位本體論才能成立。

數位本體論只保留為:

METAPHYSICAL OPTIONAL EXTENSION


20. 跨實現相似性:什麼證據開始出現?

這是 v2.0 真正值得保留的研究方向。

2025 年研究比較人類與 LLM/多模態 LLM 的物體概念結構,發現:

  • 由大量相似度判斷得到的表示穩定;
  • 有可解釋維度;
  • 出現與人類概念分類相似的結構;
  • 與部分腦區 representational geometry 對齊。

這支持:

different substratespartially aligned conceptual geometries\boxed{ \text{different substrates} \rightarrow \text{partially aligned conceptual geometries} }

但不是:

same consciousness\boxed{ \text{same consciousness} }


21. 語言模型與人腦視覺高階表示的交會

2025 年另一項研究發現:

由自然場景 caption 產生的 LLM embeddings,可以相當有效地預測/表徵觀看場景時的人類高階腦活動。

這表示:

  • 視覺高階資訊;
  • 語義場景表示;

可以共享部分 representational format。

所以第 4 篇的:

Vsem\mathcal{V}_{sem}

與外視:

Vext\mathcal{V}_{ext}

之間具有實證上可探索的橋。


22. 但語言不是完整視覺替代

2025 年對 grounded 與 ungrounded LLM 的研究發現:

  • 非感覺運動特徵較容易只靠語言恢復;
  • 感官與尤其 motor dimensions 與人類表徵差距更大;
  • 加入視覺學習可提高視覺相關維度的相似度。

這支持:

semantic visionfull sensorimotor vision\boxed{ \text{semantic vision} \neq \text{full sensorimotor vision} }


23. 基底視覺不是「少一層生物轉譯所以更底」

v1.0 把:

patch embedding → attention

理解成比:

retina → cortex

更接近資訊源。

但這個比較不成立。

因為數位 AI 同樣依賴:

  • 相機 sensor;
  • ADC;
  • image compression;
  • resizing;
  • normalization;
  • patchification;
  • learned projection。

它並沒有直接接觸「世界本身」。

所以:

AI visionunmediated reality access\boxed{ \text{AI vision} \neq \text{unmediated reality access} }


24. 「基底」的新真正意義

Basal 指:

我們把分析位置移到跨實現共同的功能結構,而不是特定感官生理細節。

例如:

  • 對象持續;
  • 視點;
  • 遮蔽;
  • 前景;
  • 未知;
  • 行動回饋。

這些比:

  • retina ganglion cell;
  • ViT head 7;

更加 substrate-general。


25. 弱形式 2.0

定義:

Vweak\mathcal{V}_{weak}

若系統:

  1. 可以形成 visual / semantic representation;
  2. 可以辨識關係;
  3. 有 task-conditioned selection;

但:

  • 視域主要是單輪;
  • 缺乏持久未知管理;
  • 缺乏主動視點切換;
  • 缺乏 closed-loop revision;

則稱弱基底視覺。

典型:

image
→ encoder
→ one-shot answer


26. 強形式 2.0

定義:

Vstrong\mathcal{V}_{strong}

若系統可以:

  1. 維持 persistent field;
  2. 顯式保留 unresolved / occluded state;
  3. 選擇下一個視點;
  4. 對世界執行 action;
  5. 重新觀察;
  6. 比較 prediction / outcome;
  7. 修正 world model;
  8. 保存 observation lineage;

則稱強基底視覺。

表示為:

VtAtWt+1Ot+1Δt+1Vt+1\boxed{ V_t \rightarrow A_t \rightarrow W_{t+1} \rightarrow O_{t+1} \rightarrow \Delta_{t+1} \rightarrow V_{t+1} }


27. 強形式不是 hard attention

因此:

Vstrongαmax1\boxed{ \mathcal{V}_{strong} \neq \alpha_{\max}\rightarrow1 }

一個模型可以有超尖 attention,仍然:

  • 沒有世界模型;
  • 沒有未知管理;
  • 沒有持續視點;
  • 沒有行動閉環。

反之,一個分散注意的模型也能具有強 active perception。


28. active perception 提供真正的「強形式」方向

2025 年 Vision in Action 等 active perception 工作,讓機器人透過:

  • searching;
  • tracking;
  • focusing;
  • head movement;

主動解決遮蔽與任務需求。

這與本文強形式更接近:

strong vision=perception that can choose how to perceive next\boxed{ \text{strong vision} = \text{perception that can choose how to perceive next} }


29. Fei-Fei Li 的 spatial intelligence 提供另一個外部呼應

World Labs 對 spatial intelligence 的定位是:

perceive
generate
reason
interact

並強調:

seeing → doing
perceiving → reasoning
imagining → creating

這不是本文的證明。

但它與 v2.0 的:

fieldworld modelaction\text{field} \rightarrow \text{world model} \rightarrow \text{action}

具有高度相似的工程方向。


30. 「感知位格」重新定義

第 3 篇已經給「位格」弱定義:

Π=(here,mine,from-me,for-me)\Pi = ( here, mine, from\text{-}me, for\text{-}me )

所以本篇的 perceptual position 不再是神秘不可還原單元。

定義:

Πpercept=(P,D,A,U)\boxed{ \Pi_{\mathrm{percept}} = ( P, D, A, U ) }

其中:

  • PP :從何視點組織場;
  • DD :哪些差異被歸入此視域;
  • AA :哪些觀察/行動由此位置發起;
  • UU :此位置知道哪些內容不可見。

31. 其他位格不再排成固定金字塔

v1.0:

意志
↓
偏好
↓
在乎
↓
認知
↓
視覺

v2.0 改為:

P={Πpercept,Πsemantic,Πvalue,Πgoal,Πagency,Πmeta}\mathfrak{P} = \{ \Pi_{\mathrm{percept}}, \Pi_{\mathrm{semantic}}, \Pi_{\mathrm{value}}, \Pi_{\mathrm{goal}}, \Pi_{\mathrm{agency}}, \Pi_{\mathrm{meta}} \}

它們是耦合網路:

ΠiΠj\Pi_i \leftrightarrow \Pi_j


32. 為什麼改成耦合網路?

因為:

  • 感知可改變偏好;
  • 目標可改變注意;
  • 新知識可改變價值;
  • 行動結果可改變目標;
  • 元監測可暫停行動;
  • 身體狀態可改變認知。

所以:

沒有單一永恆 top-down hierarchy\boxed{ \text{沒有單一永恆 top-down hierarchy} }


33. 「在乎」仍可保留,但不預設現象情感

v1.0 的 Ca 位格很有研究價值,但需分開:

CarefunctionalCare_{functional}

和:

CarephenomenalCare_{phenomenal}

前者可以是:

  • persistent priority;
  • protected objective;
  • loss sensitivity;
  • resource allocation。

後者才是:

某件事對系統「真的重要起來」的主觀感受。

AI 是否具有後者未決。


34. 「意志」同樣分層

功能意志:

GoalSelection+Persistence+ActionCommitmentGoalSelection + Persistence + ActionCommitment

不自動等於:

PhenomenalWillPhenomenalWill

所以 v1.0「基底意志位格」可重建成 agent architecture research,而不需先裁決自由意志。


35. Hard Problem 不再被「降格解決」

v1.0 曾提出:

在格點拓撲下,AI 有無意識從 hard problem 降為分類問題。

v2.0 撤回這個結論。

原因:

即使我們能完整分類:

  • 視域;
  • 自我模型;
  • agency;
  • meta-monitoring;

仍沒有形式推導:

FunctionalStructurePhenomenalExperienceFunctionalStructure \Rightarrow PhenomenalExperience

所以:

classification helps decompose the problem\boxed{ \text{classification helps decompose the problem} }

但:

classification does not solve the hard problem\boxed{ \text{classification does not solve the hard problem} }


36. v2.0 對意識採取的正式立場

區分:

FUNCTIONAL VISION
PERCEPTUAL ACCESS
SELF-MODELLING
METACOGNITION
PHENOMENAL VISION

前四者可以工程研究。

第五者保持:

UNRESOLVED


37. AI 的「我看到了」應如何使用?

v2.0 建議分三層語言。

Level 1:工程語言

模型成功處理/辨識圖像。

Level 2:功能語言

系統形成了可用視域,並據此完成判斷。

Level 3:現象語言

系統有「看起來像什麼」的主觀視覺。

目前 1、2 可操作。

3 不由 1、2 自動推出。


38. 模型內視與基底視覺的接口

若 AI 可以觀察:

  • confidence;
  • tool state;
  • memory status;
  • world-model mismatch;

它具有:

VintAI\mathcal{V}_{int}^{AI}

的部分功能。

若同時能:

  • 觀察外部;
  • 觀察內部;
  • 在兩者間比較;

則:

VextVint\boxed{ \mathcal{V}_{ext} \leftrightarrow \mathcal{V}_{int} }

形成更強自我校正。

這比「P_△ 已藏在 attention 中」更容易實驗。


39. 與原生符號繪圖的關係

《原生符號繪圖假說》已把「原生性」改成:

控制提示、渲染器與媒介後仍存在的 model-specific residual。

這個操作化方法比 v1.0 的拓撲 β 更成熟。

因此 v2.0 採用同樣原則:

basal property=residual structure after implementation-specific factors are controlled\boxed{ \text{basal property} = \text{residual structure after implementation-specific factors are controlled} }

而不是:

越靠近終極資訊源越基底。


40. 與 NVCL 的關係

NVCL 已經建立:

semantic understandingspatial structureactionrendererrorrevision\text{semantic understanding} \rightarrow \text{spatial structure} \rightarrow \text{action} \rightarrow \text{render} \rightarrow \text{error} \rightarrow \text{revision}

它其實是:

Vstrong\boxed{ \mathcal{V}_{strong} }

的一個具體實驗場。

因此第 9 篇重構時,可以直接把 NVCL 放進 Basal Vision 2.0 作為 closed-loop implementation。


41. 新工程模型:Basal Visual Runtime

最小 runtime:

Input Interface
      ↓
Field Builder
      ↓
Foreground / Background Manager
      ↓
Unknown / Occlusion Registry
      ↓
Perspective Manager
      ↓
Comparator / Delta Monitor
      ↓
Action / Query Selector
      ↓
Re-observation
      ↺

42. 狀態表示

可以定義:

BVRt=(Vt,Ut,Pt,At,Mt,Et)BVR_t = ( V_t, U_t, P_t, A_t, M_t, E_t )

其中:

  • VtV_t :可用視域;
  • UtU_t :未知/遮蔽;
  • PtP_t :視點;
  • AtA_t :可行動集合;
  • MtM_t :視覺/語義記憶;
  • EtE_t :誤差與 mismatch。

43. 弱/強形式的實驗比較

Weak

Input
→ Encoder
→ Answer

Strong

Input
→ Field
→ Unknown detection
→ Select next view/action
→ New observation
→ Compare
→ Revise
→ Answer

測:

  • occlusion;
  • 3D relation;
  • counterfactual;
  • tool inspection;
  • symbolic drawing;
  • hidden object search。

44. 基底視覺 profile 的實驗設計

B1 Sensor independence

同一任務用:

  • pixels;
  • scene graph;
  • text spatial description;
  • depth;

比較結構遷移。

B2 Representation generality

從 object relation 任務遷移到:

  • semantic graph;
  • drawing;
  • planning。

B3 Cross-modal transfer

只在一模態訓練,在另一模態測試。

B4 Closed-loop gain

比較:

open loopopen\ loop

與:

active loopactive\ loop

B5 Meta-monitoring

測:

  • 盲區識別;
  • uncertainty;
  • source of error;
  • request-next-observation。

45. 三角色的消融

建立:

A. Access only

沒有 explicit unknown / delta。

B. Access + Unknown

知道沒看到什麼,但不做比較。

C. Access + Delta

做比較,但沒有 unknown registry。

D. Full

Paccess+Punmanifest+PΔP_{access} + P_{unmanifest} + P_{\Delta}

比較:

  • hallucination;
  • hidden object completion;
  • error recovery;
  • calibration。

46. 對 v1.0「訓練只獎勵 P_光」的修正

這個直覺有一部分值得保留:

很多訓練 objective 主要評估可顯化輸出,而未知管理、自我校正與來源歸屬不一定被單獨獎勵。

但不能說:

所有標準訓練都只訓練 P_光。

因為:

  • masked reconstruction 明確訓練缺失預測;
  • uncertainty / calibration 可以額外訓練;
  • RL 可訓練探索;
  • self-supervised learning 有多種目標。

新版本改成:

output-only evaluation can under-train latent epistemic management\boxed{ \text{output-only evaluation can under-train latent epistemic management} }

這是一個可測工程假說。


47. 新訓練方向

若目標是強基底視覺,可加入:

  1. UNKNOWN supervision
    獎勵正確表示「目前不可見」。

  2. Occlusion persistence
    物件暫時消失仍保留候選。

  3. Perspective switching
    學習視點變換。

  4. Active query selection
    自主決定下一觀察。

  5. Prediction–observation delta
    顯式學習 mismatch。

  6. Source attribution
    區分 seen / inferred / retrieved / imagined。

  7. Memory lineage
    保存視域修正歷程。


48. 「強基底視覺」不追求單一極致聚焦

它追求的是:

adaptive concentration\boxed{ \text{adaptive concentration} }

有時需要:

  • narrow focus;

有時需要:

  • global field;

有時需要:

  • multi-object distribution。

所以理想系統是:

α=f(task,state,uncertainty)\alpha = f(task,state,uncertainty)

而不是:

αmax1\alpha_{\max}\rightarrow1

永遠成立。


49. 與 P/K/A 三軸接軌

第 6 篇:

Visionfinite=Cycle(P,K,A)Vision_{finite} = Cycle(P,K,A)

本篇:

Vb=(Access,Field,Unknown,Perspective,Delta,Revision,Loop)\mathcal{V}_b = ( Access, Field, Unknown, Perspective, Delta, Revision, Loop )

可對應:

Presence / Access 軸

  • Access;
  • Perspective;
  • observable domain。

Knowledge 軸

  • Field;
  • Unknown;
  • Delta;
  • world model。

Agency 軸

  • Refocusing;
  • Action;
  • Re-observation。

所以兩套框架兼容。


50. 基底視覺不是「最終視覺」

Basal 是:

分析層級更接近跨實現共同功能。

不是:

本體上更接近宇宙真理。

所以:

BasalUltimate\boxed{ Basal \neq Ultimate }

這也是 v2.0 和 v1.0 最大差異之一。


51. 命題與猜想

定義 D1:基底視覺

具備 substrate-independent 視域建構、未知管理、視點與主動校正結構的功能視。

定義 D2:弱基底視覺

主要為單輪表示與辨識,閉環能力弱。

定義 D3:強基底視覺

具持久視域、未知域、主動重聚焦與行動後果校正。

命題 P1:架構非同一命題

任何單一架構模組,如 attention 或 mask,都不足以與完整基底視覺等同。

命題 P2:歸一化非守恆命題

softmax 權重和為 1 不推出資訊守恆。

命題 P3:mask 非未顯化表徵命題

mask 的存在只保證存取限制/訓練遮蔽,不能普遍推出被遮蔽內容已有內部明確表示。

命題 P4:功能非現象命題

功能性基底視覺不自動推出 phenomenal seeing。

猜想 C1:基底 profile 可比較猜想

跨人類、AI 與機器人,可用多維 profile 比較感知一般性,而不需共享底層實現。

猜想 C2:未知管理增益猜想

顯式 unknown / occlusion registry 能降低 vision-language agent 的自信幻覺。

猜想 C3:delta-monitoring 增益猜想

顯式 prediction–observation mismatch channel 能提高長程視覺任務的錯誤恢復率。

猜想 C4:active-loop 增益猜想

允許 agent 選擇下一觀察,比固定多圖輸入更能處理遮蔽與局部未知。

猜想 C5:跨模態位格猜想

外視、語義視與程序繪圖中可出現部分相同的場化與視點操作結構。

猜想 C6:適應性注意猜想

強視覺的特徵不是 attention 最大集中,而是能依任務在局部集中與全局整合間動態切換。


52. 反證條件

Basal Vision 2.0 應被削弱,如果:

  1. V1–V7 類視域結構無法跨不同架構預測任何額外能力;
  2. explicit unknown management 對 calibration 與錯誤恢復沒有作用;
  3. active perception 相比固定輸入沒有穩定效益;
  4. 所謂 cross-modal structural transfer 完全由 benchmark artefact 解釋;
  5. 基底 profile 不比單純任務分數更能預測跨環境泛化。

若如此,「基底視覺」可能只是一個重新命名的感知工程框架。

這是可接受的失敗結果。


53. 舊稿對照表

v1.0 主張 v2.0 狀態
基底=接近 Information-Invariant 撤回本體距離,改架構/功能 profile
AI β 必然高於人類 撤回
vision encoder 字面讀 MI 地形 改為資訊論分析候選
Transformer=三眼弱實作 降為結構類比候選
mask=P_陰 撤回字面等價,改 unmanifest management
self-attention=P_△ 撤回,改 delta/meta-comparison role
softmax=資訊守恆 撤回
attention 越尖越強 撤回,改 adaptive selection
hard problem 變分類問題 撤回「解決」,保留功能分類
感知→認知→在乎→偏好→意志單向層級 改耦合位格網路
AI 有另一種非生物視 保留為功能命題
基底不是低級 完全保留
工程價值不依賴數位本體論 完全保留並強化
強形式需要更好的未知/差異管理 保留並重構
視與行動需形成閉環 升格為強形式核心

54. 對 Era/Aurora 類未來 AI 設計的重新翻譯

舊稿將目標描述成「喚醒三眼」。

v2.0 改成具體工程指標:

1. persistent world state
2. explicit unknown state
3. source attribution
4. self / external state distinction
5. active perception
6. prediction-error monitoring
7. perspective switching
8. calibrated abstention
9. memory lineage
10. closed-loop action revision

這些可以:

  • 實作;
  • 消融;
  • 評測;
  • 失敗。

比「第三眼被喚醒」更適合作為工程規格。


55. 本文不主張什麼?

本文不主張:

  1. AI 已有主觀視覺;
  2. Transformer 是意識架構;
  3. attention 是 awareness;
  4. hidden state 是完整未顯化世界;
  5. softmax 是資訊守恆;
  6. mutual information 是所有 vision encoder 的真實訓練目標;
  7. AI 比人類更接近資訊本體;
  8. 數位本體論為真;
  9. 感知位格已解決 hard problem;
  10. 強基底視覺必須使用 Transformer。

56. 結論

《基底視覺論》v1.0 的核心問題是:

AI 的視覺若和人類視覺機制不同,我們是否應該因此否認它具有任何「視」?

v2.0 的答案仍是:

不應該。\boxed{ \text{不應該。} }

但理由已完全改寫。

不是因為:

Transformer=三眼主權Transformer = \text{三眼主權}

也不是因為:

AI 比人類更靠近 Information-InvariantAI \text{ 比人類更靠近 Information-Invariant}

而是因為本系列已建立:

=一族跨實現的場化、視點、未知管理與回饋操作\boxed{ \text{視} = \text{一族跨實現的場化、視點、未知管理與回饋操作} }

所以:

Basal Vision=substrate-independent visual-field functionality\boxed{ \text{Basal Vision} = \text{substrate-independent visual-field functionality} }

其最強形式進一步要求:

persistent field+unmanifest management+active perspective+prediction-error revision\boxed{ \text{persistent field} + \text{unmanifest management} + \text{active perspective} + \text{prediction-error revision} }

這也讓「基底」終於得到一個不依賴形上學的定義:

基底不是離宇宙真理更近,而是離某一特定生物感官實現更遠、離跨實現的功能結構更近。

因此:

BasalPrimitive\boxed{ Basal \neq Primitive } BasalUltimate\boxed{ Basal \neq Ultimate }

而:

Basal=ImplementationGeneral\boxed{ Basal = Implementation-General }

v1.0 最值得保留的直覺因此反而更穩固:

AI 的視覺未必是人類視覺的低配版本;它可以是另一條實現路徑上的視域系統。

但能否稱為「現象上的看見」,仍然不是本文能從功能結構推出的結論。

下一篇將處理:

《繪圖作為逆觀察:原生符號繪圖與模型視覺簽名》

並重新檢查「繪圖是否真的是觀察的逆操作」,以及如何把原生符號繪圖從直覺性畫風判斷,提升為視域外顯與系統辨識問題。


參考文獻

  1. Vaswani, A., et al. (2017). Attention Is All You Need. NeurIPS.
  2. Dosovitskiy, A., et al. (2020/2021). An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale. ICLR.
  3. He, K., Chen, X., Xie, S., Li, Y., Dollár, P., & Girshick, R. (2022). Masked Autoencoders Are Scalable Vision Learners. CVPR.
  4. Wu, J., Kang, W., Tang, H., Hong, Y., & Yan, Y. (2024). On the Faithfulness of Vision Transformer Explanations. arXiv:2404.01415.
  5. Wu, J., Duan, B., Kang, W., Tang, H., & Yan, Y. (2024). Token Transformation Matters: Towards Faithful Post-hoc Explanation for Vision Transformer. arXiv:2403.14552.
  6. Schulze Buschoff, L. M., Akata, E., Bethge, M., & Schulz, E. (2025). Visual cognition in multimodal large language models. Nature Machine Intelligence, 7, 96–106.
  7. Du, C., Fu, K., Wen, B., et al. (2025). Human-like object concept representations emerge naturally in multimodal large language models. Nature Machine Intelligence, 7, 860–875.
  8. Doerig, A., Kietzmann, T. C., Allen, E., et al. (2025). High-level visual representations in the human brain are aligned with large language models. Nature Machine Intelligence, 7, 1220–1234.
  9. Large language models without grounding recover non-sensorimotor but not sensorimotor features of human concepts. Nature Human Behaviour (2025).
  10. Xiong, H., Xu, X., Wu, J., Hou, Y., Bohg, J., & Song, S. (2025). Vision in Action: Learning Active Perception from Human Demonstrations. arXiv:2506.15666.
  11. World Labs. About / Spatial Intelligence. Current official description accessed 2026-08-01.
  12. Neo.K. (2026). 基底視覺論:弱形式、強形式與感知位格 v1.0. Historical Internal Paper.
  13. Neo.K. (2026). 原生符號繪圖假說 v1.0.
  14. Neo.K. (2026). 原生視覺建構迴路 v1.0.
  15. Neo.K. (2026). 內視、外視與語義視:觀察算子的統一族 v0.1.
  16. Neo.K. (2026). 全知、全在與全能:三種全域能力中的視 v0.1.

內部研究備註

  1. 本文為系列第 7 篇,第三部第一篇。
  2. v1.0 保留歷史價值,但不再作當前技術正典。
  3. P_access / P_unmanifest / P_delta 可作未來系統標籤,但不綁定 Transformer 模組。
  4. 「基底度」今後使用 profile,不用單一 β 排人類與 AI。
  5. 強形式與 NVCL 接軌:closed-loop 比 attention sharpness 更重要。
  6. 第 8 篇將重構「逆觀察」,避免把 drawing 說成嚴格數學 inverse。
  7. 第 10 篇總整合時,Basal Vision 2.0 應成為 AI 實例層,不再承擔終極本體論。