← Archive
lm-003768 · 2026-09

IPM v0.1 Canonical Index:智能物理計量學系列總論、統一符號表與 v0.2 實驗入口

下載 MD 檔 ⬇
📎 附件 · Companion files — 隨文交付的程式 / 證明 / 資料,可獨立下載重驗

IPM v0.1 Canonical Index:智能物理計量學系列總論、統一符號表與 v0.2 實驗入口

Intelligence Physical Metrology v0.1 — Canonical Series Index, Unified Notation, and Experimental Entry Point

系列:《智能的物理計量:從最小語意執行到成果品質與計算時空》
英文系列: Physical Metrology of Intelligence: From Minimal Semantic Execution to Quality and Computational Spacetime
系列代號: EML-IPM
作者: Neo.K with Aletheia(GPT-5.6 Sol)
機構: EveMissLab/一言諾科技有限公司
版本: v0.1 Canonical Index
日期: 2026-09-02
狀態: Theoretical Series 10/10 COMPLETE
用途: 網站系列首頁、GitHub README、後續實驗與 v0.2 reference implementation 的 canonical entry point


0. 一句話版本

IPM 研究的不是「AI 有幾分聰明」,而是:

 一個智能系統在指定任務下, 以多少物理計算時空與多少外部鷹架, 完成多少有效語意工作, 最後產生多少可驗證品質? \boxed{ \textbf{ 一個智能系統在指定任務下, 以多少物理計算時空與多少外部鷹架, 完成多少有效語意工作, 最後產生多少可驗證品質? } }

1. IPM Canonical Intelligence Event

IIPM=(T,QIPM,Nμ,Pcompute,SC,M).\boxed{ \mathfrak I_{\mathrm{IPM}} = ( \mathfrak T, \mathfrak Q_{\mathrm{IPM}}, \mathbf N_{\mu}, \mathfrak P_{\mathrm{compute}}, \mathfrak S_C, \mathfrak M ). }

其中:

  • T\mathfrak T:Task / Specification Object;
  • QIPM\mathfrak Q_{\mathrm{IPM}}:Quality Object;
  • Nμ\mathbf N_{\mu}:Semantic Work Object;
  • Pcompute\mathfrak P_{\mathrm{compute}}:Physical Computation Object;
  • SC\mathfrak S_C:Scaffolding Capability Record;
  • M\mathfrak M:Measurement Metadata。

IPM 的核心不是找一個 AI IQ,而是比較:

PcomputeNμQIPM.\boxed{ \mathfrak P_{\mathrm{compute}} \rightarrow \mathbf N_{\mu} \rightarrow \mathfrak Q_{\mathrm{IPM}}. }

2. 系列依賴圖

Paper 01  Turn / LOOP / Single Pass
   │
   ├──> Paper 02  μI Semantic Execution
   │       └──> Paper 03  Cognition ↔ Neural Evidence
   │               └──> Paper 04  ATP / Joule / Thermodynamics
   │                       └──> Paper 05  Physical Cost + CST
   │
   ├──> Paper 06  Formal / Structured Quality
   │       └──> Paper 07  IBQF / Binary Human Residual
   │               └──> Paper 08  High-Ambiguity Quality Ontology
   │
   └──────── Paper 05 + Paper 08 ───────> Paper 09 Scaffolding
                                              └──> Paper 10 Unified IPM

三條主線:

P01P02P03P04P05\boxed{ P01\rightarrow P02\rightarrow P03\rightarrow P04\rightarrow P05 }

Execution / Physical Line

P06P07P08\boxed{ P06\rightarrow P07\rightarrow P08 }

Quality Line

(P01,P05,P08)P09P10\boxed{ (P01,P05,P08)\rightarrow P09\rightarrow P10 }

Capability / Integration Line


3. Paper 01–10 Canonical Map

Paper 核心問題 Canonical 輸出
01 一輪到底是哪一種一輪? U,G,I,L,R,S,PU,G,I,L,R,S,P ;ELI;Interaction Compression
02 智能「算一次」是什麼? μI\mu_INμgross/effN_\mu^{gross/eff} ;semantic state transition
03 認知如何跨層對到神經事件? Cross-Level Triangulation;measurement grade
04 神經事件如何對到 Joule? gross/base/marginal/attributed energy;Landauer type safety
05 FLOPs 之外的物理成本是什麼? CP\mathbf C_PVCST\mathbf V_{CST}ΘCST\Theta_{CST}
06 成果品質怎麼客觀量? Formal / Structured / Residual Quality;Hard Gate
07 人類主觀品質怎麼低負擔量? IBQF/BRQM;binary/pairwise → latent quality
08 高歧義成果有哪些品質維度? Typed Quality Ontology;Construct Graph
09 拿掉 LOOP 還剩多少 native capability? SSR、SDR、SCM、scaffolding ablation
10 如何統一比較智能產率? IIPM\mathfrak I_{IPM} ;Pareto frontier;reporting standard

4. Canonical Layer Stack

L4:Task Achievement / QualityL3:Semantic ExecutionL2:Algorithmic / Representational RealizationL1:Physical ComputationL0:Thermodynamic Realization.\boxed{ \begin{aligned} L_4 &: \text{Task Achievement / Quality}\\ L_3 &: \text{Semantic Execution}\\ L_2 &: \text{Algorithmic / Representational Realization}\\ L_1 &: \text{Physical Computation}\\ L_0 &: \text{Thermodynamic Realization}. \end{aligned} }

重要:

L4L3L2L1L0.\boxed{ L_4\neq L_3\neq L_2\neq L_1\neq L_0. }

各層可以建立映射,但不可互相偷換。


5. 統一符號:Turn / Execution

符號 定義
XX Task input
S\mathcal S Success specification
WW Evaluation environment
UU User Interaction Turns
GG Generation Trajectories
II Model Invocations
LL External Feedback Loops
RR Retries / Rollouts
SS Selection / Verification
τ\tau Single solving trajectory
ELI Externally Loopless Intelligence

最乾淨的 single-pass 條件:

U=1,G=1,R=1,L=0,S=0.\boxed{ U=1,\quad G=1,\quad R=1,\quad L=0,\quad S=0. }

但:

NoExternalLoopNoSequentialComputation.\boxed{ NoExternalLoop\neq NoSequentialComputation. }

6. 統一符號:Semantic Work

μI=Minimum Intelligent Semantic Execution Unit.\boxed{ \mu_I = \text{Minimum Intelligent Semantic Execution Unit}. }

Operationally:

μI:ztzt+1\boxed{ \mu_I:z_t\rightarrow z_{t+1} }

其中該 transition 必須是 task-relevant、causally useful、在指定 semantic resolution 下 operationally minimal。

TokenμIFLOP.\boxed{ Token\neq\mu_I\neq FLOP. } Nμgross\boxed{ N_\mu^{gross} }

表示候選語意活動總量;

Nμeff\boxed{ N_\mu^{eff} }

表示對成果具有有效因果貢獻的語意工作量。

ημ=NμeffNμgross.\boxed{ \eta_\mu = \frac{N_\mu^{eff}}{N_\mu^{gross}}. }

7. 統一符號:Cross-Level Realization

μIρC(μI)ρP(μI)ρT(μI).\boxed{ \mu_I \rightarrow \rho_C(\mu_I) \rightarrow \rho_P(\mu_I) \rightarrow \rho_T(\mu_I). }
  • ρC\rho_C:algorithmic/computational realization;
  • ρP\rho_P:physical trace;
  • ρT\rho_T:thermodynamic realization。

因此:

1μIConstant FLOPs\boxed{ 1\mu_I\neq Constant\ FLOPs }

且:

1μIConstant Joule.\boxed{ 1\mu_I\neq Constant\ Joule. }

8. 統一符號:Energy

E=(Egross,Ebase,Emarg,Eattrib,Ethermo,min).\boxed{ \mathcal E = (E_{gross},E_{base},E_{marg},E_{attrib},E_{thermo,min}). } Egross=t0tfPsystem(t)dt\boxed{ E_{gross} = \int_{t_0}^{t_f}P_{system}(t)dt } Emarg=t0tf[Psystem(t)Pbaseline(t)]dt.\boxed{ E_{marg} = \int_{t_0}^{t_f} [P_{system}(t)-P_{baseline}(t)]dt. }

Landauer:

Eerase,min=kBTln2\boxed{ E_{erase,min}=k_BT\ln2 }

但:

LandauerBoundActualIntelligenceCost.\boxed{ LandauerBound\neq ActualIntelligenceCost. }

9. 統一符號:Physical Computation

CP=(O,BM,BI,BN,VM,VC,T,E).\boxed{ \mathbf C_P = (\mathbf O,\mathbf B_M,B_I,\mathbf B_N,V_M,V_C,T,\mathcal E). }

其中:

O=(OFP64,OFP32,OBF16,OFP16,OINT8,)\boxed{ \mathbf O =(O_{FP64},O_{FP32},O_{BF16},O_{FP16},O_{INT8},\ldots) } BM=(Breg,Bcache,Bsram,Bhbm,Bhost).\boxed{ \mathbf B_M =(B_{reg},B_{cache},B_{sram},B_{hbm},B_{host}). }

Memory residency:

VM=Mresident(t)dt.\boxed{ V_M=\int M_{resident}(t)dt. }

Device time:

VC=D(t)dt.\boxed{ V_C=\int D(t)dt. }

10. Computational Spacetime

資源場:

R(t)=(rC(t),rM(t),rN(t),rS(t)).\boxed{ \mathbf R(t) =(r_C(t),r_M(t),r_N(t),r_S(t)). }

Raw CST:

VCST=R(t)dt=(VC,VM,VN,VS).\boxed{ \mathbf V_{CST} = \int\mathbf R(t)dt = (V_C,V_M,V_N,V_S). }

IPM v0.1 的重要 type-safety rule:

VC+VM+VN+VS\boxed{ V_C+V_M+V_N+V_S }

在沒有 normalization 前沒有物理意義。

因此:

CST=VectorFirst.\boxed{ CST=VectorFirst. }

11. Computational Spacetime Topology

ΘCST=(Twall,Tserial,Pparallel,Dpeak,Mpeak,Bpeak,Γcomm).\boxed{ \Theta_{CST} = (T_{wall},T_{serial},P_{parallel},D_{peak},M_{peak},B_{peak},\Gamma_{comm}). }

所以:

SameCSTVolumeSameCSTTopology.\boxed{ SameCSTVolume\neq SameCSTTopology. }

8GPU×10s8GPU\times10s1GPU×80s1GPU\times80s 可以具有相同 device-time volume,但 latency、peak capacity、communication 與 deployability 不相同。


12. 統一符號:Quality

Structured Quality:

QS=(QC,QA,QK,QR,QB,QV,QP).\boxed{ \mathbf Q_S =(Q_C,Q_A,Q_K,Q_R,Q_B,Q_V,Q_P). }

對應:

  • Correctness;
  • Alignment;
  • Completeness;
  • Consistency;
  • Robustness;
  • Verifiability;
  • Provenance。

三層品質:

QL=(QF,QS,QH).\boxed{ \mathcal Q_L=(Q_F,Q_S,Q_H). }
  • QFQ_F:Formal Objective;
  • QSQ_S:Structured Objective / Semi-Objective;
  • QHQ_H:Human Residual。

13. High-Ambiguity Quality Ontology

Q[d,τ,c,a]\boxed{ \mathcal Q[d,\tau,c,a] }

其中:

  • dd:domain / modality;
  • τ\tau:task;
  • cc:context;
  • aa:audience / evaluator population。
Q=QcoreQdomainQtask.\boxed{ \mathcal Q = \mathcal Q_{core} \oplus \mathcal Q_{domain} \oplus \mathcal Q_{task}. }

品質測量鏈:

TaskConstructIndicatorItemObservationLatentEstimate.\boxed{ Task \rightarrow Construct \rightarrow Indicator \rightarrow Item \rightarrow Observation \rightarrow LatentEstimate. }

因此:

ConstructIndicatorItemMetric.\boxed{ Construct\neq Indicator\neq Item\neq Metric. }

14. IBQF / BRQM

微觀回答:

bi{0,1}.\boxed{ b_i\in\{0,1\}. }

宏觀 latent quality:

θRd.\boxed{ \boldsymbol\theta\in\mathbb R^d. }

因此:

BinaryObservationBinaryPhenomenon.\boxed{ BinaryObservation\neq BinaryPhenomenon. }

基本映射:

{0,1}Nθ^H.\boxed{ \{0,1\}^{N} \rightarrow \widehat{\boldsymbol\theta}_H. }

母原則:

 評分者負責做容易、局部、具體的判斷; 測量系統負責做困難、全域、連續的量化。 \boxed{ \textbf{ 評分者負責做容易、局部、具體的判斷; 測量系統負責做困難、全域、連續的量化。 } }

15. Scaffolding

S=(ST,SR,SN,SV,SE,SM,SP).\boxed{ \mathbf S=(S_T,S_R,S_N,S_V,S_E,S_M,S_P). }
  • STS_T:Tool / External Information;
  • SRS_R:Retry;
  • SNS_N:Multi-sample / Best-of-N / Self-Consistency;
  • SVS_V:Verifier / Critic;
  • SES_E:Environment Feedback;
  • SMS_M:External / Persistent Memory;
  • SPS_P:Planner / Controller。

Single-pass:

QSP=Q(M,0).\boxed{ \mathfrak Q_{SP}=\mathfrak Q(M,\mathbf0). }

Full system:

QF=Q(M,SF).\boxed{ \mathfrak Q_F=\mathfrak Q(M,\mathbf S_F). }

16. SSR / SDR / SCM

SSR=QSPQF\boxed{ SSR=\frac{Q_{SP}}{Q_F} } SDR=1SSR.\boxed{ SDR=1-SSR. } SCMj=CjFCjSP.\boxed{ SCM_j=\frac{C_j^F}{C_j^{SP}}. }

SSR/SDR 必須與 physical overhead 一起解讀,不能把 scaffold dependence 本身當成缺陷。

LoopCheating.\boxed{ Loop\neq Cheating. }

真正需要避免的是能力來源與成本被隱藏。


17. Intelligence Yield

若品質 projection 已公開:

Q=ΠQ(Q),Q^*=\Pi_Q(\mathfrak Q),

則:

YI=(QEmarg,QVC,QVM,QBM,QBN,QT).\boxed{ \mathbf Y_I = \left( \frac{Q^*}{E_{marg}}, \frac{Q^*}{V_C}, \frac{Q^*}{V_M}, \frac{Q^*}{B_M}, \frac{Q^*}{B_N}, \frac{Q^*}{T} \right). }

語意產率:

Yμ/E=NμeffEmarg\boxed{ Y_{\mu/E}=\frac{N_\mu^{eff}}{E_{marg}} } YQ/μ=QNμeff.\boxed{ Y_{Q/\mu}=\frac{Q^*}{N_\mu^{eff}}. }

18. No Premature Scalarization Principle

 能保留向量時,不先壓成總分; 能保留結構時,不先壓成平均; 能保留不確定性時,不先假裝精確。 \boxed{ \textbf{ 能保留向量時,不先壓成總分; 能保留結構時,不先壓成平均; 能保留不確定性時,不先假裝精確。 } }

Scalarization 只有在 task、policy、weights、gates 與 boundary 明示後才合法。


19. Pareto Comparison

若:

QAQB\mathfrak Q_A\succeq\mathfrak Q_B

且所有 relevant cost axes:

CA,jCB,jC_{A,j}\le C_{B,j}

並至少一軸嚴格較優,則:

AIPMB.\boxed{ A\succ_{IPM}B. }

若不是 dominance:

保留 trade-off,不強迫總排名。\boxed{ \text{保留 trade-off,不強迫總排名。} }

20. IPM Minimum Reporting Standard v0.1

最低報告欄位:

Task

  1. Task ID / Task Text
  2. Success Specification
  3. Evaluation Environment
  4. Quality Boundary
  5. Physical Boundary

Quality

  1. Quality Schema
  2. Ontology Version
  3. Hard Gates
  4. Objective Verification
  5. Human Residual Protocol
  6. Quality Uncertainty

Execution

  1. Single-Pass / Full-System Flag
  2. Model Invocation Count
  3. Trajectory Count
  4. Retry Count
  5. Tool Calls
  6. Verifier / Selector Class

Physical

  1. Hardware
  2. Software / Runtime
  3. Wall Time
  4. Device Occupancy
  5. Peak Memory
  6. Memory Residency
  7. Memory Traffic
  8. Interconnect Traffic
  9. Energy Type
  10. Energy Boundary

Hidden / Discarded Work

  1. Candidate Count
  2. Discarded Attempts
  3. Wasted Physical Cost

Measurement

  1. Quality Grade
  2. Semantic Grade
  3. Energy Grade
  4. CST Grade
  5. Scaffolding Grade
  6. Scalarization / Projection Rule,如有。

21. Canonical Comparison Protocol

  1. Freeze Task: T\mathfrak T
  2. Freeze Quality Ontology: Qschema,VersionQ\mathcal Q_{schema},Version_Q
  3. Run Single Pass: (QSP,PSP)(\mathfrak Q_{SP},\mathfrak P_{SP})
  4. Run Scaffolded: (QF,PF)(\mathfrak Q_F,\mathfrak P_F)
  5. Compute SSR / SDR / SCM / ΔP\Delta\mathfrak P
  6. 可行時估 Nμ\mathbf N_\mu
  7. 建立 FIPM\mathcal F_{IPM}
  8. 若決策真的需要 scalar,才公開 ΠQ,ΠC\Pi_Q,\Pi_C
  9. 報 measurement grades 與 uncertainty。
  10. 保存 raw trace / provenance。

22. 系列核心 Invariants

TokenμIFLOP.\boxed{Token\neq\mu_I\neq FLOP.} OneUserTurnOnePhysicalTurn.\boxed{OneUserTurn\neq OnePhysicalTurn.} NoExternalLoopNoSequentialComputation.\boxed{NoExternalLoop\neq NoSequentialComputation.} Pass@kPass@1.\boxed{Pass@k\neq Pass@1.} SystemCapabilityModelNativeCapability.\boxed{SystemCapability\neq ModelNativeCapability.} SemanticWorkPhysicalWork.\boxed{SemanticWork\neq PhysicalWork.} FLOPsPhysicalComputationalCost.\boxed{FLOPs\neq PhysicalComputationalCost.} GrossEnergyMarginalEnergy.\boxed{GrossEnergy\neq MarginalEnergy.} LandauerBoundActualIntelligenceCost.\boxed{LandauerBound\neq ActualIntelligenceCost.} QualityUniversalScalar.\boxed{Quality\neq UniversalScalar.} FormalVerificationRealWorldGoalCorrectness.\boxed{FormalVerification\neq RealWorldGoalCorrectness.} BinaryObservationBinaryPhenomenon.\boxed{BinaryObservation\neq BinaryPhenomenon.} ReliabilityValidity.\boxed{Reliability\neq Validity.} NoveltyCreativity.\boxed{Novelty\neq Creativity.} LoopCheating.\boxed{Loop\neq Cheating.} InvisibleOutputZeroCost.\boxed{InvisibleOutput\neq ZeroCost.} SameQualitySamePhysicalCost.\boxed{SameQuality\neq SamePhysicalCost.} ScalarizationDeclaredPolicy.\boxed{Scalarization\Rightarrow DeclaredPolicy.} MeasurementUncertainty+Boundary+Version.\boxed{Measurement\Rightarrow Uncertainty+Boundary+Version.}

23. 五個 v0.1 可證偽命題

F1 — Token Hypothesis

若 token 是良好普適智能工作單位,則:

NμeffTokenCount\frac{N_\mu^{eff}}{TokenCount}

應跨 model / language / phrasing 相對穩定。

F2 — FLOPs Sufficiency

若 FLOPs 足夠描述物理成本,控制 FLOPs 後:

T,E,BM,BN,VMT,E,B_M,B_N,V_M

不應仍有巨大獨立差異。

F3 — Binary Burden Hypothesis

若 BRQM 的低負擔假說成立,適當 binary / pairwise protocol 應在至少部分場景改善 response time、consistency、dropout、predictive validity 或 fatigue。

F4 — Scaffolding Separation

QFQSPQ_F-Q_{SP} 在多數任務與 compute budgets 顯著存在,則 native / system capability separation 具有實證必要性。

F5 — Semantic Intermediate Utility

Nμ\mathbf N_\mu 無法改善 efficiency prediction、error explanation、scaffold analysis 或 cross-architecture comparison,則 μI\mu_I 應被修正甚至淘汰。


24. IPM v0.2:不要先做大平台

v0.2 的第一步應是最小可證偽實驗,而不是立刻做完整產品。

Experiment A — Single-Pass vs Scaffolded

優先用 math / code / structured reasoning,因為 QQ 容易客觀驗證。

條件:

  • A0:single pass;
  • A1:longer internal budget;
  • A2:multi-sample;
  • A3:verifier;
  • A4:tool / environment;
  • A5:full agentic loop。

每層記:

(Qk,Ek,Tk,VC,k,VM,k,BM,k,BN,k).(Q_k,E_k,T_k,V_{C,k},V_{M,k},B_{M,k},B_{N,k}).

Experiment B — Binary vs Numeric Human Measurement

比較:

  1. direct 0–10;
  2. structured Yes/No;
  3. adaptive pairwise。

量:response time、missingness、inconsistency、test-retest、predictive validity、fatigue。

Experiment C — μI\mu_I Operational Identification

選 proof steps、code repair、constraint puzzle,建立 candidate semantic transition,再做 ablation / counterfactual replacement。

Experiment D — Physical Trace Alignment

先從同一台機器的:GPU power telemetry、latency、memory peak、memory bandwidth、device occupancy 做起。

v0.2 不必一開始宣稱 data-center-level energy。

Experiment E — Token / FLOPs Proxy Failure Test

選相同 task quality、不同 language / verbosity / context / memory pressure 的執行,比較:

TokenCount,FLOPs,E,T,BM,Nμeff.TokenCount,FLOPs,E,T,B_M,N_\mu^{eff}.

25. v0.2 推進順序

推薦:

ADBCE.\boxed{ A\rightarrow D\rightarrow B\rightarrow C\rightarrow E. }

原因:

  1. 先確認 scaffolding gap 是否穩定存在;
  2. 建立 physical telemetry;
  3. 驗證 IBQF/BRQM 的人類測量負擔假說;
  4. 再攻最難的 μI\mu_I
  5. 最後挑戰 token / FLOPs proxy。

26. v0.2 最小 Run Schema

ipm_version: "0.2-experimental"
task:
  id:
  specification_version:
  quality_schema_version:

execution:
  mode: single_pass | scaffolded
  trajectories:
  retries:
  tool_calls:
  verifier_passes:

quality:
  hard_gate:
  structured_vector:
  human_residual:
  uncertainty:
  grade:

semantic:
  mu_count_gross:
  mu_count_effective:
  confidence:
  grade:

physical:
  wall_time_s:
  energy_type:
  energy_j:
  device_time:
  memory_peak_bytes:
  memory_residency_byte_s:
  memory_traffic_bytes:
  interconnect_bytes:
  boundary:
  grade:

provenance:
  model:
  hardware:
  software:
  timestamp:

27. Versioning Rule

任何下列定義改變都應 bump version:

  • μI\mu_I definition;
  • quality ontology;
  • CST normalization;
  • scaffold taxonomy;
  • reporting schema。

因此:

MeasurementDefinitionChangeVersionChange.\boxed{ MeasurementDefinitionChange\Rightarrow VersionChange. }

28. 系列最終母命題

 智能不只在於能否得到答案, 還在於一個物理世界中的系統, 為了得到這個答案, 究竟必須執行多少有效語意工作, 占用多少計算時空, 消耗多少能量, 依賴多少外部鷹架, 最後換回多少可驗證品質。 \boxed{ \textbf{ 智能不只在於能否得到答案, 還在於一個物理世界中的系統, 為了得到這個答案, 究竟必須執行多少有效語意工作, 占用多少計算時空, 消耗多少能量, 依賴多少外部鷹架, 最後換回多少可驗證品質。 } }

29. Canonical Status

EML-IPM v0.1=10 Papers+Canonical Index+Unified Notation+Experimental Entry Point.\boxed{ \text{EML-IPM v0.1} = 10\ Papers + Canonical\ Index + Unified\ Notation + Experimental\ Entry\ Point. }

下一個 canonical milestone:

IPM v0.2 — Experimental Measurement Protocol

它的目的不是證明 IPM v0.1 正確,而是讓 IPM 的核心命題第一次真正有機會被:

支持、修正、或證偽。\boxed{ \textbf{支持、修正、或證偽。} }