title: "表格不是魔法:矩陣原生智能的邊界、失敗模式與下一代驗證計畫" title_en: "Tables Are Not Magic: Boundaries, Failure Modes, and the Next Validation Program for Matrix-Native Intelligence" series: "矩陣原生智能與可稽核計算系列" series_en: "Matrix-Native Intelligence and Auditable Computation Series" series_id: "EML-MNIAC-2026" paper_id: "EML-MNIAC-2026-10" version: "v0.1" date: "2026-08-17" language: "zh-Hant" document_type: "系列第10篇/封頂篇/Claim Maturity Audit and Validation Roadmap" status: "Public Draft / Series Closure" author: "Neo.K(許筌崴)/EveMissLab" depends_on: - "EML-MNIAC-2026-00~09" canonical_keywords: - Matrix-Native Intelligence - Auditable Computation - Claim Maturity - Falsification - Failure Modes - Spreadsheet Intelligence - MMR - MMLC - MLF - Executable Identity - MMR-IFN - Agent Control Plane - Validation Roadmap - Research Closure
表格不是魔法
矩陣原生智能的邊界、失敗模式與下一代驗證計畫
Tables Are Not Magic:
Boundaries, Failure Modes, and the Next Validation Program for Matrix-Native Intelligence
摘要
本文是《矩陣原生智能與可稽核計算》系列第 10 篇與第一階段封頂篇。
本系列最早由一個極端、甚至容易被誤解的問題開始:
Excel、CSV、試算表與矩陣,能不能不只是 AI 的資料容器,而真正進入模型狀態、計算、智能、Agent 與人機共同操作架構?
經過前九篇整理後,這個問題已被拆成多個彼此獨立、可驗證也可失敗的子命題:
- 模型狀態能否以開放表格形式保存並重建;
- 學習能否在顯式 invariant 下進行;
- 錯誤 constraint 是否產生不可消除 structural residual;
- Spreadsheet 是否能作有限可見 runtime / state machine;
- MMR / MMLC 是否能把方向、身份、依賴、交易與稽核顯式化;
- 這些結構是否真的能改善 formula audit / repair;
- MLF 是否能作為跨 spreadsheet / sequence / tensor / graph / human view 的可追溯 canonical interchange;
- 多個 projection 是否能共同指向一個 executable state,而不產生多重真實;
- MMR / IFN 是否能被編譯成真正控制 neural interaction graph 的 sparse-attention backend;
- Spreadsheet / matrix ledger 是否能成為 Agent Operations 的人類可見 control plane。
第一階段的答案不是一個單一的 Yes / No。
目前已有較硬內部證據支持:
- CSV-native toy Transformer 的 open model state 與 deterministic reload;
- constrained linear learning 中 trace-mismatch structural floor 的解析式與數值吻合;
- Spreadsheet Ledger / EML-LQ 的有限自適應計算與顯式 invariant;
- MMR-Bench 對 synthetic structure、Formula AST、dependency graph、real XLSX / OOXML cache、multi-engine disagreement 與 signed replayable certificates 的受控驗證;
- MLF 1.0 stable core 與 Compiler 1.0.0 的 release verification;
- MMLC deterministic runtime / audit / replay 路線;
- PHOSPHOR-SHEET 的 Human / AI / Spreadsheet projection 以及 governed XLSX round trip;
- MMR-IFN 的 deterministic hierarchical graph、edge-list sparse attention、same-graph numerical equivalence 與 CPU prototype;
- Veritaxa Workbench 的 Spreadsheet Control Plane、SQLite Source of Truth、Agent Task Ledger、Evidence / Review / Audit workflow。
但第一階段也有明確未跨越的邊界:
- 沒有真實企業 workbook 大規模外部驗證;
- 沒有證明 MMR direction semantics 對一般 spreadsheet task 有普遍增益;
- 沒有證明 MLF 能提高一般模型 sample efficiency 或 accuracy;
- 沒有 MMR-IFN causal language-model training、GPU/Triton kernel、quality parity 或 large-scale scaling law;
- 沒有 D-ALAN production multi-node network;
- 沒有多工作簿 concurrent write / conflict-resolution protocol;
- 沒有 production-grade fully autonomous Agent governance;
- 沒有獨立第三方重現整個系列的主要工程結果。
本文因此提出五級 Claim Maturity Scale:
分別代表:構想、形式化、可執行原型、受控驗證、真實/外部驗證、跨環境生產與獨立重現。
本文的核心結論不是「矩陣原生智能已成立」,而是:
因此最強命題即使失敗,也不會使整個系列歸零。反之,若要讓「Matrix-Native Intelligence」從研究名稱升格為強架構主張,下一階段必須停止主要依賴 synthetic success,直接進入真實 workbook、因果式 ablation、GPU hardware profile、small-LM training、multi-writer conflict、durable Agent failure injection 與外部重現。
1. 封頂篇的任務:不是再加一層理論
本篇不再增加:
- 新矩陣;
- 新公理;
- 新縮寫;
- 新 Agent;
- 新 route family。
它只做三件事:
2. 系列最危險的失敗方式
這個研究線最容易失敗的方式不是:
某個程式跑錯。
而是:
每一個局部成功都被語言膨脹成更大的成功。
例如:
被偷換成:
或者:
被偷換成:
因此本文首先列出十個禁止偷換的等號。
3. 十個禁止等號
3.1 CSV 不等於 Runtime
CSV-native prototype 證明的是:
3.2 Excel 不等於 LLM
Excel 可以計算、更新、顯示 state。
不代表它具有語言模型的 learned representation。
3.3 Matrix 不等於 Universal Ontology
矩陣是有用表示。
不是宇宙本體論。
3.4 Direction 不等於天然語義
不是:
之類的永久本體真理。
3.5 MMLC 不等於新通用電腦
它是一種 identity / dependency / transaction / audit 組合架構。
3.6 MLF 不等於 Universal Format
它處理 matrix-originated structured interchange 的一個特定問題域。
3.7 Executable Identity 不等於形上學單一真理
它只是 runtime scope 中的 authority contract。
3.8 Sparse Graph 不等於硬體加速
這是 MMR-IFN 下一階段最重要的系統性風險。
3.9 Agent Control Plane 不等於 Solved Autonomy
治理只是自治的一個必要條件。
3.10 多 Agent 不等於多份獨立證據
共同來源、共同模型與共同偏差可能產生假共識。
4. Claim Maturity Scale
本文採五級成熟度,加一個 L0 構想層。
L0 — Intuition / Question
只有:
- 直覺;
- 問題;
- 類比。
沒有正式 claim。
L1 — Formalized Hypothesis
已有:
- 定義;
- 架構;
- 數學式;
- failure condition;
但沒有真正執行。
L2 — Executable Prototype
已有:
- code;
- workbook;
- runtime;
- reproducible local artifact。
但未形成充分 benchmark。
L3 — Controlled Internal Validation
已有:
- regression tests;
- controlled benchmark;
- synthetic / bounded fixture;
- replay;
- tamper test;
- analytic cross-check。
但沒有外部真實世界充分驗證。
L4 — Real-World / External Validation
至少有:
- natural / real workload;
- independent environment;
- external reproduction;
- blinded or held-out data;
- matched strong baseline。
L5 — Production / Independent Robustness
已有:
- prolonged deployment;
- multi-environment reproduction;
- independent users / teams;
- failure monitoring;
- scale evidence;
- documented operational reliability。
5. 一條重要規則:元件成熟度不等於強主張成熟度
MLF 1.0 可以是:
的 format / compiler。
但:
「MLF 讓 AI 更準」
可能仍只有:
同理:
MMR-IFN sparse backend 可以:
但:
「MMR-IFN 是更好的 LLM 架構」
仍只有:
所以 maturity 必須綁定:
6. 系列 Claim Maturity Matrix
| ID | 主張 | 目前級別 | 已有證據 | 尚缺 |
|---|---|---|---|---|
| C1 | 有限神經模型狀態可完整外部化為 open tabular state | L3 | toy Transformer、gradient、loss、save/reload | 大模型可用性與外部重現 |
| C2 | 正確 invariant 可與 learning 共存 | L3 | controlled linear experiment | 更一般非線性任務 |
| C3 | 錯誤 invariant 可造成 structural error floor | L3 | 解析式+數值吻合 | 更一般 constraint families |
| C4 | Spreadsheet 可作有限 visible computational environment | L3 | EML-LQ、Excel numerical workbook | 大規模/多人/性能 |
| C5 | MMR identity/direction/traversal 可形式化 | L2–L3 | MMR/MLA/MMLC runtime | 一般任務必要性 |
| C6 | MMR direction 在結構任務可能有額外資訊值 | L3(窄域) | synthetic route task / formula tasks | natural task incremental gain |
| C7 | Formula audit 可整合 AST/dependency/role/multi-engine/certificate | L3 | MMR-Bench v1.0 | real enterprise corpus |
| C8 | MLF 可保存完整結構並產生可追溯 projection | L3 | stable spec/compiler/migration/tests | 外部 ecosystem / adoption |
| C9 | MMLC 可作 deterministic auditable ledger runtime | L3 | runtime / audit / hashes / replay | independent domain benchmarks |
| C10 | 單一 executable state 可有 Human/AI/Sheet projections | L3 | PHOSPHOR-SHEET v1.0–1.2 | multi-writer / long-run deployment |
| C11 | Spreadsheet 可作 governed Agent control plane | L3 | Veritaxa / PHOSPHOR control | production workload / external team |
| C12 | IFN hierarchical routes 可編譯為 sparse neural graph | L3(後端) | 10/10 tests、same-graph equivalence | trained task quality |
| C13 | MMR-IFN 具有 useful model architecture tradeoff | L1–L2 | executable backend only | causal LM、training、GPU、quality |
| C14 | D-ALAN 多節點分歧可改善可靠度 | L1 | formal architecture | real multi-node experiment |
| C15 | Matrix-native intelligence 對一般 AI 有普遍優勢 | L0–L1 | 無充分證據 | 整套外部 benchmark |
7. C1:Open Model State 已經證明什麼?
nanocsv_llm.py 的強點不在模型品質。
它已經證明:
可以支援指定條件下:
8. C1 沒證明什麼?
它沒有證明:
- CSV 是大型 checkpoint 最佳格式;
- CSV 讀取最快;
- 人類看到權重就理解模型;
- tabular state 自動帶來 interpretability。
所以:
9. C1 下一步其實不需要「一兆參數 CSV」
真正合理下一步是:
而不是強迫 production LLM 把每個 float 寫成文字。
10. C2/C3:系列最硬的數學結果之一
對原 trace mismatch experiment:
subject to:
得到:
原實驗:
而 gradient run:
11. 這個結果的價值不依賴 Spreadsheet
即使整個「矩陣原生 AI」最後失敗,
此結果仍然獨立存在:
這是本系列一個真正不依賴品牌架構的數學產物。
12. C2/C3 的下一步
應擴展至:
- nonlinear constraints;
- multiple invariants;
- inconsistent constraint sets;
- soft constraints;
- dynamic feasible sets;
- neural network tasks。
核心問題:
13. C4:Spreadsheet Runtime 的成熟邊界
EML-LQ 與 log-coordinate workbook 已足以支持:
但成熟架構最後反而把:
- large data;
- runtime;
- persistence;
移出 workbook。
這不是失敗。
而是:
14. C4 的殺死條件
如果一個 workload:
- workbook size 過大;
- recalculation 過慢;
- concurrency 不可控;
- formula lineage 不透明;
- version conflicts 過高;
那 Spreadsheet 應退回:
而不是堅持:
15. C5/C6:MMR 的真正風險
MMR 最容易變成:
為每個方向、區域、路徑加 metadata。
如果:
在更低成本下做到同樣效果,
則:
應被降格為:
16. Direction Gain 必須是 Incremental Gain
不能只比較:
和:
真正 baseline 必須至少包含:
- AST;
- dependency graph;
- region;
- role;
- normal graph path。
然後測:
17. 如果 怎麼辦?
答案不是修改定義直到成功。
而是:
這仍然可能有價值。
18. C7:MMR-Bench 最成熟的成果是「知道自己會錯」
MMR-Bench 的第一階段歷史包含:
- direction leakage;
- clustered error majority;
- stale OOXML cache;
- MMR evaluator XOR bug;
- external evaluator 1904-date bug。
所以:
這是比單次高分更重要的成果。
19. v1.0 的硬邊界
v1.0 目前最強紀錄包括:
- 69/69 regression;
- 4/4 certificate signatures;
- 4/4 exact replay;
- 10/10 tamper rejection;
- 10/10 wrong-source rejection;
- zero formula writes;
- 20/20 exponent agreement;
- 20/20 ROUND agreement;
- 20/20 1904-gap retention。
但:
20. C7 下一步:Natural Workbook Benchmark
下一代 benchmark 必須把資料拆成兩種:
A — Natural Errors
真實人工產生:
- inconsistent formulas;
- stale caches;
- legitimate exceptions;
- copied mistakes;
- broken names;
- changed business logic。
B — Controlled Injections
可精確知道 ground truth:
- formula drift;
- cache corruption;
- reference mutation;
- region pattern error;
- date system mismatch。
兩者不能混在一起報一個 accuracy。
21. Real Workbook Data Split
為防止 template leakage,
train / development / test 不應按 cell 隨機切。
應按:
切分。
否則同一模板的不同月份可能洩漏答案。
22. C7 下一階段必須有真正 Baselines
至少比較:
- simple formula majority;
- AST pattern;
- dependency graph only;
- static spreadsheet analyzer;
- LLM-only;
- graph+LLM;
- MMR without direction;
- full MMR;
- human reviewer baseline。
23. Spreadsheet Intelligence 的真正主要指標
不能只看:
至少報:
24. Safety Gate 應比 Accuracy 更嚴格
對自動寫回,
本文建議:
作第一階段 safety target。
這是工程目標,不是數學定理。
如果做不到,
系統仍可保留:
角色。
25. C8:MLF 1.0 已經是一個 stable core,但不是通用格式勝利
MLF 1.0 已凍結:
- deterministic container;
- manifest;
- cells / regions / roles;
- source formula + AST + dependency;
- routes;
- provenance;
- conversion loss;
- four fingerprints;
- migration;
- safe import/export boundaries。
這是格式工程層的實質成果。
26. MLF 已知限制應保持在正文,不是腳註
包括:
- bounded XLSX;
- no VBA;
- no Power Query;
- no pivot execution;
- no external refresh;
- presentation round-trip not universally lossless;
- no production claim for learned dependency inference;
- naturalistic workbook fixtures still synthetic。
27. MLF 的死亡測試
如果:
在真實 pipeline 中:
- 更簡單;
- 更小;
- 更快;
- 一樣可追蹤;
那麼 MLF 不應宣稱獨立格式必要性。
它可以退回:
28. MLF 真正最值得保留的變數可能是
即:
因為很多格式轉換最大問題不是:
不能轉。
而是:
轉成功了,但沒有人記得丟了什麼。
所以:
可能比「無損萬能格式」更現實。
29. C9:MMLC Runtime 的真正價值
MMLC 現階段最有力的不是「多向」。
而是:
這些特性可以在其他表示中也有價值。
30. MMLC 必須和普通 DAG / workflow runtime 正面比較
需要問:
多向矩陣帳本到底比普通 DAG + provenance 多出什麼?
若答案只剩:
顯示比較漂亮。
那麼 execution-layer claim 應降格。
31. C10:Executable Identity 是工程原則,不是哲學
PHOSPHOR-SHEET 已經支持:
並以受控 09_Control 完成 round trip。
這是一個實際工程原型。
32. C10 還沒跨過的關卡
目前尚未證明:
- multi-workbook concurrent edit;
- multi-user RBAC;
- quorum approval;
- offline merge;
- network-partition behavior;
- long-running production sessions;
- distributed writer arbitration。
因此:
33. 下一代 Executable Identity 壓力測試
至少要製造:
- stale workbook;
- wrong session;
- duplicate command;
- executed command replay;
- wrong runtime version;
- two workbooks editing same target;
- approval revocation;
- handler removed after approval;
- workbook tampering;
- crash after side effect but before status write。
34. 最後一個案例最重要
若:
但 process 在:
寫入前 crash,
resume 後可能再執行一次。
這是:
需要:
- idempotency key;
- transaction log;
- external-side-effect receipt;
才能處理。
35. C11:Agent Control Plane 已有形,不等於自治問題被解完
Veritaxa 已有:
- 17-sheet control plane;
- SQLite Source of Truth;
- status machine;
- URL frontier;
- evidence;
- review;
- ranking;
- publishing;
- audit;
- monitoring。
這支持:
36. 但「看得見」不等於「做得對」
Agent 仍可能:
- crawl wrong source;
- misclassify;
- hallucinate extraction;
- overuse expensive tools;
- generate review backlog;
- publish bad conclusions。
所以:
37. Agent 系統下一步應測 Failure Recovery,不只正常流程
每一 stage 都應注入:
- timeout;
- process kill;
- invalid tool result;
- partial write;
- duplicate event;
- stale approval;
- network disconnect。
然後問:
38. 外部 durable-execution 世界給了很清楚的壓力
Temporal 把:
crash 後從原位置繼續
直接作為 durable execution 核心。
OpenAI Agents SDK 現行 HITL 也把:
作為 approval pause / serialize / resume boundary。
所以 MNIAC Agent 線下一階段不能只證明:
狀態表看起來完整。
而要證明:
39. C12/C13:MMR-IFN 的兩種成熟度必須拆開
Backend Claim
已有:
級受控原型。
Model Quality Claim
只有:
40. MMR-IFN v0.1 已經做對的事
它沒有宣稱:
CPU ratio 21.75 = LLM speedup。
而是明確限制:
- same route graph;
- CPU;
- unoptimized dense reference;
- inference;
- no GPU;
- no training quality。
這個邊界必須保持。
41. 2026 sparse-attention frontier 使門檻更高
現代比較對象已不只:
- Longformer;
- BigBird;
- Routing Transformer。
還包括:
- FlashAttention 類 IO-aware dense;
- hardware-aligned learned sparsity;
- adaptive content sparse methods;
- 2026 年仍在快速出現的新動態 sparse schemes。
因此:
本身已遠遠不足以構成競爭優勢。
42. 最新 frontier 帶來的一個反直覺
2026 已出現使用 compression signal 做 parameter-free adaptive sparse selection 的研究。
這代表:
甚至「不增加 learnable router」也不再是 MMR-IFN 的獨特點。
所以真正 differentiation 必須是:
43. MMR-IFN 下一代最小實驗不是直接訓練巨大模型
合理順序:
- causal correctness;
- graph batching;
- hierarchical synthetic tasks;
- flat-sparse baselines;
- hardware profile;
- small language model;
- larger context;
- only then scale.
44. 第一階段 Small-LM Gate
建議使用一個可以負擔的模型規模,
例如:
級別,
做:
- causal LM;
- 4K / 8K context;
- held-out perplexity;
- wall-clock;
- peak memory;
- route recall;
- hardware utilization。
數值只是實驗規模建議,不是架構常數。
45. Small-LM Go Condition
只有在至少滿足一種:
Quality parity + resource gain
同時:
或:
Better quality at matched budget
在相同 training budget。
才值得 scale。
46. 如果只得到「更可稽核」呢?
那仍可能成立:
不必硬升格成:
47. C14:D-ALAN 是目前最應克制的一條線
D-ALAN 有:
- independent sources;
- independent models;
- divergence matrix;
- consensus estimate;
- active causal verification;
完整理論藍圖。
但 production evidence 幾乎不存在。
所以:
48. D-ALAN 第一個真正實驗不應碰「世界真理」
先做可控 simulation。
例如:
nodes,
其中:
- 2 independent reliable sources;
- 1 noisy source;
- 2 correlated copied sources。
測:
- divergence detection;
- false consensus;
- correlated-source weighting;
- retraction;
- source failure recovery。
49. 多 Agent 的真正 baseline
不能只比:
應比較:
- one model / one source;
- five models / same source;
- same model / five sources;
- five models / five independent sources;
- five agents with correlated provenance;
- provenance-aware aggregation。
這才能知道:
增益來自模型數還是證據獨立性?
50. C15:最強總命題目前必須保持弱
系列名稱:
不等於目前已證:
目前最強可守住的版本是:
51. 這已經不是零成果
因為「最強命題未證」和:
完全不同。
前九篇已經產生多個獨立子成果。
52. 即使 Matrix-Native LLM 失敗,哪些仍成立?
至少:
- Open Model State;
- Constraint structural residual;
- Visible Spreadsheet Runtime;
- typed transaction / local audit;
- Formula AST / dependency / cache differential;
- signed spreadsheet calculation certificates;
- MLF traceable projection;
- semantic vs execution identity;
- governed multi-projection state;
- spreadsheet-native Agent control surface;
- route-certificate engineering。
53. 這是整個系列真正成功的地方
最早的問題是一個:
現在變成:
每個 可以:
- pass;
- fail;
- downgrade;
- remain open。
所以研究不再需要:
「全對或全錯」。
54. 下一代驗證計畫總覽
本文提出六條平行路線。
Track A — Real Workbook Audit
驗證 MMR / MMR-Bench。
Track B — Format / Projection Interoperability
驗證 MLF。
Track C — Multi-Projection Runtime Stress
驗證 Executable Identity。
Track D — Small-LM Structural Sparse Training
驗證 MMR-IFN。
Track E — Durable Agent Operations
驗證 Veritaxa / PHOSPHOR Agent control。
Track F — Multi-Agent Divergence
驗證 D-ALAN。
55. Track A — Real Workbook Audit
最低資料要求:
- multiple workbook families;
- natural formula exceptions;
- multiple formula languages / styles;
- stale caches;
- named ranges;
- tables;
- date systems;
- hidden sheets;
- external-link metadata;
- different business templates。
56. Track A 的兩組 Ground Truth
Natural
由 human expert review 產生。
Injected
由 controlled mutation 產生。
兩者分開報告。
57. Track A 核心指標
58. Track A Kill Criterion
如果:
控制 AST / dependency / role 後,
在 natural workbook 上沒有穩定 incremental gain,
則:
停止擴張「方向智能」主張。
59. Track B — MLF Interoperability
選來源:
- XLSX;
- CSV;
- Markdown;
- graph JSON;
- formula AST;
- selected ONNX-like graph projection。
測:
60. Track B 核心指標
- semantic fingerprint preservation;
- presentation preservation;
- dependency preservation;
- provenance preservation;
- declared loss accuracy;
- package size;
- conversion time;
- tooling complexity。
61. Track B Kill Criterion
若:
長期過高,
或:
以顯著更低複雜度達到相同可追蹤性,
則:
不再追求 broad interchange format。
62. Track C — Executable Identity Stress
建立:
刻意製造 concurrent mutations。
63. Track C 測試矩陣
| 測試 | 應有結果 |
|---|---|
| stale workbook | reject / explicit rebase |
| wrong session | reject |
| duplicate command | reject |
| terminal replay | no second execution |
| wrong version | migrate or reject |
| two writers same revision | conflict |
| invalid approval | reject |
| capability removed | reject |
| crash before action | recover |
| crash after action/before receipt | no duplicate irreversible action |
64. Track C 最硬的成功條件
在完整 failure-injection suite:
這比「UI 看起來一致」更重要。
65. Track D — MMR-IFN
分四階段。
D1 — Structural tasks
- nested retrieval;
- tree path;
- variable binding;
- workbook lineage。
D2 — Negative controls
- random dependency;
- flat locality;
- no hierarchy。
D3 — Small causal LM
- perplexity;
- quality;
- throughput;
- memory。
D4 — Hardware
- CUDA / Triton / ROCm;
- block layout;
- HBM traffic;
- kernel utilization。
66. Track D 必須新增 2026 Baselines
除傳統:
- Longformer;
- BigBird;
- Routing Transformer;
還應納入:
- FlashAttention dense;
- hardware-aligned sparse;
- adaptive content-based sparse;
- 最新 parameter-free adaptive sparse。
因為 benchmark 必須跟:
比較,
不是只打 2020 baseline。
67. Track D Kill Criterion
若在 hierarchy-rich tasks 上:
相對:
沒有:
- quality;
- resource;
- auditability;
任一可重現優勢,
則:
停止強模型競爭主張。
68. Track E — Durable Agent Operations
Veritaxa / PHOSPHOR 應從:
轉:
每個 workflow stage 都 kill process。
69. Track E 要測的不是「能不能重試」
而是:
70. Track E 測試
- process crash;
- DB locked;
- storage unavailable;
- model timeout;
- approval delayed 24h;
- tool args invalid;
- event duplicated;
- retry after partial success;
- publish interrupted;
- external source changed during run。
71. Track E 核心指標
72. Track E Kill Criterion
如果 human review backlog:
長期:
且 automation 不能降低,
則:
應下調。
不能只繼續增加 Agent 數量。
73. Track F — D-ALAN
先用 simulation。
再用 low-risk public-data tasks。
最後才考慮高風險 domain。
74. Track F 核心指標
- false consensus;
- divergence recall;
- provenance-aware calibration;
- correlated-source penalty;
- source failure isolation;
- recovery after source correction;
- evidence independence estimate。
75. Track F Kill Criterion
如果 provenance-aware multi-node:
single strong node,
則:
不應以「多 AI」作賣點。
76. 需要第三方重現
目前大量證據仍屬:
下一階段至少應選三個最容易獨立重現的 artifact:
- Ledger mismatch analytic experiment;
- MMR-Bench certificate replay;
- MMR-IFN route graph / same-graph attention equivalence。
讓第三方在乾淨環境執行。
77. 為什麼先選這三個?
因為它們:
- 資源低;
- ground truth 清楚;
- 不需要 proprietary data;
- pass/fail 清楚;
- 容易產生 machine-verifiable result。
這比一開始要求第三方訓練大模型更合理。
78. Independent Reproduction Gate
只有當外部環境可重現:
時,
claim maturity 才升到:
79. 需要「失敗註冊」
未來每個實驗不只保存:
PASS
還應保存:
FAILED_HYPOTHESIS
NO_GAIN
UNSUPPORTED
INCONCLUSIVE
ENVIRONMENT_FAILURE
BASELINE_WINS
避免:
80. 這和本系列精神一致
MMR-Bench 最有價值的歷史之一就是:
100% 假成功被自己抓掉。
所以封頂後真正應標準化的是:
81. Claim Registry 建議
每個 claim 保存:
claim_id:
statement:
scope:
maturity:
supporting_artifacts:
external_baselines:
known_counterexamples:
falsification_test:
last_verified:
next_gate:
status:
82. 不再使用「整篇論文成熟度」
一篇 paper 可以同時有:
- theorem-like result:L3;
- architecture:L2;
- speculation:L1。
所以:
太粗。
應使用:
83. 下一代最小研究包
如果只做一個總實驗包,
本文建議包含:
MNIAC-Validation-v1/
├─ claims/
├─ open-state/
├─ constraint-floor/
├─ real-workbook-bench/
├─ mlf-roundtrip/
├─ executable-identity-chaos/
├─ mmr-ifn-structural/
├─ agent-failure-injection/
├─ multi-agent-divergence/
├─ baselines/
├─ certificates/
└─ failures/
84. 每一個結果都要有三個答案
- 它證明了什麼?
- 它沒有證明什麼?
- 下一個會殺死它的測試是什麼?
這三問應成為之後所有 MNIAC 文件的固定格式。
85. 系列的「最強主張」現在應重新命名
不應是:
更精確是:
即:
結構不一定必須是矩陣,但如果原始問題具有身份、區域、方向、依賴、歷史與路由,這些結構不應只因模型偏好線性輸入就被無聲丟失。
86. 這個命題比「Excel AI」更耐久
即使未來模型完全改變:
- Transformer;
- SSM;
- graph model;
- symbolic-neural hybrid;
仍然會問:
87. 系列真正統一的核心不是 Matrix,而是 Preservation
回頭看 00–09:
- CSV 保存 model state;
- Ledger 保存 constraint;
- Spreadsheet 保存 visible state;
- MMR 保存 direction / identity;
- MMLC 保存 dependency / audit;
- MMR-Bench 保存 disagreement / certificate;
- MLF 保存 projection linkage / loss;
- Executable Identity 保存 runtime authority;
- MMR-IFN 保存 interaction graph;
- Agent Control Plane 保存 task / evidence / review history。
所以最深層共同變數其實是:
88. Preservation 的七個維度
本文提出:
其中:
- :identity preservation;
- :semantic preservation;
- :dependency preservation;
- :history / provenance preservation;
- :authority preservation;
- :route / relation preservation;
- :verification evidence preservation。
89. 「完整保存」也不是永遠最佳
保存所有東西會增加:
- storage;
- token;
- compute;
- complexity;
- privacy risk。
所以真正問題是:
而不是:
90. 因此 Projection 仍然必要
好的 projection 是:
壞的 projection 是:
這可能是整個系列最簡短的工程原則之一。
91. 最終統一模型
可以將整個 MNIAC 寫成:
其中:
- :state / canonical objects;
- :computation;
- :dependency / graph;
- :route / representation;
- :intelligence;
- :history;
- :projections;
- :authority / governance;
- :verification / audit。
92. 任何大型主張都必須回答九個問題
- State 在哪?
- Computation 誰做?
- Dependency 怎麼保存?
- Route 誰決定?
- AI 的 authority 有多大?
- History 能否 replay?
- Projection 丟了什麼?
- Verification 誰負責?
- Failure 發生時怎麼恢復?
93. 如果答不出來,就不是「AI Native」的證據
模型很強並不能自動回答這九題。
所以:
這也是 Agent 時代越來越重要的區分。
94. 系列的最終降格原則
每條研究線都應允許被降格而不是被硬救。
MMR direction 無增益
→ UI / metadata。
MMLC 無額外 execution value
→ audit schema。
MLF 太重
→ interchange profile。
MMR-IFN 不快不準
→ auditable routing tool。
Spreadsheet control 負擔太大
→ Web / API projection。
D-ALAN 無益
→ multi-source audit protocol。
95. 降格不是失敗
一個研究物件最終從:
新計算範式
降成:
很好用的 audit format
仍然可能是有價值成果。
真正的失敗是:
96. 系列封頂後不應繼續無限擴文
第一階段研究綱領已足夠完整:
下一步的主要增量應來自:
而不是再增加同義理論名稱。
97. 第一優先:真實 workbook
因為這同時驗證:
- MMR;
- MMLC;
- MMR-Bench;
- MLF import;
- Agent review surface。
是一個高信息-density test。
98. 第二優先:MMR-IFN causal small model
因為這直接測:
「矩陣結構能否真正成為模型架構優勢?」
是整個系列最強命題的主要 gate。
99. 第三優先:Executable Identity chaos test
因為 Agent 越來越能執行外部 action,
multi-projection stale write / replay / duplicate action 的成本會快速上升。
這是一個直接有工程價值的方向。
100. 第四優先:外部重現
沒有外部重現,
整個系列最多仍是:
要跨到:
必須有人在不同環境得到同一結論。
101. 最終結論:表格真的不是魔法
表格不會因為:
- 有很多格子;
- 可以放公式;
- 可以存權重;
就自動變成 AI。
矩陣也不會因為:
- 有方向;
- 有圖;
- 有路徑;
就自動變成更好的智能架構。
102. 但「不是魔法」不等於「沒有力量」
表格的力量在於:
矩陣帳本的力量在於:
MLF 的力量在於:
Executable Identity 的力量在於:
MMR-IFN 的潛力在於:
Agent Control Plane 的力量在於:
103. 這個系列真正做完的是「拆解」
最早:
現在被拆成:
104. 強命題還沒完成,但研究對象完成了
這是封頂篇最重要的一句:
這比永遠不會錯的概念框架更重要。
105. 系列最終狀態
已封頂
- 00 — Unified Program
- 01 — Open Model State
- 02 — Constrained Learning
- 03 — Spreadsheet Runtime
- 04 — MMR / MMLC
- 05 — MMR-Bench
- 06 — MLF
- 07 — Executable Identity
- 08 — MMR-IFN
- 09 — Agent Control Plane
- 10 — Boundaries / Failure / Validation
106. 第一階段不再新增核心篇
後續若繼續,
應以:
- experiment report;
- benchmark release;
- runtime release;
- external audit;
- replication report;
為主。
也就是:
107. 封頂命題
最後將整個系列壓成一個最保守、也最難被未來淘汰的命題:
而任何壓縮、投影、稀疏化與自動化都應能回答:
這就是第一階段《矩陣原生智能與可稽核計算》系列的最終封頂。
參考資料
系列內部文件
- EML-MNIAC-2026-00, 《矩陣原生智能與可稽核計算:從表格模型狀態到多重投影執行同一性的統一研究綱領》.
- EML-MNIAC-2026-01, 《模型可以住在表格裡嗎?CSV-Native Transformer 與可檢查模型狀態》.
- EML-MNIAC-2026-02, 《學習作為帳本流:不變量、投影約束與不可消除結構殘差》.
- EML-MNIAC-2026-03, 《試算表作為可見計算環境:狀態、公式、依賴、記憶與稽核》.
- EML-MNIAC-2026-04, 《從二維表格到多向矩陣帳本:方向、遍歷、平行身份與可逆追蹤》.
- EML-MNIAC-2026-05, 《試算表智能的可證偽實驗:公式修復、方向推理、AST 與真實 XLSX 差分》.
- EML-MNIAC-2026-06, 《AI Matrix Ledger Format:從試算表格式到可追溯計算結構》.
- EML-MNIAC-2026-07, 《單一狀態、多重投影:人類、AI、試算表與 Runtime 的執行同一性》.
- EML-MNIAC-2026-08, 《矩陣原生智能:從多方向路由到 MMR-IFN 稀疏注意力》.
- EML-MNIAC-2026-09, 《從智能矩陣到 Agent 控制平面:可見狀態、任務帳本與人機共同操作》.
核心工程
nanocsv_llm.py.ledger_experiment.py.- EML-LQ Spreadsheet Ledger /
build_xlsx.py. - MMR-Bench v0.1–v1.0.
- MLF 1.0 / MLF Compiler 1.0.0.
kakon77777-commits/matrix-ledger-format.- MMLC Runtime 1.0 / MMLF 1.0.
kakon77777-commits/mmlc-runtime.- PHOSPHOR-SHEET v1.0–v1.2.
kakon77777-commits/eml-phosphor.- Veritaxa Workbench v0.1–v0.4.
- MMR-IFN Transformer v0.1.
外部技術與研究近鄰
- PyTorch, Saving and Loading Models.
- ONNX, Intermediate Representation Specification.
- Microsoft Learn, Excel Recalculation.
- LLVM Project, MLIR Language Reference.
- W3C, PROV-O: The PROV Ontology.
- Microsoft Azure Architecture Center, Event Sourcing, CQRS, Materialized View patterns.
- Kubernetes Documentation, Controllers.
- Temporal, Durable Execution / Workflow Documentation.
- OpenAI Agents SDK, Human-in-the-loop and RunState documentation.
- Beltagy et al., Longformer.
- Zaheer et al., Big Bird.
- Roy et al., Routing Transformer.
- Dao et al., FlashAttention.
- Yuan et al., Native Sparse Attention, 2025.
- Kundu, Ghosh & Honavar, Parameter-free Adaptive Sparse Attention via Compression-Based Content Selection, 2026.
- Barowy et al., ExceLint: Automatically Finding Spreadsheet Formula Errors.
- Abraham & Erwig, Spreadsheet Debugging.
- Tang et al., Efficient and Compact Spreadsheet Formula Graphs.
系列狀態:第一階段封頂。