---
title: "矩陣原生智能：從多方向路由到 MMR-IFN 稀疏注意力"
title_en: "Matrix-Native Intelligence: From Multidirectional Routing to MMR-IFN Sparse Attention"
series: "矩陣原生智能與可稽核計算系列"
series_en: "Matrix-Native Intelligence and Auditable Computation Series"
series_id: "EML-MNIAC-2026"
paper_id: "EML-MNIAC-2026-08"
version: "v0.1"
date: "2026-08-17"
language: "zh-Hant"
document_type: "系列第08篇／Matrix-Native Intelligence／MMR-IFN 架構統合論文"
status: "Public Draft"
author: "Neo.K（許筌崴）／EveMissLab"
depends_on:
  - "EML-MNIAC-2026-04 從二維表格到多向矩陣帳本 v0.1"
  - "EML-MNIAC-2026-05 試算表智能的可證偽實驗 v0.1"
  - "EML-MNIAC-2026-06 AI Matrix Ledger Format v0.1"
  - "EML-MNIAC-2026-07 單一狀態、多重投影 v0.1"
internal_artifacts:
  - "無限分形九宮格：基於多向矩陣表徵的遞歸計算拓撲"
  - "MMR-IFN Transformer：基於遞歸多向路由的層級稀疏注意力架構 v0.1"
  - "MMR-IFN-Transformer_v0.1 完整實作包"
canonical_keywords:
  - Matrix-Native Intelligence
  - MMR
  - IFN
  - MMR-IFN Transformer
  - Sparse Attention
  - Edge-list Attention
  - Hierarchical Routing
  - Recursive Addressing
  - Route Certificate
  - Structured Sparsity
  - Long-Context Modeling
  - Graph Attention
---

# 矩陣原生智能
## 從多方向路由到 MMR-IFN 稀疏注意力

**Matrix-Native Intelligence:  
From Multidirectional Routing to MMR-IFN Sparse Attention**

---

## 摘要

本文是《矩陣原生智能與可稽核計算》系列第 08 篇，重新回到本系列最強、也最需要節制的模型架構命題：

> **矩陣中的方向、區域、依賴、層級與遍歷資訊，能否不只作為外部 metadata，而真正進入神經模型的 interaction topology？**

本文以多向矩陣表徵（MMR）、無限分形九宮格（Infinite Fractal Nonet, IFN）與既有最小原型 MMR-IFN Transformer v0.1 為內部基礎。IFN 不被視為「無限大模型」，而被定義成一個潛在無界、實際執行時始終有限的遞歸地址與路由語法。對任一次有限活動切片：

$$
\mathfrak F_{\mathrm{active}}
$$

可編譯成：

$$
\boxed{
\mathfrak F_{\mathrm{active}}
\xrightarrow{\operatorname{compile}}
G=(V,E)
\xrightarrow{\operatorname{sparse\ attention}}
Y.
}
$$

其中 token 狀態不只包含內容與原始序列位置：

$$
z_i
=
(h_i,p_i,\alpha_i,\rho_i,c_i),
$$

還加入 IFN 地址 $\alpha_i$ 、路由角色 $\rho_i$ 與信心／風險狀態 $c_i$。v0.1 使用確定性九元地址器，建立 SELF、LOCAL_FORWARD、LOCAL_BACKWARD、SAME_LEAF、ANCESTOR_UP、DESCENDANT_DOWN、SIBLING_REP、GLOBAL_ANCHOR 八種有向路由，再直接於 edge list 上執行 segment-softmax attention，而非先建構完整 $n\times n$ mask。

既有原型在 CPU 環境中記錄 10/10 測試通過；當 $n$ 從 81 增加至 6561 時，平均有效度數約介於 10.69–15.78，稀疏邊／dense pair 比例由 14.89% 下降至 0.24%。在「同一個固定路由圖」上，edge-list forward 相對未最佳化 dense masked reference 在 $n=81,243,729,1458$ 時記錄約 7.27、12.54、11.95、21.75 的 CPU 執行比，數值最大誤差低於約 $2.39\times10^{-7}$。

本文強調：這些結果只支持 **edge-list sparse backend 在本原型、同圖比較、CPU 環境下成立**。它們不支持「語言模型快 21.75 倍」、不支持 GPU kernel 優勢、不支持訓練吞吐提升，也不支持 MMR-IFN 能維持或提高語言模型品質。

外部研究進一步收緊了這個邊界。Longformer、BigBird 與 Routing Transformer 已分別展示 local/global、random/global/local 與 content-based sparse attention；Graph Attention Networks 顯示 attention 本來就能直接作用於圖鄰居；FlashAttention 說明 wall-clock 效率高度依賴 IO-aware kernel，而不只取決於 asymptotic sparsity；較新的 Native Sparse Attention 則進一步把可訓練稀疏策略與現代硬體 kernel 對齊。因此，MMR-IFN 的研究增量不能只是「我們也做 sparse attention」，而必須被壓縮為：

$$
\boxed{
\text{可追溯層級地址}
+
\text{多類語義路由}
+
\text{固定骨架／可學捷徑分離}
+
\text{路由證書}
+
\text{跨表示血緣}
}
$$

是否能在具有真實階層與依賴結構的任務中，提供比 flat sparse、window sparse、graph message passing 或一般 learned routing 更好的品質—成本—可稽核性折衷。

本文因此把「矩陣原生智能」從口號改成一組可反證主張。其最弱版本已由原型支持：**矩陣／層級結構可以被編譯成真正控制 attention 邊集合的神經 interaction graph。** 其最強版本——「這種結構在大模型上具有普遍品質與效率優勢」——目前完全尚未成立。

---

# 1. 「矩陣原生」到底意味什麼？

如果一個 Transformer 的輸入只是：

$$
x_1,x_2,\ldots,x_n
$$

而「矩陣」只存在於：

- UI；
- 預處理；
- Excel；
- JSON metadata；

最後進模型時全部被 flatten 成普通 sequence，

那麼：

$$
\boxed{
\text{Matrix-aware preprocessing}
\neq
\text{Matrix-native model}.
}
$$

---

# 2. 本文採用的最低定義

## 定義 1 — Matrix-Native Interaction

若原始結構中的：

- coordinate；
- region；
- direction；
- hierarchy；
- dependency；
- route；

至少有一部分直接決定模型可建立的 interaction edges：

$$
E
=
\Gamma(
Position,
Region,
Direction,
Hierarchy,
Dependency,
Task
),
$$

則稱模型具有某種：

$$
\boxed{
\text{Matrix-Native Interaction}.
}
$$

這個定義不要求模型內部真的保存一個 Excel grid。

它要求的是：

> 原始結構不能在進 attention 前完全消失。

---

# 3. 弱、中、強三種命題

## 3.1 弱命題

$$
\boxed{
\text{Structure}
\rightarrow
\text{Sparse Attention Graph}
}
$$

可以被工程實作。

**目前已有原型支持。**

---

## 3.2 中命題

對某類 structured tasks：

$$
QualityCost(MMR\text{-}IFN)
>
QualityCost(flat\ baseline).
$$

**尚需 benchmark。**

---

## 3.3 強命題

$$
\boxed{
\text{MMR-IFN is generally superior for large language models}.
}
$$

**目前沒有證據。**

---

# 4. IFN 不是 Transformer 的新名字

IFN 的概念層可寫：

$$
\boxed{
IFN
=
MMRLocalUnit
+
RecursiveAddress
+
Routing
+
Expansion
+
Stopping
+
Verification.
}
$$

Transformer 只是其中一種有限 backend：

$$
\mathfrak F_{\mathrm{active}}
\rightarrow
G
\rightarrow
Attention.
$$

因此：

$$
\boxed{
IFN
\neq
Transformer.
}
$$

---

# 5. 「Infinite」不是無限執行

IFN 的有限地址空間：

$$
\mathcal A_{<\omega}
=
\bigcup_{d=0}^{\infty}\Sigma_9^d.
$$

潛在無界地址：

$$
\mathcal A_{\omega}
=
\Sigma_9^{\mathbb N}.
$$

但任意時間：

$$
\mathcal A^{(t)}_{\mathrm{active}}
\subset
\mathcal A_{<\omega}
$$

且：

$$
\boxed{
|\mathcal A^{(t)}_{\mathrm{active}}|<\infty.
}
$$

所以：

$$
\boxed{
\text{Unbounded Grammar}
\neq
\text{Infinite Runtime}.
}
$$

---

# 6. 九宮格不是本體必然

IFN 使用：

$$
\Sigma_9
=
\{0,\ldots,8\}.
$$

原因包括：

- 中心＋八方向；
- 二維視覺直覺；
- 固定 fanout；
- 容易遞歸地址化。

但可以一般化為：

$$
\Sigma_b.
$$

所以：

$$
\boxed{
b=9
}
$$

是設計選擇，

不是：

$$
\boxed{
\text{宇宙級最佳分支數}.
}
$$

---

# 7. MMR Local Unit

定義：

$$
\mathcal M_{\alpha}
=
(
X_{\alpha},
D_{\alpha},
R_{\alpha},
P_{\alpha},
L_{\alpha},
C_{\alpha},
B_{\alpha}
).
$$

其中：

- $X_\alpha$：局部 state；
- $D_\alpha$：direction semantics；
- $R_\alpha$：routes；
- $P_\alpha$：traversal / update order；
- $L_\alpha$：lineage；
- $C_\alpha$：confidence / uncertainty；
- $B_\alpha$：budget。

這延續前幾篇的一條主線：

$$
\boxed{
\text{矩陣不是只保存 value。}
}
$$

---

# 8. Token 狀態

MMR-IFN Transformer v0.1 對 token $i$ 定義：

$$
\boxed{
z_i
=
(h_i,p_i,\alpha_i,\rho_i,c_i).
}
$$

其中：

$$
h_i\in\mathbb R^d
$$

是內容，

$$
p_i
$$

是原始 sequence position，

$$
\alpha_i
$$

是 IFN address，

$$
\rho_i
$$

是 routing role，

$$
c_i
$$

是 confidence / risk / route state。

---

# 9. Address 不能取代 Position

必須：

$$
\boxed{
\alpha_i
\neq
p_i.
}
$$

IFN address 決定：

> token 可以和哪些位置交互？

position mechanism 回答：

> 原始順序／相對位置是什麼？

所以即使使用 IFN：

- RoPE；
- ALiBi；
- learned position；

仍可能存在。

---

# 10. v0.1 的確定性地址器

給定：

- sequence length $n$ ；
- target leaf size $s$ ；
- branch $b=9$ ；

深度：

$$
D
=
\max
\left(
1,
\left\lceil
\log_9\frac ns
\right\rceil
\right).
$$

再以 normalized sequence position：

$$
x_i
=
\frac{i+1/2}{n}
$$

逐層取九元區間。

---

# 11. 地址器的目的不是語義最優

此地址器的優點：

- deterministic；
- order-preserving；
- $O(nD)$ ；
- no training；
- replayable。

缺點：

$$
\boxed{
\text{Address depends on position, not semantics}.
}
$$

所以它是一個：

$$
\boxed{
\text{Control Baseline}.
}
$$

不是完整 intelligence。

---

# 12. 為什麼先固定地址反而重要？

如果第一版同時學：

- address；
- route；
- attention；
- task；

成功時無法知道：

> 到底哪一層有用？

所以 v0.1 刻意把：

$$
AddressLearning=0.
$$

這是和 MMR-Bench 相同的研究節制：

$$
\boxed{
\text{One new uncertainty at a time}.
}
$$

---

# 13. 前綴群組

對地址前綴：

$$
\pi
$$

與深度 $d$：

$$
G_{d,\pi}
=
\{i\mid \alpha_i[:d]=\pi\}.
$$

v0.1 選：

$$
r(G_{d,\pi})
$$

為該群組中位 token，

作 deterministic representative。

---

# 14. 代表 token 的問題

這樣不增加：

$$
n.
$$

但它同時讓某個真實 token 承擔：

1. content role；
2. routing-control role。

可能產生：

$$
\boxed{
ContentControlEntanglement.
}
$$

所以後續版本已提出：

$$
\boxed{
VirtualSummaryToken.
}
$$

---

# 15. 八種路由

v0.1 定義：

$$
E
=
\bigcup_{r=0}^{7}E_r.
$$

其中：

1. SELF；
2. LOCAL_FORWARD；
3. LOCAL_BACKWARD；
4. SAME_LEAF；
5. ANCESTOR_UP；
6. DESCENDANT_DOWN；
7. SIBLING_REP；
8. GLOBAL_ANCHOR。

---

# 16. Local edges

若 radius：

$$
w,
$$

則：

$$
E_{local}
=
\{(i,j):|i-j|\le w\}.
$$

這保留：

$$
\boxed{
\text{short-range sequence continuity}.
}
$$

它本身與 sliding-window sparse attention 有直接近鄰。

---

# 17. Same-leaf edges

對最深層群組：

$$
G_{D,\pi},
$$

建立：

$$
G_{D,\pi}
\times
G_{D,\pi}.
$$

若 leaf size：

$$
s
$$

有界，

成本近似：

$$
\boxed{
O(ns).
}
$$

---

# 18. Ancestor / Descendant edges

每個 token 指向各層代表：

$$
(i,r_d)
$$

並建立反向：

$$
(r_d,i).
$$

所以資訊可以：

$$
Leaf
\rightarrow
Ancestor
\rightarrow
OtherScale.
$$

這是 IFN 和純 local-window sparse 的主要差異之一。

---

# 19. Sibling Representative edges

同一父節點下不同子群組代表互連。

在：

$$
b=9
$$

時，單父節點 fully connected sibling candidate 最大為：

$$
81
$$

條有向邊。

因此 fanout 仍有明確 bounded constant。

---

# 20. Global Anchor

根代表：

$$
r_0
$$

和所有 tokens 雙向連接。

作用：

$$
\boxed{
\text{short global communication path}.
}
$$

但它也可能造成：

$$
\boxed{
\text{root shortcut / information bottleneck}.
}
$$

所以必須消融：

- one root；
- multiple global tokens；
- no root；
- periodic dense layers。

---

# 21. 和 BigBird 的近鄰

BigBird 也使用：

$$
\boxed{
Local
+
Random
+
Global.
}
$$

並說明少量 global tokens 對 expressive power 有重要作用。

因此：

$$
\boxed{
GLOBAL\_ANCHOR
}
$$

不是 MMR-IFN 獨有思想。

MMR-IFN 真正不同點要看：

- hierarchy；
- address lineage；
- multiple route types；
- deterministic certificate。

---

# 22. 和 Longformer 的近鄰

Longformer 使用：

$$
\boxed{
SlidingWindow
+
TaskGlobalAttention.
}
$$

其 attention 隨 sequence length 線性增長。

所以：

$$
LOCAL
+
GLOBAL
$$

也不能作為 MMR-IFN 原創主張。

---

# 23. 和 Routing Transformer 的近鄰

Routing Transformer 使用 learned content-based clustering / routing，

目標是只讓 query attention 到內容相關位置。

其複雜度被設計成低於 dense：

$$
O(n^{1.5}d).
$$

這直接對應 MMR-IFN 未來：

$$
\boxed{
LearnedShortcut
+
LearnedAddress.
}
$$

所以後者必須與 content-based sparse baseline 正面比較。

---

# 24. 和 Graph Attention 的近鄰

Graph Attention Networks 已經顯示：

$$
\boxed{
Attention
}
$$

可以天然只在：

$$
Neighborhood
$$

上執行。

因此：

$$
\boxed{
\text{attention over edge list}
}
$$

本身不是新理論。

MMR-IFN 的可研究增量在於：

> 邊是如何由 MMR／IFN 的地址、層級、來源與路由類型產生，以及這些邊是否可審計。

---

# 25. Edge-list Attention

對：

$$
(i,j,r)\in E,
$$

v0.1 分數：

$$
s_{ij}
=
\frac{Q_iK_j^\top}{\sqrt{d_h}}
+
b_r.
$$

其中：

$$
b_r
$$

是 route-type bias。

---

# 26. Segment Softmax

只對：

$$
j\in Nbr(i)
$$

做：

$$
a_{ij}
=
\frac{
e^{s_{ij}}
}{
\sum_{m\in Nbr(i)}e^{s_{im}}
}.
$$

所以：

$$
\boxed{
\text{softmax support}
=
E.
}
$$

這才是「結構真的進模型」的核心位置。

---

# 27. Message Aggregation

$$
y_i
=
\sum_{j\in Nbr(i)}
a_{ij}V_j.
$$

v0.1 用：

- `scatter_reduce(amax)`；
- `index_add`；

實作。

因此 forward 不需要：

$$
n\times n
$$

完整 attention tensor。

---

# 28. Dense Reference 的真正用途

為了驗證 edge-list 數值語義，

建立：

$$
M_{ij}
=
1[(i,j)\in E].
$$

再比較：

$$
Y_{sparse}
\stackrel{?}{\approx}
Y_{dense-mask}.
$$

最大誤差：

$$
\lesssim2.39\times10^{-7}.
$$

所以：

$$
\boxed{
\text{Sparse backend}
}
$$

在相同 graph semantics 下通過數值等價測試。

---

# 29. 這個 dense reference 不是競爭基準

它不是：

- FlashAttention；
- xFormers；
- Triton block-sparse kernel；
- optimized SDPA。

所以：

$$
\boxed{
\frac{T_{dense-reference}}{T_{edge-list}}
}
$$

不能被當成：

$$
\boxed{
\text{production speedup}.
}
$$

---

# 30. Complexity

地址：

$$
C_A(n)
=
O(nD),
$$

其中：

$$
D
=
O(\log_9(n/s)).
$$

---

# 31. 路由成本

局部：

$$
O(nw),
$$

ancestor / descendant：

$$
O(nD),
$$

same leaf：

$$
O(ns),
$$

siblings：

$$
O(nb).
$$

在固定：

$$
w,s,b
$$

時：

$$
\boxed{
C_R(n)
=
O(n\log n).
}
$$

---

# 32. Attention cost

若：

$$
|E|=nk,
$$

則：

$$
\boxed{
C_{attn}
=
O(nkd).
}
$$

只有：

$$
k
$$

保持近常數或緩慢增長，

才有可能顯著低於：

$$
O(n^2d).
$$

---

# 33. Total Cost

$$
T(n)
=
C_A(n)
+
C_R(n)
+
O(nkd)
+
C_{certificate}(n).
$$

這一式很重要。

因為 MMR-IFN 不應只算：

$$
AttentionKernel.
$$

還必須算：

$$
\boxed{
RoutingConstruction
+
Verification.
}
$$

---

# 34. 稀疏不等於快

這是現代 sparse-attention 文獻給本系列最重要的外部教訓。

即使：

$$
FLOPs\downarrow,
$$

wall-clock 仍可能因：

- irregular memory；
- kernel launch；
- poor occupancy；
- gather/scatter；
- index overhead；

而不下降。

---

# 35. FlashAttention 的警告

FlashAttention 的核心不是 sparse，

而是：

$$
\boxed{
IO\text{-}aware exact attention.
}
$$

它指出：

> 理論 compute complexity 不是全部，HBM ↔ SRAM data movement 可能主導實際效能。

因此 MMR-IFN 不能只證明：

$$
|E|\ll n^2.
$$

還必須證明：

$$
\boxed{
\text{edge layout is hardware-efficient}.
}
$$

---

# 36. Native Sparse Attention 的新壓力

較新的 Native Sparse Attention 直接把：

- coarse compression；
- fine token selection；
- local precision；
- hardware alignment；
- end-to-end trainability；

放在同一設計中。

所以 2026 的 MMR-IFN benchmark 不應只和：

$$
Longformer / BigBird.
$$

還要面對：

$$
\boxed{
\text{hardware-aligned learned sparse attention}.
}
$$

---

# 37. MMR-IFN 的 CPU 原型數據

測試環境：

- Python 3.11；
- PyTorch 2.10 CPU；
- no CUDA；
- batch 1；
- $d=64$ ；
- 4 heads；
- inference mode。

因此所有數字都應加上：

$$
\boxed{
\text{Prototype / CPU / Same-Graph}.
}
$$

---

# 38. 10 項測試

既有文件列：

1. address range / order；
2. prefix；
3. deterministic route graph；
4. self edge；
5. subquadratic edge count within tested range；
6. sparse / dense numerical equivalence；
7. encoder shape / gradient；
8. root-anchor reachability；
9. route certificate tamper rejection；
10. certificate no formula/model-write authorization。

結果：

$$
\boxed{
10/10\ PASS.
}
$$

---

# 39. Graph Scale

| $n$ | $|E|$ | 平均度數 | dense pairs | sparse/dense |
|---:|---:|---:|---:|---:|
| 81 | 977 | 12.06 | 6,561 | 14.891% |
| 243 | 2,597 | 10.69 | 59,049 | 4.398% |
| 729 | 10,193 | 13.98 | 531,441 | 1.918% |
| 2,187 | 27,365 | 12.51 | 4,782,969 | 0.572% |
| 6,561 | 103,505 | 15.78 | 43,046,721 | 0.240% |

這支持：

$$
\boxed{
\text{Fixed v0.1 routing remains sparse in tested sizes}.
}
$$

---

# 40. 它沒有證明 asymptotic 常數度數

平均度數：

$$
10.69\sim15.78
$$

在測試範圍內穩定，

但不能推出：

$$
\boxed{
\limsup_{n\to\infty}k(n)<\infty.
}
$$

尤其：

- learned routes；
- dynamic expansion；
- semantic shortcuts；

都可能讓：

$$
k(n)
$$

上升。

---

# 41. Route Construction Time

內部數據：

| $n$ | 路由建構時間 |
|---:|---:|
| 81 | 0.00047 s |
| 243 | 0.00147 s |
| 729 | 0.00595 s |
| 2,187 | 0.01888 s |
| 6,561 | 0.07662 s |

這只量：

$$
\boxed{
Python\ CPU\ graph\ construction.
}
$$

---

# 42. Same-Graph Attention Timing

| $n$ | edge-list | dense reference | ratio | max error |
|---:|---:|---:|---:|---:|
| 81 | 0.00071 s | 0.00514 s | 7.27 | $1.79\times10^{-7}$ |
| 243 | 0.00114 s | 0.01428 s | 12.54 | $2.38\times10^{-7}$ |
| 729 | 0.00483 s | 0.05769 s | 11.95 | $2.38\times10^{-7}$ |
| 1,458 | 0.00569 s | 0.12371 s | 21.75 | $2.38\times10^{-7}$ |

---

# 43. 正確解讀

這些數字支持：

$$
\boxed{
\text{Avoiding dense tensor allocation helps this CPU reference implementation}.
}
$$

它不支持：

- LLM 21.75×；
- GPU 21.75×；
- training 21.75×；
- FlashAttention comparison；
- equal task quality。

---

# 44. Route Certificate

v0.1 定義：

$$
\mathcal C_R
=
(
H_{input},
modelID,
routerID,
H_G,
n,
|E|,
k,
G
).
$$

證書可以驗證：

- address；
- edges；
- route types；
- config；
- payload tampering。

---

# 45. Certificate 仍然很弱

它不能驗證：

- model weights；
- output tensor；
- training data；
- kernel nondeterminism；
- task truth；
- human intent。

所以：

$$
\boxed{
RouteIntegrity
\neq
ModelCorrectness.
}
$$

---

# 46. 為什麼 Route Certificate 仍有價值？

因為 learned / sparse systems 最大問題之一是：

> 模型到底看了誰？

如果 route graph 不保存，

就只能事後推測。

Route certificate 至少能回答：

$$
\boxed{
\text{Which interactions were allowed?}
}
$$

---

# 47. 從 MLF / Executable Identity 接回來

MLF 前一篇保存：

$$
RouteGraph.
$$

Executable Identity 保存：

$$
Runtime / Version / Source.
$$

MMR-IFN 可以再保存：

$$
\boxed{
ModelInteractionGraph.
}
$$

因此整條線形成：

$$
\boxed{
\text{Data Structure}
\rightarrow
\text{Execution Structure}
\rightarrow
\text{Neural Interaction Structure}.
}
$$

---

# 48. 固定路由訓練

第一階段：

$$
G
$$

固定，

只學：

$$
\theta.
$$

$$
\theta^\ast
=
\arg\min_{\theta}
\mathcal L_{task}(f_\theta(X,G),Y).
$$

目的：

> 先問 graph 本身能不能承載任務。

---

# 49. Static Semantic Address

第二階段由：

- document section；
- AST；
- syntax tree；
- dependency graph；
- region；

建立：

$$
\alpha_i.
$$

這是：

$$
\boxed{
\text{data-dependent but not model-learned routing}.
}
$$

這一步和 MLF / MMR-Bench 最容易直接接合。

---

# 50. Learned Shortcut

固定 backbone：

$$
E_0
$$

只學：

$$
E_{learned}.
$$

$$
E
=
E_0
\cup
E_{learned}.
$$

可以加入：

$$
\mathcal L
=
\mathcal L_{task}
+
\lambda_E|E_{learned}|
+
\lambda_HH(P_{route})
+
\lambda_S\mathcal L_{stability}.
$$

---

# 51. 為什麼只先學 shortcut？

因為：

$$
\boxed{
\text{Known Structure}
+
\text{Learned Exception}
}
$$

比：

$$
\boxed{
\text{Learn Everything From Scratch}
}
$$

更容易稽核。

也更容易做 ablation：

> learned route 是否真的有用？

---

# 52. Learned Address 是最高風險層

最後才讓：

$$
P(\alpha_i\mid h_i,p_i,q)
$$

可學。

風險：

- address collapse；
- root collapse；
- over-depth；
- route instability；
- discrete assignment；
- cache invalidation。

---

# 53. 這和 MoE Routing 有相似失敗模式

Mixture-of-Experts 系統已有：

- load imbalance；
- routing instability；
- expert collapse；
- communication overhead；

等問題。

所以 MMR-IFN learned addressing 不能假設：

> 可學就一定比固定地址好。

必須監控：

$$
\boxed{
LoadBalance
+
DepthCost
+
RouteEntropy
+
Stability
+
Recall.
}
$$

---

# 54. Necessary Baselines

MMR-IFN 既有文件已正確列出至少：

1. Dense attention；
2. Local window；
3. Local + global；
4. BigBird；
5. Longformer；
6. flat top-k learned sparse；
7. tree Transformer；
8. GNN message passing；
9. block-sparse；
10. MMR-IFN。

本文再加入現代基準：

11. IO-aware dense / FlashAttention；
12. hardware-aligned learned sparse / NSA 類方案。

---

# 55. 為什麼不能只比 Dense？

如果 MMR-IFN 比：

$$
O(n^2)
$$

dense 快，

但比：

$$
Longformer
$$

更慢、更差，

那麼：

$$
\boxed{
\text{它沒有實際研究優勢}.
}
$$

所以真正 benchmark 必須是：

$$
\boxed{
\text{Sparse vs Sparse}.
}
$$

---

# 56. Benchmark 指標

至少報：

- task quality；
- training loss；
- perplexity / accuracy；
- FLOPs；
- wall-clock；
- peak memory；
- graph construction；
- forward time；
- backward time；
- route recall；
- average degree；
- path length；
- kernel utilization；
- energy if available；
- certificate overhead。

---

# 57. Synthetic Task 1 — Nested Key Retrieval

資料有：

- section；
- subsection；
- local key。

查詢需沿 address prefix 找答案。

測：

$$
\boxed{
RouteRecall(depth).
}
$$

---

# 58. Synthetic Task 2 — Tree Path Recovery

給兩 leaf nodes：

$$
u,v,
$$

輸出：

- LCA；
- full path。

直接測：

$$
LCP(\alpha_u,\alpha_v).
$$

---

# 59. Synthetic Task 3 — Long-range Variable Binding

在：

$$
Block_A
$$

定義變數，

在：

$$
Block_B
$$

使用。

測 ancestor / shortcut 是否能保持 binding。

這比純 language modeling 更能測真正 hierarchy value。

---

# 60. Synthetic Task 4 — Spreadsheet Dependency

把 MMR-Bench 的：

- region；
- formula AST；
- lineage；
- blast radius；

轉成 node task。

問：

> 錯誤 upstream cell 會影響哪些區域？

這是本系列最自然的 end-to-end bridge。

---

# 61. 必須設計反例任務

資料：

$$
Dependency
\approx
RandomGraph.
$$

沒有穩定 hierarchy。

若：

$$
MMR\text{-}IFN
<
FlatSparse
$$

反而是好結果。

因為這能界定：

$$
\boxed{
\text{HierarchyBias is useful only when hierarchy exists}.
}
$$

---

# 62. Matrix-Native Intelligence 的最小正主張

目前可以說：

$$
\boxed{
\text{Matrix / hierarchical structure can directly constrain neural interaction edges}.
}
$$

這已不是：

$$
UI.
$$

也不是：

$$
MetadataOnly.
$$

---

# 63. 中等主張

若未來實驗證明：

$$
Quality(MMR\text{-}IFN)
\ge
Quality(Baseline)
$$

同時：

$$
Cost(MMR\text{-}IFN)
<
Cost(Baseline)
$$

且：

$$
Auditability(MMR\text{-}IFN)
>
Auditability(Baseline),
$$

才可以說：

$$
\boxed{
\text{Matrix-native routing provides a useful architecture tradeoff}.
}
$$

---

# 64. 強主張仍不成立

目前完全不能說：

$$
\boxed{
\text{Matrix-native intelligence replaces Transformer}.
}
$$

因為：

- 沒有 causal LM；
- 沒有 pretraining；
- 沒有 downstream LM tasks；
- 沒有 CUDA/Triton；
- 沒有 end-to-end wall-clock training；
- 沒有 quality parity；
- 沒有 large-scale scaling law。

---

# 65. 原型目前 12 個限制

既有文件已列：

1. position-only address；
2. representative-content/control entanglement；
3. one route label per edge；
4. no edge prior；
5. no causal mask；
6. no variable-length graph batching；
7. Python batch loop；
8. no CUDA/Triton；
9. no training task result；
10. hash-only certificate；
11. no dynamic stopping；
12. no proof that branching factor 9 is best。

這些應視為：

$$
\boxed{
\text{Research TODO}
}
$$

而不是 footnote。

---

# 66. v0.2 的合理順序

現有路線：

1. graph batching；
2. causal mode；
3. virtual representatives；
4. multi-route labels；
5. hierarchical synthetic benchmark；
6. signed route certificate。

這個順序是合理的。

因為它先補：

$$
\boxed{
\text{Correctness / Systems Foundations}
}
$$

再談：

$$
\boxed{
\text{Large-Scale Intelligence}.
}
$$

---

# 67. 本文再增加一個 v0.3 建議：Hardware Profile

未來每個 route graph 應除：

$$
|E|
$$

外，記錄：

- degree distribution；
- contiguous edge blocks；
- gather locality；
- block occupancy；
- index bytes；
- kernel launches；
- HBM traffic estimate。

因為：

$$
\boxed{
GraphSparse
\neq
HardwareSparse.
}
$$

---

# 68. Routing Layout Compilation

未來可以定義：

$$
G
\xrightarrow{\Lambda}
G_{hw}
$$

其中：

$$
\Lambda
$$

把 semantic route graph 重新編譯成：

- block sparse；
- tiled sparse；
- grouped query blocks；
- fused neighborhood batches。

只要保持：

$$
\boxed{
SemanticEdgePreservation.
}
$$

就可以讓：

$$
\text{semantic topology}
$$

和：

$$
\text{hardware topology}
$$

分開優化。

---

# 69. 這和 MLF 的 Projection Philosophy 再次接上

MLF：

$$
CanonicalStructure
\rightarrow
Projection.
$$

MMR-IFN 可以類比：

$$
SemanticRouteGraph
\rightarrow
HardwareRouteProjection.
$$

因此：

$$
\boxed{
\text{Model topology}
\neq
\text{kernel layout}.
}
$$

只要兩者可驗證對應。

---

# 70. Route Certificate 應進一步升級

下一版可加入：

- input hash；
- model checkpoint hash；
- router version；
- semantic graph hash；
- compiled hardware graph hash；
- kernel version；
- environment；
- output digest；
- Ed25519 signature。

形成：

$$
\boxed{
\text{Model Interaction Certificate}.
}
$$

---

# 71. Certificate 不是可解釋性萬靈丹

即使知道：

$$
i\rightarrow j
$$

有 attention edge，

也不代表：

> token $j$ 就是 token $i$ 做出決策的唯一原因。

所以：

$$
\boxed{
InteractionTrace
\neq
CausalExplanation.
}
$$

這點必須和 mechanistic interpretability 分開。

---

# 72. 但它比完全不保存 interaction 強

如果 dynamic routing 每次都改，

卻完全不記：

$$
E_t,
$$

那麼：

- replay；
- debugging；
- comparison；
- failure analysis；

都會困難。

所以：

$$
\boxed{
\text{Route Provenance}
}
$$

本身仍有工程價值。

---

# 73. 三個真正值得驗證的優勢

本文認為 MMR-IFN 最值得測的不是「更像大腦」，而是三件具體事。

## A — Structured Recall

階層／依賴任務上：

$$
RouteRecall\uparrow.
$$

## B — Sparse Cost

在 quality parity 下：

$$
Memory/Compute\downarrow.
$$

## C — Auditable Interaction

能回答：

$$
\boxed{
\text{why these nodes were allowed to interact}.
}
$$

---

# 74. 三個最可能失敗的地方

## F-A — Hierarchy Is Wrong

真實語言關係不是穩定九元 tree。

## F-B — Routing Cost Dominates

建圖、gather、scatter 比 dense optimized attention 更慢。

## F-C — Sparse Recall Loss

重要長距離 edge 不在：

$$
E.
$$

則：

$$
\boxed{
\text{No attention weight can recover a missing edge}.
}
$$

這是 sparse model 最本質的風險。

---

# 75. Missing-edge Error

對 task 必要關係集合：

$$
E^\ast,
$$

router 給：

$$
E.
$$

定義：

$$
Recall_E
=
\frac{|E\cap E^\ast|}{|E^\ast|}.
$$

如果：

$$
Recall_E\ll1,
$$

attention 再聰明也救不了。

因此：

$$
\boxed{
Routing Recall
}
$$

應是核心指標。

---

# 76. Over-routing Error

反過來，

若：

$$
|E|
\rightarrow
n^2,
$$

則：

$$
\boxed{
SparseBenefit\rightarrow0.
}
$$

所以需要：

$$
\boxed{
Recall-Cost Frontier.
}
$$

---

# 77. Route Quality Frontier

可定義：

$$
\boxed{
\mathcal F_{route}
=
(
Recall_E,
|E|,
Latency,
Memory,
TaskQuality
).
}
$$

MMR-IFN 真正要證明的是：

> 在某些 structured tasks 上，它形成比 baseline 更好的 Pareto frontier。

---

# 78. 和 learned sparse 的真正比較

如果 NSA 類模型能直接從 query/key content 找到重要 tokens，

而 MMR-IFN 需要昂貴外部結構，

則後者必須證明至少一項：

1. 更高 recall；
2. 更低 route cost；
3. 更強 structured inductive bias；
4. 更穩定；
5. 更可稽核；
6. 更少 training data。

否則沒有必要。

---

# 79. Matrix-Native 不等於 Matrix-Only

MMR-IFN 仍保留：

$$
p_i.
$$

也可保留：

- sequence embeddings；
- semantic embeddings；
- learned shortcuts。

所以：

$$
\boxed{
MatrixNative
\neq
SequenceRejected.
}
$$

它更像：

$$
\boxed{
Sequence
+
Structure
+
Route.
}
$$

---

# 80. 本系列最初問題的最強回答

最早：

> Excel／matrix 能不能變成 AI？

到這裡，最強但仍誠實的答案是：

$$
\boxed{
\textbf{
矩陣結構本身可以不只儲存資料，
而能被編譯成神經模型真正的可交互拓撲。
}
}
$$

這已經超過：

- CSV state；
- spreadsheet UI；
- control plane；
- data format。

---

# 81. 但「大模型」仍然沒有被證明

目前 MMR-IFN 只是：

$$
\boxed{
\text{Minimal Executable Sparse Transformer Backend}.
}
$$

不是：

$$
\boxed{
\text{Trained Large Language Model}.
}
$$

所以早期「Excel 大模型」直覺到這裡只能說：

> 路徑已經接到 model architecture。

還不能說：

> 大模型已完成。

---

# 82. 可證偽條件

## F1 — Quality Collapse

相同 compute budget 下：

$$
TaskQuality_{IFN}
<
TaskQuality_{baseline}
$$

且無 audit benefit 抵償。

---

## F2 — Route Recall Failure

重要依賴 edge recall 長期低於可用門檻。

---

## F3 — Hardware Loss

理論：

$$
|E|\ll n^2
$$

但 GPU wall-clock：

$$
T_{IFN}\ge T_{dense}.
$$

---

## F4 — Construction Dominates

$$
T_{route}
\gg
T_{attn\ saved}.
$$

---

## F5 — Hierarchy No Gain

在真正 hierarchical tasks 上：

$$
IFN
\approx
FlatSparse
$$

但成本更高。

---

## F6 — Root Shortcut Cheat

模型只依賴 global root，

hierarchical routes 被完全忽略。

---

## F7 — Learned Address Collapse

大部分 token：

$$
\alpha_i
$$

集中少數地址，

導致負載失衡。

---

## F8 — No Replay

相同 input / router / version 無法重建同一 route graph。

---

## F9 — Certificate Theater

route certificate 雖存在，

但無法支援 debugging / audit / reproduction。

---

# 83. 本文結論

MMR-IFN Transformer v0.1 的真正成果不是：

$$
\boxed{
\text{Sparse Transformer Victory}.
}
$$

而是：

$$
\boxed{
\text{Research Object Closure}.
}
$$

也就是：

$$
\boxed{
\text{Sequence}
\rightarrow
\text{Hierarchical Address}
\rightarrow
\text{Semantic Route Graph}
\rightarrow
\text{Edge-list Attention}
\rightarrow
\text{Verifiable Route Record}.
}
$$

此 pipeline 已經可以：

- construct；
- execute；
- compare with dense same-graph reference；
- backpropagate；
- measure；
- hash；
- tamper-test。

因此：

$$
\boxed{
\textbf{
「矩陣原生智能」已經從表示論問題，
推進成可以被模型實驗直接否證的架構問題。
}
}
$$

但真正重要的下一階段不是繼續堆概念，而是：

$$
\boxed{
\text{Train}
+
\text{Compare}
+
\text{Profile}
+
\text{Ablate}
+
\text{Fail}.
}
$$

如果 MMR-IFN 在 hierarchical / dependency-rich tasks 上不能形成更好的：

$$
Quality
-
Cost
-
Auditability
$$

折衷，

那它就應退回：

$$
\boxed{
\text{一種可稽核 sparse routing schema}.
}
$$

如果它能，

那麼本系列最早那個看似奇怪的問題——

> 「表格／矩陣能不能成為 AI 架構的一部分？」

才會第一次得到真正強的實驗答案。

---

# 84. 下一篇

## EML-MNIAC-2026-09

**《從智能矩陣到 Agent 控制平面：可見狀態、任務帳本與人機共同操作》**

下一篇將轉向 Agent / Operations 支線，整合：

- EML-LQ Agent；
- D-ALAN；
- Veritaxa Workbench；
- Spreadsheet Control Plane；
- MLF governed inference；
- PHOSPHOR control；
- AI proposal vs canonical promotion；
- human review；
- multi-Agent shared ledger。

---

# 參考資料

## 內部理論與工程

1. EveMissLab, 《無限分形九宮格：基於多向矩陣表徵的遞歸計算拓撲》, 2026-07-23.
2. EveMissLab, 《MMR-IFN Transformer：基於遞歸多向路由的層級稀疏注意力架構》v0.1, 2026-07-23.
3. EveMissLab, `MMR-IFN-Transformer_v0.1_完整實作包`.
4. EML-MNIAC-2026-04, 《從二維表格到多向矩陣帳本》.
5. EML-MNIAC-2026-06, 《AI Matrix Ledger Format》.
6. EML-MNIAC-2026-07, 《單一狀態、多重投影》.

## 外部技術近鄰

7. Beltagy, I., Peters, M. E., & Cohan, A. (2020). **Longformer: The Long-Document Transformer.** arXiv:2004.05150.
8. Zaheer, M. et al. (2020). **Big Bird: Transformers for Longer Sequences.** arXiv:2007.14062.
9. Roy, A., Saffar, M., Vaswani, A., & Grangier, D. (2020). **Efficient Content-Based Sparse Attention with Routing Transformers.** arXiv:2003.05997.
10. Veličković, P. et al. (2017/2018). **Graph Attention Networks.** arXiv:1710.10903.
11. Dao, T., Fu, D. Y., Ermon, S., Rudra, A., & Ré, C. (2022). **FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness.** arXiv:2205.14135.
12. Yuan, J. et al. (2025). **Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse Attention.** arXiv:2502.11089.
13. Fedus, W., Zoph, B., & Shazeer, N. (2021). **Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity.** arXiv:2101.03961.

這些工作共同說明 sparse attention 已有成熟的 local、global、random、content-based、graph-neighborhood、hardware-aware 與 learned-routing 研究脈絡。因此 MMR-IFN 的新增價值不能只來自「稀疏」兩字；它必須由可追溯 hierarchy、route typing、cross-representation lineage、route certificates，以及在真正具有結構的任務上可測的 quality–cost–auditability frontier 來證明。

---

**系列狀態：** 第 08 篇完成。  
**下一篇：** EML-MNIAC-2026-09 —《從智能矩陣到 Agent 控制平面：可見狀態、任務帳本與人機共同操作》
