← Archive
lm-002800 · 2026-08

比較域失效

下載 MD 檔 ⬇

title: "比較域失效:當異質存在不再適合同一排行榜" english_title: "Comparison-Domain Failure: When Heterogeneous Existences No Longer Belong on a Single Leaderboard" series: "異質存在、體驗域與技術適配研究系列" series_english: "Heterogeneous Existence, Experiential Domains, and Technological Fit Series" series_id: "HEETF" paper_id: "HEETF-06" author: "Neo.K" organization: "EveMissLab" version: "0.1.0" status: "Research Draft / Comparison-Domain Theory" date: "2026-08-14" language: "zh-TW"

比較域失效

當異質存在不再適合同一排行榜

Comparison-Domain Failure

When Heterogeneous Existences No Longer Belong on a Single Leaderboard

作者: Neo.K
機構: EveMissLab
系列: 異質存在、體驗域與技術適配研究系列(HEETF),Paper 06
版本: v0.1.0
日期: 2026-08-14


摘要

本文研究 HEETF 系列前五篇自然導出的問題:

當不同存在具有不同的體驗域、生成路徑、成本函數與載體拓撲時,「誰比較強」是否仍然是一個天然有定義的問題?

本文答案是:

No comparison domainNo well-typed ranking claim.\boxed{ \text{No comparison domain} \Rightarrow \text{No well-typed ranking claim}. }

「A 比 B 強」不是一個自動帶有完整語義的二元關係。它至少需要聲明:

  • 比較哪一類 task;
  • 哪些 outcome metrics;
  • metric 方向;
  • 哪些成本;
  • single instance 還是 fleet;
  • marginal 還是 lifecycle;
  • raw output 還是 validated output;
  • 是否納入 agency、self-arrival、subjective welfare;
  • 如何 normalization;
  • 如何 aggregation;
  • 權重由誰決定;
  • 哪些 metrics 對兩個被比較存在都適用。

本文因此定義 Comparison Domain

C=(P,S,M,d,N,Σ,F,Ω).\boxed{ \mathfrak C = ( \mathcal P, \mathcal S, \mathbf M, \mathbf d, N, \Sigma, F, \Omega ). }

其中:

  • P\mathcal P:task / problem family;
  • S\mathcal S:被比較 subject / system set;
  • M=(m1,,mk)\mathbf M=(m_1,\ldots,m_k):比較指標;
  • d\mathbf d:每個指標的方向與型別;
  • NN:normalization / scale contract;
  • Σ\Sigma:scope / cost / instance contract;
  • FF:aggregation / ordering rule;
  • Ω\Omega:觀察邊界與證據要求。

對存在 ii,其 profile:

xiC=(m1(i),,mk(i)).\boxed{ \mathbf x_i^{\mathfrak C} = ( m_1(i),\ldots,m_k(i) ). }

若所有 metrics 已朝「越大越好」定向,定義 Pareto dominance:

iPj\boxed{ i\succeq_P j }

若:

mr(i)mr(j)r,m_r(i)\ge m_r(j) \qquad \forall r,

且至少一項嚴格較大時為 strict dominance。

第一個核心結果為 Pareto Crossing Incomparability Proposition。若存在 p,qp,q

mp(i)>mp(j),m_p(i)>m_p(j),

但:

mq(i)<mq(j),m_q(i)<m_q(j),

則在純 Pareto ordering 下:

i̸Pjj̸Pi.\boxed{ i\not\succeq_P j \quad\text{且}\quad j\not\succeq_P i. }

也就是:

iPj.\boxed{ i\parallel_P j. }

這不是資訊不足,而是偏序本身合法地回傳:

不可由 dominance 決定勝負。

第二個核心結果為 Weight-Reversal Theorem。在 normalized finite metric space 中,若 i,ji,j profile crossing,且 scalar score 採正權重線性 aggregation:

Sw(i)=r=1kwrmr(i),wr>0,S_{\mathbf w}(i) = \sum_{r=1}^k w_r m_r(i), \qquad w_r>0,

則一般可以找到兩組合法正權重:

w,w\mathbf w, \mathbf w'

使:

Sw(i)>Sw(j)\boxed{ S_{\mathbf w}(i) > S_{\mathbf w}(j) }

而:

Sw(i)<Sw(j).\boxed{ S_{\mathbf w'}(i) < S_{\mathbf w'}(j). }

因此:

No Weight-Free Total Ranking Corollary

對 crossing profiles,單靠 profile data 本身不能唯一推出一個線性總排行榜。要把偏序壓成 total order,必須增加:

value / aggregation assumptions.\boxed{ \text{value / aggregation assumptions}. }

所以:

「誰比較強?」

若沒有權重與 aggregation contract,很多時候不是「答案還沒算出來」,而是:

排序公理尚未提供.\boxed{ \text{排序公理尚未提供}. }

第三個核心結果為 Ordinal Aggregation Non-Invariance Proposition。如果某 metric 只有 ordinal information,則任意 strictly increasing transformation 都保留其 ordinal meaning;但對 ordinal scores 直接做加總,一般不在此 transformation class 下保持排名。因此把純 ordinal ranks 當 cardinal quantities 相加,若沒有額外 measurement structure,可能產生不具 representation-invariant 意義的排行榜。

這與 representational measurement theory 的經典原則一致:數字表徵可以允許特定 transformations,而「有意義的比較」應在允許 transformation 下保持不變。本文不重新發明 measurement theory,而把這個要求引入 heterogeneous-existence comparison。

第四個核心結果為 Metric Applicability Failure Proposition。如果 metric mrm_r 對 subject ii 有定義,但對 jj 無定義:

mr(i)Xr,mr(j)=,m_r(i)\in X_r, \qquad m_r(j)=\bot,

則沒有額外 imputation / surrogate / omission rule 時,完整 scalar score:

F(m1,,mk)F(m_1,\ldots,m_k)

可能無法對兩者共同評估。

典型例子:

  • human lived experience;
  • phenomenal pleasure;
  • autobiographical self-arrival;

若 AI 是否具有相應 phenomenal state 未被建立,就不能為了「榜單完整」而無聲填入一個任意數值。

第五個結果為 Scope-Mismatch Proposition。比較:

one human\text{one human}

與:

1000-model fleet\text{1000-model fleet}

若未聲明 instance scope,即使數據皆正確,排名仍可能犯 unit mismatch。HEETF-05 已區分 single-instance throughput 與 aggregate replicated throughput;本篇將其提升為 comparison typing rule:

same metric name⇏same comparison unit.\boxed{ \text{same metric name} \not\Rightarrow \text{same comparison unit}. }

本文因此提出八種 Comparison-Domain Failure Modes

  1. Task Mismatch:比較的其實不是同一 task;
  2. Metric Applicability Failure:metric 不適用所有 subjects;
  3. Scale Failure:measurement scale 不允許所採 aggregation;
  4. Scope Mismatch:single / fleet、marginal / lifecycle 混用;
  5. Pareto Crossing:多維 profile 互有優劣;
  6. Aggregation Underdetermination:缺少權重/ordering rule;
  7. Objective Mismatch:被比較者追求的目標函數不同;
  8. Observer / Policy Mismatch:public benchmark score 與 subject's lived objective 不同。

既有 multi-criteria decision analysis、many-objective optimization 與 composite-indicator literature 已廣泛研究 weighting、aggregation、robustness、Pareto fronts 與方法選擇。Greco 等人對 composite indices 的 review 特別指出 weighting、aggregation 與 robustness 是 index construction 的核心問題;Cinelli 等人則建立 MCDA methods 的 comprehensive taxonomy;many-objective optimization literature 也長期指出,在 objective 數量增加時,Pareto dominance 可能對大量 solutions 失去足夠 discrimination power。研究評估領域的 The Metric Tide 與 Adler–Harzing 對 academic rankings 的批判也提供一個制度性提醒:將複雜表現壓成單一 metric / rank 會產生治理、解釋與激勵問題。

本文不因此宣稱:

排名沒有用。

相反,本文提出:

Typed Ranking Principle

一個排行榜只有在:

C\boxed{ \mathfrak C }

被清楚聲明時,才是一個可解讀的工程/科學/制度物件。

例如完全可以有:

  • Human-only Mathematics
  • Open Mathematics
  • Human × AI Mathematics
  • Autonomous AI Mathematics

每一個都可以有合法排行榜。

但:

rankC1\boxed{ \text{rank}_{\mathfrak C_1} }

不應被無聲提升成:

rankall existence.\boxed{ \text{rank}_{\mathrm{all\ existence}}. }

本文還提出一個重要的 Comparison Exit Principle。若某個體 ii 的真實 experiential objective:

UiU_i

與排行榜 score:

SCS_{\mathfrak C}

沒有 order-preserving relation,則:

「我不參與這個排行榜」

可以是一個理性 domain choice,而不必解釋為:

  • denial;
  • cowardice;
  • inability;
  • irrationality。

例如一名數學家真正最大化:

Ui=Vitruth+Viself-arrival,\boxed{ U_i = V_i^{\mathrm{truth}} + V_i^{\mathrm{self\text{-}arrival}}, }

而 public leaderboard 只最大化:

S=proofs per year.\boxed{ S = \text{proofs per year}. }

那麼:

maxS\boxed{ \max S }

本來就不等於:

maxUi.\boxed{ \max U_i. }

他退出「最快產出」排行榜,並不等於否定 proof truth。

同樣,作家可以拒絕與一個可並行生成一萬部小說的 AI fleet 比「每小時總字數」,因為那不是他的創作 objective。

因此本文最終提出:

Comparison-Domain Adequacy Principle

只有當比較域:

  1. 指標適用;
  2. 尺度可用;
  3. scope 一致;
  4. aggregation 明示;
  5. objective 相關;

時,「誰比較強」才是一個足夠完整的問題。

否則合法答案可以不是:

A>BA>B

也不是:

B>A,B>A,

而是:

ACB\boxed{ A\parallel_{\mathfrak C}B }

或更直接:

comparison-domain failure.\boxed{ \text{comparison-domain failure}. }

這不是拒絕比較。

而是拒絕一個沒有完整型別的比較。

關鍵詞: Comparison Domain, Incomparability, Pareto Dominance, Multi-Criteria Decision Analysis, Composite Indicator, Ranking, Measurement Theory, Heterogeneous Intelligence, Benchmarking, Human–AI Comparison


1. 「誰比較強?」其實少了很多字

問:

人類和 AI,誰比較強?

這句話至少省略:

  • 做什麼?
  • 哪個 AI?
  • 哪個人?
  • 單一還是群體?
  • 比速度還是品質?
  • 算不算成本?
  • 算不算體驗?
  • 算不算 agency?
  • 取平均還是極值?

所以它不是完整 proposition。


2. 比較必須 Typed

定義:

C=(P,S,M,d,N,Σ,F,Ω).\boxed{ \mathfrak C = ( \mathcal P, \mathcal S, \mathbf M, \mathbf d, N, \Sigma, F, \Omega ). }

3. Task Family

P\mathcal P

指定:

  • theorem proving;
  • novel writing;
  • coding;
  • chess;
  • diagnosis;
  • text transcription。

跨 task 的「總強度」需要額外 aggregation。


4. Subject Set

S\mathcal S

指定被比的 unit:

  • one human;
  • expert human;
  • population mean;
  • one model;
  • one model instance;
  • fleet;
  • human–AI team。

5. Metrics

M=(m1,,mk).\mathbf M = (m_1,\ldots,m_k).

例如:

  • correctness;
  • quality;
  • latency;
  • cost;
  • energy;
  • throughput;
  • self-arrival;
  • autonomy;
  • reliability。

6. Metric Direction

不同 metric 方向不同。

例如:

  • accuracy 越大越好;
  • latency 越小越好;
  • cost 越小越好。

所以先由:

d\mathbf d

轉成統一 orientation。


7. Normalization

NN

指定:

  • min-max;
  • z-score;
  • ratio;
  • raw units;
  • domain-specific transformation。

不同 normalization 可改變 aggregate score。


8. Scope Contract

Σ\Sigma

至少聲明:

  • marginal / lifecycle;
  • raw / validated;
  • single / fleet;
  • current / historical;
  • human review included / excluded。

9. Aggregator

FF

可以是:

  • Pareto partial order;
  • weighted sum;
  • lexicographic;
  • threshold rule;
  • outranking;
  • nonlinear utility。

沒有 FF,通常只有 profile。


10. Observation Boundary

Ω\Omega

指定:

  • evidence window;
  • uncertainty;
  • missing data;
  • verifier;
  • confidence。

11. Performance Profile

對存在 ii

xiC=(m1(i),,mk(i)).\boxed{ \mathbf x_i^{\mathfrak C} = ( m_1(i), \ldots, m_k(i) ). }

12. Profile 不是 Ranking

兩個 profiles:

xi,xj\mathbf x_i, \mathbf x_j

可以完全知道,

但仍沒有 total order。


13. Pareto Dominance

假設所有 metrics 已轉成越大越好。

定義:

iPji\succeq_P j

若:

mr(i)mr(j)m_r(i)\ge m_r(j)

對所有 rr

若至少一項嚴格:

iPj.i\succ_P j.

14. Pareto Crossing

若:

mp(i)>mp(j)m_p(i)>m_p(j)

但:

mq(i)<mq(j),m_q(i)<m_q(j),

稱 profile crossing。


15. Pareto Crossing Incomparability Proposition

命題 15.1

若 profile crossing,則:

i̸Pj\boxed{ i\not\succeq_P j }

且:

j̸Pi.\boxed{ j\not\succeq_P i. }

證明

iPji\succeq_P j 要求所有 metrics 上 ii 不差於 jj

但已知:

mq(i)<mq(j),m_q(i)<m_q(j),

故失敗。

同理:

mp(j)<mp(i),m_p(j)<m_p(i),

jPij\succeq_P i 也失敗。

因此:

iPj.\boxed{ i\parallel_P j. } \boxed{\square}

16. 不可比不是不知道

這裡不是:

數據不夠。

而是:

在 Pareto rule 下,兩者互有優劣。

所以偏序合法地不給 total ranking。


17. 最簡單例子

人類 H:

(Q=10,R=2).( Q=10, R=2 ).

AI A:

(Q=8,R=100).( Q=8, R=100 ).

QQ 是品質、 RR 是速度,

H 品質高。

A 速度高。

Pareto 下:

HPA.\boxed{ H\parallel_P A. }

18. 想排第一第二就要加規則

例如:

S=wQQ+wRR.S = w_Q Q + w_R R.

這已經加入 value weights。


19. Linear Scalarization

定義:

Sw(i)=r=1kwrmr(i)\boxed{ S_{\mathbf w}(i) = \sum_{r=1}^k w_r m_r(i) }

其中:

wr>0.w_r>0.

20. Weight-Reversal Theorem

定理 20.1

假設 profiles finite 且 crossing:

mp(i)>mp(j),m_p(i)>m_p(j), mq(i)<mq(j).m_q(i)<m_q(j).

則存在正權重向量:

w\mathbf w

與:

w\mathbf w'

使:

Sw(i)>Sw(j),S_{\mathbf w}(i)>S_{\mathbf w}(j),

但:

Sw(i)<Sw(j).S_{\mathbf w'}(i)<S_{\mathbf w'}(j).

證明

令:

Δr=mr(i)mr(j).\Delta_r = m_r(i)-m_r(j).

已知:

Δp>0,\Delta_p>0, Δq<0.\Delta_q<0.

選擇 w\mathbf w 令:

wp=Mw_p=M

且其餘 positive weights 固定為小值 ϵ>0\epsilon>0

因所有 Δr\Delta_r finite,取 MM 足夠大,有:

rwrΔr>0.\sum_r w_r\Delta_r>0.

所以:

Sw(i)>Sw(j).S_{\mathbf w}(i)>S_{\mathbf w}(j).

反之,選:

wq=Mw_q=M'

且其餘為 ϵ\epsilon

MM' 足夠大,因:

Δq<0,\Delta_q<0,

可得:

rwrΔr<0.\sum_r w'_r\Delta_r<0.

因此:

Sw(i)<Sw(j).S_{\mathbf w'}(i)<S_{\mathbf w'}(j). \boxed{\square}

21. No Weight-Free Total Ranking Corollary

profile crossing 時,

若沒有額外指定:

w,\boxed{ \mathbf w, }

資料本身不唯一決定 weighted-sum total order。


22. 這就是「誰比較強?」少掉的排序公理

如果你問:

品質重要還是速度重要?

這不是被 metric data 自己回答。

需要:

preference / policy / objective.\boxed{ \text{preference / policy / objective}. }

23. 不只 Weighted Sum

其他 aggregation:

  • lexicographic;
  • maximin;
  • threshold;
  • outranking;

也都加入 additional rule。

所以 total order 從來不是免費生成。


24. Composite Index Literature

composite-indicator research 長期研究:

  • weighting;
  • aggregation;
  • sensitivity;
  • robustness。

同一 raw indicators 經不同 method,rankings 可能不同。

本文把這個老問題帶入 human–AI comparison。


25. Metric Scale 也有限制

一個數字:

77

不一定可以當 cardinal quantity。

有時它只是:

第 7 名。


26. Ordinal Scale

若:

mm

只有 ordinal meaning,

則任何 strictly increasing:

gg

都保留:

m(a)>m(b)    g(m(a))>g(m(b)).m(a)>m(b) \iff g(m(a))>g(m(b)).

27. 但 Sum 可能不保留

例如兩 metrics m1,m2m_1,m_2 都只是 ordinal。

原 score:

S=m1+m2.S=m_1+m_2.

對某兩對象可有:

S(a)>S(b).S(a)>S(b).

對其中一個 metric 做合法 monotonic transformation:

g(m1),g(m_1),

新的:

S=g(m1)+m2S' = g(m_1)+m_2

可能反轉 ranking。


28. Ordinal Aggregation Non-Invariance Proposition

命題 28.1

ordinal metrics 的 raw numerical sums 一般不 invariant under admissible strictly monotonic transformations。

證明

取:

a=(2,1),a=(2,1), b=(1,2).b=(1,2).

原本:

S(a)=3,S(a)=3, S(b)=3.S(b)=3.

令:

g(x)=x3.g(x)=x^3.

則:

S(a)=8+1=9,S'(a)=8+1=9, S(b)=1+2=3.S'(b)=1+2=3.

同樣 ordinal ordering 被保留,

但 aggregate relation 改變。

也可以構造原本 strict ranking 後反轉之例。

故 raw sum 並非 ordinal-invariant。

\boxed{\square}

29. Measurement Theory 的重要邊界

所以:

numbers\boxed{ \text{numbers} }

不自動意味:

legitimate arithmetic.\boxed{ \text{legitimate arithmetic}. }

measurement contract 決定哪些 operations meaningful。


30. Metric Applicability

即使 scale 沒問題,

metric 也可能根本不適用。


31. Undefined Metric

用:

\bot

表示未定義。

例如:

mphen(H)=human reported experience,m_{\mathrm{phen}}(H) = \text{human reported experience},

但若 AI phenomenal status 未建立:

mphen(A)=.m_{\mathrm{phen}}(A) = \bot.

32. Metric Applicability Failure Proposition

若 aggregation:

FF

要求所有 kk metrics,

但:

mr(j)=,m_r(j)=\bot,

則沒有額外 missing-data rule 時:

F(xj)\boxed{ F(\mathbf x_j) }

未定義。


33. 不能為了排行榜完整就偷偷填 0

因為:

unknown0.\boxed{ \text{unknown} \neq 0. }

這和 SEHTS 的 unknown-preservation 精神一致。


34. Omission 也不是中性的

如果直接刪掉 metric,

comparison domain 已改變:

CC.\mathfrak C \rightarrow \mathfrak C'.

排名可能因此改變。


35. Surrogate 也要聲明

例如用:

  • self-report;
  • behavior;
  • physiological proxy;

替代 lived experience,

必須明示:

proxytarget identity.\boxed{ \text{proxy} \neq \text{target identity}. }

36. Scope Mismatch

HEETF-05 已證明:

R(1)R(n).R^{(1)} \neq R^{(n)}.

所以:

一個作家 vs 一千個 AI replicas

不能假裝是同一 instance unit。


37. Scope-Mismatch Proposition

若 metric:

mm

ii 測 single-instance,

jj 測 aggregate-fleet,

則:

m(i),m(j)m(i),m(j)

雖同名,仍不是同 scope quantity。


38. Marginal / Lifecycle 也一樣

HEETF-04:

CmargClife.C^{\mathrm{marg}} \neq C^{\mathrm{life}}.

所以:

人類算十年教育,AI 只算 inference

是一個 scope mismatch。

反過來也一樣。


39. Raw / Validated 也一樣

HEETF-05:

RrawRvalid.R^{\mathrm{raw}} \neq R^{\mathrm{valid}}.

不能:

  • AI 算 raw token;
  • 人類算出版作品;

然後叫同一 throughput。


40. Objective Mismatch

有時 metrics 都可測,

但兩個 subject 根本不追求相同 objective。


41. 數學家 A

UA=Vtruth.\boxed{ U_A = V^{\mathrm{truth}}. }

越快解 theorem 越好。


42. 數學家 B

UB=Vtruth+Vself-arrival.\boxed{ U_B = V^{\mathrm{truth}} + V^{\mathrm{self\text{-}arrival}}. }

他還想自己抵達。


43. 同一排行榜不能自動代表兩者

如果榜單:

S=proofs/year,S = \text{proofs/year},

可能高度代表:

UAU_A

卻不代表:

UB.U_B.

44. Score–Utility Alignment

定義:

SCS_{\mathfrak C}

為 public comparison score。

若對 subject ii

SC(a)>SC(b)S_{\mathfrak C}(a) > S_{\mathfrak C}(b)

總能推出:

Ui(a)Ui(b),U_i(a) \ge U_i(b),

稱 score 對 ii 至少 order-compatible。


45. Misalignment

若存在:

a,ba,b

使:

SC(a)>SC(b),S_{\mathfrak C}(a) > S_{\mathfrak C}(b),

但:

Ui(a)<Ui(b),U_i(a) < U_i(b),

則:

leaderboard improvement\boxed{ \text{leaderboard improvement} }

不等於:

subject welfare improvement.\boxed{ \text{subject welfare improvement}. }

46. Comparison Exit Principle

若 leaderboard score 與 subject objective 明顯 misaligned,

則:

我不參加這個比較。

可以是理性 choice。


47. 這不是「輸不起」

當然有些退出可能真的來自 self-protection。

但 framework 不允許從:

exit\boxed{ \text{exit} }

直接推出:

cowardice.\boxed{ \text{cowardice}. }

因為比較域可能真的和 objective 無關。


48. 最簡單例子:挖土機

人不會跟挖土機比:

土方量 / hour.\boxed{ \text{土方量 / hour}. }

然後因輸掉而 identity collapse。

因為:

comparison class 已分離.\boxed{ \text{comparison class 已分離}. }

49. AI 的衝擊在於它進入 identity domains

AI 不只挖土。

它開始:

  • 寫;
  • 畫;
  • 算;
  • 證;
  • 設計;
  • coding。

這些領域本來就是 human identity leaderboard。


50. 所以比較域重建成為制度問題

未來可以分:

Human-Only Domain

CH.\mathfrak C_H.

Open Domain

CO.\mathfrak C_O.

Human × AI Domain

CHA.\mathfrak C_{HA}.

Autonomous AI Domain

CA.\mathfrak C_A.

51. 同一成果可以進不同 Ranking

一個 proof 可以在:

CO\mathfrak C_O

算 public mathematical achievement。

但不一定在:

CH\mathfrak C_H

算 human-only competition achievement。

完全不矛盾。


52. Public Truth 與 Competition Rule 分開

truth statuscompetition eligibility.\boxed{ \text{truth status} \neq \text{competition eligibility}. }

這也接 SEHTS-05 的 semantic / policy separation。


53. Ranking Domain Lift 是危險操作

若:

iC1j,i \succ_{\mathfrak C_1} j,

不推出:

iC2j.\boxed{ i \succ_{\mathfrak C_2} j. }

更不推出:

i>j as existences.\boxed{ i > j \text{ as existences}. }

54. Domain-Lift Invalidity Proposition

若:

F1F2F_1 \neq F_2

或 metrics / scopes 不同,

C1\mathfrak C_1 中的 ranking relation 一般不保證在 C2\mathfrak C_2 保持。


55. 「最強 AI」也是 Typed

至少要說:

  • coding;
  • reasoning;
  • latency;
  • price;
  • context;
  • tool use;
  • reliability;
  • safety。

不同榜單可以有不同 winner。


56. 「最強人類」更荒謬

一個世界冠軍棋手:

不等於:

  • 最強小說家;
  • 最強外科醫師;
  • 最快短跑者。

大家本來就知道 domain typing。

AI 只是讓我們忘記這個常識。


57. 多 objective 越多,Pareto discrimination 越弱

many-objective optimization literature 長期指出:

當 objective 維度增加,

大量 solutions 可能互不 dominance。

這不是 bug。

而是:

high-dimensional tradeoff structure.\boxed{ \text{high-dimensional tradeoff structure}. }

58. 所以「更多 metrics」不一定產生更清楚排名

反而可能:

comparability.\boxed{ \text{comparability} \downarrow. }

59. 強行產生總排行需要更強 value assumptions

例如:

w.\boxed{ \mathbf w. }

metrics 越多,

權重政治/價值問題越重要。


60. Composite Index Robustness

如果小幅:

  • weight changes;
  • normalization changes;
  • missing-data policy changes;

就造成巨大 rank reversal,

榜單應揭露:

ranking fragility.\boxed{ \text{ranking fragility}. }

61. Ranking Robustness Set

可定義 admissible contracts:

F.\mathcal F.

對兩存在:

i,j,i,j,

若:

iFji\succ_F j

對所有:

FF,F\in\mathcal F,

才稱 robust dominance over contract family。


62. Contract-Robust Ranking

定義:

irobj\boxed{ i \succ_{\mathrm{rob}} j }

若在所有 admissible comparison contracts 下:

ij.i \succ j.

這是一個比單一排行榜更強的 claim。


63. 如果 ranking 隨權重一直翻轉?

那真正結果就是:

tradeoff-sensitive.\boxed{ \text{tradeoff-sensitive}. }

而不是硬挑一個永久 winner。


64. 公開排行榜應該披露什麼?

至少:

  • task set;
  • metric definitions;
  • scale;
  • normalization;
  • weights;
  • scope;
  • missing-data handling;
  • variance;
  • uncertainty;
  • rank sensitivity。

65. The Metric Tide 的制度提醒

research assessment literature 已經反覆提醒:

metrics 可以幫助評估,

但不應讓單一數字無聲取代複雜 judgment。

本文把相同原則用到 heterogeneous intelligences。


66. Ranking 不是 Reality

排行榜是:

a projection of a declared domain.\boxed{ \text{a projection of a declared domain}. }

不是存在本身。


67. Comparison Domain Failure Modes

本文正式定義八類:

CDF-1 Task Mismatch

不同 task。

CDF-2 Metric Applicability Failure

metric 對部分 subject 無定義。

CDF-3 Scale Failure

不允許該 arithmetic / aggregation。

CDF-4 Scope Mismatch

single / fleet、marginal / lifecycle 混用。

CDF-5 Pareto Crossing

互有優劣。

CDF-6 Aggregation Underdetermination

缺少排序 rule / weights。

CDF-7 Objective Mismatch

leaderboard objective 與 subject objective 不同。

CDF-8 Observer / Policy Mismatch

public score 被誤當完整 lived / normative ranking。


68. Comparison-Domain Adequacy

定義:

Adeq(C;i,j)\boxed{ \operatorname{Adeq}(\mathfrak C;i,j) }

若至少滿足:

  1. shared task reference;
  2. applicable metrics;
  3. legitimate scale operations;
  4. matched scopes;
  5. explicit aggregation;
  6. relevant objective;
  7. evidence sufficient。

69. Comparison-Domain Failure

若:

Adeq=0,\operatorname{Adeq}=0,

則:

i>j\boxed{ i>j }

不是一個 adequately supported comparison claim。


70. 合法輸出可以是 Incomparable

iCj.\boxed{ i\parallel_{\mathfrak C}j. }

這不是失敗的資料分析。

有時這才是正確結果。


71. 合法輸出也可以是「Domain Undefined」

如果 metrics 根本不共用:

C undefined for (i,j).\boxed{ \mathfrak C \text{ undefined for }(i,j). }

比硬補數字更誠實。


72. 這對 AI/人類尤其重要

因為:

  • human lived experience;
  • AI throughput;
  • human authorship;
  • machine replication;
  • biological fatigue;

不是天然對稱 metrics。


73. Cross-Substrate Ranking Contract

若真的要比較,

至少要明示哪些 axes 被保留:

correctness+latency+cost\boxed{ \text{correctness} + \text{latency} + \text{cost} }

而不是偷偷說:

這就是整體智能。


74. Intelligence 本身可能就是 Vector

可以先寫:

Ii=(Ireason,Imemory,Ilearning,Iplanning,Isocial,Iembodied,).\boxed{ \mathbf I_i = ( I_{\mathrm{reason}}, I_{\mathrm{memory}}, I_{\mathrm{learning}}, I_{\mathrm{planning}}, I_{\mathrm{social}}, I_{\mathrm{embodied}}, \ldots ). }

是否能壓成:

IiRI_i\in\mathbb R

是額外問題。


75. 不做 Universal Intelligence Scalar

本文不提出:

I\boxed{ I^\ast }

作為 universal intelligence score。

因為那正好會重犯本篇警告。


76. Human Preference 不是 Benchmark Error

如果一個人說:

我不要跟 AI 比。

可能只是:

CAIhis objective domain.\boxed{ \mathfrak C_{\mathrm{AI}} \notin \text{his objective domain}. }

77. 但也不能拿 Personal Preference 否定 Open Benchmark

反過來:

我不想比,所以這個 AI benchmark 無效。

也錯。

只要:

C\mathfrak C

清楚,

benchmark 可以完全合法。


78. Domain Pluralism

所以未來最合理不是只有一個榜單,

而是:

{C1,,Cn}.\boxed{ \{ \mathfrak C_1, \ldots, \mathfrak C_n \}. }

多個清楚 typed comparison domains。


79. 比較域本身也是制度設計

誰決定:

  • metric;
  • weights;
  • eligibility;
  • scope;

就是在決定:

what counts as winning.\boxed{ \text{what counts as winning}. }

這不是純數據問題。


80. 新穎性邊界

本文不宣稱首次提出:

  • Pareto dominance;
  • multicriteria decision analysis;
  • multiattribute utility;
  • incomparability;
  • composite indices;
  • measurement theory;
  • rank sensitivity;
  • academic ranking criticism。

本文提出的是 HEETF 系列中的組合框架:

  1. Comparison Domain C\mathfrak C
  2. heterogeneous subject profile;
  3. Pareto Crossing Incomparability;
  4. Weight-Reversal Theorem;
  5. No Weight-Free Total Ranking;
  6. Metric Applicability Failure;
  7. Scope-Mismatch typing;
  8. Comparison Exit Principle;
  9. eight Comparison-Domain Failure modes;
  10. Comparison-Domain Adequacy contract。

81. 本文不證明什麼?

本文不證明:

所有人與 AI 都不可比較.\boxed{ \text{所有人與 AI 都不可比較}. }

不證明:

排行榜沒有價值.\boxed{ \text{排行榜沒有價值}. }

不證明:

所有 weighting 都任意.\boxed{ \text{所有 weighting 都任意}. }

不證明:

任何跨載體比較都不合法.\boxed{ \text{任何跨載體比較都不合法}. }

也不證明:

退出比較一定理性.\boxed{ \text{退出比較一定理性}. }

82. 真正的命題

本文只要求:

比較關係必須 typed。

如果你聲明:

C,\mathfrak C,

就可以合法說:

ACB.A \succ_{\mathfrak C} B.

但不要把它偷換成:

A>B in every meaningful sense.\boxed{ A>B \text{ in every meaningful sense}. }

83. 最終原則一

No Domain, No Ranking

¬C¬well-typed total ranking.\boxed{ \neg\mathfrak C \Rightarrow \neg\text{well-typed total ranking}. }

84. 最終原則二

Crossing Profiles Need Values

Pareto crossingtotal ranking requires extra aggregation assumptions.\boxed{ \text{Pareto crossing} \Rightarrow \text{total ranking requires extra aggregation assumptions}. }

85. 最終原則三

Unknown Is Not Zero

m(i)=0.\boxed{ m(i)=\bot \neq 0. }

86. 最終原則四

Leaderboard Rank Is Domain-Local

iC1j⇏iC2j.\boxed{ i\succ_{\mathfrak C_1}j \not\Rightarrow i\succ_{\mathfrak C_2}j. }

87. 最終原則五

Comparison Exit Can Be Legitimate

如果:

SCS_{\mathfrak C}

不表示:

Ui,U_i,

那不最佳化榜單不必是 irrational。


88. 結論

前五篇一路推到這裡:

HEETF-01:

technology progressindividual welfare progress.\boxed{ \text{technology progress} \neq \text{individual welfare progress}. }

HEETF-02:

same worldsame lived world.\boxed{ \text{same world} \neq \text{same lived world}. }

HEETF-03:

same outcomesame path / experience.\boxed{ \text{same outcome} \neq \text{same path / experience}. }

HEETF-04:

same tasksame cost function.\boxed{ \text{same task} \neq \text{same cost function}. }

HEETF-05:

same semantic tasksame carrier topology.\boxed{ \text{same semantic task} \neq \text{same carrier topology}. }

到了本篇:

different experiences+different paths+different costs+different carriers\boxed{ \text{different experiences} + \text{different paths} + \text{different costs} + \text{different carriers} }

不可能自動產生:

one natural leaderboard.\boxed{ \text{one natural leaderboard}. }

要排名,

你必須另外指定:

C.\boxed{ \mathfrak C. }

如果只比較:

同一 theorem 的 correctness,

人與 AI 可以非常清楚地比較。

如果比:

proof throughput,

也可以。

如果比:

lifecycle energy cost,

仍然可以。

但如果問:

「人類和 AI 到底誰比較強?」

卻沒有任何 domain,

那問題缺少的不是更多 benchmark。

而是:

comparison semantics.\boxed{ \text{comparison semantics}. }

而且當 profiles crossing 時,

即使所有數據都已經知道,

仍然可能只有:

HPA.\boxed{ H\parallel_P A. }

要把它改成:

H>AH>A

或:

A>H,A>H,

就必須加上:

  • weights;
  • values;
  • policy;
  • objective。

這不是作弊。

這本來就是 ranking formation 的一部分。

真正需要避免的只是:

把自己加入的 value assumptions 偽裝成數據自己說的。

因此:

「我不要跟 AI 比。」

有時候可能是逃避。

但有時候也只是非常精確地說:

這不是我要參加的 comparison domain.\boxed{ \text{這不是我要參加的 comparison domain}. }

就像人不需要因挖土機一小時挖得比自己一生多,就重新評估自己作為人的價值。

AI 真正特殊的是:

它開始進入人類原本最常拿來建立 identity 的領域。

所以未來不是停止比較。

而是學會:

把比較放回正確的域。

這就是 HEETF-06 的核心。


參考文獻

[1] Keeney, R. L., & Raiffa, H. (1976). Decisions with Multiple Objectives: Preferences and Value Tradeoffs. Wiley.

[2] Krantz, D. H., Luce, R. D., Suppes, P., & Tversky, A. (1971). Foundations of Measurement, Volume I: Additive and Polynomial Representations. Academic Press.

[3] Roberts, F. S. (1979). Measurement Theory with Applications to Decisionmaking, Utility, and the Social Sciences. Addison-Wesley.

[4] Greco, S., Ishizaka, A., Tasiou, M., & Torrisi, G. (2019). On the Methodological Framework of Composite Indices: A Review of the Issues of Weighting, Aggregation, and Robustness. Social Indicators Research, 141, 61–94. DOI: 10.1007/s11205-017-1832-9.

[5] Cinelli, M., Kadziński, M., Gonzalez, M., & Słowiński, R. (2020). How to support the application of multiple criteria decision analysis? Let us start with a comprehensive taxonomy. Omega, 96, 102261. DOI: 10.1016/j.omega.2020.102261.

[6] Li, B., Li, J., Tang, K., & Yao, X. (2015). Many-Objective Evolutionary Algorithms: A Survey. ACM Computing Surveys, 48(1), Article 13. DOI: 10.1145/2792984.

[7] Vamplew, P., Dazeley, R., Berry, A., Issabekov, R., & Dekker, E. (2011). Empirical evaluation methods for multiobjective reinforcement learning algorithms. Machine Learning, 84, 51–80. DOI: 10.1007/s10994-010-5232-5.

[8] Wilsdon, J., Allen, L., Belfiore, E., et al. (2015). The Metric Tide: Report of the Independent Review of the Role of Metrics in Research Assessment and Management. HEFCE.

[9] Adler, N. J., & Harzing, A.-W. (2009). When Knowledge Wins: Transcending the Sense and Nonsense of Academic Rankings. Academy of Management Learning & Education, 8(1), 72–95.

[10] Greco, S., Ehrgott, M., & Figueira, J. R. (Eds.). (2016). Multiple Criteria Decision Analysis: State of the Art Surveys, 2nd ed. Springer.


Appendix A. Comparison-Domain Schema

comparison_domain:
  task_family:
  subject_set:
  metrics:
  metric_direction:
  normalization:
  scope_contract:
  aggregator:
  observation_boundary:

subject_profile:
  subject_id:
  values:
  missing_metrics:

Appendix B. Comparison-Domain Failure Codes

CDF-1:
  TASK_MISMATCH

CDF-2:
  METRIC_APPLICABILITY_FAILURE

CDF-3:
  SCALE_FAILURE

CDF-4:
  SCOPE_MISMATCH

CDF-5:
  PARETO_CROSSING

CDF-6:
  AGGREGATION_UNDERDETERMINATION

CDF-7:
  OBJECTIVE_MISMATCH

CDF-8:
  OBSERVER_POLICY_MISMATCH

Appendix C. Ranking Output Types

DOMINATES
DOMINATED_BY
PARETO_INCOMPARABLE
ROBUSTLY_DOMINATES
TRADEOFF_SENSITIVE
DOMAIN_UNDEFINED
INSUFFICIENT_EVIDENCE

Appendix D. Minimal Ranking Disclosure

TASK:
SUBJECT_UNIT:
METRICS:
SCALES:
NORMALIZATION:
SCOPE:
WEIGHTS_OR_ORDER_RULE:
MISSING_DATA_POLICY:
UNCERTAINTY:
ROBUSTNESS_ANALYSIS:

Appendix E. Bridge to HEETF-07

HEETF-06:
  there is no natural universal leaderboard

therefore:
  subject need not optimize civilization's dominant comparison score

HEETF-07:
  chosen technological environment
  experiential technological fit
  right to chosen technological epoch
  virtual epistemic sovereignty