自然語言、圖像與創意如何被量?:高歧義成果的結構化品質空間
How Can Natural Language, Images, and Creative Outputs Be Measured? A Structured Quality Space for High-Ambiguity Artifacts
系列: 《智能的物理計量:從最小語意執行到成果品質與計算時空》英文系列: Physical Metrology of Intelligence: From Minimal Semantic Execution to Quality and Computational Spacetime 系列編號: EML-IPM篇次: Paper 08 / 10文件編號: EML-IPM-08作者: Neo.K with Aletheia(GPT-5.6 Sol)機構: EveMissLab/一言諾科技有限公司版本: v0.1日期: 2026-09-02文件性質: 公開純理論論文/跨模態品質計量方法論工程狀態: 無 MVP;本文建立高歧義成果的 typed quality-space 架構,供 Paper 07 的 IBQF/BRQM 進行後續測量
摘要
IPM Paper 06 已提出:
Objectifiable First, Human Residual Last . \boxed{
\text{Objectifiable First,\ Human Residual Last}.
} Objectifiable First, Human Residual Last .
Paper 07 則進一步提出:
many low-load binary / pairwise judgments → latent quality reconstruction . \boxed{
\text{many low-load binary / pairwise judgments}
\rightarrow
\text{latent quality reconstruction}.
} many low-load binary / pairwise judgments → latent quality reconstruction .
但此時仍有一個更根本的問題:
我們究竟應該問哪些問題? \boxed{
\textbf{
我們究竟應該問哪些問題?
}
} 我們究竟應該問哪些問題?
如果「品質維度」本身定義錯誤,那麼再精密的 Bradley–Terry、IRT、Bayesian posterior 或 adaptive questionnaire,也只會更精確地測量錯的東西。
因此本文主張,高歧義成果不能直接被投影成:
Q ∈ [ 0 , 1 ] . Q\in[0,1]. Q ∈ [ 0 , 1 ] .
它們首先應被表示成一個 Typed Structured Quality Space :
Q [ d , τ , c , a ] \boxed{
\mathcal Q[d,\tau,c,a]
} Q [ d , τ , c , a ]
其中:
d d d :domain / modality,領域與模態;
τ \tau τ :task,任務;
c c c :context,情境;
a a a :audience / evaluator population,受眾或評估群體。
品質因此不是:
Q ( Y ) \boxed{
Q(Y)
} Q ( Y )
而是:
Q ( Y ∣ d , τ , c , a ) . \boxed{
Q(
Y
\mid
d,\tau,c,a
).
} Q ( Y ∣ d , τ , c , a ) .
本文進一步把品質空間拆成三層:
Q = Q c o r e ⊕ Q d o m a i n ⊕ Q t a s k . \boxed{
\mathcal Q
=
\mathcal Q_{\mathrm{core}}
\oplus
\mathcal Q_{\mathrm{domain}}
\oplus
\mathcal Q_{\mathrm{task}}.
} Q = Q core ⊕ Q domain ⊕ Q task .
第一層 Cross-Domain Core Quality 包含可跨多模態共用的品質型別:
Q c o r e = ( F , C , K , R , U , V ) \boxed{
\mathbf Q_{\mathrm{core}}
=
(
F,
C,
K,
R,
U,
V
)
} Q core = ( F , C , K , R , U , V )
其中:
F F F :Fidelity / Alignment,是否忠實完成目標;
C C C :Coherence,一致性與內部連貫;
K K K :Completeness / Coverage,重要結構是否完整;
R R R :Robustness,對合理擾動是否穩定;
U U U :Usefulness / Functional Effectiveness,是否實際有用;
V V V :Verifiability / Traceability,能否被核查或追溯。
第二層為 Domain-Specific Quality 。例如:
自然語言
Q t e x t = ( F a c t u a l i t y , R e l e v a n c e , C o h e r e n c e , C o m p l e t e n e s s , C l a r i t y , P r a g m a t i c F i t , S t y l e F i t , E v i d e n c e A d e q u a c y ) . \boxed{
\mathbf Q_{\mathrm{text}}
=
(
Factuality,
Relevance,
Coherence,
Completeness,
Clarity,
PragmaticFit,
StyleFit,
EvidenceAdequacy
).
} Q text = ( F a c t u a l i t y , R e l e v an ce , C o h er e n ce , C o m pl e t e n ess , C l a r i t y , P r a g ma t i c F i t , S t y l e F i t , E v i d e n ce A d e q u a cy ) .
圖像
Q i m a g e = ( P r o m p t A l i g n m e n t , S t r u c t u r a l I n t e g r i t y , C o m p o s i t i o n , T e c h n i c a l Q u a l i t y , A e s t h e t i c F i t , S e m a n t i c L e g i b i l i t y , A r t i f a c t F r e e d o m , C o n t e x t F i t ) . \boxed{
\mathbf Q_{\mathrm{image}}
=
(
PromptAlignment,
StructuralIntegrity,
Composition,
TechnicalQuality,
AestheticFit,
SemanticLegibility,
ArtifactFreedom,
ContextFit
).
} Q image = ( P r o m pt A l i g nm e n t , S t r u c t u r a l I n t e g r i t y , C o m p os i t i o n , T ec hni c a l Q u a l i t y , A es t h e t i c F i t , S e man t i c L e g ibi l i t y , A r t i f a c tF r ee d o m , C o n t e x tF i t ) .
音樂
Q m u s i c = ( T e m p o r a l C o h e r e n c e , R h y t h m i c I n t e g r i t y , H a r m o n i c F i t , T i m b r e Q u a l i t y , E x p r e s s i v e n e s s , N o v e l t y , S t y l e F i t , F u n c t i o n a l F i t ) . \boxed{
\mathbf Q_{\mathrm{music}}
=
(
TemporalCoherence,
RhythmicIntegrity,
HarmonicFit,
TimbreQuality,
Expressiveness,
Novelty,
StyleFit,
FunctionalFit
).
} Q music = ( T e m p or a l C o h er e n ce , R h y t hmi c I n t e g r i t y , H a r m o ni c F i t , T imb r e Q u a l i t y , E x p r ess i v e n ess , N o v e l t y , S t y l e F i t , F u n c t i o na l F i t ) .
故事
Q s t o r y = ( P l o t C a u s a l i t y , C h a r a c t e r C o n s i s t e n c y , W o r l d C o n s i s t e n c y , P a c i n g , V o i c e , E m o t i o n a l I m p a c t , N o v e l t y , T h e m a t i c I n t e g r i t y ) . \boxed{
\mathbf Q_{\mathrm{story}}
=
(
PlotCausality,
CharacterConsistency,
WorldConsistency,
Pacing,
Voice,
EmotionalImpact,
Novelty,
ThematicIntegrity
).
} Q story = ( P l o tC a u s a l i t y , C ha r a c t er C o n s i s t e n cy , W or l d C o n s i s t e n cy , P a c in g , V o i ce , E m o t i o na l I m p a c t , N o v e l t y , T h e ma t i c I n t e g r i t y ) .
第三層是 Task-Specific Quality 。例如同一張圖:
作為廣告;
作為醫療示意圖;
作為遊戲概念圖;
作為純藝術作品;
其 success conditions 不同。
因此:
I m a g e Q u a l i t y ≠ U n i v e r s a l I m a g e B e a u t y . \boxed{
ImageQuality
\neq
UniversalImageBeauty.
} I ma g e Q u a l i t y = U ni v er s a l I ma g e B e a u t y .
本文特別提出 Quality Construct Graph(品質構念圖) :
G Q = ( V Q , E Q ) \boxed{
G_Q
=
(
V_Q,E_Q
)
} G Q = ( V Q , E Q )
其中 V Q V_Q V Q 是品質構念, E Q E_Q E Q 表示構念之間的依賴、包含、衝突或條件關係。
例如:
F a c t u a l i t y → T r u s t w o r t h i n e s s , Factuality
\rightarrow
Trustworthiness, F a c t u a l i t y → T r u s tw or t hin ess ,
C o m p o s i t i o n → V i s u a l H i e r a r c h y , Composition
\rightarrow
VisualHierarchy, C o m p os i t i o n → V i s u a l H i er a r c h y ,
N o v e l t y ⇏ U s e f u l n e s s . Novelty
\not\Rightarrow
Usefulness. N o v e l t y ⇒ U se f u l n ess .
這使品質維度不再只是扁平 checklist。
本文再定義從抽象構念到可問問題的 Measurement Itemization Pipeline :
Task → Quality Construct → Observable Indicator → Binary/Pairwise Item → Latent Estimate . \boxed{
\text{Task}
\rightarrow
\text{Quality Construct}
\rightarrow
\text{Observable Indicator}
\rightarrow
\text{Binary/Pairwise Item}
\rightarrow
\text{Latent Estimate}.
} Task → Quality Construct → Observable Indicator → Binary/Pairwise Item → Latent Estimate .
例如抽象構念:
C l a r i t y Clarity C l a r i t y
不能直接問:
清晰度幾分?
而應先拆成 observable indicators:
是否需要重讀?
是否能辨識主要結論?
是否知道下一步操作?
是否存在未定義關鍵詞?
再變成 Paper 07 的低負擔二元/成對問題。
因此:
C o n s t r u c t ≠ I t e m . \boxed{
Construct
\neq
Item.
} C o n s t r u c t = I t e m .
以及:
M e t r i c ≠ C o n s t r u c t . \boxed{
Metric
\neq
Construct.
} M e t r i c = C o n s t r u c t .
本文並提出 Construct Validity Gate :
G C = G ( C o v e r a g e , D i s c r i m i n a n t V a l i d i t y , C o n v e r g e n t E v i d e n c e , C o n t e x t S t a b i l i t y ) . \boxed{
G_C
=
G(
Coverage,
DiscriminantValidity,
ConvergentEvidence,
ContextStability
).
} G C = G ( C o v er a g e , D i scr iminan t V a l i d i t y , C o n v er g e n tE v i d e n ce , C o n t e x tS t abi l i t y ) .
若一個 measurement dimension 並沒有真正對應任務所需品質,則即使測量 reliability 很高:
R e l ↑ Rel\uparrow R e l ↑
仍可能:
V a l i d i t y ↓ . Validity\downarrow. V a l i d i t y ↓ .
因此:
R e l i a b i l i t y ≠ V a l i d i t y . \boxed{
Reliability
\neq
Validity.
} R e l iabi l i t y = V a l i d i t y .
高歧義成果還存在一個重要問題:品質空間本身可能是開放的。
若大量 residual errors、rater disagreement 或新型 failure 無法被現有維度解釋,系統不應硬塞進舊維度,而應允許:
Q t → Q t + 1 = Q t ∪ { new construct } . \boxed{
\mathcal Q_t
\rightarrow
\mathcal Q_{t+1}
=
\mathcal Q_t
\cup
\{\text{new construct}\}.
} Q t → Q t + 1 = Q t ∪ { new construct } .
本文稱此為 Open Quality Ontology 。
因此品質 schema 必須:
versioned;
extensible;
context-aware;
falsifiable。
本文對創造力特別提出:
N o v e l t y ≠ C r e a t i v i t y . \boxed{
Novelty
\neq
Creativity.
} N o v e l t y = C r e a t i v i t y .
創造性成果至少需要同時考慮:
N o v e l t y ⊕ A p p r o p r i a t e n e s s / U s e f u l n e s s . \boxed{
Novelty
\oplus
Appropriateness/Usefulness.
} N o v e l t y ⊕ A pp r o p r ia t e n ess / U se f u l n ess .
一個完全新奇但毫無用途、完全不符任務的答案,不應僅因「新」而獲得高創造力品質。
同理:
A e s t h e t i c P r e f e r e n c e ≠ T e c h n i c a l Q u a l i t y , \boxed{
AestheticPreference
\neq
TechnicalQuality,
} A es t h e t i c P r e f er e n ce = T ec hni c a l Q u a l i t y ,
F l u e n c y ≠ F a c t u a l i t y , \boxed{
Fluency
\neq
Factuality,
} F l u e n cy = F a c t u a l i t y ,
P r o m p t S i m i l a r i t y ≠ I m a g e Q u a l i t y . \boxed{
PromptSimilarity
\neq
ImageQuality.
} P r o m ptS imi l a r i t y = I ma g e Q u a l i t y .
因此任何單一 proxy metric 都只能是:
π j ( Q ) \boxed{
\pi_j(\mathcal Q)
} π j ( Q )
——品質空間的一個投影,而不是品質本體。
本文最終建立 High-Ambiguity Quality Object :
Q H A = ( Q s c h e m a , G Q , Q F , Q S , θ ^ H , Σ H , B Q , V e r s i o n Q ) . \boxed{
\mathfrak Q_{HA}
=
(
\mathcal Q_{\mathrm{schema}},
G_Q,
\mathbf Q_F,
\mathbf Q_S,
\widehat{\boldsymbol{\theta}}_H,
\Sigma_H,
B_Q,
Version_Q
).
} Q H A = ( Q schema , G Q , Q F , Q S , θ H , Σ H , B Q , V er s i o n Q ) .
其中:
Q s c h e m a \mathcal Q_{\mathrm{schema}} Q schema :typed quality schema;
G Q G_Q G Q :construct graph;
Q F \mathbf Q_F Q F :formal/objective layer;
Q S \mathbf Q_S Q S :structured semi-objective layer;
θ ^ H \widehat{\boldsymbol{\theta}}_H θ H :IBQF latent human residual;
Σ H \Sigma_H Σ H :人類 residual uncertainty;
B Q B_Q B Q :evaluation boundary;
V e r s i o n Q Version_Q V er s i o n Q :品質本體版本。
如此,IPM 的品質端可以從數學/程式一路延伸到:
natural language;
image;
audio/music;
story;
interface/design;
multimodal artifacts;
而不需要假裝所有模態共享同一個「8.6/10」。
本文終端命題為:
高歧義不是不可測; 它只是不能在尚未建立品質空間之前, 被過早壓縮成一個數字。 \boxed{
\textbf{
高歧義不是不可測;
它只是不能在尚未建立品質空間之前,
被過早壓縮成一個數字。
}
} 高歧義不是不可測; 它只是不能在尚未建立品質空間之前, 被過早壓縮成一個數字。
1. 「主觀」不等於「沒有結構」
一篇小說是否感人,
一張圖是否協調,
一段文字是否自然,
確實包含主觀成分。
2. 但主觀不等於任意
如果大量讀者都能辨識:
角色前後矛盾;
圖中手指結構異常;
段落缺少主題;
音樂節拍突然斷裂;
表示其中很多品質仍具有可結構化部分。
3. 所以:
S u b j e c t i v e ≠ U n s t r u c t u r e d . \boxed{
Subjective
\neq
Unstructured.
} S u bj ec t i v e = U n s t r u c t u r e d .
4. 高歧義應先分解,再測量
Paper 07 解決的是:
怎麼問人?
Paper 08 解決:
問什麼?
5. 最危險的錯誤不是 measurement noise
而是:
construct misspecification . \boxed{
\text{construct misspecification}.
} construct misspecification .
也就是你測得很準,
但測的根本不是你想知道的品質。
6. Typed Quality Space
本文定義:
Q [ d , τ , c , a ] . \boxed{
\mathcal Q[d,\tau,c,a].
} Q [ d , τ , c , a ] .
品質至少受:
domain;
task;
context;
audience;
共同條件化。
7. 為什麼需要 domain?
因為文字的 grammar,
不是圖片的 composition。
8. 為什麼需要 task?
一篇摘要要求:
一篇小說則可以要求:
9. 為什麼需要 context?
同一 UI:
desktop;
mobile;
accessibility mode;
會產生不同使用品質。
10. 為什麼需要 audience?
專家眼中的:
C l a r i t y Clarity C l a r i t y
與 novice 眼中的:
C l a r i t y Clarity C l a r i t y
可以不同。
11. 因此:
Q u a l i t y = R e l a t i o n , not intrinsic decoration . \boxed{
Quality
=
Relation,
\quad
\text{not intrinsic decoration}.
} Q u a l i t y = R e l a t i o n , not intrinsic decoration .
12. 三層品質空間
Q = Q c o r e ⊕ Q d o m a i n ⊕ Q t a s k . \boxed{
\mathcal Q
=
\mathcal Q_{\mathrm{core}}
\oplus
\mathcal Q_{\mathrm{domain}}
\oplus
\mathcal Q_{\mathrm{task}}.
} Q = Q core ⊕ Q domain ⊕ Q task .
13. Cross-Domain Core
本文暫定六個核心型別:
Q c o r e = ( F , C , K , R , U , V ) . \boxed{
\mathbf Q_{\mathrm{core}}
=
(
F,C,K,R,U,V
).
} Q core = ( F , C , K , R , U , V ) .
14. Fidelity / Alignment
F F F
問:
有沒有做使用者真正要求的事情?
15. Coherence
C C C
問:
輸出的部分彼此能否共存?
16. Completeness
K K K
問:
必要結構是否缺失?
17. Robustness
R R R
問:
合理的小擾動會不會讓成果崩潰?
18. Usefulness / Functional Effectiveness
U U U
問:
這個成果是否真的完成使用目的?
19. Verifiability / Traceability
V V V
問:
若需要核查,能否知道它為何成立/從何而來?
20. Core 不代表 universal weights
不同 task:
w F , w C , w K , w R , w U , w V w_F,w_C,w_K,w_R,w_U,w_V w F , w C , w K , w R , w U , w V
可以不同。
21. 甚至某些維度可以不適用
純藝術作品不一定需要:
E v i d e n c e A d e q u a c y . EvidenceAdequacy. E v i d e n ce A d e q u a cy .
22. 因此:
C o r e T y p e ≠ M a n d a t o r y E q u a l W e i g h t . \boxed{
CoreType
\neq
MandatoryEqualWeight.
} C or e T y p e = M an d a t or y E q u a l W e i g h t .
23. 自然語言品質空間
Q t e x t = ( F a c t u a l i t y , R e l e v a n c e , C o h e r e n c e , C o m p l e t e n e s s , C l a r i t y , P r a g m a t i c F i t , S t y l e F i t , E v i d e n c e A d e q u a c y ) . \boxed{
\mathbf Q_{\mathrm{text}}
=
(
Factuality,
Relevance,
Coherence,
Completeness,
Clarity,
PragmaticFit,
StyleFit,
EvidenceAdequacy
).
} Q text = ( F a c t u a l i t y , R e l e v an ce , C o h er e n ce , C o m pl e t e n ess , C l a r i t y , P r a g ma t i c F i t , S t y l e F i t , E v i d e n ce A d e q u a cy ) .
24. Factuality 與 Fluency 必須分開
一段話可以:
F l u e n c y ↑ Fluency\uparrow F l u e n cy ↑
但:
F a c t u a l i t y ↓ . Factuality\downarrow. F a c t u a l i t y ↓ .
25. 所以:
F l u e n c y ≠ T r u t h . \boxed{
Fluency
\neq
Truth.
} F l u e n cy = T r u t h .
26. Relevance 與 Completeness 也不同
回答可以每句都相關,
但漏掉一半要求。
27. 因此:
R e l e v a n t ≠ C o m p l e t e . \boxed{
Relevant
\neq
Complete.
} R e l e v an t = C o m pl e t e .
28. Coherence 與 Correctness 也不同
錯誤理論可以極度自洽。
所以:
C o h e r e n c e ≠ C o r r e c t n e s s . \boxed{
Coherence
\neq
Correctness.
} C o h er e n ce = C or r ec t n ess .
29. Pragmatic Fit
同一句話對:
beginner;
expert;
emergency;
academic setting;
適切度不同。
30. 圖像品質空間
Q i m a g e = ( P r o m p t A l i g n m e n t , S t r u c t u r a l I n t e g r i t y , C o m p o s i t i o n , T e c h n i c a l Q u a l i t y , A e s t h e t i c F i t , S e m a n t i c L e g i b i l i t y , A r t i f a c t F r e e d o m , C o n t e x t F i t ) . \boxed{
\mathbf Q_{\mathrm{image}}
=
(
PromptAlignment,
StructuralIntegrity,
Composition,
TechnicalQuality,
AestheticFit,
SemanticLegibility,
ArtifactFreedom,
ContextFit
).
} Q image = ( P r o m pt A l i g nm e n t , S t r u c t u r a l I n t e g r i t y , C o m p os i t i o n , T ec hni c a l Q u a l i t y , A es t h e t i c F i t , S e man t i c L e g ibi l i t y , A r t i f a c tF r ee d o m , C o n t e x tF i t ) .
31. Prompt Alignment
圖像是否真的符合要求:
32. Structural Integrity
例如:
anatomy;
perspective;
object topology;
geometry。
33. Technical Quality
例如:
noise;
blur;
compression;
resolution;
rendering artifacts。
34. Aesthetic Fit
則是:
visual balance;
style harmony;
appeal;
intentionality。
35. Technical Quality 與 Aesthetic Quality 不同
一張極度清晰的圖,
仍可以:
A e s t h e t i c F i t ↓ . AestheticFit\downarrow. A es t h e t i c F i t ↓ .
36. 反過來
低解析、粗糙的藝術作品,
仍可能:
A e s t h e t i c I m p a c t ↑ . AestheticImpact\uparrow. A es t h e t i c I m p a c t ↑ .
37. 所以:
T e c h n i c a l Q u a l i t y ≠ A e s t h e t i c Q u a l i t y . \boxed{
TechnicalQuality
\neq
AestheticQuality.
} T ec hni c a l Q u a l i t y = A es t h e t i c Q u a l i t y .
38. NIMA 等 image-aesthetic work 的啟示
影像評估研究會刻意建模人類 rating distribution,而不只預測單一平均分。
這提醒我們:
A e s t h e t i c D i s t r i b u t i o n \boxed{
AestheticDistribution
} A es t h e t i cD i s t r ib u t i o n
本身就可能是 measurement object。
39. 圖像相似度也不是品質
若 generated image 與 prompt embedding 高相似:
S i m i l a r i t y ↑ Similarity\uparrow S imi l a r i t y ↑
不表示:
anatomy 正確;
composition 好;
artifacts 少。
40. 所以:
P r o m p t S i m i l a r i t y ≠ I m a g e Q u a l i t y . \boxed{
PromptSimilarity
\neq
ImageQuality.
} P r o m ptS imi l a r i t y = I ma g e Q u a l i t y .
41. 音樂品質空間
Q m u s i c = ( T e m p o r a l C o h e r e n c e , R h y t h m i c I n t e g r i t y , H a r m o n i c F i t , T i m b r e Q u a l i t y , E x p r e s s i v e n e s s , N o v e l t y , S t y l e F i t , F u n c t i o n a l F i t ) . \boxed{
\mathbf Q_{\mathrm{music}}
=
(
TemporalCoherence,
RhythmicIntegrity,
HarmonicFit,
TimbreQuality,
Expressiveness,
Novelty,
StyleFit,
FunctionalFit
).
} Q music = ( T e m p or a l C o h er e n ce , R h y t hmi c I n t e g r i t y , H a r m o ni c F i t , T imb r e Q u a l i t y , E x p r ess i v e n ess , N o v e l t y , S t y l e F i t , F u n c t i o na l F i t ) .
42. 音樂不是只測 waveform quality
技術無雜訊:
A u d i o C l e a n = 1 AudioClean=1 A u d i o C l e an = 1
不代表:
M u s i c Q u a l i t y = 1. MusicQuality=1. M u s i c Q u a l i t y = 1.
43. Temporal Coherence
旋律/節奏是否形成可理解的時間結構。
44. Functional Fit
背景音樂、舞曲、電影配樂、實驗音樂的 success condition 不同。
45. 故事品質空間
Q s t o r y = ( P l o t C a u s a l i t y , C h a r a c t e r C o n s i s t e n c y , W o r l d C o n s i s t e n c y , P a c i n g , V o i c e , E m o t i o n a l I m p a c t , N o v e l t y , T h e m a t i c I n t e g r i t y ) . \boxed{
\mathbf Q_{\mathrm{story}}
=
(
PlotCausality,
CharacterConsistency,
WorldConsistency,
Pacing,
Voice,
EmotionalImpact,
Novelty,
ThematicIntegrity
).
} Q story = ( P l o tC a u s a l i t y , C ha r a c t er C o n s i s t e n cy , W or l d C o n s i s t e n cy , P a c in g , V o i ce , E m o t i o na l I m p a c t , N o v e l t y , T h e ma t i c I n t e g r i t y ) .
46. Plot Causality
事件不是只「接在一起」,
而應存在可理解關係。
47. Character Consistency
角色變化可以很大,
但需要:
motivated transition . \boxed{
\text{motivated transition}.
} motivated transition .
48. 所以 consistency 不等於不准變
真正測的是:
C h a n g e W i t h o u t S u f f i c i e n t C a u s e ? \boxed{
ChangeWithoutSufficientCause?
} C han g e W i t h o u tS u f f i c i e n tC a u se ?
49. UI / Design 品質空間
Q d e s i g n = ( T a s k S u c c e s s , A f f o r d a n c e , H i e r a r c h y , C o n s i s t e n c y , A c c e s s i b i l i t y , E r r o r T o l e r a n c e , V i s u a l C o h e r e n c e , A u d i e n c e F i t ) . \boxed{
\mathbf Q_{\mathrm{design}}
=
(
TaskSuccess,
Affordance,
Hierarchy,
Consistency,
Accessibility,
ErrorTolerance,
VisualCoherence,
AudienceFit
).
} Q design = ( T a s k S u ccess , A f f or d an ce , H i er a r c h y , C o n s i s t e n cy , A ccess ibi l i t y , E r r or T o l er an ce , V i s u a l C o h er e n ce , A u d i e n ce F i t ) .
50. 美觀不等於可用
A e s t h e t i c A p p e a l ≠ U s a b i l i t y . \boxed{
AestheticAppeal
\neq
Usability.
} A es t h e t i c A pp e a l = U s abi l i t y .
51. 可用也不等於適合所有人
所以:
A c c e s s i b i l i t y Accessibility A ccess ibi l i t y
必須單獨存在。
52. 創造力品質
創造力研究長期有一個常見核心:
N o v e l t y + U s e f u l n e s s / A p p r o p r i a t e n e s s . \boxed{
Novelty
+
Usefulness/Appropriateness.
} N o v e l t y + U se f u l n ess / A pp r o p r ia t e n ess .
53. Novelty alone 不夠
亂碼非常新。
但:
U s e f u l n e s s = 0. Usefulness=0. U se f u l n ess = 0.
54. 因此:
N o v e l t y ≠ C r e a t i v i t y . \boxed{
Novelty
\neq
Creativity.
} N o v e l t y = C r e a t i v i t y .
55. 反過來
非常有用但完全常規,
也可能:
N o v e l t y ≈ 0. Novelty\approx0. N o v e l t y ≈ 0.
56. 所以 creativity 更像一個多軸區域
C c r e a t i v e = ( N o v e l t y , A p p r o p r i a t e n e s s , V a l u e , S u r p r i s e , C o h e r e n c e ) . \boxed{
\mathcal C_{\mathrm{creative}}
=
(
Novelty,
Appropriateness,
Value,
Surprise,
Coherence
).
} C creative = ( N o v e l t y , A pp r o p r ia t e n ess , V a l u e , S u r p r i se , C o h er e n ce ) .
57. 不要過早寫成
C r e a t i v i t y = N o v e l t y × U s e f u l n e s s . Creativity
=
Novelty\times Usefulness. C r e a t i v i t y = N o v e l t y × U se f u l n ess .
除非研究場景明確需要這種 projection。
58. 因為 multiplicative scalar 會偷偷加入 trade-off 規則
59. Quality Construct Graph
扁平 vector 還不夠。
有些品質是依賴的。
60. 定義:
G Q = ( V Q , E Q ) . \boxed{
G_Q=(V_Q,E_Q).
} G Q = ( V Q , E Q ) .
61. 節點
是 construct:
factuality;
clarity;
composition;
novelty。
62. 邊
可以表示:
prerequisite;
causal support;
overlap;
conflict;
conditional dependency。
63. 例如
E v i d e n c e A d e q u a c y → F a c t u a l C o n f i d e n c e . EvidenceAdequacy
\rightarrow
FactualConfidence. E v i d e n ce A d e q u a cy → F a c t u a l C o n f i d e n ce .
64. 又例如
N o v e l t y ⇏ U s e f u l n e s s . Novelty
\not\Rightarrow
Usefulness. N o v e l t y ⇒ U se f u l n ess .
65. Visual hierarchy 可以影響:
R e a d a b i l i t y , T a s k N a v i g a t i o n . Readability,
TaskNavigation. R e a d abi l i t y , T a s k N a v i g a t i o n .
66. 這使品質 schema 變成結構
而不是 checklist。
67. Construct 與 Indicator 不同
例如:
C l a r i t y Clarity C l a r i t y
是 latent construct。
68. 不能直接把它當 item
應先尋找 observable indicator。
69. Clarity indicators 可能是
是否需要重讀?
能否指出主結論?
能否預測下一步?
是否存在未定義術語?
70. 所以:
C o n s t r u c t ≠ O b s e r v a b l e I n d i c a t o r . \boxed{
Construct
\neq
ObservableIndicator.
} C o n s t r u c t = O b ser v ab l e I n d i c a t or .
71. Indicator 再變成 Paper 07 item
例如:
你第一次讀完是否能指出作者的核心結論?
Yes/No . \text{Yes/No}. Yes/No .
72. 或 pairwise
A 與 B 哪個更容易讓你找出核心結論?
A / B . A/B. A / B .
73. Measurement Itemization Pipeline
T a s k → C o n s t r u c t → I n d i c a t o r → I t e m → O b s e r v a t i o n → L a t e n t E s t i m a t e . \boxed{
Task
\rightarrow
Construct
\rightarrow
Indicator
\rightarrow
Item
\rightarrow
Observation
\rightarrow
LatentEstimate.
} T a s k → C o n s t r u c t → I n d i c a t or → I t e m → O b ser v a t i o n → L a t e n tE s t ima t e .
74. 這條鏈每一步都可能出錯
所以每一步都要能被檢查。
75. Metric 不等於 Construct
例如 BLEU、embedding similarity、CLIPScore、aesthetic model score 等,
都只是:
M e t r i c j = π j ( Q ) . \boxed{
Metric_j
=
\pi_j(
\mathcal Q
).
} M e t r i c j = π j ( Q ) .
76. 一個 metric 只能看 quality space 的某個投影
77. Metric Correlation 不等於 Construct Validity
如果 metric 與 human score 相關,
仍不代表:
它已經捕捉完整品質。
78. 因為 human score 本身也可能是錯的 projection
79. 所以需要 Construct Validity Gate
G C = G ( C o v e r a g e , D i s c r i m i n a n t V a l i d i t y , C o n v e r g e n t E v i d e n c e , C o n t e x t S t a b i l i t y ) . \boxed{
G_C
=
G(
Coverage,
DiscriminantValidity,
ConvergentEvidence,
ContextStability
).
} G C = G ( C o v er a g e , D i scr iminan t V a l i d i t y , C o n v er g e n tE v i d e n ce , C o n t e x tS t abi l i t y ) .
80. Coverage
這套 dimensions 是否漏掉重要品質?
81. Discriminant Validity
兩個宣稱不同的 dimensions 是否其實測同一件事?
82. Convergent Evidence
不同方法是否對同一 construct 有合理收斂?
83. Context Stability
換情境後,construct 是否仍保留原意?
84. Reliability 仍不等於 Validity
一個錯誤量尺可以每次都非常穩定。
所以:
R e l i a b i l i t y ≠ V a l i d i t y . \boxed{
Reliability
\neq
Validity.
} R e l iabi l i t y = V a l i d i t y .
85. Open Quality Ontology
品質空間不應被假定永遠固定。
86. 如果大量 failure 無法解釋
例如 AI 圖像出現一種新的 artifact,
現有 dimensions 沒有位置放。
87. 不應硬塞
而應允許:
Q t + 1 = Q t ∪ { q n e w } . \boxed{
\mathcal Q_{t+1}
=
\mathcal Q_t
\cup
\{q_{\mathrm{new}}\}.
} Q t + 1 = Q t ∪ { q new } .
88. Trigger
新 construct 的候選來源可以是:
residual error;
clustered rater comments;
unexplained disagreement;
adversarial failure;
new task class。
89. 這叫 Quality Ontology Expansion
90. 但不能每次看到一個怪例子就加一維
需要:
reproducibility;
discriminant value;
explanatory gain。
91. Quality Schema 必須 versioned
Q v 1 ≠ Q v 2 . \boxed{
\mathcal Q^{v_1}
\neq
\mathcal Q^{v_2}.
} Q v 1 = Q v 2 .
92. 否則跨時間 benchmark 會偷偷改量尺
93. Quality Schema Drift
如果 2026 和 2028 的「圖像品質」用不同 dimensions,
分數不可直接縱向比較。
94. 所以:
B e n c h m a r k V e r s i o n ⇒ Q u a l i t y O n t o l o g y V e r s i o n . \boxed{
BenchmarkVersion
\Rightarrow
QualityOntologyVersion.
} B e n c hma r k V er s i o n ⇒ Q u a l i t y O n t o l o g y V er s i o n .
95. Measurement Invariance
跨群體比較前,
要問:
這些 items 對不同群體是否仍在測同一 construct?
96. 例如 novice 的「清晰」
可能受背景知識影響。
97. 所以:
S a m e I t e m ≠ S a m e M e a s u r e m e n t F u n c t i o n \boxed{
SameItem
\neq
SameMeasurementFunction
} S am e I t e m = S am e M e a s u r e m e n tF u n c t i o n
across populations。
98. 這延續 Paper 07 的 context/rater modeling
99. Multi-Modal Quality
一個 artifact 可能同時有:
text;
image;
audio;
interaction。
100. 不能只把各 modality score 平均
Q m u l t i ≠ Q T + Q I + Q A 3 . \boxed{
Q_{\mathrm{multi}}
\neq
\frac{Q_T+Q_I+Q_A}{3}.
} Q multi = 3 Q T + Q I + Q A .
101. 因為跨模態存在 coupling
例如:
文字說:
按紅色按鈕。
圖片卻只有藍色按鈕。
102. 各自單獨可能合法
但:
C r o s s M o d a l C o n s i s t e n c y = 0. \boxed{
CrossModalConsistency=0.
} C r oss M o d a l C o n s i s t e n cy = 0.
103. 所以 multimodal schema 需要 coupling dimensions
Q c o u p l e = ( C r o s s M o d a l C o n s i s t e n c y , R e f e r e n c e A l i g n m e n t , T e m p o r a l S y n c , R e d u n d a n c y , C o m p l e m e n t a r i t y ) . \boxed{
\mathbf Q_{\mathrm{couple}}
=
(
CrossModalConsistency,
ReferenceAlignment,
TemporalSync,
Redundancy,
Complementarity
).
} Q couple = ( C r oss M o d a l C o n s i s t e n cy , R e f er e n ce A l i g nm e n t , T e m p or a l S y n c , R e d u n d an cy , C o m pl e m e n t a r i t y ) .
104. Redundancy 不一定是壞事
Accessibility 可能故意讓 text + icon 重複。
105. 因此要看 task function
106. Complementarity
不同模態是否各自提供新的有用資訊。
107. Quality Fiber View
本文可更抽象表示:
對每個 task/context point:
x = ( d , τ , c , a ) , x=(d,\tau,c,a), x = ( d , τ , c , a ) ,
都有一個 quality fiber:
Q x . \boxed{
\mathcal Q_x.
} Q x .
108. 整體不是單一固定向量空間
而是:
Q = ⋃ x Q x . \boxed{
\mathcal Q
=
\bigcup_x
\mathcal Q_x.
} Q = x ⋃ Q x .
109. 這比 universal score 更符合高歧義成果
110. Cross-Task Projection
如果要比較不同 task,
需建立:
Π x → y : Q x → Q y . \boxed{
\Pi_{x\rightarrow y}:
\mathcal Q_x
\rightarrow
\mathcal Q_y.
} Π x → y : Q x → Q y .
111. 但不是所有維度都可投影
例如:
P l o t C a u s a l i t y PlotCausality P l o tC a u s a l i t y
沒有直接對應到圖片 technical noise。
112. 所以跨模態比較應只比較 shared constructs
113. 這對 IPM 很重要
若比較:
文字模型和圖像模型誰更聰明?
不能直接拿兩種 domain score 相除。
114. 必須找到共同 task-level achievement
例如:
是否成功傳達指定空間關係?
115. 然後比較同一 task construct 下的:
P c o m p u t e → Q . \mathfrak P_{\mathrm{compute}}
\rightarrow
Q. P compute → Q .
116. 否則:
C r o s s D o m a i n S c o r e C o m p a r i s o n \boxed{
CrossDomainScoreComparison
} C r ossD o main S cor e C o m p a r i so n
沒有共同 measurement basis。
117. 這是跨基質與跨模態智能計量的一個大限制
118. 高歧義成果的三層測量流程
Stage 1 — Formal / Structural Extraction
先自動抽:
hard constraints;
detectable errors;
explicit requirements;
provenance。
119. Stage 2 — Quality Schema Mapping
選:
Q [ d , τ , c , a ] . \mathcal Q[d,\tau,c,a]. Q [ d , τ , c , a ] .
120. Stage 3 — Human Residual Itemization
把尚未解決的 constructs 轉成:
binary;
pairwise;
adaptive items。
121. Stage 4 — Latent Reconstruction
{ b i } → θ ^ . \{b_i\}
\rightarrow
\widehat{\boldsymbol{\theta}}. { b i } → θ .
122. Stage 5 — Residual Analysis
看:
unexplained disagreement;
residual errors;
new failure clusters。
123. Stage 6 — Ontology Revision
必要時:
Q v n → Q v n + 1 . \mathcal Q^{v_n}
\rightarrow
\mathcal Q^{v_{n+1}}. Q v n → Q v n + 1 .
124. 這使品質測量本身變成可學習系統
但:
M e a s u r e m e n t S y s t e m L e a r n i n g ≠ M o v i n g G o a l p o s t s W i t h o u t V e r s i o n i n g . \boxed{
MeasurementSystemLearning
\neq
MovingGoalpostsWithoutVersioning.
} M e a s u r e m e n tS y s t e m L e a r nin g = M o v in g G o a l p os t s W i t h o u t V er s i o nin g .
125. 必須保留 canonical historical versions
才能做 longitudinal comparison。
126. High-Ambiguity Quality Object
本文正式定義:
Q H A = ( Q s c h e m a , G Q , Q F , Q S , θ ^ H , Σ H , D R , B Q , V e r s i o n Q ) . \boxed{
\mathfrak Q_{HA}
=
(
\mathcal Q_{\mathrm{schema}},
G_Q,
\mathbf Q_F,
\mathbf Q_S,
\widehat{\boldsymbol{\theta}}_H,
\Sigma_H,
\mathcal D_R,
B_Q,
Version_Q
).
} Q H A = ( Q schema , G Q , Q F , Q S , θ H , Σ H , D R , B Q , V er s i o n Q ) .
127. Q s c h e m a \mathcal Q_{\mathrm{schema}} Q schema
目前採用哪些 constructs。
128. G Q G_Q G Q
construct 之間的關係。
129. Q F \mathbf Q_F Q F
formal/objective evidence。
130. Q S \mathbf Q_S Q S
structured semi-objective evidence。
131. θ ^ H \widehat{\boldsymbol{\theta}}_H θ H
Paper 07 產生的人類 latent residual。
132. Σ H \Sigma_H Σ H
不確定性。
133. D R \mathcal D_R D R
群體 disagreement structure。
134. B Q B_Q B Q
domain/task/context/audience boundary。
135. V e r s i o n Q Version_Q V er s i o n Q
quality ontology version。
136. Scalar 仍然只是 Projection
如果某 leaderboard 必須一個數字:
Q ∗ = Π ( Q H A ∣ P o l i c y ) . \boxed{
Q^*
=
\Pi(
\mathfrak Q_{HA}
\mid
Policy
).
} Q ∗ = Π ( Q H A ∣ P o l i cy ) .
137. Policy 要公開
例如:
correctness hard gate;
task alignment 40%;
usability 30%;
aesthetics 30%。
138. 但原始 quality object 必須保留
否則 leaderboard 抹掉 trade-off。
139. Pareto Quality Frontier
甚至品質本身也可能需要 Pareto。
一張圖:
A:
technical quality 高;
novelty 普通。
B:
technical defects 多;
creative impact 極高。
140. 若 task 沒指定 preference,
不應硬說誰「品質總體更高」。
141. 所以:
Q u a l i t y D o m i n a n c e \boxed{
QualityDominance
} Q u a l i t y D o minan ce
只有在所有 relevant dimensions 都不差時才自然成立。
142. 十八個 Canonical Invariants
Invariant 1
S u b j e c t i v e ≠ U n s t r u c t u r e d . \boxed{
Subjective
\neq
Unstructured.
} S u bj ec t i v e = U n s t r u c t u r e d .
Invariant 2
H i g h A m b i g u i t y ≠ U n m e a s u r a b l e . \boxed{
HighAmbiguity
\neq
Unmeasurable.
} H i g h A mbi g u i t y = U nm e a s u r ab l e .
Invariant 3
Q u a l i t y ≠ U n i v e r s a l S c a l a r . \boxed{
Quality
\neq
UniversalScalar.
} Q u a l i t y = U ni v er s a l S c a l a r .
Invariant 4
C o n s t r u c t ≠ I n d i c a t o r . \boxed{
Construct
\neq
Indicator.
} C o n s t r u c t = I n d i c a t or .
Invariant 5
I n d i c a t o r ≠ I t e m . \boxed{
Indicator
\neq
Item.
} I n d i c a t or = I t e m .
Invariant 6
M e t r i c ≠ C o n s t r u c t . \boxed{
Metric
\neq
Construct.
} M e t r i c = C o n s t r u c t .
Invariant 7
R e l i a b i l i t y ≠ V a l i d i t y . \boxed{
Reliability
\neq
Validity.
} R e l iabi l i t y = V a l i d i t y .
Invariant 8
F l u e n c y ≠ F a c t u a l i t y . \boxed{
Fluency
\neq
Factuality.
} F l u e n cy = F a c t u a l i t y .
Invariant 9
C o h e r e n c e ≠ C o r r e c t n e s s . \boxed{
Coherence
\neq
Correctness.
} C o h er e n ce = C or r ec t n ess .
Invariant 10
T e c h n i c a l Q u a l i t y ≠ A e s t h e t i c Q u a l i t y . \boxed{
TechnicalQuality
\neq
AestheticQuality.
} T ec hni c a l Q u a l i t y = A es t h e t i c Q u a l i t y .
Invariant 11
P r o m p t S i m i l a r i t y ≠ I m a g e Q u a l i t y . \boxed{
PromptSimilarity
\neq
ImageQuality.
} P r o m ptS imi l a r i t y = I ma g e Q u a l i t y .
Invariant 12
N o v e l t y ≠ C r e a t i v i t y . \boxed{
Novelty
\neq
Creativity.
} N o v e l t y = C r e a t i v i t y .
Invariant 13
A e s t h e t i c A p p e a l ≠ U s a b i l i t y . \boxed{
AestheticAppeal
\neq
Usability.
} A es t h e t i c A pp e a l = U s abi l i t y .
Invariant 14
S a m e I t e m ≠ S a m e M e a s u r e m e n t F u n c t i o n \boxed{
SameItem
\neq
SameMeasurementFunction
} S am e I t e m = S am e M e a s u r e m e n tF u n c t i o n
across populations without invariance evidence.
Invariant 15
M u l t i m o d a l Q u a l i t y ≠ M e a n ( M o d a l i t y S c o r e s ) . \boxed{
MultimodalQuality
\neq
Mean(ModalityScores).
} M u l t im o d a l Q u a l i t y = M e an ( M o d a l i t y S cor es ) .
Invariant 16
C r o s s D o m a i n C o m p a r i s o n ⇒ S h a r e d C o n s t r u c t B a s i s . \boxed{
CrossDomainComparison
\Rightarrow
SharedConstructBasis.
} C r ossD o main C o m p a r i so n ⇒ S ha r e d C o n s t r u c tB a s i s .
Invariant 17
O n t o l o g y R e v i s i o n ⇒ V e r s i o n i n g . \boxed{
OntologyRevision
\Rightarrow
Versioning.
} O n t o l o g y R e v i s i o n ⇒ V er s i o nin g .
Invariant 18
P r e c i s e M e a s u r e m e n t ≠ C o r r e c t C o n s t r u c t S e l e c t i o n . \boxed{
PreciseMeasurement
\neq
CorrectConstructSelection.
} P r ec i se M e a s u r e m e n t = C or r ec tC o n s t r u c tS e l ec t i o n .
143. 對 IPM 品質端的完成
Paper 06 建立:
Q F ⊕ Q S . \mathfrak Q_F
\oplus
\mathfrak Q_S. Q F ⊕ Q S .
144. Paper 07 建立:
Q H I B Q F . \mathfrak Q_H^{IBQF}. Q H I B QF .
145. Paper 08 現在補上:
Q s c h e m a \boxed{
\mathcal Q_{\mathrm{schema}}
} Q schema
——也就是究竟有哪些品質需要被量。
146. 因此完整品質端
Q I P M = ( Q s c h e m a , Q F , Q S , Q H I B Q F , V e r s i o n Q ) . \boxed{
\mathfrak Q_{\mathrm{IPM}}
=
(
\mathcal Q_{\mathrm{schema}},
\mathfrak Q_F,
\mathfrak Q_S,
\mathfrak Q_H^{IBQF},
Version_Q
).
} Q IPM = ( Q schema , Q F , Q S , Q H I B QF , V er s i o n Q ) .
147. 這讓品質側真正閉合
不是:
先決定一個總分,再找方法算。
而是:
Task → Quality Ontology → Evidence → Latent Measurement → Optional Projection . \boxed{
\text{Task}
\rightarrow
\text{Quality Ontology}
\rightarrow
\text{Evidence}
\rightarrow
\text{Latent Measurement}
\rightarrow
\text{Optional Projection}.
} Task → Quality Ontology → Evidence → Latent Measurement → Optional Projection .
148. 結論:高歧義不是不能量,而是不能太早壓縮
自然語言、圖像、音樂、故事、設計與創意,
之所以難量,
不是因為它們完全沒有結構。
而是因為:
它們的品質本體本身就是多維、情境化、任務化, 而且部分維度仍需人類感知才能完成測量。 \boxed{
\textbf{
它們的品質本體本身就是多維、情境化、任務化,
而且部分維度仍需人類感知才能完成測量。
}
} 它們的品質本體本身就是多維、情境化、任務化, 而且部分維度仍需人類感知才能完成測量。
如果我們太早問:
這張圖 0–10 幾分?
其實做了兩次壓縮。
第一次:
Q [ d , τ , c , a ] → observer internal impression . \boxed{
\mathcal Q[d,\tau,c,a]
\rightarrow
\text{observer internal impression}.
} Q [ d , τ , c , a ] → observer internal impression .
第二次:
internal impression → 0 ∼ 10. \boxed{
\text{internal impression}
\rightarrow
0\sim10.
} internal impression → 0 ∼ 10.
大量資訊在過程中消失。
Paper 07 已經移除了第二次不必要壓縮:
local binary/pairwise → latent reconstruction . \boxed{
\text{local binary/pairwise}
\rightarrow
\text{latent reconstruction}.
} local binary/pairwise → latent reconstruction .
Paper 08 則移除第一次概念混亂:
我們先明確建立:
Q s c h e m a \boxed{
\mathcal Q_{\mathrm{schema}}
} Q schema
再問:
哪些部分可形式驗證?
哪些可結構核查?
哪些真的只能靠人類感知?
人類感知又應拆成哪些局部 observables?
因此:
高歧義不是不可測; 它只是不能在尚未建立品質空間之前, 被過早壓縮成一個數字。 \boxed{
\textbf{
高歧義不是不可測;
它只是不能在尚未建立品質空間之前,
被過早壓縮成一個數字。
}
} 高歧義不是不可測; 它只是不能在尚未建立品質空間之前, 被過早壓縮成一個數字。
到這一步,IPM 的「分子側」基本完成。
我們現在已有:
Q I P M \boxed{
\mathfrak Q_{\mathrm{IPM}}
} Q IPM
作為成果品質,
以及:
P c o m p u t e \boxed{
\mathfrak P_{\mathrm{compute}}
} P compute
作為物理成本。
剩下最後兩篇要重新回到整個系列最初的問題。
Paper 09 將問:
如果把 tools、retry、verifier、best-of-N、 外部 LOOP 與大量隱藏鷹架一層層拿掉, 一個模型自己到底還剩多少智能? \boxed{
\textbf{
如果把 tools、retry、verifier、best-of-N、
外部 LOOP 與大量隱藏鷹架一層層拿掉,
一個模型自己到底還剩多少智能?
}
} 如果把 tools 、 retry 、 verifier 、 best-of-N 、 外部 LOOP 與大量隱藏鷹架一層層拿掉, 一個模型自己到底還剩多少智能?
也就是:
《拿掉 LOOP 還剩多少智能?:單次智能、鷹架依賴與隱藏計算成本》 。
文獻基礎
[1] Gatt, A., & Krahmer, E. (2018). Survey of the State of the Art in Natural Language Generation: Core Tasks, Applications and Evaluation. Journal of Artificial Intelligence Research , 61, 65–170. DOI: 10.1613/jair.5477.
[2] Liu, Y., Iter, D., Xu, Y., Wang, S., Xu, R., & Zhu, C. (2023). G-Eval: NLG Evaluation using GPT-4 with Better Human Alignment. EMNLP 2023 , 2511–2522.
[3] Hessel, J., Holtzman, A., Forbes, M., Le Bras, R., & Choi, Y. (2021). CLIPScore: A Reference-free Evaluation Metric for Image Captioning. EMNLP 2021 .
[4] Talebi, H., & Milanfar, P. (2018). NIMA: Neural Image Assessment. IEEE Transactions on Image Processing , 27(8), 3998–4011. DOI: 10.1109/TIP.2018.2831899.
[5] Brachmann, A., & Redies, C. (2017). Computational and Experimental Approaches to Visual Aesthetics. Frontiers in Computational Neuroscience , 11, 102. DOI: 10.3389/fncom.2017.00102.
[6] Silvia, P. J., Winterstein, B. P., Willse, J. T., et al. (2008). Assessing Creativity with Divergent Thinking Tasks: Exploring the Reliability and Validity of New Subjective Scoring Methods. Psychology of Aesthetics, Creativity, and the Arts , 2(2), 68–85.
[7] Said-Metwaly, S., Van den Noortgate, W., & Kyndt, E. (2017). Approaches to Measuring Creativity: A Systematic Literature Review. Creativity. Theories – Research – Applications , 4(2), 238–275.
[8] Harvey, S., & Berry, J. W. (2023). Toward a Meta-Theory of Creativity Forms: How Novelty and Usefulness Shape Creativity. Academy of Management Review .
[9] Beaty, R. E., & Johnson, D. R. (2021). Automating Creativity Assessment with SemDis: An Open Platform for Computing Semantic Distance. Behavior Research Methods , 53, 757–780. DOI: 10.3758/s13428-020-01453-w.
系列路徑
Paper 01|一輪到底是一輪什麼?:使用者回合、隱藏 LOOP 與單次智能的重新定義
Paper 02|智能到底算了一次什麼?:最小智能語意執行單位的候選理論
Paper 03|從認知到神經元:人腦如何跨層測量智能計算
Paper 04|從神經元到焦耳:智能計算的能量、熱力學與物理下界
Paper 05|計算不是只有 FLOPs:記憶體、互連、硬體占用與計算時空體積
Paper 06|成果品質到底怎麼量?:從形式化正確性到結構化智能品質
Paper 07|不要叫人類替自己的感覺打分數:IBQF 二元測量與低負擔品質評估
Paper 08|自然語言、圖像與創意如何被量?:高歧義成果的結構化品質空間
Paper 09|拿掉 LOOP 還剩多少智能?:單次智能、鷹架依賴與隱藏計算成本
Paper 10|一個答案值多少物理世界?:智能產率的統一計量框架