Series A — Algorithmic Observation, Recommendation & Platform Ecology
Paper A06 — Metric Success, Product Failure
指標成功與產品失敗:推薦平台的代理目標錯位、多方效用與長期生態治理
English Title: Metric Success, Product Failure: Proxy Objective Misalignment, Multi-Stakeholder Utility, and Long-Term Ecosystem Governance in Recommender PlatformsSeries: Algorithmic Observation, Recommendation & Platform EcologyPaper ID: A06Version: v0.1Date: 2026-08-31Status: Canonical UTF-8 SourceAuthor: Neo.K / EveMissLab
Abstract
現代推薦平台高度依賴可量化指標,例如 click-through rate、watch time、completion rate、return frequency、advertising revenue、conversion、retention 與 recommendation hit rate。這些指標具有工程可測量性與商業可操作性,因此自然成為模型優化、產品實驗與組織績效的核心。然而,平台真正關心的價值通常比任何單一 metric 更廣:使用者是否得到有價值的資訊與娛樂、是否願意長期回訪、創作者是否存在可持續生產的回報、內容供給是否保持多樣、廣告主是否獲得真實有效注意力,以及平台本身是否維持長期商業與社群資產。
本文延續 Series A 前五篇,提出 Metric–Product Separation framework。核心命題是:
M e t r i c S u c c e s s ≠ P r o d u c t S u c c e s s . \boxed{
MetricSuccess
\neq
ProductSuccess.
} M e t r i c S u ccess = P r o d u c tS u ccess .
本文將推薦平台視為 multi-stakeholder dynamic system,定義使用者效用、創作者效用、廣告主效用、平台短期收益、平台長期價值與內容供給健康度。本文分析 proxy optimization 如何經由 Goodhart-style failure、selection effects、survivorship bias、endogenous telemetry、exposure concentration 與 delayed externalities,使局部 KPI 上升但整體產品價值下降。
本文提出 Local Metric Improvement / Global Utility Degradation(LMIGUD)、Proxy Capture Ratio(PCR)、Delayed Ecosystem Cost(DEC)、Metric-Induced Behavior Shift(MIBS)、Stakeholder Utility Divergence(SUD)、Product Health Gap(PHG)與 Long-Horizon Platform Utility(LHPU)等概念。本文進一步區分 leading metrics、lagging metrics 與 hidden-state metrics,指出 creator exit、content supply concentration、user agency loss 與 advertiser attention quality 常比財務收入更晚或更難顯現。
本文最後提出一個多層治理架構:短期 ranking objective 不直接等同 platform objective;A/B test 不應只看 engagement uplift,而應配套 guardrail metrics、cohort retention、creator viability、long-tail observability、explicit-preference violation、paid-promotion conversion、surveyed satisfaction 與 delayed ecosystem audit。本文並將 Home、Following、Explore、Popular、Search 等觀察 surface 視為不同產品契約,以降低單一 proxy 對全部資訊入口的控制。
Series A 至此完成一個完整閉環:推薦系統決定可觀察世界、推論使用者偏好、由自身曝光污染行為證據、集中平台曝光、改變創作者供給,最後又可能因局部 KPI 成功而錯誤判定系統健康。本文的核心結論是:推薦平台真正需要優化的不是「更多互動」,而是多方、長期、可持續的資訊生態效用。
Keywords: recommender systems; Goodhart's law; proxy metrics; multi-stakeholder recommendation; user satisfaction; creator economy; platform governance; engagement; long-term utility; ecosystem health
1. Introduction
推薦系統工程非常擅長回答:
D i d C T R I n c r e a s e ? DidCTRIncrease? D i d C T R I n cr e a se ?
D i d W a t c h T i m e I n c r e a s e ? DidWatchTimeIncrease? D i d W a t c h T im e I n cr e a se ?
D i d C o n v e r s i o n I n c r e a s e ? DidConversionIncrease? D i d C o n v er s i o n I n cr e a se ?
D i d R e v e n u e I n c r e a s e ? DidRevenueIncrease? D i d R e v e n u e I n cr e a se ?
這些問題重要,而且不能被忽略。
但平台真正需要回答的問題通常是:
D i d T h e P r o d u c t B e c o m e B e t t e r ? \boxed{
DidTheProductBecomeBetter?
} D i d T h e P r o d u c tB eco m e B e tt er ?
兩者不是同一問題。
一個推薦改版可能:
C T R ↑ , CTR\uparrow, C T R ↑ ,
W a t c h T i m e ↑ , WatchTime\uparrow, W a t c h T im e ↑ ,
V i e w s ↑ , Views\uparrow, V i e w s ↑ ,
甚至:
R e v e n u e ↑ . Revenue\uparrow. R e v e n u e ↑ .
同時也可能:
U s e r A g e n c y ↓ , UserAgency\downarrow, U ser A g e n cy ↓ ,
L o n g T a i l O b s e r v a b i l i t y ↓ , LongTailObservability\downarrow, L o n g T ai l O b ser v abi l i t y ↓ ,
C r e a t o r E n t r y V i a b i l i t y ↓ , CreatorEntryViability\downarrow, C r e a t or E n t r y V iabi l i t y ↓ ,
I n f o r m a t i o n D i v e r s i t y ↓ . InformationDiversity\downarrow. I n f or ma t i o n D i v er s i t y ↓ .
這種情況並不矛盾。
因為前一組量:
M M M
是可觀測 proxy metrics,
後一組量:
H H H
是更廣的 product / ecosystem health state。
若:
M ≢ H , M
\not\equiv
H, M ≡ H ,
則:
max M \max M max M
不保證:
max H . \max H. max H .
本文因此提出 Series A 的收束命題:
M e t r i c S u c c e s s ≠ P r o d u c t S u c c e s s . \boxed{
MetricSuccess
\neq
ProductSuccess.
} M e t r i c S u ccess = P r o d u c tS u ccess .
2. Series A Recap
Series A 前五篇建立:
A01 — Observation
C t → O u , t ( s ) L u , t ( s ) . \mathcal{C}_t
\xrightarrow{
\mathcal{O}_{u,t}^{(s)}
}
L_{u,t}^{(s)}. C t O u , t ( s ) L u , t ( s ) .
推薦是 visibility allocation 與 observation policy。
A02 — Preference Semantics
A w a r e n e s s ≠ I n t e r e s t ≠ I n t e n t ≠ F e a s i b i l i t y . Awareness
\neq
Interest
\neq
Intent
\neq
Feasibility. A w a r e n ess = I n t er es t = I n t e n t = F e a s ibi l i t y .
A03 — Endogenous Preference Contamination
R e c o m m e n d a t i o n → E x p o s u r e → T e l e m e t r y → P r e f e r e n c e E s t i m a t e . Recommendation
\rightarrow
Exposure
\rightarrow
Telemetry
\rightarrow
PreferenceEstimate. R eco mm e n d a t i o n → E x p os u r e → T e l e m e t r y → P r e f er e n ce E s t ima t e .
平台可能用自己產生的行為資料證明自己的推論。
A04 — Platform-Induced Exposure Bubble
T o p i c D i v e r s i t y ≠ C r e a t o r D i v e r s i t y ≠ E x p o s u r e D i v e r s i t y . TopicDiversity
\neq
CreatorDiversity
\neq
ExposureDiversity. T o p i cD i v er s i t y = C r e a t or D i v er s i t y = E x p os u r eD i v er s i t y .
A05 — Creator Ecology
E x p o s u r e → E v i d e n c e → A u d i e n c e → C r e a t o r R e t u r n → F u t u r e S u p p l y . Exposure
\rightarrow
Evidence
\rightarrow
Audience
\rightarrow
CreatorReturn
\rightarrow
FutureSupply. E x p os u r e → E v i d e n ce → A u d i e n ce → C r e a t or R e t u r n → F u t u r e S u ppl y .
A06 問的是:
如果以上各層正在退化,但平台核心 KPI 仍然增加,組織是否可能認為改版成功?
答案是:
Y e s . \boxed{
Yes.
} Y es .
而這正是最難偵測的 failure mode。
3. Related Work
3.1 Accuracy is not enough
Recommender-system evaluation 很早就指出 predictive accuracy 並不足以衡量推薦品質。McNee、Riedl 與 Konstan 的經典論文 “Being Accurate is Not Enough” 強調 novelty、serendipity、diversity 等 beyond-accuracy qualities 對實際推薦價值的重要性 [2]。
因此:
P r e d i c t i o n A c c u r a c y ≠ R e c o m m e n d a t i o n U t i l i t y . PredictionAccuracy
\neq
RecommendationUtility. P r e d i c t i o n A cc u r a cy = R eco mm e n d a t i o n U t i l i t y .
3.2 User experience
Knijnenburg 等人提出 recommender systems user-experience framework,區分客觀系統面、使用者主觀感受、互動與個人/情境特徵 [3]。
這表示即使推薦系統 offline metrics 很好:
O f f l i n e S c o r e ↑ OfflineScore\uparrow O f f l in e S cor e ↑
也不能直接推出:
U s e r E x p e r i e n c e ↑ . UserExperience\uparrow. U ser E x p er i e n ce ↑ .
3.3 Multi-stakeholder recommendation
Burke、Abdollahpouri、Mobasher 等人指出,推薦系統往往存在多個利益相關者,包括 users、providers、platforms 等,推薦價值不應只以單一 consumer utility 衡量 [4][5]。
因此平台 objective 更接近:
U = f ( U c o n s u m e r , U p r o v i d e r , U p l a t f o r m , … ) . U
=
f(
U_{consumer},
U_{provider},
U_{platform},
\ldots
). U = f ( U co n s u m er , U p r o v i d er , U pl a t f or m , … ) .
3.4 Provider fairness and ecosystem effects
Two-sided fairness 與 provider fairness research 進一步指出,推薦排序決定商品或內容提供者獲得的 exposure,而只最大化 consumer-side metric 可能導致不平衡的 provider opportunities [6][7]。
3.5 Feedback loops
Series A 已引用的大量研究顯示,推薦輸出會影響後續資料、熱門度與使用者行為 [8][9]。
因此:
M e t r i c t Metric_t M e t r i c t
不是被動測量世界。
它可能部分來自:
P o l i c y t Policy_t P o l i c y t
本身。
3.6 Goodhart-style proxy failure
Goodhart-style reasoning 可概括為:
當某個 measure 被強烈用作 target 時,它與真正目標之間原本穩定的關係可能失效。
Manheim 與 Garrabrant 將 Goodhart effects 分類為 regressional、extremal、causal 與 adversarial forms [10]。
推薦平台特別容易出現:
C a u s a l G o o d h a r t CausalGoodhart C a u s a l G oo d ha r t
因為平台直接改變:
E x p o s u r e , B e h a v i o r , T e l e m e t r y . Exposure,
Behavior,
Telemetry. E x p os u r e , B e ha v i or , T e l e m e t r y .
4. Metric–Product Separation
令:
M t = ( C T R t , W a t c h T i m e t , R e v e n u e t , R e t e n t i o n t , C o n v e r s i o n t ) M_t
=
(
CTR_t,
WatchTime_t,
Revenue_t,
Retention_t,
Conversion_t
) M t = ( C T R t , W a t c h T im e t , R e v e n u e t , R e t e n t i o n t , C o n v er s i o n t )
為 observable metrics。
令:
H t = ( U s e r U t i l i t y t , C r e a t o r H e a l t h t , S u p p l y D i v e r s i t y t , A g e n c y t , A d v e r t i s e r V a l u e t , L o n g T e r m P l a t f o r m V a l u e t ) H_t
=
(
UserUtility_t,
CreatorHealth_t,
SupplyDiversity_t,
Agency_t,
AdvertiserValue_t,
LongTermPlatformValue_t
) H t = ( U ser U t i l i t y t , C r e a t or H e a l t h t , S u ppl y D i v er s i t y t , A g e n c y t , A d v er t i ser V a l u e t , L o n g T er m P l a t f or mV a l u e t )
為 product-health state。
Metric model:
M t = g ( H t , P o l i c y t , U I t , M e a s u r e m e n t t , E x t e r n a l S t a t e t ) . M_t
=
g(
H_t,
Policy_t,
UI_t,
Measurement_t,
ExternalState_t
). M t = g ( H t , P o l i c y t , U I t , M e a s u r e m e n t t , E x t er na l S t a t e t ) .
注意:
M t M_t M t
不是:
H t H_t H t
的直接同義詞。
因此:
M t is an observation of the product, not the product itself. \boxed{
M_t
\text{ is an observation of the product, not the product itself.}
} M t is an observation of the product, not the product itself.
這與 A01 的 observation distinction 在產品治理層重新出現。
5. Proxy Objective
真實平台目標:
G t . G_t. G t .
但 G t G_t G t 通常不可直接即時測量。
因此平台使用 proxy:
M t . M_t. M t .
訓練或產品決策變成:
max π E [ M t ∣ π ] . \max_\pi
E[
M_t
\mid
\pi
]. π max E [ M t ∣ π ] .
但真正希望:
max π E [ G t ∣ π ] . \max_\pi
E[
G_t
\mid
\pi
]. π max E [ G t ∣ π ] .
只有在:
C o r r ( M , G ) Corr(M,G) C or r ( M , G )
在 intervention 後仍穩定時,
前者才可靠代表後者。
Goodhart failure 發生於:
C o r r ( M , G ∣ d o ( π ) ) ↓ . Corr(
M,G
\mid
do(\pi)
)
\downarrow. C or r ( M , G ∣ d o ( π )) ↓ .
6. Causal Metric Inflation
A03 已指出:
R e c o m m e n d a t i o n → E x p o s u r e → T e l e m e t r y . Recommendation
\rightarrow
Exposure
\rightarrow
Telemetry. R eco mm e n d a t i o n → E x p os u r e → T e l e m e t r y .
因此平台可以透過 UI 或 policy 直接提升 metric。
例如:
A u t o p l a y ↑ ⇒ W a t c h E v e n t s ↑ . Autoplay\uparrow
\Rightarrow
WatchEvents\uparrow. A u t o pl a y ↑⇒ W a t c h E v e n t s ↑ .
R e p e a t E x p o s u r e ↑ ⇒ I m p r e s s i o n s ↑ . RepeatExposure\uparrow
\Rightarrow
Impressions\uparrow. R e p e a tE x p os u r e ↑⇒ I m p r ess i o n s ↑ .
H e a d C o n t e n t S h a r e ↑ ⇒ P r e d i c t e d C T R A c c u r a c y ↑ . HeadContentShare\uparrow
\Rightarrow
PredictedCTRAccuracy\uparrow. H e a d C o n t e n tS ha r e ↑⇒ P r e d i c t e d C T R A cc u r a cy ↑ .
這些數字變化未必代表:
U s e r P r e f e r e n c e A l i g n m e n t ↑ . UserPreferenceAlignment\uparrow. U ser P r e f er e n ce A l i g nm e n t ↑ .
本文稱:
M e t r i c I n f l a t i o n c a u s a l . \boxed{
MetricInflation_{causal}.
} M e t r i c I n f l a t i o n c a u s a l .
7. Local Metric Improvement / Global Utility Degradation
本文定義:
L M I G U D . \boxed{
LMIGUD.
} L M I G U D .
若存在某 policy change:
π 0 → π 1 \pi_0
\rightarrow
\pi_1 π 0 → π 1
使:
M ( π 1 ) > M ( π 0 ) M(\pi_1)>M(\pi_0) M ( π 1 ) > M ( π 0 )
但:
U g l o b a l ( π 1 ) < U g l o b a l ( π 0 ) , U^{global}(\pi_1)
<
U^{global}(\pi_0), U g l o ba l ( π 1 ) < U g l o ba l ( π 0 ) ,
則發生:
L M I G U D = 1. LMIGUD=1. L M I G U D = 1.
其中:
U g l o b a l U^{global} U g l o ba l
可以包含多 stakeholder 與長期效用。
這是:
M e t r i c S u c c e s s ∧ P r o d u c t F a i l u r e . \boxed{
MetricSuccess
\land
ProductFailure.
} M e t r i c S u ccess ∧ P r o d u c tF ai l u r e .
的形式化定義。
8. Multi-Stakeholder Utility
令 stakeholders:
S = { U s e r , C r e a t o r , A d v e r t i s e r , P l a t f o r m , S o c i e t y } . \mathcal{S}
=
\{
User,
Creator,
Advertiser,
Platform,
Society
\}. S = { U ser , C r e a t or , A d v er t i ser , P l a t f or m , S oc i e t y } .
總效用:
U t M S = ∑ s ∈ S w s U s ( t ) . U^{MS}_t
=
\sum_{s\in\mathcal{S}}
w_s
U_s(t). U t M S = s ∈ S ∑ w s U s ( t ) .
其中:
w s w_s w s
不是自然常數,而是 platform governance choice。
例如:
U u s e r = f ( R e l e v a n c e , S a t i s f a c t i o n , A g e n c y , T i m e V a l u e , I n f o r m a t i o n G a i n ) . U_{user}
=
f(
Relevance,
Satisfaction,
Agency,
TimeValue,
InformationGain
). U u ser = f ( R e l e v an ce , S a t i s f a c t i o n , A g e n cy , T im e V a l u e , I n f or ma t i o n G ain ) .
U c r e a t o r = f ( E x p o s u r e , A u d i e n c e G r o w t h , R e v e n u e , P r e d i c t a b i l i t y , V i a b i l i t y ) . U_{creator}
=
f(
Exposure,
AudienceGrowth,
Revenue,
Predictability,
Viability
). U cr e a t or = f ( E x p os u r e , A u d i e n ce G r o w t h , R e v e n u e , P r e d i c t abi l i t y , V iabi l i t y ) .
U a d v e r t i s e r = f ( Q u a l i f i e d A t t e n t i o n , C o n v e r s i o n , B r a n d L i f t , I n c r e m e n t a l i t y ) . U_{advertiser}
=
f(
QualifiedAttention,
Conversion,
BrandLift,
Incrementality
). U a d v er t i ser = f ( Q u a l i f i e d A tt e n t i o n , C o n v er s i o n , B r an d L i f t , I n cr e m e n t a l i t y ) .
U p l a t f o r m = f ( R e v e n u e , R e t e n t i o n , R i s k , F u t u r e S u p p l y , T r u s t ) . U_{platform}
=
f(
Revenue,
Retention,
Risk,
FutureSupply,
Trust
). U pl a t f or m = f ( R e v e n u e , R e t e n t i o n , R i s k , F u t u r e S u ppl y , T r u s t ) .
9. Platform Objective Is Not Ranking Objective
Ranking model 可能優化:
J r a n k = α C T R + β W a t c h T i m e + γ C o n v e r s i o n . J_{rank}
=
\alpha CTR
+
\beta WatchTime
+
\gamma Conversion. J r ank = α C T R + β W a t c h T im e + γ C o n v er s i o n .
但 platform objective:
J p l a t f o r m J_{platform} J pl a t f or m
應該更廣。
因此:
J r a n k ≠ J p l a t f o r m . \boxed{
J_{rank}
\neq
J_{platform}.
} J r ank = J pl a t f or m .
合理架構是:
J r a n k ⊂ J p l a t f o r m J_{rank}
\subset
J_{platform} J r ank ⊂ J pl a t f or m
並且:
J p l a t f o r m J_{platform} J pl a t f or m
對:
J r a n k J_{rank} J r ank
施加 constraints / guardrails。
10. Proxy Capture Ratio
本文提出 Proxy Capture Ratio:
P C R = Δ M Δ U t a r g e t + ϵ . PCR
=
\frac{
\Delta M
}{
\Delta U^{target}+\epsilon
}. P C R = Δ U t a r g e t + ϵ Δ M .
如果:
Δ M ≫ 0 \Delta M\gg0 Δ M ≫ 0
但:
Δ U t a r g e t ≈ 0 , \Delta U^{target}\approx0, Δ U t a r g e t ≈ 0 ,
則:
P C R ≫ 1. PCR\gg1. P C R ≫ 1.
代表 policy 很有效地提升 metric,但沒有等比例提升真正目標。
若:
Δ U t a r g e t < 0 , \Delta U^{target}<0, Δ U t a r g e t < 0 ,
則 proxy capture 更嚴重。
11. Engagement Is Not Satisfaction
使用者觀看很久可能代表:
真正喜歡;
任務需要;
autoplay;
sunk-cost;
search friction;
interface trapping;
情緒性 consumption;
找不到替代內容。
因此:
W a t c h T i m e ≠ S a t i s f a c t i o n . WatchTime
\neq
Satisfaction. W a t c h T im e = S a t i s f a c t i o n .
更精確:
W a t c h T i m e = f ( S a t i s f a c t i o n , F r i c t i o n , H a b i t , A u t o p l a y , A v a i l a b i l i t y , I n t e n t , C o n t e n t L e n g t h ) . WatchTime
=
f(
Satisfaction,
Friction,
Habit,
Autoplay,
Availability,
Intent,
ContentLength
). W a t c h T im e = f ( S a t i s f a c t i o n , F r i c t i o n , H abi t , A u t o pl a y , A v ai l abi l i t y , I n t e n t , C o n t e n t L e n g t h ) .
因此:
max W a t c h T i m e \max WatchTime max W a t c h T im e
不能自然等價:
max S a t i s f a c t i o n . \max Satisfaction. max S a t i s f a c t i o n .
12. Retention Is Not Loyalty
同樣:
R e t e n t i o n Retention R e t e n t i o n
可能來自:
H i g h U t i l i t y HighUtility H i g h U t i l i t y
但也可能來自:
H i g h S w i t c h i n g C o s t , N e t w o r k E f f e c t s , C o n t e n t L o c k I n , L a c k O f A l t e r n a t i v e s . HighSwitchingCost,
NetworkEffects,
ContentLockIn,
LackOfAlternatives. H i g h S w i t c hin g C os t , N e tw or k E f f ec t s , C o n t e n t L oc k I n , L a c k O f A l t er na t i v es .
因此:
R e t e n t i o n ↑ Retention\uparrow R e t e n t i o n ↑
不能單獨證明:
T r u s t ↑ . Trust\uparrow. T r u s t ↑ .
真正要觀察:
V o l u n t a r y P r e f e r e n c e VoluntaryPreference V o l u n t a r y P r e f er e n ce
與:
D e p e n d e n c e . Dependence. D e p e n d e n ce .
13. Views Are Not Qualified Attention
廣告/創作者經濟常使用:
V i e w s , I m p r e s s i o n s , W a t c h T i m e Views,
Impressions,
WatchTime V i e w s , I m p r ess i o n s , W a t c h T im e
代表 attention。
但真正具有商業價值的是:
Q u a l i f i e d A t t e n t i o n . QualifiedAttention. Q u a l i f i e d A tt e n t i o n .
定義:
Q A = A c t i v e A t t e n t i o n × A u d i e n c e F i t × I n t e n t i o n a l i t y . QA
=
ActiveAttention
\times
AudienceFit
\times
Intentionality. Q A = A c t i v e A tt e n t i o n × A u d i e n ce F i t × I n t e n t i o na l i t y .
因此:
1 , 000 , 000 1,000,000 1 , 000 , 000
個 autoplay-heavy views,
可能小於:
100 , 000 100,000 100 , 000
個 high-intent views 的 advertiser value。
14. Advertising Metric Dilution
如果:
V i e w s ↑ Views\uparrow V i e w s ↑
主要來自:
P a s s i v e E x p o s u r e ↑ , PassiveExposure\uparrow, P a ss i v e E x p os u r e ↑ ,
則:
C P V CPV C P V
可能表面改善,
但:
C o n v e r s i o n P e r Q u a l i f i e d V i e w ConversionPerQualifiedView C o n v er s i o n P er Q u a l i f i e d V i e w
可能下降。
本文定義:
A M D = 1 − Q u a l i f i e d A t t e n t i o n G r o w t h R a w V i e w G r o w t h + ϵ . AMD
=
1-
\frac{
QualifiedAttentionGrowth
}{
RawViewGrowth+\epsilon
}. A M D = 1 − R a w V i e w G r o w t h + ϵ Q u a l i f i e d A tt e n t i o n G r o w t h .
即 Advertising Metric Dilution。
若:
A M D ↑ , AMD\uparrow, A M D ↑ ,
raw traffic 與真正商業 attention 的關係正在變弱。
15. Short-Term Revenue versus Long-Term Asset
平台短期收入:
R t . R_t. R t .
平台長期資產:
A t p l a t f o r m = ( U s e r T r u s t , C r e a t o r S u p p l y , C o n t e n t G r a p h , B r a n d , A d v e r t i s e r C o n f i d e n c e , S o c i a l R e l a t i o n s ) . A_t^{platform}
=
(
UserTrust,
CreatorSupply,
ContentGraph,
Brand,
AdvertiserConfidence,
SocialRelations
). A t pl a t f or m = ( U ser T r u s t , C r e a t or S u ppl y , C o n t e n tG r a p h , B r an d , A d v er t i ser C o n f i d e n ce , S oc ia l R e l a t i o n s ) .
長期價值:
V t = R t + λ A t p l a t f o r m . V_t
=
R_t
+
\lambda
A_t^{platform}. V t = R t + λ A t pl a t f or m .
若 policy:
Δ R t > 0 \Delta R_t>0 Δ R t > 0
但:
Δ A t p l a t f o r m < 0 , \Delta A_t^{platform}<0, Δ A t pl a t f or m < 0 ,
則短期財報可能改善,
但:
V t + T V_{t+T} V t + T
下降。
16. Delayed Ecosystem Cost
本文定義 Delayed Ecosystem Cost:
D E C ( π ) = ∑ τ = 1 T γ τ [ H b a s e l i n e ( t + τ ) − H π ( t + τ ) ] + . DEC(\pi)
=
\sum_{\tau=1}^{T}
\gamma^\tau
[
H_{baseline}(t+\tau)
-
H_{\pi}(t+\tau)
]_+. D E C ( π ) = τ = 1 ∑ T γ τ [ H ba se l in e ( t + τ ) − H π ( t + τ ) ] + .
其中:
[ ⋅ ] + [\cdot]_+ [ ⋅ ] +
取 positive loss。
典型 DEC 包括:
creator exit;
newcomer decline;
content diversity loss;
user habit weakening;
advertiser ROI decline;
trust erosion。
這些通常不會在:
D a y 1 Day1 D a y 1
A/B test 中完整出現。
17. Leading, Lagging, and Hidden Metrics
17.1 Leading metrics
快速變化:
C T R , S e s s i o n L e n g t h , W a t c h E v e n t s . CTR,
SessionLength,
WatchEvents. C T R , S ess i o n L e n g t h , W a t c h E v e n t s .
17.2 Lagging metrics
較慢顯現:
C r e a t o r E x i t , C o h o r t R e t e n t i o n , A d v e r t i s e r R e n e w a l , O r g a n i c E n t r y . CreatorExit,
CohortRetention,
AdvertiserRenewal,
OrganicEntry. C r e a t or E x i t , C o h or tR e t e n t i o n , A d v er t i ser R e n e w a l , O r g ani c E n t r y .
17.3 Hidden-state metrics
難以直接測量:
T r u s t , A g e n c y , I n f o r m a t i o n U t i l i t y , O p p o r t u n i t y C o s t , C r e a t o r D i s c o u r a g e m e n t . Trust,
Agency,
InformationUtility,
OpportunityCost,
CreatorDiscouragement. T r u s t , A g e n cy , I n f or ma t i o n U t i l i t y , O pp or t u ni t y C os t , C r e a t or D i sco u r a g e m e n t .
平台如果只優化:
L e a d i n g M e t r i c s , LeadingMetrics, L e a d in g M e t r i cs ,
容易忽略:
L a g g i n g L o s s . LaggingLoss. L a g g in g L oss .
18. Metric-Induced Behavior Shift
當 creator 知道:
M e t r i c x Metric_x M e t r i c x
決定 exposure,
creator 會適應:
S t r a t e g y k → O p t i m i z e ( M e t r i c x ) . Strategy_k
\rightarrow
Optimize(Metric_x). S t r a t e g y k → O pt imi z e ( M e t r i c x ) .
本文定義:
M I B S = ∂ C r e a t o r B e h a v i o r ∂ M e t r i c I n c e n t i v e . MIBS
=
\frac{
\partial CreatorBehavior
}{
\partial MetricIncentive
}. M I B S = ∂ M e t r i c I n ce n t i v e ∂ C r e a t or B e ha v i or .
例如:
C T R CTR C T R
高度重要時,
可能增加:
T h u m b n a i l O p t i m i z a t i o n , T i t l e O p t i m i z a t i o n , T o p i c C o n v e r g e n c e . ThumbnailOptimization,
TitleOptimization,
TopicConvergence. T h u mbnai l O pt imi z a t i o n , T i tl e O pt imi z a t i o n , T o p i c C o n v er g e n ce .
這不必然是壞事。
但:
M I B S MIBS M I B S
過強時:
C o n t e n t Content C o n t e n t
開始服務 metric,而不是 audience value。
19. Platform Goodhart Loop
平台 Goodhart loop:
M e t r i c C h o s e n MetricChosen M e t r i c C h ose n
⇓ \Downarrow ⇓
R a n k i n g O p t i m i z e s M e t r i c RankingOptimizesMetric R ank in g O pt imi z es M e t r i c
⇓ \Downarrow ⇓
U I A n d C r e a t o r s A d a p t UIAndCreatorsAdapt U I A n d C r e a t or s A d a pt
⇓ \Downarrow ⇓
M e t r i c D i s t r i b u t i o n C h a n g e s MetricDistributionChanges M e t r i cD i s t r ib u t i o n C han g es
⇓ \Downarrow ⇓
M e t r i c B e c o m e s L e s s I n f o r m a t i v e A b o u t G o a l . MetricBecomesLessInformativeAboutGoal. M e t r i c B eco m es L ess I n f or ma t i v e A b o u tG o a l .
形式:
O p t i m i z e ( M ) → C h a n g e D a t a G e n e r a t i n g P r o c e s s → C o r r ( M , G ) ↓ . \boxed{
Optimize(M)
\rightarrow
ChangeDataGeneratingProcess
\rightarrow
Corr(M,G)\downarrow.
} O pt imi z e ( M ) → C han g eD a t a G e n er a t in g P r ocess → C or r ( M , G ) ↓ .
這是推薦平台最重要的 proxy risk。
20. Stakeholder Utility Divergence
本文定義:
S U D = V a r ( Δ U s : s ∈ S ) . SUD
=
Var(
\Delta U_s
:
s\in\mathcal{S}
). S U D = V a r ( Δ U s : s ∈ S ) .
若某改版:
U p l a t f o r m s h o r t ↑ U_{platform}^{short}\uparrow U pl a t f or m s h or t ↑
但:
U c r e a t o r ↓ , U_{creator}\downarrow, U cr e a t or ↓ ,
U u s e r ↓ , U_{user}\downarrow, U u ser ↓ ,
則:
S U D ↑ . SUD\uparrow. S U D ↑ .
SUD 本身不代表 policy 必然錯。
但高:
S U D SUD S U D
表示平台必須做顯式 governance trade-off,而不能稱為「全面改善」。
21. Product Health Gap
定義:
P H G = N o r m a l i z e d M e t r i c P e r f o r m a n c e − N o r m a l i z e d E c o s y s t e m H e a l t h . PHG
=
NormalizedMetricPerformance
-
NormalizedEcosystemHealth. P H G = N or ma l i z e d M e t r i c P er f or man ce − N or ma l i z e d E cosy s t e m H e a l t h .
如果:
P H G ≫ 0 , PHG\gg0, P H G ≫ 0 ,
表示 dashboard 看起來很健康,
但 ecosystem metrics 沒有同步改善。
這就是:
D a s h b o a r d R e a l i t y G a p . \boxed{
DashboardRealityGap.
} D a s hb o a r d R e a l i t y G a p .
22. Metric Portfolio
單一 metric 容易 Goodhart。
因此平台應使用:
M = { M 1 , … , M n } . \mathcal{M}
=
\{
M_1,\ldots,M_n
\}. M = { M 1 , … , M n } .
例如:
M e n g a g e m e n t , M_{engagement}, M e n g a g e m e n t ,
M s a t i s f a c t i o n , M_{satisfaction}, M s a t i s f a c t i o n ,
M a g e n c y , M_{agency}, M a g e n cy ,
M d i v e r s i t y , M_{diversity}, M d i v er s i t y ,
M c r e a t o r , M_{creator}, M cr e a t or ,
M a d v e r t i s e r , M_{advertiser}, M a d v er t i ser ,
M l o n g t e r m . M_{longterm}. M l o n g t er m .
重要的是:
M e t r i c P o r t f o l i o MetricPortfolio M e t r i c P or t f o l i o
也不能只變成更多 KPI。
必須有:
E x p l i c i t T r a d e o f f P o l i c y . \boxed{
ExplicitTradeoffPolicy.
} E x pl i c i tT r a d eo f f P o l i cy .
23. Guardrail Metrics
一個 ranking experiment 可以最大化:
P r i m a r y M e t r i c . PrimaryMetric. P r ima r y M e t r i c .
但需要 guardrails:
G = { G 1 , … , G m } . G=
\{
G_1,\ldots,G_m
\}. G = { G 1 , … , G m } .
例如:
C T R ↑ CTR\uparrow C T R ↑
只有在:
E x p l i c i t P r e f e r e n c e V i o l a t i o n ≤ θ 1 , ExplicitPreferenceViolation\leq\theta_1, E x pl i c i tP r e f er e n ce V i o l a t i o n ≤ θ 1 ,
C r e a t o r E x p o s u r e G i n i ≤ θ 2 , CreatorExposureGini\leq\theta_2, C r e a t or E x p os u r e G ini ≤ θ 2 ,
L o n g T a i l O b s e r v a b i l i t y ≥ θ 3 , LongTailObservability\geq\theta_3, L o n g T ai l O b ser v abi l i t y ≥ θ 3 ,
U s e r S a t i s f a c t i o n ≥ θ 4 UserSatisfaction\geq\theta_4 U ser S a t i s f a c t i o n ≥ θ 4
時才接受。
因此:
U p l i f t without guardrails is not sufficient evidence of improvement. \boxed{
Uplift
\text{ without guardrails is not sufficient evidence of improvement.}
} U pl i f t without guardrails is not sufficient evidence of improvement.
24. Hard Constraints versus Soft Objectives
不是所有價值都適合加權平均。
例如:
J = C T R − λ P r i v a c y V i o l a t i o n J
=
CTR
-
\lambda PrivacyViolation J = C T R − λ P r i v a cy V i o l a t i o n
可能暗示:
H i g h C T R HighCTR H i g h C T R
可以補償 privacy violation。
但某些條件應是:
P r i v a c y V i o l a t i o n = 0. PrivacyViolation=0. P r i v a cy V i o l a t i o n = 0.
同理:
E x p l i c i t B l o c k V i o l a t i o n = 0 ExplicitBlockViolation=0 E x pl i c i tB l oc k V i o l a t i o n = 0
應更接近 hard constraint。
因此推薦治理應區分:
S o f t O b j e c t i v e SoftObjective S o f tO bj ec t i v e
與:
H a r d C o n s t r a i n t . HardConstraint. H a r d C o n s t r ain t .
25. Long-Horizon Platform Utility
本文提出:
L H P U = ∑ t = 0 T γ t [ w u U u s e r , t + w c U c r e a t o r , t + w a U a d v e r t i s e r , t + w p U p l a t f o r m , t ] . \boxed{
LHPU
=
\sum_{t=0}^{T}
\gamma^t
[
w_uU_{user,t}
+
w_cU_{creator,t}
+
w_aU_{advertiser,t}
+
w_pU_{platform,t}
].
} L H P U = t = 0 ∑ T γ t [ w u U u ser , t + w c U cr e a t or , t + w a U a d v er t i ser , t + w p U pl a t f or m , t ] .
其中:
0 < γ ≤ 1. 0<\gamma\leq1. 0 < γ ≤ 1.
若:
γ → 0 , \gamma\rightarrow0, γ → 0 ,
系統近似短期 extraction。
若:
γ \gamma γ
較高,
future supply、trust 與 retention 會更重要。
26. Dynamic Platform State
平台 state:
S t = ( U s e r s t , C r e a t o r s t , C o n t e n t t , R e l a t i o n s t , A d v e r t i s e r s t , T r u s t t ) . S_t
=
(
Users_t,
Creators_t,
Content_t,
Relations_t,
Advertisers_t,
Trust_t
). S t = ( U ser s t , C r e a t or s t , C o n t e n t t , R e l a t i o n s t , A d v er t i ser s t , T r u s t t ) .
policy:
π t . \pi_t. π t .
狀態轉移:
S t + 1 = F ( S t , π t , E x t e r n a l t ) . S_{t+1}
=
F(
S_t,\pi_t,External_t
). S t + 1 = F ( S t , π t , E x t er na l t ) .
因此推薦系統不是 static ranking problem。
它是:
C o n t r o l P r o b l e m O v e r P l a t f o r m S t a t e . \boxed{
ControlProblemOverPlatformState.
} C o n t r o l P r o b l e m O v er P l a t f or m S t a t e .
27. Myopic Optimization
若平台只解:
π t ∗ = arg max π R t ( π ) , \pi_t^*
=
\arg\max_\pi
R_t(\pi), π t ∗ = arg π max R t ( π ) ,
可能得到:
M y o p i c P o l i c y . MyopicPolicy. M y o p i c P o l i cy .
更合理:
π ∗ = arg max π E [ ∑ τ = t t + T γ τ − t U ( S τ ) ] . \pi^*
=
\arg\max_\pi
E[
\sum_{\tau=t}^{t+T}
\gamma^{\tau-t}
U(S_\tau)
]. π ∗ = arg π max E [ τ = t ∑ t + T γ τ − t U ( S τ )] .
這也是 recommender systems 從 supervised ranking 走向 sequential decision-making 的重要理由之一。
28. User Time as a Scarce Resource
平台常將:
W a t c h T i m e WatchTime W a t c h T im e
視為收益。
但對 user:
T i m e Time T im e
是成本。
因此同一分鐘:
+ 1 +1 + 1
對 platform engagement,
可能是:
− 1 -1 − 1
對 user opportunity budget。
真正 user net utility:
N U u = V a l u e C o n s u m e d − T i m e C o s t − A t t e n t i o n C o s t − O p p o r t u n i t y C o s t . NU_u
=
ValueConsumed
-
TimeCost
-
AttentionCost
-
OpportunityCost. N U u = V a l u e C o n s u m e d − T im e C os t − A tt e n t i o n C os t − O pp or t u ni t y C os t .
如果:
W a t c h T i m e ↑ WatchTime\uparrow W a t c h T im e ↑
但:
N U u ↓ , NU_u\downarrow, N U u ↓ ,
則 engagement uplift 不等於 user welfare uplift。
29. Information Platforms and Freshness
對新聞、AI、金融、天氣等高時效資訊:
V a l u e ( v , t ) = V 0 e − λ Δ t . Value(v,t)
=
V_0e^{-\lambda\Delta t}. V a l u e ( v , t ) = V 0 e − λ Δ t .
若 recommender 偏好:
P o p u l a r O l d C o n t e n t PopularOldContent P o p u l a r O l d C o n t e n t
而壓制:
F r e s h N i c h e C o n t e n t , FreshNicheContent, F r es h N i c h e C o n t e n t ,
可能:
W a t c h T i m e WatchTime W a t c h T im e
保持高,
但:
I n f o r m a t i o n U t i l i t y ↓ . InformationUtility\downarrow. I n f or ma t i o n U t i l i t y ↓ .
因此 domain-specific utility 必須納入:
F r e s h n e s s . Freshness. F r es hn ess .
30. Surface-Level Product Contracts
A01 已區分:
H o m e , F o l l o w i n g , E x p l o r e , P o p u l a r , S e a r c h . Home,
Following,
Explore,
Popular,
Search. H o m e , F o l l o w in g , E x pl or e , P o p u l a r , S e a r c h .
A06 進一步認為:
每個 surface 是:
P r o d u c t C o n t r a c t . \boxed{
ProductContract.
} P r o d u c tC o n t r a c t .
例如:
F o l l o w i n g Following F o l l o w in g
承諾較高 explicit relation fidelity。
E x p l o r e Explore E x pl or e
承諾較高 novelty。
P o p u l a r Popular P o p u l a r
承諾 global popularity。
如果全部 surface 最後都被:
E n g a g e m e n t M e t r i c EngagementMetric E n g a g e m e n tM e t r i c
統一支配,
則:
S u r f a c e D i f f e r e n t i a t i o n → 0. SurfaceDifferentiation\rightarrow0. S u r f a ceD i f f er e n t ia t i o n → 0.
這是一種 product-contract collapse。
31. Multi-Surface Objective Separation
令:
J H J_H J H
為 Home objective,
J F J_F J F
為 Following objective,
J E J_E J E
為 Explore objective,
J P J_P J P
為 Popular objective,
J S J_S J S
為 Search objective。
要求:
J H ≠ J F ≠ J E ≠ J P ≠ J S . J_H
\neq
J_F
\neq
J_E
\neq
J_P
\neq
J_S. J H = J F = J E = J P = J S .
不是數學上必須完全不同,
而是 product intent 不應被單一:
J e n g a g e m e n t J_{engagement} J e n g a g e m e n t
完全吞併。
32. A/B Testing Failure
標準 A/B test:
T r e a t m e n t → Δ M e t r i c . Treatment
\rightarrow
\Delta Metric. T r e a t m e n t → Δ M e t r i c .
如果 window:
T A B T_{AB} T A B
很短,
但 ecosystem effect delay:
T e c o ≫ T A B , T_{eco}\gg T_{AB}, T eco ≫ T A B ,
則:
A / B A/B A / B
只能觀察:
S h o r t T e r m E f f e c t . ShortTermEffect. S h or tT er m E f f ec t .
因此:
N o S h o r t T e r m H a r m ⇏ N o L o n g T e r m H a r m . \boxed{
NoShortTermHarm
\not\Rightarrow
NoLongTermHarm.
} N o S h or tT er m H a r m ⇒ N o L o n g T er m H a r m .
33. Delayed Holdout
平台可以保留:
L o n g T e r m H o l d o u t C o h o r t . LongTermHoldoutCohort. L o n g T er m H o l d o u tC o h or t .
對:
30 , 60 , 90 , 180 30,
60,
90,
180 30 , 60 , 90 , 180
天觀察:
R e t e n t i o n , S a t i s f a c t i o n , C r e a t o r D i v e r s i t y , C r e a t o r E x i t , A d V a l u e . Retention,
Satisfaction,
CreatorDiversity,
CreatorExit,
AdValue. R e t e n t i o n , S a t i s f a c t i o n , C r e a t or D i v er s i t y , C r e a t or E x i t , A d V a l u e .
這提供:
D E C DEC D E C
的估計。
34. Cohort Analysis
aggregate metrics 容易掩蓋:
U s e r G r o u p D i f f e r e n c e s . UserGroupDifferences. U ser G r o u p D i f f er e n ces .
例如:
M a i n s t r e a m U s e r s ↑ MainstreamUsers\uparrow M ain s t r e am U ser s ↑
但:
N i c h e U s e r s ↓ . NicheUsers\downarrow. N i c h e U ser s ↓ .
總體:
W a t c h T i m e ↑ . WatchTime\uparrow. W a t c h T im e ↑ .
因此需要:
Δ U m a i n s t r e a m , \Delta U_{mainstream}, Δ U main s t r e am ,
Δ U n i c h e , \Delta U_{niche}, Δ U ni c h e ,
Δ U h i g h I n t e n t , \Delta U_{highIntent}, Δ U hi g h I n t e n t ,
Δ U n e w . \Delta U_{new}. Δ U n e w .
creator 同理。
35. Survivorship Bias
若:
C r e a t o r s E x i t CreatorsExit C r e a t or s E x i t
之後不再出現在 active dataset,
平台可能觀察:
A v e r a g e C r e a t o r P e r f o r m a n c e ↑ . AverageCreatorPerformance\uparrow. A v er a g e C r e a t or P er f or man ce ↑ .
這甚至可能只是:
L o w P e r f o r m a n c e C r e a t o r s R e m o v e d F r o m D e n o m i n a t o r . LowPerformanceCreatorsRemovedFromDenominator. L o w P er f or man ce C r e a t or s R e m o v e d F r o m D e n o mina t or .
因此:
A v e r a g e P e r f o r m a n c e ↑ ⇏ C r e a t o r E c o s y s t e m I m p r o v e d . \boxed{
AveragePerformance\uparrow
\not\Rightarrow
CreatorEcosystemImproved.
} A v er a g e P er f or man ce ↑ ⇒ C r e a t or E cosy s t e m I m p r o v e d .
36. Metric Decomposition
任何 metric uplift 應拆解來源。
例如:
Δ W a t c h T i m e = Δ W T b e t t e r m a t c h + Δ W T a u t o p l a y + Δ W T c o n t e n t l e n g t h + Δ W T r e p e a t + Δ W T f r i c t i o n . \Delta WatchTime
=
\Delta WT_{bettermatch}
+
\Delta WT_{autoplay}
+
\Delta WT_{contentlength}
+
\Delta WT_{repeat}
+
\Delta WT_{friction}. Δ W a t c h T im e = Δ W T b e tt er ma t c h + Δ W T a u t o pl a y + Δ W T co n t e n tl e n g t h + Δ W T r e p e a t + Δ W T f r i c t i o n .
如果只看 total:
Δ W a t c h T i m e > 0 , \Delta WatchTime>0, Δ W a t c h T im e > 0 ,
會失去 causal interpretation。
37. Product Success Criterion
本文提出最低 product success 條件:
policy π 1 \pi_1 π 1 相對 π 0 \pi_0 π 0 ,若要稱為 product improvement,至少需要:
Δ U M S ≥ 0 , \Delta U^{MS}\geq0, Δ U M S ≥ 0 ,
以及:
H a r d C o n s t r a i n t s S a t i s f i e d = 1 , HardConstraintsSatisfied=1, H a r d C o n s t r ain t s S a t i s f i e d = 1 ,
並且:
D E C ≤ θ . DEC\leq\theta. D E C ≤ θ .
不要求所有 stakeholder:
Δ U s > 0. \Delta U_s>0. Δ U s > 0.
但 trade-off 必須被明確知道。
38. Platform Health Dashboard
合理 dashboard 至少包含六層。
Layer 1 — Engagement
C T R , W a t c h T i m e , R e t u r n R a t e . CTR,
WatchTime,
ReturnRate. C T R , W a t c h T im e , R e t u r n R a t e .
Layer 2 — Satisfaction
S u r v e y , N o t I n t e r e s t e d , C o m p l a i n t , V o l u n t a r y R e t u r n . Survey,
NotInterested,
Complaint,
VoluntaryReturn. S u r v ey , N o t I n t er es t e d , C o m pl ain t , V o l u n t a r y R e t u r n .
Layer 3 — Agency
E x p l i c i t P r e f e r e n c e V i o l a t i o n , S e a r c h C o s t , F o l l o w i n g V i s i b i l i t y . ExplicitPreferenceViolation,
SearchCost,
FollowingVisibility. E x pl i c i tP r e f er e n ce V i o l a t i o n , S e a r c h C os t , F o l l o w in g V i s ibi l i t y .
Layer 4 — Exposure
G i n i , H H I , L o n g T a i l O b s e r v a b i l i t y , C o l d S t a r t O b s e r v a b i l i t y . Gini,
HHI,
LongTailObservability,
ColdStartObservability. G ini , H H I , L o n g T ai l O b ser v abi l i t y , C o l d S t a r tO b ser v abi l i t y .
Layer 5 — Creator Ecosystem
C S T R , C V R , E x i t R a t e , T T S A , S u p p l y H H I . CSTR,
CVR,
ExitRate,
TTSA,
SupplyHHI. C S T R , C V R , E x i tR a t e , T T S A , S u ppl y H H I .
Layer 6 — Commercial Quality
Q u a l i f i e d A t t e n t i o n , P a i d R O I , A d v e r t i s e r R e n e w a l , I n c r e m e n t a l i t y . QualifiedAttention,
PaidROI,
AdvertiserRenewal,
Incrementality. Q u a l i f i e d A tt e n t i o n , P ai d R O I , A d v er t i ser R e n e w a l , I n cr e m e n t a l i t y .
39. Product Health Gap Monitoring
每期計算:
P H G t = S c o r e ( M t ) − S c o r e ( H t ) . PHG_t
=
Score(M_t)
-
Score(H_t). P H G t = S cor e ( M t ) − S cor e ( H t ) .
若:
M t ↑ M_t\uparrow M t ↑
但:
H t ↓ , H_t\downarrow, H t ↓ ,
則:
P H G t ↑ . PHG_t\uparrow. P H G t ↑ .
應觸發:
M e t r i c D i v e r g e n c e R e v i e w . \boxed{
MetricDivergenceReview.
} M e t r i cD i v er g e n ce R e v i e w .
40. Metric Governance
metric governance 不應只由 model team 決定。
需要:
recommender engineers;
product;
creator ecosystem;
advertising;
trust and safety;
research;
user research。
因為:
M e t r i c C h o i c e MetricChoice M e t r i c C h o i ce
本身就是:
P l a t f o r m P o l i c y . PlatformPolicy. P l a t f or m P o l i cy .
41. Recommendation Governance as Constitutional Layer
可將推薦架構分為:
Layer 1 — Prediction
預測:
P ( c l i c k ) , P ( w a t c h ) , P ( s a v e ) . P(click),
P(watch),
P(save). P ( c l i c k ) , P ( w a t c h ) , P ( s a v e ) .
Layer 2 — Ranking
生成:
S c o r e . Score. S cor e .
Layer 3 — Observation Policy
決定:
H o m e , E x p l o r e , F o l l o w Home,
Explore,
Follow H o m e , E x pl or e , F o l l o w
如何分配 visibility。
Layer 4 — Governance
決定:
W h i c h M e t r i c s M a t t e r , W h i c h C o n s t r a i n t s C a n n o t B e V i o l a t e d , W h i c h S t a k e h o l d e r s C o u n t . WhichMetricsMatter,
WhichConstraintsCannotBeViolated,
WhichStakeholdersCount. W hi c h M e t r i cs M a tt er , W hi c h C o n s t r ain t s C ann o tB e V i o l a t e d , W hi c h S t ak e h o l d er s C o u n t .
A06 的核心是:
R a n k i n g C a n n o t D e f i n e I t s O w n S u c c e s s C r i t e r i a . \boxed{
RankingCannotDefineItsOwnSuccessCriteria.
} R ank in g C ann o t D e f in e I t s O w n S u ccess C r i t er ia .
42. A Unified Series-A Dynamic Model
綜合 A01–A06:
C t → O b s e r v a t i o n P o l i c y E x p o s u r e t \mathcal{C}_t
\xrightarrow{
ObservationPolicy
}
Exposure_t C t O b ser v a t i o n P o l i cy E x p os u r e t
⇓ \Downarrow ⇓
T e l e m e t r y t → P r e f e r e n c e I n f e r e n c e U s e r M o d e l t + 1 Telemetry_t
\xrightarrow{
PreferenceInference
}
UserModel_{t+1} T e l e m e t r y t P r e f er e n ce I n f er e n ce U ser M o d e l t + 1
⇓ \Downarrow ⇓
E x p o s u r e D i s t r i b u t i o n t + 1 ExposureDistribution_{t+1} E x p os u r eD i s t r ib u t i o n t + 1
⇓ \Downarrow ⇓
C r e a t o r R e t u r n t + 1 CreatorReturn_{t+1} C r e a t or R e t u r n t + 1
⇓ \Downarrow ⇓
C r e a t o r S u p p l y t + 2 CreatorSupply_{t+2} C r e a t or S u ppl y t + 2
⇓ \Downarrow ⇓
P l a t f o r m M e t r i c s t + 2 . PlatformMetrics_{t+2}. P l a t f or m M e t r i c s t + 2 .
然後組織使用:
P l a t f o r m M e t r i c s PlatformMetrics P l a t f or m M e t r i cs
決定下一輪:
O b s e r v a t i o n P o l i c y . ObservationPolicy. O b ser v a t i o n P o l i cy .
因此形成:
P o l i c y → O b s e r v a t i o n → B e h a v i o r → S u p p l y → M e t r i c s → P o l i c y . \boxed{
Policy
\rightarrow
Observation
\rightarrow
Behavior
\rightarrow
Supply
\rightarrow
Metrics
\rightarrow
Policy.
} P o l i cy → O b ser v a t i o n → B e ha v i or → S u ppl y → M e t r i cs → P o l i cy .
這是一個完整 platform control loop。
43. The Metric-Success Trap
Metric-Success Trap:
更高 popularity prior;
高流量內容預測更穩定;
CTR / watch time 上升;
dashboard 判定成功;
policy 繼續強化;
long-tail / creator viability 下降;
ecosystem cost 延遲出現;
到真正 business metric 下降時,供給結構已改變。
形式:
S h o r t T e r m M e t r i c G a i n → P o l i c y L o c k I n → D e l a y e d S t r u c t u r a l L o s s . \boxed{
ShortTermMetricGain
\rightarrow
PolicyLockIn
\rightarrow
DelayedStructuralLoss.
} S h or tT er m M e t r i c G ain → P o l i cy L oc k I n → D e l a y e d S t r u c t u r a l L oss .
44. Reversibility
平台 policy 應考慮:
R e v e r s i b i l i t y . Reversibility. R e v er s ibi l i t y .
UI weighting 可以很快修改。
但 creator exit:
E x i t k Exit_k E x i t k
可能不可逆。
community relation loss 也可能:
R e c o v e r y C o s t ≫ P o l i c y C h a n g e C o s t . RecoveryCost\gg PolicyChangeCost. R eco v er y C os t ≫ P o l i cy C han g e C os t .
因此:
I r r e v e r s i b l e E c o s y s t e m E f f e c t s \boxed{
IrreversibleEcosystemEffects
} I r r e v er s ib l e E cosy s t e m E f f ec t s
需要更高 experiment threshold。
45. Precaution for High-Leverage Metrics
若某 metric:
M M M
同時影響:
ranking;
employee goals;
creator strategy;
ad pricing;
product experiments;
則其:
S y s t e m L e v e r a g e SystemLeverage S y s t e m L e v er a g e
非常高。
高 leverage metric 應要求:
H i g h e r V a l i d a t i o n . HigherValidation. H i g h er V a l i d a t i o n .
46. Practical Experiment Acceptance Rule
本文提出:
接受 treatment π 1 \pi_1 π 1 若:
P r i m a r y M e t r i c U p l i f t > 0 PrimaryMetricUplift>0 P r ima r y M e t r i c U pl i f t > 0
且:
G u a r d r a i l V i o l a t i o n s = 0 GuardrailViolations=0 G u a r d r ai l V i o l a t i o n s = 0
且:
S h o r t T e r m S a t i s f a c t i o n ≥ b a s e l i n e ShortTermSatisfaction\geq baseline S h or tT er m S a t i s f a c t i o n ≥ ba se l in e
且:
E x p o s u r e H e a l t h ≥ t h r e s h o l d . ExposureHealth\geq threshold. E x p os u r eH e a l t h ≥ t h r es h o l d .
之後:
L o n g T e r m H o l d o u t LongTermHoldout L o n g T er m H o l d o u t
驗證:
D E C . DEC. D E C .
47. Case-Study Boundary
本 Series 最初由影音平台推薦體驗問題觸發,但本文不主張任何特定平台已發生全部 failure chain。
例如觀察到:
H i g h P o p u l a r i t y C o n t e n t S h a r e ↑ HighPopularityContentShare\uparrow H i g h P o p u l a r i t y C o n t e n tS ha r e ↑
只足以建立:
H y p o t h e s i s . Hypothesis. H y p o t h es i s .
若要推出:
C r e a t o r E x i t CreatorExit C r e a t or E x i t
或:
A d v e r t i s e r V a l u e D e c l i n e , AdvertiserValueDecline, A d v er t i ser V a l u eD ec l in e ,
需要後續 longitudinal evidence。
因此:
F r a m e w o r k ≠ V e r d i c t . \boxed{
Framework
\neq
Verdict.
} F r am e w or k = V er d i c t .
Series A 的價值是建立:
W h a t T o M e a s u r e , W h a t T o D i s t i n g u i s h , W h a t C a u s a l C h a i n s T o T e s t . WhatToMeasure,
WhatToDistinguish,
WhatCausalChainsToTest. W ha tT o M e a s u r e , W ha tT oD i s t in g u i s h , W ha tC a u s a l C hain s T o T es t .
48. Design Recommendations
本文建議平台:
將 ranking metrics 與 platform health metrics 分層;
不把 watch history 視為同質 preference evidence;
保留 explicit preference semantics;
分離 Popular 與 Explore;
保留 long-tail / cold-start testability;
監測 creator cohort survival;
評估 paid promotion 是否轉化為 organic viability;
對 autoplay 與 passive exposure 做 provenance;
使用 multi-stakeholder utility;
建立 long-term holdout 與 delayed ecosystem audit;
對高不可逆性 policy 使用更高 deployment threshold;
允許 users 選擇 observation policy。
49. Limitations
第一,多 stakeholder utility 的權重:
w s w_s w s
具有治理與價值判斷,不存在唯一客觀最優解。
第二,部分 hidden-state metrics,例如 trust、agency 與 information utility,測量成本高且可能有 survey bias。
第三,長期 holdout 會降低短期 experiment velocity,平台需要在學習速度與長期識別能力之間取捨。
第四,Goodhart effects 並不表示所有 proxy 都無效。工程系統仍必須使用 metrics;問題是需要理解其邊界與資料生成機制。
第五,Series A 建立的是 general framework,不提供對特定平台演算法內部權重的證明。
50. Conclusion
Series A 從一個看似簡單的問題開始:
為什麼推薦頁一直給我不想看的東西?
最後得到的不是一個單純 accuracy 問題,而是一個完整的平台控制系統。
第一層:
R e c o m m e n d a t i o n = O b s e r v a t i o n P o l i c y . \boxed{
Recommendation
=
ObservationPolicy.
} R eco mm e n d a t i o n = O b ser v a t i o n P o l i cy .
第二層:
A w a r e n e s s ≠ I n t e r e s t ≠ I n t e n t . \boxed{
Awareness
\neq
Interest
\neq
Intent.
} A w a r e n ess = I n t er es t = I n t e n t .
第三層:
P l a t f o r m I n d u c e d B e h a v i o r ≠ U s e r P r e f e r e n c e E v i d e n c e . \boxed{
PlatformInducedBehavior
\neq
UserPreferenceEvidence.
} P l a t f or m I n d u ce d B e ha v i or = U ser P r e f er e n ce E v i d e n ce .
第四層:
T o p i c D i v e r s i t y ≠ E x p o s u r e D i v e r s i t y . \boxed{
TopicDiversity
\neq
ExposureDiversity.
} T o p i cD i v er s i t y = E x p os u r eD i v er s i t y .
第五層:
N o E x p o s u r e ≠ L o w Q u a l i t y . \boxed{
NoExposure
\neq
LowQuality.
} N o E x p os u r e = L o w Q u a l i t y .
第六層:
M e t r i c S u c c e s s ≠ P r o d u c t S u c c e s s . \boxed{
MetricSuccess
\neq
ProductSuccess.
} M e t r i c S u ccess = P r o d u c tS u ccess .
完整 Series A dynamics:
O b s e r v a t i o n → P r e f e r e n c e I n t e r p r e t a t i o n → E n d o g e n o u s C o n t a m i n a t i o n → E x p o s u r e C o n c e n t r a t i o n → C r e a t o r S u p p l y R e s p o n s e → M e t r i c G o v e r n a n c e . \boxed{
Observation
\rightarrow
PreferenceInterpretation
\rightarrow
EndogenousContamination
\rightarrow
ExposureConcentration
\rightarrow
CreatorSupplyResponse
\rightarrow
MetricGovernance.
} O b ser v a t i o n → P r e f er e n ce I n t er p r e t a t i o n → E n d o g e n o u s C o n t amina t i o n → E x p os u r e C o n ce n t r a t i o n → C r e a t or S u ppl y R es p o n se → M e t r i c G o v er nan ce .
最終平台應優化的不是:
max C T R \max CTR max C T R
也不是:
max W a t c h T i m e . \max WatchTime. max W a t c h T im e .
而是:
max L o n g T e r m M u l t i S t a k e h o l d e r U t i l i t y \boxed{
\max
LongTermMultiStakeholderUtility
} max L o n g T er m M u l t i S t ak e h o l d er U t i l i t y
subject to:
U s e r A g e n c y , S a f e t y , C r e a t o r T e s t a b i l i t y , I n f o r m a t i o n Q u a l i t y , C o m m e r c i a l I n t e g r i t y . \boxed{
UserAgency,
Safety,
CreatorTestability,
InformationQuality,
CommercialIntegrity.
} U ser A g e n cy , S a f e t y , C r e a t or T es t abi l i t y , I n f or ma t i o n Q u a l i t y , C o mm er c ia l I n t e g r i t y .
推薦系統不是一個單純替內容排序的函數。
它是一個:
D y n a m i c V i s i b i l i t y A l l o c a t i o n S y s t e m . \boxed{
DynamicVisibilityAllocationSystem.
} D y nami c V i s ibi l i t y A l l oc a t i o n S y s t e m .
而 visibility allocation 會改變:
B e h a v i o r , P r e f e r e n c e D a t a , P o p u l a r i t y , C r e a t o r S t r a t e g y , C o n t e n t S u p p l y , R e v e n u e , P l a t f o r m F u t u r e . Behavior,
PreferenceData,
Popularity,
CreatorStrategy,
ContentSupply,
Revenue,
PlatformFuture. B e ha v i or , P r e f er e n ceD a t a , P o p u l a r i t y , C r e a t or S t r a t e g y , C o n t e n tS u ppl y , R e v e n u e , P l a t f or m F u t u r e .
因此,推薦平台真正的工程挑戰不是:
如何讓模型更懂得讓使用者點擊?
而是:
H o w C a n A P l a t f o r m A l l o c a t e O b s e r v a t i o n W i t h o u t D e s t r o y i n g T h e E c o l o g y T h a t M a k e s O b s e r v a t i o n V a l u a b l e ? \boxed{
HowCanAPlatformAllocateObservationWithoutDestroyingTheEcologyThatMakesObservationValuable?
} H o w C an A P l a t f or m A l l oc a t e O b ser v a t i o nW i t h o u t D es t r oy in g T h e E co l o g y T ha tM ak es O b ser v a t i o nV a l u ab l e ?
這是 Series A 的最終命題。
References
[0] Neo.K / EveMissLab. “Recommendation as an Observation Operator.” Series A, Paper A01, v0.1, 2026.
[1] Neo.K / EveMissLab. “Recommendation, Cold Start, and Creator Ecological Collapse.” Series A, Paper A05, v0.1, 2026.
[2] McNee, S. M., Riedl, J., & Konstan, J. A. “Being Accurate is Not Enough: How Accuracy Metrics Have Hurt Recommender Systems.” CHI Extended Abstracts, 2006.
[3] Knijnenburg, B. P., Willemsen, M. C., Gantner, Z., Soncu, H., & Newell, C. “Explaining the User Experience of Recommender Systems.” User Modeling and User-Adapted Interaction, 22, 441–504, 2012. DOI: 10.1007/s11257-011-9118-4.
[4] Burke, R., Abdollahpouri, H., Mobasher, B., & Gupta, T. “Towards Multi-Stakeholder Utility Evaluation of Recommender Systems.” UMAP Extended Proceedings, 2016.
[5] Abdollahpouri, H., Burke, R., & Mobasher, B. “Recommender Systems as Multistakeholder Environments.” UMAP, 2017.
[6] Patro, G. K., Biswas, A., Ganguly, N., Gummadi, K. P., & Chakraborty, A. “FairRec: Two-Sided Fairness for Personalized Recommendations in Two-Sided Platforms.” Proceedings of The Web Conference 2020, 2020. arXiv:2002.10764.
[7] Wu, Y., Cao, J., Xu, G., & Tan, Y. “TFROM: A Two-sided Fairness-Aware Recommendation Model for Both Customers and Providers.” SIGIR 2021, pp. 1013–1022. DOI: 10.1145/3404835.3462882.
[8] Chaney, A. J. B., Stewart, B. M., & Engelhardt, B. E. “How Algorithmic Confounding in Recommendation Systems Increases Homogeneity and Decreases Utility.” RecSys 2018, pp. 224–232. DOI: 10.1145/3240323.3240370.
[9] Mansoury, M., Abdollahpouri, H., Pechenizkiy, M., Mobasher, B., & Burke, R. “Feedback Loop and Bias Amplification in Recommender Systems.” CIKM 2020, pp. 2145–2148. DOI: 10.1145/3340531.3412152.
[10] Manheim, D., & Garrabrant, S. “Categorizing Variants of Goodhart's Law.” arXiv:1803.04585, 2018.
[11] Jannach, D., & Jugovac, M. “Measuring the Business Value of Recommender Systems.” ACM Transactions on Management Information Systems, 10(4), 2019.
[12] Adomavicius, G., Bockstedt, J. C., Curley, S. P., & Zhang, J. “Recommender Systems, Ground Truth, and Preference Pollution.” AI Magazine, 43(2), 177–189, 2022. DOI: 10.1002/aaai.12055.
[13] Sun, Y., & Sun, B. “From Exposure to Followers: A Stock-and-Flow Closed-Loop Framework of Creator Dynamics.” Information Processing & Management, 63(5), 104677, 2026. DOI: 10.1016/j.ipm.2026.104677.
[14] Zhao, W., Feng, H., Feng, N., & Li, M. “Does Paying for Visibility Pay Off? The Impact of Sponsored Recommendations on UGC Platforms.” Decision Support Systems, 206, 114668, 2026. DOI: 10.1016/j.dss.2026.114668.
Series A Closure
Series A 正篇完成:
A01 — Recommendation as an Observation Operator
A02 — Explicit Preference versus Inferred Preference
A03 — Passive Exposure and Endogenous Preference Contamination
A04 — The Platform-Induced Exposure Bubble
A05 — Recommendation, Cold Start, and Creator Ecological Collapse
A06 — Metric Success, Product Failure
後續工程化文件:
TW-A — User-Controllable Recommendation & Observation Architecture