Series A — Algorithmic Observation, Recommendation & Platform Ecology
Paper A03 — Passive Exposure and Endogenous Preference Contamination
被動曝光與內生偏好污染:自動播放、遙測來源與推薦回饋迴圈的因果框架
English Title: Passive Exposure and Endogenous Preference Contamination: A Causal Framework for Autoplay, Telemetry Provenance, and Recommender Feedback LoopsSeries: Algorithmic Observation, Recommendation & Platform EcologyPaper ID: A03Version: v0.1Date: 2026-08-31Status: Canonical UTF-8 SourceAuthor: Neo.K / EveMissLab
Abstract
推薦系統依賴使用者行為資料估計偏好,但行為資料本身常由推薦系統先前的曝光決策所產生。當平台主動決定哪些內容出現在首頁、資訊流或動態頁,並進一步透過 autoplay、預載播放、停留計時或歷史紀錄將被動曝光轉為「觀看」事件時,後續模型可能將平台自身造成的行為誤解為使用者自主偏好。這會形成一種內生性問題:推薦系統先製造資料,再以這些資料作為自己原先判斷正確的證據。
本文延續 Paper A01 的 Recommendation-as-Observation-Operator framework 與 Paper A02 的 Multidimensional Preference-State Model,提出 Endogenous Preference Contamination framework。本文將 impression、passive autoplay、attentional exposure、active play、intentional search、save、follow 等事件分離,並以 structural causal model 描述平台政策、曝光、被動播放、使用者潛在狀態、遙測紀錄與下一輪推薦之間的因果關係。
本文提出 Recommendation-Induced Preference Artifact(RIPA),指推薦系統造成的曝光或被動互動被模型錯誤吸收為內在偏好證據;並進一步定義 Observation-Generated Evidence、Endogenous Evidence Ratio、Passive Exposure Contamination Index、Self-Confirmation Ratio 與 Re-exposure Amplification Factor。本文主張,watch history 不應被視為同質事件集合,而應保存完整 event provenance,包括 surface、rank、trigger、autoplay、visibility、audio state、duration、user initiation、seek behavior 與 downstream action。
本文並提出一組修正原則:將曝光生成機制與偏好推論分離;對低 intentionality 事件使用低權重或零偏好更新;將顯式使用者行為與平台誘發行為區分;使用隨機探索、propensity correction、causal adjustment 或受控 holdout 估計 counterfactual preference;以及禁止將 autoplay-generated history 直接等同 active consumption。此框架提供後續研究平台誘導曝光繭房、冷啟動失敗、創作者生態集中與 KPI 自我證成的因果基礎。
Keywords: recommender systems; feedback loop; passive exposure; autoplay; endogeneity; causal inference; exposure bias; preference pollution; telemetry provenance; implicit feedback
1. Introduction
推薦系統通常被描述為:
H i s t o r y → P r e f e r e n c e E s t i m a t e → R e c o m m e n d a t i o n . History
\rightarrow
PreferenceEstimate
\rightarrow
Recommendation. H i s t or y → P r e f er e n ce E s t ima t e → R eco mm e n d a t i o n .
但真實平台往往更接近:
R e c o m m e n d a t i o n t → E x p o s u r e t → B e h a v i o r t → H i s t o r y t + 1 → R e c o m m e n d a t i o n t + 1 . Recommendation_t
\rightarrow
Exposure_t
\rightarrow
Behavior_t
\rightarrow
History_{t+1}
\rightarrow
Recommendation_{t+1}. R eco mm e n d a t i o n t → E x p os u r e t → B e ha v i o r t → H i s t or y t + 1 → R eco mm e n d a t i o n t + 1 .
這個差異非常重要。
如果 B e h a v i o r t Behavior_t B e ha v i o r t 是使用者在自由可觀察環境下自主產生的行為,歷史資料可以提供偏好證據。但若 B e h a v i o r t Behavior_t B e ha v i o r t 的一部分由平台本身的介面、排序、預載、自動播放與曝光策略生成,那麼:
B e h a v i o r t = f ( U s e r S t a t e t , P l a t f o r m P o l i c y t , E x p o s u r e t , I n t e r f a c e t ) . Behavior_t
=
f(
UserState_t,
PlatformPolicy_t,
Exposure_t,
Interface_t
). B e ha v i o r t = f ( U ser S t a t e t , P l a t f or m P o l i c y t , E x p os u r e t , I n t er f a c e t ) .
此時將所有行為直接寫入 preference history,等於把:
P l a t f o r m P o l i c y t PlatformPolicy_t P l a t f or m P o l i c y t
的效果混入:
U s e r P r e f e r e n c e t . UserPreference_t. U ser P r e f er e n c e t .
這是一種內生性。
推薦系統不再只是觀察使用者,而是在部分製造自己之後要觀察的資料。
Chaney、Stewart 與 Engelhardt 將這類現象描述為 algorithmic confounding:部署中的推薦系統影響使用者後續行為,而系統再利用這些被自身影響的資料訓練下一代模型,可能造成行為同質化與效用下降 [3]。Mansoury 等人的研究則展示 feedback loop 會放大 popularity bias、降低 aggregate diversity 並改變使用者 taste representation [4]。Adomavicius 等人進一步提出 preference pollution,指出推薦會改變後續被視為 ground truth 的偏好資料 [6]。
本文聚焦其中一個更細緻但在現代影音與資訊流平台極重要的問題:
P a s s i v e E x p o s u r e → R e c o r d e d B e h a v i o r → I n f e r r e d P r e f e r e n c e \boxed{
Passive\ Exposure
\rightarrow
Recorded\ Behavior
\rightarrow
Inferred\ Preference
} P a ss i v e E x p os u r e → R ecor d e d B e ha v i or → I n f er r e d P r e f er e n ce
特別是當 autoplay 或類 autoplay 行為被寫入觀看歷史時,平台可能形成:
The system creates evidence and then learns from its own evidence. \boxed{
\text{The system creates evidence and then learns from its own evidence.}
} The system creates evidence and then learns from its own evidence.
本文將此稱為 Endogenous Preference Contamination。
2. Relation to Papers A01 and A02
Paper A01 建立:
C t → O u , t ( s ) L u , t ( s ) → finite attention W u , t e x p . \mathcal{C}_t
\xrightarrow{
\mathcal{O}_{u,t}^{(s)}
}
L_{u,t}^{(s)}
\xrightarrow{
\text{finite attention}
}
\mathcal{W}_{u,t}^{exp}. C t O u , t ( s ) L u , t ( s ) finite attention W u , t e x p .
其中 O u , t ( s ) \mathcal{O}_{u,t}^{(s)} O u , t ( s ) 是平台的 observation operator。
Paper A02 再建立:
L u , t → I n t e r a c t i o n E u , t → P z u , v , t , L_{u,t}
\xrightarrow{
Interaction
}
\mathcal{E}_{u,t}
\xrightarrow{
\mathcal{P}
}
\mathbf{z}_{u,v,t}, L u , t I n t er a c t i o n E u , t P z u , v , t ,
其中 P \mathcal{P} P 是 preference-state inference operator,而:
z u , v , t = ( A w a r , I n t r , I n t e n t , F e a s , R e l , H o r i z o n , D e c l ) . \mathbf{z}_{u,v,t}
=
(
Awar,
Intr,
Intent,
Feas,
Rel,
Horizon,
Decl
). z u , v , t = ( A w a r , I n t r , I n t e n t , F e a s , R e l , H or i z o n , D ec l ) .
A03 指出一個新的問題:
E u , t \mathcal{E}_{u,t} E u , t
並不是天然外生的。
實際上:
E u , t = h ( O u , t , U I t , U s e r S t a t e t ) . \mathcal{E}_{u,t}
=
h(
\mathcal{O}_{u,t},
UI_t,
UserState_t
). E u , t = h ( O u , t , U I t , U ser S t a t e t ) .
因此:
P \mathcal{P} P
所收到的證據,本身已經受到:
O u , t \mathcal{O}_{u,t} O u , t
影響。
完整閉環為:
O t → E x p o s u r e t → T e l e m e t r y t → P t → z t + 1 → O t + 1 \boxed{
\mathcal{O}_t
\rightarrow
Exposure_t
\rightarrow
Telemetry_t
\rightarrow
\mathcal{P}_t
\rightarrow
\mathbf{z}_{t+1}
\rightarrow
\mathcal{O}_{t+1}
} O t → E x p os u r e t → T e l e m e t r y t → P t → z t + 1 → O t + 1
這就是本篇研究的核心。
3. Related Work
3.1 Algorithmic confounding
Chaney 等人指出,推薦系統常以已經受到先前推薦政策影響的資料進行訓練與評估,造成 feedback loop。其模擬結果顯示,這種 algorithmic confounding 可使使用者行為逐漸同質化而沒有帶來相應效用提升 [3]。
此研究提供本文的重要基礎:
O b s e r v e d B e h a v i o r ≠ B e h a v i o r W i t h o u t R e c o m m e n d a t i o n P o l i c y . ObservedBehavior
\neq
BehaviorWithoutRecommendationPolicy. O b ser v e d B e ha v i or = B e ha v i or W i t h o u tR eco mm e n d a t i o n P o l i cy .
3.2 Feedback-loop bias amplification
Mansoury 等人透過模擬推薦器與使用者反覆互動,展示 feedback loop 可以放大 popularity bias,降低 aggregate diversity,並使使用者的推薦經驗逐漸同質化 [4]。
因此:
B i a s t → E x p o s u r e t → I n t e r a c t i o n t → B i a s t + 1 Bias_t
\rightarrow
Exposure_t
\rightarrow
Interaction_t
\rightarrow
Bias_{t+1} B ia s t → E x p os u r e t → I n t er a c t i o n t → B ia s t + 1
可以是一個正回饋系統。
3.3 User feedback-loop bias
Pan 等人提出 user feedback-loop bias,並以 temporal exposure probability 與 inverse propensity scoring 修正偏差 [5]。此結果說明 exposure probability 不能被忽略,因為使用者歷史並不是在均勻 exposure condition 下自然生成。
3.4 Preference pollution
Adomavicius 等人將推薦前後的使用者回饋視為連續 feedback loop,指出 recommender systems 會影響後續被視為 ground truth 的 preference data,造成 non-representativeness 與 preference pollution [6]。
本文將其中一部分進一步細分:
Preference Pollution ⊃ Passive-Exposure-Induced Contamination . \text{Preference Pollution}
\supset
\text{Passive-Exposure-Induced Contamination}. Preference Pollution ⊃ Passive-Exposure-Induced Contamination .
3.5 Causal correction
Krauth、Wang 與 Jordan 提出 Causal Adjustment for Feedback Loops,主張若模型推論 intervention distributions,而非僅學習 observational distributions,就能在理論上避免部分 feedback-loop 問題 [7]。
更近期研究也使用 causal graph 同時建模 item exposure 與 user satisfaction,指出使用者行為可能受到廣告、推廣與曝光機制影響,而不完全反映真實 user interest [8]。
這些工作共同支持本文的核心方向:
E x p o s u r e m e c h a n i s m must be modeled separately from P r e f e r e n c e . \boxed{ Exposure\ mechanism \text{ must be modeled separately from } Preference. } E x p os u r e m ec hani s m must be modeled separately from P r e f er e n ce .
4. Event Taxonomy
本文先將常被平台統稱為「觀看」的事件拆開。
令事件類型集合:
T = { I M P , P A , A E , A P , S C , S E E K , L I K E , S A V E , F O L L O W } . \mathcal{T}
=
\{
IMP,
PA,
AE,
AP,
SC,
SEEK,
LIKE,
SAVE,
FOLLOW
\}. T = { I M P , P A , A E , A P , S C , S E E K , L I K E , S A V E , F O LL O W } .
其中:
I M P = Impression , IMP=\text{Impression}, I M P = Impression ,
表示內容出現在可見 surface 中。
P A = Passive Autoplay , PA=\text{Passive Autoplay}, P A = Passive Autoplay ,
表示沒有明確 user-initiation 的自動播放。
A E = Attentional Exposure , AE=\text{Attentional Exposure}, A E = Attentional Exposure ,
表示內容在視窗中停留足夠時間,但不代表主動觀看。
A P = Active Play , AP=\text{Active Play}, A P = Active Play ,
表示使用者主動啟動播放。
S C = Search-initiated Consumption , SC=\text{Search-initiated Consumption}, S C = Search-initiated Consumption ,
表示使用者先主動搜尋再進入內容。
S E E K = Seek / Scrub , SEEK=\text{Seek / Scrub}, S E E K = Seek / Scrub ,
表示使用者主動改變播放位置。
L I K E , S A V E , F O L L O W LIKE,\ SAVE,\ FOLLOW L I K E , S A V E , F O LL O W
則代表更高語意密度的明示或準明示行為。
這些事件不能簡化為:
W a t c h = 1. Watch=1. W a t c h = 1.
5. Intentionality Weight
對每個事件 e e e 定義 intentionality:
I ( e ) ∈ [ 0 , 1 ] . I(e)\in[0,1]. I ( e ) ∈ [ 0 , 1 ] .
例如可以有:
I ( I M P ) ≈ 0 , I(IMP)\approx0, I ( I M P ) ≈ 0 ,
I ( P A ) ≈ 0 , I(PA)\approx0, I ( P A ) ≈ 0 ,
I ( A E ) > 0 , I(AE)>0, I ( A E ) > 0 ,
I ( A P ) > I ( A E ) , I(AP)>I(AE), I ( A P ) > I ( A E ) ,
I ( S C ) > I ( A P ) , I(SC)>I(AP), I ( S C ) > I ( A P ) ,
而:
I ( S A V E ) , I ( F O L L O W ) I(SAVE),I(FOLLOW) I ( S A V E ) , I ( F O LL O W )
通常具有更高的 user-declared intentionality。
本文不規定固定數值,因為不同平台與情境會不同;核心要求是:
I ( e ) must be preserved as semantics, not erased by event aggregation. \boxed{ I(e) \text{ must be preserved as semantics, not erased by event aggregation.} } I ( e ) must be preserved as semantics, not erased by event aggregation.
如果所有事件最後只存:
video_id
watch_seconds
timestamp
則大量重要因果資訊已在資料層消失。
6. Telemetry Provenance
本文定義 canonical interaction event:
e = ( u , v , s , r , g , a , v i s , a u d , d , s e e k , a c t , t ) . e
=
(
u,
v,
s,
r,
g,
a,
vis,
aud,
d,
seek,
act,
t
). e = ( u , v , s , r , g , a , v i s , a u d , d , see k , a c t , t ) .
其中:
u u u :user;
v v v :item;
s s s :surface;
r r r :rank position;
g g g :trigger;
a a a :autoplay flag;
v i s vis v i s :visibility state;
a u d aud a u d :audio state;
d d d :duration;
s e e k seek see k :seek activity;
a c t act a c t :downstream action;
t t t :timestamp。
trigger 至少應區分:
{ manual , search , follow , recommendation , autoplay , external , notification } . \{
\text{manual},
\text{search},
\text{follow},
\text{recommendation},
\text{autoplay},
\text{external},
\text{notification}
\}. { manual , search , follow , recommendation , autoplay , external , notification } .
因此:
H i s t o r y History H i s t or y
不應只是 item ID sequence,而應為:
H i s t o r y = S e q u e n c e o f P r o v e n a n c e d E v e n t s . \boxed{
History
=
Sequence\ of\ Provenanced\ Events.
} H i s t or y = S e q u e n ce o f P r o v e nan ce d E v e n t s .
7. Structural Causal Model
令:
Z t Z_t Z t
為使用者真實但不可完全觀測的 preference state。
令:
Π t \Pi_t Π t
為平台推薦政策。
令:
E t E_t E t
為 exposure。
令:
A t A_t A t
為 autoplay / passive interface action。
令:
B t B_t B t
為使用者可觀測行為。
令:
T t T_t T t
為 telemetry record。
令:
Z ^ t + 1 \hat{Z}_{t+1} Z ^ t + 1
為模型更新後的使用者狀態估計。
因果關係可寫為:
Π t → E t , \Pi_t
\rightarrow
E_t, Π t → E t ,
E t → A t , E_t
\rightarrow
A_t, E t → A t ,
( Z t , E t , A t ) → B t , (Z_t,E_t,A_t)
\rightarrow
B_t, ( Z t , E t , A t ) → B t ,
( E t , A t , B t ) → T t , (E_t,A_t,B_t)
\rightarrow
T_t, ( E t , A t , B t ) → T t ,
T t → Z ^ t + 1 , T_t
\rightarrow
\hat{Z}_{t+1}, T t → Z ^ t + 1 ,
以及:
Z ^ t + 1 → Π t + 1 . \hat{Z}_{t+1}
\rightarrow
\Pi_{t+1}. Z ^ t + 1 → Π t + 1 .
因此:
Π t → T t → Z ^ t + 1 → Π t + 1 \boxed{
\Pi_t
\rightarrow
T_t
\rightarrow
\hat{Z}_{t+1}
\rightarrow
\Pi_{t+1}
} Π t → T t → Z ^ t + 1 → Π t + 1
是一條內生閉環。
若模型忽略:
Π t , E t , A t , \Pi_t,
E_t,
A_t, Π t , E t , A t ,
而直接推論:
T t ⇒ Z t , T_t
\Rightarrow
Z_t, T t ⇒ Z t ,
就會把平台造成的 telemetry variation 錯誤歸因到使用者 preference。
8. Endogenous Preference Contamination
本文定義 Endogenous Preference Contamination:
若某 preference update:
Δ Z ^ u , v , t \Delta\hat{Z}_{u,v,t} Δ Z ^ u , v , t
主要由平台自身的 exposure 或 interface action 造成,而不是由使用者自主 intent 造成,且模型沒有保留或修正該因果來源,則此更新受到內生偏好污染。
形式化地,令:
T t o b s T_t^{obs} T t o b s
是實際 telemetry。
令:
T t d o ( Π = π 0 ) T_t^{do(\Pi=\pi_0)} T t d o ( Π = π 0 )
表示在基準 policy π 0 \pi_0 π 0 下的 counterfactual telemetry。
如果:
Δ T t = T t o b s − T t d o ( Π = π 0 ) \Delta T_t
=
T_t^{obs}
-
T_t^{do(\Pi=\pi_0)} Δ T t = T t o b s − T t d o ( Π = π 0 )
很大,而模型仍將全部:
T t o b s T_t^{obs} T t o b s
歸因為 preference evidence,則污染風險上升。
9. Recommendation-Induced Preference Artifact
本文提出:
R I P A = R e c o m m e n d a t i o n - I n d u c e d P r e f e r e n c e A r t i f a c t \boxed{ RIPA = Recommendation\text{-}Induced\ Preference\ Artifact } R I P A = R eco mm e n d a t i o n - I n d u ce d P r e f er e n ce A r t i f a c t
RIPA 指:
推薦系統先透過曝光、排序、自動播放或重複呈現提高某內容的可觀察性,再將因此產生的弱互動解讀為使用者原本就具有的偏好。
典型鏈條:
R e c o m m e n d a t i o n ( x ) Recommendation(x) R eco mm e n d a t i o n ( x )
⇓ \Downarrow ⇓
P a s s i v e E x p o s u r e ( x ) PassiveExposure(x) P a ss i v e E x p os u r e ( x )
⇓ \Downarrow ⇓
T e l e m e t r y W a t c h ( x ) TelemetryWatch(x) T e l e m e t r y W a t c h ( x )
⇓ \Downarrow ⇓
P r e f e r e n c e E s t i m a t e ( x ) ↑ PreferenceEstimate(x)\uparrow P r e f er e n ce E s t ima t e ( x ) ↑
⇓ \Downarrow ⇓
R e c o m m e n d a t i o n ( x ) ↑ Recommendation(x)\uparrow R eco mm e n d a t i o n ( x ) ↑
這形成:
E x p o s u r e → P s e u d o P r e f e r e n c e → M o r e E x p o s u r e \boxed{
Exposure
\rightarrow
PseudoPreference
\rightarrow
MoreExposure
} E x p os u r e → P se u d o P r e f er e n ce → M or e E x p os u r e
10. Observation-Generated Evidence
Paper A01 將推薦視為 observation operator。
因此本文定義:
O G E ( e ) = 1 OGE(e)=1 O GE ( e ) = 1
若事件 e e e 的存在高度依賴平台觀察策略本身。
例如:
首頁卡片 impression;
自動播放;
預覽片段;
自動連播;
高 rank placement;
notification-triggered open。
相反地:
O G E ( e ) ≈ 0 OGE(e)\approx0 O GE ( e ) ≈ 0
可能出現在:
使用者輸入精確 query;
使用者直接進入 creator page;
外部連結直接開啟;
bookmark / saved item 主動回訪。
這不是說 OGE 事件沒有價值,而是它們不能與 user-originated evidence 混為一談。
11. The Autoplay-History Problem
自動播放本身不是必然有害。
問題在於:
A u t o p l a y → H i s t o r y → P r e f e r e n c e U p d a t e . Autoplay
\rightarrow
History
\rightarrow
PreferenceUpdate. A u t o pl a y → H i s t or y → P r e f er e n ce U p d a t e .
如果:
P A ( v ) PA(v) P A ( v )
只要持續數秒就被寫為:
W a t c h H i s t o r y ( v ) = 1 , WatchHistory(v)=1, W a t c h H i s t or y ( v ) = 1 ,
而 downstream recommender 又把:
W a t c h H i s t o r y ( v ) WatchHistory(v) W a t c h H i s t or y ( v )
視為 positive implicit feedback,則:
P A PA P A
被語意轉換成:
I n t e r e s t . Interest. I n t er es t .
這個轉換沒有受到使用者明確行為支持。
因此本文提出:
Proposition 1 — Autoplay Non-Equivalence
A u t o p l a y E x p o s u r e ≠ A c t i v e C o n s u m p t i o n \boxed{
AutoplayExposure
\neq
ActiveConsumption
} A u t o pl a y E x p os u r e = A c t i v e C o n s u m pt i o n
Proposition 2 — History Non-Homogeneity
H i s t o r y ≠ H o m o g e n e o u s P r e f e r e n c e E v i d e n c e \boxed{
History
\neq
Homogeneous\ Preference\ Evidence
} H i s t or y = H o m o g e n eo u s P r e f er e n ce E v i d e n ce
歷史資料必須保留其生成機制。
12. Passive Exposure Contamination Index
令事件集合為:
E u , t . \mathcal{E}_{u,t}. E u , t .
令模型對事件 e e e 的 preference update contribution 為:
w p ( e ) . w_p(e). w p ( e ) .
定義 Passive Exposure Contamination Index:
P E C I = ∑ e ∈ E w p ( e ) ⋅ 1 [ I ( e ) < θ I ] ∑ e ∈ E ∣ w p ( e ) ∣ + ϵ . PECI
=
\frac{
\sum_{e\in\mathcal{E}}
w_p(e)
\cdot
\mathbf{1}[I(e)<\theta_I]
}{
\sum_{e\in\mathcal{E}}
|w_p(e)|
+\epsilon
}. P E C I = ∑ e ∈ E ∣ w p ( e ) ∣ + ϵ ∑ e ∈ E w p ( e ) ⋅ 1 [ I ( e ) < θ I ] .
其中:
θ I \theta_I θ I
為 intentionality threshold。
若:
P E C I → 1 , PECI\rightarrow1, P E C I → 1 ,
代表大量 preference update 來自低 intentionality 事件。
13. Endogenous Evidence Ratio
定義:
E E R = N o b s e r v a t i o n − g e n e r a t e d e v i d e n c e N a l l p r e f e r e n c e e v i d e n c e . EER
=
\frac{
N_{\mathrm{observation-generated\ evidence}}
}{
N_{\mathrm{all\ preference\ evidence}}
}. E E R = N all preference evidence N observation − generated evidence .
高 E E R EER E E R 不一定代表錯誤。
例如首頁推薦系統本來就會依賴大量 observation-generated interactions。
真正問題是:
E E R ↑ EER\uparrow E E R ↑
同時:
C a u s a l C o r r e c t i o n ≈ 0. CausalCorrection\approx0. C a u s a l C or r ec t i o n ≈ 0.
因此可以定義:
R i s k e n d o = E E R ⋅ ( 1 − C c ) , Risk_{endo}
=
EER
\cdot
(1-C_c), R i s k e n d o = E E R ⋅ ( 1 − C c ) ,
其中:
C c ∈ [ 0 , 1 ] C_c\in[0,1] C c ∈ [ 0 , 1 ]
表示 causal correction coverage。
14. Self-Confirmation Ratio
若系統在時間 t t t 推薦內容類別 x x x ,並因自身曝光造成 telemetry 增加,再於 t + 1 t+1 t + 1 將此作為推薦 x x x 的主要證據,就形成 self-confirmation。
定義:
S C R = P ( R e c o m m e n d t + 1 ( x ) ∣ E v i d e n c e t ( x ) was policy-induced ) . SCR
=
P(
Recommend_{t+1}(x)
\mid
Evidence_t(x)\ \text{was policy-induced}
). S C R = P ( R eco mm e n d t + 1 ( x ) ∣ E v i d e n c e t ( x ) was policy-induced ) .
實務上可以比較:
S C R o b s e r v e d SCR_{observed} S C R o b ser v e d
與 randomized holdout 中的:
S C R b a s e l i n e . SCR_{baseline}. S C R ba se l in e .
如果:
S C R o b s e r v e d ≫ S C R b a s e l i n e , SCR_{observed}\gg SCR_{baseline}, S C R o b ser v e d ≫ S C R ba se l in e ,
表示系統可能對自身誘發訊號過度學習。
15. Re-exposure Amplification Factor
令:
q t ( v ) = P ( v is exposed at t ) . q_t(v)
=
P(
v\text{ is exposed at }t
). q t ( v ) = P ( v is exposed at t ) .
若第一次弱曝光後:
q t + 1 ( v ) q_{t+1}(v) q t + 1 ( v )
因低 intentionality telemetry 而顯著提高,定義:
R A F ( v ) = q t + 1 ( v ) q t ( v ) + ϵ . RAF(v)
=
\frac{
q_{t+1}(v)
}{
q_t(v)+\epsilon
}. R A F ( v ) = q t ( v ) + ϵ q t + 1 ( v ) .
若使用者沒有 active engagement,但:
R A F ( v ) ≫ 1 , RAF(v)\gg1, R A F ( v ) ≫ 1 ,
系統可能把 passive exposure 誤當 positive signal。
16. Skip Is Not Always Negative, but Repeated Skip Matters
同樣地,本文也不主張:
Q u i c k S k i p = D i s l i k e . QuickSkip
=
Dislike. Q u i c k S k i p = D i s l ik e .
使用者可能因當時無時間、畫面位置或其他任務快速滑過。
但是若同一內容或 topic 反覆:
E x p o s u r e → S k i p Exposure
\rightarrow
Skip E x p os u r e → S k i p
且沒有 active re-entry,系統應逐步累積:
E v i d e n c e n o n i n t e n t . Evidence_{nonintent}. E v i d e n c e n o nin t e n t .
尤其:
R e p e a t e d S k i p + N o S e a r c h + N o S a v e + N o C r e a t o r V i s i t RepeatedSkip
+
NoSearch
+
NoSave
+
NoCreatorVisit R e p e a t e d S k i p + N o S e a r c h + N o S a v e + N o C r e a t or V i s i t
應降低:
P ( I n t e n t > 0 ) . P(Intent>0). P ( I n t e n t > 0 ) .
這與 A02 的 KBD / temporal declaration 相容。
17. Corrective Architecture
17.1 Separate exposure log and preference log
不要:
history = watched_items
而應至少分為:
exposure_ledger
interaction_ledger
preference_evidence_ledger
explicit_declaration_ledger
其中:
E x p o s u r e L e d g e r ≠ P r e f e r e n c e E v i d e n c e L e d g e r . ExposureLedger
\neq
PreferenceEvidenceLedger. E x p os u r e L e d g er = P r e f er e n ce E v i d e n ce L e d g er .
17.2 Provenance-aware aggregation
所有 aggregated feature 應保留:
s o u r c e _ t y p e . source\_type. so u r ce _ t y p e .
例如:
W a t c h S e c o n d s = W a t c h S e c o n d s m a n u a l + W a t c h S e c o n d s a u t o p l a y + W a t c h S e c o n d s s e a r c h + W a t c h S e c o n d s f o l l o w . WatchSeconds
=
WatchSeconds_{manual}
+
WatchSeconds_{autoplay}
+
WatchSeconds_{search}
+
WatchSeconds_{follow}. W a t c h S eco n d s = W a t c h S eco n d s man u a l + W a t c h S eco n d s a u t o pl a y + W a t c h S eco n d s se a r c h + W a t c h S eco n d s f o l l o w .
而不是只有:
W a t c h S e c o n d s t o t a l . WatchSeconds_{total}. W a t c h S eco n d s t o t a l .
17.3 Intentionality gates
可設定:
Δ p ( e ) = 0 \Delta p(e)=0 Δ p ( e ) = 0
若:
I ( e ) < θ 0 I(e)<\theta_0 I ( e ) < θ 0
且沒有 downstream active signal。
或者使用連續加權:
Δ p ( e ) = I ( e ) ⋅ g ( e ) . \Delta p(e)
=
I(e)
\cdot
g(e). Δ p ( e ) = I ( e ) ⋅ g ( e ) .
17.4 Explicit-declaration precedence
若 A02 的有效 declaration:
D e c l ( u , v , t ) Decl(u,v,t) D ec l ( u , v , t )
與 autoplay telemetry 衝突,則:
D e c l > P A Decl
>
PA D ec l > P A
在其他條件相近時應成立。
17.5 Exposure-aware learning
偏好模型應估計:
P ( I n t e r a c t i o n ∣ E x p o s u r e ) , P(
Interaction
\mid
Exposure
), P ( I n t er a c t i o n ∣ E x p os u r e ) ,
而不是將:
N o I n t e r a c t i o n NoInteraction N o I n t er a c t i o n
直接當成 uniform negative。
可以使用:
inverse propensity scoring;
causal adjustment;
randomized exploration;
interleaving;
controlled holdout;
doubly robust estimation。
17.6 Counterfactual preference estimation
真正想知道的是:
P ( A c t i v e E n g a g e m e n t ( v ) ∣ d o ( E x p o s u r e P o l i c y = π ) ) . P(
ActiveEngagement(v)
\mid
do(ExposurePolicy=\pi)
). P ( A c t i v e E n g a g e m e n t ( v ) ∣ d o ( E x p os u r e P o l i cy = π )) .
而不只是:
P ( A c t i v e E n g a g e m e n t ( v ) ∣ O b s e r v e d E x p o s u r e P o l i c y = π t ) . P(
ActiveEngagement(v)
\mid
ObservedExposurePolicy=\pi_t
). P ( A c t i v e E n g a g e m e n t ( v ) ∣ O b ser v e d E x p os u r e P o l i cy = π t ) .
18. Randomized Opening of the Loop
完全依賴 production recommender 的資料會讓:
P o l i c y Policy P o l i cy
與:
O b s e r v e d P r e f e r e n c e ObservedPreference O b ser v e d P r e f er e n ce
越來越糾纏。
因此需要少量受控 randomization。
例如保留:
ϵ \epsilon ϵ
比例的候選位置,從符合最低 relevance / safety constraint 的內容池中隨機抽取。
這可估計:
P ( E n g a g e m e n t ∣ R a n d o m E x p o s u r e ) P(
Engagement
\mid
RandomExposure
) P ( E n g a g e m e n t ∣ R an d o m E x p os u r e )
並與:
P ( E n g a g e m e n t ∣ P o l i c y E x p o s u r e ) P(
Engagement
\mid
PolicyExposure
) P ( E n g a g e m e n t ∣ P o l i cy E x p os u r e )
比較。
這不是要求平台大量亂推內容,而是為 causal calibration 保留最低必要的識別能力。
19. Preference-State Update with Provenance
承接 A02:
z u , v , t = ( A w a r , I n t r , I n t e n t , F e a s , R e l , H o r i z o n , D e c l ) . \mathbf{z}_{u,v,t}
=
(
Awar,
Intr,
Intent,
Feas,
Rel,
Horizon,
Decl
). z u , v , t = ( A w a r , I n t r , I n t e n t , F e a s , R e l , H or i z o n , D ec l ) .
本文將更新改為:
z t + 1 = F ( z t , e t , I ( e t ) , O G E ( e t ) , D e c l t ) . \mathbf{z}_{t+1}
=
F(
\mathbf{z}_t,
e_t,
I(e_t),
OGE(e_t),
Decl_t
). z t + 1 = F ( z t , e t , I ( e t ) , O GE ( e t ) , D ec l t ) .
而不是:
z t + 1 = F ( z t , w a t c h t i m e t ) . \mathbf{z}_{t+1}
=
F(
\mathbf{z}_t,
watchtime_t
). z t + 1 = F ( z t , w a t c h t im e t ) .
例如 autoplay 事件可能提高:
A w a r Awar A w a r
因為使用者確實接觸到內容,
但不一定應提高:
I n t r Intr I n t r
或:
I n t e n t . Intent. I n t e n t .
這是一個關鍵語意分離:
P a s s i v e E x p o s u r e ⇒ A w a r e n e s s E v i d e n c e \boxed{
PassiveExposure
\Rightarrow
AwarenessEvidence
} P a ss i v e E x p os u r e ⇒ A w a r e n ess E v i d e n ce
但:
P a s s i v e E x p o s u r e ⇏ I n t e r e s t E v i d e n c e \boxed{
PassiveExposure
\not\Rightarrow
InterestEvidence
} P a ss i v e E x p os u r e ⇒ I n t er es tE v i d e n ce
20. Autoplay Can Update Awareness without Updating Interest
這提供 autoplay 更合理的資料用途。
若:
P A ( u , v , t ) = 1 , PA(u,v,t)=1, P A ( u , v , t ) = 1 ,
可以更新:
A w a r ( u , v , t + 1 ) > A w a r ( u , v , t ) . Awar(u,v,t+1)
>
Awar(u,v,t). A w a r ( u , v , t + 1 ) > A w a r ( u , v , t ) .
因為使用者確實已被暴露於 v v v 。
但:
I n t r ( u , v , t + 1 ) ≈ I n t r ( u , v , t ) Intr(u,v,t+1)
\approx
Intr(u,v,t) I n t r ( u , v , t + 1 ) ≈ I n t r ( u , v , t )
除非後續出現:
A c t i v e P l a y , S e a r c h , S e e k , S a v e , F o l l o w , L i k e ActivePlay,
Search,
Seek,
Save,
Follow,
Like A c t i v e P l a y , S e a r c h , S ee k , S a v e , F o l l o w , L ik e
等更強 signal。
這直接避免 A02 所述:
A w a r e n e s s → I n t e r e s t Awareness
\rightarrow
Interest A w a r e n ess → I n t er es t
的錯誤坍縮。
21. Repeated Exposure and Familiarity Effects
重複曝光可能造成 familiarity、mere-exposure effect 或 recognition。
即使不討論心理學上的偏好改變,推薦系統在資料層至少應知道:
D w e l l t Dwell_t D w e l l t
可能因:
P r i o r E x p o s u r e < t PriorExposure_{<t} P r i or E x p os u r e < t
而增加。
因此:
D w e l l Dwell D w e l l
不是完全獨立的偏好 proxy。
可將:
D w e l l R e s i d u a l = D w e l l O b s e r v e d − D w e l l ^ ( E x p o s u r e C o u n t , S u r f a c e , R a n k ) DwellResidual
=
DwellObserved
-
\hat{Dwell}(ExposureCount,Surface,Rank) D w e l l R es i d u a l = D w e l l O b ser v e d − D w e l l ^ ( E x p os u r e C o u n t , S u r f a ce , R ank )
作為比 raw dwell 更乾淨的訊號之一。
22. Product-Level Failure Mode
如果產品團隊分別優化:
A u t o p l a y R a t e ↑ , AutoplayRate\uparrow, A u t o pl a y R a t e ↑ ,
W a t c h E v e n t s ↑ , WatchEvents\uparrow, W a t c h E v e n t s ↑ ,
H i s t o r y C o v e r a g e ↑ , HistoryCoverage\uparrow, H i s t or y C o v er a g e ↑ ,
R e c o m m e n d a t i o n C T R ↑ , RecommendationCTR\uparrow, R eco mm e n d a t i o n C T R ↑ ,
每個局部 KPI 都可能改善。
但整體系統可能:
A u t o p l a y → W a t c h E v e n t → P r e f e r e n c e E s t i m a t e → R e p e a t e d R e c o m m e n d a t i o n . Autoplay
\rightarrow
WatchEvent
\rightarrow
PreferenceEstimate
\rightarrow
RepeatedRecommendation. A u t o pl a y → W a t c h E v e n t → P r e f er e n ce E s t ima t e → R e p e a t e d R eco mm e n d a t i o n .
結果:
D a s h b o a r d S u c c e s s = 1 DashboardSuccess=1 D a s hb o a r d S u ccess = 1
但:
U s e r U t i l i t y ↓ . UserUtility\downarrow. U ser U t i l i t y ↓ .
因此 A03 為 A06 的:
M e t r i c S u c c e s s ≠ P r o d u c t S u c c e s s MetricSuccess
\neq
ProductSuccess M e t r i c S u ccess = P r o d u c tS u ccess
提供微觀資料生成機制。
23. Evaluation Metrics
23.1 Passive Exposure Contamination Index
已定義:
P E C I . PECI. P E C I .
23.2 Endogenous Evidence Ratio
已定義:
E E R . EER. E E R .
23.3 Self-Confirmation Ratio
已定義:
S C R . SCR. S C R .
23.4 Re-exposure Amplification Factor
已定義:
R A F . RAF. R A F .
23.5 Intentionality Calibration Error
令模型給事件 e e e 的 preference evidence strength 為:
w ^ ( e ) , \hat{w}(e), w ^ ( e ) ,
而經實驗估計的 intentional preference contribution 為:
w ∗ ( e ) . w^*(e). w ∗ ( e ) .
則:
I C E = E [ ∣ w ^ ( e ) − w ∗ ( e ) ∣ ] . ICE
=
\mathbb{E}
[
|\hat{w}(e)-w^*(e)|
]. I C E = E [ ∣ w ^ ( e ) − w ∗ ( e ) ∣ ] .
23.6 Counterfactual Preference Error
若 randomized or causal estimate 提供:
p u , v c f , p^{cf}_{u,v}, p u , v c f ,
模型估計為:
p ^ u , v , \hat{p}_{u,v}, p ^ u , v ,
則:
C P E = E [ ∣ p ^ u , v − p u , v c f ∣ ] . CPE
=
\mathbb{E}
[
|\hat{p}_{u,v}-p^{cf}_{u,v}|
]. C P E = E [ ∣ p ^ u , v − p u , v c f ∣ ] .
24. Empirical Protocol
24.1 Autoplay versus manual-play cohort
建立兩組 exposure:
G A = autoplay , G_A=\text{autoplay}, G A = autoplay ,
G M = manual play . G_M=\text{manual play}. G M = manual play .
控制 item、rank、session context 後比較:
P ( F u t u r e S e a r c h ∣ G A ) P(
FutureSearch
\mid
G_A
) P ( F u t u r e S e a r c h ∣ G A )
與:
P ( F u t u r e S e a r c h ∣ G M ) . P(
FutureSearch
\mid
G_M
). P ( F u t u r e S e a r c h ∣ G M ) .
若兩者差異顯著,則不能將初始 watch duration 同質解讀。
24.2 History-write audit
記錄:
A u t o p l a y D u r a t i o n AutoplayDuration A u t o pl a y D u r a t i o n
與:
H i s t o r y W r i t e HistoryWrite H i s t or y W r i t e
之間的 threshold。
檢查:
P ( H i s t o r y W r i t e = 1 ∣ N o U s e r I n i t i a t i o n ) . P(
HistoryWrite=1
\mid
NoUserInitiation
). P ( H i s t or y W r i t e = 1 ∣ N o U ser I ni t ia t i o n ) .
24.3 Re-exposure audit
對第一次 passive exposure 的內容,測量:
R A F . RAF. R A F .
並將其與第一次 active engagement 的內容比較。
24.4 Explicit-declaration conflict test
使用 A02 的:
S N O O Z E , K B D , M A J O R _ O N L Y SNOOZE,
KBD,
MAJOR\_ONLY S N O O Z E , K B D , M A J O R _ O N L Y
等 declaration,測試 autoplay telemetry 是否會錯誤覆寫。
24.5 Randomized holdout
對小比例內容使用受控曝光,建立 counterfactual calibration set,用以估計:
C P E . CPE. C P E .
25. Case-Study Boundary
本文所討論的 autoplay-history contamination 是一種可檢驗的系統機制,而不是對任何特定平台未公開後端的事實斷言。
觀察到:
A u t o p l a y Autoplay A u t o pl a y
與:
H i s t o r y E n t r y HistoryEntry H i s t or y E n t r y
同時存在,只能證明前端與帳戶紀錄之間存在某種關係。
若要進一步主張:
H i s t o r y E n t r y → R e c o m m e n d a t i o n W e i g h t HistoryEntry
\rightarrow
RecommendationWeight H i s t or y E n t r y → R eco mm e n d a t i o nW e i g h t
具有特定大小,仍需要:
account-level experiment;
recommendation distribution before / after;
controlled interactions;
large-sample logging;
或平台公開技術文件。
因此本文將 Bilibili、YouTube、TikTok 等平台視為可套用此框架的實驗場,而不以單一使用者經驗推出其內部模型參數。
26. Design Requirements
一個 provenance-safe recommendation system 至少應滿足:
Requirement 1
E x p o s u r e E v e n t ≠ P r e f e r e n c e E v e n t . ExposureEvent
\neq
PreferenceEvent. E x p os u r e E v e n t = P r e f er e n ce E v e n t .
Requirement 2
A u t o p l a y E v e n t ≠ A c t i v e P l a y E v e n t . AutoplayEvent
\neq
ActivePlayEvent. A u t o pl a y E v e n t = A c t i v e P l a y E v e n t .
Requirement 3
所有 preference-relevant events 必須包含:
P r o v e n a n c e . Provenance. P r o v e nan ce .
Requirement 4
模型更新必須知道:
I n t e n t i o n a l i t y . Intentionality. I n t e n t i o na l i t y .
Requirement 5
平台需要至少一種估計:
C o u n t e r f a c t u a l P r e f e r e n c e CounterfactualPreference C o u n t er f a c t u a l P r e f er e n ce
的方法。
Requirement 6
明示使用者 declaration 不應被低 intentionality telemetry 無限制覆寫。
Requirement 7
資料管線必須允許回溯:
Why did this preference score change? \text{Why did this preference score change?} Why did this preference score change?
27. Extended Closed-Loop Model
綜合 A01–A03,可得:
C t → O t L t → E x p o s u r e / U I E t → T e l e m e t r y T t → P z ^ t + 1 → R a n k i n g O t + 1 \boxed{
\mathcal{C}_t
\xrightarrow{
\mathcal{O}_t
}
L_t
\xrightarrow{
Exposure/UI
}
\mathcal{E}_t
\xrightarrow{
Telemetry
}
T_t
\xrightarrow{
\mathcal{P}
}
\hat{\mathbf{z}}_{t+1}
\xrightarrow{
Ranking
}
\mathcal{O}_{t+1}
} C t O t L t E x p os u r e / U I E t T e l e m e t r y T t P z ^ t + 1 R ank in g O t + 1
如果:
T e l e m e t r y Telemetry T e l e m e t r y
沒有 provenance,則:
O t \mathcal{O}_t O t
的效果會偷偷流入:
z ^ t + 1 . \hat{\mathbf{z}}_{t+1}. z ^ t + 1 .
此時推薦系統的狀態估計變成:
z ^ t + 1 = U s e r P r e f e r e n c e + P l a t f o r m H i s t o r y + I n t e r f a c e A r t i f a c t . \hat{\mathbf{z}}_{t+1}
=
UserPreference
+
PlatformHistory
+
InterfaceArtifact. z ^ t + 1 = U ser P r e f er e n ce + P l a t f or m H i s t or y + I n t er f a ce A r t i f a c t .
卻被系統誤以為:
z ^ t + 1 = U s e r P r e f e r e n c e . \hat{\mathbf{z}}_{t+1}
=
UserPreference. z ^ t + 1 = U ser P r e f er e n ce .
這正是 Endogenous Preference Contamination。
28. Implications for the Next Papers
A03 建立後,A04 的 Platform-Induced Exposure Bubble 可被更精確表示。
若某些熱門內容具有:
q t ( v ) ≫ q t ( w ) , q_t(v)\gg q_t(w), q t ( v ) ≫ q t ( w ) ,
且高 exposure 自己又生成更多 preference evidence:
q t ( v ) ↑ ⇒ E v i d e n c e t ( v ) ↑ ⇒ q t + 1 ( v ) ↑ , q_t(v)\uparrow
\Rightarrow
Evidence_t(v)\uparrow
\Rightarrow
q_{t+1}(v)\uparrow, q t ( v ) ↑⇒ E v i d e n c e t ( v ) ↑⇒ q t + 1 ( v ) ↑ ,
則 popularity concentration 不再只是排序偏好,而是一個動態 amplification process。
A05 則會進一步問:
當新人內容沒有足夠 exposure,因此沒有足夠 evidence 時,模型是否把「沒有證據」錯誤解釋成「證據顯示沒有人喜歡」?
因此 A03 是從個體 preference contamination 通往 creator ecology 的橋樑。
29. Limitations
第一,intentionality 不是完全可觀測變數。manual click 也可能是誤觸,autoplay 也可能最終轉為真正興趣,因此不能使用僵硬 binary mapping。
第二,causal correction 需要額外假設、randomization 或可靠 propensity model;在大型 production systems 中可能有成本。
第三,部分平台目標本來就包含 discovery,因此 observation-generated evidence 不應全部被丟棄。本文主張的是分離其語意,而不是將其視為無效。
第四,使用者偏好本身也可能被推薦真正改變。本文區分「推薦改變真實偏好」與「推薦只改變 telemetry 卻被誤判為偏好」;前者是更深層的 preference formation 問題,不在本篇完全處理。
第五,本文主要處理個人化推薦資料生成。廣告競價、商業 boost、內容審核與安全政策等其他 exposure mechanisms 可在後續擴展。
30. Conclusion
本文提出 Endogenous Preference Contamination framework。
核心問題可表示為:
R e c o m m e n d a t i o n → E x p o s u r e → T e l e m e t r y → P r e f e r e n c e E s t i m a t e → R e c o m m e n d a t i o n \boxed{
Recommendation
\rightarrow
Exposure
\rightarrow
Telemetry
\rightarrow
PreferenceEstimate
\rightarrow
Recommendation
} R eco mm e n d a t i o n → E x p os u r e → T e l e m e t r y → P r e f er e n ce E s t ima t e → R eco mm e n d a t i o n
當 telemetry 沒有保留其生成來源時:
P l a t f o r m I n d u c e d B e h a v i o r ≈ U s e r P r e f e r e n c e E v i d e n c e \boxed{
PlatformInducedBehavior
\approx
UserPreferenceEvidence
} P l a t f or m I n d u ce d B e ha v i or ≈ U ser P r e f er e n ce E v i d e n ce
會成為一個危險的錯誤等價。
本文因此提出:
R e c o m m e n d a t i o n - I n d u c e d P r e f e r e n c e A r t i f a c t \boxed{
Recommendation\text{-}Induced\ Preference\ Artifact
} R eco mm e n d a t i o n - I n d u ce d P r e f er e n ce A r t i f a c t
以及:
O b s e r v a t i o n - G e n e r a t e d E v i d e n c e \boxed{
Observation\text{-}Generated\ Evidence
} O b ser v a t i o n - G e n er a t e d E v i d e n ce
並以:
P E C I , E E R , S C R , R A F , I C E , C P E PECI,
EER,
SCR,
RAF,
ICE,
CPE P E C I , E E R , S C R , R A F , I C E , C P E
等指標衡量資料污染與自我證成程度。
最重要的語意原則是:
P a s s i v e E x p o s u r e ⇒ A w a r e n e s s E v i d e n c e \boxed{
PassiveExposure
\Rightarrow
AwarenessEvidence
} P a ss i v e E x p os u r e ⇒ A w a r e n ess E v i d e n ce
但:
P a s s i v e E x p o s u r e ⇏ I n t e r e s t E v i d e n c e \boxed{
PassiveExposure
\not\Rightarrow
InterestEvidence
} P a ss i v e E x p os u r e ⇒ I n t er es tE v i d e n ce
同樣地:
A u t o p l a y ≠ A c t i v e C o n s u m p t i o n \boxed{
Autoplay
\neq
ActiveConsumption
} A u t o pl a y = A c t i v e C o n s u m pt i o n
以及:
H i s t o r y ≠ H o m o g e n e o u s P r e f e r e n c e E v i d e n c e \boxed{
History
\neq
HomogeneousPreferenceEvidence
} H i s t or y = H o m o g e n eo u s P r e f er e n ce E v i d e n ce
推薦系統若希望真正理解使用者,不能只收集更多行為,而必須知道行為是如何被產生的。
因此,對推薦資料而言:
P r o v e n a n c e is not metadata about preference evidence; P r o v e n a n c e is part of the preference evidence itself. \boxed{ Provenance \text{ is not metadata about preference evidence;} \quad Provenance \text{ is part of the preference evidence itself.} } P r o v e nan ce is not metadata about preference evidence; P r o v e nan ce is part of the preference evidence itself.
這完成 Series A 的第三層理論基礎:
A01:誰決定使用者看見什麼;
A02:系統如何理解使用者與內容的關係;
A03:系統本身如何污染它用來理解使用者的資料。
下一步 A04 將把這個閉環提升到整個平台曝光分布,處理從個人資訊繭房轉化為全平台熱門度繭房的機制。
References
[0] Neo.K / EveMissLab. “Recommendation as an Observation Operator: A Formal Framework for Content Availability, Observability, Discoverability, and User Agency.” Series A, Paper A01, v0.1, 2026.
[1] Neo.K / EveMissLab. “Explicit Preference versus Inferred Preference: A Multidimensional User-State Model from Awareness to Willingness to Allocate Resources.” Series A, Paper A02, v0.1, 2026.
[2] Hu, Y., Koren, Y., & Volinsky, C. “Collaborative Filtering for Implicit Feedback Datasets.” 2008 Eighth IEEE International Conference on Data Mining, pp. 263–272, 2008. DOI: 10.1109/ICDM.2008.22.
[3] Chaney, A. J. B., Stewart, B. M., & Engelhardt, B. E. “How Algorithmic Confounding in Recommendation Systems Increases Homogeneity and Decreases Utility.” Proceedings of the 12th ACM Conference on Recommender Systems, pp. 224–232, 2018. DOI: 10.1145/3240323.3240370. arXiv:1710.11214.
[4] Mansoury, M., Abdollahpouri, H., Pechenizkiy, M., Mobasher, B., & Burke, R. “Feedback Loop and Bias Amplification in Recommender Systems.” Proceedings of the 29th ACM International Conference on Information & Knowledge Management, 2020. arXiv:2007.13019.
[5] Pan, W., Cui, S., Wen, H., Chen, K., Zhang, C., & Wang, F. “Correcting the User Feedback-Loop Bias for Recommendation Systems.” arXiv:2109.06037, 2021.
[6] Adomavicius, G., Bockstedt, J. C., Curley, S. P., & Zhang, J. “Recommender Systems, Ground Truth, and Preference Pollution.” AI Magazine, 43(2), 177–189, 2022. DOI: 10.1002/aaai.12055.
[7] Krauth, K., Wang, Y., & Jordan, M. I. “Breaking Feedback Loops in Recommender Systems with Causal Inference.” arXiv:2207.01616, 2022.
[8] Liao, J., Yang, M., Zhou, W., Zhang, H., & Wen, J. “Modeling Item Exposure and User Satisfaction for Debiased Recommendation with Causal Inference.” Information Sciences, 676, 120834, 2024. DOI: 10.1016/j.ins.2024.120834.
[9] Zoralioglu, Y., & Yalcin, E. “Dynamic Feedback Loops in Recommender Systems: Analyzing Fairness, Popularity Bias, and User Group Disparities.” Journal of Intelligent Information Systems, 2026. DOI: 10.1007/s10844-026-01025-y.
Series Continuation
A04 — The Platform-Induced Exposure Bubble
A05 — Recommendation, Cold Start, and Creator Ecological Collapse
A06 — Metric Success, Product Failure