VWDC-06 — Transport-Aware World Decisions, Policy Transfer, and Reality-Gap Robust Control
轉移感知世界決策、策略搬運與現實差距穩健控制:部署證書、Reality Regret、Probe/Fallback 與自致分佈漂移
Bridge Series: Visual–World Domain Computation (VWDC) — Paper 06
Depends on: VWDC-01–05, GVSS-01–10, WDC-01–08, WDC Runtime Whitepaper, frozen RRT-20
Author: Neo.K / EveMissLab
Version: v0.1
Date: 2026-08-17
Status: Formal transport-aware decision/control paper. Action-wise deployment certificates, real-action regret under transport intervals, robust lower-bound action selection, unsupported-action non-identifiability, perfect-probe/fallback thresholds, fixed-policy discounted sim-to-real value bounds, transferred-policy regret bounds, deployment-induced occupancy-shift debt, RTC admissibility under policy shift, supported-action fallback gates, online local calibration concentration, and finite-horizon deploy/probe/fallback Bellman statements are proved under explicit hypotheses. Robust MDP/RL, domain randomization, sim-to-real policy transfer, offline-RL pessimism, safe policy adaptation, and digital-twin control are established neighboring research and are not claimed as VWDC inventions. No strong novelty claim is made.
Keywords: sim-to-real policy transfer, robust control, digital twin, policy deployment, reality gap, transport contract, safe fallback, probe value, unsupported action, occupancy shift, reality regret, WDC
Abstract
VWDC-05 introduced a claim-specific Reality Transport Contract:
That contract tells us when a world-derived quantity may support a reality claim.
VWDC-06 asks the operational question:
When may a decision or policy optimized inside a WDC world be executed in reality?
A simulation-optimal action is not automatically reality-optimal.
A reality-facing controller therefore maintains action-wise transported value estimates:
and Reality Transport Contract uncertainty radii:
The contract asserts, on supported state/action region:
The controller's deployment action set is:
The first result is a direct deployment certificate.
Let:
If:
then:
for every reality action-value vector consistent with the RTC intervals.
Thus deployment can be certified when the simulated decision margin dominates transport debt.
When this condition fails, the controller has no certificate that the simulated optimum survives the reality gap.
1. Action-wise reality transport
Let:
be the current deployment state/context.
Let:
The world model supplies:
The RTC supplies a supported region:
and transport error:
2. Supported action
Definition VWDC06-D1
Action is RTC-supported at state if:
and the relevant RTC is current.
3. Transport interval
Definition VWDC06-D2
For supported pair:
define:
The contract claims:
4. Simulated optimum
5. Reality optimum
These need not agree.
6. VWDC06-T1 — Transport-margin deployment certificate
Theorem VWDC06-T1
Assume every candidate action is RTC-supported and satisfies:
If some action satisfies:
then is the unique reality-optimal action for every reality value vector consistent with the intervals.
Proof
For every :
and:
Apply the strict margin condition.
7. Interpretation
The criterion is stronger than:
simulated margin is large.
It requires:
simulated margin exceeds the transport uncertainty of both winner and competitors.
Thus:
can still yield no deployment certificate.
8. Nominal deployment region
Define:
Inside this region, nominal simulation and transported reality decisions agree under the RTC.
9. Uncertified region
Outside:
the simulated optimum can still be correct.
It is merely not certified by the current RTC.
Possible actions:
- robust deploy;
- external probe;
- safe fallback;
- human review;
- stop.
10. VWDC06-T2 — Reality action regret under transport intervals
Theorem VWDC06-T2
Let:
and:
Under action-wise interval bounds:
the real regret of deploying satisfies:
Hence:
Proof
because:
11. Transport debt becomes decision debt
VWDC-05's external-validity debt is therefore not only epistemic.
It directly upper-bounds possible action regret.
12. Robust lower-bound deployment
Define:
Choose:
13. VWDC06-T3 — Robust lower-bound guarantee
Theorem VWDC06-T3
For every reality value vector consistent with the RTC:
Proof
For selected action:
By definition:
This is a robust performance floor, not proof of reality optimality.
14. Safe fallback
Let:
be a fallback action with independently certified reality value lower bound:
15. VWDC06-T4 — Supported-action fallback gate
Theorem VWDC06-T4
If:
then every supported simulation action has a lower certified floor below the fallback guarantee, and choosing the fallback maximizes the available certified lower bound.
Proof
Immediate comparison of certified lower bounds.
This is a pessimistic risk policy.
It need not maximize expected reward.
16. Unsupported actions
Suppose:
No RTC interval is available.
17. VWDC06-N1 — Unsupported-action reality value is not identified by supported RTC data alone
Proposition VWDC06-N1
Without structural/generalization assumptions, there exist two reality models:
that agree on every RTC-supported state/action pair and differ arbitrarily on:
Therefore supported validation data alone do not identify:
Proof
Define the two reality models identically on supported pairs.
Assign different rewards/transitions to the unsupported pair.
No supported observation distinguishes them.
18. Consequence
If the simulator's best action is unsupported, there is no distribution-free guarantee that it is safe or valuable in reality.
Therefore:
19. Explicit unsupported-action risk policies
Possible policies:
PROBE_FIRST
SAFE_FALLBACK
HUMAN_REVIEW
STRUCTURAL_EXTRAPOLATION_WITH_BOUND
ROBUST_WORST_CASE
FORBID
The policy must be explicit.
20. Probe action
An external probe can reduce action-value uncertainty before deployment.
Let:
be probe cost.
21. Two-regime risky-action example
Compare:
- risky action ;
- safe fallback value normalized to .
Risky action gives:
in good regime with probability and:
in bad regime with probability:
where:
Without probe, expected risky value:
Optimal no-probe value is:
22. Perfect regime probe
A perfect probe reveals good/bad regime.
Then controller chooses risky action only in good regime.
Probe value before cost:
After cost:
23. VWDC06-T5 — Perfect-probe versus fallback threshold
Theorem VWDC06-T5
The perfect probe is strictly better than the best immediate action iff:
Proof
Probe advantage is:
Use identity:
for:
24. Interpretation
Probe value is highest near the decision boundary.
If risky action is almost certainly good or almost certainly bad, the value of perfect information is small.
25. Human review
Human review can be modeled as a noisy/costly probe.
Its value depends on:
- accuracy;
- decision consequence;
- latency;
- cost.
No universal preference over machine probe is claimed.
26. Current safe Sim2Real precedent
Current 2026 Sim2Real work explicitly adjusts deployment risk based on uncertainty about the target environment context.
Other work uses uncertainty estimates to adapt policies or focus real/sim exploration.
VWDC's RTC acts as a contract-level version of the same broad deployment concern.
27. Policy transfer
A policy:
is optimized in world model:
Reality executes:
under:
Policy value can differ due to:
- reward mismatch;
- transition mismatch;
- observation mismatch;
- action execution mismatch.
28. Discounted MDP setting
Assume:
World and reality share state/action spaces for this theorem.
Rewards:
For every:
assume:
Transition kernels satisfy:
29. VWDC06-T6 — Fixed-policy discounted Sim2Real value bound
Theorem VWDC06-T6
For any fixed policy :
Proof
For any state:
The last term uses the bounded-function TV inequality for:
Take the supremum and rearrange.
30. Horizon amplification
Transition mismatch is amplified by approximately:
Long-horizon policies can therefore be much more sensitive to small local dynamics mismatch.
31. Uniform policy transport error
Define:
Then:
for every fixed policy under the theorem's assumptions.
32. World-optimal versus reality-optimal policy
Let:
33. VWDC06-T7 — Transferred-policy reality regret bound
Theorem VWDC06-T7
If the fixed-policy transport bound:
holds uniformly for all policies under consideration, then:
Proof
because:
34. Interpretation
A bounded world-to-reality value error yields a direct policy-transfer regret certificate.
It does not say zero-shot deployment is always safe.
It says policy regret is controlled only inside the theorem's discrepancy assumptions.
35. Current robust-transfer precedent
Recent work studies:
- constrained Sim2Real policy transfer with adaptive risk;
- robust transfer under uncertainty sets;
- safe domain randomization with OOD detection;
- deployment-time continual adaptation;
- residual policy correction.
These are direct engineering neighbors.
VWDC does not claim robust Sim2Real RL as new.
36. Pessimism
When model/reality uncertainty is large, a controller can optimize conservative lower bounds rather than nominal simulated values.
This is classical robust/offline RL logic.
VWDC ties the pessimism radius to RTC support and transport debt.
37. Unsupported-state/action pessimism
For unsupported pairs, a conservative lower bound can be:
- domain-specific safety floor;
- worst-case allowed value;
- / forbidden.
Choice is a risk-policy decision.
38. Deployment distribution
An RTC is validated under a context/state-action occupancy distribution:
Deploying policy induces:
These need not match.
39. Local transport loss
Let:
Old validated expected transport loss:
Deployment loss:
40. VWDC06-T8 — Policy-induced occupancy-shift debt
Theorem VWDC06-T8
Therefore:
Proof
Bounded-function total-variation inequality.
41. Self-invalidating deployment
A policy optimized in simulation can intentionally visit states rarely or never visited by the validation policy.
Thus deployment can increase:
and invalidate its own pre-deployment certainty.
This is a form of policy-induced external-validity shift.
42. RTC deployment tolerance
Let decision policy tolerate transport debt:
If:
the old contract remains sufficient under this conservative shift bound.
43. VWDC06-T9 — RTC certification loss under occupancy shift
Theorem VWDC06-T9
If the only available certification is:
and the right-hand side exceeds decision tolerance , then the existing RTC no longer certifies deployment at tolerance .
Proof
The available upper bound is insufficient to establish:
This is loss of certification, not proof that deployment is actually invalid.
44. Distinguish invalidity from uncertified status
Use:
VALIDATED
CERTIFIED_FOR_POLICY
UNCERTIFIED_AFTER_SHIFT
INVALIDATED_BY_EVIDENCE
A failed certificate does not prove the policy will fail.
45. Policy-induced support expansion
If:
places mass on unsupported RTC cells, the transport guarantee must account for uncovered mass as in VWDC-05.
46. Unsupported occupancy debt
Let:
If transport loss is bounded by :
47. Safe exploration
Before entering an unsupported cell, controller can:
- probe;
- fallback;
- ask human;
- run shadow mode;
- collect external validation.
48. Shadow deployment
A policy can produce recommendations without controlling the real system.
External observations update RTC without action execution.
This can reduce uncertainty with lower intervention risk.
49. Intervention probe
A real probe intentionally executes a bounded action to learn local transport behavior.
It can be higher information but higher risk.
50. Safe probe contract
Probe itself requires:
- permitted action envelope;
- cost/risk bound;
- measurement plan;
- stop condition;
- rollback/fallback.
51. Human gate
For high-stakes unsupported actions:
can be a mandatory deployment gate.
This is policy governance, not an assertion that humans are infallible.
52. Safe fallback override
A safe fallback can override higher simulated value when:
- transport lower bound is weak;
- unsupported mass is high;
- reality drift is detected;
- downstream loss is asymmetric.
53. Online reality feedback
After deployment/probe, collect:
Update:
- action-value estimates;
- local discrepancy;
- support status;
- RTC version;
- policy.
54. Local bounded reward calibration
Fix one state/action cell.
Observed reality reward:
World predicted mean:
Reality empirical mean:
55. VWDC06-T10 — Online local reward calibration certificate
Theorem VWDC06-T10
For independent reality observations in one fixed cell, with true reality mean :
With probability at least :
Thus a transport discrepancy estimate:
must add both world-prediction and reality-measurement uncertainty before being used as a contract radius.
Proof
Hoeffding inequality.
56. Online feedback does not automatically solve structural mismatch
More observations shrink statistical uncertainty in visited cells.
They do not identify unsupported counterfactual/action regions without exploration/assumptions.
57. Dynamic policy adaptation precedent
Current Sim2Real work adapts deployed policies according to inferred target-environment context or uncertainty.
Safe continual domain adaptation also updates policies after deployment while trying to preserve safety.
VWDC uses RTC-local validity as the governance object around such adaptation.
58. Policy update versioning
Every deployed policy has:
Every policy update creates:
with:
- parent policy;
- update data;
- RTC used;
- safety/fallback status.
59. RTC-policy pair
A deployment certificate applies to:
Changing policy can change occupancy.
Changing RTC can change certified actions.
Version both.
60. Closed-loop feedback
The transport layer is reflexive.
61. Policy deployment can create data
Executed actions create:
- new observations;
- new supported cells;
- possible safety incidents;
- policy-selection bias.
Record the deployment policy.
62. Deployment-selection bias
Reality data collected under current policy are not a neutral sample from all state/action pairs.
Future calibration should account for visitation policy.
This connects directly to GVSS-10 routing-selection bias.
63. Off-policy correction
If deployment propensities are known and support holds, causal/off-policy methods may reweight reality observations.
VWDC-06 does not rederive off-policy evaluation.
64. Unsupported actions remain unsupported
No estimator can recover action outcomes in regions never observed without assumptions.
Online feedback only helps where data enter.
65. Reality action mask
Maintain:
66. Status examples
CERTIFIED
SUPPORTED
EXTRAPOLATED
UNSUPPORTED
FORBIDDEN
Risk tier determines which statuses are deployable.
67. Risk-tier policy
Example:
LOW_STAKES
Allow supported/extrapolated actions within debt threshold.
HIGH_STAKES
Require certified actions or human gate.
SAFETY_CRITICAL
Require domain-specific safety certificate beyond RTC alone.
68. RTC is not a safety proof
A reality transport contract bounds a declared model/measurement discrepancy.
It does not replace:
- physical safety constraints;
- formal verification;
- regulatory requirements;
- operator training.
69. Safety filter
A runtime may apply a separate certified safety filter:
Only actions passing both transport and safety gates execute.
70. Deployment gate
71. Authority gate
Some actions require:
- human approval;
- legal authorization;
- operational role.
Decision optimality does not imply authority.
72. Robust deployment action
A robust controller solves:
For independent action intervals, this reduces to maximizing lower bounds.
73. Coupled uncertainty
If RTC errors across actions are coupled, interval-wise lower-bound optimization can be conservative or incomplete.
Use structured ambiguity set.
VWDC-06 does not solve general robust MDP ambiguity geometry.
74. Nominal versus robust decision
If VWDC06-T1 holds, nominal and robust decisions coincide.
If intervals overlap, robust and nominal actions can differ.
75. Value of external probe
An external probe has value when expected reduction in deployment decision risk exceeds:
- probe cost;
- delay;
- intervention risk.
This is VWDC-04 value-of-branch applied to reality-facing decisions.
76. Value of human review
Same principle:
Human review is not free information.
77. Probe/fallback region
If risky action could dominate fallback but is not certified, the controller compares:
- probe;
- fallback;
- robust alternative;
- human review.
78. Perfect probe theorem meaning
VWDC06-T5 says the perfect probe is valuable only when there is meaningful decision ambiguity.
It is not valuable simply because uncertainty exists.
79. Deployment abstention
STOP/ABSTAIN is a valid action.
A controller can say:
NO ACTION CURRENTLY HAS ACCEPTABLE RTC + SAFETY CERTIFICATE
80. Current foundation-agent reality gap
Recent 2026 work explicitly frames agent deployment as facing noisy inputs, stochastic transitions, execution constraints, and distribution shifts absent from clean benchmark environments.
This broader observation reinforces the need to separate simulation/benchmark optimality from deployment reliability.
81. Reality gap taxonomy
For deployment:
Do not collapse all failures to one scalar if diagnosis matters.
82. Observation gap
Reality sensor process differs from simulated observation model.
83. Dynamics gap
State transitions differ.
84. Action gap
Commanded and executed action differ.
85. Reward/utility gap
Simulator objective differs from actual deployment utility.
86. Measurement gap
External evaluator/sensor differs from target latent quantity.
87. Support gap
RTC lacks validated data for state/action region.
88. Policy-shift gap
Deployment policy induces a new occupancy distribution.
89. Gap-localized fallback
Fallback can be selected by which gap dominates.
Examples:
- observation gap → human/sensor fallback;
- dynamics gap → robust conservative controller;
- action gap → actuator-safe mode;
- support gap → probe/abstain;
- policy-shift gap → revalidation.
90. Sim2Real policy adaptation
Current approaches include:
- domain randomization;
- offline domain randomization;
- residual adaptation;
- context inference;
- risk-sensitive dynamic adaptation;
- safe continual adaptation.
VWDC does not prescribe one.
RTC tells the Governor where each adaptation result is externally supported.
91. Offline domain randomization precedent
Recent provable offline domain randomization uses limited real data to fit distributions over simulator parameters and studies policy transfer guarantees.
This is a direct neighbor to reality-tethered policy transfer.
92. Real-Sim-Real loop precedent
Current Real-Sim-Real frameworks iteratively refine simulator parameters with real-world data and retrain/adapt policies.
VWDC's loop:
is a governance-level analogue.
93. Context-adaptive transfer
Recent work conditions/adapts policies on inferred deployment context rather than relying only on one robust policy.
This corresponds to a state-dependent RTC and local transport debt.
94. Safe adaptation
A policy can adapt only inside a safe/authority envelope.
Policy performance improvement does not excuse violating hard constraints.
95. Value transport and safety transport
A world model may accurately transport reward value while poorly transporting safety-event probability.
Maintain separate RTCs:
96. VWDC06-N2 — Reward transport validity does not imply safety transport validity
Counterexample
Simulator predicts task reward exactly but omits a rare real-world hazard state.
Reward RTC is accurate on observed reward.
Safety-event model is wrong.
Therefore:
97. Multi-contract gate
A deployed action can require:
for:
- reward;
- safety;
- resource;
- fairness;
- operational constraints.
98. Contract conflict
Different RTCs can recommend different actions.
Use multiobjective/constrained control.
Do not average incompatible safety constraints into utility silently.
99. Constrained deployment
Example:
subject to:
100. Unsupported safety action
If safety risk is unsupported, high simulated reward cannot compensate under a hard safety policy.
101. Policy-transfer acceptance packet
policy_id
world_model_version
rtc_versions
state_action_scope
simulated_value
reality_value_bound
safety_contracts
occupancy_shift_bound
unsupported_mass
fallback_policy
human_gate
expiry
102. Per-decision packet
decision_id
state_context
candidate_actions
world_values
transport_intervals
supported_status
selected_action
selection_mode
probe_or_human_result
rtc_version
safety_gate
authority_gate
103. Deployment modes
CERTIFIED_DEPLOY
ROBUST_DEPLOY
SHADOW
PROBE
SAFE_FALLBACK
HUMAN_REQUIRED
STOPPED
104. CERTIFIED_DEPLOY
Use when transport-margin certificate and all safety/authority gates pass.
105. ROBUST_DEPLOY
Use when nominal ranking is uncertain but robust lower-bound action meets policy threshold.
106. SHADOW
Generate action without executing.
Collect reality observations and compare.
107. PROBE
Execute bounded information-gathering action.
108. SAFE_FALLBACK
Use predeclared conservative action/controller.
109. HUMAN_REQUIRED
Escalate when action consequence/risk policy requires human authority or information.
110. STOPPED
No feasible action satisfies deployment policy.
111. Policy-shift monitor
Track empirical:
versus validation occupancy:
Trigger RTC review when shift exceeds policy threshold.
112. Occupancy estimation uncertainty
Empirical TV estimates are themselves uncertain.
The runtime should not treat estimated shift as exact.
VWDC-06 records this as future statistical refinement.
113. State-action drift map
Instead of one global TV distance, maintain local visitation ratio/coverage.
This helps identify where revalidation is needed.
114. High-debt region
If policy repeatedly approaches high-debt region, external validation value increases.
VWDC-04 can schedule targeted validation.
115. Policy deployment as active validation
Carefully bounded deployment can serve both:
- task execution;
- transport data collection.
This is dual control.
No novelty claim.
116. Data collection must be safe
Exploration for calibration does not justify hazardous actions.
Use domain constraints and human/authority gates.
117. Safe continual adaptation precedent
Current robotics work studies post-Sim2Real continual domain adaptation while minimizing safety risk and preserving previously learned safe behavior.
This supports separating adaptation from unrestricted online exploration.
118. Safe policy updates precedent
Current work also studies policy updates that preserve certified safety properties over previously encountered task distributions.
VWDC distinguishes such safety guarantees from RTC external-validity guarantees.
119. Deployment policy drift
A new policy can visit a new distribution even if environment dynamics do not change.
Thus RTC invalidation can be endogenous.
120. Environment drift
Reality itself can also change.
Need separate:
and:
121. Closed-loop regime
The complete loop is:
122. Policy update can invalidate world comparison
If policy update changes actions, old world-vs-reality trajectory comparisons may no longer match the new visitation distribution.
Version validation by policy.
123. Counterfactual policy evaluation
Before deploying a new policy, WDC can simulate it.
But simulation cannot certify unsupported reality regions without transport assumptions.
124. External off-policy data
Logged real-world data can evaluate candidate policies only under support/causal assumptions.
This is established off-policy/causal inference territory.
125. No-free-policy-transfer theorem intuition
A world policy can exploit precisely those state/action regions where the world model is wrong.
This is why pessimism/robustness becomes relevant.
VWDC's unsupported-region gate operationalizes the same concern.
126. Reward hacking analogue
A simulator-trained policy can exploit simulator artifacts to obtain high simulated value with poor reality performance.
RTC debt should increase where such exploitation is plausible/observed.
127. Model exploitation detector
Compare:
- simulated value gain;
- transport uncertainty;
- external outcomes.
Large simulation gain concentrated in high-debt regions is a warning.
128. Deployment audit
After policy rollout, report:
- predicted return;
- real observed return;
- transport residuals;
- visited unsupported mass;
- interventions/probes;
- fallbacks;
- human overrides.
129. Reality regret
Empirical reality regret is usually unknown because the unchosen real optimal action/policy is counterfactual.
Use:
- bounds;
- randomized evaluation;
- experiments;
- structural assumptions.
Do not report actual regret as observed unless identified.
130. VWDC06-N3 — One realized deployment does not identify counterfactual policy regret
Only the deployed action outcome is observed.
Without counterfactual identification, the value of unchosen actions is not directly observed.
Therefore:
This is a standard causal/off-policy boundary.
131. Regret certificate versus empirical regret
VWDC06-T2/T7 provide model/RTC-based regret bounds.
They are not direct measurements of realized counterfactual regret.
132. Policy-safe region
Define:
as states where every action selected by satisfies the required RTC/safety gates.
133. Exit condition
If reality enters:
policy automatically transitions to:
- fallback;
- human;
- probe;
- stop.
134. Fallback invariant
A safe fallback should itself have a maintained RTC/safety certificate.
Fallback is not magically safe because it is called fallback.
135. Human fallback invariant
Human authority does not eliminate reality uncertainty.
Log human decision and outcome for RTC update.
136. Deployment contract hierarchy
Suggested:
POLICY_CONTRACT
-> RTC_REWARD
-> RTC_SAFETY
-> SAFETY_FILTER
-> AUTHORITY_POLICY
-> FALLBACK_CONTRACT
137. Contract composition
A policy is deployable only if all required subordinate contracts are current.
138. VWDC06-T11 — Contract-conjunction monotonicity
Theorem VWDC06-T11
If deployment requires all contracts in set:
to pass, then adding another mandatory contract cannot enlarge the certified action set.
Proof
The feasible set is an intersection.
Adding a constraint intersects with another set.
Set intersection cannot enlarge.
139. Conservative contract expansion
Stricter governance can reduce deployable actions.
It may increase safety/trust but reduce performance/coverage.
140. Contract debt vector
Define:
141. Deployment Pareto frontier
Actions/policies can trade:
- expected reward;
- transport debt;
- safety risk;
- external probing cost;
- human burden.
Nondominated options form deployment frontier.
142. VWDC06-T12 — Deployment Pareto necessity
Theorem VWDC06-T12
Every optimum of a scalar deployment objective strictly increasing in declared costs/risks and strictly decreasing in declared benefits lies on the nondominated deployment frontier.
Proof
Standard dominance argument.
143. Current literature boundary — Sim2Real transfer
VWDC-06 does not claim as inventions:
- domain randomization;
- robust RL;
- safe RL;
- constrained MDPs;
- sim-to-real policy adaptation;
- residual policy learning;
- context inference;
- offline domain randomization.
144. Current literature boundary — pessimism
Pessimism under uncertainty/distribution shift is established in offline and robust RL.
VWDC uses RTC-supported lower bounds as the decision-governance interface.
145. Current literature boundary — safe adaptation
Safe continual adaptation and provably safe policy updates are active current research.
VWDC does not claim these techniques.
It keeps transport validity separate from formal safety guarantees.
146. Current literature boundary — digital-twin control
Real-Sim-Real and adaptive twin loops already use real feedback to improve simulators/policies.
VWDC formalizes the contract/provenance layer around these loops.
147. Candidate VWDC-specific synthesis
Subject to broader literature audit, candidate bridge-specific synthesis is:
- RTC action-value intervals used directly as deployment certificates;
- transport debt converted into explicit one-step reality-regret bounds;
- supported versus unsupported action semantics as a deployment gate;
- probe/fallback/human/stop represented as peer deployment actions;
- fixed-policy Sim2Real value discrepancy tied to RTC reward/transition gap;
- policy-induced occupancy shift treated as an endogenous RTC-certification failure mode;
- versioned RTC-policy pairs with online reality feedback;
- multi-contract gating separating reward transport, safety transport, safety filters, and authority.
No strong novelty claim is made in v0.1.
148. What VWDC-06 proves
Under explicit hypotheses, VWDC-06 proves:
- a simulated action whose lower transported value exceeds all competitors' upper transported values is reality-optimal throughout the RTC ambiguity set;
- deploying the world-optimal one-step action has real regret bounded by the transport errors of the real and simulated winners;
- maximizing RTC lower bounds yields the largest available certified action-value floor;
- a guaranteed fallback dominates the available robust lower-bound certificate when its guarantee is larger;
- unsupported action value is not identified from supported validation data alone without structural assumptions;
- a perfect probe in the two-regime risky-action model is worthwhile exactly under the stated threshold;
- fixed-policy discounted world/reality value difference is bounded by reward and transition mismatch;
- the world-optimal policy has reality regret at most twice a uniform fixed-policy transport bound;
- policy-induced occupancy shift changes expected transport loss by at most a TV-distance term;
- an RTC loses certification at a decision tolerance when its only available deployment-shift upper bound exceeds that tolerance;
- local reality reward means admit Hoeffding calibration intervals under i.i.d. bounded observations;
- reward transport validity does not imply safety transport validity;
- one realized deployment does not identify counterfactual policy regret;
- adding mandatory deployment contracts cannot enlarge the certified action set;
- every strictly monotone scalar deployment optimum lies on the nondominated deployment frontier.
149. What VWDC-06 does not prove
It does not prove:
- RTC intervals are always calibrated;
- unsupported actions are always unsafe;
- pessimistic deployment maximizes expected reward;
- a generic safety fallback is absolutely safe;
- reward and transition TV errors are easy to estimate;
- fixed-policy MDP bounds remain tight for long horizons;
- policy occupancy TV can be estimated without uncertainty;
- online adaptation remains safe without separate safety guarantees;
- deployment data are unconfounded;
- human review is infallible;
- an RTC replaces regulation or formal safety verification.
150. Proposed VWDC-07
The next paper should close the loop around deployment-triggered learning and governance:
Chinese:
閉環現實回饋、安全策略適應與轉移感知持續世界
Main questions:
- How should deployed observations update the world model and RTC jointly?
- How should policy updates preserve old certified safety/transport regions?
- When should feedback create a new world/reality regime version?
- How should old RTCs be retained/superseded?
- How should continual adaptation avoid self-confirming deployment bias?
- When should the runtime roll back a policy/world update?
- How should real incidents propagate through model, RTC, and branch evidence?
- Can a stability criterion be defined for the full World→RTC→Policy→Reality loop?
151. References
- Gengyue Han, Yiheng Feng, Transferable Reinforcement Learning via Probabilistic Latent Embeddings and Dynamic Policy Adaptation for Sim-to-Real Deployment, arXiv:2605.27659, 2026.
- Lu Shi et al., An Real-Sim-Real (RSR) Loop Framework for Generalizable Robotic Policy Transfer with Differentiable Simulation, arXiv:2503.10118, 2025.
- Mohamad H. Danesh et al., Safe Domain Randomization via Uncertainty-Aware Out-of-Distribution Detection and Policy Adaptation, arXiv:2507.06111, 2025.
- Josip Josifovski et al., Safe Continual Domain Adaptation after Sim2Real Transfer of Reinforcement Learning Policies in Robotics, arXiv:2503.10949, 2025.
- Maksim Anisimov, Francesco Belardinelli, Matthew Wicker, SafeAdapt: Provably Safe Policy Updates in Deep Reinforcement Learning, arXiv:2604.09452, 2026.
- Zeyuan Tang et al., Uncertainty-Aware Robotic World Model Makes Offline Model-Based Reinforcement Learning More Powerful, arXiv:2504.16680, revised 2026.
- Provable Sim-to-Real Transfer via Offline Domain Randomization, arXiv:2506.10133, 2025.
- The Sim-to-Real Gap of Foundation Model Agents, arXiv:2606.07017, 2026.
- Can Context Bridge the Reality Gap? Sim-to-Real Transfer through Dynamics-Grounded Context Adaptation, arXiv:2511.04249, revised 2026.
- VWDC-01–05, GVSS-01–10, WDC-01–08, WDC Runtime Whitepaper, and frozen RRT-20, internal series artifacts, 2026.
152. Conclusion
VWDC-05 creates a reality transport contract.
VWDC-06 makes deployment obey it.
The central one-step certificate is:
When it holds, the simulation choice survives every reality value allowed by the RTC.
When it does not, transport debt becomes a real decision ambiguity.
Even then the simulated optimum has a reality-regret bound:
Unsupported actions have no distribution-free reality guarantee.
A probe can be worth more than immediate deployment or fallback near the decision boundary.
Long-horizon policy transfer amplifies reward/dynamics mismatch.
And deployment can invalidate its own RTC by changing the state-action occupancy distribution:
Therefore policy deployment is not the end of validation.
It changes the reality data distribution and feeds new evidence back into the world and transport layers.
The canonical VWDC-06 principle is:
This establishes transport-aware decision and policy control for reality-facing WDC systems.
Canonical-source policy
This file is the canonical UTF-8 source artifact.
- Canonical inline mathematics uses
$...$. - Canonical display mathematics uses
$$...$$. - No Unicode mathematical-symbol conversion is used as source normalization.
- No
unicode_escaperound trip is used. - Backslashes and delimiters are preserved literally.
- Validation is required before release.
- This paper does not merge or rename GVSS, WDC, VWDC, or RRT.