VWDC-05 — Reality-Tethered World Calibration, Transport Contracts, and External Validity
現實繫結世界校準、轉移契約與外部有效性:條件驗證、模型差異、版本漂移與反事實假設邊界
Bridge Series: Visual–World Domain Computation (VWDC) — Paper 05
Depends on: VWDC-01–04, GVSS-01–10, WDC-01–08, WDC Runtime Whitepaper, frozen RRT-20
Author: Neo.K / EveMissLab
Version: v0.1
Date: 2026-08-17
Status: Formal transport-calibration paper. Contract localization, marginal-cancellation no-go, conditional-discrepancy aggregation, partial-coverage transport bounds, counterfactual joint non-identifiability, parameter–discrepancy confounding, inverse-variance external evidence fusion, common-bias fusion no-go, distribution-shift transport bounds, version-drift invalidation, calibration/transport error composition, and claim-scope monotonicity are proved under explicit hypotheses. Causal transportability, digital-twin calibration, model discrepancy analysis, subtrace-conditional validation, and counterfactual validation are established neighboring research and are not claimed as VWDC inventions. No strong novelty claim is made.
Keywords: digital twin calibration, transportability, external validity, model discrepancy, conditional validation, counterfactual validation, sim-to-real, reality gap, transport contract, world model, WDC
Abstract
VWDC-03 and VWDC-04 separated two error classes for a world-derived reality claim:
where:
- is world-internal estimation/simulation uncertainty;
- is world-to-reality transport discrepancy.
VWDC-04 proved that internal branching can drive:
while leaving:
unchanged.
Therefore simulation precision alone cannot certify external validity.
VWDC-05 makes the transport layer itself a first-class, calibrated, versioned object.
The central object is a Reality Transport Contract:
where:
- is the exact world/reality target quantity or estimand;
- is the state/action/task validity region;
- is the world-model/runtime version;
- is the reality-regime/measurement version;
- is the world-to-reality transport map;
- is validation level/protocol;
- is the explicit assumption set;
- is the residual transport-discrepancy bound or uncertainty object;
- is external validation evidence;
- specifies drift/revalidation conditions;
- is provenance.
The contract does not say:
this digital twin is valid.
It says:
this particular world-derived quantity may be transported to this particular reality target, over this declared region, under these assumptions, with this residual discrepancy and this version scope.
That distinction is the foundation of VWDC-05.
1. Target-quantity relativity
A world model can be accurate for one quantity and inaccurate for another.
Let:
be a world-derived quantity.
Let:
be the corresponding reality quantity.
A transport map is:
The subscript is essential.
There is no requirement that one transport map validates all world outputs.
2. Reality Transport Contract
Definition VWDC05-D1
A transport contract is:
The contract is valid only inside all declared scopes.
3. Scope dimensions
At minimum record:
claim_or_estimand
world_model_version
world_runtime_version
reality_regime
measurement_process
state_region
action_region
task_region
validation_level
assumption_set
residual_discrepancy
validation_data
expiry_condition
4. Global-twin validity no-go
VWDC05-N1 — No scalar twin-validity score determines all transport claims
Counterexample
A simulator exactly predicts position:
but systematically mispredicts temperature:
A single scalar "twin accuracy" cannot imply both quantity-specific transport guarantees.
Therefore:
Transport validity is quantity and region relative.
5. Validation region
Let:
denote a task/state/action context.
Let world and reality conditional target means be:
Define local transport discrepancy:
6. Marginal discrepancy
Under reality target distribution:
marginal means are:
7. VWDC05-N2 — Marginal validation can hide conditional transport failure
Counterexample
Let:
Reality:
World:
Then:
Marginal mean error is zero.
But conditional errors are:
Thus:
This is the core reason to validate local/conditional behavior rather than only aggregate outputs.
8. Conditional validation precedent
Current subtrace-conditional validation research repeatedly initializes simulations from observed system states, fixes selected stochastic primitives to observed realizations, and validates conditional output distributions.
It is explicitly designed to detect misspecified input models that can be invisible under marginal-output validation.
VWDC adopts this as a direct validation precedent.
9. Conditional discrepancy aggregation
Suppose a finite partition:
Reality target weights:
Local signed discrepancies:
10. VWDC05-T1 — Aggregate transport discrepancy bound
Theorem VWDC05-T1
Proof
Triangle inequality and convexity of the weighted average.
11. Interpretation
Marginal error can cancel.
A local maximum bound cannot.
Therefore transport validation should report at least:
- signed marginal discrepancy;
- weighted absolute discrepancy;
- worst validated local discrepancy.
12. Local transport debt
Define:
Do not collapse this object too early.
13. Partial validation coverage
Let:
be validated cells.
Let reality target probability mass outside covered cells be:
Assume target discrepancy/loss is bounded in:
14. VWDC05-T2 — Partial-coverage transport bound
Theorem VWDC05-T2
If for each covered cell:
transport discrepancy is bounded by:
then total expected discrepancy satisfies:
Proof
Split expectation over covered and uncovered regions.
Use known cell bounds on and the global boundedness outside .
15. Coverage debt
Define:
A highly accurate contract over a tiny region can still have large target-domain debt.
16. Validation coverage map
Recommended:
region_id
state_scope
action_scope
task_scope
target_mass
sample_count
local_discrepancy
confidence
validation_level
support_status
17. Validation levels
VWDC proposes the following engineering ladder.
V0 INTERNAL_ONLY
V1 MARGINAL_EXTERNAL
V2 CONDITIONAL_SUBTRACE
V3 INTERVENTIONAL
V4 STRUCTURAL_COUNTERFACTUAL
These are VWDC labels, not a claim to replace existing validation taxonomies.
18. V0 — Internal only
Checks:
- code;
- transition invariants;
- replay;
- internal consistency.
No external validity claim.
19. V1 — Marginal external
Checks:
or selected marginal moments/distributions.
Useful but can hide conditional mismatch.
20. V2 — Conditional/subtrace
Checks:
against observed conditional behavior.
This can localize discrepancy.
21. V3 — Interventional
Checks world predictions against observed effects of real interventions/actions in matching scope.
Stronger for intervention claims.
22. V4 — Structural/counterfactual
Claims about unobserved joint potential outcomes, latent couplings, or structural mechanisms.
Some components can remain assumption indexed even after strong marginal/interventional validation.
23. Counterfactual estimands
Let binary potential outcomes:
Suppose both marginals are known exactly:
24. VWDC05-N3 — Correct potential-outcome marginals do not identify individual benefit probability
Counterexample
World A — perfectly coupled
Then:
Both marginals are Bernoulli .
World B — opposite coupling
Then:
Again both marginals are Bernoulli .
Therefore the two potential-outcome marginals do not identify:
25. Counterfactual assumption label
Joint/counterfactual claims should record:
MARGINAL_IDENTIFIED
JOINT_ASSUMPTION_INDEXED
STRUCTURAL_ASSUMPTION_INDEXED
SENSITIVITY_BOUNDED
Do not report an assumption-indexed quantity as empirically identified.
26. DTCF relation
Current Digital Twin Counterfactual Framework work distinguishes marginally testable causal quantities from copula/joint-dependent counterfactual quantities whose identification still relies on unobservable within-unit dependence assumptions.
VWDC uses the same boundary.
27. Model calibration
A simulator often has calibration parameters:
and model discrepancy:
Observed reality data can be represented schematically as:
28. VWDC05-N4 — Parameter–discrepancy confounding
Counterexample
Let:
and noiseless observation:
Then every pair:
for any real fits the observation exactly.
Therefore:
without further constraints.
This is a minimal instance of calibration/discrepancy confounding.
29. Current online calibration precedent
Recent 2026 Bayesian digital-twin calibration work explicitly addresses model discrepancy, parameter–discrepancy confounding, gradual drift, and abrupt regime changes.
It reinforces the need to version and continuously revalidate calibration contracts.
30. Calibration target
VWDC requires calibration objective to state whether it targets:
- predictive accuracy;
- physical parameter interpretation;
- intervention response;
- transport bound;
- latent-state inference.
A parameter value calibrated for prediction need not be physically identifiable.
31. External validation dataset
Let external datasets be:
Each has:
- measurement process;
- target region;
- time;
- bias assumptions;
- variance;
- provenance.
32. Independent unbiased scalar estimators
Suppose dataset yields:
with:
and variance:
Assume estimators are independent.
33. VWDC05-T3 — Minimum-variance linear external-evidence fusion
Theorem VWDC05-T3
Among unbiased linear estimators:
the minimum-variance weights are:
The fused variance is:
Proof
Lagrange multiplier minimization of:
subject to:
This is classical inverse-variance weighting.
34. Shared external bias
Suppose:
with common bias:
Then averaging many datasets reduces:
but not .
35. VWDC05-N5 — Many external datasets do not remove shared measurement bias
If:
then even as:
the fused estimator remains biased by:
Therefore:
The same dependence logic used for world branches applies to external measurements.
36. External evidence independence profile
Track:
- measurement device family;
- institution/source;
- preprocessing;
- calibration standard;
- time period;
- target population;
- evaluator/labeler;
- shared pipeline.
37. Reality regime
Let:
identify the target reality regime.
Examples:
- operating season;
- hardware configuration;
- patient population;
- policy environment;
- sensor calibration state.
Transport validity is regime relative.
38. Distribution drift
Let old validation context distribution be:
and current target distribution:
Let discrepancy function:
39. VWDC05-T4 — Distribution-shift transport bound
Theorem VWDC05-T4
Using total variation convention:
for:
Proof
Standard bounded-function total-variation inequality.
40. Drift-adjusted discrepancy
If old validated expected discrepancy is:
then current expected discrepancy obeys:
provided the discrepancy function itself remains unchanged.
41. Structural drift caveat
If the conditional discrepancy mechanism also changes:
covariate-distribution TV alone is insufficient.
Require model/regime drift testing.
42. World-model version drift
A new world-model version:
can change:
- predictions;
- latent representation;
- action semantics;
- calibration parameters;
- discrepancy pattern.
An old transport contract is not automatically inherited.
43. Reality-regime drift
Likewise:
can invalidate transport without any world-model update.
44. VWDC05-N6 — Validation of one version does not validate another version without an invariance argument
Counterexample
Old model:
matches reality exactly on validation region.
New model:
returns the negated prediction.
The old validation data contain no evidence that:
is valid.
Therefore:
45. Contract status
Recommended:
VALID
CONDITIONALLY_VALID
STALE
REVALIDATION_REQUIRED
INVALID
SUPERSEDED
46. Expiry triggers
A transport contract should expire or enter review when:
- world-model version changes;
- evaluator changes;
- measurement process changes;
- target distribution drifts beyond threshold;
- reality regime changes;
- validation age exceeds policy;
- anomaly/drift test rejects stationarity.
47. Continual validation precedent
Current adaptive digital-twin work explicitly combines drift detection, model updating, and statistical validation to decide when and whether a digital twin update improves predictive performance.
This is direct precedent for contract expiry/revalidation.
48. World estimate and transport
Let:
estimate world target:
Assume:
Let transport map be:
-Lipschitz.
Transport discrepancy:
49. External measurement/calibration error
Suppose reality target estimate:
satisfies:
50. VWDC05-T5 — Three-part calibrated claim bound
Theorem VWDC05-T5
Proof
Insert:
and:
and apply triangle inequality plus Lipschitzness.
51. Three error classes
The bound separates:
More simulation primarily attacks the first.
Better transport calibration attacks the second.
Better external measurement attacks the third.
52. Measurement precision no-go
Perfect external sensor precision:
does not remove transport/model discrepancy.
Likewise perfect simulation precision does not remove external measurement bias.
53. Regional transport map
A practical contract can be:
Different state/action/task regions can require different corrections.
54. Local discrepancy map
Store:
This is the transport analogue of GVSS provider capability maps.
55. Local unsupported region
If no external validation exists in region:
status is:
UNSUPPORTED
unless an explicit transfer assumption links it to validated regions.
56. VWDC05-N7 — Unsupported region transport is not empirically identified by neighboring validation alone
Without smoothness/invariance assumptions, two reality systems can agree on all validated regions and differ arbitrarily on an unvalidated region.
Therefore:
57. Spatial/task smoothness
A model can extrapolate transport discrepancy across nearby regions if assuming:
This assumption must be recorded and validated where possible.
58. VWDC05-T6 — Lipschitz regional extrapolation bound
Theorem VWDC05-T6
If:
is -Lipschitz over context metric , and region has nearest validated point , then:
Proof
Lipschitz inequality.
This converts geometric coverage gaps into explicit extrapolation debt.
59. Validation density
Define:
Then extrapolation debt grows with:
60. Validation design
VWDC-04 can now choose external validation actions where:
is high.
61. Subtrace validation localization
Observed checkpoints/subtraces can define regions:
Conditional tests update only the contract cells they actually validate.
Do not globally upgrade the twin because one subtrace passed.
62. Marginal test pass
A passed V1 test can update:
MARGINAL_VALIDATED
not:
STRUCTURALLY_VALIDATED
63. Conditional test pass
A passed V2 test can reduce local discrepancy in its conditional scope.
It does not identify unobserved joint potential-outcome couplings.
64. Interventional test pass
A real intervention can validate:
against world predictions over matched scope.
Still do not extrapolate to unmatched actions/states without transport assumptions.
65. Structural claim
A structural model asserts invariances/mechanisms beyond directly observed distributions.
These assumptions can support transport.
They should remain explicit.
66. Classical causal transportability boundary
Causal transportability theory uses structured knowledge about differences between source and target populations/environments to decide whether causal effects can be transported and what data are required.
VWDC does not claim causal transportability as new.
The Reality Transport Contract is a runtime engineering wrapper around claim-specific transport assumptions, validation, discrepancy, and versioning.
67. Selection diagram relation
A future implementation can attach a causal/selection diagram to:
This can encode which mechanisms differ between world and reality target.
VWDC-05 does not develop do-calculus.
68. Transport contract acceptance
A claim transport is accepted only if:
- target quantity is named;
- source/target semantics align;
- state/action/task region is supported;
- validation level meets policy;
- residual discrepancy is below claim tolerance;
- structural assumptions are accepted;
- versions are current;
- provenance is complete.
69. Claim tolerance
Let downstream decision tolerate error:
Transport is decision-admissible when:
This does not mean the model is "true."
It means residual debt is within declared decision tolerance.
70. VWDC05-T7 — Claim-scope restriction cannot increase worst-case validated discrepancy
Theorem VWDC05-T7
Let:
be two validity regions.
Define:
Then:
Proof
Supremum over a subset cannot exceed supremum over the full set.
Narrower claim scope can support a stronger guarantee.
71. Scope expansion debt
Expanding a contract to more:
- states;
- actions;
- populations;
- time regimes;
requires new evidence or stronger assumptions.
72. Generalization label
Recommended:
DIRECTLY_VALIDATED
INTERPOLATED
EXTRAPOLATED
STRUCTURALLY_TRANSPORTED
UNSUPPORTED
73. Contract inheritance
A child contract can inherit validated information from a parent only if:
- same quantity semantics;
- compatible versions;
- subset region;
- no weaker validation requirement;
- discrepancy bound remains valid.
74. VWDC05-T8 — Safe subset inheritance
Theorem VWDC05-T8
If contract is valid on region with uniform discrepancy:
then for any:
the same discrepancy bound remains valid on .
Proof
Immediate set restriction.
75. Superset inheritance no-go
Validation on:
does not establish the same bound on:
This is the external-validity analogue of unsupported routing regions in GVSS-10.
76. Calibration artifact
Define:
77. Contract artifact
Define:
78. Contract fingerprint
Hash:
- world model;
- world runtime;
- external datasets;
- measurement process;
- target quantity;
- state/action/task region;
- validation protocol;
- assumption set;
- discrepancy model.
79. Contract versioning
Never overwrite:
Create:
and retain parent/supersession relation.
80. Contract expiry
Possible rules:
ON_WORLD_MODEL_CHANGE
ON_REALITY_REGIME_CHANGE
ON_MEASUREMENT_CHANGE
ON_DRIFT_TEST
AFTER_TIME_WINDOW
ON_VALIDATION_FAILURE
81. Online calibration
Streaming external data can update:
The contract remains valid only if update/validation policy passes.
82. Gradual drift
Use forgetting/windowing/state-space discrepancy models if stationarity fails gradually.
83. Abrupt drift
A changepoint/reset policy may create a new reality regime:
Do not force new data into the old contract.
84. Continual calibration boundary
Online Bayesian calibration under gradual/abrupt drift is established current research.
VWDC uses the output as versioned calibration evidence.
85. Calibration does not equal validation
Fitting the twin to observed data uses data to reduce discrepancy.
Validation should reserve independent or appropriately corrected evidence where possible.
Otherwise calibration and validation can become circular.
86. VWDC05-N8 — Perfect in-sample calibration does not prove out-of-sample transport
Counterexample
Fit arbitrary interpolator to finite validation points with zero residual error.
Choose a target point outside those points where the interpolator differs arbitrarily from reality.
Therefore:
87. Holdout / prospective evidence
When possible, use:
- holdout external data;
- future temporal data;
- independent sensors/sites;
- interventions;
to challenge the transport contract.
88. Reality Gap Analysis relation
Recent digital-twin reality-gap work explicitly treats continuous integration of new sensor data, context mismatch detection, and recalibration across a twin lifecycle.
VWDC places such modules inside the transport contract lifecycle.
89. Simulation-based inference correction relation
Recent work on simulation-based inference under model misspecification uses scarce calibration observations to transport/correct simulation-trained posterior estimates toward reality-supported posteriors.
This is another current example of explicit sim-to-real correction rather than assuming synthetic and real distributions match.
90. Model discrepancy map
Possible representation:
The model form can be:
- Gaussian process;
- neural residual;
- basis expansion;
- piecewise constant cells;
- robust interval.
VWDC does not prescribe one.
91. Identifiability profile
Each calibrated parameter/claim should record:
DIRECTLY_IDENTIFIED
IDENTIFIED_UP_TO_EQUIVALENCE
PRIOR_REGULARIZED
DISCREPANCY_CONFOUNDED
ASSUMPTION_INDEXED
92. Parameter meaning versus prediction
A calibration can improve prediction while making latent parameter interpretation unreliable.
Separate:
from:
93. Multiple fidelity sources
External evidence can include:
- real measurement;
- hardware-in-loop;
- high-fidelity simulator;
- low-fidelity simulator;
- expert annotation.
Their evidential status is not identical.
94. Evidence fidelity label
REAL_MEASUREMENT
HARDWARE_IN_LOOP
HIGH_FIDELITY_SIM
LOW_FIDELITY_SIM
EXPERT_LABEL
DERIVED
95. High-fidelity simulation is not reality
Even a more expensive simulator remains a model unless directly tethered to external measurement.
96. Measurement model
Reality observation itself may be:
Transport validation therefore compares two modeled observation processes, not omniscient reality state.
97. Sensor calibration
Measurement uncertainty belongs in:
A bad sensor can make a good twin look wrong or a bad twin look correct.
98. Measurement-process version
If sensor calibration/process changes:
the transport contract requires review.
99. External-label evaluator
Human/expert labels are measurement processes with:
- disagreement;
- bias;
- protocol;
- version.
Do not treat them as perfect truth automatically.
100. Target-semantic mismatch
World and reality quantities must have matching semantics.
Example:
- simulator "collision" event;
- real-world safety incident.
If definitions differ, numerical agreement is not transport.
101. VWDC05-N9 — Numeric agreement without semantic alignment does not establish transport
Two quantities can share the same numeric values while denoting different events/constructs.
Therefore:
102. Semantic contract
Record:
world_quantity_definition
reality_quantity_definition
units
time_window
aggregation_rule
population
intervention_semantics
103. Units
All transport maps should state unit conversion.
Silent unit mismatch is contract failure.
104. Time alignment
World local time:
must map to reality time:
under an explicit synchronization rule.
105. Intervention alignment
World action:
and real intervention:
must be semantically aligned before an interventional transport claim.
106. Policy transport
A world-derived optimal policy may fail in reality even if one-step outcome predictions calibrate well.
Policy transport requires sequential/state-distribution validation.
107. Distribution shift under deployed policy
Deploying a policy changes the visited reality distribution.
A contract validated under historical policy may face new state distribution after deployment.
108. Closed-loop transport
Transport contracts for control should include:
- policy;
- induced state distribution;
- action support;
- feedback effects.
109. Off-policy support relation
If a real action/state region was never externally observed, policy effects there are unsupported without structural assumptions.
This mirrors contextual-bandit positivity.
110. VWDC05-N10 — Observational support gaps block nonparametric interventional validation
If real validation data contain no support for action in state region , then without causal/structural assumptions the conditional interventional response in that region is not identified from those observational data alone.
This is a classical positivity/causal-identification boundary.
111. Transportability literature
Causal transportability formalizes when experimental findings can be moved across populations using explicit structural knowledge about which mechanisms differ.
VWDC inherits the principle:
transport requires assumptions about invariance/difference, not merely source accuracy.
112. Contract graph
A runtime can store:
whose nodes are transport contracts and whose edges are:
- derived from;
- supersedes;
- subset of;
- structurally transported from;
- invalidated by.
113. Contract dependency
If world model:
is invalidated, every RTC depending on enters review.
This uses VWDC-03 dependency invalidation.
114. Contract replay
After recalibrating a world model, validation tests can be replayed from archived external evidence when semantics remain compatible.
New reality data may still be required for current-regime validity.
115. Validation provenance
Every validation result records:
dataset
data_time
data_domain
measurement_version
world_model_version
validation_code_version
test_statistic
threshold
result
scope
116. Contract audit
An auditor should be able to answer:
- What exact reality claim is allowed?
- In what region?
- Under which model version?
- Which external data support it?
- What assumptions remain?
- What discrepancy remains?
- What invalidates the contract?
117. Allowed claim scope
Examples:
WORLD_ONLY
REALITY_MARGINAL
REALITY_CONDITIONAL
REALITY_INTERVENTIONAL
COUNTERFACTUAL_ASSUMPTION_INDEXED
118. Claim renderer
User-facing output should display scope.
Example:
Validated for one-step marginal prediction on region Z under model v3 and sensor process s2; intervention and individual counterfactual claims are unsupported.
119. Do not hide assumption-indexed claims
A numeric output can be accompanied by:
assumption_set:
- shared_latent_rank_invariance
- copula_family_gaussian
rather than presenting it as empirically identified.
120. Sensitivity interval
For assumption-indexed counterfactual quantity:
report:
This makes structural dependence explicit.
121. External validation branch
VWDC-04 selects external experiments.
VWDC-05 updates:
with their results.
122. Calibration/validation loop
123. Reality-tether strength
A contract can expose a vector:
Do not collapse unless a policy demands it.
124. Recency
Older validation can become less relevant under nonstationary reality.
Track age and drift indicators.
125. Region-weighted current validity
For deployment distribution:
compute weighted current debt over cells rather than raw average across validation dataset.
126. Deployment-shift audit
Compare:
and:
Large difference increases reweighting/extrapolation debt.
127. Calibration sample selection
VWDC-04 can actively choose external measurements in high-debt/high-mass regions.
This becomes active reality tethering.
128. Reality-tethered world
Definition VWDC05-D2
A world instance/model is reality tethered for claim if there exists a current:
whose allowed scope contains the intended claim/deployment context.
This is claim relative.
129. Not globally tethered
A world can be tethered for:
- temperature forecast;
and untethered for:
- failure counterfactual.
No contradiction.
130. Twin promotion
Possible runtime status:
SIMULATION_ONLY
EXTERNALLY_COMPARED
REALITY_TETHERED_MARGINAL
REALITY_TETHERED_CONDITIONAL
REALITY_TETHERED_INTERVENTIONAL
STRUCTURAL_COUNTERFACTUAL_ASSUMPTION_INDEXED
131. Digital twin naming boundary
VWDC does not legislate who may use the industry term "digital twin."
It defines stricter runtime evidence statuses for WDC claims.
132. Calibration versus reality tether
A calibrated twin is not necessarily transport-valid for every target claim.
Tether requires claim-specific external validity.
133. Decision-support threshold
If transport debt is below:
claim can be admitted for that decision support policy.
Higher-stakes policy can demand smaller:
134. Risk-tiered transport
Recommended:
LOW_STAKES
OPERATIONAL
HIGH_STAKES
SAFETY_CRITICAL
Each can impose different validation/transport thresholds.
135. Safety-critical boundary
A formal RTC is not a substitute for domain-specific regulatory validation, safety engineering, or legal requirements.
136. Benchmark A — marginal cancellation
Use the two-stratum counterexample.
Verify marginal error zero while conditional errors are nonzero.
137. Benchmark B — partial coverage
Leave target mass unvalidated.
Verify uncovered mass contributes explicit worst-case debt.
138. Benchmark C — counterfactual coupling
Construct two joint potential-outcome laws with identical marginals and different probability of benefit.
139. Benchmark D — calibration confounding
Use:
Verify nonunique decomposition.
140. Benchmark E — external fusion
Simulate independent estimators with known variance.
Verify inverse-variance weighting.
141. Benchmark F — shared external bias
Add common bias.
Verify averaging does not remove it.
142. Benchmark G — distribution shift
Create bounded discrepancy function and two context distributions.
Verify total-variation transport bound.
143. Benchmark H — version invalidation
Validate version 0.
Swap to deliberately incorrect version 1.
Ensure RTC enters REVALIDATION_REQUIRED.
144. Benchmark I — regional extrapolation
Use Lipschitz discrepancy function.
Validate sparse grid.
Check nearest-neighbor extrapolation bound.
145. Benchmark J — three-part error
Inject known:
- world error;
- transport discrepancy;
- measurement error.
Verify triangle/Lipschitz bound.
146. Benchmark K — semantic mismatch
Create two numerically identical but differently defined outcome variables.
Ensure transport rejected.
147. Benchmark L — support gap
Remove real observations for one action/state cell.
Ensure status is UNSUPPORTED absent structural assumptions.
148. Current literature boundary — transportability
Causal transportability and external-validity theory are established.
VWDC does not claim selection diagrams, do-calculus, or transport formulas as new.
149. Current literature boundary — Bayesian calibration
Computer-model calibration and model-discrepancy methods are established.
VWDC does not claim calibration/discrepancy modeling as new.
150. Current literature boundary — digital twin validation
Digital-twin validation, continual updating, uncertainty quantification, and sim-to-real gap analysis are established active research areas.
VWDC provides a runtime contract vocabulary connecting them to branching-world evidence.
151. Candidate VWDC-specific synthesis
Subject to broader literature audit, the bridge-specific synthesis is:
- making world-to-reality transport a first-class versioned contract attached to a specific claim/estimand;
- localizing transport validity over state/action/task regions rather than using a global twin-validity label;
- integrating marginal, conditional/subtrace, interventional, and structural/counterfactual validation into explicit claim scopes;
- separating world-estimation, transport, and measurement error in one runtime claim bound;
- representing unsupported validation mass and regional extrapolation as explicit transport debt;
- attaching contract expiry to both world-model and reality-regime drift;
- carrying joint/counterfactual non-identifiability as an assumption-indexed status rather than an implicit simulator output;
- connecting VWDC-04 external-experiment selection to a persistent reality-tether contract lifecycle.
No strong novelty claim is made in v0.1.
152. What VWDC-05 proves
Under explicit hypotheses, VWDC-05 proves:
- marginal agreement can coexist with arbitrarily important conditional mismatch;
- aggregate signed transport discrepancy is bounded by weighted absolute and maximum local discrepancy;
- partial validation coverage yields an explicit uncovered-mass worst-case term;
- identical potential-outcome marginals do not identify individual benefit probability;
- calibration parameters and model discrepancy can be non-identifiable from field fit alone;
- inverse-variance weights minimize variance among independent unbiased linear external-evidence fusions;
- adding many external datasets does not remove shared measurement bias;
- expected bounded discrepancy under distribution shift changes by at most when the discrepancy function itself is stable;
- validation of one world/reality version does not imply validation of another without an invariance argument;
- world-estimation, transport, and reality-measurement error compose additively under Lipschitz/triangle assumptions;
- unsupported regions are not empirically identified without transfer assumptions;
- Lipschitz regional discrepancy yields an explicit nearest-validation extrapolation bound;
- restricting claim scope cannot increase worst-case discrepancy;
- a uniform validated bound safely inherits to subsets of the validated region;
- perfect in-sample calibration does not prove out-of-support external validity;
- numeric agreement without semantic alignment does not establish valid transport;
- observational support gaps block nonparametric intervention validation without additional assumptions.
153. What VWDC-05 does not prove
It does not prove:
- one transport map works for every quantity;
- conditional validation identifies every latent mechanism;
- interventional validation identifies all individual counterfactuals;
- model-discrepancy functions are identifiable without assumptions;
- external datasets are unbiased or independent;
- total-variation covariate shift captures structural drift;
- Lipschitz discrepancy is valid in every domain;
- continual calibration automatically preserves external validity;
- a digital twin satisfying one RTC is globally reality-valid;
- a VWDC transport contract replaces domain regulation or safety validation.
154. Proposed VWDC-06
The next paper should use the calibrated transport layer inside decision and control:
Chinese:
轉移感知世界決策、策略搬運與現實差距穩健控制
Main questions:
- When may a policy optimized in WDC be deployed to reality?
- How should local transport debt alter action choice?
- What is robust policy regret under bounded reality gap?
- How should unsupported action/state regions trigger safe fallback or data collection?
- How should reality feedback update policy and transport contract jointly?
- When does policy deployment shift invalidate its own validation distribution?
- Can a transport-aware Governor choose between simulation policy, conservative policy, human review, and external probe?
- What guarantees remain under partial transport validity?
155. References
- Olav Laudy, The Digital Twin Counterfactual Framework: A Validation Architecture for Simulated Potential Outcomes, arXiv:2604.01325, 2026.
- Mohammadmahdi Ghasemloo, David J. Eckman, Yaxian Li, Subtrace-Conditional Validation of Simulation Models and Digital Twins, arXiv:2607.17088, 2026.
- Judea Pearl, Elias Bareinboim, External Validity: From Do-Calculus to Transportability Across Populations, arXiv:1503.01603.
- Yanqi Xu et al., Online Bayesian Calibration under Gradual and Abrupt Changes, arXiv:2605.06612, 2026.
- Pierre-Louis Ruhlmann et al., Flow Matching Calibration for Simulation-Based Inference under Model Misspecification, arXiv:2509.23385, revised 2026.
- Sizhe Ma, Katherine A. Flanigan, Mario Bergés, Bridging the Reality Gap in Digital Twins with Context-Aware, Physics-Guided Deep Learning, arXiv:2505.11847, 2025.
- Clement Ruah et al., How to Bridge the Sim-to-Real Gap in Digital Twin-Aided Telecommunication Networks, arXiv:2507.07067, 2025.
- A Continual Validation, Updating, and Decision-Making Framework for Adaptive Digital Twins, arXiv:2607.18164, 2026.
- VWDC-01–04, GVSS-01–10, WDC-01–08, WDC Runtime Whitepaper, and frozen RRT-20, internal series artifacts, 2026.
156. Conclusion
VWDC-04 asks which external validation experiment deserves budget.
VWDC-05 turns the result into a durable, scoped, versioned Reality Transport Contract.
The contract is not:
the twin is valid.
It is:
Marginal agreement can conceal conditional failure.
Conditional agreement does not identify every counterfactual joint quantity.
Perfect calibration fit can confound physical parameters with model discrepancy.
More external datasets do not remove shared measurement bias.
Old validation does not automatically survive model or reality drift.
And the complete comparison to reality contains at least three distinct error terms:
The canonical VWDC-05 principle is:
This establishes the reality-tether layer required before WDC policies can be responsibly transported beyond their simulated worlds.
Canonical-source policy
This file is the canonical UTF-8 source artifact.
- Canonical inline mathematics uses
$...$. - Canonical display mathematics uses
$$...$$. - No Unicode mathematical-symbol conversion is used as source normalization.
- No
unicode_escaperound trip is used. - Backslashes and delimiters are preserved literally.
- Validation is required before release.
- This paper does not merge or rename GVSS, WDC, VWDC, or RRT.