GVSS-04 — Reflexive Visual Navigation
反身視覺空間導航:自適應表示、閉環約束搜尋與可達性感知生成
Series: Global Visual Space & Generative Navigation — Paper 04
Bridge: GVSS × Reflexive Representation Theory (RRT)
Author: Neo.K / EveMissLab
Version: v0.1
Date: 2026-08-17
Status: Formal bridge paper. The abstract state-space, reachability, action-separation, vector-defect, protected-regression, representation-chart, evaluator-gating, lineage, and Pareto statements are proved under the explicit hypotheses below. Closed-loop text-to-image refinement, diffusion guidance/control, multimodal verification, test-time search, and agentic image generation are established neighboring research directions and are not claimed as GVSS inventions. No strong novelty claim is made.
Keywords: Global Visual State Space, reflexive visual navigation, closed-loop generation, constraint compiler, generative reachability, test-time refinement, visual verifier, diffusion control, agentic image generation, representation reflexivity, AI art direction
Abstract
The first three papers of the Global Visual Space & Generative Navigation series established three different objects.
Paper 01 defines the complete finite raster state space
for a fixed digital image specification
Paper 02 interprets visual generation as constrained navigation:
Paper 03 separates complete representability from practical generative reachability and introduces an operational reachable domain
under bounded search resources.
GVSS-04 adds the missing fourth layer:
How should an intelligent controller change its own visual representation, constraint program, generator binding, workflow, observer, verifier, and search budget in response to the images that its current regime produces?
The central object is a visual generation regime
where:
- is the bound generator/model stack;
- is the current intent compiler;
- is the compiled constraint program;
- contains constraint weights;
- contains references, adapters, LoRAs, controls, style providers, and related bindings;
- is the search/sampling/refinement policy;
- is the visual observer/evaluation interface;
- is the verifier/evaluator configuration;
- is the remaining resource budget;
- is the lineage/provenance state.
Generation produces
Evaluation returns the existing eight-axis runtime vector
corresponding to:
- Prompt / Constraint Adherence;
- Technical Image Quality;
- Human Preference / Aesthetic Proxy;
- Style Consistency;
- Diversity / Coverage;
- Anti-Homogenization;
- Character / Subject Consistency;
- Reference / Control Consistency.
Given target thresholds
define the deficit vector
A reflexive visual controller chooses an action
from an action vocabulary containing, for example,
The regime then updates:
This is the GVSS specialization of representation reflexivity.
The first structural theorem separates three commonly conflated forms of refinement.
Let
denote the base reachable set of a fixed generator/provider binding.
Let
denote the accepted visual set defined by a constraint program.
Define the effective target domain
Then, under fixed binding semantics:
- RESAMPLE changes the sampled trajectory but leaves both and unchanged;
- RECOMPILE can change while leaving the base generator reachable set unchanged;
- REBIND can change itself.
Thus:
This is the central diagnostic separation of GVSS-04.
Practical reachability is made policy dependent.
For controller/search policy and budget , define
as the set of images that can be produced or reached by the declared policy within the resource budget.
If:
and unused budget may simply be discarded, then
If policy can simulate every bounded execution of , then
Therefore practical reachable visual space is not a property of the foundation generator alone.
It depends on:
The closed-loop error geometry is represented by a vector recurrence
where is a nonnegative action-dependent transfer matrix and is newly injected model/compiler/evaluator defect.
Unrolling yields
with the empty matrix product interpreted as identity.
Thus a refinement action can reduce one deficit while transporting or amplifying another.
This gives a formal version of the existing runtime rule that one metric improvement must not silently destroy protected dimensions.
If accepted refinement steps satisfy
for a protected metric coordinate , then after accepted refinement steps
The paper also formalizes the role of the Global Style Map.
Any finite-dimensional style map is a representation chart
A Euclidean style distance
is not automatically intrinsic.
For an invertible but non-isometric coordinate change
generally
Hence:
Style coordinates are navigation representations, not declarations of the unique geometry of visual meaning.
Evaluation itself also changes practical reachability.
For verifier and threshold , define
The verifier-gated reachable set is
If
then:
A stricter verifier reduces accepted reachability.
But a stricter or more strongly optimized verifier does not necessarily improve true human intent if the evaluator is misaligned.
A two-image counterexample is sufficient:
but proxy evaluator
Optimizing selects the worse image according to true utility .
Therefore:
The loop is only as valid as the representation, evaluator, and protected constraints used to drive it.
GVSS-04 consequently defines the Visual Reflexive Pareto Frontier over a vector such as
The objective is not one universal image score.
It is a nondominated frontier over visual deficit, compute, verification, switching, homogenization, and lineage debt.
The central conclusion is:
1. Position in the series
Paper 01 asks:
Where can digital images exist?
Answer:
Paper 02 asks:
How can desired visual regions be specified?
Answer:
Paper 03 asks:
Which regions can a generator practically reach?
Answer:
Paper 04 asks:
How can an intelligent search system alter its own navigation regime when generation fails?
This introduces reflexivity into GVSS.
2. Classical and current neighboring work
Modern text-to-image generation already supports many forms of conditional control.
Latent Diffusion Models use cross-attention and latent-space generation to support flexible conditioning.
Classifier-Free Guidance changes the tradeoff between fidelity and diversity by combining conditional and unconditional scores.
ControlNet introduces additional structural conditioning such as edges, depth, segmentation, and pose.
Prompt-to-Prompt manipulates cross-attention to control semantic edits while preserving image structure.
Therefore GVSS-04 does not claim conditional visual control as new.
3. Current closed-loop generation precedent
By 2025--2026, visual generation research contains explicit closed-loop and test-time refinement systems.
Examples include:
- test-time prompt refinement using an MLLM to inspect generated images and rewrite prompts;
- iterative image-generation loops using a VLM critic;
- VisionDirector-style structured-goal extraction, semantic verification, staged editing, and rollback;
- Agentic Retoucher-style perception-reasoning-action retouching;
- agentic image generators that integrate planning, search, memory, and feedback;
- verifier-guided test-time search over diffusion noise candidates.
Thus the contribution of GVSS-04 is not the existence of a generation/evaluation loop.
The intended contribution is the state-space and regime-separation synthesis:
4. Digital image state space
Let fixed raster specification be
Let:
Then:
Its cardinality is:
This is inherited from GVSS Paper 01.
5. Structured visual domain
The complete raster space contains:
- natural images;
- artificial images;
- meaningful images;
- noise-like states;
- invalid project assets;
- perceptually redundant states.
Let:
denote a declared structured/meaningful domain.
RRT does not make this subset intrinsic.
Its definition depends on the observer and task.
6. Generator binding
Definition GVSS04-D1
A generator binding is
where identifies the base generator and identifies provider-side control assets such as:
- model profile;
- LoRA;
- adapter;
- reference;
- ControlNet;
- image-to-image source;
- backend-specific workflow;
- sampler family.
The pair determines a base generation regime.
7. Base reachable set
Definition GVSS04-D2
Let:
be the declared base reachable set of generator binding under the allowed finite-precision control surface.
This can be interpreted as:
- exact algorithmic reachability;
- positive-probability support;
- effective support above a probability threshold;
- empirical bounded reachable set.
The interpretation must be stated.
8. Constraint program
Let human intent be:
The compiler produces:
where:
- is a structured constraint program;
- contains weights/priorities.
9. Accepted constraint set
Definition GVSS04-D3
For compiled constraint program , define:
Soft constraints can instead contribute an energy/reward.
The hard-set notation is used for structural theorems.
10. Effective target domain
Definition GVSS04-D4
A target can fail because:
- is empty/internally inconsistent;
- is nonempty but does not intersect ;
- the intersection is nonempty but the search policy fails to find it within budget;
- the compiler does not represent the human intent correctly;
- the evaluator rejects valid target images or accepts invalid ones.
These failures require different actions.
11. Search policy
Definition GVSS04-D5
A search policy
maps history, current constraints, bound generator, and remaining budget into generation/refinement actions.
The history can include:
- previous seeds;
- previous images;
- scores;
- diagnoses;
- failed constraints;
- provider switches;
- human feedback.
12. Bounded practical reachability
Definition GVSS04-D6
Let:
be the set of images that can be produced within budget by executions permitted under policy and generator binding .
The budget can include:
13. GVSS04-T1 — Budget monotonicity
Theorem GVSS04-T1
Suppose the budget feasibility relation is nested and the policy may leave unused resources unspent.
If:
then:
Proof
Every execution feasible under remains feasible under .
Therefore every image reachable with remains reachable with .
This is a resource-monotonicity statement.
It does not say extra budget is always useful.
14. Policy simulation order
Definition GVSS04-D7
Write:
if policy can reproduce every execution trajectory permitted by under budget .
15. GVSS04-T2 — Policy reachability dominance
Theorem GVSS04-T2
If:
then:
Proof
Every execution is also a execution by simulation.
Thus Search Intelligence has a mathematically distinct role from model weights.
16. Practical reachable domain
RRT/GVSS therefore recommends:
Do not assign practical reachability to the base model alone when orchestration/search differs.
17. Visual generation regime
Definition GVSS04-D8
The complete visual regime is:
This state is intentionally larger than a prompt.
18. Generation step
Here:
is the stochastic generation kernel induced by the bound workflow.
For deterministic seed/workflow mappings, can be a Dirac measure conditional on seed.
19. Observer step
The observer produces:
A report can include:
- metric vector;
- hard-gate results;
- failure labels;
- region masks;
- natural-language critique;
- human preference;
- reference comparison.
20. Runtime metric vector
The existing runtime uses:
To avoid collision with provider set , the paper will write:
21. Deficit vector
Given desired thresholds:
define:
The vector:
is not claimed to measure all visual failure.
It measures failure under the declared evaluator coordinates.
22. Hard gates
Some conditions are not scalar preferences.
Let:
record hard-gate validity.
Examples:
- prohibited artifact;
- missing mandatory subject;
- wrong image dimensions;
- project identity violation;
- impossible reference mismatch.
A high composite score cannot override a failed hard gate unless the policy explicitly permits it.
23. Reflexive visual controller
Definition GVSS04-D9
The action changes zero or more coordinates of .
24. Canonical action vocabulary
The bridge adopts the existing runtime vocabulary:
Not every backend must support every action.
25. RESAMPLE semantics
RESAMPLE keeps:
- generator binding;
- compiled constraints;
- provider set;
- evaluator regime;
fixed, while changing sampling state such as:
- seed;
- noise;
- stochastic branch;
- candidate multiplicity.
This is trajectory adaptation inside one regime.
26. RECOMPILE semantics
RECOMPILE changes:
while keeping generator/provider binding fixed.
It is used when:
the current representation of human intent is suspected to be wrong or incomplete.
This is constraint-representation adaptation.
27. REBIND semantics
REBIND changes:
Examples:
- model checkpoint;
- LoRA;
- adapter;
- ControlNet;
- reference strategy;
- model profile;
- provider combination.
This can change the generator's reachable visual domain.
28. REPAIR semantics
REPAIR preserves selected global content while applying a local or structured image edit.
It can be modeled as an image-conditioned transition kernel:
29. SWITCH_BACKEND semantics
SWITCH_BACKEND changes execution mechanism/provider.
This may change:
- model family;
- control vocabulary;
- sampler;
- resolution constraints;
- cost;
- verifier interface.
It is a strong regime change.
30. HUMAN_REVIEW semantics
HUMAN_REVIEW introduces an external observer:
This can refine the joint observation but costs human attention/time.
It does not guarantee correctness.
31. GVSS04-T3 — Action-level reachability separation
Theorem GVSS04-T3
Let current generator binding be , constraint program be , and target domain:
Under the canonical action semantics:
RESAMPLE
If and remain fixed,
Only the sampled trajectory/candidate changes.
RECOMPILE
If remains fixed while:
then:
while:
REBIND
If:
then:
may differ from:
Proof
Each conclusion follows directly from which arguments of are changed by the declared action.
32. Diagnostic consequence
A system should not use REBIND merely because one seed failed.
Likewise, repeated RESAMPLE cannot repair a compiler that consistently represents the wrong intent.
This motivates action escalation.
33. Failure ladder
GVSS-04 recommends a first-pass hierarchy:
This extends but does not overwrite the failure classes already developed in Paper 02/runtime documents.
34. Failure escalation rule
Do not escalate from:
to:
without evidence that the current search policy has exhausted an appropriate bounded search region.
Do not escalate from:
to:
without a broader reachability certificate.
This is the GVSS specialization of RRT failure escalation.
35. Vector defect model
The runtime has multiple quality coordinates.
Therefore scalar error is too coarse.
Let:
36. Action-dependent transfer matrix
For action , define nonnegative matrix:
The matrix describes how previous deficit coordinates survive or cross-couple after the action.
37. Defect injection
Let:
represent new defect caused by:
- model stochasticity;
- prompt/compiler error;
- edit artifacts;
- reference drift;
- evaluator instability;
- provider mismatch.
38. GVSS04-T4 — Vector visual-defect recursion
Theorem GVSS04-T4
Suppose:
componentwise, with nonnegative.
Then:
where the empty product is identity.
Proof
Induction using nonnegativity of all transfer matrices.
39. Interpretation
An action can improve one coordinate while amplifying another.
Example:
- stronger style binding may reduce style deficit;
- but can increase homogenization;
- strong reference control may increase identity consistency;
- but reduce diversity;
- aggressive local repair can improve anatomy;
- but alter texture/style.
This is why one composite score is insufficient.
40. Protected metric coordinates
The existing runtime gives special protection to:
- Prompt;
- Style;
- Character;
- Reference.
Let:
be the protected index set.
41. GVSS04-T5 — Cumulative protected-regression bound
Theorem GVSS04-T5
Suppose every accepted refinement satisfies, for protected coordinate :
Then after accepted steps:
Proof
Sum the one-step inequalities telescopically.
The theorem does not guarantee improvement.
It only bounds accepted cumulative regression.
42. Monotonic improvement is too strong
Requiring every metric to increase at every step can make useful tradeoffs impossible.
The runtime therefore should combine:
- hard regression bounds;
- Pareto acceptance;
- budget;
- protected coordinates.
43. Visual Reflexive Pareto Frontier
Definition GVSS04-D10
Define visual regime cost vector:
The exact coordinates are application dependent.
The Visual Reflexive Pareto Frontier is the nondominated regime/action set.
44. GVSS04-T6 — Pareto necessity
Theorem GVSS04-T6
Every optimum of a scalar objective strictly increasing in all declared cost/deficit coordinates lies on the Visual Reflexive Pareto Frontier.
Proof
Standard dominance argument.
45. Test-time compute as navigation budget
Recent text-to-image work uses test-time compute through:
- multiple noise candidates;
- verifiers;
- iterative correction;
- edit loops;
- prompt refinement.
GVSS interprets this compute as a larger bounded navigation budget.
46. Candidate-parallel search
One policy can allocate budget to parallel candidates:
A verifier selects one.
This increases search breadth.
47. Iterative search
Another policy uses:
This changes search depth and conditions future actions on past results.
Recent work shows that iterative refinement can outperform compute-matched parallel sampling on some compositional benchmarks.
GVSS-04 treats both as different search policies on the same visual-state problem.
48. Policy-dependent practical reachability
Parallel sampling and iterative refinement can expose different bounded reachable regions.
Thus:
and:
need not coincide.
No universal inclusion is claimed.
49. GVSS04-N1 — Same model does not imply same practical reachable set
Two systems can use the same foundation model but different:
- prompt compilers;
- control providers;
- verifiers;
- search policies;
- retry budgets.
Then their practical reachable sets can differ.
This follows from the definition of:
50. Intent compiler as representation
The human does not provide a pixel address.
The compiler constructs a lower-dimensional representation of intent.
Define:
This is more than prompt rewriting.
It can choose:
- constraints;
- weights;
- references;
- providers;
- control models;
- workflow.
51. Intent representation defect
Suppose a declared semantic comparison space for intent is available.
Let:
be the intent reconstructed from compiled constraints.
Define:
This is only meaningful after a metric/observer on intent has been declared.
52. Recompile trigger
RECOMPILE is appropriate when evidence suggests:
or prompt-adherence deficit is high while base generator capability is not yet implicated.
The system should not interpret every low image score as a model-capability failure.
53. Generator reachability suspicion
Repeated failure can justify REBIND only under a declared bounded-search policy.
A practical heuristic is:
- prompt/constraint representation stable;
- evaluator stable;
- multiple diverse trajectories tried;
- protected metrics remain persistently below threshold.
This remains an empirical diagnostic, not a mathematical impossibility proof.
54. Reachability proof versus reachability suspicion
Paper 03 distinguishes operational reachability from ideal support.
GVSS-04 adds another distinction:
Most real image systems can only estimate practical reachability.
55. Style map as a chart
Let:
The Global Artist Keyword Style Map can be viewed as an empirical chart in which artists/styles are indexed by coordinates or feature descriptions.
56. GVSS04-N2 — Style chart is not the whole visual space
A finite-dimensional style map does not replace:
It is a low-dimensional representation chosen for:
- retrieval;
- navigation;
- interpolation;
- prompt compilation;
- style diagnosis.
The map can be useful even if it is highly non-injective.
57. GVSS04-T7 — Coordinate-distance non-intrinsicness
Theorem GVSS04-T7
Let:
Let be invertible but not an Euclidean isometry.
Then there exist such that:
Proof
Because is not an isometry, by definition there exists vector with:
Choose .
Therefore raw Euclidean style distance is representation dependent unless an invariance/normalization rule is supplied.
58. Example
In two style coordinates let:
A one-unit difference on axis 1 becomes a 100-unit difference.
The underlying images did not change.
Only the coordinate chart did.
59. Style chart refinement
Future style maps may move from:
to:
or a nonlinear graph representation.
GVSS does not collapse if the chart changes.
The chart is a navigation representation inside the larger theory.
60. Verifier-gated reachability
Definition GVSS04-D11
Given evaluator/verifier:
and threshold :
Define:
61. GVSS04-T8 — Evaluator-threshold monotonicity
Theorem GVSS04-T8
If:
then:
Proof
implies:
A stricter gate cannot increase the accepted set.
62. Evaluator-guided search
A verifier can be used for:
- ranking parallel candidates;
- stopping;
- choosing which constraint to repair;
- choosing a seed;
- choosing a provider;
- choosing whether to edit or regenerate.
Therefore changes the effective navigation policy even if is fixed.
63. Evaluator misalignment
Let:
represent the intended downstream utility.
The evaluator is a proxy:
A proxy can be misaligned.
64. GVSS04-N3 — Closed-loop proxy optimization no-go
Consider two images:
Let:
Let proxy evaluator:
A controller that optimizes chooses .
Therefore:
Even a perfect closed loop can optimize the wrong objective.
65. Metric pluralism
This is one reason the runtime uses:
A single learned preference score should not silently replace:
- prompt adherence;
- technical validity;
- identity;
- reference consistency;
- diversity;
- anti-homogenization.
66. GenEval / VQAScore boundary
Fine-grained compositional benchmarks and VQA-based image-text evaluation provide stronger evaluation than one holistic similarity number on some tasks.
They remain task-dependent proxies.
GVSS-04 treats evaluation as a representation/observer layer rather than an oracle.
67. Human review as observer fusion
Let automated observer be:
Human observer:
Joint report:
This can contain at least as much raw report information as either coordinate alone.
It also costs human time.
68. Human review is not infallible
Human evaluators can disagree.
Project intent can be underspecified.
A single human approval does not define universal aesthetic truth.
The human observer is another declared information source.
69. Visual lineage
The existing runtime creates lineage child packets rather than overwriting parent packets.
GVSS-04 makes this a formal provenance rule.
70. Lineage record
Definition GVSS04-D12
Every refinement artifact has:
For pure one-parent refinement:
references exactly one earlier artifact.
71. GVSS04-T9 — Single-parent lineage acyclicity
Theorem GVSS04-T9
Suppose every child artifact has a unique ID greater than its parent's creation index and records exactly one earlier parent.
Then the lineage graph is acyclic.
Every artifact has a finite parent path ending at a root artifact.
Proof
Every edge strictly decreases creation index when followed from child to parent.
A directed cycle would require the index to strictly decrease and return to its starting value.
Impossible.
Finite descent ends at a node with no parent.
72. Merge workflows
If an artifact can merge multiple parents, lineage becomes a DAG rather than a tree.
The same acyclicity proof holds if every parent is older than the child.
73. Provenance-preserving refinement
A child should record:
- parent;
- action;
- prompt/compiler delta;
- provider delta;
- seed/workflow;
- evaluator result;
- backend;
- cost.
This allows later audit of which regime change produced which improvement/regression.
74. Reflexive visual state
The minimal closed-loop state can now be written:
is fixed for one raster specification.
Most other coordinates can change.
75. Reflexive update
The controller is reflexive because the report generated under the current visual regime changes the future visual regime.
76. Open-loop generator
An open-loop system uses:
except for predetermined sampling state.
No report-dependent regime adaptation occurs.
77. Closed-loop generator
A closed-loop system permits:
The representation/control regime can therefore move.
78. Reflexive visual navigation
Definition GVSS04-D13
A system performs Reflexive Visual Navigation (RVN) if:
- it generates or edits visual states under regime ;
- an observer evaluates the resulting state;
- the report can change a future representation/control coordinate of ;
- the change is recorded with cost and provenance.
79. RVN is not ordinary rerolling
Repeated random sampling under a fixed prompt/model can be closed-loop only if results influence later policy.
Blindly generating 100 independent samples is search, but not representation-reflexive adaptation.
80. RVN is not prompt engineering alone
Prompt rewriting is one RVN action.
Other actions can change:
- model;
- adapter;
- reference;
- control architecture;
- evaluator;
- backend;
- human review.
The theory is deliberately broader than prompt optimization.
81. RVN is not model fine-tuning alone
Fine-tuning changes or .
RVN can operate without any training by changing:
- search;
- constraints;
- bindings;
- editing;
- verification.
Many current closed-loop T2I systems are training-free.
82. Search breadth versus search depth
Parallel best-of- increases breadth.
Iterative critic/edit/refine increases depth.
Provider rebind changes the local domain being searched.
The runtime controller can allocate budget among all three.
83. Test-time scaling
Define test-time budget:
Increasing it can permit:
- more candidate seeds;
- more verifier calls;
- more edit rounds;
- more provider trials.
By GVSS04-T1, the feasible bounded reachability set cannot shrink if old executions remain allowed.
The expected quality need not monotonically increase under a bad policy.
84. Reachability-aware controller
A controller should maintain hypotheses about failure cause.
Example state:
The next action can depend on this failure belief.
This is an engineering proposal.
85. Bayesian diagnostic controller
A probabilistic implementation could maintain:
The controller then selects action maximizing expected reduction of relevant deficit minus cost.
GVSS-04 does not derive the optimal Bayesian controller.
86. Rule-based controller
The existing runtime uses rule-based triggers.
For example:
- low prompt adherence -> RECOMPILE;
- persistent low style/character/reference scores -> REBIND;
- local technical defect -> REPAIR;
- near threshold -> RESAMPLE.
This is a practical first implementation.
87. Learned controller
A learned policy can replace rules.
But it should still expose:
- action;
- budget;
- metric changes;
- lineage.
Otherwise failure diagnosis becomes opaque.
88. Controller no-go
A more intelligent controller cannot reach an image outside the semantic/algorithmic reachability of every provider it is allowed to bind.
Search intelligence can enlarge practical bounded reachability.
It cannot manufacture expressivity absent from the entire admissible generator family.
This is the visual version of RRT language-search incompleteness.
89. Multi-provider reachable domain
Let allowed bindings be:
With zero-cost switching idealization:
Real switching cost can reduce practical bounded coverage.
90. GVSS04-T10 — Union upper bound for multi-provider reachability
Theorem GVSS04-T10
Any policy restricted to provider family:
can only produce images in:
Proof
Every produced image is produced under one currently bound generator in .
No controller can exceed the union of its admissible generator semantic reachability without adding a new provider/edit operation that enlarges the family.
91. Repair operators enlarge the action family
If REPAIR can map an image to a state not directly reachable from the base generator, the relevant reachable system must include repair kernels too.
Therefore:
not necessarily the T2I generator alone.
92. Runtime reachable closure
Let action kernels be:
The runtime reachable closure is all states reachable by finite admissible compositions under budget.
This is the proper object for a multi-tool visual agent.
93. Model Intelligence versus Search Intelligence
Paper 03 distinguishes generator capability from the ability to find useful regions.
GVSS-04 sharpens this into:
The observer and evaluator affect which trajectories are pursued.
Search intelligence is therefore observer dependent.
94. Evaluator-induced blindness
If evaluator assigns low scores to an actually valuable visual mode, the controller may systematically avoid that region.
Thus the practical reachable accepted set can be smaller than raw runtime reachability.
This is a representation-induced blind spot.
95. Anti-homogenization axis
The runtime explicitly contains:
This is theoretically important.
A controller that maximizes only prompt adherence/aesthetic preference can collapse toward common high-reward modes.
Anti-homogenization acts as a diversity debt constraint.
96. Style consistency versus diversity
These objectives can conflict:
can accompany:
There is no universal optimum without a project-specific objective.
This motivates Pareto evaluation.
97. Unrealized Visual Frontier under a controller
Paper 03 defines:
GVSS-04 can make it policy dependent:
Thus better search intelligence can enlarge the discovered frontier even without changing the base model.
98. Frontier discovery is observer relative
The neighborhood:
depends on a distance/observer.
Therefore "novel" remains metric-relative.
RRT-05 style intrinsicness warnings apply directly.
99. Style chart and novelty chart are different
A style map may be useful for one kind of novelty.
A semantic scene-graph metric may reveal another.
A perceptual embedding may reveal another.
No one chart is the full GVSS geometry.
100. Closed-loop novelty search
A future controller can explicitly optimize for:
- target satisfaction;
- distance from reference modes;
- anti-homogenization;
- style coherence.
This creates a constrained novelty navigation problem.
GVSS-04 does not claim an optimal novelty objective.
101. Current literature: Latent Diffusion
Latent Diffusion Models show that a powerful image generator can operate in a learned latent representation rather than raw pixels and support flexible conditioning through cross-attention.
This is a clear example that the representation used for navigation need not be the raster state space itself.
GVSS treats:
as the ambient representable image space, not necessarily the computational search coordinate system.
102. Current literature: Classifier-Free Guidance
Classifier-Free Guidance explicitly changes the fidelity/diversity tradeoff at inference.
This is a prior example of a control parameter altering the practical sampling geometry.
GVSS-04 treats guidance scale as part of the regime.
103. Current literature: ControlNet
ControlNet adds structural visual controls such as:
- edges;
- depth;
- pose;
- segmentation.
This supports Paper 02's claim that a "prompt" is not the complete visual coordinate/control system.
104. Current literature: Prompt-to-Prompt
Prompt-to-Prompt controls cross-attention to perform semantic edits while preserving more of an existing image structure.
This is a precursor to REPAIR / representation-preserving local transition ideas.
105. Current literature: Test-time Prompt Refinement
TIR uses a multimodal model to:
- inspect generated image;
- diagnose prompt/image mismatch;
- rewrite prompt;
- regenerate;
- verify again.
This is directly a RECOMPILE-style loop.
GVSS adds the possibility that the correct action may instead be RESAMPLE, REBIND, or REPAIR.
106. Current literature: Iterative Refinement
Recent iterative compositional image generation uses a VLM critic to propose corrections and can outperform compute-matched parallel sampling on several compositional benchmarks.
This supports the distinction between:
and:
107. Current literature: VisionDirector
VisionDirector decomposes long instructions into structured goals, uses multimodal verification, chooses staged edits, and supports rollback.
This is especially close to the runtime logic developed before GVSS-04.
The overlap must be cited explicitly.
108. Current literature: Agentic Retoucher
Agentic Retoucher formulates editing as a perception-reasoning-action loop and performs targeted local refinement.
This is a current direct precedent for REPAIR as a closed-loop visual action.
109. Current literature: Qwen-Image-Agent
Qwen-Image-Agent identifies a Context Gap between user context and the generation context needed by T2I systems.
It integrates planning, reasoning, search, memory, and feedback.
This is very close to the GVSS idea that user intent must be compiled into a richer generation regime rather than treated as a complete prompt coordinate.
GVSS-04 therefore avoids claiming the broad "intent/context compiler" problem as uniquely new.
110. Current literature: Verifier-guided Test-Time Scaling
Recent test-time scaling work searches over diffusion/flow noise samples and uses reward/verifier models to allocate compute.
This is a direct neighboring approach to budgeted reachable-space search.
111. Evaluation boundary
GenEval evaluates object-focused compositional properties such as:
- count;
- color;
- position;
- co-occurrence.
VQAScore/GenAI-Bench evaluate complex image-text alignment through visual question answering.
These demonstrate that different observers reveal different failure modes.
GVSS formalizes this as observer/evaluator dependence rather than choosing one universal metric.
112. What is classical / neighboring
GVSS-04 does not claim as inventions:
- diffusion sampling;
- latent diffusion;
- classifier-free guidance;
- ControlNet;
- cross-attention editing;
- best-of-N sampling;
- image reward models;
- VLM image criticism;
- prompt refinement;
- iterative image refinement;
- test-time scaling;
- agentic image generation;
- multi-agent visual refinement;
- Pareto optimization.
113. Candidate GVSS-specific synthesis
Subject to deeper novelty audit, the bridge-specific synthesis is:
- embedding the existing GVSS state-space and reachable-set hierarchy into an RRT-style regime state;
- formally separating RESAMPLE, RECOMPILE, and REBIND by which part of the target/reachable geometry they can change;
- defining policy- and budget-dependent practical visual reachability;
- carrying the existing eight-axis runtime evaluation vector into a vector defect recursion;
- interpreting finite-dimensional style maps as representation charts rather than intrinsic visual geometry;
- treating evaluator thresholds as gates on practical accepted reachability;
- joining visual generation, evaluation, action selection, cost, and lineage into one reflexive visual state;
- defining the Visual Reflexive Pareto Frontier.
No strong novelty claim is made in v0.1.
114. What GVSS-04 proves
Under its explicit definitions/hypotheses, GVSS-04 proves:
- bounded practical reachability is monotone in nested budget;
- a policy that can simulate another policy weakly dominates its bounded reachable set;
- RESAMPLE, RECOMPILE, and REBIND act at distinct geometric levels under the declared action semantics;
- nonnegative vector visual-defect recursions unroll into transported historical defects plus injected defects;
- protected-coordinate regression bounds accumulate additively;
- every strictly monotone scalar optimum lies on the declared visual Pareto frontier;
- Euclidean style-coordinate distance is not intrinsic under arbitrary invertible coordinate changes;
- stricter evaluator thresholds shrink the accepted reachable set;
- proxy-evaluator optimization can worsen true intent under evaluator misalignment;
- single-parent refinement lineage is acyclic when parent indices precede children;
- a multi-provider controller cannot exceed the union of the reachable domains of its admissible runtime action family without adding a new generative/edit operator.
115. What GVSS-04 does not prove
It does not prove:
- a complete intrinsic metric on visual meaning;
- the exact reachable set of any proprietary image model;
- that more test-time compute always improves expected quality;
- that a VLM evaluator represents human intent exactly;
- that iterative refinement always dominates parallel sampling;
- that REBIND always fixes persistent failure;
- that a finite style map captures all visual style;
- that the current eight metric dimensions are complete;
- that human review is infallible;
- that the Unrealized Visual Frontier can be globally measured over all human visual history.
116. Engineering correspondence to the existing runtime
The existing runtime sequence
is the operational implementation boundary of RVN.
GVSS-04 supplies the theoretical interpretation of each stage.
117. Run
Run samples from:
118. Verify
Verify applies observer/evaluator:
119. Diagnose
Diagnose estimates which failure layer is responsible.
120. SelectAction
SelectAction chooses which regime coordinate to change.
121. Refine
Refine applies the action-specific transition.
122. Repeat
The resulting child regime/artifact becomes the state for the next iteration.
123. Closed-loop stopping
Possible stop reasons:
- ACCEPT: target/gates satisfied;
- STOP: budget or policy limit reached;
- HUMAN_REVIEW: automated diagnosis insufficient.
Stopping is not equivalent to proof that no better image exists.
124. Bounded rationality
The runtime acts under:
Therefore the output is best interpreted as:
the result selected under one finite search policy and budget,
not:
the globally optimal image in .
125. Visual search debt
A more complex controller can improve search but adds:
- inference calls;
- VLM calls;
- image edits;
- provider switches;
- latency;
- state management.
This is search debt.
126. Verification debt
More evaluators can expose more failure modes but add:
- compute;
- disagreement;
- calibration burden;
- false-rejection risk.
This is verification debt.
127. Switching debt
REBIND / SWITCH_BACKEND can enlarge capability but add:
- format translation;
- style drift;
- seed discontinuity;
- provider cost;
- reproducibility loss.
This is switching debt.
128. Provenance debt
A runtime that overwrites prompts/workflows/images loses causal information about why an improvement occurred.
Lineage logging reduces provenance debt.
129. Visual regime cost vector
A fuller cost vector can be:
130. No universal scalar objective
Different projects weight:
- style coherence;
- novelty;
- character consistency;
- cost;
- turnaround time;
differently.
Therefore the bridge keeps the vector explicit.
131. RRT relation
The RRT closure principles specialize as follows.
RRT defect transport
becomes:
RRT information/representation order
becomes:
- observer/evaluator refinement;
- style-chart representation;
- human/automated observer fusion.
RRT cost debt
becomes:
- compute;
- verification;
- switching;
- homogenization.
RRT escalation certificate
becomes:
only when lower-level failure is sufficiently diagnosed.
RRT provenance
becomes lineage-preserving packet evolution.
132. Why this is not RRT-21
RRT is closed at RRT-20.
GVSS-04 is a domain specialization.
It imports frozen RRT meta-laws without reopening RRT numbering.
133. Canonical bridge formula
The entire bridge can be written:
134. Canonical visual state
This is the default formal state for later GVSS closed-loop papers.
135. Canonical failure separation
The failure can live in:
- sampling;
- constraints;
- compiler;
- search;
- generator reachability;
- evaluator;
- intent.
136. Canonical runtime principle
Do not rebind the whole model when a seed reroll is enough.
Do not reroll indefinitely when the compiled intent is wrong.
137. Canonical reachability principle
138. Canonical observer principle
139. Canonical provenance principle
140. Proposed GVSS-05
The next paper should not broaden the domain again.
It should sharpen failure diagnosis.
Proposed title:
Chinese:
視覺失敗分層與可達性診斷:從 Seed Failure 到 Generator-Boundary Failure
Main questions:
- How many failures justify escalation from RESAMPLE to RECOMPILE?
- When is persistent failure evidence of provider reachability mismatch?
- Can evaluator disagreement distinguish observer failure from generator failure?
- How should failure beliefs update across iterations?
- What stopping rule prevents infinite rerolling?
- Can diagnostic policies be benchmarked independently of generators?
141. References
- Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, Björn Ommer, High-Resolution Image Synthesis with Latent Diffusion Models, arXiv:2112.10752 / CVPR 2022.
- Jonathan Ho, Tim Salimans, Classifier-Free Diffusion Guidance, arXiv:2207.12598.
- Lvmin Zhang, Anyi Rao, Maneesh Agrawala, Adding Conditional Control to Text-to-Image Diffusion Models, arXiv:2302.05543.
- Amir Hertz et al., Prompt-to-Prompt Image Editing with Cross Attention Control, arXiv:2208.01626.
- Dhruba Ghosh, Hanna Hajishirzi, Ludwig Schmidt, GenEval: An Object-Focused Framework for Evaluating Text-to-Image Alignment, arXiv:2310.11513.
- Jiazheng Xu et al., ImageReward: Learning and Evaluating Human Preferences for Text-to-Image Generation, arXiv:2304.05977.
- Zhiqiu Lin et al., Evaluating Text-to-Visual Generation with Image-to-Text Generation, arXiv:2404.01291.
- Mohammad Abdul Hafeez Khan et al., Test-time Prompt Refinement for Text-to-Image Models, arXiv:2507.22076.
- Shantanu Jaiswal et al., Iterative Refinement Improves Compositional Image Generation, arXiv:2601.15286.
- Meng Chu et al., VisionDirector: Vision-Language Guided Closed-Loop Refinement for Generative Image Synthesis, arXiv:2512.19243.
- Shaocheng Shen et al., Agentic Retoucher for Text-To-Image Generation, arXiv:2601.02046.
- Qwen-Image-Agent: Bridging the Context Gap in Real-World Image Generation, arXiv:2606.26907.
- Vignesh Sundaresha et al., An Efficient Test-Time Scaling Approach for Image Generation, arXiv:2512.08985.
- GVSS Paper 01, The Global Visual State Space Hypothesis, internal series artifact, 2026.
- GVSS Paper 02, Visual Generation as Constraint-Domain Solving, internal series artifact, 2026.
- GVSS Paper 03, Generative Reachability and Unrealized Visuals, internal series artifact, 2026.
- RRT-20, Reflexive Representation Theory: Unified Closure, Meta-Theorems, Limits, and Research Program, internal series artifact, 2026.
142. Conclusion
GVSS Paper 01 turns digital images into points in:
Paper 02 turns generation into constrained search.
Paper 03 turns model capability into bounded reachability.
GVSS-04 turns the search regime itself into a dynamical state.
The controller no longer asks only:
Which image should I sample?
It can ask:
Which representation of intent should I use?
Which constraint should I relax or strengthen?
Which provider should I bind?
Which observer should I trust?
Which region should I repair?
How much budget should I spend?
When should I stop?
The key state is:
The key recursion is:
The key geometric separation is:
The key reachability claim is:
And the bridge principle is:
This establishes Reflexive Visual Navigation as the fourth formal layer of the GVSS sequence.
Canonical-source policy
This file is the canonical UTF-8 source artifact.
- Canonical inline mathematics uses
$...$. - Canonical display mathematics uses
$$...$$. - No Unicode mathematical-symbol conversion is used as source normalization.
- No
unicode_escaperound trip is used. - Backslashes and delimiters are preserved literally.
- Validation is required before release.
- This paper does not reopen RRT numbering.