NeurIPS 2026

SLVR: Structured Latent Visual Reasoning via Human-like Reasoning Flows

1Carnegie Mellon University 2Meta *Corresponding author
SLVR teaser

SLVR turns visual reasoning into a typed latent flow: plan what to inspect, ground the relevant region, select visual evidence, integrate it internally, and only then decode the answer—without generating a textual chain of thought at inference time.

Abstract

Multimodal large language models often answer visual reasoning questions by relying on linguistic priors rather than task-relevant visual evidence. Textual chain-of-thought reasoning can encourage evidence-seeking, but autoregressive rationale generation increases inference cost. Latent reasoning avoids explicit rationale generation, yet existing approaches provide limited control over what intermediate states encode.

We propose Structured Latent Visual Reasoning (SLVR), a training framework that bridges explicit chain-of-thought and latent reasoning by organizing multimodal reasoning into typed latent stages for planning, grounding, evidence selection, and reasoning integration. SLVR first trains the model to rely on the image through Visual Dependency Induction (VDI), then applies stage-specific latent supervision using plans, boxes, visual evidence, and final rationales. Built on Qwen2.5-VL-7B, SLVR improves consistently across multimodal reasoning benchmarks while avoiding the decoding overhead of textual CoT.

Method

From visual dependency to structured latent reasoning

SLVR pipeline
01

Visual Dependency Induction

Mask answer-revealing textual cues and contrast the correct answer with visually plausible distractors, forcing the model to recover the missing fact from the image.

02

Typed Latent Reasoning

Organize hidden computation into functional latent stages for planning, grounding, evidence selection, and integration, with supervision matched to each stage.

Results

Consistent gains across visual reasoning benchmarks

+9.4
MMVP
66.7 → 76.1
+14.2
BLINK Relation
38.8 → 53.0
+10.6
V* Spatial
73.7 → 84.3
+9.7
MathVista
63.7 → 73.4
ModelV* OverallMMVPBLINK CountBLINK JigsawBLINK RelationBLINK IQMathVistaChartQA
Qwen2.5-VL-7B78.566.766.752.038.826.063.786.2
SLVR (Ours)85.276.172.462.353.034.573.489.6

Absolute comparison against the Qwen2.5-VL-7B base model. See the paper for all prior-method baselines and subset results.

Ablation

Each stage contributes

Base66.7MMVP
→
+ VDI71.2MMVP
→
+ Plan73.0MMVP
→
+ Ground74.2MMVP
→
+ Evidence75.2MMVP
→
SLVR76.1MMVP

The final latent budget uses 64 planning tokens, 16 evidence tokens, and 16 integration tokens.

Citation

BibTeX

@inproceedings{gao2026slvr,
  title     = {SLVR: Structured Latent Visual Reasoning via Human-like Reasoning Flows},
  author    = {Gao, Albert and Xue, Bing and Zanette, Andrea},
  booktitle = {Advances in Neural Information Processing Systems},
  year      = {2026}
}