Pith. sign in

REVIEW 11 cited by

Thinking with Generated Images

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2505.22525 v1 pith:PKWXCOSB submitted 2025-05-28 cs.CV cs.AIcs.CL

classification cs.CVcs.AIcs.CL
keywords visualmodelsreasoningthinkinggeneratedgenerationimagesvision
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We present Thinking with Generated Images, a novel paradigm that fundamentally transforms how large multimodal models (LMMs) engage with visual reasoning by enabling them to natively think across text and vision modalities through spontaneous generation of intermediate visual thinking steps. Current visual reasoning with LMMs is constrained to either processing fixed user-provided images or reasoning solely through text-based chain-of-thought (CoT). Thinking with Generated Images unlocks a new dimension of cognitive capability where models can actively construct intermediate visual thoughts, critique their own visual hypotheses, and refine them as integral components of their reasoning process. We demonstrate the effectiveness of our approach through two complementary mechanisms: (1) vision generation with intermediate visual subgoals, where models decompose complex visual tasks into manageable components that are generated and integrated progressively, and (2) vision generation with self-critique, where models generate an initial visual hypothesis, analyze its shortcomings through textual reasoning, and produce refined outputs based on their own critiques. Our experiments on vision generation benchmarks show substantial improvements over baseline approaches, with our models achieving up to 50% (from 38% to 57%) relative improvement in handling complex multi-object scenarios. From biochemists exploring novel protein structures, and architects iterating on spatial designs, to forensic analysts reconstructing crime scenes, and basketball players envisioning strategic plays, our approach enables AI models to engage in the kind of visual imagination and iterative refinement that characterizes human creative, analytical, and strategic thinking. We release our open-source suite at https://github.com/GAIR-NLP/thinking-with-generated-images.

Discussion (0). Sign in to comment.

Forward citations

Cited by 11 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Bridging Interleaved Multi-Modal Reasoning as a Unified Decision Process

    cs.AI 2026-07 conditional novelty 7.0 of 10

    BRAID jointly optimizes text and image generation in interleaved multi-modal reasoning by casting the full trajectory as a unified MDP with shared advantages and a VLM process reward.

  2. Demystifying Video Reasoning

    cs.CV 2026-03 conditional novelty 7.0 of 10

    Video diffusion models reason along the denoising trajectory (Chain-of-Steps), not primarily across frames, and this mechanism can be nudged by ensembling latent trajectories.

  3. VC-Tooler: Learning Compositional and Adaptive Visual Tool Use

    cs.CV 2026-08 conditional novelty 6.0 of 10

    VC-Tooler trains Qwen3-VL-8B on hierarchically synthesized tool trajectories via SFT then GRPO with a judge-based tool reward, reaching open-source SOTA on V* (95.8) and VTC-Bench (35.3) and transferring to 11 unseen ...

  4. FaithEyes: Towards Faithful Tool Use via Multi-Agent Process-Image Verification

    cs.CV 2026-07 conditional novelty 6.0 of 10

    A self-judging multi-agent VLM injects process-image helpfulness verdicts into tool observations and scales tool rewards by the helpful-call ratio, improving accuracy and tool faithfulness.

  5. OPLD: On-Policy Latent Distillation for Multimodal Reasoning

    cs.CV 2026-07 conditional novelty 6.0 of 10

    On-policy distillation of a CoT-privileged teacher’s token preferences and latent trajectories yields a CoT-free student that beats prior visual-latent methods on several multimodal benchmarks.

  6. BVS: Bayesian Visual Search with Multimodal Large Language Model for Fine-grained Perception

    cs.CV 2026-07 conditional novelty 6.0 of 10

    BVS combines early-stop attention rollout priors with a scale-aware non-stationary kernel and GP-UCB to locate tiny objects in UHR images more accurately and with fewer MLLM queries than prior visual-search methods.

  7. Bridging Modal Isolation in Interleaved Thinking: Supervising Modality Transitions via Stepwise Reinforcement

    cs.CV 2026-06 unverdicted novelty 6.0 of 10

    Supervising text–image handoffs with Reflective SFT and Flow-GRPO (MoTiF) reduces modal isolation and raises accuracy on four visual puzzle benchmarks versus end-task-only training.

  8. MMRad-22K: A Structured Multimodal Evidence Dataset for Chest X-ray Report Generation

    cs.CV 2026-02 conditional novelty 6.0 of 10

    A 21,994-case chest X-ray dataset with interleaved regional text and image crops helps LVLMs generate more clinically accurate reports than text-only chain-of-thought.

  9. MentisOculi: Revealing the Limits of Reasoning with Mental Imagery

    cs.AI 2026-02 conditional novelty 6.0 of 10

    Visual thoughts — latent tokens, interleaved images, or video rollouts — do not currently improve multi-step reasoning over text-only baselines in frontier models.

  10. IRG-MotionLLM: Interleaving Motion Generation, Assessment and Refinement for Text-to-Motion Generation

    cs.CV 2025-12 conditional novelty 6.0 of 10

    Interleaving motion generation with text-motion assessment and refinement improves alignment between generated human motion and goal text.

  11. UniDrive-WM: Unified Understanding, Planning and Generation World Model for Autonomous Driving

    cs.CV 2026-01 conditional novelty 5.0 of 10

    A unified VLM for autonomous driving that couples trajectory planning with future-frame image generation improves open- and closed-loop planning metrics on Bench2Drive and nuScenes.

Pith tools