Pith. sign in

REVIEW 4 cited by

Training-Free Text-Guided Image Editing with Visual Autoregressive Model

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2503.23897 v1 pith:6VBHKHWP submitted 2025-03-31 cs.CV cs.AI

classification cs.CVcs.AI
keywords imageeditinginversionmodificationstext-guidedvisualachievesautoregressive
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Text-guided image editing is an essential task that enables users to modify images through natural language descriptions. Recent advances in diffusion models and rectified flows have significantly improved editing quality, primarily relying on inversion techniques to extract structured noise from input images. However, inaccuracies in inversion can propagate errors, leading to unintended modifications and compromising fidelity. Moreover, even with perfect inversion, the entanglement between textual prompts and image features often results in global changes when only local edits are intended. To address these challenges, we propose a novel text-guided image editing framework based on VAR (Visual AutoRegressive modeling), which eliminates the need for explicit inversion while ensuring precise and controlled modifications. Our method introduces a caching mechanism that stores token indices and probability distributions from the original image, capturing the relationship between the source prompt and the image. Using this cache, we design an adaptive fine-grained masking strategy that dynamically identifies and constrains modifications to relevant regions, preventing unintended changes. A token reassembling approach further refines the editing process, enhancing diversity, fidelity, and control. Our framework operates in a training-free manner and achieves high-fidelity editing with faster inference speeds, processing a 1K resolution image in as fast as 1.2 seconds. Extensive experiments demonstrate that our method achieves performance comparable to, or even surpassing, existing diffusion- and rectified flow-based approaches in both quantitative metrics and visual quality. The code will be released.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. SynVAR: Synergizing Spatial and Semantic Alignment in Visual Autoregressive Model

    cs.CV 2026-08 conditional novelty 6.0 of 10

    SynVAR improves compositional generation of VAR models by injecting spatial priors, constraining early self-attention, and enhancing high-frequency details.

  2. UniGen-AR: Unifying Visual Generation with Auto-Regressive Modeling

    cs.CV 2026-07 conditional novelty 6.0 of 10

    An MLLM-conditioned next-scale VAR decoder handles 15+ unified visual generation tasks with competitive quality and substantially lower latency than diffusion baselines.

  3. Adaptive Visual Autoregressive Acceleration via Dual-Linkage Entropy Analysis

    cs.CV 2026-02 conditional novelty 5.0 of 10

    A training-free entropy-guided token-pruning framework accelerates VAR image generation up to 2.9× with negligible benchmark loss by activating pruning at an adaptive entropy-growth inflection point and adjusting rati...

  4. Mastering Regional 3DGS: Locating, Initializing, and Editing with Diverse 2D Priors

    cs.CV 2025-07 conditional novelty 5.0 of 10

    A 3D Gaussian Splatting editing pipeline that combines 2D diffusion localization, depth-based point seeding, and sequential view refinement to achieve up to 4x faster local edits.

Pith tools