Pith. sign in

REVIEW 3 cited by

Fine-Tuning Visual Autoregressive Models for Subject-Driven Generation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2504.02612 v2 pith:DGZSQ4DL submitted 2025-04-03 cs.CV

classification cs.CV
keywords modelsgenerationpracticalsubject-drivenautoregressivecomputationaldetailsdiffusion-based
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Recent advances in text-to-image generative models have enabled numerous practical applications, including subject-driven generation, which fine-tunes pretrained models to capture subject semantics from only a few examples. While diffusion-based models produce high-quality images, their extensive denoising steps result in significant computational overhead, limiting real-world applicability. Visual autoregressive (VAR) models, which predict next-scale tokens rather than spatially adjacent ones, offer significantly faster inference suitable for practical deployment. In this paper, we propose the first VAR-based approach for subject-driven generation. However, naive fine-tuning VAR leads to computational overhead, language drift, and reduced diversity. To address these challenges, we introduce selective layer tuning to reduce complexity and prior distillation to mitigate language drift. Additionally, we found that the early stages have a greater influence on the generation of subject than the latter stages, which merely synthesize minor details. Based on this finding, we propose scale-wise weighted tuning, which prioritizes coarser resolutions for promoting the model to focus on the subject-relevant information instead of local details. Extensive experiments validate that our method significantly outperforms diffusion-based baselines across various metrics and demonstrates its practical usage.

Discussion (0). Sign in to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Layout-Conditioned Autoregressive Text-to-Image Generation via Structured Masking

    cs.CV 2025-09 conditional novelty 6.0 of 10

    SMARLI achieves strong layout control in autoregressive text-to-image generation via structured attention masks and GRPO post-training with a CLIP-based layout reward.

  2. FreeLoRA: Enabling Training-Free LoRA Fusion for Autoregressive Multi-Subject Personalization

    cs.CV 2025-07 conditional novelty 6.0 of 10

    A method for multi-subject image personalization that fuses independently trained LoRA modules at inference time on visual autoregressive models.

  3. Translation of Text Embedding via Delta Vector to Suppress Strongly Entangled Content in Text-to-Image Diffusion Models

    cs.CV 2025-08 conditional novelty 5.0 of 10

    Subtracting the text embedding of an unwanted concept from a target word's embedding, with cross-attention keys and values steered in opposite directions, suppresses strongly entangled content in Stable Diffusion and ...

Pith tools