Pith. sign in

REVIEW 3 cited by

DreamTuner: Single Image is Enough for Subject-Driven Generation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2312.13691 v1 pith:OA4YVMX4 submitted 2023-12-21 cs.CV

classification cs.CV
keywords generationimagesubjectsubject-drivendetailsdreamturnerfeaturesmodels
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Diffusion-based models have demonstrated impressive capabilities for text-to-image generation and are expected for personalized applications of subject-driven generation, which require the generation of customized concepts with one or a few reference images. However, existing methods based on fine-tuning fail to balance the trade-off between subject learning and the maintenance of the generation capabilities of pretrained models. Moreover, other methods that utilize additional image encoders tend to lose important details of the subject due to encoding compression. To address these challenges, we propose DreamTurner, a novel method that injects reference information from coarse to fine to achieve subject-driven image generation more effectively. DreamTurner introduces a subject-encoder for coarse subject identity preservation, where the compressed general subject features are introduced through an attention layer before visual-text cross-attention. We then modify the self-attention layers within pretrained text-to-image models to self-subject-attention layers to refine the details of the target subject. The generated image queries detailed features from both the reference image and itself in self-subject-attention. It is worth emphasizing that self-subject-attention is an effective, elegant, and training-free method for maintaining the detailed features of customized subjects and can serve as a plug-and-play solution during inference. Finally, with additional subject-driven fine-tuning, DreamTurner achieves remarkable performance in subject-driven image generation, which can be controlled by a text or other conditions such as pose. For further details, please visit the project page at https://dreamtuner-diffusion.github.io/.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. StyleBlend: Enhancing Style-Specific Content Creation in Text-to-Image Diffusion Models

    cs.CV 2025-02 conditional novelty 6.0 of 10

    StyleBlend learns few-shot artistic style as separate layout and texture components and blends them during diffusion sampling to improve text-aligned, style-specific image generation.

  2. Dimension-Reduction Attack! Video Generative Models are Experts on Controllable Image Synthesis

    cs.CV 2025-05 conditional novelty 5.0 of 10

    A video diffusion model, HunyuanVideo-I2V, is adapted with mixup transitions, frame-skip position embeddings, and attention masking to outperform image-only models on several controllable image generation benchmarks.

  3. Training Free Stylized Abstraction

    cs.CV 2025-05 conditional novelty 5.0 of 10

    A training-free framework coupling VLLM-based identity distillation with cross-domain rectified flow inversion generates identity-preserving stylized abstractions from a single reference image, evaluated by a new GPT-...

Pith tools