Pith. sign in

REVIEW 15 cited by

Emu: Enhancing Image Generation Models Using Photogenic Needles in a Haystack

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2309.15807 v1 pith:2YR67UG4 submitted 2023-09-27 cs.CV

classification cs.CV
keywords modelsimagesgenerationmodelpre-trainedvisualaestheticappealing
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

Training text-to-image models with web scale image-text pairs enables the generation of a wide range of visual concepts from text. However, these pre-trained models often face challenges when it comes to generating highly aesthetic images. This creates the need for aesthetic alignment post pre-training. In this paper, we propose quality-tuning to effectively guide a pre-trained model to exclusively generate highly visually appealing images, while maintaining generality across visual concepts. Our key insight is that supervised fine-tuning with a set of surprisingly small but extremely visually appealing images can significantly improve the generation quality. We pre-train a latent diffusion model on $1.1$ billion image-text pairs and fine-tune it with only a few thousand carefully selected high-quality images. The resulting model, Emu, achieves a win rate of $82.9\%$ compared with its pre-trained only counterpart. Compared to the state-of-the-art SDXLv1.0, Emu is preferred $68.4\%$ and $71.3\%$ of the time on visual appeal on the standard PartiPrompts and our Open User Input benchmark based on the real-world usage of text-to-image models. In addition, we show that quality-tuning is a generic approach that is also effective for other architectures, including pixel diffusion and masked generative transformer models.

Discussion (0). Sign in to comment.

Forward citations

Cited by 15 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Towards Physics of Multimodal Pretraining: Knowledge Flow, Modality Synergy, Early Unification, and Recipes

    cs.CV 2026-08 conditional novelty 6.0 of 10

    Multimodal pretraining transfers asymmetrically: language boosts vision, understanding boosts generation, generation is mostly neutral, and early unified training prevents vision laziness.

  2. PIPBench: A Profile-Inclusive Framework for Personalized Image Generation Evaluation

    cs.CV 2026-07 conditional novelty 6.0 of 10

    PIPBench is a profile-inclusive benchmark with 1,369 test cases from 251 real users and synthetic agents that evaluates personalized image generation methods, revealing that current approaches struggle to jointly inte...

  3. Transferability Between Understanding and Generation in Unified Multimodal Models

    cs.CV 2026-07 conditional novelty 6.0 of 10

    Cross-task capability transfer in UMMs is architecture-dependent and can be exploited by training understanding to improve generation while avoiding distribution shift.

  4. Evaluating and comparing gender bias across four text-to-image models

    cs.CY 2025-09 conditional novelty 6.0 of 10

    Across 30 professions and 6,000 images, DALL-E 3 over-represented women, Stable Diffusion XL and Cascade over-represented men in high-status roles, and Emu was more balanced.

  5. LuxDiT: Lighting Estimation with Video Diffusion Transformer

    cs.GR 2025-09 conditional novelty 6.0 of 10

    A video diffusion transformer fine-tuned on synthetic and real data predicts HDR environment maps from images/videos, cutting peak light-direction error by roughly 45% on sunny outdoor scenes versus DiffusionLight.

  6. ShortFT: Diffusion Model Alignment via Shortcut-based Fine-Tuning

    cs.CV 2025-07 conditional novelty 6.0 of 10

    ShortFT fine-tunes Stable Diffusion by backpropagating reward gradients through a distilled few-step shortcut denoising chain, improving alignment scores over DRaFT-LV and DRTune.

  7. Understanding Trade offs When Conditioning Synthetic Data

    cs.CV 2025-07 conditional novelty 6.0 of 10

    Diverse layout-plus-prompt conditioning of diffusion models generates synthetic data that improves few-shot object detection mAP by up to 177% over real-data-only training, while prompt-only conditioning wins when con...

  8. Populate-A-Scene: Affordance-Aware Human Video Generation

    cs.CV 2025-07 conditional novelty 6.0 of 10

    A fine-tuned text-to-video model inserts a person into a scene and generates an interaction video without bounding boxes or pose input, and its attention maps reveal a latent sense of affordance.

  9. FAIL: Flow Matching Adversarial Imitation Learning for Image Generation

    cs.CV 2026-02 conditional novelty 5.0 of 10

    Post-training of flow matching can be framed as adversarial imitation learning, and the proposed FAIL methods improve FLUX's generation quality using 13K expert images without preference pairs.

  10. UniLDiff: Unlocking the Power of Diffusion Priors for All-in-One Image Restoration

    cs.CV 2025-07 conditional novelty 5.0 of 10

    UniLDiff combines degradation-aware attention fusion with a detail-aware expert decoder to achieve state-of-the-art perceptual quality on unified image restoration benchmarks.

  11. VISTA: A Visual Analytics Framework to Enhance Foundation Model-Generated Data Labels

    cs.CV 2025-07 conditional novelty 5.0 of 10

    A human-in-the-loop visual analytics system that detects, summarizes, and corrects label and alignment errors in foundation-model-generated image segmentation data, improving downstream open-vocabulary segmentation pe...

  12. Diff2Flow: Training Flow Matching Models via Diffusion Model Alignment

    cs.CV 2025-06 conditional novelty 5.0 of 10

    A reparameterization recipe that lets pre-trained Stable Diffusion checkpoints be finetuned as flow matching models, giving faster convergence and better performance under parameter-efficient constraints.

  13. MGVQ: Could VQ-VAE Beat VAE? A Generalizable Tokenizer with Multi-group Quantization

    cs.CV 2025-07 conditional novelty 4.0 of 10

    Splitting quantization across multiple small sub-codebooks with nested masking raises VQ-VAE reconstruction fidelity, giving MGVQ rFID 0.49 and PSNR 24.70 on ImageNet at 16 times downsampling.

  14. On the Comprehensibility of Multi-structured Financial Documents using LLMs and Pre-processing Tools

    cs.IR 2025-06 conditional novelty 4.0 of 10

    Preprocessing financial PDFs into text, tables, and chart data with existing tools improves LLM question-answering accuracy over direct GPT-4o image input in a small private evaluation.

  15. Smoothed Preference Optimization via ReNoise Inversion for Aligning Diffusion Models with Varied Human Preferences

    cs.CV 2025-06 conditional novelty 4.0 of 10

    SmPO-Diffusion improves diffusion-model preference alignment with reward-model soft labels and ReNoise inversion, reporting higher human-preference scores and up to 26x lower training cost than Diffusion-KTO.

Pith tools