Pith. sign in

REVIEW 18 cited by

DiffusionDB: A Large-scale Prompt Gallery Dataset for Text-to-Image Generative Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2210.14896 v4 pith:WSGSFIJH submitted 2022-10-26 cs.CV cs.AIcs.HCcs.LG

classification cs.CVcs.AIcs.HCcs.LG
keywords promptsdiffusiondbmodelsdatasetimagesmodelpromptusers
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

With recent advancements in diffusion models, users can generate high-quality images by writing text prompts in natural language. However, generating images with desired details requires proper prompts, and it is often unclear how a model reacts to different prompts or what the best prompts are. To help researchers tackle these critical challenges, we introduce DiffusionDB, the first large-scale text-to-image prompt dataset totaling 6.5TB, containing 14 million images generated by Stable Diffusion, 1.8 million unique prompts, and hyperparameters specified by real users. We analyze the syntactic and semantic characteristics of prompts. We pinpoint specific hyperparameter values and prompt styles that can lead to model errors and present evidence of potentially harmful model usage, such as the generation of misinformation. The unprecedented scale and diversity of this human-actuated dataset provide exciting research opportunities in understanding the interplay between prompts and generative models, detecting deepfakes, and designing human-AI interaction tools to help users more easily use these models. DiffusionDB is publicly available at: https://poloclub.github.io/diffusiondb.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 18 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Navigating the Open-Source Model Ecosystem: An Empirical Study of Creator Practices in Artistic Image Generation

    cs.HC 2026-07 accept novelty 7.0 of 10

    A novel 6M-image Pixiv dataset shows open-source image generation has long-tail model usage, slow life cycles with version inertia, and surging multi-LoRA customization linked to higher engagement.

  2. Beyond Prompts: Unconditional 3D Inversion for Out-of-Distribution Shapes

    cs.CV 2026-04 unverdicted novelty 7.0 of 10

    Text-to-3D models lose prompt sensitivity for out-of-distribution shapes due to sink traps but retain geometric diversity via unconditional priors, enabling a decoupled inversion method for robust editing.

  3. KroQuant: Kronecker-Structured Block Transforms for Efficient Post-Training Quantization of Diffusion Transformers

    cs.LG 2026-07 conditional novelty 6.0 of 10

    Using learned 32×32 Kronecker block transforms as an online activation smoother improves W4A4 image quality of PixArt-Sigma, SANA, and FLUX.1-schnell over SVDQuant and LoRaQ, with a kernel up to 14% faster than SmoothQuant.

  4. Direct Diffusion Score Preference Optimization via Stepwise Contrastive Policy-Pair Supervision

    cs.CV 2025-12 conditional novelty 6.0 of 10

    Diffusion image models can be aligned without human labels by supervising every denoising step with score targets from original versus degraded prompts.

  5. Elastic ViTs from Pretrained Models without Retraining

    cs.CV 2025-10 conditional novelty 6.0 of 10

    A single-shot, label-free, retraining-free structured pruning method generates elastic ViTs at any sparsity by reweighting gradient-based importance scores with block correlations learned by an evolutionary strategy.

  6. TetriServe: Efficiently Serving Mixed DiT Workloads

    cs.LG 2025-10 conditional novelty 6.0 of 10

    TetriServe's step-level, deadline-aware sequence parallelism improves SLO attainment for mixed-resolution diffusion transformer serving by up to 32% over fixed-SP systems.

  7. Maestro: Self-Improving Text-to-Image Generation via Agent Orchestration

    cs.AI 2025-09 conditional novelty 6.0 of 10

    A multi-agent prompt-refinement system using pairwise AI judging and targeted edit signals outperforms prior automated methods on complex text-to-image tasks.

  8. Directly Aligning the Full Diffusion Trajectory with Fine-Grained Human Preference

    cs.AI 2025-09 conditional novelty 6.0 of 10

    Direct-Align and SRPO fine-tune FLUX using ground-truth-noise recovery and text-conditional relative rewards, improving human-evaluated realism and aesthetics roughly 3x.

  9. Understanding and evaluating computer vision models through the lens of counterfactuals

    cs.CV 2025-08 conditional novelty 6.0 of 10

    Counterfactual-based methods for concept attribution in classifiers and for dynamic bias evaluation and mitigation in text-to-image models.

  10. 4KAgent: Agentic Any Image to 4K Super-Resolution

    cs.CV 2025-07 reject novelty 6.0 of 10

    An agentic pipeline that plans and executes image restoration from a toolbox of pretrained models to upscale arbitrary images to 4K, reporting state-of-the-art results on many benchmarks.

  11. Instella-T2I: Pushing the Limits of 1D Discrete Latent Space Image Generation

    cs.CV 2025-06 conditional novelty 6.0 of 10

    1D binary image latents reduce a 1024x1024 image to 128 discrete tokens and support text-to-image generation with diffusion and autoregressive models.

  12. Ambient Diffusion Omni: Training Good Models with Bad Data

    cs.GR 2025-06 conditional novelty 6.0 of 10

    Ambient Diffusion Omni trains diffusion models on mixed-quality data by learning when corrupted images can be treated as clean, improving generation quality and diversity.

  13. SEED: A Benchmark Dataset for Sequential Facial Attribute Editing with Diffusion Models

    cs.CV 2025-05 reject novelty 6.0 of 10

    SEED is a 91,526-image benchmark of diffusion-generated sequential facial edits with sequence, mask, and prompt annotations, and FAITH adds DWT high-frequency cues to a transformer for edit-sequence detection.

  14. SketchFlex: Facilitating Spatial-Semantic Coherence in Text-to-Image Generation with Region-Based Sketches

    cs.HC 2025-02 conditional novelty 6.0 of 10

    SketchFlex combines sketch-aware prompt recommendation with decompose-and-recompose shape refinement to help novices generate multi-object images from rough region sketches.

  15. Argus-Unified: Towards A Compact and Economical Unified Model for Image Understanding and Generation

    cs.CV 2026-07 conditional novelty 5.0 of 10

    A compact unified model that reuses a frozen VLM encoder and hybrid continuous/discrete tokens reaches competitive image understanding and generation with 15.6M training images and about $2,000 in compute.

  16. Attention Dynamics in Diffusion Models: A Visual Analytics Framework for Human-AI Collaboration

    cs.CV 2026-06 conditional novelty 4.5 of 10

    A DAAM-based visual analytics workflow links step-resolved token attention trajectories, phase summaries, and spatial competition maps for Stable Diffusion-class models on a 60-prompt benchmark.

  17. NTIRE 2025 challenge on Text to Image Generation Model Quality Assessment

    cs.CV 2025-05 conditional novelty 4.0 of 10

    The NTIRE 2025 challenge report compares 20 methods for fine-grained text-to-image quality assessment, introduces the EvalMuse-Structure dataset, and finds every participating team outperformed the baselines.

  18. DejAIvu: Identifying and Explaining AI Art on the Web in Real-Time with Saliency Maps

    cs.CV 2025-02 conditional novelty 4.0 of 10

    DejAIvu is a browser extension that classifies images as AI-generated or human-made in real time and highlights the evidence with saliency heatmaps.

Pith tools