Pith. sign in

REVIEW 16 cited by

PixArt-\Sigma: Weak-to-Strong Training of Diffusion Transformer for 4K Text-to-Image Generation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2403.04692 v2 pith:ZPMPJEWY submitted 2024-03-07 cs.CV

classification cs.CV
keywords pixart-sigmaimagetrainingdatadiffusionimagesmodel
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

In this paper, we introduce PixArt-\Sigma, a Diffusion Transformer model~(DiT) capable of directly generating images at 4K resolution. PixArt-\Sigma represents a significant advancement over its predecessor, PixArt-\alpha, offering images of markedly higher fidelity and improved alignment with text prompts. A key feature of PixArt-\Sigma is its training efficiency. Leveraging the foundational pre-training of PixArt-\alpha, it evolves from the `weaker' baseline to a `stronger' model via incorporating higher quality data, a process we term "weak-to-strong training". The advancements in PixArt-\Sigma are twofold: (1) High-Quality Training Data: PixArt-\Sigma incorporates superior-quality image data, paired with more precise and detailed image captions. (2) Efficient Token Compression: we propose a novel attention module within the DiT framework that compresses both keys and values, significantly improving efficiency and facilitating ultra-high-resolution image generation. Thanks to these improvements, PixArt-\Sigma achieves superior image quality and user prompt adherence capabilities with significantly smaller model size (0.6B parameters) than existing text-to-image diffusion models, such as SDXL (2.6B parameters) and SD Cascade (5.1B parameters). Moreover, PixArt-\Sigma's capability to generate 4K images supports the creation of high-resolution posters and wallpapers, efficiently bolstering the production of high-quality visual content in industries such as film and gaming.

Discussion (0). Sign in to comment.

Forward citations

Cited by 16 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. UniMC: Taming Diffusion Transformer for Unified Keypoint-Guided Multi-Class Image Generation

    cs.CV 2025-07 conditional novelty 7.0 of 10

    UniMC uses tokenized instance conditions (class, box, keypoints) and a timestep-aware modulator in a DiT backbone to control multi-class human and animal image generation, trained and evaluated on the new HAIG-2.9M dataset.

  2. AI-generated Images Challenge Visual Trust in High-risk Scenarios

    cs.CV 2026-07 conditional novelty 6.0 of 10

    On SafeIMG, a new safety-focused benchmark of 1,131 GPT Image 2 images, the best VLM detects 49.5% of generated images and the best specialized detector 33.1%, versus 81.7% for humans.

  3. SnapGen++: Unleashing Diffusion Transformers for Efficient High-Fidelity Image Generation on Edge Devices

    cs.CV 2026-01 conditional novelty 6.0 of 10

    A compact elastic diffusion transformer with adaptive sparse attention and knowledge-guided distribution-matching distillation achieves 4-step 1K image generation on a phone in roughly 1.8 seconds.

  4. DMQ: Dissecting Outliers of Diffusion Models for Post-Training Quantization

    cs.CV 2025-07 conditional novelty 6.0 of 10

    A post-training quantization method that combines learned channel scaling and power-of-two scaling keeps diffusion image quality high at 4-bit weight, 6-bit activation precision.

  5. ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation

    cs.CV 2025-06 conditional novelty 6.0 of 10

    A 91K GPT-4o-generated image and editing dataset, and a fine-tuned open model Janus-4o, report improved text-to-image scores and new editing ability.

  6. Multi-Group Proportional Representation for Text-to-Image Models

    cs.CV 2025-05 conditional novelty 6.0 of 10

    The authors apply the MPR metric (an integral probability metric) to text-to-image generation, derive tractable forms for linear and decision-tree function classes, and use it as a fine-tuning objective that reduces i...

  7. Differentiable Solver Search for Fast Diffusion Sampling

    cs.CV 2025-05 conditional novelty 6.0 of 10

    A differentiable search over solver coefficients and sampling timesteps produces a fast diffusion sampler that outperforms DPM-Solver++ and UniPC at 5 to 10 steps.

  8. Unified Backbone Refinement for Diffusion Models via Internal-Latent Analysis

    cs.CV 2026-07 conditional novelty 5.0 of 10

    Early abrupt deviations in deep diffusion latents track artifacts; EMA detection plus backbone-specific suppression (DUNE) reduces them without retraining.

  9. APT: Improving Diffusion Models for High Resolution Image Generation with Adaptive Path Tracing

    cs.CV 2025-07 conditional novelty 5.0 of 10

    A training-free add-on for latent diffusion models that fixes patch statistics and re-schedules noise, improving detail in high-resolution images while reducing sampling steps.

  10. Inversion-DPO: Precise and Efficient Post-Training for Diffusion Models

    cs.CV 2025-07 reject novelty 5.0 of 10

    Inversion-DPO uses DDIM inversion to convert winning and losing images into noise trajectories, yielding a simpler DPO loss for diffusion model alignment that trains faster and improves text-to-image and compositional...

  11. Structured Captions Improve Prompt Adherence in Text-to-Image Models (Re-LAION-Caption 19M)

    cs.CV 2025-07 conditional novelty 5.0 of 10

    Structured four-field captions produced small but consistent gains in VQA-based text-image alignment over shuffled versions of the same captions when fine-tuning PixArt-Sigma and Stable Diffusion 2.

  12. Normalized Attention Guidance: Universal Negative Guidance for Diffusion Models

    cs.CV 2025-05 conditional novelty 5.0 of 10

    Normalized Attention Guidance (NAG) stabilizes attention-space extrapolation with L1 normalization and refinement, restoring negative prompting in few-step diffusion models across architectures and modalities.

  13. MixDiffusion: Mixing Diffusion-based Uni-condition Text-to-Image Generation Models for Multi-condition Image Synthesis

    cs.CV 2026-07 conditional novelty 4.0 of 10

    MixDiffusion derives a joint noise prediction as the sum of per-condition noise estimates minus the base model, enabling multi-condition control without training.

  14. Re-Thinking the Automatic Evaluation of Image-Text Alignment in Text-to-Image Models

    cs.CL 2025-06 conditional novelty 4.0 of 10

    Current image-text alignment metrics, including CLIPScore and DSGScore, produce unstable model rankings under random seeds and are highly sensitive to tiny image perturbations.

  15. How Much To Guide: Revisiting Adaptive Guidance in Classifier-Free Guidance Text-to-Vision Diffusion Models

    cs.CV 2025-06 conditional novelty 4.0 of 10

    Using classifier-free guidance for only the first 30-50% of denoising steps preserves generation quality while saving 20-30% inference time across image and video diffusion models.

  16. OpenUni: A Simple Baseline for Unified Multimodal Understanding and Generation

    cs.CV 2025-05 conditional novelty 4.0 of 10

    A lightweight open-source connector between a frozen multimodal LLM and a diffusion model yields a unified model that matches larger systems on image generation and understanding benchmarks, with the caveat that headl...

Pith tools