Pith. sign in

REVIEW 11 cited by

PIXART-{\delta}: Fast and Controllable Image Generation with Latent Consistency Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2401.05252 v1 pith:YPXVWYCP submitted 2024-01-10 cs.CV

classification cs.CV
keywords pixart-deltaimagesalphagenerationhigh-qualityimagemodel
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

This technical report introduces PIXART-{\delta}, a text-to-image synthesis framework that integrates the Latent Consistency Model (LCM) and ControlNet into the advanced PIXART-{\alpha} model. PIXART-{\alpha} is recognized for its ability to generate high-quality images of 1024px resolution through a remarkably efficient training process. The integration of LCM in PIXART-{\delta} significantly accelerates the inference speed, enabling the production of high-quality images in just 2-4 steps. Notably, PIXART-{\delta} achieves a breakthrough 0.5 seconds for generating 1024x1024 pixel images, marking a 7x improvement over the PIXART-{\alpha}. Additionally, PIXART-{\delta} is designed to be efficiently trainable on 32GB V100 GPUs within a single day. With its 8-bit inference capability (von Platen et al., 2023), PIXART-{\delta} can synthesize 1024px images within 8GB GPU memory constraints, greatly enhancing its usability and accessibility. Furthermore, incorporating a ControlNet-like module enables fine-grained control over text-to-image diffusion models. We introduce a novel ControlNet-Transformer architecture, specifically tailored for Transformers, achieving explicit controllability alongside high-quality image generation. As a state-of-the-art, open-source image generation model, PIXART-{\delta} offers a promising alternative to the Stable Diffusion family of models, contributing significantly to text-to-image synthesis.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 11 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. A Decomposable Probe for Few-Step Diffusion Models: Prompt, Latent, and Score Selectivity across Backbone Families and Distillation Paradigms

    cs.CV 2026-07 conditional novelty 6.5 of 10

    A three-layer perturbation probe shows latent selectivity is a near-binary rectified-flow fingerprint that survives ADD distillation, while score selectivity tracks distillation objective across 23 T2I models.

  2. TerraDiT: Point-Conditioned Diffusion Transformer for Satellite Image Synthesis

    cs.CV 2026-03 conditional novelty 6.0 of 10

    GeoDiT is a point-conditioned diffusion transformer that generates satellite imagery from sparse labeled points and claims to beat existing remote sensing generators on FID and SSIM.

  3. Interpretable and Reliable Detection of AI-Generated Images via Grounded Reasoning in MLLMs

    cs.CV 2025-06 conditional novelty 6.0 of 10

    Fine-tuning Qwen-2.5-VL on the new FakeXplained dataset of 8,772 AI-generated images with box-and-caption artifact annotations yields an explainable detector with 98.1% accuracy and 37.8% IoU.

  4. ComposeAnything: Composite Object Priors for Text-to-Image Generation

    cs.CV 2025-05 conditional novelty 6.0 of 10

    ComposeAnything generates composite object priors from LLM-generated 2.5D layouts and guides diffusion denoising, improving compositional fidelity in text-to-image generation.

  5. FastFace: Tuning Identity Preservation in Distilled Diffusion via Guidance and Attention

    cs.CV 2025-05 conditional novelty 6.0 of 10

    An inference-time framework of decoupled classifier-free guidance and attention manipulation improves identity preservation and prompt alignment when pretrained face ID adapters are used with few-step distilled diffus...

  6. ControlTac: Force- and Position-Controlled Tactile Data Augmentation with a Single Reference Image

    cs.CV 2025-05 conditional novelty 6.0 of 10

    A two-stage diffusion framework generates realistic tactile images from one reference image, conditioned on target contact force and position, and the generated images improve downstream force estimation, pose estimat...

  7. SpecMaskFoley: Steering Pretrained Spectral Masked Generative Transformer Toward Synchronized Video-to-audio Synthesis via ControlNet

    cs.SD 2025-05 conditional novelty 6.0 of 10

    A ControlNet branch plus a frequency-aware feature aligner lets a pretrained masked generative TTA model produce video-synchronized foley, beating several from-scratch models on VGGSound.

  8. UltraImageGen: Efficient Ultra-High-Resolution Image Generation with Hierarchical Local Attention

    cs.CV 2025-10 conditional novelty 5.0 of 10

    A pretrained FLUX diffusion model is adapted with local-window attention plus low-resolution global guidance, allowing 4K text-to-image generation from 1K-only training data at about 2x lower cost.

  9. NanoControl: A Lightweight Framework for Precise and Efficient Control in Diffusion Transformer

    cs.CV 2025-08 conditional novelty 5.0 of 10

    NanoControl injects condition-specific key-value pairs into every attention block of Flux via a LoRA-style branch, claiming state-of-the-art controllability at 0.024% extra parameters and 0.029% extra FLOPs.

  10. DC-AR: Efficient Masked Autoregressive Image Generation with Deep Compression Hybrid Tokenizer

    cs.CV 2025-07 conditional novelty 5.0 of 10

    DC-AR generates 512x512 images in 12 masked autoregressive steps plus 20 diffusion refinement steps, using a 32x compressed 2D tokenizer, and reports gFID 5.49 on MJHQ-30K.

  11. LiftVSR: Lifting Image Diffusion to Video Super-Resolution via Hybrid Temporal Modeling with Only 4$\times$RTX 4090s

    cs.CV 2025-06 conditional novelty 5.0 of 10

    LiftVSR combines short-segment dynamic temporal attention, a long-term attention memory cache, and Diffusion Forcing style asymmetric sampling to achieve strong perceptual video super-resolution scores with dramatical...

Pith tools