Pith. sign in

REVIEW 7 cited by

ViDiT-Q: Efficient and Accurate Quantization of Diffusion Transformers for Image and Video Generation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2406.02540 v3 pith:KOB76DQG submitted 2024-06-04 cs.CV

classification cs.CV
keywords quantizationvideochallengesdiffusiongenerationmemorytransformersvidit-q
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Diffusion transformers have demonstrated remarkable performance in visual generation tasks, such as generating realistic images or videos based on textual instructions. However, larger model sizes and multi-frame processing for video generation lead to increased computational and memory costs, posing challenges for practical deployment on edge devices. Post-Training Quantization (PTQ) is an effective method for reducing memory costs and computational complexity. When quantizing diffusion transformers, we find that existing quantization methods face challenges when applied to text-to-image and video tasks. To address these challenges, we begin by systematically analyzing the source of quantization error and conclude with the unique challenges posed by DiT quantization. Accordingly, we design an improved quantization scheme: ViDiT-Q (Video & Image Diffusion Transformer Quantization), tailored specifically for DiT models. We validate the effectiveness of ViDiT-Q across a variety of text-to-image and video models, achieving W8A8 and W4A8 with negligible degradation in visual quality and metrics. Additionally, we implement efficient GPU kernels to achieve practical 2-2.5x memory saving and a 1.4-1.7x end-to-end latency speedup.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 7 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. TASQ: Temporal-Adaptive Bit Sparsification Quantization for Diffusion Models

    cs.CV 2026-08 conditional novelty 7.0 of 10

    A temporal-spatial LSB mask over one shared weight buffer lets diffusion models use lower bit precision in less sensitive denoising stages, cutting compute by 25-50% on bit-serial hardware with no loss in image quality.

  2. QuantWAMs: Calibrating at the Right Granularity for World Action Models

    cs.AI 2026-07 conditional novelty 6.0 of 10

    Aligning PTQ decisions to WAM structure, closed-loop rollouts, and the joint video–action objective yields W4A4 policies within 0.2–0.7 pp of FP16 on simulation benchmarks with ~29% block memory.

  3. DMQ: Dissecting Outliers of Diffusion Models for Post-Training Quantization

    cs.CV 2025-07 conditional novelty 6.0 of 10

    A post-training quantization method that combines learned channel scaling and power-of-two scaling keeps diffusion image quality high at 4-bit weight, 6-bit activation precision.

  4. Q-VDiT: Towards Accurate Quantization and Distillation of Video-Generation Diffusion Transformers

    cs.CV 2025-05 conditional novelty 6.0 of 10

    Q-VDiT quantizes video diffusion transformers to 3-4 bit weights by adding a learned rank-1 error correction (TQE) and a temporal distribution distillation loss (TMD), nearly doubling VBench scene consistency at W3A6 ...

  5. MARch\'e: Fast Masked Autoregressive Image Generation with Cache-Aware Attention

    cs.LG 2025-05 conditional novelty 6.0 of 10

    MARche accelerates masked autoregressive image generation by caching stable token projections and refreshing only attention-selected tokens, reaching up to 1.72x speedup with some loss in FID.

  6. HyperVAttention: Efficient Sparse Attention with Spatio-Temporal Clustering for Video Diffusion

    cs.CV 2026-07 conditional novelty 5.5 of 10

    Training-free sparse attention for video DiTs cuts latency up to ~2.1× via 3D local-window clustering, hybrid step updates, and hardware-aware cluster merging while improving fidelity over prior sparse methods.

  7. Awakening Diffusion Transformers: Eliciting Stronger Generation and Understanding via Massive Activation Modulation

    cs.CV 2026-07 conditional novelty 5.0 of 10

    Massive activations in DiTs are timestep-driven detail channels; suppressing them guides finer sampling and AdaLN-modulating them yields more discriminative dense features.

Pith tools