REVIEW 4 cited by
TR-DQ: Time-Rotation Diffusion Quantization
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Diffusion models have been widely adopted in image and video generation. However, their complex network architecture leads to high inference overhead for its generation process. Existing diffusion quantization methods primarily focus on the quantization of the model structure while ignoring the impact of time-steps variation during sampling. At the same time, most current approaches fail to account for significant activations that cannot be eliminated, resulting in substantial performance degradation after quantization. To address these issues, we propose Time-Rotation Diffusion Quantization (TR-DQ), a novel quantization method incorporating time-step and rotation-based optimization. TR-DQ first divides the sampling process based on time-steps and applies a rotation matrix to smooth activations and weights dynamically. For different time-steps, a dedicated hyperparameter is introduced for adaptive timing modeling, which enables dynamic quantization across different time steps. Additionally, we also explore the compression potential of Classifier-Free Guidance (CFG-wise) to establish a foundation for subsequent work. TR-DQ achieves state-of-the-art (SOTA) performance on image generation and video generation tasks and a 1.38-1.89x speedup and 1.97-2.58x memory reduction in inference compared to existing quantization methods.
Forward citations
Cited by 4 Pith papers
-
GVD: Guiding Video Diffusion Model for Scalable Video Distillation
GVD guides a pre-trained video diffusion model with clustering-derived features to distill video datasets, outperforming prior methods on MiniUCF and HMDB51 while retaining over 70% of full-data accuracy using under 4...
-
FPSAttention: Training-Aware FP8 and Sparsity Co-Design for Fast Video Diffusion
FPSAttention co-designs FP8 quantization and sparsity with training, achieving 4.96x end-to-end video generation speedup on Wan2.1 with roughly preserved quality.
-
ICM-Fusion: In-Context Meta-Optimized LoRA Fusion for Multi-Task Adaptation
ICM-Fusion uses a conditional VAE plus task-vector guidance to fuse multiple LoRA adapters into one model, reporting marginal average gains on vision and language benchmarks and larger gains in a few-shot long-tail setup.
-
HadaNorm: Diffusion Transformer Quantization through Mean-Centered Transformations
A quantization transform combining dynamic per-token mean centering, channel scaling, and Hadamard transforms reduces 4-bit activation quantization error in diffusion transformers, achieving a CLIP score of 31.69 on P...
Discussion (0). Continue with ORCID to comment.