REVIEW 5 cited by
ZigMa: A DiT-style Zigzag Mamba Diffusion Model
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
abstract
The diffusion model has long been plagued by scalability and quadratic complexity issues, especially within transformer-based structures. In this study, we aim to leverage the long sequence modeling capability of a State-Space Model called Mamba to extend its applicability to visual data generation. Firstly, we identify a critical oversight in most current Mamba-based vision methods, namely the lack of consideration for spatial continuity in the scan scheme of Mamba. Secondly, building upon this insight, we introduce a simple, plug-and-play, zero-parameter method named Zigzag Mamba, which outperforms Mamba-based baselines and demonstrates improved speed and memory utilization compared to transformer-based baselines. Lastly, we integrate Zigzag Mamba with the Stochastic Interpolant framework to investigate the scalability of the model on large-resolution visual datasets, such as FacesHQ $1024\times 1024$ and UCF101, MultiModal-CelebA-HQ, and MS COCO $256\times 256$ . Code will be released at https://taohu.me/zigma/
Forward citations
Cited by 5 Pith papers
-
Exploring Diffusion Transformer Designs via Grafting
Grafting uses activation distillation and lightweight fine-tuning to edit pretrained diffusion transformers into hybrid architectures with near-baseline quality at under 2% pretraining compute.
-
StyleRWKV: High-Quality and High-Efficiency Style Transfer with RWKV-like Architecture
StyleRWKV applies recurrent RWKV-style attention with deformable shifting and skip scanning to achieve fast, high-quality arbitrary style transfer.
-
QuarterMap: Efficient Post-Training Token Pruning for Visual State Space Models
QuarterMap prunes spatial activations before VMamba's four-directional scan and upsamples after, yielding up to 1.11x throughput with under 1% accuracy loss on ImageNet classification.
-
MaIR: A Locality- and Continuity-Preserving Mamba for Image Restoration
MaIR combines a stripe-based S-shaped scanning strategy and a sequence-shuffle attention block to improve Mamba-based image restoration, reporting new best PSNR on 14 benchmarks.
-
LinGen: Towards High-Resolution Minute-Length Text-to-Video Generation with Linear Computational Complexity
LinGen replaces self-attention in diffusion transformers with a linear-complexity MATE block, enabling 512p 68-second video generation on a single H100 with quality comparable to Gen-3, LumaLabs, and Kling.
Discussion (0). Continue with ORCID to comment.