REVIEW 5 cited by
MambaVC: Learned Visual Compression with Selective State Spaces
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Learned visual compression is an important and active task in multimedia. Existing approaches have explored various CNN- and Transformer-based designs to model content distribution and eliminate redundancy, where balancing efficacy (i.e., rate-distortion trade-off) and efficiency remains a challenge. Recently, state-space models (SSMs) have shown promise due to their long-range modeling capacity and efficiency. Inspired by this, we take the first step to explore SSMs for visual compression. We introduce MambaVC, a simple, strong and efficient compression network based on SSM. MambaVC develops a visual state space (VSS) block with a 2D selective scanning (2DSS) module as the nonlinear activation function after each downsampling, which helps to capture informative global contexts and enhances compression. On compression benchmark datasets, MambaVC achieves superior rate-distortion performance with lower computational and memory overheads. Specifically, it outperforms CNN and Transformer variants by 9.3% and 15.6% on Kodak, respectively, while reducing computation by 42% and 24%, and saving 12% and 71% of memory. MambaVC shows even greater improvements with high-resolution images, highlighting its potential and scalability in real-world applications. We also provide a comprehensive comparison of different network designs, underscoring MambaVC's advantages. Code is available at https://github.com/QinSY123/2024-MambaVC.
Forward citations
Cited by 5 Pith papers
-
Next-Frame Decoding for Ultra-Low-Bitrate Image Compression with Video Diffusion Priors
Ultra-low-bitrate image decoding is cast as one-step next-frame prediction from a compact anchor using adapted video diffusion priors, yielding large perceptual bitrate savings versus DiffC.
-
DCVC-MB: Neural B-Frame Video Compression using State Space Models
DCVC-MB, a neural B-frame video codec using Mamba state-space fusion, reports BD-rate savings up to 8.98% over prior neural codecs and up to 30.45% over VTM-19.0-LDP.
-
Understanding Rate-Distortion Performance in Distributed Transformer Inference
Deeper transformer layers produce intermediate representations that are harder to lossy-compress, and the paper links this to growing covariance and Rademacher complexity.
-
Linear Attention Modeling for Learned Image Compression
LALIC replaces transformer and Mamba blocks in learned image compression with bidirectional RWKV linear-attention blocks, reporting BD-rate gains over VTM-9.1 while keeping decoder latency moderate.
-
CMamba: Learned Image Compression with State Space Models
A hybrid CNN and Mamba (state space model) image compression codec reports BD-Rate savings of 14.95% to 18.83% over VVC with fewer parameters, FLOPs, and lower decoding time than the prior best learned method.
Discussion (0). Continue with ORCID to comment.