REVIEW 11 cited by
SiMBA: Simplified Mamba-Based Architecture for Vision and Multivariate Time series
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Transformers have widely adopted attention networks for sequence mixing and MLPs for channel mixing, playing a pivotal role in achieving breakthroughs across domains. However, recent literature highlights issues with attention networks, including low inductive bias and quadratic complexity concerning input sequence length. State Space Models (SSMs) like S4 and others (Hippo, Global Convolutions, liquid S4, LRU, Mega, and Mamba), have emerged to address the above issues to help handle longer sequence lengths. Mamba, while being the state-of-the-art SSM, has a stability issue when scaled to large networks for computer vision datasets. We propose SiMBA, a new architecture that introduces Einstein FFT (EinFFT) for channel modeling by specific eigenvalue computations and uses the Mamba block for sequence modeling. Extensive performance studies across image and time-series benchmarks demonstrate that SiMBA outperforms existing SSMs, bridging the performance gap with state-of-the-art transformers. Notably, SiMBA establishes itself as the new state-of-the-art SSM on ImageNet and transfer learning benchmarks such as Stanford Car and Flower as well as task learning benchmarks as well as seven time series benchmark datasets. The project page is available on this website ~\url{https://github.com/badripatro/Simba}.
Forward citations
Cited by 11 Pith papers
-
Black-Mamba: Biologically-Inspired Leaky Accumulation for Conceptual Knowledge under Distribution Drift
Gating test-time memory writes on leaky accumulated surprisal preserves most of the adaptation benefit while roughly halving the number of updates.
-
DCVC-MB: Neural B-Frame Video Compression using State Space Models
DCVC-MB, a neural B-frame video codec using Mamba state-space fusion, reports BD-rate savings up to 8.98% over prior neural codecs and up to 30.45% over VTM-19.0-LDP.
-
Partial Ring Scan: Revisiting Scan Order in Vision State Space Models
Ring-based scanning with selective channel routing improves accuracy, speed, and rotation robustness of vision state-space models.
-
CLIMP: Contrastive Language-Image Mamba Pretraining
A fully Mamba-based (VMamba + Mamba LLM) CLIP model matches or beats transformer baselines on retrieval and OOD benchmarks, and natively supports high resolutions and dense captions.
-
Mamba-X: An End-to-End Vision Mamba Accelerator for Edge Computing Devices
A dedicated accelerator for Vision Mamba using a Kogge-Stone systolic scan array and hybrid 8-bit quantization achieves 2.3x end-to-end speedup and 11.5x energy-efficiency gain over an edge GPU with less than 1% top-1...
-
Training-free Token Reduction for Vision Mamba
MTR uses Mamba's timescale parameter Δ as a token importance score to merge unimportant tokens, giving training-free inference speedups with small accuracy loss.
-
SparseSSM: Efficient Selective Structured State Space Models Can Be Pruned in One-Shot
SparseSSM extends OBS-style second-order pruning to Mamba's discretized, time-shared state-transition matrix, pruning 50% of its weights in one pass without fine-tuning.
-
FLDmamba: Integrating Fourier and Laplace Transform Decomposition with Mamba for Enhanced Time Series Prediction
FLDmamba combines a learnable Fourier filter on Mamba's step size with a damped-sinusoid output layer and reports superior long-term forecasting accuracy on standard benchmarks.
-
Exploring State-Space-Model based Language Model in Music Generation
SiMBA, a Mamba-based decoder, converges faster and matches or slightly trails a Transformer baseline in a single-codebook text-to-music system under limited compute.
-
Burst Image Super-Resolution via Multi-Cross Attention Encoding and Multi-Scan State-Space Decoding
A Transformer-Mamba hybrid with multi-cross attention and multi-scan state-space fusion reports state-of-the-art PSNR and SSIM on synthetic burst super-resolution benchmarks.
-
Straightforward Bayesian A/B testing with Dirichlet posteriors
The submission is internally inconsistent: the abstract promises a Bayesian A/B testing method, but the full text is a different computer vision paper, leaving the claimed result unevaluable.
Discussion (0). Continue with ORCID to comment.