REVIEW 6 cited by
How to Train Your HiPPO: State Space Models with Generalized Orthogonal Basis Projections
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Linear time-invariant state space models (SSM) are a classical model from engineering and statistics, that have recently been shown to be very promising in machine learning through the Structured State Space sequence model (S4). A core component of S4 involves initializing the SSM state matrix to a particular matrix called a HiPPO matrix, which was empirically important for S4's ability to handle long sequences. However, the specific matrix that S4 uses was actually derived in previous work for a particular time-varying dynamical system, and the use of this matrix as a time-invariant SSM had no known mathematical interpretation. Consequently, the theoretical mechanism by which S4 models long-range dependencies actually remains unexplained. We derive a more general and intuitive formulation of the HiPPO framework, which provides a simple mathematical interpretation of S4 as a decomposition onto exponentially-warped Legendre polynomials, explaining its ability to capture long dependencies. Our generalization introduces a theoretically rich class of SSMs that also lets us derive more intuitive S4 variants for other bases such as the Fourier basis, and explains other aspects of training S4, such as how to initialize the important timescale parameter. These insights improve S4's performance to 86% on the Long Range Arena benchmark, with 96% on the most difficult Path-X task.
Forward citations
Cited by 6 Pith papers
-
Rivaling Transformers: Multi-Scale Structured State-Space Mixtures for Agentic 6G O-RAN
A 0.70M-parameter multi-scale state-space mixture predicts next-step RSRP on an O-RAN testbed with RMSE 0.29 dB and R2=0.993, running 3-10x faster than the tested Transformers.
-
LIDAR: Lightweight Adaptive Cue-Aware Fusion Vision Mamba for Multimodal Segmentation of Structural Cracks
LIDAR is a lightweight adaptive fusion Vision Mamba network that reports state-of-the-art multimodal crack segmentation accuracy with 5.35M parameters on a light-field depth dataset.
-
Few-Shot Object Detection via Spatial-Channel State Space Model
A Mamba-based channel sequence model combined with spatial attention improves few-shot object detection on VOC and COCO.
-
FlexiD-Fuse: Flexible number of inputs multi-modal medical image fusion based on diffusion model
FlexiD-Fuse adapts a denoising diffusion model and an expectation-maximization step to fuse either two or three medical images with one shared network, and reports better scores than fixed-count baselines on standard metrics.
-
Quantizing Small-Scale State-Space Models for Edge AI
Quantization-aware training with a frozen state matrix lifts sequential MNIST accuracy from 40% under post-training quantization to 96%, and a heterogeneous precision scheme cuts memory by 6 times.
-
W4S4: WaLRUS Meets S4 for Long-Range Sequence Modeling
W4S4 initializes S4 state space models with WaLRUS wavelet frames and reports better delay reconstruction and classification accuracy than HiPPO-based S4, with frozen (A,B).
Discussion (0). Continue with ORCID to comment.