Pith. sign in

REVIEW 2 cited by

State-space Models with Layer-wise Nonlinearity are Universal Approximators with Exponential Decaying Memory

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2309.13414 v3 pith:EUAYCO5D submitted 2023-09-23 cs.LG cs.AImath.DS

classification cs.LGcs.AImath.DS
keywords modelsstate-spaceactivationlayer-wisenonlinearcapacitydecayingexponential
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

State-space models have gained popularity in sequence modelling due to their simple and efficient network structures. However, the absence of nonlinear activation along the temporal direction limits the model's capacity. In this paper, we prove that stacking state-space models with layer-wise nonlinear activation is sufficient to approximate any continuous sequence-to-sequence relationship. Our findings demonstrate that the addition of layer-wise nonlinear activation enhances the model's capacity to learn complex sequence patterns. Meanwhile, it can be seen both theoretically and empirically that the state-space models do not fundamentally resolve the issue of exponential decaying memory. Theoretical results are justified by numerical verifications.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. DSSMs: State Space Models with Explicit Memory via Delay Differential Equations

    cs.LG 2026-07 conditional novelty 6.5 of 10

    Delay State Space Models augment diagonal SSMs with explicit delayed feedback, stable discrete parameterization, and FFT training, improving delayed-retrieval tasks and matching or beating S4D on most standard sequenc...

  2. Adjoint sharding for very long context training of state space models

    cs.LG 2025-01 reject novelty 5.0 of 10

    The paper derives an adjoint-based gradient sharding algorithm for SSMs and claims up to 3X memory reduction, but provides no experimental evidence for the central empirical claims.

Pith tools