Pith. sign in

arXiv preprint arXiv:2002.09402 , year =

5 Pith papers cite this work, alongside 31 external citations. Polarity classification is still indexing.

5 Pith papers citing it
31 external citations · external index

citation-role summary

background 1

citation-polarity summary

fields

cs.LG 5

years

2026 4 2025 1

roles

background 1

polarities

background 1

representative citing papers

The Recurrent Transformer: Greater Effective Depth and Efficient Decoding

cs.LG · 2026-04-23 · unverdicted · novelty 6.0

Recurrent Transformers add per-layer recurrent memory via self-attention on own activations plus a tiling algorithm that reduces training memory traffic, yielding better C4 pretraining cross-entropy than parameter-matched standard transformers with fewer layers.

Sessa: Selective State Space Attention

cs.LG · 2026-04-20 · unverdicted · novelty 5.0

Sessa integrates attention within recurrent paths to achieve power-law memory tails and flexible non-decaying selective retrieval, outperforming baselines on long-context tasks.

Pretraining Recurrent Networks without Recurrence

cs.LG · 2026-06-04 · conditional · novelty 4.0

SMT trains nonlinear RNNs by imitating one-step memory-transition labels generated by a Transformer, replacing BPTT's unrolled credit assignment with time-parallel supervised learning.

citing papers explorer

Showing 5 of 5 citing papers.