Learning to Remember, Learn, and Forget in Attention-Based Models

Djohan Bonnet; Elidona Shiqerukaj; Emre Neftci; Jamie Lohoff; Jan Finkbeiner

arxiv: 2602.09075 · v4 · pith:BTT56JT5new · submitted 2026-02-09 · 💻 cs.LG · cs.AI

Learning to Remember, Learn, and Forget in Attention-Based Models

Djohan Bonnet , Jamie Lohoff , Jan Finkbeiner , Elidona Shiqerukaj , Emre Neftci This is my paper

classification 💻 cs.LG cs.AI

keywords palimpsaattentionlearningmemorymodelsassociativecapacitygated

0 comments

read the original abstract

In-Context Learning (ICL) in transformers acts as an online associative memory and is believed to underpin their high performance on complex sequence processing tasks. However, in gated linear attention models, this memory has a fixed capacity and is prone to interference, especially for long sequences. We propose Palimpsa, a self-attention model that views ICL as a continual learning problem that must address a stability-plasticity dilemma. Palimpsa uses Bayesian metaplasticity, where the plasticity of each attention state is tied to an importance state grounded by a prior distribution that captures accumulated knowledge. We demonstrate that various gated linear attention models emerge as specific architecture choices and posterior approximations, and that Mamba2 is a special case of Palimpsa where forgetting dominates. This theoretical link enables the transformation of any non-metaplastic model into a metaplastic one, significantly expanding its memory capacity. Our experiments show that Palimpsa consistently outperforms baselines on the Multi-Query Associative Recall (MQAR) benchmark and on Commonsense Reasoning tasks.

This paper has not been read by Pith yet.

discussion (0)

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

Memory by Design: Probabilistic Sequence Layers
stat.ML 2026-05 unverdicted novelty 6.0

The design-model framework unifies sub-quadratic sequence models as Bayesian filters and introduces a covariance-tracking Bayesian Layer that improves retrieval robustness beyond training regimes on MQAR and RULER benchmarks.