REVIEW 4 major objections 1 minor 1 cited by
Memory-Augmented Transformers: A Systematic Review from Neuroscience Principles to Enhanced Model Architectures
T0 review · 4 major / 1 minor · reviewed 2026-08-05 · deepseek-v4-flash
Pith's one-line read A three-axis taxonomy organizes Memory-Augmented Transformers and shows a shift from static caches to adaptive test-time learning.
desk verdict A coherent survey abstract atop an unreadable body: worth sending back for a clean copy and a methods section, not yet citable. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The organizing device is the taxonomy itself: each memory-augmented Transformer is located by a triple—objective, representation, integration. The four memory operations (reading, writing, forgetting, capacity management) are the analytic lens that turns the taxonomy from a static list into a diagnosis of what each architecture can and cannot do. The neuroscience mapping supplies the design rationale: multi-timescale memory justifies separate memory stores, selective attention justifies content-based retrieval, and consolidation justifies writing and forgetting rules—so the taxonomy is meant to be explanatory, not merely descriptive.
What would settle it
Take a defined literature window, for example memory-augmented Transformer papers published 2023–2025, collect the full set with a fixed search query, and try to assign each paper to exactly one cell of the three-axis taxonomy. If more than a small fraction resist placement, or the full set shows no chronological trend toward test-time learning, the central claims fail.
Extended reading notes
Core claim
The paper's central claim is that the diverse memory mechanisms bolted onto Transformers form a coherent design space, not a collection of unrelated patches. It proposes a three-axis taxonomy: functional objectives (context extension, reasoning, knowledge integration, adaptation), memory representations (parameter-encoded, state-based, explicit, hybrid), and integration mechanisms (attention fusion, gated control, associative retrieval). Using reading, writing, forgetting, and capacity management as the core memory operations, it argues that the field is shifting from fixed external caches toward adaptive systems that update their own memory at test time. It further claims that neuroscience
Load-bearing premise
The framework stands or falls on whether the neuroscience constructs genuinely correspond to the engineering categories and whether the surveyed papers were selected systematically rather than because they fit the story.
Editorial extensions
If this is right
- Architectures that currently look incomparable—a long-context cache, a gated state model, an external associative memory—become comparable through their position on the three axes.
- The claimed shift to adaptive, test-time learning sets a concrete design target: future models should write and update memory during inference, not only at training time.
- Scalability and interference are named as the two bottlenecks any new memory design must address, so progress on them can be measured directly.
- Hierarchical buffering and surprise-gated updates are named as emerging mechanisms, giving practitioners concrete starting points for new architectures.
- The neuroscience link provides a vocabulary for generating new designs, such as treating a Transformer's memory as a multi-timescale system with distinct stores.
Reading between the lines
- Editorial inference: the taxonomy's usefulness depends on whether it can classify papers independently; a natural test is to take a held-out set of recent memory-augmented Transformer papers and see whether two annotators place them in the same taxonomy cells.
- Editorial inference: if the shift-to-test-time-learning claim is right, evaluation practice should change—benchmarks that only test static retrieval will miss the capabilities that differentiate the newer architectures.
- Editorial inference: the paper asserts rather than demonstrates its 'systematic' coverage; a reader should look for a reproducible search protocol, because without one the trend claim could reflect selection bias.
- Editorial inference: the memory-operations framework likely generalizes beyond Transformers to any architecture with an explicit memory component, so the roadmap could also apply to state-space models or recurrent networks.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The manuscript (arXiv:2508.10824) claims to provide a unified framework bridging neuroscience principles with engineering advances in Memory-Augmented Transformers, organized along three taxonomic dimensions (functional objectives, memory representations, integration mechanisms), and to identify a shift from static caches toward adaptive, test-time learning systems. The abstract is self-consistent and the claims are appropriately modest for a review. However, the supplied full text is heavily corrupted (mojibake) and unreadable; it also contains an inserted fragment from a different paper (arXiv:2508.10821v3 [q-bio.QM]). Consequently, none of the substantive content—sections, tables, equations, references, or the survey corpus—can be verified. The 'systematic' nature of the review and the empirical trend claim cannot be audited from the provided text.
Significance. If the claims hold, the three-axis taxonomy could provide a useful common vocabulary for memory-augmented transformers and a roadmap toward continual-learning architectures. The synthesis of neuroscience constructs and engineering designs could be valuable if the mapping is substantive. However, because the body is unreadable, the significance is conditional. The paper does not derive new mathematical results or ship code; its contribution, if any, is organizational and interpretive. Credit is due for a self-consistent abstract and clearly stated dimensions, but no machine-checkable evidence is available.
major comments (4)
- [Full text (after abstract)] The entire body of the manuscript is encoded mojibake; no section, table, equation, or reference can be read or verified. For a systematic review, the central claims—the taxonomy and the trend toward adaptive test-time learning—rest on the ability to check which papers were surveyed and how they were mapped. This is load-bearing and prevents any audit. I cannot determine whether the summaries are accurate or whether the conclusions follow from the cited literature.
- [Visible fragment in full text] The text includes the line 'arXiv:2508.10821v3 [q-bio.QM] 20 May 2026', which is an identifier for a different paper in quantitative biology. This indicates that the PDF/source has been corrupted or misassembled. As supplied, the document is not self-consistent; it contains content that does not belong to the claimed review. This is a material issue because it makes it impossible to attribute any of the surrounding text to the authors' intended manuscript.
- [Search protocol/inclusion criteria] The abstract calls the review 'systematic,' but no search strategy, inclusion/exclusion criteria, database sources, or PRISMA-style flow is visible in the supplied text. The trend claim ('a shift from static caches toward adaptive, test-time learning systems') is an empirical statement about the surveyed literature; without a reproducible corpus and coding protocol, it may be a cherry-picked narrative. This is an audibility problem, not necessarily an error, but it must be fixed for the review to be evaluable.
- [Neuroscience-to-engineering mapping] The 'unified framework' is the paper's main contribution, but the body that would justify the correspondence between biological constructs (multi-timescale memory, selective attention, consolidation) and architectural families (parameter-encoded, state-based, explicit, hybrid) is unreadable. This raises a correctness risk: the mapping could be substantive constraints or post-hoc labels. I cannot adjudicate without a legible exposition of how each architecture family instantiates each biological principle, including any criteria for correspondence.
minor comments (1)
- [Abstract] The abstract is readable and clear; the title matches the abstract. No further presentation issues can be assessed because the remainder is illegible. If the corruption is a rendering artifact, please resubmit a clean version.
Circularity Check
No significant circularity: the paper is a survey/taxonomy with no fitted parameters, derived equations, or load-bearing self-citations, so there is no derivation chain that reduces to its inputs.
full rationale
The manuscript is a systematic review and taxonomy paper. Its central claims are the proposed three-axis taxonomy (functional objectives, memory representations, integration mechanisms) and an interpretive trend statement that 'analysis ... reveals a shift from static caches toward adaptive, test-time learning systems.' Neither claim is derived from equations or fitted constants; the taxonomy is a classification scheme, and the trend is an inductive reading of the surveyed literature. There is no self-definitional step, no fitted input renamed as prediction, and no self-citation chain invoked to force a conclusion. The abstract and the legible fragments contain no derivation machinery for the trend claim, and the full text is largely mojibake, which makes the 'systematic' methodology unauditable; however, an inability to verify the survey corpus is an evidence/auditability concern, not a circularity of the kind defined here. The visible arXiv identifier (2508.10821v3 [q-bio.QM]) embedded in the text further indicates document corruption, but this does not constitute a logical loop. Accordingly, no specific circular step can be quoted with the required reduction, and the appropriate finding is no significant circularity.
Assumptions & free parameters
assumptions (2)
- domain assumption Cited papers report their results accurately
- domain assumption The neuroscience-to-engineering correspondence is substantive
Cite this review
Pith. "Pith review of Memory-Augmented Transformers: A Systematic Review from Neuroscience Principles to Enhanced Model Architectures." pith.science (2026). https://pith.science/paper/VDNDJEAA
@misc{pith2026250810824,
author = {Pith},
title = {Pith review of: Memory-Augmented Transformers: A Systematic Review from Neuroscience Principles to Enhanced Model Architectures},
year = {2026},
howpublished = {\url{https://pith.science/paper/VDNDJEAA}},
note = {Machine review of arXiv:2508.10824}
}
read the original abstract
Memory is fundamental to intelligence, enabling learning, reasoning, and adaptability across biological and artificial systems. While Transformer architectures excel at sequence modeling, they face critical limitations in long-range context retention, continual learning, and knowledge integration. This review presents a unified framework bridging neuroscience principles, including dynamic multi-timescale memory, selective attention, and consolidation, with engineering advances in Memory-Augmented Transformers. We organize recent progress through three taxonomic dimensions: functional objectives (context extension, reasoning, knowledge integration, adaptation), memory representations (parameter-encoded, state-based, explicit, hybrid), and integration mechanisms (attention fusion, gated control, associative retrieval). Our analysis of core memory operations (reading, writing, forgetting, and capacity management) reveals a shift from static caches toward adaptive, test-time learning systems. We identify persistent challenges in scalability and interference, alongside emerging solutions including hierarchical buffering and surprise-gated updates. This synthesis provides a roadmap toward cognitively-inspired, lifelong-learning Transformer architectures.
Forward citations
Cited by 1 Pith paper
-
Context-Aware Multi-Turn Visual-Textual Reasoning in LVLMs via Dynamic Memory and Adaptive Visual Guidance
The proposed CAMVR framework is not supported by verifiable evidence, and the manuscript itself labels its experimental results as fabricated.
Reference graph
Works this paper leans on
-
[1]
Adult hippocampal neurogenesis and cognitive flexibility—linking memory and mood
Christoph Anacker and Ren \'e Hen. Adult hippocampal neurogenesis and cognitive flexibility—linking memory and mood. Nature Reviews Neuroscience, 18 0 (6): 0 335--346, 2017
2017
-
[2]
Global workspace theory (gwt) and prefrontal cortex: Recent developments
Bernard J Baars, Natalie Geld, and Robert Kozma. Global workspace theory (gwt) and prefrontal cortex: Recent developments. Frontiers in psychology, 12: 0 749868, 2021
2021
-
[3]
Working memory: looking back and looking forward
Alan Baddeley. Working memory: looking back and looking forward. Nature reviews neuroscience, 4 0 (10): 0 829--839, 2003
2003
-
[4]
Fast adaptation to rule switching using neuronal surprise
Martin LLR Barry and Wulfram Gerstner. Fast adaptation to rule switching using neuronal surprise. PLoS computational biology, 20 0 (2): 0 e1011839, 2024
2024
-
[5]
Neuromodulators and long-term synaptic plasticity in learning and memory: A steered-glutamatergic perspective
Amjad H Bazzari and H Rheinallt Parri. Neuromodulators and long-term synaptic plasticity in learning and memory: A steered-glutamatergic perspective. Brain sciences, 9 0 (11): 0 300, 2019
2019
-
[6]
Titans: Learning to memorize at test time
Ali Behrouz, Peilin Zhong, and Vahab Mirrokni. Titans: Learning to memorize at test time. arXiv preprint arXiv:2501.00663, 2024
arXiv 2024
-
[7]
Atlas: Learning to optimally memorize the context at test time
Ali Behrouz, Zeman Li, Praneeth Kacham, Majid Daliri, Yuan Deng, Peilin Zhong, Meisam Razaviyayn, and Vahab Mirrokni. Atlas: Learning to optimally memorize the context at test time. arXiv preprint arXiv:2505.23735, 2025
arXiv 2025
-
[8]
Longformer: The long-document transformer
Iz Beltagy, Matthew E Peters, and Arman Cohan. Longformer: The long-document transformer. arXiv preprint arXiv:2004.05150, 2020
arXiv 2004
Show all 124 references
-
[9]
Memory layers at scale
Vincent-Pierre Berges, Barlas O g uz, Daniel Haziza, Wen-tau Yih, Luke Zettlemoyer, and Gargi Ghosh. Memory layers at scale. arXiv preprint arXiv:2412.09764, 2024
2024 arXiv
-
[10]
Improving language models by retrieving from trillions of tokens
Sebastian Borgeaud, Arthur Mensch, Jordan Hoffmann, Trevor Cai, Eliza Rutherford, Katie Millican, George Bm Van Den Driessche, Jean-Baptiste Lespiau, Bogdan Damoc, Aidan Clark, et al. Improving language models by retrieving from trillions of tokens. In International conference...
2022
-
[11]
The neuroanatomical, neurophysiological and psychological basis of memory: Current models and their origins
Eduardo Camina and Francisco G \"u ell. The neuroanatomical, neurophysiological and psychological basis of memory: Current models and their origins. Frontiers in pharmacology, 8: 0 438, 2017
2017
-
[12]
An evolved universal transformer memory
Edoardo Cetin, Qi Sun, Tianyu Zhao, and Yujin Tang. An evolved universal transformer memory. In The Thirteenth International Conference on Learning Representations, 2025
2025
-
[13]
Walking down the memory maze: Beyond context limit through interactive reading
Howard Chen, Ramakanth Pasunuru, Jason Weston, and Asli Celikyilmaz. Walking down the memory maze: Beyond context limit through interactive reading. arXiv preprint arXiv:2310.05029, 2023
2023 arXiv
-
[14]
Mem0: Building production-ready ai agents with scalable long-term memory
Prateek Chhikara, Dev Khant, Saket Aryan, Taranjeet Singh, and Deshraj Yadav. Mem0: Building production-ready ai agents with scalable long-term memory. arXiv preprint arXiv:2504.19413, 2025
2025 arXiv
-
[15]
Interactions between attention and memory
Marvin M Chun and Nicholas B Turk-Browne. Interactions between attention and memory. Current opinion in neurobiology, 17 0 (2): 0 177--184, 2007
2007
-
[16]
Attention-dependent coupling with forebrain and brainstem neuromodulatory nuclei differs across the lifespan
Nicholas G Cicero, Elizabeth Riley, Khena M Swallow, Eve De Rosa, and Adam Anderson. Attention-dependent coupling with forebrain and brainstem neuromodulatory nuclei differs across the lifespan. GeroScience, pp.\ 1--20, 2025
2025
-
[17]
What are the differences between long-term, short-term, and working memory? Progress in brain research, 169: 0 323--338, 2008
Nelson Cowan. What are the differences between long-term, short-term, and working memory? Progress in brain research, 169: 0 323--338, 2008
2008
-
[18]
Transformer-xl: Attentive language models beyond a fixed-length context
Zihang Dai, Zhilin Yang, Yiming Yang, Jaime Carbonell, Quoc V Le, and Ruslan Salakhutdinov. Transformer-xl: Attentive language models beyond a fixed-length context. arXiv preprint arXiv:1901.02860, 2019
1901 arXiv
-
[19]
The global neuronal workspace model of conscious access: from neuronal architectures to clinical applications
Stanislas Dehaene, Jean-Pierre Changeux, and Lionel Naccache. The global neuronal workspace model of conscious access: from neuronal architectures to clinical applications. Characterizing consciousness: From cognition to the clinic?, pp.\ 55--84, 2011
2011
-
[20]
Pronouns reactivate conceptual representations in human hippocampal neurons
Doris E Dijksterhuis, Matthew W Self, Jessy K Possel, Judith C Peters, ECW van Straaten, Sander Idema, Johannes C Baaijen, Sandra MA van der Salm, Erik J Aarnoutse, Nicole CE van Klink, et al. Pronouns reactivate conceptual representations in human hippocampal neurons. Science...
2024
-
[21]
Interaction between the amygdala and the medial temporal lobe memory system predicts better memory for emotional events
Florin Dolcos, Kevin S LaBar, and Roberto Cabeza. Interaction between the amygdala and the medial temporal lobe memory system predicts better memory for emotional events. Neuron, 42 0 (5): 0 855--863, 2004
2004
-
[22]
Rethinking memory in ai: Taxonomy, operations, topics, and future directions
Yiming Du, Wenyu Huang, Danna Zheng, Zhaowei Wang, Sebastien Montella, Mirella Lapata, Kam-Fai Wong, and Jeff Z Pan. Rethinking memory in ai: Taxonomy, operations, topics, and future directions. arXiv preprint arXiv:2505.00675, 2025
2025
-
[23]
Memory-augmented transformers can implement linear first-order optimization methods
Sanchayan Dutta and Suvrit Sra. Memory-augmented transformers can implement linear first-order optimization methods. arXiv preprint arXiv:2410.07263, 2024
2024 arXiv
-
[24]
Efficient llm inference using dynamic input pruning and cache-aware masking
Marco Federici, Davide Belli, Mart Van Baalen, Amir Jalalirad, Andrii Skliar, Bence Major, Markus Nagel, and Paul Whatmough. Efficient llm inference using dynamic input pruning and cache-aware masking. arXiv preprint arXiv:2412.01380, 2024
2024 arXiv
-
[25]
Human-like episodic memory for infinite context llms
Zafeirios Fountas, Martin A Benfeghoul, Adnan Oomerjee, Fenia Christopoulou, Gerasimos Lampouras, Haitham Bou-Ammar, and Jun Wang. Human-like episodic memory for infinite context llms. arXiv preprint arXiv:2407.09450, 2024
2024
-
[26]
Experiencing surprise: The temporal dynamics of its impact on memory
Darya Frank, Alex Kafkas, and Daniela Montaldi. Experiencing surprise: The temporal dynamics of its impact on memory. Journal of Neuroscience, 42 0 (33): 0 6435--6444, 2022
2022
-
[27]
An efficient context-dependent memory framework for llm-centric agents
Pengyu Gao, Jinming Zhao, Xinyue Chen, and Long Yilin. An efficient context-dependent memory framework for llm-centric agents. In Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technolo...
2025
-
[28]
Memotr: Long-term memory-augmented transformer for multi-object tracking
Ruopeng Gao and Limin Wang. Memotr: Long-term memory-augmented transformer for multi-object tracking. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pp.\ 9901--9910, 2023
2023
-
[29]
Top-down modulation: bridging selective attention and working memory
Adam Gazzaley and Anna C Nobre. Top-down modulation: bridging selective attention and working memory. Trends in cognitive sciences, 16 0 (2): 0 129--135, 2012
2012
-
[30]
zip2zip: Inference-time adaptive vocabularies for language models via token compression
Saibo Geng, Nathan Ranchin, Maxime Peyrard, Chris Wendler, Michael Gastpar, Robert West, et al. zip2zip: Inference-time adaptive vocabularies for language models via token compression. arXiv preprint arXiv:2506.01084, 2025
2025
-
[31]
The role of the ca3 hippocampal subregion in spatial memory: a process oriented behavioral assessment
Paul E Gilbert and Andrea M Brushfield. The role of the ca3 hippocampal subregion in spatial memory: a process oriented behavioral assessment. Progress in Neuro-Psychopharmacology and Biological Psychiatry, 33 0 (5): 0 774--781, 2009
2009
-
[32]
Neural turing machines
Alex Graves, Greg Wayne, and Ivo Danihelka. Neural turing machines. arXiv preprint arXiv:1410.5401, 2014
2014 arXiv
-
[33]
Hybrid computing using a neural network with dynamic external memory
Alex Graves, Greg Wayne, Malcolm Reynolds, Tim Harley, Ivo Danihelka, Agnieszka Grabska-Barwi \'n ska, Sergio G \'o mez Colmenarejo, Edward Grefenstette, Tiago Ramalho, John Agapiou, et al. Hybrid computing using a neural network with dynamic external memory. Nature, 538 0 (76...
2016
-
[34]
Hipporag: Neurobiologically inspired long-term memory for large language models
Bernal Jim \'e nez Guti \'e rrez, Yiheng Shu, Yu Gu, Michihiro Yasunaga, and Yu Su. Hipporag: Neurobiologically inspired long-term memory for large language models. In The Thirty-eighth Annual Conference on Neural Information Processing Systems, 2024
2024
-
[35]
Hierarchical process memory: memory as an integral component of information processing
Uri Hasson, Janice Chen, and Christopher J Honey. Hierarchical process memory: memory as an integral component of information processing. Trends in cognitive sciences, 19 0 (6): 0 304--313, 2015
2015
-
[36]
Memory matters: The need to improve long-term memory in llm-agents
Kostas Hatalis, Despina Christou, Joshua Myers, Steven Jones, Keith Lambert, Adam Amos-Binks, Zohreh Dannenhauer, and Dustin Dannenhauer. Memory matters: The need to improve long-term memory in llm-agents. In Proceedings of the AAAI Symposium Series, volume 2, pp.\ 277--280, 2023
2023
-
[37]
Hmt: Hierarchical memory transformer for efficient long context language processing
Zifan He, Yingqi Cao, Zongyue Qin, Neha Prakriya, Yizhou Sun, and Jason Cong. Hmt: Hierarchical memory transformer for efficient long context language processing. arXiv preprint arXiv:2405.06067, 2024 a
2024 arXiv
-
[38]
Human-inspired perspectives: A survey on ai long-term memory
Zihong He, Weizhe Lin, Hao Zheng, Fan Zhang, Matt W Jones, Laurence Aitchison, Xuhai Xu, Miao Liu, Per Ola Kristensson, and Junxiao Shen. Human-inspired perspectives: A survey on ai long-term memory. arXiv preprint arXiv:2411.00489, 2024 b
2024 arXiv
-
[39]
Replay bursts in humans coincide with activation of the default mode and parietal alpha networks
Cameron Higgins, Yunzhe Liu, Diego Vidaurre, Zeb Kurth-Nelson, Ray Dolan, Timothy Behrens, and Mark Woolrich. Replay bursts in humans coincide with activation of the default mode and parietal alpha networks. Neuron, 109 0 (5): 0 882--893, 2021
2021
-
[40]
Transformerfam: Feedback attention is working memory
Dongseong Hwang, Weiran Wang, Zhuoyuan Huo, Khe Chai Sim, and Pedro Moreno Mengibar. Transformerfam: Feedback attention is working memory. arXiv preprint arXiv:2404.09173, 2024
2024 arXiv
-
[41]
Memory os of ai agent
Jiazheng Kang, Mingming Ji, Zhe Zhao, and Ting Bai. Memory os of ai agent. arXiv preprint arXiv:2506.06326, 2025 a
2025 arXiv
-
[42]
Lm2: Large memory models
Jikun Kang, Wenqi Wu, Filippos Christianos, Alex J Chan, Fraser Greenlee, George Thomas, Marvin Purtorab, and Andy Toulis. Lm2: Large memory models. arXiv preprint arXiv:2502.06049, 2025 b
2025 arXiv
-
[43]
Distinguishing examples while building concepts in hippocampal and artificial networks
Louis Kang and Taro Toyoizumi. Distinguishing examples while building concepts in hippocampal and artificial networks. Nature Communications, 15 0 (1): 0 647, 2024
2024
-
[44]
Mechanisms of systems memory consolidation during sleep
Jens G Klinzing, Niels Niethard, and Jan Born. Mechanisms of systems memory consolidation during sleep. Nature neuroscience, 22 0 (10): 0 1598--1610, 2019
2019
-
[45]
Memreasoner: A memory-augmented llm architecture for multi-hop reasoning
Ching-Yun Ko, Sihui Dai, Payel Das, Georgios Kollias, Subhajit Chaudhury, and Aurelie Lozano. Memreasoner: A memory-augmented llm architecture for multi-hop reasoning. In The First Workshop on System-2 Reasoning at Scale, NeurIPS'24, 2024
2024
-
[46]
Flexible working memory through selective gating and attentional tagging
Wouter Kruijne, Sander M Bohte, Pieter R Roelfsema, and Christian NL Olivers. Flexible working memory through selective gating and attentional tagging. Neural Computation, 33 0 (1): 0 1--40, 2021
2021
-
[47]
Semantic memory: A review of methods, models, and current challenges
Abhilasha A Kumar. Semantic memory: A review of methods, models, and current challenges. Psychonomic bulletin & review, 28 0 (1): 0 40--80, 2021
2021
-
[48]
Self-attentive associative memory
Hung Le, Truyen Tran, and Svetha Venkatesh. Self-attentive associative memory. In International conference on machine learning, pp.\ 5682--5691. PMLR, 2020
2020
-
[49]
Reasoning under 1 billion: Memory-augmented reinforcement learning for large language models
Hung Le, Dai Do, Dung Nguyen, and Svetha Venkatesh. Reasoning under 1 billion: Memory-augmented reinforcement learning for large language models. arXiv preprint arXiv:2504.02273, 2025
2025 arXiv
-
[50]
Matter: Memory-augmented transformer using heterogeneous knowledge sources
Dongkyu Lee, Chandana Satya Prakash, Jack FitzGerald, and Jens Lehmann. Matter: Memory-augmented transformer using heterogeneous knowledge sources. arXiv preprint arXiv:2406.04670, 2024
2024 arXiv
-
[51]
An update on memory reconsolidation updating
Jonathan LC Lee, Karim Nader, and Daniela Schiller. An update on memory reconsolidation updating. Trends in cognitive sciences, 21 0 (7): 0 531--545, 2017
2017
-
[52]
u ttler, Mike Lewis, Wen-tau Yih, Tim Rockt \
Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich K \"u ttler, Mike Lewis, Wen-tau Yih, Tim Rockt \"a schel, et al. Retrieval-augmented generation for knowledge-intensive nlp tasks. Advances in neural information processing...
2020
-
[53]
Falcon: Feedback-driven adaptive long/short-term memory reinforced coding optimization system
Zeyuan Li, Yangfan He, Lewei He, Jianhui Wang, Tianyu Shi, Bin Lei, Yuchen Li, and Qiuwu Chen. Falcon: Feedback-driven adaptive long/short-term memory reinforced coding optimization system. arXiv preprint arXiv:2410.21349, 2024
2024 arXiv
-
[54]
Self-evolving agents with reflective and memory-augmented abilities
Xuechen Liang, Yangfan He, Yinghui Xia, Xinyuan Song, Jianhui Wang, Meiling Tao, Li Sun, Xinhang Yuan, Jiayi Su, Keqin Li, et al. Self-evolving agents with reflective and memory-augmented abilities. arXiv preprint arXiv:2409.00872, 2024
2024 arXiv
-
[55]
Advances and challenges in foundation agents: From brain-inspired intelligence to evolutionary, collaborative, and safe systems
Bang Liu, Xinfeng Li, Jiayi Zhang, Jinlin Wang, Tanjin He, Sirui Hong, Hongzhang Liu, Shaokun Zhang, Kaitao Song, Kunlun Zhu, et al. Advances and challenges in foundation agents: From brain-inspired intelligence to evolutionary, collaborative, and safe systems. arXiv preprint ...
2025 arXiv
-
[56]
Think-in-memory: Recalling and post-thinking enable llms with long-term memory
Lei Liu, Xiaoyan Yang, Yue Shen, Binbin Hu, Zhiqiang Zhang, Jinjie Gu, and Guannan Zhang. Think-in-memory: Recalling and post-thinking enable llms with long-term memory. arXiv preprint arXiv:2311.08719, 2023
2023 arXiv
-
[57]
Memlong: Memory-augmented retrieval for long text modeling
Weijie Liu, Zecheng Tang, Juntao Li, Kehai Chen, and Min Zhang. Memlong: Memory-augmented retrieval for long text modeling. arXiv preprint arXiv:2408.16967, 2024
2024 arXiv
-
[58]
Experience replay is associated with efficient nonlocal learning
Yunzhe Liu, Marcelo G Mattar, Timothy EJ Behrens, Nathaniel D Daw, and Raymond J Dolan. Experience replay is associated with efficient nonlocal learning. Science, 372 0 (6544): 0 eabf1357, 2021
2021
-
[59]
Acquiring new memories in neocortex of hippocampal-lesioned mice
Wenhan Luo, Di Yun, Yi Hu, Miaomiao Tian, Jiajun Yang, Yifan Xu, Yong Tang, Yang Zhan, Hong Xie, and Ji-Song Guan. Acquiring new memories in neocortex of hippocampal-lesioned mice. Nature communications, 13 0 (1): 0 1601, 2022
2022
-
[60]
Memory-augmented graph neural networks: A brain-inspired review
Guixiang Ma, Vy A Vo, Theodore L Willke, and Nesreen K Ahmed. Memory-augmented graph neural networks: A brain-inspired review. IEEE Transactions on Artificial Intelligence, 5 0 (5): 0 2011--2025, 2023
2011
-
[61]
The stability-plasticity dilemma: Investigating the continuum from catastrophic forgetting to age-limited learning effects, 2013
Martial Mermillod, Aur \'e lia Bugaiska, and Patrick Bonin. The stability-plasticity dilemma: Investigating the continuum from catastrophic forgetting to age-limited learning effects, 2013
2013
-
[62]
The magical number seven, plus or minus two: Some limits on our capacity for processing information
George A Miller. The magical number seven, plus or minus two: Some limits on our capacity for processing information. Psychological review, 63 0 (2): 0 81, 1956
1956
-
[63]
o ksal, Ayyoob Imani, Mohsen Fayyaz, and Hinrich Sch \
Ali Modarressi, Abdullatif K \"o ksal, Ayyoob Imani, Mohsen Fayyaz, and Hinrich Sch \"u tze. Memllm: Finetuning llms to use an explicit read-write memory. arXiv preprint arXiv:2404.11672, 2024
2024 arXiv
-
[64]
Episodic memory and beyond: the hippocampus and neocortex in transformation
Morris Moscovitch, Roberto Cabeza, Gordon Winocur, and Lynn Nadel. Episodic memory and beyond: the hippocampus and neocortex in transformation. Annual review of psychology, 67 0 (1): 0 105--134, 2016
2016
-
[65]
Memory: Neurobiological mechanisms and assessment
Swaleha Mujawar, Jaideep Patil, Bhushan Chaudhari, and Daniel Saldanha. Memory: Neurobiological mechanisms and assessment. Industrial psychiatry journal, 30 0 (Suppl 1): 0 S311--S314, 2021
2021
-
[66]
Most brain activity is ‘background noise’—and that’s upending our understanding of consciousness, 2021
Thomas Nail. Most brain activity is ‘background noise’—and that’s upending our understanding of consciousness, 2021
2021
-
[67]
Dynamic memory compression: Retrofitting llms for accelerated inference
Piotr Nawrot, Adrian a \'n cucki, Marcin Chochowski, David Tarjan, and Edoardo M Ponti. Dynamic memory compression: Retrofitting llms for accelerated inference. arXiv preprint arXiv:2403.09636, 2024
2024 arXiv
-
[68]
Functional role of gamma and theta oscillations in episodic memory
Erika Nyhus and Tim Curran. Functional role of gamma and theta oscillations in episodic memory. Neuroscience & Biobehavioral Reviews, 34 0 (7): 0 1023--1035, 2010
2010
-
[69]
Memgpt: Towards llms as operating systems
Charles Packer, Vivian Fang, Shishir\_G Patil, Kevin Lin, Sarah Wooders, and Joseph\_E Gonzalez. Memgpt: Towards llms as operating systems. 2023
2023
-
[70]
Abc: Attention with bounded-memory control
Hao Peng, Jungo Kasai, Nikolaos Pappas, Dani Yogatama, Zhaofeng Wu, Lingpeng Kong, Roy Schwartz, and Noah A Smith. Abc: Attention with bounded-memory control. arXiv preprint arXiv:2110.02488, 2021
2021 arXiv
-
[71]
Position: Episodic memory is the missing piece for long-term llm agents
Mathis Pink, Qinyuan Wu, Vy Ai Vo, Javier Turek, Jianing Mu, Alexander Huth, and Mariya Toneva. Position: Episodic memory is the missing piece for long-term llm agents. arXiv preprint arXiv:2502.06975, 2025
2025 arXiv
-
[72]
Train short, test long: Attention with linear biases enables input length extrapolation
Ofir Press, Noah A Smith, and Mike Lewis. Train short, test long: Attention with linear biases enables input length extrapolation. arXiv preprint arXiv:2108.12409, 2021
2021 arXiv
-
[73]
Neuromodulation of the feedforward dentate gyrus-ca3 microcircuit
Luke Y Prince, Travis J Bacon, Cezar M Tigaret, and Jack R Mellor. Neuromodulation of the feedforward dentate gyrus-ca3 microcircuit. Frontiers in synaptic neuroscience, 8: 0 32, 2016
2016
-
[74]
Layerwise recurrent router for mixture-of-experts
Zihan Qiu, Zeyu Huang, Shuang Cheng, Yizhi Zhou, Zili Wang, Ivan Titov, and Jie Fu. Layerwise recurrent router for mixture-of-experts. arXiv preprint arXiv:2408.06793, 2024
2024 arXiv
-
[75]
A multisensory perspective of working memory
Michel Quak, Raquel Elea London, and Durk Talsma. A multisensory perspective of working memory. Frontiers in human neuroscience, 9: 0 197, 2015
2015
-
[76]
Compressive transformers for long-range sequence modelling
Jack W Rae, Anna Potapenko, Siddhant M Jayakumar, and Timothy P Lillicrap. Compressive transformers for long-range sequence modelling. arXiv preprint arXiv:1911.05507, 2019
1911 arXiv
-
[77]
A default mode of brain function
Marcus E Raichle, Ann Mary MacLeod, Abraham Z Snyder, William J Powers, Debra A Gusnard, and Gordon L Shulman. A default mode of brain function. Proceedings of the national academy of sciences, 98 0 (2): 0 676--682, 2001
2001
-
[78]
Adaptive knowledge consolidation: A dynamic approach to mitigating catastrophic forgetting in text-based neural networks
J Ranjith and Santhi Baskaran. Adaptive knowledge consolidation: A dynamic approach to mitigating catastrophic forgetting in text-based neural networks. 2024
2024
-
[79]
Hierarchical dynamics as a macroscopic organizing principle of the human brain
Ryan V Raut, Abraham Z Snyder, and Marcus E Raichle. Hierarchical dynamics as a macroscopic organizing principle of the human brain. Proceedings of the National Academy of Sciences, 117 0 (34): 0 20890--20897, 2020
2020
-
[80]
Dissociating distinct cortical networks associated with subregions of the human medial temporal lobe using precision neuroimaging
Daniel Reznik, Robert Trampel, Nikolaus Weiskopf, Menno P Witter, and Christian F Doeller. Dissociating distinct cortical networks associated with subregions of the human medial temporal lobe using precision neuroimaging. Neuron, 111 0 (17): 0 2756--2772, 2023
2023
-
[81]
Associative recurrent memory transformer
Ivan Rodkin, Yuri Kuratov, Aydar Bulatov, and Mikhail Burtsev. Associative recurrent memory transformer. arXiv preprint arXiv:2407.04841, 2024
2024 arXiv
-
[82]
The mechanisms for pattern completion and pattern separation in the hippocampus
Edmund T Rolls. The mechanisms for pattern completion and pattern separation in the hippocampus. Frontiers in systems neuroscience, 7: 0 74, 2013
2013
-
[83]
Working memory and neural oscillations: alpha--gamma versus theta--gamma codes for distinct wm information? Trends in cognitive sciences, 18 0 (1): 0 16--25, 2014
Fr \'e d \'e ric Roux and Peter J Uhlhaas. Working memory and neural oscillations: alpha--gamma versus theta--gamma codes for distinct wm information? Trends in cognitive sciences, 18 0 (1): 0 16--25, 2014
2014
-
[84]
Deep learning needs a prefrontal cortex
Jacob Russin, Randall C O’Reilly, and Yoshua Bengio. Deep learning needs a prefrontal cortex. Work Bridging AI Cogn Sci, 107 0 (603-616): 0 1, 2020
2020
-
[85]
Open-ended instructable embodied agents with memory-augmented large language models
Gabriel Sarch, Yue Wu, Michael J Tarr, and Katerina Fragkiadaki. Open-ended instructable embodied agents with memory-augmented large language models. arXiv preprint arXiv:2310.15127, 2023
2023 arXiv
-
[86]
Retrieval-augmented decision transformer: External memory for in-context rl
Thomas Schmied, Fabian Paischer, Vihang Patil, Markus Hofmarcher, Razvan Pascanu, and Sepp Hochreiter. Retrieval-augmented decision transformer: External memory for in-context rl. arXiv preprint arXiv:2410.07071, 2024
2024 arXiv
-
[87]
Stress and multiple memory systems: from ‘thinking’to ‘doing’
Lars Schwabe and Oliver T Wolf. Stress and multiple memory systems: from ‘thinking’to ‘doing’. Trends in cognitive sciences, 17 0 (2): 0 60--68, 2013
2013
-
[88]
Cognitive memory in large language models
Lianlei Shan, Shixian Luo, Zezhou Zhu, Yu Yuan, and Yong Wu. Cognitive memory in large language models. arXiv preprint arXiv:2504.02441, 2025
2025 arXiv
-
[89]
The neural representations underlying asymmetric cross-modal prediction of words
Liang Shi, Chuqi Liu, Xiaojing Peng, Yifei Cao, Daniel A Levy, and Gui Xue. The neural representations underlying asymmetric cross-modal prediction of words. Human Brain Mapping, 44 0 (6): 0 2418--2435, 2023
2023
-
[90]
Prediction errors disrupt hippocampal representations and update episodic memories
Alyssa H Sinclair, Grace M Manalili, Iva K Brunec, R Alison Adcock, and Morgan D Barense. Prediction errors disrupt hippocampal representations and update episodic memories. Proceedings of the National Academy of Sciences, 118 0 (51): 0 e2117625118, 2021
2021
-
[91]
Memory consolidation
Larry R Squire, Lisa Genzel, John T Wixted, and Richard G Morris. Memory consolidation. Cold Spring Harbor perspectives in biology, 7 0 (8): 0 a021766, 2015
2015
-
[92]
Hippocampal pattern completion is linked to gamma power increases and alpha power decreases during recollection
Bernhard P Staresina, Sebastian Michelmann, Mathilde Bonnefond, Ole Jensen, Nikolai Axmacher, and Juergen Fell. Hippocampal pattern completion is linked to gamma power increases and alpha power decreases during recollection. elife, 5: 0 e17397, 2016
2016
-
[93]
Transformer-squared: Self-adaptive llms
Qi Sun, Edoardo Cetin, and Yujin Tang. Transformer-squared: Self-adaptive llms. In The Thirteenth International Conference on Learning Representations, 2025
2025
-
[94]
Associative transformer
Yuwei Sun, Hideya Ochiai, Zhirong Wu, Stephen Lin, and Ryota Kanai. Associative transformer. arXiv preprint arXiv:2309.12862, 2023
2023 arXiv
-
[95]
Hippocampal--prefrontal interactions during spatial d ecision-making
Lucas CS Tavares and Adriano BL Tort. Hippocampal--prefrontal interactions during spatial d ecision-making. Hippocampus, 32 0 (1): 0 38--54, 2022
2022
-
[96]
Transformer memory as a differentiable search index
Yi Tay, Vinh Tran, Mostafa Dehghani, Jianmo Ni, Dara Bahri, Harsh Mehta, Zhen Qin, Kai Hui, Zhe Zhao, Jai Gupta, et al. Transformer memory as a differentiable search index. Advances in Neural Information Processing Systems, 35: 0 21831--21843, 2022
2022
-
[97]
The hippocampal indexing theory and episodic memory: updating the index
Timothy J Teyler and Jerry W Rudy. The hippocampal indexing theory and episodic memory: updating the index. Hippocampus, 17 0 (12): 0 1158--1169, 2007
2007
-
[98]
Attention is all you need
Ashish Vaswani. Attention is all you need. Advances in neural information processing systems, 30: 0 I, 2017
2017
-
[99]
Hierarchical reasoning model
Guan Wang, Jin Li, Yuhao Sun, Xing Chen, Changling Liu, Yue Wu, Meng Lu, Sen Song, and Yasin Abbasi Yadkori. Hierarchical reasoning model. arXiv preprint arXiv:2506.21734, 2025 a
2025 arXiv
-
[100]
Schrodinger's memory: Large language models
Wei Wang and Qing Li. Schrodinger's memory: Large language models. arXiv preprint arXiv:2409.10482, 2024
2024 arXiv
-
[101]
Augmenting language models with long-term memory
Weizhi Wang, Li Dong, Hao Cheng, Xiaodong Liu, Xifeng Yan, Jianfeng Gao, and Furu Wei. Augmenting language models with long-term memory. Advances in Neural Information Processing Systems, 36: 0 74530--74543, 2023
2023
-
[102]
R3mem: Bridging memory retention and retrieval via reversible compression
Xiaoqiang Wang, Suyuchen Wang, Yun Zhu, and Bang Liu. R3mem: Bridging memory retention and retrieval via reversible compression. arXiv preprint arXiv:2502.15957, 2025 b
2025 arXiv
-
[103]
Beyond the limits: A survey of techniques to extend the context length in large language models
Xindi Wang, Mahsa Salmani, Parsa Omidi, Xiangyu Ren, Mehdi Rezagholizadeh, and Armaghan Eshaghi. Beyond the limits: A survey of techniques to extend the context length in large language models. arXiv preprint arXiv:2402.02244, 2024 a
2024 arXiv
-
[104]
Memoryllm: Towards self-updatable large language models
Yu Wang, Yifan Gao, Xiusi Chen, Haoming Jiang, Shiyang Li, Jingfeng Yang, Qingyu Yin, Zheng Li, Xian Li, Bing Yin, et al. Memoryllm: Towards self-updatable large language models. arXiv preprint arXiv:2402.04624, 2024 b
2024 arXiv
-
[105]
M+: Extending memoryllm with scalable long-term memory
Yu Wang, Dmitry Krotov, Yuanzhe Hu, Yifan Gao, Wangchunshu Zhou, Julian McAuley, Dan Gutfreund, Rogerio Feris, and Zexue He. M+: Extending memoryllm with scalable long-term memory. arXiv preprint arXiv:2502.00592, 2025 c
2025 arXiv
-
[106]
Jarvis-1: Open-world multi-task agents with memory-augmented multimodal language models
Zihao Wang, Shaofei Cai, Anji Liu, Yonggang Jin, Jinbing Hou, Bowei Zhang, Haowei Lin, Zhaofeng He, Zilong Zheng, Yaodong Yang, et al. Jarvis-1: Open-world multi-task agents with memory-augmented multimodal language models. IEEE Transactions on Pattern Analysis and Machine Int...
2024
-
[107]
Stateful memory-augmented transformers for efficient dialogue modeling
Qingyang Wu and Zhou Yu. Stateful memory-augmented transformers for efficient dialogue modeling. arXiv preprint arXiv:2209.07634, 2022
2022 arXiv
-
[108]
Memformer: A memory-augmented transformer for sequence modeling
Qingyang Wu, Zhenzhong Lan, Kun Qian, Jing Gu, Alborz Geramifard, and Zhou Yu. Memformer: A memory-augmented transformer for sequence modeling. arXiv preprint arXiv:2010.06891, 2020
2010 arXiv
-
[109]
The kanerva machine: A generative distributed memory
Yan Wu, Greg Wayne, Alex Graves, and Timothy Lillicrap. The kanerva machine: A generative distributed memory. arXiv preprint arXiv:1804.01756, 2018
2018 arXiv
-
[110]
From human memory to ai memory: A survey on memory mechanisms in the era of llms
Yaxiong Wu, Sheng Liang, Chen Zhang, Yichao Wang, Yongyue Zhang, Huifeng Guo, Ruiming Tang, and Yong Liu. From human memory to ai memory: A survey on memory mechanisms in the era of llms. arXiv preprint arXiv:2504.15965, 2025
2025 arXiv
-
[111]
Memorizing transformers
Yuhuai Wu, Markus N Rabe, DeLesley Hutchins, and Christian Szegedy. Memorizing transformers. arXiv preprint arXiv:2203.08913, 2022 a
2022 arXiv
-
[112]
An efficient memory-augmented transformer for knowledge-intensive nlp tasks
Yuxiang Wu, Yu Zhao, Baotian Hu, Pasquale Minervini, Pontus Stenetorp, and Sebastian Riedel. An efficient memory-augmented transformer for knowledge-intensive nlp tasks. arXiv preprint arXiv:2210.16773, 2022 b
2022 arXiv
-
[113]
A-mem: Agentic memory for llm agents
Wujiang Xu, Zujie Liang, Kai Mei, Hang Gao, Juntao Tan, and Yongfeng Zhang. A-mem: Agentic memory for llm agents. arXiv preprint arXiv:2502.12110, 2025
2025 arXiv
-
[114]
Adaptive computation with elastic input sequence
Fuzhao Xue, Valerii Likhosherstov, Anurag Arnab, Neil Houlsby, Mostafa Dehghani, and Yang You. Adaptive computation with elastic input sequence. In International Conference on Machine Learning, pp.\ 38971--38988. PMLR, 2023
2023
-
[115]
Memory3: Language modeling with explicit memory
Hongkang Yang, Zehao Lin, Wenjin Wang, Hao Wu, Zhiyu Li, Bo Tang, Wenqiang Wei, Jinbo Wang, Zeyun Tang, Shichao Song, et al. Memory3: Language modeling with explicit memory. arXiv preprint arXiv:2407.01178, 2024
2024 arXiv
-
[116]
Longer context, deeper thinking: Uncovering the role of long-context ability in reasoning
Wang Yang, Zirui Liu, Hongye Jin, Qingyu Yin, Vipin Chaudhary, and Xiaotian Han. Longer context, deeper thinking: Uncovering the role of long-context ability in reasoning. arXiv preprint arXiv:2505.17315, 2025
2025 arXiv
-
[117]
Score: Story coherence and retrieval enhancement for ai narratives
Qiang Yi, Yangfan He, Jianhui Wang, Xinyuan Song, Shiyao Qian, Xinhang Yuan, Miao Zhang, Li Sun, Keqin Li, Kuan Lu, et al. Score: Story coherence and retrieval enhancement for ai narratives. arXiv preprint arXiv:2503.23512, 2025
2025
-
[118]
Malt diffusion: Memory-augmented latent transformers for any-length video generation
Sihyun Yu, Meera Hahn, Dan Kondratyuk, Jinwoo Shin, Agrim Gupta, Jos \'e Lezama, Irfan Essa, David Ross, and Jonathan Huang. Malt diffusion: Memory-augmented latent transformers for any-length video generation. arXiv preprint arXiv:2502.12632, 2025
2025 arXiv
-
[119]
Peripheral memory for llms: Integration of sequential memory banks with adaptive querying
Songlin Zhai, Yuan Meng, Yongrui Chen, Yiwei Wang, and Guilin Qi. Peripheral memory for llms: Integration of sequential memory banks with adaptive querying. In Forty-second International Conference on Machine Learning, 2025
2025
-
[120]
A survey on the memory mechanism of large language model based agents
Zeyu Zhang, Xiaohe Bo, Chen Ma, Rui Li, Xu Chen, Quanyu Dai, Jieming Zhu, Zhenhua Dong, and Ji-Rong Wen. A survey on the memory mechanism of large language model based agents. arXiv preprint arXiv:2404.13501, 2024
2024 arXiv
-
[121]
Memorybank: Enhancing large language models with long-term memory
Wanjun Zhong, Lianghong Guo, Qiqi Gao, He Ye, and Yanlin Wang. Memorybank: Enhancing large language models with long-term memory. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 38, pp.\ 19724--19731, 2024 a
2024
-
[122]
Random tree model of meaningful memory
Weishun Zhong, Tankut Can, Antonis Georgiou, Ilya Shnayderman, Mikhail Katkov, and Misha Tsodyks. Random tree model of meaningful memory. bioRxiv, pp.\ 2024--12, 2024 b
2024
-
[123]
Mlkv: Multi-layer key-value heads for memory efficient transformer decoding
Zayd Muhammad Kawakibi Zuhri, Muhammad Farid Adilazuarda, Ayu Purwarianti, and Alham Fikri Aji. Mlkv: Multi-layer key-value heads for memory efficient transformer decoding. arXiv preprint arXiv:2406.09297, 2024
2024 arXiv
-
[124]
write newline
" write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 gl...
Reviewed August 5, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.