Pith. sign in

REVIEW 4 major objections 1 minor 1 cited by

Memory-Augmented Transformers: A Systematic Review from Neuroscience Principles to Enhanced Model Architectures

T0 review · 4 major / 1 minor · reviewed 2026-08-05 · deepseek-v4-flash

Pith's one-line read A three-axis taxonomy organizes Memory-Augmented Transformers and shows a shift from static caches to adaptive test-time learning.

desk verdict A coherent survey abstract atop an unreadable body: worth sending back for a clean copy and a methods section, not yet citable. read the letter →

arxiv 2508.10824 v2 pith:VDNDJEAA submitted 2025-08-14 cs.LG cs.CL

classification cs.LGcs.CL
keywords memory-augmentedTransformersneuroscience-inspiredAIthree-axistaxonomytest-timelearninglifelongsequencemodelingconsolidationsurvey
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This review tries to establish that the scattered engineering of memory-augmented Transformers is not a bag of tricks: recent work can be organized along three axes—what the memory is for, how it is represented, and how it is integrated—and those axes, read through neuroscience ideas about multi-timescale memory, selective attention, and consolidation, show a field-level trajectory from static caches toward adaptive, test-time learning. A sympathetic reader would care because if the taxonomy holds, it gives researchers a shared vocabulary for comparing architectures and points the design space toward lifelong-learning Transformers, with hierarchical buffering and surprise-gated updates as explicit next steps. The paper's contribution is organizational and diagnostic, not a new model: it claims to show where the field has been and where it is heading.

What carries the argument

The organizing device is the taxonomy itself: each memory-augmented Transformer is located by a triple—objective, representation, integration. The four memory operations (reading, writing, forgetting, capacity management) are the analytic lens that turns the taxonomy from a static list into a diagnosis of what each architecture can and cannot do. The neuroscience mapping supplies the design rationale: multi-timescale memory justifies separate memory stores, selective attention justifies content-based retrieval, and consolidation justifies writing and forgetting rules—so the taxonomy is meant to be explanatory, not merely descriptive.

What would settle it

Take a defined literature window, for example memory-augmented Transformer papers published 2023–2025, collect the full set with a fixed search query, and try to assign each paper to exactly one cell of the three-axis taxonomy. If more than a small fraction resist placement, or the full set shows no chronological trend toward test-time learning, the central claims fail.

Watch

Extended reading notes

Core claim

The paper's central claim is that the diverse memory mechanisms bolted onto Transformers form a coherent design space, not a collection of unrelated patches. It proposes a three-axis taxonomy: functional objectives (context extension, reasoning, knowledge integration, adaptation), memory representations (parameter-encoded, state-based, explicit, hybrid), and integration mechanisms (attention fusion, gated control, associative retrieval). Using reading, writing, forgetting, and capacity management as the core memory operations, it argues that the field is shifting from fixed external caches toward adaptive systems that update their own memory at test time. It further claims that neuroscience

Load-bearing premise

The framework stands or falls on whether the neuroscience constructs genuinely correspond to the engineering categories and whether the surveyed papers were selected systematically rather than because they fit the story.

Editorial extensions

If this is right

  • Architectures that currently look incomparable—a long-context cache, a gated state model, an external associative memory—become comparable through their position on the three axes.
  • The claimed shift to adaptive, test-time learning sets a concrete design target: future models should write and update memory during inference, not only at training time.
  • Scalability and interference are named as the two bottlenecks any new memory design must address, so progress on them can be measured directly.
  • Hierarchical buffering and surprise-gated updates are named as emerging mechanisms, giving practitioners concrete starting points for new architectures.
  • The neuroscience link provides a vocabulary for generating new designs, such as treating a Transformer's memory as a multi-timescale system with distinct stores.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Editorial inference: the taxonomy's usefulness depends on whether it can classify papers independently; a natural test is to take a held-out set of recent memory-augmented Transformer papers and see whether two annotators place them in the same taxonomy cells.
  • Editorial inference: if the shift-to-test-time-learning claim is right, evaluation practice should change—benchmarks that only test static retrieval will miss the capabilities that differentiate the newer architectures.
  • Editorial inference: the paper asserts rather than demonstrates its 'systematic' coverage; a reader should look for a reproducible search protocol, because without one the trend claim could reflect selection bias.
  • Editorial inference: the memory-operations framework likely generalizes beyond Transformers to any architecture with an explicit memory component, so the roadmap could also apply to state-space models or recurrent networks.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 1 minor

Summary. The manuscript (arXiv:2508.10824) claims to provide a unified framework bridging neuroscience principles with engineering advances in Memory-Augmented Transformers, organized along three taxonomic dimensions (functional objectives, memory representations, integration mechanisms), and to identify a shift from static caches toward adaptive, test-time learning systems. The abstract is self-consistent and the claims are appropriately modest for a review. However, the supplied full text is heavily corrupted (mojibake) and unreadable; it also contains an inserted fragment from a different paper (arXiv:2508.10821v3 [q-bio.QM]). Consequently, none of the substantive content—sections, tables, equations, references, or the survey corpus—can be verified. The 'systematic' nature of the review and the empirical trend claim cannot be audited from the provided text.

Significance. If the claims hold, the three-axis taxonomy could provide a useful common vocabulary for memory-augmented transformers and a roadmap toward continual-learning architectures. The synthesis of neuroscience constructs and engineering designs could be valuable if the mapping is substantive. However, because the body is unreadable, the significance is conditional. The paper does not derive new mathematical results or ship code; its contribution, if any, is organizational and interpretive. Credit is due for a self-consistent abstract and clearly stated dimensions, but no machine-checkable evidence is available.

major comments (4)
  1. [Full text (after abstract)] The entire body of the manuscript is encoded mojibake; no section, table, equation, or reference can be read or verified. For a systematic review, the central claims—the taxonomy and the trend toward adaptive test-time learning—rest on the ability to check which papers were surveyed and how they were mapped. This is load-bearing and prevents any audit. I cannot determine whether the summaries are accurate or whether the conclusions follow from the cited literature.
  2. [Visible fragment in full text] The text includes the line 'arXiv:2508.10821v3 [q-bio.QM] 20 May 2026', which is an identifier for a different paper in quantitative biology. This indicates that the PDF/source has been corrupted or misassembled. As supplied, the document is not self-consistent; it contains content that does not belong to the claimed review. This is a material issue because it makes it impossible to attribute any of the surrounding text to the authors' intended manuscript.
  3. [Search protocol/inclusion criteria] The abstract calls the review 'systematic,' but no search strategy, inclusion/exclusion criteria, database sources, or PRISMA-style flow is visible in the supplied text. The trend claim ('a shift from static caches toward adaptive, test-time learning systems') is an empirical statement about the surveyed literature; without a reproducible corpus and coding protocol, it may be a cherry-picked narrative. This is an audibility problem, not necessarily an error, but it must be fixed for the review to be evaluable.
  4. [Neuroscience-to-engineering mapping] The 'unified framework' is the paper's main contribution, but the body that would justify the correspondence between biological constructs (multi-timescale memory, selective attention, consolidation) and architectural families (parameter-encoded, state-based, explicit, hybrid) is unreadable. This raises a correctness risk: the mapping could be substantive constraints or post-hoc labels. I cannot adjudicate without a legible exposition of how each architecture family instantiates each biological principle, including any criteria for correspondence.
minor comments (1)
  1. [Abstract] The abstract is readable and clear; the title matches the abstract. No further presentation issues can be assessed because the remainder is illegible. If the corruption is a rendering artifact, please resubmit a clean version.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the paper is a survey/taxonomy with no fitted parameters, derived equations, or load-bearing self-citations, so there is no derivation chain that reduces to its inputs.

full rationale

The manuscript is a systematic review and taxonomy paper. Its central claims are the proposed three-axis taxonomy (functional objectives, memory representations, integration mechanisms) and an interpretive trend statement that 'analysis ... reveals a shift from static caches toward adaptive, test-time learning systems.' Neither claim is derived from equations or fitted constants; the taxonomy is a classification scheme, and the trend is an inductive reading of the surveyed literature. There is no self-definitional step, no fitted input renamed as prediction, and no self-citation chain invoked to force a conclusion. The abstract and the legible fragments contain no derivation machinery for the trend claim, and the full text is largely mojibake, which makes the 'systematic' methodology unauditable; however, an inability to verify the survey corpus is an evidence/auditability concern, not a circularity of the kind defined here. The visible arXiv identifier (2508.10821v3 [q-bio.QM]) embedded in the text further indicates document corruption, but this does not constitute a logical loop. Accordingly, no specific circular step can be quoted with the required reduction, and the appropriate finding is no significant circularity.

Assumptions & free parameters 0 free parameters · 2 assumptions · 0 invented entities

As a review, the paper introduces no fitted numbers and no new physical or architectural entities. Its central claims rest on the accuracy of the primary literature it synthesizes and on the validity of the neuroscience mapping. With the full text unreadable, neither can be audited.

assumptions (2)
  • domain assumption Cited papers report their results accurately
    The survey's summaries and trend claim inherit the correctness of the primary sources it synthesizes; no independent replication is possible from the provided text.
  • domain assumption The neuroscience-to-engineering correspondence is substantive
    The 'unified framework' in the abstract depends on multi-timescale memory, selective attention, and consolidation mapping onto the named architectural families as real design counterparts rather than as loose metaphor.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Memory-Augmented Transformers: A Systematic Review from Neuroscience Principles to Enhanced Model Architectures." pith.science (2026). https://pith.science/paper/VDNDJEAA

@misc{pith2026250810824,
  author       = {Pith},
  title        = {Pith review of: Memory-Augmented Transformers: A Systematic Review from Neuroscience Principles to Enhanced Model Architectures},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/VDNDJEAA}},
  note         = {Machine review of arXiv:2508.10824}
}
read the original abstract

Memory is fundamental to intelligence, enabling learning, reasoning, and adaptability across biological and artificial systems. While Transformer architectures excel at sequence modeling, they face critical limitations in long-range context retention, continual learning, and knowledge integration. This review presents a unified framework bridging neuroscience principles, including dynamic multi-timescale memory, selective attention, and consolidation, with engineering advances in Memory-Augmented Transformers. We organize recent progress through three taxonomic dimensions: functional objectives (context extension, reasoning, knowledge integration, adaptation), memory representations (parameter-encoded, state-based, explicit, hybrid), and integration mechanisms (attention fusion, gated control, associative retrieval). Our analysis of core memory operations (reading, writing, forgetting, and capacity management) reveals a shift from static caches toward adaptive, test-time learning systems. We identify persistent challenges in scalability and interference, alongside emerging solutions including hierarchical buffering and surprise-gated updates. This synthesis provides a roadmap toward cognitively-inspired, lifelong-learning Transformer architectures.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Context-Aware Multi-Turn Visual-Textual Reasoning in LVLMs via Dynamic Memory and Adaptive Visual Guidance

    cs.CV 2025-09 reject novelty 3.0 of 10

    The proposed CAMVR framework is not supported by verifiable evidence, and the manuscript itself labels its experimental results as fabricated.

Reference graph

Works this paper leans on

124 extracted references · 33 canonical work pages · cited by 1 Pith paper

  1. [1]

    Adult hippocampal neurogenesis and cognitive flexibility—linking memory and mood

    Christoph Anacker and Ren \'e Hen. Adult hippocampal neurogenesis and cognitive flexibility—linking memory and mood. Nature Reviews Neuroscience, 18 0 (6): 0 335--346, 2017

  2. [2]

    Global workspace theory (gwt) and prefrontal cortex: Recent developments

    Bernard J Baars, Natalie Geld, and Robert Kozma. Global workspace theory (gwt) and prefrontal cortex: Recent developments. Frontiers in psychology, 12: 0 749868, 2021

  3. [3]

    Working memory: looking back and looking forward

    Alan Baddeley. Working memory: looking back and looking forward. Nature reviews neuroscience, 4 0 (10): 0 829--839, 2003

  4. [4]

    Fast adaptation to rule switching using neuronal surprise

    Martin LLR Barry and Wulfram Gerstner. Fast adaptation to rule switching using neuronal surprise. PLoS computational biology, 20 0 (2): 0 e1011839, 2024

  5. [5]

    Neuromodulators and long-term synaptic plasticity in learning and memory: A steered-glutamatergic perspective

    Amjad H Bazzari and H Rheinallt Parri. Neuromodulators and long-term synaptic plasticity in learning and memory: A steered-glutamatergic perspective. Brain sciences, 9 0 (11): 0 300, 2019

  6. [6]

    Titans: Learning to memorize at test time

    Ali Behrouz, Peilin Zhong, and Vahab Mirrokni. Titans: Learning to memorize at test time. arXiv preprint arXiv:2501.00663, 2024

  7. [7]

    Atlas: Learning to optimally memorize the context at test time

    Ali Behrouz, Zeman Li, Praneeth Kacham, Majid Daliri, Yuan Deng, Peilin Zhong, Meisam Razaviyayn, and Vahab Mirrokni. Atlas: Learning to optimally memorize the context at test time. arXiv preprint arXiv:2505.23735, 2025

  8. [8]

    Longformer: The long-document transformer

    Iz Beltagy, Matthew E Peters, and Arman Cohan. Longformer: The long-document transformer. arXiv preprint arXiv:2004.05150, 2020

Show all 124 references
  1. [9]

    Memory layers at scale

    Vincent-Pierre Berges, Barlas O g uz, Daniel Haziza, Wen-tau Yih, Luke Zettlemoyer, and Gargi Ghosh. Memory layers at scale. arXiv preprint arXiv:2412.09764, 2024

  2. [10]

    Improving language models by retrieving from trillions of tokens

    Sebastian Borgeaud, Arthur Mensch, Jordan Hoffmann, Trevor Cai, Eliza Rutherford, Katie Millican, George Bm Van Den Driessche, Jean-Baptiste Lespiau, Bogdan Damoc, Aidan Clark, et al. Improving language models by retrieving from trillions of tokens. In International conference...

  3. [11]

    The neuroanatomical, neurophysiological and psychological basis of memory: Current models and their origins

    Eduardo Camina and Francisco G \"u ell. The neuroanatomical, neurophysiological and psychological basis of memory: Current models and their origins. Frontiers in pharmacology, 8: 0 438, 2017

  4. [12]

    An evolved universal transformer memory

    Edoardo Cetin, Qi Sun, Tianyu Zhao, and Yujin Tang. An evolved universal transformer memory. In The Thirteenth International Conference on Learning Representations, 2025

  5. [13]

    Walking down the memory maze: Beyond context limit through interactive reading

    Howard Chen, Ramakanth Pasunuru, Jason Weston, and Asli Celikyilmaz. Walking down the memory maze: Beyond context limit through interactive reading. arXiv preprint arXiv:2310.05029, 2023

  6. [14]

    Mem0: Building production-ready ai agents with scalable long-term memory

    Prateek Chhikara, Dev Khant, Saket Aryan, Taranjeet Singh, and Deshraj Yadav. Mem0: Building production-ready ai agents with scalable long-term memory. arXiv preprint arXiv:2504.19413, 2025

  7. [15]

    Interactions between attention and memory

    Marvin M Chun and Nicholas B Turk-Browne. Interactions between attention and memory. Current opinion in neurobiology, 17 0 (2): 0 177--184, 2007

  8. [16]

    Attention-dependent coupling with forebrain and brainstem neuromodulatory nuclei differs across the lifespan

    Nicholas G Cicero, Elizabeth Riley, Khena M Swallow, Eve De Rosa, and Adam Anderson. Attention-dependent coupling with forebrain and brainstem neuromodulatory nuclei differs across the lifespan. GeroScience, pp.\ 1--20, 2025

  9. [17]

    What are the differences between long-term, short-term, and working memory? Progress in brain research, 169: 0 323--338, 2008

    Nelson Cowan. What are the differences between long-term, short-term, and working memory? Progress in brain research, 169: 0 323--338, 2008

  10. [18]

    Transformer-xl: Attentive language models beyond a fixed-length context

    Zihang Dai, Zhilin Yang, Yiming Yang, Jaime Carbonell, Quoc V Le, and Ruslan Salakhutdinov. Transformer-xl: Attentive language models beyond a fixed-length context. arXiv preprint arXiv:1901.02860, 2019

  11. [19]

    The global neuronal workspace model of conscious access: from neuronal architectures to clinical applications

    Stanislas Dehaene, Jean-Pierre Changeux, and Lionel Naccache. The global neuronal workspace model of conscious access: from neuronal architectures to clinical applications. Characterizing consciousness: From cognition to the clinic?, pp.\ 55--84, 2011

  12. [20]

    Pronouns reactivate conceptual representations in human hippocampal neurons

    Doris E Dijksterhuis, Matthew W Self, Jessy K Possel, Judith C Peters, ECW van Straaten, Sander Idema, Johannes C Baaijen, Sandra MA van der Salm, Erik J Aarnoutse, Nicole CE van Klink, et al. Pronouns reactivate conceptual representations in human hippocampal neurons. Science...

  13. [21]

    Interaction between the amygdala and the medial temporal lobe memory system predicts better memory for emotional events

    Florin Dolcos, Kevin S LaBar, and Roberto Cabeza. Interaction between the amygdala and the medial temporal lobe memory system predicts better memory for emotional events. Neuron, 42 0 (5): 0 855--863, 2004

  14. [22]

    Rethinking memory in ai: Taxonomy, operations, topics, and future directions

    Yiming Du, Wenyu Huang, Danna Zheng, Zhaowei Wang, Sebastien Montella, Mirella Lapata, Kam-Fai Wong, and Jeff Z Pan. Rethinking memory in ai: Taxonomy, operations, topics, and future directions. arXiv preprint arXiv:2505.00675, 2025

  15. [23]

    Memory-augmented transformers can implement linear first-order optimization methods

    Sanchayan Dutta and Suvrit Sra. Memory-augmented transformers can implement linear first-order optimization methods. arXiv preprint arXiv:2410.07263, 2024

  16. [24]

    Efficient llm inference using dynamic input pruning and cache-aware masking

    Marco Federici, Davide Belli, Mart Van Baalen, Amir Jalalirad, Andrii Skliar, Bence Major, Markus Nagel, and Paul Whatmough. Efficient llm inference using dynamic input pruning and cache-aware masking. arXiv preprint arXiv:2412.01380, 2024

  17. [25]

    Human-like episodic memory for infinite context llms

    Zafeirios Fountas, Martin A Benfeghoul, Adnan Oomerjee, Fenia Christopoulou, Gerasimos Lampouras, Haitham Bou-Ammar, and Jun Wang. Human-like episodic memory for infinite context llms. arXiv preprint arXiv:2407.09450, 2024

  18. [26]

    Experiencing surprise: The temporal dynamics of its impact on memory

    Darya Frank, Alex Kafkas, and Daniela Montaldi. Experiencing surprise: The temporal dynamics of its impact on memory. Journal of Neuroscience, 42 0 (33): 0 6435--6444, 2022

  19. [27]

    An efficient context-dependent memory framework for llm-centric agents

    Pengyu Gao, Jinming Zhao, Xinyue Chen, and Long Yilin. An efficient context-dependent memory framework for llm-centric agents. In Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technolo...

  20. [28]

    Memotr: Long-term memory-augmented transformer for multi-object tracking

    Ruopeng Gao and Limin Wang. Memotr: Long-term memory-augmented transformer for multi-object tracking. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pp.\ 9901--9910, 2023

  21. [29]

    Top-down modulation: bridging selective attention and working memory

    Adam Gazzaley and Anna C Nobre. Top-down modulation: bridging selective attention and working memory. Trends in cognitive sciences, 16 0 (2): 0 129--135, 2012

  22. [30]

    zip2zip: Inference-time adaptive vocabularies for language models via token compression

    Saibo Geng, Nathan Ranchin, Maxime Peyrard, Chris Wendler, Michael Gastpar, Robert West, et al. zip2zip: Inference-time adaptive vocabularies for language models via token compression. arXiv preprint arXiv:2506.01084, 2025

  23. [31]

    The role of the ca3 hippocampal subregion in spatial memory: a process oriented behavioral assessment

    Paul E Gilbert and Andrea M Brushfield. The role of the ca3 hippocampal subregion in spatial memory: a process oriented behavioral assessment. Progress in Neuro-Psychopharmacology and Biological Psychiatry, 33 0 (5): 0 774--781, 2009

  24. [32]

    Neural turing machines

    Alex Graves, Greg Wayne, and Ivo Danihelka. Neural turing machines. arXiv preprint arXiv:1410.5401, 2014

  25. [33]

    Hybrid computing using a neural network with dynamic external memory

    Alex Graves, Greg Wayne, Malcolm Reynolds, Tim Harley, Ivo Danihelka, Agnieszka Grabska-Barwi \'n ska, Sergio G \'o mez Colmenarejo, Edward Grefenstette, Tiago Ramalho, John Agapiou, et al. Hybrid computing using a neural network with dynamic external memory. Nature, 538 0 (76...

  26. [34]

    Hipporag: Neurobiologically inspired long-term memory for large language models

    Bernal Jim \'e nez Guti \'e rrez, Yiheng Shu, Yu Gu, Michihiro Yasunaga, and Yu Su. Hipporag: Neurobiologically inspired long-term memory for large language models. In The Thirty-eighth Annual Conference on Neural Information Processing Systems, 2024

  27. [35]

    Hierarchical process memory: memory as an integral component of information processing

    Uri Hasson, Janice Chen, and Christopher J Honey. Hierarchical process memory: memory as an integral component of information processing. Trends in cognitive sciences, 19 0 (6): 0 304--313, 2015

  28. [36]

    Memory matters: The need to improve long-term memory in llm-agents

    Kostas Hatalis, Despina Christou, Joshua Myers, Steven Jones, Keith Lambert, Adam Amos-Binks, Zohreh Dannenhauer, and Dustin Dannenhauer. Memory matters: The need to improve long-term memory in llm-agents. In Proceedings of the AAAI Symposium Series, volume 2, pp.\ 277--280, 2023

  29. [37]

    Hmt: Hierarchical memory transformer for efficient long context language processing

    Zifan He, Yingqi Cao, Zongyue Qin, Neha Prakriya, Yizhou Sun, and Jason Cong. Hmt: Hierarchical memory transformer for efficient long context language processing. arXiv preprint arXiv:2405.06067, 2024 a

  30. [38]

    Human-inspired perspectives: A survey on ai long-term memory

    Zihong He, Weizhe Lin, Hao Zheng, Fan Zhang, Matt W Jones, Laurence Aitchison, Xuhai Xu, Miao Liu, Per Ola Kristensson, and Junxiao Shen. Human-inspired perspectives: A survey on ai long-term memory. arXiv preprint arXiv:2411.00489, 2024 b

  31. [39]

    Replay bursts in humans coincide with activation of the default mode and parietal alpha networks

    Cameron Higgins, Yunzhe Liu, Diego Vidaurre, Zeb Kurth-Nelson, Ray Dolan, Timothy Behrens, and Mark Woolrich. Replay bursts in humans coincide with activation of the default mode and parietal alpha networks. Neuron, 109 0 (5): 0 882--893, 2021

  32. [40]

    Transformerfam: Feedback attention is working memory

    Dongseong Hwang, Weiran Wang, Zhuoyuan Huo, Khe Chai Sim, and Pedro Moreno Mengibar. Transformerfam: Feedback attention is working memory. arXiv preprint arXiv:2404.09173, 2024

  33. [41]

    Memory os of ai agent

    Jiazheng Kang, Mingming Ji, Zhe Zhao, and Ting Bai. Memory os of ai agent. arXiv preprint arXiv:2506.06326, 2025 a

  34. [42]

    Lm2: Large memory models

    Jikun Kang, Wenqi Wu, Filippos Christianos, Alex J Chan, Fraser Greenlee, George Thomas, Marvin Purtorab, and Andy Toulis. Lm2: Large memory models. arXiv preprint arXiv:2502.06049, 2025 b

  35. [43]

    Distinguishing examples while building concepts in hippocampal and artificial networks

    Louis Kang and Taro Toyoizumi. Distinguishing examples while building concepts in hippocampal and artificial networks. Nature Communications, 15 0 (1): 0 647, 2024

  36. [44]

    Mechanisms of systems memory consolidation during sleep

    Jens G Klinzing, Niels Niethard, and Jan Born. Mechanisms of systems memory consolidation during sleep. Nature neuroscience, 22 0 (10): 0 1598--1610, 2019

  37. [45]

    Memreasoner: A memory-augmented llm architecture for multi-hop reasoning

    Ching-Yun Ko, Sihui Dai, Payel Das, Georgios Kollias, Subhajit Chaudhury, and Aurelie Lozano. Memreasoner: A memory-augmented llm architecture for multi-hop reasoning. In The First Workshop on System-2 Reasoning at Scale, NeurIPS'24, 2024

  38. [46]

    Flexible working memory through selective gating and attentional tagging

    Wouter Kruijne, Sander M Bohte, Pieter R Roelfsema, and Christian NL Olivers. Flexible working memory through selective gating and attentional tagging. Neural Computation, 33 0 (1): 0 1--40, 2021

  39. [47]

    Semantic memory: A review of methods, models, and current challenges

    Abhilasha A Kumar. Semantic memory: A review of methods, models, and current challenges. Psychonomic bulletin & review, 28 0 (1): 0 40--80, 2021

  40. [48]

    Self-attentive associative memory

    Hung Le, Truyen Tran, and Svetha Venkatesh. Self-attentive associative memory. In International conference on machine learning, pp.\ 5682--5691. PMLR, 2020

  41. [49]

    Reasoning under 1 billion: Memory-augmented reinforcement learning for large language models

    Hung Le, Dai Do, Dung Nguyen, and Svetha Venkatesh. Reasoning under 1 billion: Memory-augmented reinforcement learning for large language models. arXiv preprint arXiv:2504.02273, 2025

  42. [50]

    Matter: Memory-augmented transformer using heterogeneous knowledge sources

    Dongkyu Lee, Chandana Satya Prakash, Jack FitzGerald, and Jens Lehmann. Matter: Memory-augmented transformer using heterogeneous knowledge sources. arXiv preprint arXiv:2406.04670, 2024

  43. [51]

    An update on memory reconsolidation updating

    Jonathan LC Lee, Karim Nader, and Daniela Schiller. An update on memory reconsolidation updating. Trends in cognitive sciences, 21 0 (7): 0 531--545, 2017

  44. [52]

    u ttler, Mike Lewis, Wen-tau Yih, Tim Rockt \

    Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich K \"u ttler, Mike Lewis, Wen-tau Yih, Tim Rockt \"a schel, et al. Retrieval-augmented generation for knowledge-intensive nlp tasks. Advances in neural information processing...

  45. [53]

    Falcon: Feedback-driven adaptive long/short-term memory reinforced coding optimization system

    Zeyuan Li, Yangfan He, Lewei He, Jianhui Wang, Tianyu Shi, Bin Lei, Yuchen Li, and Qiuwu Chen. Falcon: Feedback-driven adaptive long/short-term memory reinforced coding optimization system. arXiv preprint arXiv:2410.21349, 2024

  46. [54]

    Self-evolving agents with reflective and memory-augmented abilities

    Xuechen Liang, Yangfan He, Yinghui Xia, Xinyuan Song, Jianhui Wang, Meiling Tao, Li Sun, Xinhang Yuan, Jiayi Su, Keqin Li, et al. Self-evolving agents with reflective and memory-augmented abilities. arXiv preprint arXiv:2409.00872, 2024

  47. [55]

    Advances and challenges in foundation agents: From brain-inspired intelligence to evolutionary, collaborative, and safe systems

    Bang Liu, Xinfeng Li, Jiayi Zhang, Jinlin Wang, Tanjin He, Sirui Hong, Hongzhang Liu, Shaokun Zhang, Kaitao Song, Kunlun Zhu, et al. Advances and challenges in foundation agents: From brain-inspired intelligence to evolutionary, collaborative, and safe systems. arXiv preprint ...

  48. [56]

    Think-in-memory: Recalling and post-thinking enable llms with long-term memory

    Lei Liu, Xiaoyan Yang, Yue Shen, Binbin Hu, Zhiqiang Zhang, Jinjie Gu, and Guannan Zhang. Think-in-memory: Recalling and post-thinking enable llms with long-term memory. arXiv preprint arXiv:2311.08719, 2023

  49. [57]

    Memlong: Memory-augmented retrieval for long text modeling

    Weijie Liu, Zecheng Tang, Juntao Li, Kehai Chen, and Min Zhang. Memlong: Memory-augmented retrieval for long text modeling. arXiv preprint arXiv:2408.16967, 2024

  50. [58]

    Experience replay is associated with efficient nonlocal learning

    Yunzhe Liu, Marcelo G Mattar, Timothy EJ Behrens, Nathaniel D Daw, and Raymond J Dolan. Experience replay is associated with efficient nonlocal learning. Science, 372 0 (6544): 0 eabf1357, 2021

  51. [59]

    Acquiring new memories in neocortex of hippocampal-lesioned mice

    Wenhan Luo, Di Yun, Yi Hu, Miaomiao Tian, Jiajun Yang, Yifan Xu, Yong Tang, Yang Zhan, Hong Xie, and Ji-Song Guan. Acquiring new memories in neocortex of hippocampal-lesioned mice. Nature communications, 13 0 (1): 0 1601, 2022

  52. [60]

    Memory-augmented graph neural networks: A brain-inspired review

    Guixiang Ma, Vy A Vo, Theodore L Willke, and Nesreen K Ahmed. Memory-augmented graph neural networks: A brain-inspired review. IEEE Transactions on Artificial Intelligence, 5 0 (5): 0 2011--2025, 2023

  53. [61]

    The stability-plasticity dilemma: Investigating the continuum from catastrophic forgetting to age-limited learning effects, 2013

    Martial Mermillod, Aur \'e lia Bugaiska, and Patrick Bonin. The stability-plasticity dilemma: Investigating the continuum from catastrophic forgetting to age-limited learning effects, 2013

  54. [62]

    The magical number seven, plus or minus two: Some limits on our capacity for processing information

    George A Miller. The magical number seven, plus or minus two: Some limits on our capacity for processing information. Psychological review, 63 0 (2): 0 81, 1956

  55. [63]

    o ksal, Ayyoob Imani, Mohsen Fayyaz, and Hinrich Sch \

    Ali Modarressi, Abdullatif K \"o ksal, Ayyoob Imani, Mohsen Fayyaz, and Hinrich Sch \"u tze. Memllm: Finetuning llms to use an explicit read-write memory. arXiv preprint arXiv:2404.11672, 2024

  56. [64]

    Episodic memory and beyond: the hippocampus and neocortex in transformation

    Morris Moscovitch, Roberto Cabeza, Gordon Winocur, and Lynn Nadel. Episodic memory and beyond: the hippocampus and neocortex in transformation. Annual review of psychology, 67 0 (1): 0 105--134, 2016

  57. [65]

    Memory: Neurobiological mechanisms and assessment

    Swaleha Mujawar, Jaideep Patil, Bhushan Chaudhari, and Daniel Saldanha. Memory: Neurobiological mechanisms and assessment. Industrial psychiatry journal, 30 0 (Suppl 1): 0 S311--S314, 2021

  58. [66]

    Most brain activity is ‘background noise’—and that’s upending our understanding of consciousness, 2021

    Thomas Nail. Most brain activity is ‘background noise’—and that’s upending our understanding of consciousness, 2021

  59. [67]

    Dynamic memory compression: Retrofitting llms for accelerated inference

    Piotr Nawrot, Adrian a \'n cucki, Marcin Chochowski, David Tarjan, and Edoardo M Ponti. Dynamic memory compression: Retrofitting llms for accelerated inference. arXiv preprint arXiv:2403.09636, 2024

  60. [68]

    Functional role of gamma and theta oscillations in episodic memory

    Erika Nyhus and Tim Curran. Functional role of gamma and theta oscillations in episodic memory. Neuroscience & Biobehavioral Reviews, 34 0 (7): 0 1023--1035, 2010

  61. [69]

    Memgpt: Towards llms as operating systems

    Charles Packer, Vivian Fang, Shishir\_G Patil, Kevin Lin, Sarah Wooders, and Joseph\_E Gonzalez. Memgpt: Towards llms as operating systems. 2023

  62. [70]

    Abc: Attention with bounded-memory control

    Hao Peng, Jungo Kasai, Nikolaos Pappas, Dani Yogatama, Zhaofeng Wu, Lingpeng Kong, Roy Schwartz, and Noah A Smith. Abc: Attention with bounded-memory control. arXiv preprint arXiv:2110.02488, 2021

  63. [71]

    Position: Episodic memory is the missing piece for long-term llm agents

    Mathis Pink, Qinyuan Wu, Vy Ai Vo, Javier Turek, Jianing Mu, Alexander Huth, and Mariya Toneva. Position: Episodic memory is the missing piece for long-term llm agents. arXiv preprint arXiv:2502.06975, 2025

  64. [72]

    Train short, test long: Attention with linear biases enables input length extrapolation

    Ofir Press, Noah A Smith, and Mike Lewis. Train short, test long: Attention with linear biases enables input length extrapolation. arXiv preprint arXiv:2108.12409, 2021

  65. [73]

    Neuromodulation of the feedforward dentate gyrus-ca3 microcircuit

    Luke Y Prince, Travis J Bacon, Cezar M Tigaret, and Jack R Mellor. Neuromodulation of the feedforward dentate gyrus-ca3 microcircuit. Frontiers in synaptic neuroscience, 8: 0 32, 2016

  66. [74]

    Layerwise recurrent router for mixture-of-experts

    Zihan Qiu, Zeyu Huang, Shuang Cheng, Yizhi Zhou, Zili Wang, Ivan Titov, and Jie Fu. Layerwise recurrent router for mixture-of-experts. arXiv preprint arXiv:2408.06793, 2024

  67. [75]

    A multisensory perspective of working memory

    Michel Quak, Raquel Elea London, and Durk Talsma. A multisensory perspective of working memory. Frontiers in human neuroscience, 9: 0 197, 2015

  68. [76]

    Compressive transformers for long-range sequence modelling

    Jack W Rae, Anna Potapenko, Siddhant M Jayakumar, and Timothy P Lillicrap. Compressive transformers for long-range sequence modelling. arXiv preprint arXiv:1911.05507, 2019

  69. [77]

    A default mode of brain function

    Marcus E Raichle, Ann Mary MacLeod, Abraham Z Snyder, William J Powers, Debra A Gusnard, and Gordon L Shulman. A default mode of brain function. Proceedings of the national academy of sciences, 98 0 (2): 0 676--682, 2001

  70. [78]

    Adaptive knowledge consolidation: A dynamic approach to mitigating catastrophic forgetting in text-based neural networks

    J Ranjith and Santhi Baskaran. Adaptive knowledge consolidation: A dynamic approach to mitigating catastrophic forgetting in text-based neural networks. 2024

  71. [79]

    Hierarchical dynamics as a macroscopic organizing principle of the human brain

    Ryan V Raut, Abraham Z Snyder, and Marcus E Raichle. Hierarchical dynamics as a macroscopic organizing principle of the human brain. Proceedings of the National Academy of Sciences, 117 0 (34): 0 20890--20897, 2020

  72. [80]

    Dissociating distinct cortical networks associated with subregions of the human medial temporal lobe using precision neuroimaging

    Daniel Reznik, Robert Trampel, Nikolaus Weiskopf, Menno P Witter, and Christian F Doeller. Dissociating distinct cortical networks associated with subregions of the human medial temporal lobe using precision neuroimaging. Neuron, 111 0 (17): 0 2756--2772, 2023

  73. [81]

    Associative recurrent memory transformer

    Ivan Rodkin, Yuri Kuratov, Aydar Bulatov, and Mikhail Burtsev. Associative recurrent memory transformer. arXiv preprint arXiv:2407.04841, 2024

  74. [82]

    The mechanisms for pattern completion and pattern separation in the hippocampus

    Edmund T Rolls. The mechanisms for pattern completion and pattern separation in the hippocampus. Frontiers in systems neuroscience, 7: 0 74, 2013

  75. [83]

    Working memory and neural oscillations: alpha--gamma versus theta--gamma codes for distinct wm information? Trends in cognitive sciences, 18 0 (1): 0 16--25, 2014

    Fr \'e d \'e ric Roux and Peter J Uhlhaas. Working memory and neural oscillations: alpha--gamma versus theta--gamma codes for distinct wm information? Trends in cognitive sciences, 18 0 (1): 0 16--25, 2014

  76. [84]

    Deep learning needs a prefrontal cortex

    Jacob Russin, Randall C O’Reilly, and Yoshua Bengio. Deep learning needs a prefrontal cortex. Work Bridging AI Cogn Sci, 107 0 (603-616): 0 1, 2020

  77. [85]

    Open-ended instructable embodied agents with memory-augmented large language models

    Gabriel Sarch, Yue Wu, Michael J Tarr, and Katerina Fragkiadaki. Open-ended instructable embodied agents with memory-augmented large language models. arXiv preprint arXiv:2310.15127, 2023

  78. [86]

    Retrieval-augmented decision transformer: External memory for in-context rl

    Thomas Schmied, Fabian Paischer, Vihang Patil, Markus Hofmarcher, Razvan Pascanu, and Sepp Hochreiter. Retrieval-augmented decision transformer: External memory for in-context rl. arXiv preprint arXiv:2410.07071, 2024

  79. [87]

    Stress and multiple memory systems: from ‘thinking’to ‘doing’

    Lars Schwabe and Oliver T Wolf. Stress and multiple memory systems: from ‘thinking’to ‘doing’. Trends in cognitive sciences, 17 0 (2): 0 60--68, 2013

  80. [88]

    Cognitive memory in large language models

    Lianlei Shan, Shixian Luo, Zezhou Zhu, Yu Yuan, and Yong Wu. Cognitive memory in large language models. arXiv preprint arXiv:2504.02441, 2025

  81. [89]

    The neural representations underlying asymmetric cross-modal prediction of words

    Liang Shi, Chuqi Liu, Xiaojing Peng, Yifei Cao, Daniel A Levy, and Gui Xue. The neural representations underlying asymmetric cross-modal prediction of words. Human Brain Mapping, 44 0 (6): 0 2418--2435, 2023

  82. [90]

    Prediction errors disrupt hippocampal representations and update episodic memories

    Alyssa H Sinclair, Grace M Manalili, Iva K Brunec, R Alison Adcock, and Morgan D Barense. Prediction errors disrupt hippocampal representations and update episodic memories. Proceedings of the National Academy of Sciences, 118 0 (51): 0 e2117625118, 2021

  83. [91]

    Memory consolidation

    Larry R Squire, Lisa Genzel, John T Wixted, and Richard G Morris. Memory consolidation. Cold Spring Harbor perspectives in biology, 7 0 (8): 0 a021766, 2015

  84. [92]

    Hippocampal pattern completion is linked to gamma power increases and alpha power decreases during recollection

    Bernhard P Staresina, Sebastian Michelmann, Mathilde Bonnefond, Ole Jensen, Nikolai Axmacher, and Juergen Fell. Hippocampal pattern completion is linked to gamma power increases and alpha power decreases during recollection. elife, 5: 0 e17397, 2016

  85. [93]

    Transformer-squared: Self-adaptive llms

    Qi Sun, Edoardo Cetin, and Yujin Tang. Transformer-squared: Self-adaptive llms. In The Thirteenth International Conference on Learning Representations, 2025

  86. [94]

    Associative transformer

    Yuwei Sun, Hideya Ochiai, Zhirong Wu, Stephen Lin, and Ryota Kanai. Associative transformer. arXiv preprint arXiv:2309.12862, 2023

  87. [95]

    Hippocampal--prefrontal interactions during spatial d ecision-making

    Lucas CS Tavares and Adriano BL Tort. Hippocampal--prefrontal interactions during spatial d ecision-making. Hippocampus, 32 0 (1): 0 38--54, 2022

  88. [96]

    Transformer memory as a differentiable search index

    Yi Tay, Vinh Tran, Mostafa Dehghani, Jianmo Ni, Dara Bahri, Harsh Mehta, Zhen Qin, Kai Hui, Zhe Zhao, Jai Gupta, et al. Transformer memory as a differentiable search index. Advances in Neural Information Processing Systems, 35: 0 21831--21843, 2022

  89. [97]

    The hippocampal indexing theory and episodic memory: updating the index

    Timothy J Teyler and Jerry W Rudy. The hippocampal indexing theory and episodic memory: updating the index. Hippocampus, 17 0 (12): 0 1158--1169, 2007

  90. [98]

    Attention is all you need

    Ashish Vaswani. Attention is all you need. Advances in neural information processing systems, 30: 0 I, 2017

  91. [99]

    Hierarchical reasoning model

    Guan Wang, Jin Li, Yuhao Sun, Xing Chen, Changling Liu, Yue Wu, Meng Lu, Sen Song, and Yasin Abbasi Yadkori. Hierarchical reasoning model. arXiv preprint arXiv:2506.21734, 2025 a

  92. [100]

    Schrodinger's memory: Large language models

    Wei Wang and Qing Li. Schrodinger's memory: Large language models. arXiv preprint arXiv:2409.10482, 2024

  93. [101]

    Augmenting language models with long-term memory

    Weizhi Wang, Li Dong, Hao Cheng, Xiaodong Liu, Xifeng Yan, Jianfeng Gao, and Furu Wei. Augmenting language models with long-term memory. Advances in Neural Information Processing Systems, 36: 0 74530--74543, 2023

  94. [102]

    R3mem: Bridging memory retention and retrieval via reversible compression

    Xiaoqiang Wang, Suyuchen Wang, Yun Zhu, and Bang Liu. R3mem: Bridging memory retention and retrieval via reversible compression. arXiv preprint arXiv:2502.15957, 2025 b

  95. [103]

    Beyond the limits: A survey of techniques to extend the context length in large language models

    Xindi Wang, Mahsa Salmani, Parsa Omidi, Xiangyu Ren, Mehdi Rezagholizadeh, and Armaghan Eshaghi. Beyond the limits: A survey of techniques to extend the context length in large language models. arXiv preprint arXiv:2402.02244, 2024 a

  96. [104]

    Memoryllm: Towards self-updatable large language models

    Yu Wang, Yifan Gao, Xiusi Chen, Haoming Jiang, Shiyang Li, Jingfeng Yang, Qingyu Yin, Zheng Li, Xian Li, Bing Yin, et al. Memoryllm: Towards self-updatable large language models. arXiv preprint arXiv:2402.04624, 2024 b

  97. [105]

    M+: Extending memoryllm with scalable long-term memory

    Yu Wang, Dmitry Krotov, Yuanzhe Hu, Yifan Gao, Wangchunshu Zhou, Julian McAuley, Dan Gutfreund, Rogerio Feris, and Zexue He. M+: Extending memoryllm with scalable long-term memory. arXiv preprint arXiv:2502.00592, 2025 c

  98. [106]

    Jarvis-1: Open-world multi-task agents with memory-augmented multimodal language models

    Zihao Wang, Shaofei Cai, Anji Liu, Yonggang Jin, Jinbing Hou, Bowei Zhang, Haowei Lin, Zhaofeng He, Zilong Zheng, Yaodong Yang, et al. Jarvis-1: Open-world multi-task agents with memory-augmented multimodal language models. IEEE Transactions on Pattern Analysis and Machine Int...

  99. [107]

    Stateful memory-augmented transformers for efficient dialogue modeling

    Qingyang Wu and Zhou Yu. Stateful memory-augmented transformers for efficient dialogue modeling. arXiv preprint arXiv:2209.07634, 2022

  100. [108]

    Memformer: A memory-augmented transformer for sequence modeling

    Qingyang Wu, Zhenzhong Lan, Kun Qian, Jing Gu, Alborz Geramifard, and Zhou Yu. Memformer: A memory-augmented transformer for sequence modeling. arXiv preprint arXiv:2010.06891, 2020

  101. [109]

    The kanerva machine: A generative distributed memory

    Yan Wu, Greg Wayne, Alex Graves, and Timothy Lillicrap. The kanerva machine: A generative distributed memory. arXiv preprint arXiv:1804.01756, 2018

  102. [110]

    From human memory to ai memory: A survey on memory mechanisms in the era of llms

    Yaxiong Wu, Sheng Liang, Chen Zhang, Yichao Wang, Yongyue Zhang, Huifeng Guo, Ruiming Tang, and Yong Liu. From human memory to ai memory: A survey on memory mechanisms in the era of llms. arXiv preprint arXiv:2504.15965, 2025

  103. [111]

    Memorizing transformers

    Yuhuai Wu, Markus N Rabe, DeLesley Hutchins, and Christian Szegedy. Memorizing transformers. arXiv preprint arXiv:2203.08913, 2022 a

  104. [112]

    An efficient memory-augmented transformer for knowledge-intensive nlp tasks

    Yuxiang Wu, Yu Zhao, Baotian Hu, Pasquale Minervini, Pontus Stenetorp, and Sebastian Riedel. An efficient memory-augmented transformer for knowledge-intensive nlp tasks. arXiv preprint arXiv:2210.16773, 2022 b

  105. [113]

    A-mem: Agentic memory for llm agents

    Wujiang Xu, Zujie Liang, Kai Mei, Hang Gao, Juntao Tan, and Yongfeng Zhang. A-mem: Agentic memory for llm agents. arXiv preprint arXiv:2502.12110, 2025

  106. [114]

    Adaptive computation with elastic input sequence

    Fuzhao Xue, Valerii Likhosherstov, Anurag Arnab, Neil Houlsby, Mostafa Dehghani, and Yang You. Adaptive computation with elastic input sequence. In International Conference on Machine Learning, pp.\ 38971--38988. PMLR, 2023

  107. [115]

    Memory3: Language modeling with explicit memory

    Hongkang Yang, Zehao Lin, Wenjin Wang, Hao Wu, Zhiyu Li, Bo Tang, Wenqiang Wei, Jinbo Wang, Zeyun Tang, Shichao Song, et al. Memory3: Language modeling with explicit memory. arXiv preprint arXiv:2407.01178, 2024

  108. [116]

    Longer context, deeper thinking: Uncovering the role of long-context ability in reasoning

    Wang Yang, Zirui Liu, Hongye Jin, Qingyu Yin, Vipin Chaudhary, and Xiaotian Han. Longer context, deeper thinking: Uncovering the role of long-context ability in reasoning. arXiv preprint arXiv:2505.17315, 2025

  109. [117]

    Score: Story coherence and retrieval enhancement for ai narratives

    Qiang Yi, Yangfan He, Jianhui Wang, Xinyuan Song, Shiyao Qian, Xinhang Yuan, Miao Zhang, Li Sun, Keqin Li, Kuan Lu, et al. Score: Story coherence and retrieval enhancement for ai narratives. arXiv preprint arXiv:2503.23512, 2025

  110. [118]

    Malt diffusion: Memory-augmented latent transformers for any-length video generation

    Sihyun Yu, Meera Hahn, Dan Kondratyuk, Jinwoo Shin, Agrim Gupta, Jos \'e Lezama, Irfan Essa, David Ross, and Jonathan Huang. Malt diffusion: Memory-augmented latent transformers for any-length video generation. arXiv preprint arXiv:2502.12632, 2025

  111. [119]

    Peripheral memory for llms: Integration of sequential memory banks with adaptive querying

    Songlin Zhai, Yuan Meng, Yongrui Chen, Yiwei Wang, and Guilin Qi. Peripheral memory for llms: Integration of sequential memory banks with adaptive querying. In Forty-second International Conference on Machine Learning, 2025

  112. [120]

    A survey on the memory mechanism of large language model based agents

    Zeyu Zhang, Xiaohe Bo, Chen Ma, Rui Li, Xu Chen, Quanyu Dai, Jieming Zhu, Zhenhua Dong, and Ji-Rong Wen. A survey on the memory mechanism of large language model based agents. arXiv preprint arXiv:2404.13501, 2024

  113. [121]

    Memorybank: Enhancing large language models with long-term memory

    Wanjun Zhong, Lianghong Guo, Qiqi Gao, He Ye, and Yanlin Wang. Memorybank: Enhancing large language models with long-term memory. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 38, pp.\ 19724--19731, 2024 a

  114. [122]

    Random tree model of meaningful memory

    Weishun Zhong, Tankut Can, Antonis Georgiou, Ilya Shnayderman, Mikhail Katkov, and Misha Tsodyks. Random tree model of meaningful memory. bioRxiv, pp.\ 2024--12, 2024 b

  115. [123]

    Mlkv: Multi-layer key-value heads for memory efficient transformer decoding

    Zayd Muhammad Kawakibi Zuhri, Muhammad Farid Adilazuarda, Ayu Purwarianti, and Alham Fikri Aji. Mlkv: Multi-layer key-value heads for memory efficient transformer decoding. arXiv preprint arXiv:2406.09297, 2024

  116. [124]

    write newline

    " write newline "" before.all 'output.state := FUNCTION n.dashify 't := "" t empty not t #1 #1 substring "-" = t #1 #2 substring "--" = not "--" * t #2 global.max substring 't := t #1 #1 substring "-" = "-" * t #2 global.max substring 't := while if t #1 #1 substring * t #2 gl...

Pith tools

Reviewed August 5, 2026 · model on record in the stance chip above.