Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-12T01:14:01.584159Z
Paper Citation Record · LEDGER
As of 4 August 2026, this Paper Citation Record lists 100 of 103 outbound references and 3 inbound Pith citation observations for arXiv:2605.08301.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-12T01:14:01.584159Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-04T06:34:03.388597+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-03T02:03:22.318771Z
A source-named dated measurement, never combined with another source.
Source: cited_works
100 of 103 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 5ff7aafc-a2be-45bc-bef2-5cb6822ee13b · outbound
Priming: Hybrid State Space Models From Pre-trained Transformers GQA : Training generalized multi-query transformer models from multi-head checkpoints
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 51e8d96c-ee7e-4d16-a986-214fc6daf6c7 · outbound
Priming: Hybrid State Space Models From Pre-trained Transformers Training-free long-context scaling of large language models
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation b5a230cb-4f44-423f-bf26-dde20eb946b5 · outbound
Priming: Hybrid State Space Models From Pre-trained Transformers Zoology: Measuring and improving recall in efficient language models
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation d32423a0-737b-4ff9-9aa2-8ef27a780df7 · outbound
Priming: Hybrid State Space Models From Pre-trained Transformers Hybrid Architectures for Language Models: Systematic Analysis and Design Insights
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 937d8321-0e24-474d-b17f-f5c9de20b1f2 · outbound
Priming: Hybrid State Space Models From Pre-trained Transformers Chan, James Demmel, June Donato, Jack Dongarra, Victor Eijkhout, Roldan Pozo, Charles Romine, and Henk van der Vorst
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 6f681267-1086-4cd4-bc16-65167b3d0426 · outbound
Priming: Hybrid State Space Models From Pre-trained Transformers NVIDIA Nemotron Nano 2: An Accurate and Efficient Hybrid Mamba-Transformer Reasoning Model
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation bfc6085b-37c1-430f-a58a-b651690c2097 · outbound
Priming: Hybrid State Space Models From Pre-trained Transformers o ppel, Markus Spanring, Andreas Auer, Oleksandra Prudnikova, Michael Kopp, G \
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 0cb57d44-57fe-44b8-bfb0-692de7f97051 · outbound
Priming: Hybrid State Space Models From Pre-trained Transformers Transformers to ssms: Distilling quadratic knowledge to subquadratic models
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 81eea1c2-39bc-401c-bff5-fbec4135112c · outbound
Priming: Hybrid State Space Models From Pre-trained Transformers NVIDIA Nemotron 3: Efficient and Open Intelligence
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation b9810ab0-561f-44cc-be02-b1746864dd73 · outbound
Priming: Hybrid State Space Models From Pre-trained Transformers Nemotron 3 nano: Open, efficient mixture-of- experts hybrid mamba-transformer model for agentic reasoning
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 5d2e8985-7c82-48f3-9dfd-757aa98d9d5e · outbound
Priming: Hybrid State Space Models From Pre-trained Transformers Qwen3-Coder-Next Technical Report
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 73390f12-bb59-4489-a813-743d78cc460c · outbound
Priming: Hybrid State Space Models From Pre-trained Transformers MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation d60bf531-4344-49c9-921a-7ce4e222f91f · outbound
Priming: Hybrid State Space Models From Pre-trained Transformers Evaluating Large Language Models Trained on Code
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 017c088b-7c25-41f6-81e3-d757de190138 · outbound
Priming: Hybrid State Space Models From Pre-trained Transformers Baker, Benjamin Burns, Daniel Adu-Ampratwum, Xuhui Huang, Xia Ning, Song Gao, Yu Su, and Huan Sun
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 3020041d-d596-407a-8758-7b23be2f8263 · outbound
Priming: Hybrid State Space Models From Pre-trained Transformers Training Verifiers to Solve Math Word Problems
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 7e8955e8-2cc8-4860-8c1b-c175cfb36fbe · outbound
Priming: Hybrid State Space Models From Pre-trained Transformers Albert, Pranesh Srinivasan, Haining Pan, Philippe Faist, Brian A Rohr, Michael J
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation a78944fe-cdae-49a4-9fc0-fb2af640fb0c · outbound
Priming: Hybrid State Space Models From Pre-trained Transformers Transformer-xl: Attentive language models beyond a fixed-length context
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation bd33fdd7-d439-45d6-a3c7-359a5313b4ca · outbound
Priming: Hybrid State Space Models From Pre-trained Transformers Transformers are ssms: Generalized models and efficient algorithms through structured state space duality
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 01c2b27f-cbdc-4fdc-bcb6-ca9670a9df3c · outbound
Priming: Hybrid State Space Models From Pre-trained Transformers Flashattention: Fast and memory-efficient exact attention with io-awareness
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 96e8e322-c921-43a0-b314-a6df5715f1a1 · outbound
Priming: Hybrid State Space Models From Pre-trained Transformers Griffin: Mixing Gated Linear Recurrences with Local Attention for Efficient Language Models
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation b3df9085-4f0a-4da3-9d68-d551fe8ab05c · outbound
Priming: Hybrid State Space Models From Pre-trained Transformers Fewer truncations improve language modeling
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 363e801b-12a3-4f15-b0e3-bfc658e1c6b7 · outbound
Priming: Hybrid State Space Models From Pre-trained Transformers Hymba: A hybrid-head architecture for small language models
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 279ab5d2-cf54-4277-b35d-d2a7da0b70c5 · outbound
Priming: Hybrid State Space Models From Pre-trained Transformers A mathematical framework for transformer circuits
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 3cf975c1-395e-4bc3-9f16-7eedb95fba61 · outbound
Priming: Hybrid State Space Models From Pre-trained Transformers AREAL : A large-scale asynchronous reinforcement learning system for language reasoning
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 426662d7-f5b7-4f1d-bd8b-15a064ad7429 · outbound
Priming: Hybrid State Space Models From Pre-trained Transformers arXiv preprint arXiv:2512.12167 (2025) 44
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 1cabd8dc-b39d-40ab-81f7-05a3ffbb71d6 · outbound
Priming: Hybrid State Space Models From Pre-trained Transformers Zamba: A Compact 7B SSM Hybrid Model
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation ce0bfaac-6b8b-4da5-921e-5945e98fe755 · outbound
Priming: Hybrid State Space Models From Pre-trained Transformers RADLADS : Rapid attention distillation to linear attention decoders at scale
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation f20a10db-f0f4-42aa-ac4f-4b9203b03277 · outbound
Priming: Hybrid State Space Models From Pre-trained Transformers The Llama 3 Herd of Models
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 6ff6264f-8d7a-43a0-afb5-fd5dddbf66dd · outbound
Priming: Hybrid State Space Models From Pre-trained Transformers Mamba: Linear-Time Sequence Modeling with Selective State Spaces
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 281c1068-8fdb-4e87-819f-1a4ac6d0763f · outbound
Priming: Hybrid State Space Models From Pre-trained Transformers Combining recurrent, convolutional, and continuous-time models with linear state space layers
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation e531da0d-4253-4a6d-bc90-599e5a9b4b53 · outbound
Priming: Hybrid State Space Models From Pre-trained Transformers Efficiently modeling long sequences with structured state spaces
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 16c30fae-2d50-47a4-baec-22e9372604a0 · outbound
Priming: Hybrid State Space Models From Pre-trained Transformers Jet-nemotron: Efficient language model with post neural architecture search
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 3bf65d22-8989-4c53-97ea-5578daaa2961 · outbound
Priming: Hybrid State Space Models From Pre-trained Transformers A survey of model reduction by balanced truncation and some new results
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 8c850831-affb-47d6-bb81-6e593771b881 · outbound
Priming: Hybrid State Space Models From Pre-trained Transformers Measuring massive multitask language understanding
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation f8cf7fb8-eb70-48aa-8284-7c31f7af2c0f · outbound
Priming: Hybrid State Space Models From Pre-trained Transformers RULER : What s the real context size of your long-context language models? In First Conference on Language Modeling
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 1ae39118-3125-4b44-8c67-34f172dcf609 · outbound
Priming: Hybrid State Space Models From Pre-trained Transformers DeepSpeed Ulysses: System Optimizations for Enabling Training of Extreme Long Sequence Transformer Models
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation c30619b5-d517-4600-96a3-288bb4aaf694 · outbound
Priming: Hybrid State Space Models From Pre-trained Transformers Livecodebench: Holistic and contamination free evaluation of large language models for code
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation e2ef0674-9eb1-41a4-aeda-305953802a16 · outbound
Priming: Hybrid State Space Models From Pre-trained Transformers Kakade, and eran malach
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 47c771e0-f2e4-4a40-bd43-0864906837c8 · outbound
Priming: Hybrid State Space Models From Pre-trained Transformers SWE -bench: Can language models resolve real-world github issues? In The Twelfth International Conference on Learning Representations
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 490f224e-a9ff-4fa4-88b5-de3383575342 · outbound
Priming: Hybrid State Space Models From Pre-trained Transformers Unresolved cited work
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 2e58b95e-be58-4d01-9c25-cfc01b0401a0 · outbound
Priming: Hybrid State Space Models From Pre-trained Transformers Transformers are rnns: Fast autoregressive transformers with linear attention
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 3a2a1bfb-6e45-4952-a779-28f2d39b268f · outbound
Priming: Hybrid State Space Models From Pre-trained Transformers Reformer: The efficient transformer
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 47bb123b-44ac-4f48-944c-bf5a5a20f7b3 · outbound
Priming: Hybrid State Space Models From Pre-trained Transformers BABIL ong: Testing the limits of LLM s with long context reasoning-in-a-haystack
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation aea44e9c-66a0-4197-a0ae-7236307006a8 · outbound
Priming: Hybrid State Space Models From Pre-trained Transformers Gonzalez, Hao Zhang, and Ion Stoica
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation f8e55d16-425b-49ef-8448-f5994b95eb4e · outbound
Priming: Hybrid State Space Models From Pre-trained Transformers Hwang, Jiangjiang Yang, Ronan Le Bras, Oyvind Tafjord, Christopher Wilhelm, Luca Soldaini, Noah A
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 988de457-b21d-4390-a0c1-2fa629a902f9 · outbound
Priming: Hybrid State Space Models From Pre-trained Transformers Liger: Linearizing large language models to gated recurrent structures
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 28d1534a-5c58-4db3-ac42-84b8dc67bb81 · outbound
Priming: Hybrid State Space Models From Pre-trained Transformers Distilling to hybrid attention models via kl-guided layer selection
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation a00fa614-b8a5-4108-a1d7-86958c0cce91 · outbound
Priming: Hybrid State Space Models From Pre-trained Transformers Jamba: A Hybrid Transformer-Mamba Language Model
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 49ceeb95-dd09-4998-8c1a-9b81a21cbe86 · outbound
Priming: Hybrid State Space Models From Pre-trained Transformers Truthfulqa: Measuring how models mimic human falsehoods
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 5b1386f6-2a5e-4c76-bd15-7c542f625b06 · outbound
Priming: Hybrid State Space Models From Pre-trained Transformers On the stochastic realization problem
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation d8b7613b-0b3d-4d70-beda-a375c2958f6a · outbound
Priming: Hybrid State Space Models From Pre-trained Transformers Ringattention with blockwise transformers for near-infinite context
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 5188ea24-2971-405a-a5f6-d4d51d710546 · outbound
Priming: Hybrid State Space Models From Pre-trained Transformers Is your code generated by chatgpt really correct? rigorous evaluation of large language models for code generation
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 161b96d5-0ac4-425c-8728-ed406cbc5ae8 · outbound
Priming: Hybrid State Space Models From Pre-trained Transformers PICASO : Permutation-invariant context composition with state space models
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 565abb02-4b3a-4a55-8172-2a0bf9caab66 · outbound
Priming: Hybrid State Space Models From Pre-trained Transformers Unresolved cited work
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 06284579-59f1-4c7b-9b11-24059eedf42f · outbound
Priming: Hybrid State Space Models From Pre-trained Transformers Error propagation properties of recursive least-squares adaptation algorithms
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation a4f5633b-649f-4f16-a69d-e599b577a8cd · outbound
Priming: Hybrid State Space Models From Pre-trained Transformers When Not to Trust Language Models: Investigating Effectiveness of Parametric and Non-Parametric Memories
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 882de4f2-0ee7-421c-b007-f59a54d3144f · outbound
Priming: Hybrid State Space Models From Pre-trained Transformers AMC/AIME : MAA invitational competitions
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation df7df56e-506a-4215-8e78-6fb315c5543d · outbound
Priming: Hybrid State Space Models From Pre-trained Transformers Linearizing large language models
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation f562990e-7dfa-487a-bf7a-e7a8c67249af · outbound
Priming: Hybrid State Space Models From Pre-trained Transformers Landmark attention: Random-access infinite context length for transformers
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation e83cdf94-e5b8-4fbf-aaaf-6faf1ac9127b · outbound
Priming: Hybrid State Space Models From Pre-trained Transformers Leave No Context Behind: Efficient Infinite Context Transformers with Infini-attention
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 3c0df41a-7846-4aab-9c6f-176c801b3f07 · outbound
Priming: Hybrid State Space Models From Pre-trained Transformers Unresolved cited work
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 60347e00-26d8-4a50-ad55-ce99597ea7b9 · outbound
Priming: Hybrid State Space Models From Pre-trained Transformers Expansion span: Combining fading memory and retrieval in hybrid state space models
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation c39bae94-aef0-409d-ac04-5f7881ed9c77 · outbound
Priming: Hybrid State Space Models From Pre-trained Transformers In-context Learning and Induction Heads
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation f5f4758e-6e05-4ffe-8dd6-90b914375b6f · outbound
Priming: Hybrid State Space Models From Pre-trained Transformers Resurrecting recurrent neural networks for long sequences
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation aee01451-f1bc-4d82-8a51-f6d9ca9c3f38 · outbound
Priming: Hybrid State Space Models From Pre-trained Transformers Marconi: Prefix caching for the era of hybrid LLM s
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation dcf34717-e677-46d3-8a55-f08bc8ed5e90 · outbound
Priming: Hybrid State Space Models From Pre-trained Transformers Patil, Huanzhi Mao, Charlie Cheng-Jie Ji, Fanjia Yan, Vishnu Suresh, Ion Stoica, and Joseph E
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 272b97d0-e517-4bce-a3ea-da3b4cea7187 · outbound
Priming: Hybrid State Space Models From Pre-trained Transformers Time-Varying Systems and Computations
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 6436c4c4-57c5-4314-a932-a569b495373c · outbound
Priming: Hybrid State Space Models From Pre-trained Transformers Ya RN : Efficient context window extension of large language models
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation a8d63189-fa8e-4492-9239-41c59c345cb4 · outbound
Priming: Hybrid State Space Models From Pre-trained Transformers Gated KalmaNet: A Fading Memory Layer Through Test-Time Ridge Regression
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 9fa358c9-7797-4344-8374-0f00905acf27 · outbound
Priming: Hybrid State Space Models From Pre-trained Transformers Generalizing verifiable instruction following
Reference 70
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation a5ae0a2e-c954-49b4-9b55-9223e8962699 · outbound
Priming: Hybrid State Space Models From Pre-trained Transformers Ulysses sequence parallelism in the hugging face ecosystem
Reference 71
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 10b9c153-478a-4296-b8cb-82b16c474648 · outbound
Priming: Hybrid State Space Models From Pre-trained Transformers Unresolved cited work
Reference 72
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 6a65d74d-56b4-496f-bb0e-1a5633daa46f · outbound
Priming: Hybrid State Space Models From Pre-trained Transformers Samba: Simple hybrid state space models for efficient unlimited context language modeling
Reference 73
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 7ef43862-d7e3-4c5b-9a2e-a8a6ef7afcd8 · outbound
Priming: Hybrid State Space Models From Pre-trained Transformers Keisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, and Yejin Choi
Reference 74
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 13a2ee2d-c848-4f4c-859d-419c44e8051d · outbound
Priming: Hybrid State Space Models From Pre-trained Transformers Sandberg and A
Reference 75
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 9bd85cc4-e5dc-4a41-a84e-71edc2ef8922 · outbound
Priming: Hybrid State Space Models From Pre-trained Transformers Unresolved cited work
Reference 76
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation dee51b14-2b7a-4c12-a90b-8989670958c9 · outbound
Priming: Hybrid State Space Models From Pre-trained Transformers Unresolved cited work
Reference 77
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 52ba9a51-c1c8-4f15-b6a8-98bc4b2fce2d · outbound
Priming: Hybrid State Space Models From Pre-trained Transformers Flashattention-3: Fast and accurate attention with asynchrony and low-precision
Reference 78
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 3b4d2990-9ea2-48e9-9dd1-197840eb6047 · outbound
Priming: Hybrid State Space Models From Pre-trained Transformers Hybridflow: A flexible and efficient rlhf framework
Reference 79
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 326614f8-3d6d-40da-93d7-9b8f4ef33b72 · outbound
Priming: Hybrid State Space Models From Pre-trained Transformers Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism
Reference 80
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 8922d645-bb58-49af-94e9-c1af7a027842 · outbound
Priming: Hybrid State Space Models From Pre-trained Transformers Attention Sink in Transformers: A Survey on Utilization, Interpretation, and Mitigation
Reference 81
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation ec2e9b3b-4ffb-49a6-8046-950a486a530f · outbound
Priming: Hybrid State Space Models From Pre-trained Transformers Retentive Network: A Successor to Transformer for Large Language Models
Reference 82
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation b68dc4de-f625-4eae-8803-362998f0c8d6 · outbound
Priming: Hybrid State Space Models From Pre-trained Transformers Challenging big-bench tasks and whether chain-of-thought can solve them
Reference 83
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 82074d65-a63b-463c-be2d-c2229ed33159 · outbound
Priming: Hybrid State Space Models From Pre-trained Transformers Scicode: A research coding benchmark curated by scientists
Reference 84
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation cb3c5521-07db-497f-b58f-4e30e9442b2d · outbound
Priming: Hybrid State Space Models From Pre-trained Transformers Attention is all you need
Reference 85
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 66e2f336-3a1d-42fa-a17c-4159881331e2 · outbound
Priming: Hybrid State Space Models From Pre-trained Transformers Michelangelo: Long Context Evaluations Beyond Haystacks via Latent Structure Queries
Reference 86
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation bb8b7f7b-f4c7-4ec3-b938-630a75f25846 · outbound
Priming: Hybrid State Space Models From Pre-trained Transformers MesaNet: Sequence Modeling by Locally Optimal Test-Time Training
Reference 87
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 98579f67-ec15-4123-ad7b-0588d2c509bd · outbound
Priming: Hybrid State Space Models From Pre-trained Transformers An Empirical Study of Mamba-based Language Models
Reference 88
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 31c73377-1c84-489b-98cd-dae9d5eac527 · outbound
Priming: Hybrid State Space Models From Pre-trained Transformers The mamba in the llama: Distilling and accelerating hybrid models
Reference 89
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation d67adf58-18a2-4995-9480-5d8e621832d2 · outbound
Priming: Hybrid State Space Models From Pre-trained Transformers M1: Towards scalable test-time compute with mamba reasoning models
Reference 90
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 5356abb7-9ac4-4a42-9bb1-bd0936c3871f · outbound
Priming: Hybrid State Space Models From Pre-trained Transformers Unresolved cited work
Reference 91
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 71612217-148e-4283-bb0c-1cfb27be1269 · outbound
Priming: Hybrid State Space Models From Pre-trained Transformers Model soups: averaging weights of multiple fine-tuned models improves accuracy without increasing inference time
Reference 92
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 412068f2-b15f-4266-acb1-1420bb73e247 · outbound
Priming: Hybrid State Space Models From Pre-trained Transformers Effective long-context scaling of foundation models
Reference 93
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 8bc08899-e54e-4ec7-b1c7-61b869b9ce6e · outbound
Priming: Hybrid State Space Models From Pre-trained Transformers Gated linear attention transformers with hardware-efficient training
Reference 94
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 8a41a46f-e508-4d72-8d45-09fdfb443138 · outbound
Priming: Hybrid State Space Models From Pre-trained Transformers Parallelizing linear transformers with the delta rule over sequence length
Reference 95
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 647af476-cda2-4fbc-abbe-77fa6825e88c · outbound
Priming: Hybrid State Space Models From Pre-trained Transformers Gated delta networks: Improving mamba2 with delta rule
Reference 96
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 0a75c586-7ea7-4804-82b0-6ca2995911c0 · outbound
Priming: Hybrid State Space Models From Pre-trained Transformers Ape: Faster and longer context-augmented generation via adaptive parallel encoding
Reference 97
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation e6b20170-6593-4a3f-93b6-389b03d502ba · outbound
Priming: Hybrid State Space Models From Pre-trained Transformers Helmet: How to evaluate long-context language models effectively and thoroughly
Reference 98
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation b03622f6-866b-47b2-848d-461d9e794b13 · outbound
Priming: Hybrid State Space Models From Pre-trained Transformers Native sparse attention: Hardware-aligned and natively trainable sparse attention
Reference 99
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 4f4d1986-d94c-4de3-a448-3c909abf6af2 · outbound
Priming: Hybrid State Space Models From Pre-trained Transformers Stacked Residuals of Dynamic Layers for Time Series Anomaly Detection
Reference 100
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.
Observation 6e95951f-c6f1-41fe-b0cd-289351c35abe · inbound
The Capability Convergence Hypothesis: Capability from Access Structure, Not Scale Priming: Hybrid State Space Models From Pre-trained Transformers
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7241869c-1acf-401b-a44d-95865c955a0c · inbound
The Capability Convergence Hypothesis: Capability from Access Structure, Not Scale Priming: Hybrid State Space Models From Pre-trained Transformers
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8d8a1feb-65df-46c1-ba5e-724df7bd875e · inbound
Memory for Large Language Models Priming: Hybrid State Space Models From Pre-trained Transformers
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.