Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-07-11T03:21:01.687528Z
Paper Citation Record · LEDGER
As of 10 August 2026, this Paper Citation Record lists 56 of 56 outbound references and 0 inbound Pith citation observations for arXiv:2607.06601.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-07-11T03:21:01.687528Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
56 of 56 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation c6f1d44a-58a3-41cf-9474-5a12fd4a3c62 · outbound
TriRoute: Unified Learned Routing for Joint Adaptive Attention, Experts, and KV-Cache Allocation GQA: Training generalized multi-query transformer models from multi-head checkpoints
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation a16e609c-2b80-4a1f-87fb-deb0281fda50 · outbound
TriRoute: Unified Learned Routing for Joint Adaptive Attention, Experts, and KV-Cache Allocation CoLT5: Faster long-range transformers with conditional computation.Empirical Methods in Natural Language Processing (EMNLP), 2023
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 8cf0eac4-4837-4025-bc5b-debbfb13014b · outbound
TriRoute: Unified Learned Routing for Joint Adaptive Attention, Experts, and KV-Cache Allocation Longformer: The Long-Document Transformer
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation b7702268-9035-4e0d-a534-eb861ed63f8e · outbound
TriRoute: Unified Learned Routing for Joint Adaptive Attention, Experts, and KV-Cache Allocation Estimating or Propagating Gradients Through Stochastic Neurons for Conditional Computation
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation bc05f8c9-fe22-45da-a4e6-80438af1be29 · outbound
TriRoute: Unified Learned Routing for Joint Adaptive Attention, Experts, and KV-Cache Allocation Pythia: A suite for analyzing large language models across training and scaling
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 3b0a26b1-31fe-4e36-854a-cf17470d6d34 · outbound
TriRoute: Unified Learned Routing for Joint Adaptive Attention, Experts, and KV-Cache Allocation PIQA: Reasoning about physical commonsense in natural language
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation c12fb47b-8289-42ca-b7e2-90b16f86ea63 · outbound
TriRoute: Unified Learned Routing for Joint Adaptive Attention, Experts, and KV-Cache Allocation Generating Long Sequences with Sparse Transformers
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 1a129485-6300-444e-ba64-894f24bef20b · outbound
TriRoute: Unified Learned Routing for Joint Adaptive Attention, Experts, and KV-Cache Allocation Unified scaling laws for routed language models
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 4ee84a84-2d37-43f6-8e2f-590d3752ad65 · outbound
TriRoute: Unified Learned Routing for Joint Adaptive Attention, Experts, and KV-Cache Allocation Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation a2cc56b8-e4e2-46f3-b6e9-a57a31b824c9 · outbound
TriRoute: Unified Learned Routing for Joint Adaptive Attention, Experts, and KV-Cache Allocation Training Verifiers to Solve Math Word Problems
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 93ac1217-3f18-4a5a-8422-1ebd631c2aa1 · outbound
TriRoute: Unified Learned Routing for Joint Adaptive Attention, Experts, and KV-Cache Allocation SwitchHead: Accelerating Transformers with Mixture-of-Experts Attention
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation ec68947d-eccb-489f-8947-d8a121a03339 · outbound
TriRoute: Unified Learned Routing for Joint Adaptive Attention, Experts, and KV-Cache Allocation DeepSeekMoE: Towards Ultimate Expert Specialization in Mixture-of-Experts Language Models
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation f27b0295-16d3-4eca-a599-62b153aec92e · outbound
TriRoute: Unified Learned Routing for Joint Adaptive Attention, Experts, and KV-Cache Allocation Fu, Stefano Ermon, Atri Rudra, and Christopher Ré
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation c87a8819-21cc-48ac-9b77-812de7079b25 · outbound
TriRoute: Unified Learned Routing for Joint Adaptive Attention, Experts, and KV-Cache Allocation DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 30ed05dd-f00c-404d-939d-fd0b287cbb43 · outbound
TriRoute: Unified Learned Routing for Joint Adaptive Attention, Experts, and KV-Cache Allocation Uni- versal transformers
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 041e54fe-3af4-4cc1-8138-80c06638c52e · outbound
TriRoute: Unified Learned Routing for Joint Adaptive Attention, Experts, and KV-Cache Allocation LLM.int8(): 8-bit matrix multiplication for transformers at scale
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 154b7390-4c5d-4d04-b6be-58eeeb3c2292 · outbound
TriRoute: Unified Learned Routing for Joint Adaptive Attention, Experts, and KV-Cache Allocation QLoRA: Efficient finetuningofquantizedLLMs
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 598e965d-9d08-4985-ab32-4b9b2ac4af9b · outbound
TriRoute: Unified Learned Routing for Joint Adaptive Attention, Experts, and KV-Cache Allocation Depth-adaptive transformer
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 83ca5649-5439-469e-a9ec-eb14d5177954 · outbound
TriRoute: Unified Learned Routing for Joint Adaptive Attention, Experts, and KV-Cache Allocation Reducing transformer depth on demand with structured dropout
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 5ddb967a-5b0f-4a43-b651-12448e90d709 · outbound
TriRoute: Unified Learned Routing for Joint Adaptive Attention, Experts, and KV-Cache Allocation Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity.Journal of Machine Learning Research (JMLR), 23(120):1–39
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 27c5362e-215d-4238-bcf4-7944a43492d5 · outbound
TriRoute: Unified Learned Routing for Joint Adaptive Attention, Experts, and KV-Cache Allocation GPTQ: Accurate post- training quantization for generative pre-trained transformers
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 23a0a73d-f166-400d-ab5c-b204819a0901 · outbound
TriRoute: Unified Learned Routing for Joint Adaptive Attention, Experts, and KV-Cache Allocation MegaBlocks: Efficient sparse training with mixture-of-experts
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 9c4179aa-6bcf-4687-bac6-b8be76258c17 · outbound
TriRoute: Unified Learned Routing for Joint Adaptive Attention, Experts, and KV-Cache Allocation The Pile: An 800GB Dataset of Diverse Text for Language Modeling
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation c62f7489-23b5-419c-82aa-bd80143974bd · outbound
TriRoute: Unified Learned Routing for Joint Adaptive Attention, Experts, and KV-Cache Allocation Adaptive Computation Time for Recurrent Neural Networks
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 173ff941-c038-470c-b100-ab2a3fd97b21 · outbound
TriRoute: Unified Learned Routing for Joint Adaptive Attention, Experts, and KV-Cache Allocation Training Compute-Optimal Large Language Models
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 447465bc-217c-40b2-b1e1-c99de9d7778d · outbound
TriRoute: Unified Learned Routing for Joint Adaptive Attention, Experts, and KV-Cache Allocation Mahoney, Yakun Sophia Shao, Kurt Keutzer, and Amir Gholami
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation d8079ba3-b840-4b15-b822-0e11de00c7b1 · outbound
TriRoute: Unified Learned Routing for Joint Adaptive Attention, Experts, and KV-Cache Allocation Categorical reparameterization with Gumbel-Softmax
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 4d7f81cf-1db6-4dc2-bc20-210a94521bac · outbound
TriRoute: Unified Learned Routing for Joint Adaptive Attention, Experts, and KV-Cache Allocation Mixtral of Experts
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation f455abdb-917b-4b0c-8a6a-cf324a73ccb5 · outbound
TriRoute: Unified Learned Routing for Joint Adaptive Attention, Experts, and KV-Cache Allocation Scaling Laws for Neural Language Models
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 6b209d01-b2f7-43ec-80d0-7e88dae93dc2 · outbound
TriRoute: Unified Learned Routing for Joint Adaptive Attention, Experts, and KV-Cache Allocation Reformer: The efficient transformer
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 59880c2d-1f14-4510-a899-7a49269fdf5e · outbound
TriRoute: Unified Learned Routing for Joint Adaptive Attention, Experts, and KV-Cache Allocation GShard: Scaling giant models with conditional computation and automatic sharding
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation d2ac180e-cd6c-4a02-848a-411688e49dd3 · outbound
TriRoute: Unified Learned Routing for Joint Adaptive Attention, Experts, and KV-Cache Allocation BASE layers: Simplifying training of large, sparse models
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 1c18df4a-df0c-441f-a228-00bc96886f8c · outbound
TriRoute: Unified Learned Routing for Joint Adaptive Attention, Experts, and KV-Cache Allocation KIVI: A tuning-free asymmetric 2bit quantization for KV cache
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 81774670-d3d7-4251-aa82-f834564c41d1 · outbound
TriRoute: Unified Learned Routing for Joint Adaptive Attention, Experts, and KV-Cache Allocation Decoupled weight decay regularization
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 723db302-6504-4e02-97b7-aeef251d373a · outbound
TriRoute: Unified Learned Routing for Joint Adaptive Attention, Experts, and KV-Cache Allocation Maddison, Andriy Mnih, and Yee Whye Teh
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 572dd9e7-eb83-45d0-8494-342a99c2ec90 · outbound
TriRoute: Unified Learned Routing for Joint Adaptive Attention, Experts, and KV-Cache Allocation The LAMBADA dataset: Word prediction requiring a broad discourse context
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 0323126c-2c9c-4f04-844f-889b0cfff867 · outbound
TriRoute: Unified Learned Routing for Joint Adaptive Attention, Experts, and KV-Cache Allocation DeepSpeed-MoE: Advancing mixture-of- experts inference and training to power next-generation AI scale
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 47498db2-1298-4c4a-a40d-b0f60a7031c0 · outbound
TriRoute: Unified Learned Routing for Joint Adaptive Attention, Experts, and KV-Cache Allocation Mixture-of-Depths: Dynamically allocating compute in transformer-based language models
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation b7899cb2-a85c-438f-9a69-eea508631924 · outbound
TriRoute: Unified Learned Routing for Joint Adaptive Attention, Experts, and KV-Cache Allocation Hash layers for large sparse models
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation a3f626d7-f5e9-4783-a00a-1bdfc5dcd1c3 · outbound
TriRoute: Unified Learned Routing for Joint Adaptive Attention, Experts, and KV-Cache Allocation WinoGrande: An adversarial winograd schema challenge at scale.Communications of the ACM, 64(9):99–106
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation d366ebe4-852a-4557-976f-5bacdf5bb857 · outbound
TriRoute: Unified Learned Routing for Joint Adaptive Attention, Experts, and KV-Cache Allocation Confident adaptive language modeling
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 93ad1f44-0cdb-4b4e-828a-69cb4bd8bb58 · outbound
TriRoute: Unified Learned Routing for Joint Adaptive Attention, Experts, and KV-Cache Allocation Fast Transformer Decoding: One Write-Head is All You Need
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 98fa3c18-b935-48db-8bdb-833a3b0cd04e · outbound
TriRoute: Unified Learned Routing for Joint Adaptive Attention, Experts, and KV-Cache Allocation GLU Variants Improve Transformer
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 252b5bfc-ac9d-4412-a291-f49745633493 · outbound
TriRoute: Unified Learned Routing for Joint Adaptive Attention, Experts, and KV-Cache Allocation Outrageously large neural networks: The sparsely-gated mixture-of-experts layer
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation df4ee790-2647-4e2a-9d75-146d73252b30 · outbound
TriRoute: Unified Learned Routing for Joint Adaptive Attention, Experts, and KV-Cache Allocation RoFormer: Enhanced transformer with rotary position embedding.Neurocomputing, 568:127063
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation db412bd9-19d0-4b06-910b-a2f04c56fb64 · outbound
TriRoute: Unified Learned Routing for Joint Adaptive Attention, Experts, and KV-Cache Allocation Adaptive attention span in transformers
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation b05a2f47-3dca-4638-80a6-e14c5c59f5de · outbound
TriRoute: Unified Learned Routing for Joint Adaptive Attention, Experts, and KV-Cache Allocation RedPajama: An open source recipe to reproduce LLaMA training dataset
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 591d3d97-a3bc-40a2-a3f8-184120383b4d · outbound
TriRoute: Unified Learned Routing for Joint Adaptive Attention, Experts, and KV-Cache Allocation Williams
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation e4dc1fcc-2c38-4d1c-adc2-fb6971e3ef57 · outbound
TriRoute: Unified Learned Routing for Joint Adaptive Attention, Experts, and KV-Cache Allocation Efficient streaming language models with attention sinks
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 1c4f6a4c-4090-4364-94fe-a5ab1d4e26e5 · outbound
TriRoute: Unified Learned Routing for Joint Adaptive Attention, Experts, and KV-Cache Allocation Big bird: Transformers for longer sequences
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 94cf9ac2-63a2-4f92-93bf-5a894983ab02 · outbound
TriRoute: Unified Learned Routing for Joint Adaptive Attention, Experts, and KV-Cache Allocation HellaSwag: Can a machine really finish your sentence? InAssociation for Computational Linguistics (ACL), 2019
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 0e15645e-fb27-4c74-b7ed-9562502f5ce2 · outbound
TriRoute: Unified Learned Routing for Joint Adaptive Attention, Experts, and KV-Cache Allocation Root mean square layer normalization
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 1bbaa7e6-ed5d-4778-b9a0-4476bc91cad6 · outbound
TriRoute: Unified Learned Routing for Joint Adaptive Attention, Experts, and KV-Cache Allocation Mixture of attention heads: Selecting attention heads per token
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation e8afb211-eee2-4214-a981-7b401ce35bd6 · outbound
TriRoute: Unified Learned Routing for Joint Adaptive Attention, Experts, and KV-Cache Allocation H2O: Heavy-hitter oracle for efficient generative inference of large language models
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 190c1693-34d1-4084-aae5-d3ebda45542e · outbound
TriRoute: Unified Learned Routing for Joint Adaptive Attention, Experts, and KV-Cache Allocation Dai, Zhifeng Chen, Quoc V
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation a43205cb-f2bf-4834-9110-62da85b0a662 · outbound
TriRoute: Unified Learned Routing for Joint Adaptive Attention, Experts, and KV-Cache Allocation ST-MoE: Designing Stable and Transferable Sparse Expert Models
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
No inbound Pith citation observations are available.