Pith. sign in

Paper Citation Record · LEDGER

TriRoute: Unified Learned Routing for Joint Adaptive Attention, Experts, and KV-Cache Allocation

As of 10 August 2026, this Paper Citation Record lists 56 of 56 outbound references and 0 inbound Pith citation observations for arXiv:2607.06601.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.06601 v1

Coverage vector

measured 56 of 56 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-11T03:21:01.687528Z

measured 56 of 56 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

56 of 56 outbound references displayed

  • verified exact14
  • verified fuzzy39
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch3

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c6f1d44a-58a3-41cf-9474-5a12fd4a3c62 · outbound

This paper cites GQA: Training generalized multi-query transformer models from multi-head checkpoints.

TriRoute: Unified Learned Routing for Joint Adaptive Attention, Experts, and KV-Cache Allocation GQA: Training generalized multi-query transformer models from multi-head checkpoints

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-07-11T03:27:49.460673Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-11T03:21:01.687528Z digest=sha256:a5c7cd85ab4d7788c76a8d488844f72cf7a952d3b0a4db927f06503c292c5c5c

Observation a16e609c-2b80-4a1f-87fb-deb0281fda50 · outbound

This paper cites CoLT5: Faster long-range transformers with conditional computation.Empirical Methods in Natural Language Processing (EMNLP), 2023.

TriRoute: Unified Learned Routing for Joint Adaptive Attention, Experts, and KV-Cache Allocation CoLT5: Faster long-range transformers with conditional computation.Empirical Methods in Natural Language Processing (EMNLP), 2023

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-07-11T03:27:49.286954Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-11T03:21:01.687528Z digest=sha256:5e1c53a02b2ad2992e14a9203e9ad44e98b937a57f3f169850efba07a0252c04

Observation 8cf0eac4-4837-4025-bc5b-debbfb13014b · outbound

This paper cites Longformer: The Long-Document Transformer.

TriRoute: Unified Learned Routing for Joint Adaptive Attention, Experts, and KV-Cache Allocation Longformer: The Long-Document Transformer

Reference 3

Resolution
metadata mismatch
local_arxiv, observed 2026-07-11T03:27:46.701443Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-11T03:21:01.687528Z digest=sha256:15bfc6c6f26ed5276ad00992f71eafd5e69f2ca3d3cb7197825b764bc5f59e7b

Observation b7702268-9035-4e0d-a534-eb861ed63f8e · outbound

This paper cites Estimating or Propagating Gradients Through Stochastic Neurons for Conditional Computation.

TriRoute: Unified Learned Routing for Joint Adaptive Attention, Experts, and KV-Cache Allocation Estimating or Propagating Gradients Through Stochastic Neurons for Conditional Computation

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-07-11T03:27:46.831362Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-11T03:21:01.687528Z digest=sha256:43f0e2cb35031ad69fc8397f61ec35260f2567c9f4f47b9f883f58e40af37cdb

Observation bc05f8c9-fe22-45da-a4e6-80438af1be29 · outbound

This paper cites Pythia: A suite for analyzing large language models across training and scaling.

TriRoute: Unified Learned Routing for Joint Adaptive Attention, Experts, and KV-Cache Allocation Pythia: A suite for analyzing large language models across training and scaling

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-07-11T03:27:49.564973Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-11T03:21:01.687528Z digest=sha256:2488f5be898f8742de0279b48e033e907890b65fcf1f73cf7d4f92a92ed2bbab

Observation 3b0a26b1-31fe-4e36-854a-cf17470d6d34 · outbound

This paper cites PIQA: Reasoning about physical commonsense in natural language.

TriRoute: Unified Learned Routing for Joint Adaptive Attention, Experts, and KV-Cache Allocation PIQA: Reasoning about physical commonsense in natural language

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-07-11T03:27:49.257627Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-11T03:21:01.687528Z digest=sha256:9b01f17e7e95db100eeb575aac3b63a04f521e5742517d8e63a27ef732aba1fb

Observation c12fb47b-8289-42ca-b7e2-90b16f86ea63 · outbound

This paper cites Generating Long Sequences with Sparse Transformers.

TriRoute: Unified Learned Routing for Joint Adaptive Attention, Experts, and KV-Cache Allocation Generating Long Sequences with Sparse Transformers

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-07-11T03:27:46.912589Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-11T03:21:01.687528Z digest=sha256:dc81ac3cb9e65efe8a3a520264f15c2b6127ac4ad498951bbe61d7b3ab353690

Observation 1a129485-6300-444e-ba64-894f24bef20b · outbound

This paper cites Unified scaling laws for routed language models.

TriRoute: Unified Learned Routing for Joint Adaptive Attention, Experts, and KV-Cache Allocation Unified scaling laws for routed language models

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-07-11T03:27:49.317762Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-11T03:21:01.687528Z digest=sha256:6ac10bdd080cd1faa550016a7c6a15353ebdc360ac08937cf04efc54a40e9bff

Observation 4ee84a84-2d37-43f6-8e2f-590d3752ad65 · outbound

This paper cites Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge.

TriRoute: Unified Learned Routing for Joint Adaptive Attention, Experts, and KV-Cache Allocation Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-07-11T03:27:46.940199Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-11T03:21:01.687528Z digest=sha256:216e03d9c4b3428c7dd6e2495ce92b2aea8406e9620ead1da008a27e79d8e394

Observation a2cc56b8-e4e2-46f3-b6e9-a57a31b824c9 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

TriRoute: Unified Learned Routing for Joint Adaptive Attention, Experts, and KV-Cache Allocation Training Verifiers to Solve Math Word Problems

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-07-11T03:27:47.082510Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-11T03:21:01.687528Z digest=sha256:51d040cd7970189bb7e6b7e37861e7d743b17912c456acf1e5653a885f3a3857

Observation 93ac1217-3f18-4a5a-8422-1ebd631c2aa1 · outbound

This paper cites SwitchHead: Accelerating Transformers with Mixture-of-Experts Attention.

TriRoute: Unified Learned Routing for Joint Adaptive Attention, Experts, and KV-Cache Allocation SwitchHead: Accelerating Transformers with Mixture-of-Experts Attention

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-07-11T03:27:46.677614Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-11T03:21:01.687528Z digest=sha256:48cb670c49b550e5956141688d88d656c4f217bbec214f9abff367c9dcd3b355

Observation ec68947d-eccb-489f-8947-d8a121a03339 · outbound

This paper cites DeepSeekMoE: Towards Ultimate Expert Specialization in Mixture-of-Experts Language Models.

TriRoute: Unified Learned Routing for Joint Adaptive Attention, Experts, and KV-Cache Allocation DeepSeekMoE: Towards Ultimate Expert Specialization in Mixture-of-Experts Language Models

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-07-11T03:27:46.995209Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-11T03:21:01.687528Z digest=sha256:0a03847ac8d5077fef29b089aee5846d2ab7c38cb315e493a089cee997b267b7

Observation f27b0295-16d3-4eca-a599-62b153aec92e · outbound

This paper cites Fu, Stefano Ermon, Atri Rudra, and Christopher Ré.

TriRoute: Unified Learned Routing for Joint Adaptive Attention, Experts, and KV-Cache Allocation Fu, Stefano Ermon, Atri Rudra, and Christopher Ré

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-07-11T03:27:49.197139Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-11T03:21:01.687528Z digest=sha256:68f55447b1d74fdcbe03b8647fa7687986bf2f7ab8d28501e9fbb10bb79e0588

Observation c87a8819-21cc-48ac-9b77-812de7079b25 · outbound

This paper cites DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model.

TriRoute: Unified Learned Routing for Joint Adaptive Attention, Experts, and KV-Cache Allocation DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-07-11T03:27:47.022883Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-11T03:21:01.687528Z digest=sha256:82d12a987b1463efeb01df0c02fb0c16f4714a3f7d6f724aadef1f771c04aa90

Observation 30ed05dd-f00c-404d-939d-fd0b287cbb43 · outbound

This paper cites Uni- versal transformers.

TriRoute: Unified Learned Routing for Joint Adaptive Attention, Experts, and KV-Cache Allocation Uni- versal transformers

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-07-11T03:27:49.776327Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-11T03:21:01.687528Z digest=sha256:2b77a8a6c9816f847295a4e7ba8f65fd9d28950a5a809d532e25c0c505b10613

Observation 041e54fe-3af4-4cc1-8138-80c06638c52e · outbound

This paper cites LLM.int8(): 8-bit matrix multiplication for transformers at scale.

TriRoute: Unified Learned Routing for Joint Adaptive Attention, Experts, and KV-Cache Allocation LLM.int8(): 8-bit matrix multiplication for transformers at scale

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-07-11T03:27:49.848528Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-11T03:21:01.687528Z digest=sha256:5e54212ea02deaf6bb5c648b58bcdec0b80a00d8b99230c119d3ec0b4aeecf03

Observation 154b7390-4c5d-4d04-b6be-58eeeb3c2292 · outbound

This paper cites QLoRA: Efficient finetuningofquantizedLLMs.

TriRoute: Unified Learned Routing for Joint Adaptive Attention, Experts, and KV-Cache Allocation QLoRA: Efficient finetuningofquantizedLLMs

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-07-11T03:27:49.975482Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-11T03:21:01.687528Z digest=sha256:9d4c6396f02ee56ba9a9861f240c5d061c79b12d3a20e2788044ba813e01cca6

Observation 598e965d-9d08-4985-ab32-4b9b2ac4af9b · outbound

This paper cites Depth-adaptive transformer.

TriRoute: Unified Learned Routing for Joint Adaptive Attention, Experts, and KV-Cache Allocation Depth-adaptive transformer

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-07-11T03:27:49.226857Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-11T03:21:01.687528Z digest=sha256:4fd99349a44881fac9aa1324fe46137f67794f9fd117e0c0d69d8a3040ae4b01

Observation 83ca5649-5439-469e-a9ec-eb14d5177954 · outbound

This paper cites Reducing transformer depth on demand with structured dropout.

TriRoute: Unified Learned Routing for Joint Adaptive Attention, Experts, and KV-Cache Allocation Reducing transformer depth on demand with structured dropout

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-07-11T03:27:49.163885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-11T03:21:01.687528Z digest=sha256:ce8a90f0e868eccbf9b5a4e108e0cefa01032574ff83a3d84db811c0ce2c7bc7

Observation 5ddb967a-5b0f-4a43-b651-12448e90d709 · outbound

This paper cites Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity.Journal of Machine Learning Research (JMLR), 23(120):1–39.

TriRoute: Unified Learned Routing for Joint Adaptive Attention, Experts, and KV-Cache Allocation Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity.Journal of Machine Learning Research (JMLR), 23(120):1–39

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-07-11T03:27:50.163282Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-11T03:21:01.687528Z digest=sha256:e410313adcd0ddfaa72b10ba2fc3f5b676fb69022cfaeb52481ede244616d0f8

Observation 27c5362e-215d-4238-bcf4-7944a43492d5 · outbound

This paper cites GPTQ: Accurate post- training quantization for generative pre-trained transformers.

TriRoute: Unified Learned Routing for Joint Adaptive Attention, Experts, and KV-Cache Allocation GPTQ: Accurate post- training quantization for generative pre-trained transformers

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-07-11T03:27:50.035466Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-11T03:21:01.687528Z digest=sha256:d5787e3bcf4a62337a3d03af9563fef51a0ce5c612d358866eb925209d125304

Observation 23a0a73d-f166-400d-ab5c-b204819a0901 · outbound

This paper cites MegaBlocks: Efficient sparse training with mixture-of-experts.

TriRoute: Unified Learned Routing for Joint Adaptive Attention, Experts, and KV-Cache Allocation MegaBlocks: Efficient sparse training with mixture-of-experts

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-07-11T03:27:49.893346Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-11T03:21:01.687528Z digest=sha256:5aa32f0c1251e2636e61149a032abed226507a13ea11cb60340d10aeacf72225

Observation 9c4179aa-6bcf-4687-bac6-b8be76258c17 · outbound

This paper cites The Pile: An 800GB Dataset of Diverse Text for Language Modeling.

TriRoute: Unified Learned Routing for Joint Adaptive Attention, Experts, and KV-Cache Allocation The Pile: An 800GB Dataset of Diverse Text for Language Modeling

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-07-11T03:27:46.756821Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-11T03:21:01.687528Z digest=sha256:2ac27e1ae92413b996e2a68d97d1d5e015f39eebede55bcb343c00582e6b06af

Observation c62f7489-23b5-419c-82aa-bd80143974bd · outbound

This paper cites Adaptive Computation Time for Recurrent Neural Networks.

TriRoute: Unified Learned Routing for Joint Adaptive Attention, Experts, and KV-Cache Allocation Adaptive Computation Time for Recurrent Neural Networks

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-07-11T03:27:46.968390Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-11T03:21:01.687528Z digest=sha256:4a123ef9d02b6cee00065a662ea33f16841400002bf96f1d49e4209cce1b4e78

Observation 173ff941-c038-470c-b100-ab2a3fd97b21 · outbound

This paper cites Training Compute-Optimal Large Language Models.

TriRoute: Unified Learned Routing for Joint Adaptive Attention, Experts, and KV-Cache Allocation Training Compute-Optimal Large Language Models

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-07-11T03:27:47.112092Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-11T03:21:01.687528Z digest=sha256:03b3ccaf6d665a7868cbfbd72d393f4ea14c3d5dd03601d80f4d55222021aa11

Observation 447465bc-217c-40b2-b1e1-c99de9d7778d · outbound

This paper cites Mahoney, Yakun Sophia Shao, Kurt Keutzer, and Amir Gholami.

TriRoute: Unified Learned Routing for Joint Adaptive Attention, Experts, and KV-Cache Allocation Mahoney, Yakun Sophia Shao, Kurt Keutzer, and Amir Gholami

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-07-11T03:27:49.822401Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-11T03:21:01.687528Z digest=sha256:92068aba8916499d92c8344f81ed41d9f616af3ae20ad28933f42351b720f84a

Observation d8079ba3-b840-4b15-b822-0e11de00c7b1 · outbound

This paper cites Categorical reparameterization with Gumbel-Softmax.

TriRoute: Unified Learned Routing for Joint Adaptive Attention, Experts, and KV-Cache Allocation Categorical reparameterization with Gumbel-Softmax

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-07-11T03:27:49.588414Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-11T03:21:01.687528Z digest=sha256:7015f2ee145a59aead75a6c0f0337ac3587d82486f2298856b83760eb851ecb2

Observation 4d7f81cf-1db6-4dc2-bc20-210a94521bac · outbound

This paper cites Mixtral of Experts.

TriRoute: Unified Learned Routing for Joint Adaptive Attention, Experts, and KV-Cache Allocation Mixtral of Experts

Reference 28

Resolution
metadata mismatch
local_arxiv, observed 2026-07-11T03:27:47.053111Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-11T03:21:01.687528Z digest=sha256:00ef04f3c6ee0de30d20ed59cb1bd8b9b43ce9b7153644a8e42509b764d34c06

Observation f455abdb-917b-4b0c-8a6a-cf324a73ccb5 · outbound

This paper cites Scaling Laws for Neural Language Models.

TriRoute: Unified Learned Routing for Joint Adaptive Attention, Experts, and KV-Cache Allocation Scaling Laws for Neural Language Models

Reference 29

Resolution
metadata mismatch
local_arxiv, observed 2026-07-11T03:27:46.858674Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-11T03:21:01.687528Z digest=sha256:13f7eb46ca30aa59831bebfd36be7ae8663739514196ca717680f1beee9dc8a3

Observation 6b209d01-b2f7-43ec-80d0-7e88dae93dc2 · outbound

This paper cites Reformer: The efficient transformer.

TriRoute: Unified Learned Routing for Joint Adaptive Attention, Experts, and KV-Cache Allocation Reformer: The efficient transformer

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-07-11T03:27:49.484593Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-11T03:21:01.687528Z digest=sha256:958854121d90589353d9d8b89b523edf320c5f935acad9455b803c6c4ed61b6c

Observation 59880c2d-1f14-4510-a899-7a49269fdf5e · outbound

This paper cites GShard: Scaling giant models with conditional computation and automatic sharding.

TriRoute: Unified Learned Routing for Joint Adaptive Attention, Experts, and KV-Cache Allocation GShard: Scaling giant models with conditional computation and automatic sharding

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-07-11T03:27:49.438218Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-11T03:21:01.687528Z digest=sha256:2baacb621fbeb31b5f198edfdc9ffa2dd251fc234dcf7e9b5d563c23f7be4c41

Observation d2ac180e-cd6c-4a02-848a-411688e49dd3 · outbound

This paper cites BASE layers: Simplifying training of large, sparse models.

TriRoute: Unified Learned Routing for Joint Adaptive Attention, Experts, and KV-Cache Allocation BASE layers: Simplifying training of large, sparse models

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-07-11T03:27:49.361501Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-11T03:21:01.687528Z digest=sha256:91ac631cad829d874e784d56bd846078aae68f1970b7331bee71c28573f5615a

Observation 1c18df4a-df0c-441f-a228-00bc96886f8c · outbound

This paper cites KIVI: A tuning-free asymmetric 2bit quantization for KV cache.

TriRoute: Unified Learned Routing for Joint Adaptive Attention, Experts, and KV-Cache Allocation KIVI: A tuning-free asymmetric 2bit quantization for KV cache

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-07-11T03:27:49.869405Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-11T03:21:01.687528Z digest=sha256:3ac0d59ea25e52c6811379b7aaae4c2d5ac0a29a82b1127162df53f5a3a8e87f

Observation 81774670-d3d7-4251-aa82-f834564c41d1 · outbound

This paper cites Decoupled weight decay regularization.

TriRoute: Unified Learned Routing for Joint Adaptive Attention, Experts, and KV-Cache Allocation Decoupled weight decay regularization

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-07-11T03:27:49.951476Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-11T03:21:01.687528Z digest=sha256:8f35a2247759abe0bab57eea1264677409c5b10320de3716244942fc97b169fe

Observation 723db302-6504-4e02-97b7-aeef251d373a · outbound

This paper cites Maddison, Andriy Mnih, and Yee Whye Teh.

TriRoute: Unified Learned Routing for Joint Adaptive Attention, Experts, and KV-Cache Allocation Maddison, Andriy Mnih, and Yee Whye Teh

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-07-11T03:27:49.919615Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-11T03:21:01.687528Z digest=sha256:a735190aa97d4d35dc7daa560492427228d9711ddac58a5324b4c3defe1d56f1

Observation 572dd9e7-eb83-45d0-8494-342a99c2ec90 · outbound

This paper cites The LAMBADA dataset: Word prediction requiring a broad discourse context.

TriRoute: Unified Learned Routing for Joint Adaptive Attention, Experts, and KV-Cache Allocation The LAMBADA dataset: Word prediction requiring a broad discourse context

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-07-11T03:27:50.005465Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-11T03:21:01.687528Z digest=sha256:89c9f72cf9cb7db2cdcd61d97707aa2b72b0ffb015a6aa24cf1045985eb00cd8

Observation 0323126c-2c9c-4f04-844f-889b0cfff867 · outbound

This paper cites DeepSpeed-MoE: Advancing mixture-of- experts inference and training to power next-generation AI scale.

TriRoute: Unified Learned Routing for Joint Adaptive Attention, Experts, and KV-Cache Allocation DeepSpeed-MoE: Advancing mixture-of- experts inference and training to power next-generation AI scale

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-07-11T03:27:50.063823Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-11T03:21:01.687528Z digest=sha256:011efe9343e4b427db2e271531bf88f1c423d5c4117d0e87a09a7a352a7b68ab

Observation 47498db2-1298-4c4a-a40d-b0f60a7031c0 · outbound

This paper cites Mixture-of-Depths: Dynamically allocating compute in transformer-based language models.

TriRoute: Unified Learned Routing for Joint Adaptive Attention, Experts, and KV-Cache Allocation Mixture-of-Depths: Dynamically allocating compute in transformer-based language models

Reference 38

Resolution
verified exact
local_arxiv, observed 2026-07-11T03:27:46.733714Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-11T03:21:01.687528Z digest=sha256:4604a9889bdac82f1b6bed068ed046f74bdbc82d5ca847f9f9c1c820ac921d5b

Observation b7899cb2-a85c-438f-9a69-eea508631924 · outbound

This paper cites Hash layers for large sparse models.

TriRoute: Unified Learned Routing for Joint Adaptive Attention, Experts, and KV-Cache Allocation Hash layers for large sparse models

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-07-11T03:27:49.337262Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-11T03:21:01.687528Z digest=sha256:2d48ab10610ab7198e9a04936fc2354720427ad7e73778521a1c02ad957dc4fb

Observation a3f626d7-f5e9-4783-a00a-1bdfc5dcd1c3 · outbound

This paper cites WinoGrande: An adversarial winograd schema challenge at scale.Communications of the ACM, 64(9):99–106.

TriRoute: Unified Learned Routing for Joint Adaptive Attention, Experts, and KV-Cache Allocation WinoGrande: An adversarial winograd schema challenge at scale.Communications of the ACM, 64(9):99–106

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-07-11T03:27:49.640514Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-11T03:21:01.687528Z digest=sha256:5ec6c6c0151dcccdb2df7d86008a48f0c7ae302b45c82f626838dc8107bafe5b

Observation d366ebe4-852a-4557-976f-5bacdf5bb857 · outbound

This paper cites Confident adaptive language modeling.

TriRoute: Unified Learned Routing for Joint Adaptive Attention, Experts, and KV-Cache Allocation Confident adaptive language modeling

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-07-11T03:27:49.515000Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-11T03:21:01.687528Z digest=sha256:e15c14d34efc29e22104edac1a1ff16ecbf46e42aa4f75eaaa881e4a26de8d95

Observation 93ad1f44-0cdb-4b4e-828a-69cb4bd8bb58 · outbound

This paper cites Fast Transformer Decoding: One Write-Head is All You Need.

TriRoute: Unified Learned Routing for Joint Adaptive Attention, Experts, and KV-Cache Allocation Fast Transformer Decoding: One Write-Head is All You Need

Reference 42

Resolution
verified exact
local_arxiv, observed 2026-07-11T03:27:46.805207Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-11T03:21:01.687528Z digest=sha256:f89aede0d6c5378dc31f9a4425ce51cd5f772caa86ed1d354eaac115885088d5

Observation 98fa3c18-b935-48db-8bdb-833a3b0cd04e · outbound

This paper cites GLU Variants Improve Transformer.

TriRoute: Unified Learned Routing for Joint Adaptive Attention, Experts, and KV-Cache Allocation GLU Variants Improve Transformer

Reference 43

Resolution
verified exact
local_arxiv, observed 2026-07-11T03:27:46.780834Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-11T03:21:01.687528Z digest=sha256:d5a218c72e758eefad8b790006499ee5c1a270b53178d023fd8e7fdd281dd12c

Observation 252b5bfc-ac9d-4412-a291-f49745633493 · outbound

This paper cites Outrageously large neural networks: The sparsely-gated mixture-of-experts layer.

TriRoute: Unified Learned Routing for Joint Adaptive Attention, Experts, and KV-Cache Allocation Outrageously large neural networks: The sparsely-gated mixture-of-experts layer

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-07-11T03:27:49.540517Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-11T03:21:01.687528Z digest=sha256:c6796d76ba16952d7b3af0bcd47ffdf56d30021a8d376dda61bcba83b65bb7ec

Observation df4ee790-2647-4e2a-9d75-146d73252b30 · outbound

This paper cites RoFormer: Enhanced transformer with rotary position embedding.Neurocomputing, 568:127063.

TriRoute: Unified Learned Routing for Joint Adaptive Attention, Experts, and KV-Cache Allocation RoFormer: Enhanced transformer with rotary position embedding.Neurocomputing, 568:127063

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-07-11T03:27:49.613213Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-11T03:21:01.687528Z digest=sha256:d17f72d762e57175169f02a194067bd7dab629f7c06d14a7908066602976cc0c

Observation db412bd9-19d0-4b06-910b-a2f04c56fb64 · outbound

This paper cites Adaptive attention span in transformers.

TriRoute: Unified Learned Routing for Joint Adaptive Attention, Experts, and KV-Cache Allocation Adaptive attention span in transformers

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-07-11T03:27:49.803006Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-11T03:21:01.687528Z digest=sha256:67d67c1eba2b679a78a95b73b2f0832e3e40da0a6fac75ecf8d50a0f115b02f8

Observation b05a2f47-3dca-4638-80a6-e14c5c59f5de · outbound

This paper cites RedPajama: An open source recipe to reproduce LLaMA training dataset.

TriRoute: Unified Learned Routing for Joint Adaptive Attention, Experts, and KV-Cache Allocation RedPajama: An open source recipe to reproduce LLaMA training dataset

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-07-11T03:27:49.386197Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-11T03:21:01.687528Z digest=sha256:db808360d519ca7916b356e2212ea5103c6ef348a38c0394eb3939f8fdb693bc

Observation 591d3d97-a3bc-40a2-a3f8-184120383b4d · outbound

This paper cites Williams.

TriRoute: Unified Learned Routing for Joint Adaptive Attention, Experts, and KV-Cache Allocation Williams

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-07-11T03:27:50.130733Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-11T03:21:01.687528Z digest=sha256:e8f437ba50b914e7d10fe6b55cda1eae8825e4bdd58eaf3fb05e8dbd9cd73156

Observation e4dc1fcc-2c38-4d1c-adc2-fb6971e3ef57 · outbound

This paper cites Efficient streaming language models with attention sinks.

TriRoute: Unified Learned Routing for Joint Adaptive Attention, Experts, and KV-Cache Allocation Efficient streaming language models with attention sinks

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-07-11T03:27:50.093759Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-11T03:21:01.687528Z digest=sha256:a351da1531e5905159afdd6e27e82fb197066c7b14451acba15878ae3c8d379a

Observation 1c4f6a4c-4090-4364-94fe-a5ab1d4e26e5 · outbound

This paper cites Big bird: Transformers for longer sequences.

TriRoute: Unified Learned Routing for Joint Adaptive Attention, Experts, and KV-Cache Allocation Big bird: Transformers for longer sequences

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-07-11T03:27:49.705698Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-11T03:21:01.687528Z digest=sha256:c448233e4bf56d8d1eef1c1904b3ec4ceb5a27b490d9db676961fb1c1acc7637

Observation 94cf9ac2-63a2-4f92-93bf-5a894983ab02 · outbound

This paper cites HellaSwag: Can a machine really finish your sentence? InAssociation for Computational Linguistics (ACL), 2019.

TriRoute: Unified Learned Routing for Joint Adaptive Attention, Experts, and KV-Cache Allocation HellaSwag: Can a machine really finish your sentence? InAssociation for Computational Linguistics (ACL), 2019

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-07-11T03:27:49.411597Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-11T03:21:01.687528Z digest=sha256:815968ec39de7ff4d1f2d7358ebb053089fb03fcc0f4ad7ed7f989595d9fe31d

Observation 0e15645e-fb27-4c74-b7ed-9562502f5ce2 · outbound

This paper cites Root mean square layer normalization.

TriRoute: Unified Learned Routing for Joint Adaptive Attention, Experts, and KV-Cache Allocation Root mean square layer normalization

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-07-11T03:27:49.664028Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-11T03:21:01.687528Z digest=sha256:e8e6fbef91c109d9fdc32ae040eee796ab9318d95ab971add8a1297d9f97ff07

Observation 1bbaa7e6-ed5d-4778-b9a0-4476bc91cad6 · outbound

This paper cites Mixture of attention heads: Selecting attention heads per token.

TriRoute: Unified Learned Routing for Joint Adaptive Attention, Experts, and KV-Cache Allocation Mixture of attention heads: Selecting attention heads per token

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-07-11T03:27:49.684979Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-11T03:21:01.687528Z digest=sha256:e36f80a90a9c18a1e9eda0083dc2ce726993bb2709cab5c3066b2bcef00dbeaf

Observation e8afb211-eee2-4214-a981-7b401ce35bd6 · outbound

This paper cites H2O: Heavy-hitter oracle for efficient generative inference of large language models.

TriRoute: Unified Learned Routing for Joint Adaptive Attention, Experts, and KV-Cache Allocation H2O: Heavy-hitter oracle for efficient generative inference of large language models

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-07-11T03:27:49.728178Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-11T03:21:01.687528Z digest=sha256:dd20551c8d4e93fefd61b02f3cd966b289503ab54b6a12482746def456f36306

Observation 190c1693-34d1-4084-aae5-d3ebda45542e · outbound

This paper cites Dai, Zhifeng Chen, Quoc V.

TriRoute: Unified Learned Routing for Joint Adaptive Attention, Experts, and KV-Cache Allocation Dai, Zhifeng Chen, Quoc V

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-07-11T03:27:49.751235Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-11T03:21:01.687528Z digest=sha256:f32b1ca78a00f3055543f72cfce667b7089f936bdc10f8d68ae7129a4bc6cea7

Observation a43205cb-f2bf-4834-9110-62da85b0a662 · outbound

This paper cites ST-MoE: Designing Stable and Transferable Sparse Expert Models.

TriRoute: Unified Learned Routing for Joint Adaptive Attention, Experts, and KV-Cache Allocation ST-MoE: Designing Stable and Transferable Sparse Expert Models

Reference 56

Resolution
verified exact
local_arxiv, observed 2026-07-11T03:27:46.886389Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-11T03:21:01.687528Z digest=sha256:c3183cd5082742c9def93590995757be7b6c3e0277194296bf902b8c7f357962

Pith citing papers

No inbound Pith citation observations are available.