Pith. sign in

Paper Citation Record · LEDGER

TriRoute: Unified Learned Routing for Joint Adaptive Attention, Experts, and KV-Cache Allocation

As of 10 August 2026, this Paper Citation Record lists 56 of 56 outbound references and 0 inbound Pith citation observations for arXiv:2607.06601.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.06601 v1

Coverage vector

measured 56 of 56 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-11T03:21:01.687528Z

measured 56 of 56 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

56 of 56 outbound references displayed

  • verified exact14
  • verified fuzzy39
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch3

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c6f1d44a-58a3-41cf-9474-5a12fd4a3c62 · outbound

This paper cites GQA: Training generalized multi-query transformer models from multi-head checkpoints.

TriRoute: Unified Learned Routing for Joint Adaptive Attention, Experts, and KV-Cache Allocation GQA: Training generalized multi-query transformer models from multi-head checkpoints

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-07-11T03:27:49.460673Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-11T03:21:01.687528Z digest=sha256:f15c21866ace14bf37c78d0ddae3ee74d477de27071a8f3877039db6b0701672

Observation a16e609c-2b80-4a1f-87fb-deb0281fda50 · outbound

This paper cites CoLT5: Faster long-range transformers with conditional computation.Empirical Methods in Natural Language Processing (EMNLP), 2023.

TriRoute: Unified Learned Routing for Joint Adaptive Attention, Experts, and KV-Cache Allocation CoLT5: Faster long-range transformers with conditional computation.Empirical Methods in Natural Language Processing (EMNLP), 2023

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-07-11T03:27:49.286954Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-11T03:21:01.687528Z digest=sha256:e42cc242c06f407c2c5c7bb7facc733030805c4ec8163b17134eb04bf10480f8

Observation 8cf0eac4-4837-4025-bc5b-debbfb13014b · outbound

This paper cites Longformer: The Long-Document Transformer.

TriRoute: Unified Learned Routing for Joint Adaptive Attention, Experts, and KV-Cache Allocation Longformer: The Long-Document Transformer

Reference 3

Resolution
metadata mismatch
local_arxiv, observed 2026-07-11T03:27:46.701443Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-11T03:21:01.687528Z digest=sha256:4083c176b11b16622f587b311d9cc3d21e6ae2a28ae9ee43bc99994bd57206f4

Observation b7702268-9035-4e0d-a534-eb861ed63f8e · outbound

This paper cites Estimating or Propagating Gradients Through Stochastic Neurons for Conditional Computation.

TriRoute: Unified Learned Routing for Joint Adaptive Attention, Experts, and KV-Cache Allocation Estimating or Propagating Gradients Through Stochastic Neurons for Conditional Computation

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-07-11T03:27:46.831362Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-11T03:21:01.687528Z digest=sha256:f2f1f930604fec791b04e0f4c49c0a7f423c1e9075769d12c85c3c7164b65ba2

Observation bc05f8c9-fe22-45da-a4e6-80438af1be29 · outbound

This paper cites Pythia: A suite for analyzing large language models across training and scaling.

TriRoute: Unified Learned Routing for Joint Adaptive Attention, Experts, and KV-Cache Allocation Pythia: A suite for analyzing large language models across training and scaling

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-07-11T03:27:49.564973Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-11T03:21:01.687528Z digest=sha256:15356d095ef42bff0981924fb0c281409a10995d23bd804c7c0f20363072805d

Observation 3b0a26b1-31fe-4e36-854a-cf17470d6d34 · outbound

This paper cites PIQA: Reasoning about physical commonsense in natural language.

TriRoute: Unified Learned Routing for Joint Adaptive Attention, Experts, and KV-Cache Allocation PIQA: Reasoning about physical commonsense in natural language

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-07-11T03:27:49.257627Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-11T03:21:01.687528Z digest=sha256:4b0116bb98a719dbd39c31cd89cd1cb5d889236266fe2d29fd3262cc3e707023

Observation c12fb47b-8289-42ca-b7e2-90b16f86ea63 · outbound

This paper cites Generating Long Sequences with Sparse Transformers.

TriRoute: Unified Learned Routing for Joint Adaptive Attention, Experts, and KV-Cache Allocation Generating Long Sequences with Sparse Transformers

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-07-11T03:27:46.912589Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-11T03:21:01.687528Z digest=sha256:0cb2a3756d53b094d66e4a1d82c362e3bbfe83fa04055e79438c0e9390d38453

Observation 1a129485-6300-444e-ba64-894f24bef20b · outbound

This paper cites Unified scaling laws for routed language models.

TriRoute: Unified Learned Routing for Joint Adaptive Attention, Experts, and KV-Cache Allocation Unified scaling laws for routed language models

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-07-11T03:27:49.317762Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-11T03:21:01.687528Z digest=sha256:46d32c3676975c9510709b388b2ee3b29e8a0c80417b2d80417338b5835bed44

Observation 4ee84a84-2d37-43f6-8e2f-590d3752ad65 · outbound

This paper cites Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge.

TriRoute: Unified Learned Routing for Joint Adaptive Attention, Experts, and KV-Cache Allocation Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-07-11T03:27:46.940199Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-11T03:21:01.687528Z digest=sha256:18ee2915c8245906a40de7f81591358d3d59888ee8884d18a0831030f6950c16

Observation a2cc56b8-e4e2-46f3-b6e9-a57a31b824c9 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

TriRoute: Unified Learned Routing for Joint Adaptive Attention, Experts, and KV-Cache Allocation Training Verifiers to Solve Math Word Problems

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-07-11T03:27:47.082510Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-11T03:21:01.687528Z digest=sha256:ea0e2ff20679c6b23a06a2369a1ffb56191e6065a1ee91bb1354e95df4ec7415

Observation 93ac1217-3f18-4a5a-8422-1ebd631c2aa1 · outbound

This paper cites SwitchHead: Accelerating Transformers with Mixture-of-Experts Attention.

TriRoute: Unified Learned Routing for Joint Adaptive Attention, Experts, and KV-Cache Allocation SwitchHead: Accelerating Transformers with Mixture-of-Experts Attention

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-07-11T03:27:46.677614Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-11T03:21:01.687528Z digest=sha256:4c533863e8755bcde9a897c74dd0bcc2d0722283817c1b96d100c22668ad8dd6

Observation ec68947d-eccb-489f-8947-d8a121a03339 · outbound

This paper cites DeepSeekMoE: Towards Ultimate Expert Specialization in Mixture-of-Experts Language Models.

TriRoute: Unified Learned Routing for Joint Adaptive Attention, Experts, and KV-Cache Allocation DeepSeekMoE: Towards Ultimate Expert Specialization in Mixture-of-Experts Language Models

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-07-11T03:27:46.995209Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-11T03:21:01.687528Z digest=sha256:86c846abdea01ca711776ae86d9aa635cb74aee1d6328216814f96bfd9864047

Observation f27b0295-16d3-4eca-a599-62b153aec92e · outbound

This paper cites Fu, Stefano Ermon, Atri Rudra, and Christopher Ré.

TriRoute: Unified Learned Routing for Joint Adaptive Attention, Experts, and KV-Cache Allocation Fu, Stefano Ermon, Atri Rudra, and Christopher Ré

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-07-11T03:27:49.197139Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-11T03:21:01.687528Z digest=sha256:45397cac57003259cad8fe7a0d4462e293bccdd419b636d3476175c475fa83b8

Observation c87a8819-21cc-48ac-9b77-812de7079b25 · outbound

This paper cites DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model.

TriRoute: Unified Learned Routing for Joint Adaptive Attention, Experts, and KV-Cache Allocation DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-07-11T03:27:47.022883Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-11T03:21:01.687528Z digest=sha256:736d873bcdf5d37b0105d2cb71d0286bc9785bfc2fad7017db7ab4fb2f95a537

Observation 30ed05dd-f00c-404d-939d-fd0b287cbb43 · outbound

This paper cites Uni- versal transformers.

TriRoute: Unified Learned Routing for Joint Adaptive Attention, Experts, and KV-Cache Allocation Uni- versal transformers

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-07-11T03:27:49.776327Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-11T03:21:01.687528Z digest=sha256:dbcfbca1093b3b634d6cea9ce14eefbfe67c1ed6fe617d68494d2dd20886d72e

Observation 041e54fe-3af4-4cc1-8138-80c06638c52e · outbound

This paper cites LLM.int8(): 8-bit matrix multiplication for transformers at scale.

TriRoute: Unified Learned Routing for Joint Adaptive Attention, Experts, and KV-Cache Allocation LLM.int8(): 8-bit matrix multiplication for transformers at scale

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-07-11T03:27:49.848528Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-11T03:21:01.687528Z digest=sha256:0c3e4aaad28ba452c411670414d2167dc8042e3b8e18e459728d63683377a7ac

Observation 154b7390-4c5d-4d04-b6be-58eeeb3c2292 · outbound

This paper cites QLoRA: Efficient finetuningofquantizedLLMs.

TriRoute: Unified Learned Routing for Joint Adaptive Attention, Experts, and KV-Cache Allocation QLoRA: Efficient finetuningofquantizedLLMs

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-07-11T03:27:49.975482Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-11T03:21:01.687528Z digest=sha256:e8994dd4208057100e8a037a702fb6b0146e04188b310d0ce26b3b636e5672f5

Observation 598e965d-9d08-4985-ab32-4b9b2ac4af9b · outbound

This paper cites Depth-adaptive transformer.

TriRoute: Unified Learned Routing for Joint Adaptive Attention, Experts, and KV-Cache Allocation Depth-adaptive transformer

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-07-11T03:27:49.226857Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-11T03:21:01.687528Z digest=sha256:78c7015640a1e7c5d5b486b0c50917133679b7c2c0addff6a766c4823980c868

Observation 83ca5649-5439-469e-a9ec-eb14d5177954 · outbound

This paper cites Reducing transformer depth on demand with structured dropout.

TriRoute: Unified Learned Routing for Joint Adaptive Attention, Experts, and KV-Cache Allocation Reducing transformer depth on demand with structured dropout

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-07-11T03:27:49.163885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-11T03:21:01.687528Z digest=sha256:86fc9c0386ba58bc079df2b7af14af2af8c4c8ee8cdc25a9449cbf3a89f3ed74

Observation 5ddb967a-5b0f-4a43-b651-12448e90d709 · outbound

This paper cites Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity.Journal of Machine Learning Research (JMLR), 23(120):1–39.

TriRoute: Unified Learned Routing for Joint Adaptive Attention, Experts, and KV-Cache Allocation Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity.Journal of Machine Learning Research (JMLR), 23(120):1–39

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-07-11T03:27:50.163282Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-11T03:21:01.687528Z digest=sha256:baeb21a34f381cbdbcaa33f387e7259875ebf059f0a10d4dcd6f09e661ad2b18

Observation 27c5362e-215d-4238-bcf4-7944a43492d5 · outbound

This paper cites GPTQ: Accurate post- training quantization for generative pre-trained transformers.

TriRoute: Unified Learned Routing for Joint Adaptive Attention, Experts, and KV-Cache Allocation GPTQ: Accurate post- training quantization for generative pre-trained transformers

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-07-11T03:27:50.035466Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-11T03:21:01.687528Z digest=sha256:46f09e7de813e207dd9b3a0ff308f1191950ba648ff5fe764e58748b2db72502

Observation 23a0a73d-f166-400d-ab5c-b204819a0901 · outbound

This paper cites MegaBlocks: Efficient sparse training with mixture-of-experts.

TriRoute: Unified Learned Routing for Joint Adaptive Attention, Experts, and KV-Cache Allocation MegaBlocks: Efficient sparse training with mixture-of-experts

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-07-11T03:27:49.893346Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-11T03:21:01.687528Z digest=sha256:ac8550fbd81715965eaef8dd5c2e3dd7ed4c46b119261c2ed991051e09009276

Observation 9c4179aa-6bcf-4687-bac6-b8be76258c17 · outbound

This paper cites The Pile: An 800GB Dataset of Diverse Text for Language Modeling.

TriRoute: Unified Learned Routing for Joint Adaptive Attention, Experts, and KV-Cache Allocation The Pile: An 800GB Dataset of Diverse Text for Language Modeling

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-07-11T03:27:46.756821Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-11T03:21:01.687528Z digest=sha256:57855e7f3dc6015a591555d8a8f5d69c68dbe7e16a7666936d7ca31c19730dd3

Observation c62f7489-23b5-419c-82aa-bd80143974bd · outbound

This paper cites Adaptive Computation Time for Recurrent Neural Networks.

TriRoute: Unified Learned Routing for Joint Adaptive Attention, Experts, and KV-Cache Allocation Adaptive Computation Time for Recurrent Neural Networks

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-07-11T03:27:46.968390Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-11T03:21:01.687528Z digest=sha256:0157a43839e112d4979387e4a890282f68966986eaa4c2b89050aa9125d40a93

Observation 173ff941-c038-470c-b100-ab2a3fd97b21 · outbound

This paper cites Training Compute-Optimal Large Language Models.

TriRoute: Unified Learned Routing for Joint Adaptive Attention, Experts, and KV-Cache Allocation Training Compute-Optimal Large Language Models

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-07-11T03:27:47.112092Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-11T03:21:01.687528Z digest=sha256:748f0bf293abaa80c6700597321952619ee555653ab7ccaa4ce4aa5e2374b9ea

Observation 447465bc-217c-40b2-b1e1-c99de9d7778d · outbound

This paper cites Mahoney, Yakun Sophia Shao, Kurt Keutzer, and Amir Gholami.

TriRoute: Unified Learned Routing for Joint Adaptive Attention, Experts, and KV-Cache Allocation Mahoney, Yakun Sophia Shao, Kurt Keutzer, and Amir Gholami

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-07-11T03:27:49.822401Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-11T03:21:01.687528Z digest=sha256:33a0d9f6c12417b9d59e6984a92a6f850f510edc5eaad3333afcd8cf2bc8ad86

Observation d8079ba3-b840-4b15-b822-0e11de00c7b1 · outbound

This paper cites Categorical reparameterization with Gumbel-Softmax.

TriRoute: Unified Learned Routing for Joint Adaptive Attention, Experts, and KV-Cache Allocation Categorical reparameterization with Gumbel-Softmax

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-07-11T03:27:49.588414Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-11T03:21:01.687528Z digest=sha256:3cd3e6cff285432ade182eed869ccfd38a391ca05868a1d46e0991cab65198d7

Observation 4d7f81cf-1db6-4dc2-bc20-210a94521bac · outbound

This paper cites Mixtral of Experts.

TriRoute: Unified Learned Routing for Joint Adaptive Attention, Experts, and KV-Cache Allocation Mixtral of Experts

Reference 28

Resolution
metadata mismatch
local_arxiv, observed 2026-07-11T03:27:47.053111Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-11T03:21:01.687528Z digest=sha256:5281ce901a409a51abb41aba57fc8585a2d130691548e8d1c36c07ff2abf7b09

Observation f455abdb-917b-4b0c-8a6a-cf324a73ccb5 · outbound

This paper cites Scaling Laws for Neural Language Models.

TriRoute: Unified Learned Routing for Joint Adaptive Attention, Experts, and KV-Cache Allocation Scaling Laws for Neural Language Models

Reference 29

Resolution
metadata mismatch
local_arxiv, observed 2026-07-11T03:27:46.858674Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-11T03:21:01.687528Z digest=sha256:f18f5f2b5d0201c1b815470945c32de5bce525cf7d0a89af823cd9839478d5df

Observation 6b209d01-b2f7-43ec-80d0-7e88dae93dc2 · outbound

This paper cites Reformer: The efficient transformer.

TriRoute: Unified Learned Routing for Joint Adaptive Attention, Experts, and KV-Cache Allocation Reformer: The efficient transformer

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-07-11T03:27:49.484593Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-11T03:21:01.687528Z digest=sha256:0cb66197a3e84b124ac9f6fd217e7bc62047fdfa40c233c208db7fa2528725fb

Observation 59880c2d-1f14-4510-a899-7a49269fdf5e · outbound

This paper cites GShard: Scaling giant models with conditional computation and automatic sharding.

TriRoute: Unified Learned Routing for Joint Adaptive Attention, Experts, and KV-Cache Allocation GShard: Scaling giant models with conditional computation and automatic sharding

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-07-11T03:27:49.438218Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-11T03:21:01.687528Z digest=sha256:6c2cbc2d339b8643e2fbd43d2740bf7777dcc13fb4859c9f04d9f9c1869e5616

Observation d2ac180e-cd6c-4a02-848a-411688e49dd3 · outbound

This paper cites BASE layers: Simplifying training of large, sparse models.

TriRoute: Unified Learned Routing for Joint Adaptive Attention, Experts, and KV-Cache Allocation BASE layers: Simplifying training of large, sparse models

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-07-11T03:27:49.361501Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-11T03:21:01.687528Z digest=sha256:8a41a204eac75b031ed84a57a1c133d21fe59e90c401cc6106391af0dffc4f21

Observation 1c18df4a-df0c-441f-a228-00bc96886f8c · outbound

This paper cites KIVI: A tuning-free asymmetric 2bit quantization for KV cache.

TriRoute: Unified Learned Routing for Joint Adaptive Attention, Experts, and KV-Cache Allocation KIVI: A tuning-free asymmetric 2bit quantization for KV cache

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-07-11T03:27:49.869405Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-11T03:21:01.687528Z digest=sha256:75d3ade7bce8dc1bc978c533b78b18651bf77f3e094f35e62cfac889a7e17697

Observation 81774670-d3d7-4251-aa82-f834564c41d1 · outbound

This paper cites Decoupled weight decay regularization.

TriRoute: Unified Learned Routing for Joint Adaptive Attention, Experts, and KV-Cache Allocation Decoupled weight decay regularization

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-07-11T03:27:49.951476Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-11T03:21:01.687528Z digest=sha256:62fa188903f29ea9e74c3096a2d41de5e574851f1cd908da0e9911a846ceddec

Observation 723db302-6504-4e02-97b7-aeef251d373a · outbound

This paper cites Maddison, Andriy Mnih, and Yee Whye Teh.

TriRoute: Unified Learned Routing for Joint Adaptive Attention, Experts, and KV-Cache Allocation Maddison, Andriy Mnih, and Yee Whye Teh

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-07-11T03:27:49.919615Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-11T03:21:01.687528Z digest=sha256:042659b117e2d647936bd03aa53b4c696742e21a51eb0512de090e1505f4db98

Observation 572dd9e7-eb83-45d0-8494-342a99c2ec90 · outbound

This paper cites The LAMBADA dataset: Word prediction requiring a broad discourse context.

TriRoute: Unified Learned Routing for Joint Adaptive Attention, Experts, and KV-Cache Allocation The LAMBADA dataset: Word prediction requiring a broad discourse context

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-07-11T03:27:50.005465Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-11T03:21:01.687528Z digest=sha256:54931b2214dabd89a37b32257ab446e28b063facdddbdd2a1f0be2fb05bfa302

Observation 0323126c-2c9c-4f04-844f-889b0cfff867 · outbound

This paper cites DeepSpeed-MoE: Advancing mixture-of- experts inference and training to power next-generation AI scale.

TriRoute: Unified Learned Routing for Joint Adaptive Attention, Experts, and KV-Cache Allocation DeepSpeed-MoE: Advancing mixture-of- experts inference and training to power next-generation AI scale

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-07-11T03:27:50.063823Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-11T03:21:01.687528Z digest=sha256:9c13e33b29272f3cbf22014a5306fcf21c60a0f684f3038a04e9659bd3253ab5

Observation 47498db2-1298-4c4a-a40d-b0f60a7031c0 · outbound

This paper cites Mixture-of-Depths: Dynamically allocating compute in transformer-based language models.

TriRoute: Unified Learned Routing for Joint Adaptive Attention, Experts, and KV-Cache Allocation Mixture-of-Depths: Dynamically allocating compute in transformer-based language models

Reference 38

Resolution
verified exact
local_arxiv, observed 2026-07-11T03:27:46.733714Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-11T03:21:01.687528Z digest=sha256:4b80e3bc9ece931cdc54b1ca797538e719fd93cd4901aba1f7824979a202ff71

Observation b7899cb2-a85c-438f-9a69-eea508631924 · outbound

This paper cites Hash layers for large sparse models.

TriRoute: Unified Learned Routing for Joint Adaptive Attention, Experts, and KV-Cache Allocation Hash layers for large sparse models

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-07-11T03:27:49.337262Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-11T03:21:01.687528Z digest=sha256:4fb2922cbe26319b3824eb79fe9870393b0acec2617a74c731358406128f11d7

Observation a3f626d7-f5e9-4783-a00a-1bdfc5dcd1c3 · outbound

This paper cites WinoGrande: An adversarial winograd schema challenge at scale.Communications of the ACM, 64(9):99–106.

TriRoute: Unified Learned Routing for Joint Adaptive Attention, Experts, and KV-Cache Allocation WinoGrande: An adversarial winograd schema challenge at scale.Communications of the ACM, 64(9):99–106

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-07-11T03:27:49.640514Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-11T03:21:01.687528Z digest=sha256:05eed115dcd65531ecd7d2fd57d1a2b20e5bceac959ec1e200e3d3e557597f12

Observation d366ebe4-852a-4557-976f-5bacdf5bb857 · outbound

This paper cites Confident adaptive language modeling.

TriRoute: Unified Learned Routing for Joint Adaptive Attention, Experts, and KV-Cache Allocation Confident adaptive language modeling

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-07-11T03:27:49.515000Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-11T03:21:01.687528Z digest=sha256:3ae816355aa7acacdb754b0851f5336dba71bfd8beec81039179b27cf6336551

Observation 93ad1f44-0cdb-4b4e-828a-69cb4bd8bb58 · outbound

This paper cites Fast Transformer Decoding: One Write-Head is All You Need.

TriRoute: Unified Learned Routing for Joint Adaptive Attention, Experts, and KV-Cache Allocation Fast Transformer Decoding: One Write-Head is All You Need

Reference 42

Resolution
verified exact
local_arxiv, observed 2026-07-11T03:27:46.805207Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-11T03:21:01.687528Z digest=sha256:39beadfccee876d34529ba943765f26354d16cf27e2e59ad01c8633e7a018210

Observation 98fa3c18-b935-48db-8bdb-833a3b0cd04e · outbound

This paper cites GLU Variants Improve Transformer.

TriRoute: Unified Learned Routing for Joint Adaptive Attention, Experts, and KV-Cache Allocation GLU Variants Improve Transformer

Reference 43

Resolution
verified exact
local_arxiv, observed 2026-07-11T03:27:46.780834Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-11T03:21:01.687528Z digest=sha256:7ed63e0c4ddf21388fe4b854405f4695487913736cb3f8acd54144e2ba8ff7d2

Observation 252b5bfc-ac9d-4412-a291-f49745633493 · outbound

This paper cites Outrageously large neural networks: The sparsely-gated mixture-of-experts layer.

TriRoute: Unified Learned Routing for Joint Adaptive Attention, Experts, and KV-Cache Allocation Outrageously large neural networks: The sparsely-gated mixture-of-experts layer

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-07-11T03:27:49.540517Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-11T03:21:01.687528Z digest=sha256:8295c1e6ace61e487d5eaebf416c3e1a714428361e031047d342c0108ad61612

Observation df4ee790-2647-4e2a-9d75-146d73252b30 · outbound

This paper cites RoFormer: Enhanced transformer with rotary position embedding.Neurocomputing, 568:127063.

TriRoute: Unified Learned Routing for Joint Adaptive Attention, Experts, and KV-Cache Allocation RoFormer: Enhanced transformer with rotary position embedding.Neurocomputing, 568:127063

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-07-11T03:27:49.613213Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-11T03:21:01.687528Z digest=sha256:3f95c7785ac849b096f055b6ce754c95372307d69b22843bdc0dc97fc475ade7

Observation db412bd9-19d0-4b06-910b-a2f04c56fb64 · outbound

This paper cites Adaptive attention span in transformers.

TriRoute: Unified Learned Routing for Joint Adaptive Attention, Experts, and KV-Cache Allocation Adaptive attention span in transformers

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-07-11T03:27:49.803006Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-11T03:21:01.687528Z digest=sha256:82cc3720586f5dcdc14086262d327619bdadbfe92eb1e17bca9c2cdc5c8dba73

Observation b05a2f47-3dca-4638-80a6-e14c5c59f5de · outbound

This paper cites RedPajama: An open source recipe to reproduce LLaMA training dataset.

TriRoute: Unified Learned Routing for Joint Adaptive Attention, Experts, and KV-Cache Allocation RedPajama: An open source recipe to reproduce LLaMA training dataset

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-07-11T03:27:49.386197Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-11T03:21:01.687528Z digest=sha256:005e4fee747ae4553a27aeaf18eb274d5c1516b5f9c83262c71bee3708e1da2d

Observation 591d3d97-a3bc-40a2-a3f8-184120383b4d · outbound

This paper cites Williams.

TriRoute: Unified Learned Routing for Joint Adaptive Attention, Experts, and KV-Cache Allocation Williams

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-07-11T03:27:50.130733Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-11T03:21:01.687528Z digest=sha256:dddc93793426ea1a0fb4278b57923a9eb5f9101c89a663796464a738058a6938

Observation e4dc1fcc-2c38-4d1c-adc2-fb6971e3ef57 · outbound

This paper cites Efficient streaming language models with attention sinks.

TriRoute: Unified Learned Routing for Joint Adaptive Attention, Experts, and KV-Cache Allocation Efficient streaming language models with attention sinks

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-07-11T03:27:50.093759Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-11T03:21:01.687528Z digest=sha256:112cce2a1958cb9a8ff425defb9849db288fbf3b3837bcc58b6acc4de7dc2a9b

Observation 1c4f6a4c-4090-4364-94fe-a5ab1d4e26e5 · outbound

This paper cites Big bird: Transformers for longer sequences.

TriRoute: Unified Learned Routing for Joint Adaptive Attention, Experts, and KV-Cache Allocation Big bird: Transformers for longer sequences

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-07-11T03:27:49.705698Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-11T03:21:01.687528Z digest=sha256:fa62d8919198f76115e51cd0f368084740996461c603473666f75e9a7f8d4042

Observation 94cf9ac2-63a2-4f92-93bf-5a894983ab02 · outbound

This paper cites HellaSwag: Can a machine really finish your sentence? InAssociation for Computational Linguistics (ACL), 2019.

TriRoute: Unified Learned Routing for Joint Adaptive Attention, Experts, and KV-Cache Allocation HellaSwag: Can a machine really finish your sentence? InAssociation for Computational Linguistics (ACL), 2019

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-07-11T03:27:49.411597Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-11T03:21:01.687528Z digest=sha256:f15fca566304e96cef04aa8fe0db4493f7d2477316312d0a7b920e0f5fca0b35

Observation 0e15645e-fb27-4c74-b7ed-9562502f5ce2 · outbound

This paper cites Root mean square layer normalization.

TriRoute: Unified Learned Routing for Joint Adaptive Attention, Experts, and KV-Cache Allocation Root mean square layer normalization

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-07-11T03:27:49.664028Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-11T03:21:01.687528Z digest=sha256:c849bc37059697c52d57811f4ddff37aa3062dfdd7dd97888aefe19becbda9c8

Observation 1bbaa7e6-ed5d-4778-b9a0-4476bc91cad6 · outbound

This paper cites Mixture of attention heads: Selecting attention heads per token.

TriRoute: Unified Learned Routing for Joint Adaptive Attention, Experts, and KV-Cache Allocation Mixture of attention heads: Selecting attention heads per token

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-07-11T03:27:49.684979Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-11T03:21:01.687528Z digest=sha256:2b77d2ec2306441da2ff9ce134690d477ac7082131db60a511df25e5ed0c8721

Observation e8afb211-eee2-4214-a981-7b401ce35bd6 · outbound

This paper cites H2O: Heavy-hitter oracle for efficient generative inference of large language models.

TriRoute: Unified Learned Routing for Joint Adaptive Attention, Experts, and KV-Cache Allocation H2O: Heavy-hitter oracle for efficient generative inference of large language models

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-07-11T03:27:49.728178Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-11T03:21:01.687528Z digest=sha256:1b193c07416a9f28f37145611a4c66af9e41718d1ef8a250e51262e9e6ae0e40

Observation 190c1693-34d1-4084-aae5-d3ebda45542e · outbound

This paper cites Dai, Zhifeng Chen, Quoc V.

TriRoute: Unified Learned Routing for Joint Adaptive Attention, Experts, and KV-Cache Allocation Dai, Zhifeng Chen, Quoc V

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-07-11T03:27:49.751235Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-11T03:21:01.687528Z digest=sha256:2475ebcdc898e058d16902069cae91cb077216d830be0697638687a2daddf92d

Observation a43205cb-f2bf-4834-9110-62da85b0a662 · outbound

This paper cites ST-MoE: Designing Stable and Transferable Sparse Expert Models.

TriRoute: Unified Learned Routing for Joint Adaptive Attention, Experts, and KV-Cache Allocation ST-MoE: Designing Stable and Transferable Sparse Expert Models

Reference 56

Resolution
verified exact
local_arxiv, observed 2026-07-11T03:27:46.886389Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-11T03:21:01.687528Z digest=sha256:f8ad64401b25c0557862d77cf5b93bc57003a9cb8ced5a00825c479ba8214592

Pith citing papers

No inbound Pith citation observations are available.