Pith. sign in

Paper Citation Record · LEDGER

RAD: Redundancy-Aware Distillation for Hybrid Models via Self-Speculative Decoding

As of 17 August 2026, this Paper Citation Record lists 51 of 51 outbound references and 4 inbound Pith citation observations for arXiv:2505.22135.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.22135 v1

Coverage vector

measured 51 of 51 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:20:31.702035Z

measured 55 of 55 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-30T18:22:45.702572Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-30T18:25:00.050125Z

Reference resolution

51 of 51 outbound references displayed

  • verified exact1
  • verified fuzzy18
  • unresolved32
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation b1d6a737-a5a7-4442-9199-b5d0eb899580 · outbound

This paper cites LongBench: A bilingual, multitask benchmark for long context understanding.

RAD: Redundancy-Aware Distillation for Hybrid Models via Self-Speculative Decoding LongBench: A bilingual, multitask benchmark for long context understanding

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:20:36.060030Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T13:20:26.790813Z digest=sha256:1d3fafe1308f062ab4c9821505f2bf63ae51ca0c862980c56e1b510a6010408e

Observation 3d0bfba0-6fbf-47d1-865a-ba680d1ea225 · outbound

This paper cites Puzzle: Distillation-Based NAS for Inference-Optimized LLMs.

RAD: Redundancy-Aware Distillation for Hybrid Models via Self-Speculative Decoding Puzzle: Distillation-Based NAS for Inference-Optimized LLMs

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T13:20:26.989808Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:20:26.989808Z digest=sha256:a009a5c1150877cbd1f2728117b51d629882111921dd23ca59c91acceb8d07be

Observation 7c9f6219-ef5a-41ae-815f-b24b6efb8679 · outbound

This paper cites On attention redundancy: A comprehensive study.

RAD: Redundancy-Aware Distillation for Hybrid Models via Self-Speculative Decoding On attention redundancy: A comprehensive study

Reference 3

Resolution
verified exact
doi, observed 2026-08-07T13:20:35.760682Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T13:20:27.103525Z digest=sha256:e8e566e80ba5a5f9e601546903f5e60ae3d46727b83f5ba3f4dbd05e5514b9a0

Observation 37df9508-f6c9-4ea0-a92a-8dd431c00551 · outbound

This paper cites Xing, J Zico Kolter, and Albert Gu.

RAD: Redundancy-Aware Distillation for Hybrid Models via Self-Speculative Decoding Xing, J Zico Kolter, and Albert Gu

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:20:35.505508Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T13:20:27.248509Z digest=sha256:cf59c5fd50c236d21820fe880cf4c7056abbac97ccac5f1096826076da233603

Observation b38696a6-12dd-4325-b989-cdd99ed49a36 · outbound

This paper cites Nemotron-H: A Family of Accurate and Efficient Hybrid Mamba-Transformer Models.

RAD: Redundancy-Aware Distillation for Hybrid Models via Self-Speculative Decoding Nemotron-H: A Family of Accurate and Efficient Hybrid Mamba-Transformer Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T13:20:27.356000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:20:27.356000Z digest=sha256:dba52c89c07caa9db200d13f27716cfe41d370684df58f396e8155b21e2e8e25

Observation ff0ac070-64ca-4102-9d9f-9fcba4d176da · outbound

This paper cites Accelerating Large Language Model Decoding with Speculative Sampling.

RAD: Redundancy-Aware Distillation for Hybrid Models via Self-Speculative Decoding Accelerating Large Language Model Decoding with Speculative Sampling

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T13:20:27.450977Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:20:27.450977Z digest=sha256:f8e2f0ec6a087a7df9afd0ec801f1e11093cd43968b7fe204cf89770c110f9be

Observation 363d877d-4e92-41a7-ac51-4bbce35aff6f · outbound

This paper cites GenQA: Generating Millions of Instructions from a Handful of Prompts.

RAD: Redundancy-Aware Distillation for Hybrid Models via Self-Speculative Decoding GenQA: Generating Millions of Instructions from a Handful of Prompts

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T13:20:27.547714Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:20:27.547714Z digest=sha256:d6c0bc98ec5bcf85f38f737150a117f2825560a2c97d1a957e87dc572df2b0ec

Observation 59bef50d-e274-414d-ab60-f4a43793b7e6 · outbound

This paper cites Streamlining redundant layers to compress large language models.

RAD: Redundancy-Aware Distillation for Hybrid Models via Self-Speculative Decoding Streamlining redundant layers to compress large language models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T13:20:27.628183Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:20:27.628183Z digest=sha256:bf7bad5d5b668c1dab826278ac16328f54052f578042ad819ce8b3d189dcdf20

Observation 8753613a-08b4-4139-ba3d-9870820ea493 · outbound

This paper cites Stuffed mamba: State collapse and state capacity of rnn-based long-context modeling, 2024.

RAD: Redundancy-Aware Distillation for Hybrid Models via Self-Speculative Decoding Stuffed mamba: State collapse and state capacity of rnn-based long-context modeling, 2024

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T13:20:27.719920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:20:27.719920Z digest=sha256:81f7a6c78412dad15b8bdb5cacbf25b2295251b6c635088fc44e5e7cd070a856

Observation 8c5c7f89-39fa-4751-8e52-9cbae0213827 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

RAD: Redundancy-Aware Distillation for Hybrid Models via Self-Speculative Decoding Training Verifiers to Solve Math Word Problems

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T13:20:27.892094Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:20:27.892094Z digest=sha256:eb3fb0f09b8f61af5c4496079371c25241752244bdfc87ca2cede904edaee778

Observation 384e185f-0837-47e5-b562-9587a6e48c20 · outbound

This paper cites Analyzing redundancy in pretrained transformer models.

RAD: Redundancy-Aware Distillation for Hybrid Models via Self-Speculative Decoding Analyzing redundancy in pretrained transformer models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T13:20:28.023710Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:20:28.023710Z digest=sha256:e31a488a1c7919d0a749831b9cf9358f5799e3cf67227a1bb8ed7f7922309848

Observation 583648ac-4682-4bdd-9aff-17ef39464327 · outbound

This paper cites Transformers are SSMs: Generalized models and efficient algorithms through structured state space duality.

RAD: Redundancy-Aware Distillation for Hybrid Models via Self-Speculative Decoding Transformers are SSMs: Generalized models and efficient algorithms through structured state space duality

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:20:35.041855Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T13:20:28.294265Z digest=sha256:aa4a4ca183747bb142689d5481960c551997d32088a5c0f921b4382aea408441

Observation cc73098e-1da3-4970-97b0-e7bf48a4461d · outbound

This paper cites Born again neural networks.

RAD: Redundancy-Aware Distillation for Hybrid Models via Self-Speculative Decoding Born again neural networks

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:20:34.845967Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T13:20:28.378302Z digest=sha256:d36f3ab444c35f7192360591bbc5ad954f78bad2100b9c52b327057ddec174d1

Observation 71bc3d20-18fb-46f7-8624-67b907939de3 · outbound

This paper cites A framework for few-shot language model evaluation, September 2021.

RAD: Redundancy-Aware Distillation for Hybrid Models via Self-Speculative Decoding A framework for few-shot language model evaluation, September 2021

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T13:20:28.465318Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:20:28.465318Z digest=sha256:e01cf567e9bb5b6a51ae6478410d7ad6fa1996d932438868b26cbc982672f265

Observation 10680adf-502b-468a-802e-fa5f7ad8fe13 · outbound

This paper cites Zamba: A Compact 7B SSM Hybrid Model.

RAD: Redundancy-Aware Distillation for Hybrid Models via Self-Speculative Decoding Zamba: A Compact 7B SSM Hybrid Model

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T13:20:28.522002Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:20:28.522002Z digest=sha256:5b2b3bb58b1e298e3fb78dd1da5553bcb7696b21330ea90723a56c1ca80bed96

Observation 6ae875eb-0f81-4bda-88bb-554d104e26fd · outbound

This paper cites RADLADS: Rapid attention distillation to linear attention decoders at scale, 2025.

RAD: Redundancy-Aware Distillation for Hybrid Models via Self-Speculative Decoding RADLADS: Rapid attention distillation to linear attention decoders at scale, 2025

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T13:20:28.654716Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:20:28.654716Z digest=sha256:e67f12864f8adadeeceaf38096ca9cb95665c9ab7135f653e82d409d68a9d3bd

Observation 8c2edfd2-3690-4b7b-a3bf-b2f719c17cbf · outbound

This paper cites an unresolved cited work.

RAD: Redundancy-Aware Distillation for Hybrid Models via Self-Speculative Decoding Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:20:34.597696Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T13:20:28.773262Z digest=sha256:efc4a4dab0ba1cadfc8d40efba83ddb58ff3ec14144205ee9df10e19ee3c3fb0

Observation 505045e6-a7dd-4d84-98fe-510186747f4a · outbound

This paper cites Mamba: Linear-Time Sequence Modeling with Selective State Spaces.

RAD: Redundancy-Aware Distillation for Hybrid Models via Self-Speculative Decoding Mamba: Linear-Time Sequence Modeling with Selective State Spaces

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T13:20:28.862191Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:20:28.862191Z digest=sha256:8d0d7092179a7c23d1c847fc1537b4969768f68a6c015a3527d5bc2b273f8c36

Observation 792ca933-5ff7-49ee-826d-caf3319ca0ea · outbound

This paper cites Combining recurrent, convolutional, and continuous-time models with linear state space layers.

RAD: Redundancy-Aware Distillation for Hybrid Models via Self-Speculative Decoding Combining recurrent, convolutional, and continuous-time models with linear state space layers

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T13:20:28.974218Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:20:28.974218Z digest=sha256:507081f2d5a750d492ee9410fefe21d8173bacd11d0d9a4b1825cc2df3176ee0

Observation 15bcccc1-c910-4cdc-baed-262a36cc83da · outbound

This paper cites Efficiently modeling long sequences with structured state spaces.

RAD: Redundancy-Aware Distillation for Hybrid Models via Self-Speculative Decoding Efficiently modeling long sequences with structured state spaces

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T13:20:29.130322Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:20:29.130322Z digest=sha256:5c70559f569817b7777261d51f615671d4b6c635df7fd777a311c44cccee41f0

Observation 04ea52e9-2045-47f9-9246-077abb6e75e4 · outbound

This paper cites CRUXEval: A Benchmark for Code Reasoning, Understanding and Execution.

RAD: Redundancy-Aware Distillation for Hybrid Models via Self-Speculative Decoding CRUXEval: A Benchmark for Code Reasoning, Understanding and Execution

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T13:20:29.178468Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:20:29.178468Z digest=sha256:57250ca09d5197373973edd202af12f4d18027b8b485f87f6abe66d518b0f05c

Observation f7e90854-c69a-4f0a-a85f-233c2ebb1f1e · outbound

This paper cites What Matters in Transformers? Not All Attention is Needed.

RAD: Redundancy-Aware Distillation for Hybrid Models via Self-Speculative Decoding What Matters in Transformers? Not All Attention is Needed

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T13:20:29.224634Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:20:29.224634Z digest=sha256:0a457aabbf47f29c3f4a23492e02f7501fd53de8a5e49688a492169bf4d0dc5e

Observation fc3a4e75-a77d-4ce1-9059-a8989c4d24dd · outbound

This paper cites Query-key normal- ization for transformers.

RAD: Redundancy-Aware Distillation for Hybrid Models via Self-Speculative Decoding Query-key normal- ization for transformers

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:20:34.313400Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T13:20:29.315730Z digest=sha256:5f3ce8f431e8118d32f40f5975c47ae5d880d4c63ff837653a08368da924afa6

Observation abe84c62-093f-46ab-8521-57e4a50b91b5 · outbound

This paper cites Distilling the Knowledge in a Neural Network.

RAD: Redundancy-Aware Distillation for Hybrid Models via Self-Speculative Decoding Distilling the Knowledge in a Neural Network

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T13:20:29.459790Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:20:29.459790Z digest=sha256:c1121e46d90ff31af15176e63a5f1f83bba440bfbc44311ca663fadbf285ba59

Observation e47c749d-e5bd-42bd-983c-2e4954533252 · outbound

This paper cites Kakade, and Eran Malach.

RAD: Redundancy-Aware Distillation for Hybrid Models via Self-Speculative Decoding Kakade, and Eran Malach

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:20:34.069698Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T13:20:29.576545Z digest=sha256:b8c497ea4a1f54f19247d0878f24941f974339b20bdfb8d6878af0550854b5d6

Observation 09b28262-294c-4137-b3e3-0ad36e27f9bb · outbound

This paper cites Fast inference from transformers via speculative decoding.

RAD: Redundancy-Aware Distillation for Hybrid Models via Self-Speculative Decoding Fast inference from transformers via speculative decoding

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:20:33.771256Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T13:20:29.652714Z digest=sha256:33d8a028ea6f3428ecfa4c384d40f8e493ad30058fa4f3ff53e33c4143807875

Observation 75c86232-ff5e-4dd8-a5a0-b90bd121581f · outbound

This paper cites Jamba: A Hybrid Transformer-Mamba Language Model.

RAD: Redundancy-Aware Distillation for Hybrid Models via Self-Speculative Decoding Jamba: A Hybrid Transformer-Mamba Language Model

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T13:20:29.796292Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:20:29.796292Z digest=sha256:676e25be1564ae43bc76f49941463d5cdd9578e4e5707d09f32a61bbcd037093

Observation dda98709-d7ab-4e42-859f-828c0bf10a46 · outbound

This paper cites ZeroEval: A unified framework for evaluating language models, 2024.

RAD: Redundancy-Aware Distillation for Hybrid Models via Self-Speculative Decoding ZeroEval: A unified framework for evaluating language models, 2024

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:20:33.495066Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T13:20:29.900421Z digest=sha256:b8c918a5606d5724070b617d59eedd8bd09199fc40f16b7e34b3b35aafd346cd

Observation f53e4d36-c4fd-4546-b2f3-19cc08572a2b · outbound

This paper cites Longhorn: State Space Models are Amortized Online Learners.

RAD: Redundancy-Aware Distillation for Hybrid Models via Self-Speculative Decoding Longhorn: State Space Models are Amortized Online Learners

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T13:20:29.988143Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:20:29.988143Z digest=sha256:b4decd0cd9349267c1c76b1b0ec10ba771db4626240f5736feee70cde05b722d

Observation 9ad9e62b-be82-4353-ae17-d2abbba95ad4 · outbound

This paper cites ShortGPT: Layers in large language models are more redundant than you expect, 2024.

RAD: Redundancy-Aware Distillation for Hybrid Models via Self-Speculative Decoding ShortGPT: Layers in large language models are more redundant than you expect, 2024

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:20:33.250157Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T13:20:30.103410Z digest=sha256:d9d7d8d3ff97577c6291dd1a571f708d867f528d320a2b96e02df30db329c414

Observation 889c22b1-ed36-471e-a1aa-f6dedd4dd1b1 · outbound

This paper cites Are sixteen heads really better than one? In Advances in Neural Information Processing Systems, volume 32, 2019.

RAD: Redundancy-Aware Distillation for Hybrid Models via Self-Speculative Decoding Are sixteen heads really better than one? In Advances in Neural Information Processing Systems, volume 32, 2019

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:20:33.150611Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T13:20:30.170022Z digest=sha256:0c9754bef5087dabfdb0468cbb3f72549fb24f1e83a4021947f1ba9aeccaeade

Observation 7f12d9e6-c704-420a-933c-0f6509bec4c7 · outbound

This paper cites Compact Language Models via Pruning and Knowledge Distillation.

RAD: Redundancy-Aware Distillation for Hybrid Models via Self-Speculative Decoding Compact Language Models via Pruning and Knowledge Distillation

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T13:20:30.255508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:20:30.255508Z digest=sha256:99bb521ef262fb881b97a6d151e239cd680082d60376bddce0a1e0f906faf844

Observation 9ee686fd-fb30-4a8d-ae89-6a29e4acba5a · outbound

This paper cites Bayesian Optimization: Open source constrained global optimization tool for Python, 2014–.

RAD: Redundancy-Aware Distillation for Hybrid Models via Self-Speculative Decoding Bayesian Optimization: Open source constrained global optimization tool for Python, 2014–

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:20:33.044723Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T13:20:30.313713Z digest=sha256:17bf4545e95cc244681b13e6d52a5b3564f634647db590654df9925d2adff099

Observation d6b35deb-f0e0-445e-9e10-a7ed99dbd88e · outbound

This paper cites Infinity instruct.

RAD: Redundancy-Aware Distillation for Hybrid Models via Self-Speculative Decoding Infinity instruct

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:20:32.936382Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T13:20:30.359409Z digest=sha256:0ce7450480cd7d5b0f84c81acd4c36f20281c81de274cba334641f5ca5d25872

Observation 72ced6b8-b35b-453a-b537-264ad8ee29d0 · outbound

This paper cites RWKV-7 "Goose" with Expressive Dynamic State Evolution.

RAD: Redundancy-Aware Distillation for Hybrid Models via Self-Speculative Decoding RWKV-7 "Goose" with Expressive Dynamic State Evolution

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T13:20:30.438474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:20:30.438474Z digest=sha256:69fad284b7e75542137e3f6e02b1b9876f092750e99929b099942b08d04b931e

Observation 019756d4-33e6-40b8-8be0-e74ef74e8d78 · outbound

This paper cites Compressive Transformers for Long-Range Sequence Modelling.

RAD: Redundancy-Aware Distillation for Hybrid Models via Self-Speculative Decoding Compressive Transformers for Long-Range Sequence Modelling

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T13:20:30.491759Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:20:30.491759Z digest=sha256:1d664d6ac308b932222b54d77f034718223661f6fd11c4b82474e0e5b8eabf0e

Observation f1eaacc1-e7b1-4749-876b-ef971285a3f2 · outbound

This paper cites Samba: Simple Hybrid State Space Models for Efficient Unlimited Context Language Modeling.

RAD: Redundancy-Aware Distillation for Hybrid Models via Self-Speculative Decoding Samba: Simple Hybrid State Space Models for Efficient Unlimited Context Language Modeling

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T13:20:30.608844Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:20:30.608844Z digest=sha256:71e6e5ed8fac99627834a347f7dd799389dbc08290a6070a4b378a755b3f0014

Observation cea3d994-b524-4a3b-b4c1-63a1a747d507 · outbound

This paper cites TAID: Temporally adaptive interpo- lated distillation for efficient knowledge transfer in language models.

RAD: Redundancy-Aware Distillation for Hybrid Models via Self-Speculative Decoding TAID: Temporally adaptive interpo- lated distillation for efficient knowledge transfer in language models

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:20:32.852697Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T13:20:30.666488Z digest=sha256:590d2e29ec1a0c9f240fe27d0e4424823642457a73e3fd58a81ee467dd87b102

Observation ec1d226a-ee4e-4caa-8427-5d7dc4317c79 · outbound

This paper cites TriForce: Lossless Acceleration of Long Sequence Generation with Hierarchical Speculative Decoding.

RAD: Redundancy-Aware Distillation for Hybrid Models via Self-Speculative Decoding TriForce: Lossless Acceleration of Long Sequence Generation with Hierarchical Speculative Decoding

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T13:20:30.755203Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:20:30.755203Z digest=sha256:f16621e98703736861c5f8d649eb34eb2ace98549bedda8bbfd6d23ad7185bf2

Observation 610db4f9-7003-4e02-a179-f5376252a4f0 · outbound

This paper cites Transformer Layers as Painters.

RAD: Redundancy-Aware Distillation for Hybrid Models via Self-Speculative Decoding Transformer Layers as Painters

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T13:20:30.828566Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:20:30.828566Z digest=sha256:0a72bc24b90abc91771c24377ee33d63f81a073fa56a0513054d8fe1b579748a

Observation ab900abb-5524-4b8b-8cf8-d262fa74d6a1 · outbound

This paper cites OpenHermes 2.5: An open dataset of synthetic data for generalist LLM assistants, 2023.

RAD: Redundancy-Aware Distillation for Hybrid Models via Self-Speculative Decoding OpenHermes 2.5: An open dataset of synthetic data for generalist LLM assistants, 2023

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:20:32.764491Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T13:20:30.967987Z digest=sha256:4929dc89a6499a556b59b2941017a08d71f9ef7afe9a08cce701a166d4489876

Observation b904818a-7db9-4b9f-9bc3-172616b79537 · outbound

This paper cites Attention is all you need.

RAD: Redundancy-Aware Distillation for Hybrid Models via Self-Speculative Decoding Attention is all you need

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T13:20:31.059850Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:20:31.059850Z digest=sha256:474782ae8b8e1c7df0d0cba5ffa8b85dcb76b586507634a746b54ae234a14abe

Observation 663bb92a-f0a8-4a5e-95a5-02ed4cb3d922 · outbound

This paper cites Rush, and Tri Dao.

RAD: Redundancy-Aware Distillation for Hybrid Models via Self-Speculative Decoding Rush, and Tri Dao

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:20:32.646348Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T13:20:31.148458Z digest=sha256:67ab79cd0ce88f64a7d8ccde719ac0bb1fcaed8c0651721c46924dfebd28c974

Observation 35d6f519-5f1e-4140-8bb9-330f78d21323 · outbound

This paper cites an unresolved cited work.

RAD: Redundancy-Aware Distillation for Hybrid Models via Self-Speculative Decoding Unresolved cited work

Reference 44

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:20:32.538808Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T13:20:31.282968Z digest=sha256:ddcba73f400c08c8f9c6fa32265868018f8bebe53254786d9d7471cb19d4203f

Observation 03735d46-1613-48b5-970a-3f4fed6f6db8 · outbound

This paper cites Parallelizing linear transformers with the delta rule over sequence length.

RAD: Redundancy-Aware Distillation for Hybrid Models via Self-Speculative Decoding Parallelizing linear transformers with the delta rule over sequence length

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:20:32.403849Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T13:20:31.400726Z digest=sha256:35b6f63f56579742586f06fa656858c355061d7e28619db01b073ff7fc4cf236

Observation 1aa69c5a-6049-4b7c-bf0f-4ca029f51e12 · outbound

This paper cites Gated delta networks: Improving Mamba2 with delta rule.

RAD: Redundancy-Aware Distillation for Hybrid Models via Self-Speculative Decoding Gated delta networks: Improving Mamba2 with delta rule

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:20:32.287914Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T13:20:31.483657Z digest=sha256:23c4219e3bc75a73194fbcb36ccb29ec335491db4e7705fd1bfc3cecf3e40cec

Observation f5f05cd6-5fe4-4479-b4c4-e5edecafde18 · outbound

This paper cites KV cache compression, but what must we give in return? a comprehensive benchmark of long context capable approaches.

RAD: Redundancy-Aware Distillation for Hybrid Models via Self-Speculative Decoding KV cache compression, but what must we give in return? a comprehensive benchmark of long context capable approaches

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:20:32.147055Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T13:20:31.572565Z digest=sha256:30dd3c9950ed90c7a5e95adc3056c44dfeccd3adb66d96ff39ce38f743d09f38

Observation 893308ac-ba7d-42ad-97ec-976addc14d2c · outbound

This paper cites Root mean square layer normalization.

RAD: Redundancy-Aware Distillation for Hybrid Models via Self-Speculative Decoding Root mean square layer normalization

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T13:20:31.661089Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:20:31.661089Z digest=sha256:b3e9191fba47b452e540d9d2ea26905126c24217d08209c3033ee845c2b086ac

Observation 2bdbaa1d-2463-486a-abd3-f3156652c70c · outbound

This paper cites Draft & Verify: Lossless Large Language Model Acceleration via Self-Speculative Decoding.

RAD: Redundancy-Aware Distillation for Hybrid Models via Self-Speculative Decoding Draft & Verify: Lossless Large Language Model Acceleration via Self-Speculative Decoding

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T13:20:31.702035Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:20:31.702035Z digest=sha256:831e2713bfa5c93cd9d3edb9c96989d257e3a47c144b70c6d6d30f68f94c2035

Observation 06fca4ef-0e7b-4b0e-b5cd-da60bcd63155 · outbound

This paper cites an unresolved cited work.

RAD: Redundancy-Aware Distillation for Hybrid Models via Self-Speculative Decoding Unresolved cited work

Reference 398

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:20:35.277833Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T13:20:28.145383Z digest=sha256:add00a31eeac73a50cfd522d71d05f3cc6c0b1601d0abb97f5d9e50d219d811a

Observation b9a285ad-0920-4b57-a372-21208309bb95 · outbound

This paper cites doi: 10.18653/v1/2024.acl-long.172.

RAD: Redundancy-Aware Distillation for Hybrid Models via Self-Speculative Decoding doi: 10.18653/v1/2024.acl-long.172

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T13:20:26.876727Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:20:26.876727Z digest=sha256:3a1d91d4ec765f13b48ed86e3e2efd7e3b1d64d3ac1cc26ee429635f82f92907

Pith citing papers

Observation 1979501e-7fdf-4a9e-a710-451b31732a64 · inbound

Long-Context Aware Upcycling: A New Frontier for Hybrid LLM Scaling cites this paper.

Long-Context Aware Upcycling: A New Frontier for Hybrid LLM Scaling RAD: Redundancy-Aware Distillation for Hybrid Models via Self-Speculative Decoding

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-11T22:01:10.802659Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-08T03:39:37.485602Z digest=sha256:4cf2e5a662d3ad48eef8c4e84629f3970569b768d578462b3e5f1b2d2033a22c

Observation 697489d4-87b9-4603-9952-c15806008e35 · inbound

Component-Aware Self-Speculative Decoding in Hybrid Language Models cites this paper.

Component-Aware Self-Speculative Decoding in Hybrid Language Models RAD: Redundancy-Aware Distillation for Hybrid Models via Self-Speculative Decoding

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-11T16:06:40.749250Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-09T18:40:17.017805Z digest=sha256:4523dcb03ee0249cb94dda2fa96f4707d02d1c5170a8e86d51c742f75733d691

Observation bd2fc1fd-38a6-409a-8940-04dabbe38e7e · inbound

Post-Trained MoE Can Skip Half Experts via Self-Distillation cites this paper.

Post-Trained MoE Can Skip Half Experts via Self-Distillation RAD: Redundancy-Aware Distillation for Hybrid Models via Self-Speculative Decoding

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-20T12:03:15.184871Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-20T12:00:35.496822Z digest=sha256:57f3d4bc4bcf4ceeab9de80c033910a95fd1ef54f62e8ae6f47e98cb6c1fafec

Observation 6c0ae147-ab2a-483c-9629-e5dca0897884 · inbound

Post-Trained MoE Can Skip Half Experts via Self-Distillation cites this paper.

Post-Trained MoE Can Skip Half Experts via Self-Distillation RAD: Redundancy-Aware Distillation for Hybrid Models via Self-Speculative Decoding

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-06-30T18:25:00.051504Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-30T18:22:45.702572Z digest=sha256:be3f17d216087938b4413b50516313079d92b263c01478499746d20983f6b465