Pith. sign in

Paper Citation Record · LEDGER

Jakiro: Boosting Speculative Decoding with Decoupled Multi-Head via MoE

As of 10 August 2026, this Paper Citation Record lists 31 of 31 outbound references and 3 inbound Pith citation observations for arXiv:2502.06282.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.06282 v1

Coverage vector

measured 31 of 31 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-08T16:12:29.003107Z

measured 34 of 34 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T21:08:29.233848Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-12T08:40:41.670402Z

Reference resolution

31 of 31 outbound references displayed

  • verified exact0
  • verified fuzzy3
  • unresolved27
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation f884bab6-ac51-4e5e-b762-2bf828d90301 · outbound

This paper cites Hydra: Sequentially-Dependent Draft Heads for Medusa Decoding.

Jakiro: Boosting Speculative Decoding with Decoupled Multi-Head via MoE Hydra: Sequentially-Dependent Draft Heads for Medusa Decoding

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-08T16:12:28.600149Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T16:12:28.600149Z digest=sha256:945e5f2e4397f03bd786865bb5a28c71abb3a9ad196d13137946bf6c5c2614ee

Observation 8e729bbe-b16a-4451-b2e9-224ae6d3c0c7 · outbound

This paper cites Accelerating Large Language Model Decoding with Speculative Sampling.

Jakiro: Boosting Speculative Decoding with Decoupled Multi-Head via MoE Accelerating Large Language Model Decoding with Speculative Sampling

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-08T16:12:28.719796Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T16:12:28.719796Z digest=sha256:cab5eb353a1ad04ed46701aa559151738c0f4245fd67b71ac3820756c2d263ea

Observation 3e338f42-b959-4484-9153-5ae547bfbe16 · outbound

This paper cites DoLa: Decoding by Contrasting Layers Improves Factuality in Large Language Models.

Jakiro: Boosting Speculative Decoding with Decoupled Multi-Head via MoE DoLa: Decoding by Contrasting Layers Improves Factuality in Large Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-08T16:12:28.731398Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T16:12:28.731398Z digest=sha256:822ebd942cba66b1714e644388d30ab5ab16bf36ac4b0f67c9f72548d3fcbe0e

Observation 6bce880c-a622-479c-9517-99ccbb98a1ad · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Jakiro: Boosting Speculative Decoding with Decoupled Multi-Head via MoE Training Verifiers to Solve Math Word Problems

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-08T16:12:28.736999Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T16:12:28.736999Z digest=sha256:64dce05fa44dbe978a4c0280342e3f657e3a309b4d9483d9b2b4495d31033d57

Observation e132727a-6662-497e-bfb9-f076ae469a48 · outbound

This paper cites URL https://doi.org/ 10.18653/v1/2022.acl-long.489.

Jakiro: Boosting Speculative Decoding with Decoupled Multi-Head via MoE URL https://doi.org/ 10.18653/v1/2022.acl-long.489

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-08T16:12:28.747736Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T16:12:28.747736Z digest=sha256:3422bf83faaa9e0fc6b0e1e0ba143827aa43fc37020a26d51d44a07ddf98952b

Observation 39a44266-bcfe-4f2c-9020-62900aded5a1 · outbound

This paper cites The Llama 3 Herd of Models.

Jakiro: Boosting Speculative Decoding with Decoupled Multi-Head via MoE The Llama 3 Herd of Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-08T16:12:28.767844Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T16:12:28.767844Z digest=sha256:96c2c22099ec8720b4c7fcf4779bcebfab216c35a727d1483b1efffa0a83992a

Observation b21138b6-a5d9-4917-8968-b42ce5c6347e · outbound

This paper cites Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity.

Jakiro: Boosting Speculative Decoding with Decoupled Multi-Head via MoE Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-08T16:12:28.803450Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T16:12:28.803450Z digest=sha256:b0a1ef8ce60a243631d3bc2908add898da484cbbc681ffda9e472ba74e360876

Observation 4b6f7776-1fa1-4f32-a2f2-823cce307eb1 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Jakiro: Boosting Speculative Decoding with Decoupled Multi-Head via MoE DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-08T16:12:28.851712Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T16:12:28.851712Z digest=sha256:5f76bbcc2e82dda448a48165a0f2ef6504f6b9b81b76c79498a9643d7d9a54b7

Observation af491f6d-2acc-400c-969b-9c9ff6549d97 · outbound

This paper cites Mixtral of Experts.

Jakiro: Boosting Speculative Decoding with Decoupled Multi-Head via MoE Mixtral of Experts

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-08T16:12:28.862322Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T16:12:28.862322Z digest=sha256:d007054495b36f8323c05d9b398fb0f819e9136e7403ca583bc12824f44ad58d

Observation 474ecf65-4c8e-4552-9d5c-53b895f8eedf · outbound

This paper cites Crafting papers on machine learning.

Jakiro: Boosting Speculative Decoding with Decoupled Multi-Head via MoE Crafting papers on machine learning

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-08T16:12:28.872124Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T16:12:28.872124Z digest=sha256:d2fb27bd800a778933f27393b527d86a837cdc15a4701e65d77336c387d33040

Observation da351026-4d11-4b83-a666-09db98255bc9 · outbound

This paper cites Contrastive Decoding: Open-ended Text Generation as Optimization.

Jakiro: Boosting Speculative Decoding with Decoupled Multi-Head via MoE Contrastive Decoding: Open-ended Text Generation as Optimization

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-08T16:12:28.877138Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T16:12:28.877138Z digest=sha256:be0aa971f01cf6677140a7197204b5321e984c615943cd2597993829d389de4f

Observation da4b6d6f-363f-48f0-8e87-2471d32ebda5 · outbound

This paper cites EAGLE: Speculative Sampling Requires Rethinking Feature Uncertainty.

Jakiro: Boosting Speculative Decoding with Decoupled Multi-Head via MoE EAGLE: Speculative Sampling Requires Rethinking Feature Uncertainty

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-08T16:12:28.882739Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T16:12:28.882739Z digest=sha256:c87de681b9909aeb23a8b0477122e4d7202c82b6ddac34f403fce486abe24436

Observation b17e9eeb-c3b0-4415-9ecd-2f57792989f9 · outbound

This paper cites SpecInfer: Accelerating Generative Large Language Model Serving with Tree-based Speculative Inference and Verification.

Jakiro: Boosting Speculative Decoding with Decoupled Multi-Head via MoE SpecInfer: Accelerating Generative Large Language Model Serving with Tree-based Speculative Inference and Verification

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-08T16:12:28.888564Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T16:12:28.888564Z digest=sha256:0bfc7298cf4d2dead1aaf06e17bff750402d8f700a07c85769f5ecd62dfad602

Observation 4d9080f2-e0a0-4cff-903e-92c91b7322af · outbound

This paper cites PaSS: Parallel Speculative Sampling.

Jakiro: Boosting Speculative Decoding with Decoupled Multi-Head via MoE PaSS: Parallel Speculative Sampling

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-08T16:12:28.893367Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T16:12:28.893367Z digest=sha256:978bd9ad416b82d44a80be14d803d95151c3ff1f42a548d4b4bab8cecc0f68d5

Observation d4edc6f5-7857-4b8a-8bb7-3bd0c4337de6 · outbound

This paper cites Abstractive Text Summarization Using Sequence-to-Sequence RNNs and Beyond.

Jakiro: Boosting Speculative Decoding with Decoupled Multi-Head via MoE Abstractive Text Summarization Using Sequence-to-Sequence RNNs and Beyond

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-08T16:12:28.914574Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T16:12:28.914574Z digest=sha256:b4b78e85554ca48dd85935258358916cc7059f13b4ab7a9531535fec95b95fbc

Observation 2fcb2dfe-9410-4471-b1d6-9ea6841abe07 · outbound

This paper cites Fast Transformer Decoding: One Write-Head is All You Need.

Jakiro: Boosting Speculative Decoding with Decoupled Multi-Head via MoE Fast Transformer Decoding: One Write-Head is All You Need

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-08T16:12:28.967060Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T16:12:28.967060Z digest=sha256:3564ab738c6eae0e2b76bdad4bcc9c3647e1684b7f834932ff38740573a7aed9

Observation 61f01c38-99fe-4aff-bc26-7a9c44931861 · outbound

This paper cites A Comprehensive Survey of Hallucination Mitigation Techniques in Large Language Models.

Jakiro: Boosting Speculative Decoding with Decoupled Multi-Head via MoE A Comprehensive Survey of Hallucination Mitigation Techniques in Large Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-08T16:12:28.983368Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T16:12:28.983368Z digest=sha256:678d3d37beb2d614d5253f0206ed0d85e552c1d508c3f47c396e33f39a9ded99

Observation 7dd7cfd9-a3be-4322-b895-9ca44f433103 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Jakiro: Boosting Speculative Decoding with Decoupled Multi-Head via MoE LLaMA: Open and Efficient Foundation Language Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-08T16:12:28.988275Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T16:12:28.988275Z digest=sha256:dba050eedda2baaf272ea78dee77377df6178afb292e94f17f0ebe56c3c959fd

Observation 8f738c5d-bf30-4c14-9d3d-1108f790f33a · outbound

This paper cites Unlocking Efficiency in Large Language Model Inference: A Comprehensive Survey of Speculative Decoding.

Jakiro: Boosting Speculative Decoding with Decoupled Multi-Head via MoE Unlocking Efficiency in Large Language Model Inference: A Comprehensive Survey of Speculative Decoding

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-08T16:12:28.993285Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T16:12:28.993285Z digest=sha256:13bae900e499b816a0fef090530a778154c150f5ee8225e848f3cb3b0b35f83e

Observation 570e4780-b17d-4ada-a99c-8c36d0044fc8 · outbound

This paper cites Generation meets verification: Accelerating large lan- guage model inference with smart parallel auto-correct decoding.

Jakiro: Boosting Speculative Decoding with Decoupled Multi-Head via MoE Generation meets verification: Accelerating large lan- guage model inference with smart parallel auto-correct decoding

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T16:12:29.494735Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T16:12:28.998289Z digest=sha256:7eeb86ef2e2f27016da9a1522b30ad8297bb47600011724a82c243a26b546986

Observation 9955b05b-532c-4eca-98a7-0352464f36d9 · outbound

This paper cites Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena.

Jakiro: Boosting Speculative Decoding with Decoupled Multi-Head via MoE Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-08T16:12:29.003107Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T16:12:29.003107Z digest=sha256:d05080f89f9ace05290ca27df87215d249c6470e1218985211d8eb5877a85197

Observation ecc449cf-0676-4e5f-8f50-3ca8d813c063 · outbound

This paper cites OpenAI o1 System Card.

Jakiro: Boosting Speculative Decoding with Decoupled Multi-Head via MoE OpenAI o1 System Card

Reference 1991

Resolution
unresolved
no resolver link, observed 2026-08-08T16:12:28.856639Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T16:12:28.856639Z digest=sha256:52a82f2d3652dea38af6c09bad5cc75eeea662fa964ebfbba348d34a85245f52

Observation 49ccf0ad-a702-4ade-b574-6480e58bd482 · outbound

This paper cites 1162/neco.1994.6.2.181.

Jakiro: Boosting Speculative Decoding with Decoupled Multi-Head via MoE 1162/neco.1994.6.2.181

Reference 1994

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T16:12:29.519818Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T16:12:28.866995Z digest=sha256:35579db84063265aa9e8197761c8605ac39c56a3c81fc965424f6d0d94c12496

Observation c6174ef1-abbd-4bb5-946a-8c8cad82e328 · outbound

This paper cites Contrastive Decoding Improves Reasoning in Large Language Models.

Jakiro: Boosting Speculative Decoding with Decoupled Multi-Head via MoE Contrastive Decoding Improves Reasoning in Large Language Models

Reference 2016

Resolution
unresolved
no resolver link, observed 2026-08-08T16:12:28.944851Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T16:12:28.944851Z digest=sha256:d6345d3ce87b6392a0772bfeef6cb286c73cb97648c7cd1020ea4faa902c6386

Observation a21879fa-9f4a-431d-874f-8d328aec1310 · outbound

This paper cites Unchosen Experts Can Contribute Too: Unleashing MoE Models' Power by Self-Contrast.

Jakiro: Boosting Speculative Decoding with Decoupled Multi-Head via MoE Unchosen Experts Can Contribute Too: Unleashing MoE Models' Power by Self-Contrast

Reference 2017

Resolution
metadata mismatch
local_arxiv, observed 2026-08-08T16:12:29.136455Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T16:12:28.972293Z digest=sha256:6a5584984402224bdaca0bc02e4cada7917967d96ec5a3612729488f2ad7d96a

Observation 1fe85545-1999-42b1-9176-8a732949e3b6 · outbound

This paper cites Instantaneous Grammatical Error Correction with Shallow Aggressive Decoding.

Jakiro: Boosting Speculative Decoding with Decoupled Multi-Head via MoE Instantaneous Grammatical Error Correction with Shallow Aggressive Decoding

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-08T16:12:28.978206Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T16:12:28.978206Z digest=sha256:1e46222c0fd0cbe67ce662b67ffd841afa19dad75d8ce12edcafa30e93529b74

Observation 57464e7c-6901-4ed4-8eea-db7ab24109e1 · outbound

This paper cites Better & Faster Large Language Models via Multi-token Prediction.

Jakiro: Boosting Speculative Decoding with Decoupled Multi-Head via MoE Better & Faster Large Language Models via Multi-token Prediction

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-08T16:12:28.826267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T16:12:28.826267Z digest=sha256:20cc507cf53f27ef6eeff3e64bd9e15997db69c25ca82f6043906e8082e99722

Observation c8ef77de-1e96-46c5-b767-7d30a30e71a6 · outbound

This paper cites Stablemoe: Stable routing strategy for mixture of experts.

Jakiro: Boosting Speculative Decoding with Decoupled Multi-Head via MoE Stablemoe: Stable routing strategy for mixture of experts

Reference 2021

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T16:12:29.535269Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-08T16:12:28.742751Z digest=sha256:c7cc088e4e150b49baf4ac6f3696a1a593c26baf2e067a50e29692d4f1c941a9

Observation c9cea444-63b4-48fe-82cf-0ee6ddb737e8 · outbound

This paper cites Medusa: Simple LLM Inference Acceleration Framework with Multiple Decoding Heads.

Jakiro: Boosting Speculative Decoding with Decoupled Multi-Head via MoE Medusa: Simple LLM Inference Acceleration Framework with Multiple Decoding Heads

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-08T16:12:28.678565Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T16:12:28.678565Z digest=sha256:d5b3865991d82d25eb23ca5d3126e374c6156ce9ff3873de361e5cce36fac8d2

Observation 886a97bc-2c02-4d90-8038-ce236e2d5a4a · outbound

This paper cites Evaluating Large Language Models Trained on Code.

Jakiro: Boosting Speculative Decoding with Decoupled Multi-Head via MoE Evaluating Large Language Models Trained on Code

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-08T16:12:28.725560Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T16:12:28.725560Z digest=sha256:28ff447b8ce54a9dd585fd853727eb0a4e1ae114b3859b6d35ea114c85d4382a

Observation fe40e5e0-e9ef-4347-887d-1329cbbe5dab · outbound

This paper cites Why Exposure Bias Matters: An Imitation Learning Perspective of Error Accumulation in Language Generation.

Jakiro: Boosting Speculative Decoding with Decoupled Multi-Head via MoE Why Exposure Bias Matters: An Imitation Learning Perspective of Error Accumulation in Language Generation

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-08T16:12:28.631982Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T16:12:28.631982Z digest=sha256:21811197253ca6c2439b46cd688ce61246e3e7aa5af74c8d0d641b350b95277e

Pith citing papers

Observation 29553447-b171-4959-9441-10f2afb7c03a · inbound

Towards Reasoning Era: A Survey of Long Chain-of-Thought for Reasoning Large Language Models cites this paper.

Towards Reasoning Era: A Survey of Long Chain-of-Thought for Reasoning Large Language Models Jakiro: Boosting Speculative Decoding with Decoupled Multi-Head via MoE

Reference 285

Resolution
verified exact
arxiv_id, observed 2026-05-12T08:40:41.672843Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-12T08:40:40.910461Z digest=sha256:65655c0c60cdfdbd37ec3bafc31408d962fcc8543038b9cf51c1cec9b7a25ba8

Observation 68fa8e58-4876-40ce-9b54-5dfa87a0838b · inbound

DraftExpert: Expansion-Aware Self-Speculative Decoding for End-Device MoE Inference cites this paper.

DraftExpert: Expansion-Aware Self-Speculative Decoding for End-Device MoE Inference Jakiro: Boosting Speculative Decoding with Decoupled Multi-Head via MoE

Reference 14

Resolution
unresolved
no resolver link, observed 2026-07-31T15:02:37.528129Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T15:02:37.528129Z digest=sha256:8608a744b420eeaf925468cc85fcf9e0b48e97955c0d254a4bf17c5320b865a9

Observation 65d5aa5c-6c44-4655-8b74-163e69737c36 · inbound

BALANCE: Hybrid Autoregressive-Speculative LLM Inference in Wireless Edge Networks cites this paper.

BALANCE: Hybrid Autoregressive-Speculative LLM Inference in Wireless Edge Networks Jakiro: Boosting Speculative Decoding with Decoupled Multi-Head via MoE

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T21:08:29.233848Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T21:08:29.233848Z digest=sha256:024b4fdec1c51ae31932514e971751d170553b9431dc11583d0c954bb4f1d119