Pith. sign in

Paper Citation Record · LEDGER

Break the Sequential Dependency of LLM Inference Using Lookahead Decoding

As of 22 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 61 inbound Pith citation observations for arXiv:2402.02057.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2402.02057 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 61 of 61 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 61 of 61 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T11:10:15.392505Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T19:20:06.719647Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 2dc9cb81-ef4c-445f-91e8-f222bfd2cd1d · inbound

SAM Decoding: Speculative Decoding via Suffix Automaton cites this paper.

SAM Decoding: Speculative Decoding via Suffix Automaton Break the Sequential Dependency of LLM Inference Using Lookahead Decoding

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-12T19:32:52.187415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:32:52.187415Z digest=sha256:a5b22289025b0dfef42f4fa679812e08f32e28f933b8dfea2d052b3bfcc76a5d

Observation 04046587-add2-4c55-89a5-e26b4f72f2f5 · inbound

New Faithfulness-Centric Interpretability Paradigms for Natural Language Processing cites this paper.

New Faithfulness-Centric Interpretability Paradigms for Natural Language Processing Break the Sequential Dependency of LLM Inference Using Lookahead Decoding

Reference 249

Resolution
unresolved
no resolver link, observed 2026-08-12T11:42:30.982273Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T11:42:30.982273Z digest=sha256:dfa0c9753c498113b1294c58528a82f5b7902c1bd4aa01850cca7721d291fd21

Observation dc103ce0-8319-43b2-a2e8-2c5a98db2277 · inbound

PLD+: Accelerating LLM inference by leveraging Language Model Artifacts cites this paper.

PLD+: Accelerating LLM inference by leveraging Language Model Artifacts Break the Sequential Dependency of LLM Inference Using Lookahead Decoding

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-12T04:26:59.566128Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T04:26:59.566128Z digest=sha256:72010dce203cfcd677dcb7b661950fb9a7bcf810214f541bfbdbb6b2f79ea746

Observation c0a50ac7-e1dd-4a20-affb-db20d9e524ce · inbound

ZipAR: Parallel Auto-regressive Image Generation through Spatial Locality cites this paper.

ZipAR: Parallel Auto-regressive Image Generation through Spatial Locality Break the Sequential Dependency of LLM Inference Using Lookahead Decoding

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-11T21:54:15.645271Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:54:15.645271Z digest=sha256:8de9da1750f691fe175121f1dd8f6c47861d789877c4a87872a8c637d086a0ff

Observation 77b3ad3e-6f36-40e3-9058-6991421ce36b · inbound

Small Language Models (SLMs) Can Still Pack a Punch: A survey (updated 2026) cites this paper.

Small Language Models (SLMs) Can Still Pack a Punch: A survey (updated 2026) Break the Sequential Dependency of LLM Inference Using Lookahead Decoding

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-23T05:52:37.392964Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-23T05:47:48.488826Z digest=sha256:5e9e89a1d4a8938038350525d9ed4d089fce55120af22944df43868e384c21b3

Observation 09b45938-ff0d-42b4-a847-3a7b3579bf19 · inbound

Reward-Guided Speculative Decoding for Efficient LLM Reasoning cites this paper.

Reward-Guided Speculative Decoding for Efficient LLM Reasoning Break the Sequential Dependency of LLM Inference Using Lookahead Decoding

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-09T20:42:44.096964Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T20:42:44.096964Z digest=sha256:2cd7f6e31e6968ac8dcff2ddb2a8b010c4ba7e479950f575b17bc080c69f37f8

Observation 95d78aee-3486-4971-bc3b-734391255c7f · inbound

CopySpec: Accelerating LLMs with Speculative Copy-and-Paste Without Compromising Quality cites this paper.

CopySpec: Accelerating LLMs with Speculative Copy-and-Paste Without Compromising Quality Break the Sequential Dependency of LLM Inference Using Lookahead Decoding

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T23:17:51.164390Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:17:51.164390Z digest=sha256:106cc367322f38abbfb4ea8f8c06bf8209d3b68d3794ff7c25d045ad9c78d46a

Observation d631aeb9-8513-4d5d-9f62-1aa7901ec040 · inbound

SplitReason: Learning To Offload Reasoning cites this paper.

SplitReason: Learning To Offload Reasoning Break the Sequential Dependency of LLM Inference Using Lookahead Decoding

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-16T11:10:15.392505Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:10:15.392505Z digest=sha256:2e62a3cff8a4732de7b05c8cf2ae8fded8139ca4257e86e3cee5ab4347c6ae90

Observation 5e4d4656-9eb0-4a15-8dab-280faae347f2 · inbound

PipeSpec: Breaking Stage Dependencies in Hierarchical LLM Decoding cites this paper.

PipeSpec: Breaking Stage Dependencies in Hierarchical LLM Decoding Break the Sequential Dependency of LLM Inference Using Lookahead Decoding

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-16T04:21:45.427866Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T04:21:45.427866Z digest=sha256:72eb61ab2cce269d9bf2a6d5bc38514d58ca76c8a0cb0293616c0c88c4020c19

Observation 585b9b35-7c6c-4cd5-a061-e62d9f26b8eb · inbound

Scaling Laws for Speculative Decoding cites this paper.

Scaling Laws for Speculative Decoding Break the Sequential Dependency of LLM Inference Using Lookahead Decoding

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-15T23:17:34.805214Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:17:34.805214Z digest=sha256:0e2f618ecf0d39cc44c79609ed5a52c9648466be1ca3e113f4c2ae465f6a0d0f

Observation 5f30d67c-2053-4895-8e18-d40dd1e87e45 · inbound

On Next-Token Prediction in LLMs: How End Goals Determine the Consistency of Decoding Algorithms cites this paper.

On Next-Token Prediction in LLMs: How End Goals Determine the Consistency of Decoding Algorithms Break the Sequential Dependency of LLM Inference Using Lookahead Decoding

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T21:05:57.701011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T21:05:57.701011Z digest=sha256:d6ba57d4f0ba8b248aff64af294f69b4d36d0e450fe0020d5a81d0982938b870

Observation 40c1840c-6422-42a9-9344-4dacfb33782b · inbound

Speculative Decoding Reimagined for Multimodal Large Language Models cites this paper.

Speculative Decoding Reimagined for Multimodal Large Language Models Break the Sequential Dependency of LLM Inference Using Lookahead Decoding

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T15:42:52.964039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:42:52.964039Z digest=sha256:ceccf4ed65f9ef3c4224f9a00e48e33b093de78f885a3851cf1ed046b711c34b

Observation dca665e3-b590-4a66-965f-6dca443a955e · inbound

CoDec: Prefix-Shared Decoding Kernel for LLMs cites this paper.

CoDec: Prefix-Shared Decoding Kernel for LLMs Break the Sequential Dependency of LLM Inference Using Lookahead Decoding

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T14:47:55.240002Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:47:55.240002Z digest=sha256:dba53ac066902b18e4d0188bb75b7cbdc756fd287c9b55f47de28309a8c7afba

Observation 95399fe2-9958-48b3-81f7-48399dedba3f · inbound

AutoChemSchematic AI: Agentic Physics-Aware Automation for Chemical Manufacturing Scale-Up cites this paper.

AutoChemSchematic AI: Agentic Physics-Aware Automation for Chemical Manufacturing Scale-Up Break the Sequential Dependency of LLM Inference Using Lookahead Decoding

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:26.878891Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:26.878891Z digest=sha256:5261b36bda9daa58c5fea776b3b1235281ef1f45c2622cd889fab85c25718441

Observation 5edc1148-de69-4e6f-9427-6c4c22e66d0d · inbound

Mamba Drafters for Speculative Decoding cites this paper.

Mamba Drafters for Speculative Decoding Break the Sequential Dependency of LLM Inference Using Lookahead Decoding

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T11:55:12.326986Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:55:12.326986Z digest=sha256:a66fb67b10bfcb6379ce6ff1705c78cffdd79f9abf1eaac2848ec53fd104ca57

Observation 32c05481-ad55-4ea9-aadc-1b1f9b46b55c · inbound

Consultant Decoding: Yet Another Synergistic Mechanism cites this paper.

Consultant Decoding: Yet Another Synergistic Mechanism Break the Sequential Dependency of LLM Inference Using Lookahead Decoding

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T11:30:55.107331Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:30:55.107331Z digest=sha256:eac07f75a3c747610e95ee96109f8cc4ada645325ac2557665faa5692818c36d

Observation b544baf3-a714-4c4d-b7e8-c31ffbd3353b · inbound

S$^4$C: Speculative Sampling with Syntactic and Semantic Coherence for Efficient Inference of Large Language Models cites this paper.

S$^4$C: Speculative Sampling with Syntactic and Semantic Coherence for Efficient Inference of Large Language Models Break the Sequential Dependency of LLM Inference Using Lookahead Decoding

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T00:22:44.539645Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:22:44.539645Z digest=sha256:3a687ba24878fb39b7afda9a801b69c6cafe381c9d851dc111b407b723f6631d

Observation d3024e64-8314-4b8a-999c-d93988f42377 · inbound

Scaling Speculative Decoding with Lookahead Reasoning cites this paper.

Scaling Speculative Decoding with Lookahead Reasoning Break the Sequential Dependency of LLM Inference Using Lookahead Decoding

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-15T18:32:07.971833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:32:07.971833Z digest=sha256:91817211837b963a25be56c11f2d05b0cb664758a7d035a6bf1f79429f424623

Observation cb2996ac-adaf-43d5-9628-bfbb5126691a · inbound

XSpecMesh: Quality-Preserving Auto-Regressive Mesh Generation Acceleration via Multi-Head Speculative Decoding cites this paper.

XSpecMesh: Quality-Preserving Auto-Regressive Mesh Generation Acceleration via Multi-Head Speculative Decoding Break the Sequential Dependency of LLM Inference Using Lookahead Decoding

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-06T10:29:56.326027Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T10:29:56.326027Z digest=sha256:c1cb5bb33e214127abc1dd4c45406cfae76c477fc6f4df96bbb902ae8fe01ee6

Observation e6b9765c-75ed-4e60-9767-555d16cdd50d · inbound

Leveraging Generative Models for Real-Time Query-Driven Text Summarization in Large-Scale Web Search cites this paper.

Leveraging Generative Models for Real-Time Query-Driven Text Summarization in Large-Scale Web Search Break the Sequential Dependency of LLM Inference Using Lookahead Decoding

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T16:49:49.664355Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:49:49.664355Z digest=sha256:c8e4b6bcc652c9e00eeb68c7dc18f563c82df66c6095ff1f8a8126c8108dfe81

Observation c386ed0c-2dca-44de-98e1-0ab2d484f2ae · inbound

Scaling Up, Speeding Up: A Benchmark of Speculative Decoding for Efficient LLM Test-Time Scaling cites this paper.

Scaling Up, Speeding Up: A Benchmark of Speculative Decoding for Efficient LLM Test-Time Scaling Break the Sequential Dependency of LLM Inference Using Lookahead Decoding

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-05T13:48:21.420932Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:48:21.420932Z digest=sha256:bc55a8729676b50e557861e0555fdcb0d2d1d0e534f480f06291902c09e733ed

Observation 5aa796af-2c72-4e44-94bc-67574961d506 · inbound

FluentAvatar: Flicker-Free Talking-Head Animation via Phoneme-Guided Autoregressive Modeling cites this paper.

FluentAvatar: Flicker-Free Talking-Head Animation via Phoneme-Guided Autoregressive Modeling Break the Sequential Dependency of LLM Inference Using Lookahead Decoding

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-18T16:42:43.871069Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-18T16:42:25.803856Z digest=sha256:bb0bfd895d4f3ee87dce65cbaddc95f0fec58b12c926eb689af3f366f306c09b

Observation d876c74f-08e6-47df-871a-30eda85d946a · inbound

Structuring The Future: Diffusion LLM Speculative Decoding via Calibrated Draft Graphs cites this paper.

Structuring The Future: Diffusion LLM Speculative Decoding via Calibrated Draft Graphs Break the Sequential Dependency of LLM Inference Using Lookahead Decoding

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-04T15:58:13.252756Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T15:58:13.252756Z digest=sha256:60b59cb9667acf592ea795671e7dca0edc8b93dae5d8061e1f28c0cf8d28af96

Observation c277a0cf-1183-4a1e-9961-cf16eaf8edc9 · inbound

HiSpec: Hierarchical Speculative Decoding for LLMs cites this paper.

HiSpec: Hierarchical Speculative Decoding for LLMs Break the Sequential Dependency of LLM Inference Using Lookahead Decoding

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-04T13:15:30.143680Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T13:15:30.143680Z digest=sha256:b918a38f4a9977e8d48b041612fcd03d0fe4663602d10eef47edb2b577d1ce63

Observation e87571e5-1335-4918-8a8d-04f5a91143c3 · inbound

SelfJudge: Faster Speculative Decoding via Self-Supervised Judge Verification cites this paper.

SelfJudge: Faster Speculative Decoding via Self-Supervised Judge Verification Break the Sequential Dependency of LLM Inference Using Lookahead Decoding

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T15:51:39.133988Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:51:39.133988Z digest=sha256:6e4ea0d74d316525dffa9afbc9d3a0721f88ed0d5bcdb9bc528fdbd2a5e15661

Observation 8c6ed277-5f95-4911-bc79-30a7c17631f3 · inbound

TokenTiming: A Dynamic Alignment Method for Universal Speculative Decoding Model Pairs cites this paper.

TokenTiming: A Dynamic Alignment Method for Universal Speculative Decoding Model Pairs Break the Sequential Dependency of LLM Inference Using Lookahead Decoding

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T06:30:59.768495Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-18T06:27:08.991426Z digest=sha256:c50d5b57d1ea8b9d9ca65a1510a1056f9efb453611cdcc039eec672905ce22eb

Observation 80266967-58e9-41da-b603-9312dc9620b6 · inbound

VVS: Accelerating Speculative Decoding for Visual Autoregressive Generation via Partial Verification Skipping cites this paper.

VVS: Accelerating Speculative Decoding for Visual Autoregressive Generation via Partial Verification Skipping Break the Sequential Dependency of LLM Inference Using Lookahead Decoding

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-17T21:35:17.556880Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-17T21:34:14.371914Z digest=sha256:88af8bb8900136363b8b11ae5b2771e3eda6a5e393a71d2c572d495863ee0e86

Observation 6c35afb9-f427-46b6-9e0e-4251eb1ff016 · inbound

Seer: Online Context Learning for Fast Synchronous LLM Reinforcement Learning cites this paper.

Seer: Online Context Learning for Fast Synchronous LLM Reinforcement Learning Break the Sequential Dependency of LLM Inference Using Lookahead Decoding

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:40:14.499075Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-17T20:38:30.169363Z digest=sha256:8f1a15bc79ad5cb8996435cbd9ab9ba68cf5e0ee8b92819941b14e9b78251a48

Observation 481e23c1-0330-41b3-8a55-bb43383367f4 · inbound

Parallel Decoder Transformer: Planner-Conditioned Latent Coordination for Model-Intrinsic Parallel Generation cites this paper.

Parallel Decoder Transformer: Planner-Conditioned Latent Coordination for Model-Intrinsic Parallel Generation Break the Sequential Dependency of LLM Inference Using Lookahead Decoding

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-03T17:22:04.351657Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:22:04.351657Z digest=sha256:902e38db92ef5ab741887261f388bc2c188e94f24947444344a26572e1e97909

Observation 7249766e-3642-4230-9cc2-f6d02175e045 · inbound

Double: Breaking the Acceleration Limit via Double Retrieval Speculative Parallelism cites this paper.

Double: Breaking the Acceleration Limit via Double Retrieval Speculative Parallelism Break the Sequential Dependency of LLM Inference Using Lookahead Decoding

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-16T16:41:06.261854Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-16T16:39:26.611704Z digest=sha256:5072edb6f2b377541cffcedc5caf39b6f0b2e2b88af1489c79267d0f3ca0ab48

Observation d7f34214-710f-4744-bca7-da9fc428576a · inbound

When RL Meets Adaptive Speculative Training: A Unified Training-Serving System cites this paper.

When RL Meets Adaptive Speculative Training: A Unified Training-Serving System Break the Sequential Dependency of LLM Inference Using Lookahead Decoding

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-16T06:37:28.926667Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-16T06:33:41.860803Z digest=sha256:b953bc33df8b845e8ebdb456565e4f25af87db3b46e25379db8c7b0071778db1

Observation 161bbc7d-5dd7-417e-917d-614dc66196c9 · inbound

Multi-Drafter Speculative Decoding with Alignment Feedback cites this paper.

Multi-Drafter Speculative Decoding with Alignment Feedback Break the Sequential Dependency of LLM Inference Using Lookahead Decoding

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T22:50:51.608091Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-10T19:31:20.953726Z digest=sha256:22a3c5602f535e6dabf3f871f39f64a7019abcc9de65afa42271ce4272af244f

Observation cad5c742-0010-4ab2-bb08-eeee2d1f2243 · inbound

ECHO: Elastic Speculative Decoding with Sparse Gating for High-Concurrency Scenarios cites this paper.

ECHO: Elastic Speculative Decoding with Sparse Gating for High-Concurrency Scenarios Break the Sequential Dependency of LLM Inference Using Lookahead Decoding

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-15T14:10:03.124996Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-15T14:09:59.938676Z digest=sha256:e65e2d2fa92c96c442899218bb0859ca3186a12c9836394ce795c0a9e4cb35a6

Observation 9f8f78f0-47cf-44a7-8162-5a6ebfcda569 · inbound

SpecBound: Adaptive Bounded Self-Speculation with Layer-wise Confidence Calibration cites this paper.

SpecBound: Adaptive Bounded Self-Speculation with Layer-wise Confidence Calibration Break the Sequential Dependency of LLM Inference Using Lookahead Decoding

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T09:00:58.706352Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-10T16:22:05.049712Z digest=sha256:b8cb56cf459b18f2b3932d8ccae0fb72803f58ba91fb68bd4f552b2bc2c8bbfc

Observation 18aa11fd-0887-4c7b-b36d-29c82d93012d · inbound

Calibrated Speculative Decoding: Frequency-Guided Candidate Selection for Efficient Inference cites this paper.

Calibrated Speculative Decoding: Frequency-Guided Candidate Selection for Efficient Inference Break the Sequential Dependency of LLM Inference Using Lookahead Decoding

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-10T13:30:26.836168Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-10T13:26:05.910987Z digest=sha256:3ceb87407fcc9d90e9c9641c516186ea010c7c6a6fb4d43d90368f43731c8777

Observation 10e278ea-df8d-4c25-aa1d-c2c913b7ff78 · inbound

Reference-Augmented Learning for Precise Tracking Policy of Tendon-Driven Continuum Robots cites this paper.

Reference-Augmented Learning for Precise Tracking Policy of Tendon-Driven Continuum Robots Break the Sequential Dependency of LLM Inference Using Lookahead Decoding

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-07-01T08:55:34.764013Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-01T08:53:21.013613Z digest=sha256:cf9863ae73517b45181482b46fead239261505770892bc7271e08473383b85ec

Observation 62095661-8c01-44df-8209-e04431b7cf4a · inbound

NVLLM: A 3D NAND-Centric Architecture Enabling Edge on-Device LLM Inference cites this paper.

NVLLM: A 3D NAND-Centric Architecture Enabling Edge on-Device LLM Inference Break the Sequential Dependency of LLM Inference Using Lookahead Decoding

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-12T00:46:13.421697Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-07T14:14:33.939889Z digest=sha256:48c7221b63463c35c22deae1779f44b4128ecb61f09d902557a8a07238c298d0

Observation 15ef1948-d9e8-4303-888c-333cc1e5568f · inbound

Continuous Latent Diffusion Language Model cites this paper.

Continuous Latent Diffusion Language Model Break the Sequential Dependency of LLM Inference Using Lookahead Decoding

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-11T20:11:10.714208Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-08T10:04:09.646578Z digest=sha256:5f7f529fb7485668b7deafde9c2e02f3ffa8e1e4978421b0fd581676585c38e5

Observation 0d58b655-de0b-4c97-8463-a009ac734da2 · inbound

SpecBlock: Block-Iterative Speculative Decoding with Dynamic Tree Drafting cites this paper.

SpecBlock: Block-Iterative Speculative Decoding with Dynamic Tree Drafting Break the Sequential Dependency of LLM Inference Using Lookahead Decoding

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-11T03:30:57.267208Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-11T02:27:45.880587Z digest=sha256:07cee32fdc7dd4d5884414abef3d22479d0417cd748bdba93bc6ac739972875d

Observation 66ef52a5-aef0-4a1c-a1f4-4a15bf6cdf36 · inbound

SpecBlock: Block-Iterative Speculative Decoding with Dynamic Tree Drafting cites this paper.

SpecBlock: Block-Iterative Speculative Decoding with Dynamic Tree Drafting Break the Sequential Dependency of LLM Inference Using Lookahead Decoding

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-22T10:51:26.023476Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-22T10:49:25.448580Z digest=sha256:f7814df139a84de45f71eaff6bfe074044487c920451d3f1e31a9842077dd5ff

Observation 92ef94dd-642f-4c83-9ab2-1ab4232b359b · inbound

CATS: Cascaded Adaptive Tree Speculation for Memory-Limited LLM Inference Acceleration cites this paper.

CATS: Cascaded Adaptive Tree Speculation for Memory-Limited LLM Inference Acceleration Break the Sequential Dependency of LLM Inference Using Lookahead Decoding

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-13T04:07:13.051959Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-13T04:06:24.409303Z digest=sha256:c54523f34ac5a42412ab0a11c33da23ace4d42378c08ff79043e519f8104f46b

Observation 59756922-f0fc-4822-b19c-ba12820ebacb · inbound

Performance-Driven Policy Optimization for Speculative Decoding with Adaptive Windowing cites this paper.

Performance-Driven Policy Optimization for Speculative Decoding with Adaptive Windowing Break the Sequential Dependency of LLM Inference Using Lookahead Decoding

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-19T16:42:39.741861Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-19T16:40:17.370100Z digest=sha256:d21e2ca22911166c540b6723f01e150f0845f15179da040115c16e0d1083091b

Observation b5927a08-cf39-4b7b-984f-aceebefb64a3 · inbound

An Interpretable Latency Model for Speculative Decoding in LLM Serving cites this paper.

An Interpretable Latency Model for Speculative Decoding in LLM Serving Break the Sequential Dependency of LLM Inference Using Lookahead Decoding

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-06-30T21:25:04.280613Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-30T21:23:39.721900Z digest=sha256:c822e4d76b00443219e0f6180e0e69727fe26a3d52784db366b424f35bc2a182

Observation 35c9858a-8f70-41ad-a659-b32f0896c7d2 · inbound

Draft Less, Retrieve More: Hybrid Tree Construction for Speculative Decoding cites this paper.

Draft Less, Retrieve More: Hybrid Tree Construction for Speculative Decoding Break the Sequential Dependency of LLM Inference Using Lookahead Decoding

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-20T07:03:06.430726Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-20T06:58:07.996335Z digest=sha256:ed3c83d7c88ea290466ec6848968ed81ffb1af5a0b3a447fd092425400045a14

Observation 4514b904-9626-4883-9502-53c081bc6c97 · inbound

Cassandra: Enabling Reasoning LLMs at Edge via Self-Speculative Decoding cites this paper.

Cassandra: Enabling Reasoning LLMs at Edge via Self-Speculative Decoding Break the Sequential Dependency of LLM Inference Using Lookahead Decoding

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-07-01T16:25:49.800642Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-01T16:19:54.910430Z digest=sha256:7e6bf58f8eb169c0fd7d54dd53bdc9192594eab545cfd468cd5f885e338d7f91

Observation a0f701da-1319-4f55-bcd5-2867ec4ac1b4 · inbound

Bastion: Budget-Aware Speculative Decoding with Tree-structured Block Diffusion Drafting cites this paper.

Bastion: Budget-Aware Speculative Decoding with Tree-structured Block Diffusion Drafting Break the Sequential Dependency of LLM Inference Using Lookahead Decoding

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-06-29T08:53:15.729175Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-29T08:51:11.662084Z digest=sha256:cd53f82530e6cc03d1d4fc2c4246678b06098170d3b4cecf2a3707452fe67b39

Observation 0fbb1dca-ea4b-4b0d-97a4-2c7dd80da296 · inbound

Streaming Communication in Multi-Agent Reasoning cites this paper.

Streaming Communication in Multi-Agent Reasoning Break the Sequential Dependency of LLM Inference Using Lookahead Decoding

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-07-02T08:06:47.912337Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-28T06:24:07.052856Z digest=sha256:8bd3bb6d435bf6d07935cbf8ec05728a6b4d91b6dd9a2e79bdd96456de100f49

Observation ec146b83-2ad8-44c3-8b30-57a8f444efe9 · inbound

Streaming Communication in Multi-Agent Reasoning cites this paper.

Streaming Communication in Multi-Agent Reasoning Break the Sequential Dependency of LLM Inference Using Lookahead Decoding

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-04T04:58:39.126792Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T04:58:39.126792Z digest=sha256:15098720d71224eba342921c0e0cf4606620977fafc8ad1dbb2be5f4a3f94cf8

Observation 699eda5f-81cf-411f-b685-706dc01c62be · inbound

CLP: Collocation-Length Prediction for Zero-Loss Adaptive Multi-Token Inference cites this paper.

CLP: Collocation-Length Prediction for Zero-Loss Adaptive Multi-Token Inference Break the Sequential Dependency of LLM Inference Using Lookahead Decoding

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-07-03T04:37:36.698018Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-27T13:49:06.097882Z digest=sha256:04fe16f4942b0b17878c099cbd1050332b37e67e38f8c880efb5e4902fadff92

Observation 99224362-d687-4528-8979-9d822902a361 · inbound

The Hitchhiker's Guide to Agentic AI: From Foundations to Systems cites this paper.

The Hitchhiker's Guide to Agentic AI: From Foundations to Systems Break the Sequential Dependency of LLM Inference Using Lookahead Decoding

Reference 157

Resolution
verified exact
arxiv_id, observed 2026-07-04T11:09:46.520199Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-26T08:09:57.542558Z digest=sha256:e50db38655675f3442c26e200187cf37a74014e6e97e0c8767f8ec9e103c71a3

Observation 4c1a0877-baac-4beb-b2f5-d869f41fdb0a · inbound

The Hitchhiker's Guide to Agentic AI: From Foundations to Systems cites this paper.

The Hitchhiker's Guide to Agentic AI: From Foundations to Systems Break the Sequential Dependency of LLM Inference Using Lookahead Decoding

Reference 157

Resolution
unresolved
no resolver link, observed 2026-08-02T10:27:18.314340Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T10:27:18.314340Z digest=sha256:5251f13ff6bab46d7f17f9a8642bf85a58ef5416ab0ddef40755802395b1d3d6

Observation cc8ff558-417d-4586-90b9-a81634a57f18 · inbound

Speculative Decoding at Temperature Zero: A Scoped Safety-Invariance Screen with a 48,072-Sample Expansion cites this paper.

Speculative Decoding at Temperature Zero: A Scoped Safety-Invariance Screen with a 48,072-Sample Expansion Break the Sequential Dependency of LLM Inference Using Lookahead Decoding

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-07-04T16:59:58.336242Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-26T00:05:06.647295Z digest=sha256:bc43560293983dfb86c773270f814febd7e3058eaa8283643bc5f9bfef6aef32

Observation 66173187-ef08-445e-adba-c83020c6cd3b · inbound

Efficient and Trainable Language Model Test-Time Scaling via Local Branch Routing cites this paper.

Efficient and Trainable Language Model Test-Time Scaling via Local Branch Routing Break the Sequential Dependency of LLM Inference Using Lookahead Decoding

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-07-04T19:20:06.721558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-25T21:29:40.564852Z digest=sha256:3b5e82b75f0709200a8284ccf5cb343d891e2e74f1fd6cc083d3295aaddd798d

Observation f76e1e01-6c36-4a4f-bbe0-e07acf9f7e7d · inbound

Efficient and Trainable Language Model Test-Time Scaling via Local Branch Routing cites this paper.

Efficient and Trainable Language Model Test-Time Scaling via Local Branch Routing Break the Sequential Dependency of LLM Inference Using Lookahead Decoding

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-07-01T06:35:29.632224Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-01T06:33:45.824727Z digest=sha256:a84decb335c2d0b32f827e4b2ecd9a9b2c5e77410162dd7746112e6cc929e22d

Observation 24e0f59b-fd67-45d2-9a86-e7bd98e77307 · inbound

The Context-Ready Transformer cites this paper.

The Context-Ready Transformer Break the Sequential Dependency of LLM Inference Using Lookahead Decoding

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-07-01T18:25:58.302117Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-06-29T01:59:43.093422Z digest=sha256:c1887ce65664ee5eb305ea24fdd82990133361ab29d449608dea079b524b5d81

Observation 9fd9a7c0-1059-4670-95b1-72526efb4668 · inbound

From Tensor Buffer to Distributed Memory Hierarchy: A Survey of KV Cache Management for LLM Serving cites this paper.

From Tensor Buffer to Distributed Memory Hierarchy: A Survey of KV Cache Management for LLM Serving Break the Sequential Dependency of LLM Inference Using Lookahead Decoding

Reference 33

Resolution
unresolved
no resolver link, observed 2026-07-12T09:50:23.266920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T09:50:23.266920Z digest=sha256:4d608642bcb52ae508f1fc0d97c31543a2f78a4eacecac27c4f4e430330c3a6b

Observation c117637a-c16a-402b-9c74-8ad8c0857f32 · inbound

WAR: Workload-Aware Rollouts for Synchronous Agentic Reinforcement Learning cites this paper.

WAR: Workload-Aware Rollouts for Synchronous Agentic Reinforcement Learning Break the Sequential Dependency of LLM Inference Using Lookahead Decoding

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-01T18:29:31.653276Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:29:31.653276Z digest=sha256:5c376a1afe9021642166539e90f38df74ecc749d00dd7cc84cd0d0a368641ae5

Observation cfb1070d-3e61-4d60-8d5d-1a763a6b9243 · inbound

KAP: Bridging the Knowledge Selection-Runtime Consumption Gap in LLM Systems cites this paper.

KAP: Bridging the Knowledge Selection-Runtime Consumption Gap in LLM Systems Break the Sequential Dependency of LLM Inference Using Lookahead Decoding

Reference 8

Resolution
unresolved
no resolver link, observed 2026-07-31T19:52:40.513351Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T19:52:40.513351Z digest=sha256:3feec2a4be4e2dc9dc87cf14bd6cb111143abbc79194f8a1914061e7bd4bfd1e

Observation f093adaf-d518-4b78-a85c-5c89bf3c73b0 · inbound

Beyond KV Reconstruction: Functional Reconstruction for MLA Draft Models in Speculative Decoding cites this paper.

Beyond KV Reconstruction: Functional Reconstruction for MLA Draft Models in Speculative Decoding Break the Sequential Dependency of LLM Inference Using Lookahead Decoding

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-01T10:55:32.040437Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T10:55:32.040437Z digest=sha256:2400f825ccee146490a08bf91c2c12c998f401043b4c2b49e2184af45c14a7d2

Observation d2753568-a074-4b6d-869b-10a38879b878 · inbound

ParVL: Parallel Scaling and Expandable Compute Allocation for Multimodal LLMs cites this paper.

ParVL: Parallel Scaling and Expandable Compute Allocation for Multimodal LLMs Break the Sequential Dependency of LLM Inference Using Lookahead Decoding

Reference 114

Resolution
unresolved
no resolver link, observed 2026-08-05T04:16:08.547924Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T04:16:08.547924Z digest=sha256:1d63fd4332aa654b3360a9a961139a3bba6d2dbf8c08a48f785adcf6a494d2f8

Observation 00e86942-c22d-4fd4-af08-19eccf7a23d2 · inbound

PaDoc: Layout-Grounded Parallel Decoding for Document Parsing cites this paper.

PaDoc: Layout-Grounded Parallel Decoding for Document Parsing Break the Sequential Dependency of LLM Inference Using Lookahead Decoding

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T14:06:38.812461Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:06:38.812461Z digest=sha256:98c757e9c98a7f113cf1f21a528005adf574effae70d4e5f1a6ec160fd00dd39