Pith. sign in

Paper Citation Record · LEDGER

Lossless Acceleration of Large Language Models with Hierarchical Drafting based on Temporal Locality in Speculative Decoding

As of 10 August 2026, this Paper Citation Record lists 47 of 47 outbound references and 2 inbound Pith citation observations for arXiv:2502.05609.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.05609 v1

Coverage vector

measured 47 of 47 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-08T18:42:15.171177Z

measured 49 of 49 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:14:21.520298Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-15T09:19:54.344822Z

Reference resolution

47 of 47 outbound references displayed

  • verified exact2
  • verified fuzzy2
  • unresolved42
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 6fcee68a-10e9-4b7e-8ece-4dabff19c560 · outbound

This paper cites Chandra, and Marc Snir.

Lossless Acceleration of Large Language Models with Hierarchical Drafting based on Temporal Locality in Speculative Decoding Chandra, and Marc Snir

Reference 1

Resolution
verified exact
raw_fallback, observed 2026-08-08T18:42:16.067589Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-08T18:42:14.933330Z digest=sha256:fad3bfdf91b18d8f782bbe6e59b414637cc32f2e055aa4155ea070d1ebc0328a

Observation 251c2da6-7c25-48db-a5ed-4acad6964e8a · outbound

This paper cites Hydra: Sequentially-Dependent Draft Heads for Medusa Decoding.

Lossless Acceleration of Large Language Models with Hierarchical Drafting based on Temporal Locality in Speculative Decoding Hydra: Sequentially-Dependent Draft Heads for Medusa Decoding

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-08T18:42:14.939397Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T18:42:14.939397Z digest=sha256:e63fecccbd07668d67971c6c45ca4367840c87353b763b0066e992c96f0ede0c

Observation d953af77-e298-46da-a533-f88ee0ab0a78 · outbound

This paper cites an unresolved cited work.

Lossless Acceleration of Large Language Models with Hierarchical Drafting based on Temporal Locality in Speculative Decoding Unresolved cited work

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-08T18:42:14.945269Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T18:42:14.945269Z digest=sha256:a87b69e3a22eaa592ef5d60184db6187092359e7dd5e8ef6e39ad7b821f3d02d

Observation 09d5c4a8-a42c-4640-a985-0af2a6fc60ed · outbound

This paper cites an unresolved cited work.

Lossless Acceleration of Large Language Models with Hierarchical Drafting based on Temporal Locality in Speculative Decoding Unresolved cited work

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-08T18:42:14.950561Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T18:42:14.950561Z digest=sha256:826b3ec428d5f422779a2357458d9690fc6a2d7180fd4d23253920a5ed77913d

Observation 60cc7ea2-2e09-45f6-8899-1afbe06e216a · outbound

This paper cites an unresolved cited work.

Lossless Acceleration of Large Language Models with Hierarchical Drafting based on Temporal Locality in Speculative Decoding Unresolved cited work

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-08T18:42:14.955430Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T18:42:14.955430Z digest=sha256:ded9b46391f2889a62c7e8a3481f226c9d9898e8e4e4970a9b283aff22fecbfb

Observation baafcb4f-0178-4261-a8ec-d64720030fdf · outbound

This paper cites Lee, Deming Chen, and Tri Dao.

Lossless Acceleration of Large Language Models with Hierarchical Drafting based on Temporal Locality in Speculative Decoding Lee, Deming Chen, and Tri Dao

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T18:42:16.331099Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-08T18:42:14.960598Z digest=sha256:7e6830c1c72e650de7409743c261cb26106d0225f6d986dd7752475f46fee7d0

Observation c76597ee-6577-4fa9-9c03-372268744932 · outbound

This paper cites Accelerating Large Language Model Decoding with Speculative Sampling.

Lossless Acceleration of Large Language Models with Hierarchical Drafting based on Temporal Locality in Speculative Decoding Accelerating Large Language Model Decoding with Speculative Sampling

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-08T18:42:14.971410Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T18:42:14.971410Z digest=sha256:f0eb878c414ceeab2eb93e5294f2c128f275fefe8efd63c63ccbfb4f7370a91d

Observation 670823e7-eebd-484e-9071-4045e9c4792f · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Lossless Acceleration of Large Language Models with Hierarchical Drafting based on Temporal Locality in Speculative Decoding Training Verifiers to Solve Math Word Problems

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-08T18:42:14.976169Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T18:42:14.976169Z digest=sha256:4bb371a7d41cf6a907ea41eaeded6d4d534ea0e4196130636f90342f4106a97c

Observation 14c3f965-d832-4d0c-96c3-deb3d8fd0645 · outbound

This paper cites an unresolved cited work.

Lossless Acceleration of Large Language Models with Hierarchical Drafting based on Temporal Locality in Speculative Decoding Unresolved cited work

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-08T18:42:14.981480Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T18:42:14.981480Z digest=sha256:b636c770b1ff5420a6a8223eb290eb3569b0a206831fed02ef7335c798fac2a9

Observation e4978392-a1b8-466b-bacc-d8d26494bdf4 · outbound

This paper cites an unresolved cited work.

Lossless Acceleration of Large Language Models with Hierarchical Drafting based on Temporal Locality in Speculative Decoding Unresolved cited work

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-08T18:42:14.986619Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T18:42:14.986619Z digest=sha256:8175c5edf2abcd52dd51d2e9dce9b1852cfcec248c9981c23a6df36d63bafc97

Observation 8cf1dc37-df03-459b-bf5f-aabbd06a6305 · outbound

This paper cites an unresolved cited work.

Lossless Acceleration of Large Language Models with Hierarchical Drafting based on Temporal Locality in Speculative Decoding Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-08-08T18:42:16.303475Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-08T18:42:14.991341Z digest=sha256:a49767bafc593b46de348a3fcc7f3c342341d2fdf088b31fc17c1d9160b4dbfe

Observation 954fa007-48cb-48de-bfec-04714bfc71da · outbound

This paper cites Mahoney, and Kurt Keutzer.

Lossless Acceleration of Large Language Models with Hierarchical Drafting based on Temporal Locality in Speculative Decoding Mahoney, and Kurt Keutzer

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-08T18:42:14.996032Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T18:42:14.996032Z digest=sha256:e9721d32d3dc384ec66b0aba12ae99089aee48f889cd54d3c7ad79de9c6ac94b

Observation 90510e72-7108-4df1-9a38-02a9a0930e78 · outbound

This paper cites an unresolved cited work.

Lossless Acceleration of Large Language Models with Hierarchical Drafting based on Temporal Locality in Speculative Decoding Unresolved cited work

Reference 14

Resolution
unresolved
raw_fallback, observed 2026-08-08T18:42:16.287816Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-08T18:42:15.000509Z digest=sha256:b1da1503e46d87cec2746d04630aa5aa62910fe8542f4f3aef90f341125d264f

Observation a7480ac7-6fd0-45ab-8a37-8284513cd686 · outbound

This paper cites an unresolved cited work.

Lossless Acceleration of Large Language Models with Hierarchical Drafting based on Temporal Locality in Speculative Decoding Unresolved cited work

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-08T18:42:15.005433Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T18:42:15.005433Z digest=sha256:c5e2d62870c82cc082a977bead66199ff969d6b000298eb0adca5fa575a4797c

Observation 82638e70-4319-4cad-85ea-af7a65b3dd55 · outbound

This paper cites SqueezeLLM: Dense-and-Sparse Quantization.

Lossless Acceleration of Large Language Models with Hierarchical Drafting based on Temporal Locality in Speculative Decoding SqueezeLLM: Dense-and-Sparse Quantization

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-08T18:42:15.010420Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T18:42:15.010420Z digest=sha256:87d9ce452767dfc2a23dcddbabeef63db6011bec50891213a171bcabc6f19153

Observation 0c8a4623-4745-413e-b7a0-dfcf9be94b3a · outbound

This paper cites an unresolved cited work.

Lossless Acceleration of Large Language Models with Hierarchical Drafting based on Temporal Locality in Speculative Decoding Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-08-08T18:42:16.272293Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-08T18:42:15.015668Z digest=sha256:5c80dd22b5cfe261c3dad9703699be7a10452de46cd7b080b9f57b5acd908384

Observation cf31d636-0f36-4330-af52-74241fc7a3e4 · outbound

This paper cites o pf, Yannic Kilcher, Dimitri von R \.

Lossless Acceleration of Large Language Models with Hierarchical Drafting based on Temporal Locality in Speculative Decoding o pf, Yannic Kilcher, Dimitri von R \

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T18:42:16.256804Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-08T18:42:15.020330Z digest=sha256:16c34528de7f3c1cdaa8b322535065173f416401b62b21a6a4ac510faa896cc2

Observation 8b955bfc-3e22-4409-b480-66fe5f6e8881 · outbound

This paper cites Parikh, Chris Alberti, Danielle Epstein, Illia Polosukhin, Jacob Devlin, Kenton Lee, Kristina Toutanova, Llion Jones, Matthew Kelcey, Ming - Wei Chang, Andrew M.

Lossless Acceleration of Large Language Models with Hierarchical Drafting based on Temporal Locality in Speculative Decoding Parikh, Chris Alberti, Danielle Epstein, Illia Polosukhin, Jacob Devlin, Kenton Lee, Kristina Toutanova, Llion Jones, Matthew Kelcey, Ming - Wei Chang, Andrew M

Reference 19

Resolution
malformed identifier
no resolver link, observed 2026-08-08T18:42:15.025693Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T18:42:15.025693Z digest=sha256:a52ce32d65670e398d89fe9f2bd9f6f2eb4402586ef5cdead6946490fb12b8bc

Observation 9d4ed3c2-8be4-47ba-b86a-2dee408a07e0 · outbound

This paper cites an unresolved cited work.

Lossless Acceleration of Large Language Models with Hierarchical Drafting based on Temporal Locality in Speculative Decoding Unresolved cited work

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-08T18:42:15.030651Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T18:42:15.030651Z digest=sha256:5f1f64da37045a2af30df4f1bbd9c10c167c1ef5d2a31e288f6d2d88119ed2a6

Observation fca670cc-9c10-4e89-8a60-9d2a83ea4154 · outbound

This paper cites an unresolved cited work.

Lossless Acceleration of Large Language Models with Hierarchical Drafting based on Temporal Locality in Speculative Decoding Unresolved cited work

Reference 21

Resolution
unresolved
raw_fallback, observed 2026-08-08T18:42:16.241440Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-08T18:42:15.035477Z digest=sha256:4cbc7efde84eb0b0a23add4c5a6eca8269c6ba6cc636e76cbd5d074d671354bf

Observation d7bdff32-cf04-466d-8f82-aa628ca87153 · outbound

This paper cites EAGLE-2: Faster Inference of Language Models with Dynamic Draft Trees.

Lossless Acceleration of Large Language Models with Hierarchical Drafting based on Temporal Locality in Speculative Decoding EAGLE-2: Faster Inference of Language Models with Dynamic Draft Trees

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-08T18:42:15.040214Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T18:42:15.040214Z digest=sha256:58483384863a2fece03387a98308cda7a58f472f43a6669742a635315173ace7

Observation ca73317a-63e4-477b-a85e-bff1a4e6daed · outbound

This paper cites an unresolved cited work.

Lossless Acceleration of Large Language Models with Hierarchical Drafting based on Temporal Locality in Speculative Decoding Unresolved cited work

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-08T18:42:15.045181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T18:42:15.045181Z digest=sha256:d0311a020bb3af11ea3ae29c0df175d3439f1cc6a5501e5c787930dedfba5b99

Observation 4bdb2ed7-8c12-4188-954e-406a4bf8c2b0 · outbound

This paper cites Turning Trash into Treasure: Accelerating Inference of Large Language Models with Token Recycling.

Lossless Acceleration of Large Language Models with Hierarchical Drafting based on Temporal Locality in Speculative Decoding Turning Trash into Treasure: Accelerating Inference of Large Language Models with Token Recycling

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-08T18:42:15.049947Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T18:42:15.049947Z digest=sha256:959193bca803a5796f381a099072dc6974cec965eedb2e12c263f507c6e8df64

Observation dcf80767-3a84-4ffa-9d87-fb86454e1142 · outbound

This paper cites an unresolved cited work.

Lossless Acceleration of Large Language Models with Hierarchical Drafting based on Temporal Locality in Speculative Decoding Unresolved cited work

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-08T18:42:15.055141Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T18:42:15.055141Z digest=sha256:ebb898157f23505e8b935fe7425002b781d455e71ddcacfa6cf3cb5097355748

Observation ab8bb7e3-eff3-4515-bc4e-1f0c07082ad7 · outbound

This paper cites an unresolved cited work.

Lossless Acceleration of Large Language Models with Hierarchical Drafting based on Temporal Locality in Speculative Decoding Unresolved cited work

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-08T18:42:15.060135Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T18:42:15.060135Z digest=sha256:648a25ec76c53735605c6ff2ebbc1a09810e45697437746d9510d0f756d3ebdd

Observation ac4b6173-0408-4188-ac64-aa22f44d662f · outbound

This paper cites an unresolved cited work.

Lossless Acceleration of Large Language Models with Hierarchical Drafting based on Temporal Locality in Speculative Decoding Unresolved cited work

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-08T18:42:15.067987Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T18:42:15.067987Z digest=sha256:383338f6855caacc16517e3f97a59a2e2badbfe9d42f8176d90fe173f793f50b

Observation 475681fe-c2b1-4353-9692-9fb55cde2687 · outbound

This paper cites Patterson.

Lossless Acceleration of Large Language Models with Hierarchical Drafting based on Temporal Locality in Speculative Decoding Patterson

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-08T18:42:15.074195Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T18:42:15.074195Z digest=sha256:e9d1630e82f8e4048c59b5e89afa683ad27bfc47587810c5b13eae18b3dad0e7

Observation 04ad2649-815c-4d96-b615-01a27522e5da · outbound

This paper cites an unresolved cited work.

Lossless Acceleration of Large Language Models with Hierarchical Drafting based on Temporal Locality in Speculative Decoding Unresolved cited work

Reference 29

Resolution
unresolved
raw_fallback, observed 2026-08-08T18:42:16.215431Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-08T18:42:15.078918Z digest=sha256:8eec55fd2e3fedf628992c19fe55fd6664a73a1eabae7a3b13203b2d5fe7b017

Observation 750f93c6-96bb-4af7-90ee-807665b9c908 · outbound

This paper cites an unresolved cited work.

Lossless Acceleration of Large Language Models with Hierarchical Drafting based on Temporal Locality in Speculative Decoding Unresolved cited work

Reference 30

Resolution
unresolved
raw_fallback, observed 2026-08-08T18:42:16.199635Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-08T18:42:15.083484Z digest=sha256:24c0fa47fc7df704a1cf8ef8222d9225c065ef4facb1e328c3eee3f0b079def3

Observation 19db47fb-a138-44bf-902b-d0bd7f42a06a · outbound

This paper cites an unresolved cited work.

Lossless Acceleration of Large Language Models with Hierarchical Drafting based on Temporal Locality in Speculative Decoding Unresolved cited work

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-08T18:42:15.087872Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T18:42:15.087872Z digest=sha256:03bc6d4d3adcf168bb016ac8ba39737bdb2aa99882c9224288202189d9041691

Observation a7dfa254-2731-40c5-9f29-43e033f19d55 · outbound

This paper cites an unresolved cited work.

Lossless Acceleration of Large Language Models with Hierarchical Drafting based on Temporal Locality in Speculative Decoding Unresolved cited work

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-08T18:42:15.092516Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T18:42:15.092516Z digest=sha256:d2d2d705b14b25bd5856e8be02761e82f9ebadab20111eebd73c228ada441c48

Observation 7227293a-b44b-4ecc-ba54-3809cf6d4b2e · outbound

This paper cites Fast Transformer Decoding: One Write-Head is All You Need.

Lossless Acceleration of Large Language Models with Hierarchical Drafting based on Temporal Locality in Speculative Decoding Fast Transformer Decoding: One Write-Head is All You Need

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-08T18:42:15.096988Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T18:42:15.096988Z digest=sha256:c66165f12e74a55f02756f6f13b2c7040c93b1f98278b08ee8c4b68a15f5fa83

Observation b67b68e3-2ae1-4a48-8e95-5edaf4344bc0 · outbound

This paper cites an unresolved cited work.

Lossless Acceleration of Large Language Models with Hierarchical Drafting based on Temporal Locality in Speculative Decoding Unresolved cited work

Reference 34

Resolution
unresolved
raw_fallback, observed 2026-08-08T18:42:16.173396Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-08T18:42:15.102132Z digest=sha256:2420757561183abde614f7ae05ae752b9be45beb56d1ccc55840f2277e0fbe2b

Observation bb22f7e5-b7cf-4918-b335-5a5512537478 · outbound

This paper cites Accelerating LLM Inference with Staged Speculative Decoding.

Lossless Acceleration of Large Language Models with Hierarchical Drafting based on Temporal Locality in Speculative Decoding Accelerating LLM Inference with Staged Speculative Decoding

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-08T18:42:15.107139Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T18:42:15.107139Z digest=sha256:a856ca472530d7170cb9d7d855d92a41084fb5d406c073b67f1995d44734ee97

Observation af83926c-f22c-4a20-a3d2-11d54cd64e2f · outbound

This paper cites an unresolved cited work.

Lossless Acceleration of Large Language Models with Hierarchical Drafting based on Temporal Locality in Speculative Decoding Unresolved cited work

Reference 36

Resolution
unresolved
raw_fallback, observed 2026-08-08T18:42:16.157729Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-08T18:42:15.112127Z digest=sha256:f702b439d52b823e79bcd3984451e93dd72dca68d98f97df5a2664bef21717bb

Observation 9556c128-31dc-43f0-b39a-2902ce0b6315 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Lossless Acceleration of Large Language Models with Hierarchical Drafting based on Temporal Locality in Speculative Decoding Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-08T18:42:15.116726Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T18:42:15.116726Z digest=sha256:f8f551e30c37c2aef984152d1bffba1e4baa20755c42691fb2d3557905f2a394

Observation 9916496e-31e5-436a-84ad-1e1024ee5dcc · outbound

This paper cites an unresolved cited work.

Lossless Acceleration of Large Language Models with Hierarchical Drafting based on Temporal Locality in Speculative Decoding Unresolved cited work

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-08T18:42:15.121906Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T18:42:15.121906Z digest=sha256:051e7d0b7027b4ef9353ad5a275b1c5ef3344783fab1b7b1b09c1f9a03b0811d

Observation ed87d95a-db03-4aeb-8224-43ca8e6a7508 · outbound

This paper cites an unresolved cited work.

Lossless Acceleration of Large Language Models with Hierarchical Drafting based on Temporal Locality in Speculative Decoding Unresolved cited work

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-08T18:42:15.127040Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T18:42:15.127040Z digest=sha256:8f94416706712a4b5b981cb4b2e96350376391c6fd8c9003d87ea6733c74ed8a

Observation 9a8efe3b-b0eb-4639-9683-81eae12dbf4c · outbound

This paper cites Inference with Reference: Lossless Acceleration of Large Language Models.

Lossless Acceleration of Large Language Models with Hierarchical Drafting based on Temporal Locality in Speculative Decoding Inference with Reference: Lossless Acceleration of Large Language Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-08T18:42:15.131740Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T18:42:15.131740Z digest=sha256:5f18de82afa58051b7d3bbf22cfafc626e722966b5d3937aee6161119485a3b3

Observation 32ca7678-766f-4aa7-a3f1-79a0a1eae9de · outbound

This paper cites an unresolved cited work.

Lossless Acceleration of Large Language Models with Hierarchical Drafting based on Temporal Locality in Speculative Decoding Unresolved cited work

Reference 41

Resolution
unresolved
raw_fallback, observed 2026-08-08T18:42:16.141465Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-08T18:42:15.136772Z digest=sha256:0c414ca4675c17197b2673cbcb7f67436942bdcf6e3bc24fd59f382ac6d53e66

Observation cb0dfb7d-c82f-4609-a07d-062b762e778b · outbound

This paper cites Towards Fast Multilingual LLM Inference: Speculative Decoding and Specialized Drafters.

Lossless Acceleration of Large Language Models with Hierarchical Drafting based on Temporal Locality in Speculative Decoding Towards Fast Multilingual LLM Inference: Speculative Decoding and Specialized Drafters

Reference 42

Resolution
verified exact
local_arxiv, observed 2026-08-08T18:42:15.237786Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-08T18:42:15.141623Z digest=sha256:02dad95a143958d1a16b7e769eaefc103e5d3a69134d5d3477103578dae162a5

Observation 35c70148-7de9-450d-9a49-60a1986827d5 · outbound

This paper cites TinyLlama: An Open-Source Small Language Model.

Lossless Acceleration of Large Language Models with Hierarchical Drafting based on Temporal Locality in Speculative Decoding TinyLlama: An Open-Source Small Language Model

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-08T18:42:15.146922Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T18:42:15.146922Z digest=sha256:81d143e4ad399563d50aa5affc009384a3899880dd7454ad221f7798ad703f43

Observation 69f869ef-6115-4a00-b32a-d38f0bca5285 · outbound

This paper cites Barrett, Zhangyang Wang, and Beidi Chen.

Lossless Acceleration of Large Language Models with Hierarchical Drafting based on Temporal Locality in Speculative Decoding Barrett, Zhangyang Wang, and Beidi Chen

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-08T18:42:15.151791Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T18:42:15.151791Z digest=sha256:096f68938ff6a91ba604a1f082c8ef53ca2c6ebf843457b29150cc515614e3cb

Observation c74c0f5c-fb70-4011-ad59-a031f2182df3 · outbound

This paper cites Xing, Hao Zhang, Joseph E.

Lossless Acceleration of Large Language Models with Hierarchical Drafting based on Temporal Locality in Speculative Decoding Xing, Hao Zhang, Joseph E

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-08T18:42:15.156291Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T18:42:15.156291Z digest=sha256:58d5d47ffd5c2ab54c80f6bba6eeeb32eea22e3b289a625818988a3f10c86164

Observation e59a4ecd-2b8d-44cb-839c-1ad196c2bdbf · outbound

This paper cites an unresolved cited work.

Lossless Acceleration of Large Language Models with Hierarchical Drafting based on Temporal Locality in Speculative Decoding Unresolved cited work

Reference 46

Resolution
unresolved
raw_fallback, observed 2026-08-08T18:42:16.104927Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-08T18:42:15.161034Z digest=sha256:67b8e2477df64eeae7e3e66461b4e798ec71bbf7440b33941a86f1d9e102d7ba

Observation 65223a9c-9ede-4605-b468-21b56f5ba059 · outbound

This paper cites URL: " 'urlintro :=.

Lossless Acceleration of Large Language Models with Hierarchical Drafting based on Temporal Locality in Speculative Decoding URL: " 'urlintro :=

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-08T18:42:15.165690Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T18:42:15.165690Z digest=sha256:5dfd565a7becd3085d7f2158eead0d188d69a7f32237259af2059e9bf3bfa666

Observation 69fb4074-f375-4de7-8ed6-14b61166c723 · outbound

This paper cites write newline.

Lossless Acceleration of Large Language Models with Hierarchical Drafting based on Temporal Locality in Speculative Decoding write newline

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-08T18:42:15.171177Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T18:42:15.171177Z digest=sha256:1b08509621e969af06b684bceb825daf829343afcb95cc2acac6c53647df2ee3

Pith citing papers

Observation 512295ff-e0ef-409b-94d9-2f00cf25795f · inbound

Faster and Better LLMs via Latency-Aware Test-Time Scaling cites this paper.

Faster and Better LLMs via Latency-Aware Test-Time Scaling Lossless Acceleration of Large Language Models with Hierarchical Drafting based on Temporal Locality in Speculative Decoding

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:21.520298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:14:21.520298Z digest=sha256:c2e862b4a679ab97d6f942a506c73888e2a32a75e3645779570c0a2117271d38

Observation 61a54cbb-ae6f-49ce-8249-a1a734c7471d · inbound

HeiSD: Hybrid Speculative Decoding for Embodied Vision-Language-Action Models with Kinematic Awareness cites this paper.

HeiSD: Hybrid Speculative Decoding for Embodied Vision-Language-Action Models with Kinematic Awareness Lossless Acceleration of Large Language Models with Hierarchical Drafting based on Temporal Locality in Speculative Decoding

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-15T09:19:54.348149Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-15T09:15:50.963123Z digest=sha256:0230d7d325b95bb64f1be72c5f60876e1485930212b53f3c267a388882837a1b