Pith. sign in

Paper Citation Record · LEDGER

Lossless Acceleration of Large Language Models with Hierarchical Drafting based on Temporal Locality in Speculative Decoding

As of 8 August 2026, this Paper Citation Record lists 47 of 47 outbound references and 2 inbound Pith citation observations for arXiv:2502.05609.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.05609 v1

Coverage vector

measured 47 of 47 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-08T18:42:15.171177Z

measured 49 of 49 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:14:21.520298Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-15T09:19:54.344822Z

Reference resolution

47 of 47 outbound references displayed

  • verified exact2
  • verified fuzzy2
  • unresolved42
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 6fcee68a-10e9-4b7e-8ece-4dabff19c560 · outbound

This paper cites Chandra, and Marc Snir.

Lossless Acceleration of Large Language Models with Hierarchical Drafting based on Temporal Locality in Speculative Decoding Chandra, and Marc Snir

Reference 1

Resolution
verified exact
raw_fallback, observed 2026-08-08T18:42:16.067589Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-08T18:42:14.933330Z digest=sha256:a32d1ada6b1095ab7815ffb8b2eac7ebff9fe2e88d9691574bf6b01074466fea

Observation 251c2da6-7c25-48db-a5ed-4acad6964e8a · outbound

This paper cites Hydra: Sequentially-Dependent Draft Heads for Medusa Decoding.

Lossless Acceleration of Large Language Models with Hierarchical Drafting based on Temporal Locality in Speculative Decoding Hydra: Sequentially-Dependent Draft Heads for Medusa Decoding

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-08T18:42:14.939397Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T18:42:14.939397Z digest=sha256:1fd362a99d596781635f2d5cbe832b70d98a8d295f7c3368472eec218284d3d5

Observation d953af77-e298-46da-a533-f88ee0ab0a78 · outbound

This paper cites an unresolved cited work.

Lossless Acceleration of Large Language Models with Hierarchical Drafting based on Temporal Locality in Speculative Decoding Unresolved cited work

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-08T18:42:14.945269Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T18:42:14.945269Z digest=sha256:3256224878bd8a525d26578951188082cd8b10f25ea6f01e77d5925e8b329403

Observation 09d5c4a8-a42c-4640-a985-0af2a6fc60ed · outbound

This paper cites an unresolved cited work.

Lossless Acceleration of Large Language Models with Hierarchical Drafting based on Temporal Locality in Speculative Decoding Unresolved cited work

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-08T18:42:14.950561Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T18:42:14.950561Z digest=sha256:2656b711f10a99c1875d19cb4692242886067968fa30aee8167779aebc192286

Observation 60cc7ea2-2e09-45f6-8899-1afbe06e216a · outbound

This paper cites an unresolved cited work.

Lossless Acceleration of Large Language Models with Hierarchical Drafting based on Temporal Locality in Speculative Decoding Unresolved cited work

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-08T18:42:14.955430Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T18:42:14.955430Z digest=sha256:93d849d4b56bbf5f6f6d2f9158026249d63bbdc96136bfe768255baa61b3eaed

Observation baafcb4f-0178-4261-a8ec-d64720030fdf · outbound

This paper cites Lee, Deming Chen, and Tri Dao.

Lossless Acceleration of Large Language Models with Hierarchical Drafting based on Temporal Locality in Speculative Decoding Lee, Deming Chen, and Tri Dao

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T18:42:16.331099Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-08T18:42:14.960598Z digest=sha256:a022bf63c721e90ccddc456131a62b2b4c75930145916412a9153e9def4276ff

Observation c76597ee-6577-4fa9-9c03-372268744932 · outbound

This paper cites Accelerating Large Language Model Decoding with Speculative Sampling.

Lossless Acceleration of Large Language Models with Hierarchical Drafting based on Temporal Locality in Speculative Decoding Accelerating Large Language Model Decoding with Speculative Sampling

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-08T18:42:14.971410Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T18:42:14.971410Z digest=sha256:3f8e36759dacb6fee6dded8acec57a99458b57ce11d22dfc9695e9bae97da203

Observation 670823e7-eebd-484e-9071-4045e9c4792f · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Lossless Acceleration of Large Language Models with Hierarchical Drafting based on Temporal Locality in Speculative Decoding Training Verifiers to Solve Math Word Problems

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-08T18:42:14.976169Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T18:42:14.976169Z digest=sha256:000e3fd619369dba7617d866826f0c03b1190772daf12d051e3e03def66f8c7d

Observation 14c3f965-d832-4d0c-96c3-deb3d8fd0645 · outbound

This paper cites an unresolved cited work.

Lossless Acceleration of Large Language Models with Hierarchical Drafting based on Temporal Locality in Speculative Decoding Unresolved cited work

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-08T18:42:14.981480Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T18:42:14.981480Z digest=sha256:902b16e526579a86385cc86c673329992e3ce6d369d1308e89f270474ec5d517

Observation e4978392-a1b8-466b-bacc-d8d26494bdf4 · outbound

This paper cites an unresolved cited work.

Lossless Acceleration of Large Language Models with Hierarchical Drafting based on Temporal Locality in Speculative Decoding Unresolved cited work

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-08T18:42:14.986619Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T18:42:14.986619Z digest=sha256:d13ac15faa771d429a485f10d25c8b630b8e949bb7f55da0b00dbf6d4c06df6f

Observation 8cf1dc37-df03-459b-bf5f-aabbd06a6305 · outbound

This paper cites an unresolved cited work.

Lossless Acceleration of Large Language Models with Hierarchical Drafting based on Temporal Locality in Speculative Decoding Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-08-08T18:42:16.303475Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-08T18:42:14.991341Z digest=sha256:8a4ee0338f18f13d8c325d5705337148228faaf9924c9845a7e21ae473671033

Observation 954fa007-48cb-48de-bfec-04714bfc71da · outbound

This paper cites Mahoney, and Kurt Keutzer.

Lossless Acceleration of Large Language Models with Hierarchical Drafting based on Temporal Locality in Speculative Decoding Mahoney, and Kurt Keutzer

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-08T18:42:14.996032Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T18:42:14.996032Z digest=sha256:3cdf1a73ddc6336fd29f5f81c8cbf27c2d4e57b3eed00580161803bb83d22932

Observation 90510e72-7108-4df1-9a38-02a9a0930e78 · outbound

This paper cites an unresolved cited work.

Lossless Acceleration of Large Language Models with Hierarchical Drafting based on Temporal Locality in Speculative Decoding Unresolved cited work

Reference 14

Resolution
unresolved
raw_fallback, observed 2026-08-08T18:42:16.287816Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-08T18:42:15.000509Z digest=sha256:70884dfec63d5bb7f810124ac9b37e5d56a956b7cdb256820c1a497974ee3ad5

Observation a7480ac7-6fd0-45ab-8a37-8284513cd686 · outbound

This paper cites an unresolved cited work.

Lossless Acceleration of Large Language Models with Hierarchical Drafting based on Temporal Locality in Speculative Decoding Unresolved cited work

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-08T18:42:15.005433Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T18:42:15.005433Z digest=sha256:ddce834912414045966cbfb7f64efe543854f62ce31131c21c2d4a4f8e4684fc

Observation 82638e70-4319-4cad-85ea-af7a65b3dd55 · outbound

This paper cites SqueezeLLM: Dense-and-Sparse Quantization.

Lossless Acceleration of Large Language Models with Hierarchical Drafting based on Temporal Locality in Speculative Decoding SqueezeLLM: Dense-and-Sparse Quantization

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-08T18:42:15.010420Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T18:42:15.010420Z digest=sha256:edf6651320f5115b5bf0b44e2d6049d1d7319fd9aaab2dbb86c69c185db471b6

Observation 0c8a4623-4745-413e-b7a0-dfcf9be94b3a · outbound

This paper cites an unresolved cited work.

Lossless Acceleration of Large Language Models with Hierarchical Drafting based on Temporal Locality in Speculative Decoding Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-08-08T18:42:16.272293Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-08T18:42:15.015668Z digest=sha256:2fd10efc8702b7d83febd2bd8f113cb6ae1404640b94b6bf4e28ead9a9b45000

Observation cf31d636-0f36-4330-af52-74241fc7a3e4 · outbound

This paper cites o pf, Yannic Kilcher, Dimitri von R \.

Lossless Acceleration of Large Language Models with Hierarchical Drafting based on Temporal Locality in Speculative Decoding o pf, Yannic Kilcher, Dimitri von R \

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T18:42:16.256804Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-08T18:42:15.020330Z digest=sha256:f84ac962eed7d06222219862db48786d6d9423938dbec9491cbc3756634aba35

Observation 8b955bfc-3e22-4409-b480-66fe5f6e8881 · outbound

This paper cites Parikh, Chris Alberti, Danielle Epstein, Illia Polosukhin, Jacob Devlin, Kenton Lee, Kristina Toutanova, Llion Jones, Matthew Kelcey, Ming - Wei Chang, Andrew M.

Lossless Acceleration of Large Language Models with Hierarchical Drafting based on Temporal Locality in Speculative Decoding Parikh, Chris Alberti, Danielle Epstein, Illia Polosukhin, Jacob Devlin, Kenton Lee, Kristina Toutanova, Llion Jones, Matthew Kelcey, Ming - Wei Chang, Andrew M

Reference 19

Resolution
malformed identifier
no resolver link, observed 2026-08-08T18:42:15.025693Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T18:42:15.025693Z digest=sha256:9de84e499bd0c523364958764abd027dede34a0a6f7b75970676776a13c671bf

Observation 9d4ed3c2-8be4-47ba-b86a-2dee408a07e0 · outbound

This paper cites an unresolved cited work.

Lossless Acceleration of Large Language Models with Hierarchical Drafting based on Temporal Locality in Speculative Decoding Unresolved cited work

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-08T18:42:15.030651Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T18:42:15.030651Z digest=sha256:9f0ffeb5ddd44f1cfade188ef65e8f0e9c0d2d76c7c1c3e86592ee9916e1075b

Observation fca670cc-9c10-4e89-8a60-9d2a83ea4154 · outbound

This paper cites an unresolved cited work.

Lossless Acceleration of Large Language Models with Hierarchical Drafting based on Temporal Locality in Speculative Decoding Unresolved cited work

Reference 21

Resolution
unresolved
raw_fallback, observed 2026-08-08T18:42:16.241440Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-08T18:42:15.035477Z digest=sha256:0d6ccc6ea7377c8e9b5848c2ee64a20f051e8727d1181d07b2038866e1905745

Observation d7bdff32-cf04-466d-8f82-aa628ca87153 · outbound

This paper cites EAGLE-2: Faster Inference of Language Models with Dynamic Draft Trees.

Lossless Acceleration of Large Language Models with Hierarchical Drafting based on Temporal Locality in Speculative Decoding EAGLE-2: Faster Inference of Language Models with Dynamic Draft Trees

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-08T18:42:15.040214Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T18:42:15.040214Z digest=sha256:4ca928c20e12558e8dee0ae3f4366772830fa1069158430872cb3413b8e09928

Observation ca73317a-63e4-477b-a85e-bff1a4e6daed · outbound

This paper cites an unresolved cited work.

Lossless Acceleration of Large Language Models with Hierarchical Drafting based on Temporal Locality in Speculative Decoding Unresolved cited work

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-08T18:42:15.045181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T18:42:15.045181Z digest=sha256:461943c069623a8ca6a63a49d96ec2555fe6a5818d522264b72fea2fe57ea920

Observation 4bdb2ed7-8c12-4188-954e-406a4bf8c2b0 · outbound

This paper cites Turning Trash into Treasure: Accelerating Inference of Large Language Models with Token Recycling.

Lossless Acceleration of Large Language Models with Hierarchical Drafting based on Temporal Locality in Speculative Decoding Turning Trash into Treasure: Accelerating Inference of Large Language Models with Token Recycling

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-08T18:42:15.049947Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T18:42:15.049947Z digest=sha256:ed0bf8b44c4c64f1c3a438ec2dc287fe14ab8ac42e4b04d345a334cf8b62e0d9

Observation dcf80767-3a84-4ffa-9d87-fb86454e1142 · outbound

This paper cites an unresolved cited work.

Lossless Acceleration of Large Language Models with Hierarchical Drafting based on Temporal Locality in Speculative Decoding Unresolved cited work

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-08T18:42:15.055141Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T18:42:15.055141Z digest=sha256:e95121335e8cb11ddd38a2c4104064006b499330f3eef82fe074d5067ab28b4a

Observation ab8bb7e3-eff3-4515-bc4e-1f0c07082ad7 · outbound

This paper cites an unresolved cited work.

Lossless Acceleration of Large Language Models with Hierarchical Drafting based on Temporal Locality in Speculative Decoding Unresolved cited work

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-08T18:42:15.060135Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T18:42:15.060135Z digest=sha256:a623cfdf1d660c34890e0bb36ad50a1086d4a820615c42bf9706ac3716969e4b

Observation ac4b6173-0408-4188-ac64-aa22f44d662f · outbound

This paper cites an unresolved cited work.

Lossless Acceleration of Large Language Models with Hierarchical Drafting based on Temporal Locality in Speculative Decoding Unresolved cited work

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-08T18:42:15.067987Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T18:42:15.067987Z digest=sha256:c8ec731e72949aa8b9049498429435efb50c0a3caab366bbc9f4f3a10a79d9c1

Observation 475681fe-c2b1-4353-9692-9fb55cde2687 · outbound

This paper cites Patterson.

Lossless Acceleration of Large Language Models with Hierarchical Drafting based on Temporal Locality in Speculative Decoding Patterson

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-08T18:42:15.074195Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T18:42:15.074195Z digest=sha256:84dcbea7df91ee0d9c3204a5a6b1c004f39626c76d311340cd90ae723eca2ce5

Observation 04ad2649-815c-4d96-b615-01a27522e5da · outbound

This paper cites an unresolved cited work.

Lossless Acceleration of Large Language Models with Hierarchical Drafting based on Temporal Locality in Speculative Decoding Unresolved cited work

Reference 29

Resolution
unresolved
raw_fallback, observed 2026-08-08T18:42:16.215431Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-08T18:42:15.078918Z digest=sha256:4aa9cd82450c95d60b23957f0cb444b5ca1941f7b29db2ff4e95c12e3ccd1791

Observation 750f93c6-96bb-4af7-90ee-807665b9c908 · outbound

This paper cites an unresolved cited work.

Lossless Acceleration of Large Language Models with Hierarchical Drafting based on Temporal Locality in Speculative Decoding Unresolved cited work

Reference 30

Resolution
unresolved
raw_fallback, observed 2026-08-08T18:42:16.199635Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-08T18:42:15.083484Z digest=sha256:d6b562656c11c6d3db021f960305626b0da5d4e6dea2ae2ee4005a05603a4f65

Observation 19db47fb-a138-44bf-902b-d0bd7f42a06a · outbound

This paper cites an unresolved cited work.

Lossless Acceleration of Large Language Models with Hierarchical Drafting based on Temporal Locality in Speculative Decoding Unresolved cited work

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-08T18:42:15.087872Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T18:42:15.087872Z digest=sha256:906f42e401175e4f4f74bc5ded6435573319bf065e32bd6145a2dc13a2767fe2

Observation a7dfa254-2731-40c5-9f29-43e033f19d55 · outbound

This paper cites an unresolved cited work.

Lossless Acceleration of Large Language Models with Hierarchical Drafting based on Temporal Locality in Speculative Decoding Unresolved cited work

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-08T18:42:15.092516Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T18:42:15.092516Z digest=sha256:eb74a288717877aeb9bb13a9c259205ebf5219ed193c4754ccc0468e2379b21b

Observation 7227293a-b44b-4ecc-ba54-3809cf6d4b2e · outbound

This paper cites Fast Transformer Decoding: One Write-Head is All You Need.

Lossless Acceleration of Large Language Models with Hierarchical Drafting based on Temporal Locality in Speculative Decoding Fast Transformer Decoding: One Write-Head is All You Need

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-08T18:42:15.096988Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T18:42:15.096988Z digest=sha256:0946bda2e665162c3bd4b99ef68ecf883945117bcd3e075c63bb8c2f2eb6fcbe

Observation b67b68e3-2ae1-4a48-8e95-5edaf4344bc0 · outbound

This paper cites an unresolved cited work.

Lossless Acceleration of Large Language Models with Hierarchical Drafting based on Temporal Locality in Speculative Decoding Unresolved cited work

Reference 34

Resolution
unresolved
raw_fallback, observed 2026-08-08T18:42:16.173396Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-08T18:42:15.102132Z digest=sha256:6918263e3e276da84758e6fc166d546aa4aa3589428b4322334ac332311d1a62

Observation bb22f7e5-b7cf-4918-b335-5a5512537478 · outbound

This paper cites Accelerating LLM Inference with Staged Speculative Decoding.

Lossless Acceleration of Large Language Models with Hierarchical Drafting based on Temporal Locality in Speculative Decoding Accelerating LLM Inference with Staged Speculative Decoding

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-08T18:42:15.107139Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T18:42:15.107139Z digest=sha256:cd8ce0cc6c5e3f43fdd0828648fc9805ed1ecfbdcfca9fc372f3fe3032925cd6

Observation af83926c-f22c-4a20-a3d2-11d54cd64e2f · outbound

This paper cites an unresolved cited work.

Lossless Acceleration of Large Language Models with Hierarchical Drafting based on Temporal Locality in Speculative Decoding Unresolved cited work

Reference 36

Resolution
unresolved
raw_fallback, observed 2026-08-08T18:42:16.157729Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-08T18:42:15.112127Z digest=sha256:d930ba012b9ad52676948a63c9f22c4c08eb1e80a20706485fa2da16350c9427

Observation 9556c128-31dc-43f0-b39a-2902ce0b6315 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Lossless Acceleration of Large Language Models with Hierarchical Drafting based on Temporal Locality in Speculative Decoding Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-08T18:42:15.116726Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T18:42:15.116726Z digest=sha256:c95bdefb9b5da79261e421e594505cf41d72845d3eb7b8543c40421d8dc10515

Observation 9916496e-31e5-436a-84ad-1e1024ee5dcc · outbound

This paper cites an unresolved cited work.

Lossless Acceleration of Large Language Models with Hierarchical Drafting based on Temporal Locality in Speculative Decoding Unresolved cited work

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-08T18:42:15.121906Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T18:42:15.121906Z digest=sha256:2bf9ba6ad20f37d966917847913c448e41285b7bf0e9836e206d14af072d6461

Observation ed87d95a-db03-4aeb-8224-43ca8e6a7508 · outbound

This paper cites an unresolved cited work.

Lossless Acceleration of Large Language Models with Hierarchical Drafting based on Temporal Locality in Speculative Decoding Unresolved cited work

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-08T18:42:15.127040Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T18:42:15.127040Z digest=sha256:a6196776f84955acfb9bbde749c8fbbc7001ce6243d1dee50e430729edfa7236

Observation 9a8efe3b-b0eb-4639-9683-81eae12dbf4c · outbound

This paper cites Inference with Reference: Lossless Acceleration of Large Language Models.

Lossless Acceleration of Large Language Models with Hierarchical Drafting based on Temporal Locality in Speculative Decoding Inference with Reference: Lossless Acceleration of Large Language Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-08T18:42:15.131740Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T18:42:15.131740Z digest=sha256:369b643538527da0df5ad859d37aacf059b00938a552f42ebb788d5b59be2263

Observation 32ca7678-766f-4aa7-a3f1-79a0a1eae9de · outbound

This paper cites an unresolved cited work.

Lossless Acceleration of Large Language Models with Hierarchical Drafting based on Temporal Locality in Speculative Decoding Unresolved cited work

Reference 41

Resolution
unresolved
raw_fallback, observed 2026-08-08T18:42:16.141465Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-08T18:42:15.136772Z digest=sha256:affd560beee73cc3be4451340b8b110a01fc2c7f1ea8f8ef3581990e98c44242

Observation cb0dfb7d-c82f-4609-a07d-062b762e778b · outbound

This paper cites Towards Fast Multilingual LLM Inference: Speculative Decoding and Specialized Drafters.

Lossless Acceleration of Large Language Models with Hierarchical Drafting based on Temporal Locality in Speculative Decoding Towards Fast Multilingual LLM Inference: Speculative Decoding and Specialized Drafters

Reference 42

Resolution
verified exact
local_arxiv, observed 2026-08-08T18:42:15.237786Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-08T18:42:15.141623Z digest=sha256:1b68885ba10e54c1cfef568534ef89c60693451d00dc9631f6bb22d412a59406

Observation 35c70148-7de9-450d-9a49-60a1986827d5 · outbound

This paper cites TinyLlama: An Open-Source Small Language Model.

Lossless Acceleration of Large Language Models with Hierarchical Drafting based on Temporal Locality in Speculative Decoding TinyLlama: An Open-Source Small Language Model

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-08T18:42:15.146922Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T18:42:15.146922Z digest=sha256:63de3c96c5303d62e7d2bf1608e9253e9cec7e5e36fadd79ebff6970dba75c56

Observation 69f869ef-6115-4a00-b32a-d38f0bca5285 · outbound

This paper cites Barrett, Zhangyang Wang, and Beidi Chen.

Lossless Acceleration of Large Language Models with Hierarchical Drafting based on Temporal Locality in Speculative Decoding Barrett, Zhangyang Wang, and Beidi Chen

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-08T18:42:15.151791Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T18:42:15.151791Z digest=sha256:3b8ce538d24531e545acf550dece0eb46aa6613cd05d0a0ac2ef6f201f127448

Observation c74c0f5c-fb70-4011-ad59-a031f2182df3 · outbound

This paper cites Xing, Hao Zhang, Joseph E.

Lossless Acceleration of Large Language Models with Hierarchical Drafting based on Temporal Locality in Speculative Decoding Xing, Hao Zhang, Joseph E

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-08T18:42:15.156291Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T18:42:15.156291Z digest=sha256:18f93e27e37ecf110704149963e97100aa441e27eb3bf8414d39de691951f138

Observation e59a4ecd-2b8d-44cb-839c-1ad196c2bdbf · outbound

This paper cites an unresolved cited work.

Lossless Acceleration of Large Language Models with Hierarchical Drafting based on Temporal Locality in Speculative Decoding Unresolved cited work

Reference 46

Resolution
unresolved
raw_fallback, observed 2026-08-08T18:42:16.104927Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-08T18:42:15.161034Z digest=sha256:a651683ffcd1ab7e3dc943cbf7eaf98cac84a8f5532bf78d1be5c1a8f92e24cb

Observation 65223a9c-9ede-4605-b468-21b56f5ba059 · outbound

This paper cites URL: " 'urlintro :=.

Lossless Acceleration of Large Language Models with Hierarchical Drafting based on Temporal Locality in Speculative Decoding URL: " 'urlintro :=

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-08T18:42:15.165690Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T18:42:15.165690Z digest=sha256:3a54115282e05cad77e5a9da843148627ad1bf814c89a3755e01e8966c3fc269

Observation 69fb4074-f375-4de7-8ed6-14b61166c723 · outbound

This paper cites write newline.

Lossless Acceleration of Large Language Models with Hierarchical Drafting based on Temporal Locality in Speculative Decoding write newline

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-08T18:42:15.171177Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T18:42:15.171177Z digest=sha256:2659af884ab6f3af703e68cf66c8acffbe0c1536214af5c6ab97598e64be1375

Pith citing papers

Observation 512295ff-e0ef-409b-94d9-2f00cf25795f · inbound

Faster and Better LLMs via Latency-Aware Test-Time Scaling cites this paper.

Faster and Better LLMs via Latency-Aware Test-Time Scaling Lossless Acceleration of Large Language Models with Hierarchical Drafting based on Temporal Locality in Speculative Decoding

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T14:14:21.520298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:14:21.520298Z digest=sha256:918ce7a78bb7326174fd5c50e66a0c80454969afc915af4c9fe301d03067bb1f

Observation 61a54cbb-ae6f-49ce-8249-a1a734c7471d · inbound

HeiSD: Hybrid Speculative Decoding for Embodied Vision-Language-Action Models with Kinematic Awareness cites this paper.

HeiSD: Hybrid Speculative Decoding for Embodied Vision-Language-Action Models with Kinematic Awareness Lossless Acceleration of Large Language Models with Hierarchical Drafting based on Temporal Locality in Speculative Decoding

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-15T09:19:54.348149Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-15T09:15:50.963123Z digest=sha256:9f25f8bb1530df07c067da9821379a94f4e781e4edd63033746a511f6bc67fb7