Pith. sign in

Paper Citation Record · LEDGER

On multi-token prediction for efficient LLM inference

As of 21 August 2026, this Paper Citation Record lists 12 of 12 outbound references and 0 inbound Pith citation observations for arXiv:2502.09419.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.09419 v1

Coverage vector

measured 12 of 12 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T21:38:35.138928Z

measured 12 of 12 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

12 of 12 outbound references displayed

  • verified exact0
  • verified fuzzy1
  • unresolved11
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 7e1c89d7-4123-4a29-b3d4-e9397a08b79a · outbound

This paper cites Llama 3 model card.

On multi-token prediction for efficient LLM inference Llama 3 model card

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T21:38:34.611805Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T21:38:34.611805Z digest=sha256:bb4c1ad1180ddc80b6fbdf098df568df5def8600ec7140f0d0c4326c6cf9f78d

Observation 63302489-244a-41a6-ad7d-cbf2bcfede47 · outbound

This paper cites Faster Language Models with Better Multi-Token Prediction Using Tensor Decomposition.

On multi-token prediction for efficient LLM inference Faster Language Models with Better Multi-Token Prediction Using Tensor Decomposition

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T21:38:34.615153Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T21:38:34.615153Z digest=sha256:7a0c66746619a25a5b44d9712796fd9dabcd94998d651d963e5d6bdc52dd45a4

Observation cd1dd56f-b51d-4fc3-8f37-6ac15e6d8f99 · outbound

This paper cites Pythia: A suite for analyzing large language models across training and scaling.

On multi-token prediction for efficient LLM inference Pythia: A suite for analyzing large language models across training and scaling

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T21:38:34.617991Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T21:38:34.617991Z digest=sha256:5b8bb8f2a53ef1e9f45e5421de9c535d9a6f812908196c6d895f9bffb2216965

Observation e64287f3-1e88-4e5b-b48b-0f970b9dfa8c · outbound

This paper cites Medusa: Simple LLM Inference Acceleration Framework with Multiple Decoding Heads.

On multi-token prediction for efficient LLM inference Medusa: Simple LLM Inference Acceleration Framework with Multiple Decoding Heads

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T21:38:34.620399Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T21:38:34.620399Z digest=sha256:7f0a2ca4bc19b27f2def97a8e824ef24a06397aab87e7dbbba732c885aacb8c5

Observation abc25476-d4c4-4961-afff-c60c89533ed5 · outbound

This paper cites Overview of the IWSLT 2017 evaluation campaign.

On multi-token prediction for efficient LLM inference Overview of the IWSLT 2017 evaluation campaign

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T21:38:35.309185Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-07T21:38:34.623866Z digest=sha256:7b10040e9481283deef8199f0b4c53b454af56efd9da772ec6dcf087788d876f

Observation d54e0028-cae2-44b5-8d96-da2099414230 · outbound

This paper cites Better & Faster Large Language Models via Multi-token Prediction.

On multi-token prediction for efficient LLM inference Better & Faster Large Language Models via Multi-token Prediction

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T21:38:34.626760Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T21:38:34.626760Z digest=sha256:47f7326c39bac082ae16ad5bcea722671a031002ca33df6ffff264b1670ef971

Observation 6a4f7d3e-5c02-4737-8617-3a3bfe9b68ae · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

On multi-token prediction for efficient LLM inference LoRA: Low-Rank Adaptation of Large Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T21:38:34.707368Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T21:38:34.707368Z digest=sha256:f7e5dc4572bc29fadfaa3e365a5b1f7ec8c34b4a0e9ffb80c75e22c2954ab6d4

Observation 4eee505b-e77f-4694-be24-ab55ae4e2279 · outbound

This paper cites SqueezeLLM: Dense-and-Sparse Quantization.

On multi-token prediction for efficient LLM inference SqueezeLLM: Dense-and-Sparse Quantization

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T21:38:34.760608Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T21:38:34.760608Z digest=sha256:5600521fa6ea30d25a365f9ba43839cecbae56e0b4caddf5a661d3e649a6c70c

Observation dbb0d9f6-7754-44b0-88da-a21342b2432c · outbound

This paper cites write newline.

On multi-token prediction for efficient LLM inference write newline

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T21:38:34.861556Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T21:38:34.861556Z digest=sha256:b12cb36438f29f80b5f7c656aa184ba444c488e5efbdb69a2359b3b56155b59d

Observation 1e0ca3e2-175a-48f9-8738-c2b90a54b7d9 · outbound

This paper cites @esa (Ref.

On multi-token prediction for efficient LLM inference @esa (Ref

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T21:38:34.929669Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T21:38:34.929669Z digest=sha256:9b679037d05208494d2f7d85566b0bd6e709d41262f2d79f79363e28442b528d

Observation 17263872-a3ca-4f77-882a-ed1c88d1b384 · outbound

This paper cites an unresolved cited work.

On multi-token prediction for efficient LLM inference Unresolved cited work

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T21:38:35.011306Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T21:38:35.011306Z digest=sha256:c12347a79c4ac47de2e2ab27da7433d41df0565a0a9ecd3b46e3832d052dcefc

Observation 87a04758-82ba-4bb8-84f4-32199623e755 · outbound

This paper cites an unresolved cited work.

On multi-token prediction for efficient LLM inference Unresolved cited work

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T21:38:35.138928Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T21:38:35.138928Z digest=sha256:8296238a9311631fb404164f36d2560b014be45c7295a3641ec83bf567ee4b31

Pith citing papers

No inbound Pith citation observations are available.