Pith. sign in

Paper Citation Record · LEDGER

Multi-Token Prediction Needs Registers

As of 19 August 2026, this Paper Citation Record lists 24 of 24 outbound references and 2 inbound Pith citation observations for arXiv:2505.10518.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.10518 v1

Coverage vector

measured 24 of 24 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T21:14:53.320867Z

measured 26 of 26 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T15:25:35.492948Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T12:25:29.113950Z

Reference resolution

24 of 24 outbound references displayed

  • verified exact0
  • verified fuzzy8
  • unresolved16
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 7910becd-abd1-4bd4-994b-db9a7640127f · outbound

This paper cites GPT-4 Technical Report.

Multi-Token Prediction Needs Registers GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-15T21:14:53.200999Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:14:53.200999Z digest=sha256:c697568cbe30562d7c30de37686fc8bdebaef30061e12a6ce1b9444307be97d8

Observation 15f53ac5-f242-466b-b4cb-829056dbc784 · outbound

This paper cites Imagenet: A large-scale hierarchical image database.

Multi-Token Prediction Needs Registers Imagenet: A large-scale hierarchical image database

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-15T21:14:53.223634Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:14:53.223634Z digest=sha256:24a3862a40c0ce9dc1809017e12ea6f514fb55e214559813a5d01b2a0633314d

Observation 30c0227c-7069-48da-8646-37d69ae58bae · outbound

This paper cites The mystery of the pathological path-star task for language models.

Multi-Token Prediction Needs Registers The mystery of the pathological path-star task for language models

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:14:53.699784Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T21:14:53.228803Z digest=sha256:f8c82f7b1975f97e6e72e32ad166ca882f5845ddfbe587dfc09b1233cfa8e14a

Observation b6e43809-b614-46a3-a43e-2a0880c0d702 · outbound

This paper cites Gemma: Open Models Based on Gemini Research and Technology.

Multi-Token Prediction Needs Registers Gemma: Open Models Based on Gemini Research and Technology

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T21:14:53.233986Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:14:53.233986Z digest=sha256:0cdfa4f368cd0a0fc50c8a3a9f5eb86f3877bd75886d021c2233e8d636438dd1

Observation 2e8701e5-e279-4a06-b335-db83514250b9 · outbound

This paper cites The Llama 3 Herd of Models.

Multi-Token Prediction Needs Registers The Llama 3 Herd of Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T21:14:53.239228Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:14:53.239228Z digest=sha256:6916be633a5dbd03889d1fd16325381a8575ed624f4e2afb10f996029251b27b

Observation 4175f793-3e84-477d-9427-f3fee9f615db · outbound

This paper cites Bitune: Leveraging Bidirectional Attention to Improve Decoder-Only LLMs.

Multi-Token Prediction Needs Registers Bitune: Leveraging Bidirectional Attention to Improve Decoder-Only LLMs

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T21:14:53.249225Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:14:53.249225Z digest=sha256:37fada5baab90e3bc8872ab9c4e8367e32147a4b15fc80acb812d4c66ec09444

Observation c3a709bc-9955-4a5d-beec-00fcc6e10eab · outbound

This paper cites DeepSeek-V3 Technical Report.

Multi-Token Prediction Needs Registers DeepSeek-V3 Technical Report

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T21:14:53.254657Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:14:53.254657Z digest=sha256:cb303fece7a8e87356b3e8e19c6a354708e57b0713d108fc775822b74a067432

Observation 9549d15b-ccdf-4a70-b3b7-d2ac80cdf80a · outbound

This paper cites Decoupled Weight Decay Regularization.

Multi-Token Prediction Needs Registers Decoupled Weight Decay Regularization

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-15T21:14:53.259447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:14:53.259447Z digest=sha256:7e2b62dfafed3a4021910b8508d5225bf6e4d6c198ea3ef760ba8ffc012d7e12

Observation 406b36be-375a-4119-9042-829a2c2d1c8a · outbound

This paper cites PaSS: Parallel Speculative Sampling.

Multi-Token Prediction Needs Registers PaSS: Parallel Speculative Sampling

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-15T21:14:53.265049Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:14:53.265049Z digest=sha256:0d01f168bd11f67d763d783ac5ffd17c1dc7bff7b77b9505100ba71030e753a9

Observation 1669f754-baec-43be-9045-64374ce8370e · outbound

This paper cites RandAR: Decoder-only Autoregressive Visual Generation in Random Orders.

Multi-Token Prediction Needs Registers RandAR: Decoder-only Autoregressive Visual Generation in Random Orders

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-15T21:14:53.270468Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:14:53.270468Z digest=sha256:fcca7b2413d5aabb8aa51f04a5e95e240cea39cf5061954ae53580d931fdc09a

Observation 0451d1ba-3046-4107-b525-36d553a8e648 · outbound

This paper cites Prophetnet: Predicting future n-gram for sequence-to-sequencepre-training.

Multi-Token Prediction Needs Registers Prophetnet: Predicting future n-gram for sequence-to-sequencepre-training

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:14:53.682727Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T21:14:53.275573Z digest=sha256:44c08ff13998ced2649c7804de926b8e70b03eaf020e9981e9510733572d43da

Observation d8258642-b711-4a59-a160-597793173e49 · outbound

This paper cites Autoregressive Model Beats Diffusion: Llama for Scalable Image Generation.

Multi-Token Prediction Needs Registers Autoregressive Model Beats Diffusion: Llama for Scalable Image Generation

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-15T21:14:53.280605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:14:53.280605Z digest=sha256:1cc55979de246db12f44e29823afd79ef1af047d0706dcc74f314978c205c7ea

Observation 8b622e83-4041-413b-8f09-9bd8ca2a4fcf · outbound

This paper cites Guiding Language Model Reasoning with Planning Tokens.

Multi-Token Prediction Needs Registers Guiding Language Model Reasoning with Planning Tokens

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-15T21:14:53.290784Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:14:53.290784Z digest=sha256:b44b9547368a87fce2567ba544fb86e3b24046b73a258c174deb474b60127475

Observation e87206bf-d455-4993-b885-b439341fbbd6 · outbound

This paper cites Randomized Autoregressive Visual Generation.

Multi-Token Prediction Needs Registers Randomized Autoregressive Visual Generation

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-15T21:14:53.301061Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:14:53.301061Z digest=sha256:4be0936a3b1746ebe56351a1863a977cfc301ac531fc619746c63e7f1af9dc12

Observation b1acb46c-63d7-4814-9c32-369a041b63fa · outbound

This paper cites In order to facilitate reproducibility, we plan to release the code for MuToR’s implementation as well.

Multi-Token Prediction Needs Registers In order to facilitate reproducibility, we plan to release the code for MuToR’s implementation as well

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:14:53.648553Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T21:14:53.305941Z digest=sha256:ae28f0c39414102dbe75e5d5259ca74503d37485bec1364d438a540547cf19b9

Observation 85b4505f-344d-4bd0-a56e-58be0233362d · outbound

This paper cites Both language models are loaded and finetuned in bfloat16 precision.

Multi-Token Prediction Needs Registers Both language models are loaded and finetuned in bfloat16 precision

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:14:53.632056Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T21:14:53.310872Z digest=sha256:37cd3c757eca7fc6ab941cdd49f8d6b49f06944d271c0775f26af33dec2f4994

Observation 6a11299b-71e2-4f53-912c-eb62730c5d27 · outbound

This paper cites Both Next-Token baseline and MuToR are trained for 360K update steps, using AdamW optimizer withβ1 = 0.9,β 2 = 0.95 and weight decay = 0.05.

Multi-Token Prediction Needs Registers Both Next-Token baseline and MuToR are trained for 360K update steps, using AdamW optimizer withβ1 = 0.9,β 2 = 0.95 and weight decay = 0.05

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:14:53.615779Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T21:14:53.316107Z digest=sha256:0740800ddeb4993bff96e5b1cd82fee43473239a445b4856a028a66ce03dd41c

Observation 4b709c17-52d2-4266-ae62-1b5dbcd2e616 · outbound

This paper cites This indicates that the auxiliary loss provides valuable supervision that the model leverages during pretraining, to improve its learned representations.

Multi-Token Prediction Needs Registers This indicates that the auxiliary loss provides valuable supervision that the model leverages during pretraining, to improve its learned representations

Reference 1024

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:14:53.600205Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T21:14:53.320867Z digest=sha256:1e1056355444c161185594be0f9ddfe49704e34a6c769762a09f40df30cf95c9

Observation 2445add7-740b-403c-a7f7-2f5bf655e5e4 · outbound

This paper cites Classifier-Free Diffusion Guidance.

Multi-Token Prediction Needs Registers Classifier-Free Diffusion Guidance

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-15T21:14:53.244229Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:14:53.244229Z digest=sha256:195de5809e6fa79decfcac0edc29fd86233df3a55a17faa702fcb370f12d33cd

Observation c10772db-3457-4792-b5ba-99503f972cfa · outbound

This paper cites Semformer: Transformer language models with semantic planning.

Multi-Token Prediction Needs Registers Semformer: Transformer language models with semantic planning

Reference 2019

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:14:53.664891Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T21:14:53.295723Z digest=sha256:0cb91b0923bde6c754383f25a2f29c0a18d7b2b678e1e468ecd851779c5b8e25

Observation fce1d8de-d9e9-4d92-ad9e-41458987f1cb · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Multi-Token Prediction Needs Registers Training Verifiers to Solve Math Word Problems

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-15T21:14:53.217969Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:14:53.217969Z digest=sha256:805089e98d292048b02067b474c0243f4057d4a54e337843045a8aaa7e18ffae

Observation 1ceeec81-11a8-4bfa-a33f-bad853c5936c · outbound

This paper cites Memory Transformer.

Multi-Token Prediction Needs Registers Memory Transformer

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-15T21:14:53.207001Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:14:53.207001Z digest=sha256:af596bf0e185e26753072d343b9bbd1ac7de5c6bcbcffe98576bd1ac5587cfab

Observation b2e8daba-03a5-4dd3-bb4d-4301844b6f03 · outbound

This paper cites Dialogsum: A real-life scenario dialogue summarization dataset.

Multi-Token Prediction Needs Registers Dialogsum: A real-life scenario dialogue summarization dataset

Reference 2024

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:14:53.726738Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T21:14:53.212443Z digest=sha256:4ceb903d3cd12a573212033700befba29393b1f6cc5a0876dfc78f7c4db41647

Observation 65102be2-7037-4829-87fe-9524e9531c9e · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Multi-Token Prediction Needs Registers LLaMA: Open and Efficient Foundation Language Models

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-15T21:14:53.286035Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:14:53.286035Z digest=sha256:ba56422b60e9c231b24844d39e25fc747b7f4b7613c6164ff45cdff3fb46d67b

Pith citing papers

Observation efb07af1-884e-43c0-9ec7-f8dbe9355320 · inbound

Distilled Pretraining: A modern lens of Data, In-Context Learning and Test-Time Scaling cites this paper.

Distilled Pretraining: A modern lens of Data, In-Context Learning and Test-Time Scaling Multi-Token Prediction Needs Registers

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-08-05T12:25:29.120281Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-05T12:25:25.612205Z digest=sha256:1663044647c2fc914ce4581c7380a98ee515ff21f055173a895a5c8039d8c696

Observation b15b3e1e-1ce3-4422-a3d1-93d64aaf5815 · inbound

AdaMTP: An Adaptive Training Paradigm for Multi-Token Prediction cites this paper.

AdaMTP: An Adaptive Training Paradigm for Multi-Token Prediction Multi-Token Prediction Needs Registers

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-15T15:25:35.492948Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:25:35.492948Z digest=sha256:b3be2c51089720fcf1cd32d72803fa0f4c3140d8c330382f31f0436293b0ceb2