Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T15:25:35.727421Z
Paper Citation Record · LEDGER
As of 16 August 2026, this Paper Citation Record lists 56 of 56 outbound references and 0 inbound Pith citation observations for arXiv:2608.00434.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T15:25:35.727421Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
56 of 56 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 6ca8042f-fe68-43f0-95c0-6a4cae5f0e81 · outbound
AdaMTP: An Adaptive Training Paradigm for Multi-Token Prediction arXiv preprint arXiv:2505.17505 , year=
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b15b3e1e-1ce3-4422-a3d1-93d64aaf5815 · outbound
AdaMTP: An Adaptive Training Paradigm for Multi-Token Prediction Multi-Token Prediction Needs Registers
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4af58e76-e53f-4e04-abac-96ec96e0b318 · outbound
AdaMTP: An Adaptive Training Paradigm for Multi-Token Prediction Faster Language Models with Better Multi-Token Prediction Using Tensor Decomposition
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fddb3267-0753-401d-a3bc-791dfcf6b3f7 · outbound
AdaMTP: An Adaptive Training Paradigm for Multi-Token Prediction arXiv preprint arXiv:2510.14751 , year=
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ab5214b2-cd6d-42a0-b2d6-dabf286e9e62 · outbound
AdaMTP: An Adaptive Training Paradigm for Multi-Token Prediction Better & Faster Large Language Models via Multi-token Prediction
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cb42265e-5eeb-47ad-bdfb-2e05c2109b78 · outbound
AdaMTP: An Adaptive Training Paradigm for Multi-Token Prediction Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ff1c6df5-af5c-4786-b8cb-c9c05f5430d5 · outbound
AdaMTP: An Adaptive Training Paradigm for Multi-Token Prediction DeepSeek-V3 Technical Report
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8c78a3ba-b27a-42dc-8fae-fe74713e6b4b · outbound
AdaMTP: An Adaptive Training Paradigm for Multi-Token Prediction arXiv preprint arXiv:2512.24617 , year=
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 59b27054-2001-4abb-bfbd-55df76adbf8b · outbound
AdaMTP: An Adaptive Training Paradigm for Multi-Token Prediction Efficient Joint Prediction of Multiple Future Tokens
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7901a405-6c62-489d-a70e-1292971f33c1 · outbound
AdaMTP: An Adaptive Training Paradigm for Multi-Token Prediction arXiv preprint arXiv:2509.18362 , year=
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9e6ae26b-e742-4625-9c4d-5d6be734e0cd · outbound
AdaMTP: An Adaptive Training Paradigm for Multi-Token Prediction Medusa: Simple LLM Inference Acceleration Framework with Multiple Decoding Heads
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fae43951-a2c7-4818-9d31-503092c9033b · outbound
AdaMTP: An Adaptive Training Paradigm for Multi-Token Prediction arXiv preprint arXiv:2508.19228 , year=
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 11058976-3e5b-4c53-aeca-19444d6933a2 · outbound
AdaMTP: An Adaptive Training Paradigm for Multi-Token Prediction Your LLM Knows the Future: Uncovering Its Multi-Token Prediction Potential
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a0425c5c-cd91-4de9-9fb7-6485cfb8a0b1 · outbound
AdaMTP: An Adaptive Training Paradigm for Multi-Token Prediction MiMo: Unlocking the Reasoning Potential of Language Model -- From Pretraining to Posttraining
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 49e34df3-9ec8-4be1-bd75-ee6fc9069bf1 · outbound
AdaMTP: An Adaptive Training Paradigm for Multi-Token Prediction Findings of the Association for Computational Linguistics: EMNLP 2020 , pages=
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 8305692f-b842-403e-ae79-19833b7bd8f4 · outbound
AdaMTP: An Adaptive Training Paradigm for Multi-Token Prediction journal of machine learning research , volume=
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 97eec831-accd-4d0d-81d9-b292476d3de8 · outbound
AdaMTP: An Adaptive Training Paradigm for Multi-Token Prediction SqueezeLLM: Dense-and-Sparse Quantization
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b7a82594-e9ac-4fc4-80ea-f36e7ceb11bb · outbound
AdaMTP: An Adaptive Training Paradigm for Multi-Token Prediction Proceedings of machine learning and systems , volume=
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7216b7b8-fb1a-45ea-98d3-c39d1f76900a · outbound
AdaMTP: An Adaptive Training Paradigm for Multi-Token Prediction Advances in neural information processing systems , volume=
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cdb2b6dc-7b8a-43fd-968b-b0030e0584f8 · outbound
AdaMTP: An Adaptive Training Paradigm for Multi-Token Prediction Advances in Neural Information Processing Systems , volume=
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 9197c0e0-8b5d-4333-9d4e-92b4ba92d94f · outbound
AdaMTP: An Adaptive Training Paradigm for Multi-Token Prediction International conference on machine learning , pages=
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cf1ec5e0-6419-4e95-9a30-a9d11bec1ba8 · outbound
AdaMTP: An Adaptive Training Paradigm for Multi-Token Prediction A Simple and Effective Pruning Approach for Large Language Models
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0b6e21e7-5dcb-461b-ba01-d2f755056a13 · outbound
AdaMTP: An Adaptive Training Paradigm for Multi-Token Prediction MiniLLM: On-Policy Distillation of Large Language Models
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c0b25a40-0d95-4f6e-8990-46ce33cdf93d · outbound
AdaMTP: An Adaptive Training Paradigm for Multi-Token Prediction Distilling the Knowledge in a Neural Network
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 684808a8-7cb9-42a2-99f7-92cb59bca5ca · outbound
AdaMTP: An Adaptive Training Paradigm for Multi-Token Prediction Instruction Tuning with GPT-4
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9352eff0-d483-4feb-b47b-8d8aa54f07d0 · outbound
AdaMTP: An Adaptive Training Paradigm for Multi-Token Prediction Findings of the Association for Computational Linguistics: ACL 2023 , pages=
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 47e84ab0-6e18-4f92-9d22-cee4dac4cd70 · outbound
AdaMTP: An Adaptive Training Paradigm for Multi-Token Prediction Proceedings of the 61st annual meeting of the association for computational linguistics (volume 1: long papers) , pages=
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2754cfba-0c70-4c6b-b757-e3c30c72157f · outbound
AdaMTP: An Adaptive Training Paradigm for Multi-Token Prediction International conference on machine learning , pages=
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 68b7cbdc-223d-4165-9389-3710a91fcd55 · outbound
AdaMTP: An Adaptive Training Paradigm for Multi-Token Prediction Advances in neural information processing systems , volume=
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 81b89411-4885-4e9d-b25b-3f9bfe3fb54e · outbound
AdaMTP: An Adaptive Training Paradigm for Multi-Token Prediction Generating Long Sequences with Sparse Transformers
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 07f1da7e-6c0b-4e28-a59a-6fc269ebb59f · outbound
AdaMTP: An Adaptive Training Paradigm for Multi-Token Prediction MoBA: Mixture of Block Attention for Long-Context LLMs
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4341a82d-73e1-40df-a8f2-61b47b70508f · outbound
AdaMTP: An Adaptive Training Paradigm for Multi-Token Prediction DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4578af6f-bcc6-41b3-8bf3-a61977e00d1f · outbound
AdaMTP: An Adaptive Training Paradigm for Multi-Token Prediction Advances in neural information processing systems , volume=
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 93a6d0a6-88c6-4d2c-a717-34668cd039cd · outbound
AdaMTP: An Adaptive Training Paradigm for Multi-Token Prediction FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c4b3350b-d095-4e51-a2ee-e117f5dcb70f · outbound
AdaMTP: An Adaptive Training Paradigm for Multi-Token Prediction Proceedings of the 29th symposium on operating systems principles , pages=
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7f4fc9a8-475d-48af-bd96-b1dbe02f3a2f · outbound
AdaMTP: An Adaptive Training Paradigm for Multi-Token Prediction Advances in neural information processing systems , volume=
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 417c9b87-6b3d-4572-9a70-c8128b0923aa · outbound
AdaMTP: An Adaptive Training Paradigm for Multi-Token Prediction International Conference on Machine Learning , pages=
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9308dce7-57e6-4084-bbbc-044c91fdb25e · outbound
AdaMTP: An Adaptive Training Paradigm for Multi-Token Prediction Accelerating Large Language Model Decoding with Speculative Sampling
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation affac131-8332-409d-85cf-12b8e3c40c94 · outbound
AdaMTP: An Adaptive Training Paradigm for Multi-Token Prediction CCF International Conference on Natural Language Processing and Chinese Computing , pages=
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 0d9a0d3b-0c2d-4d8f-bcde-c13d645c6218 · outbound
AdaMTP: An Adaptive Training Paradigm for Multi-Token Prediction DistillSpec: Improving Speculative Decoding via Knowledge Distillation
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eacf1cdc-5280-41d9-8bb3-33c2eec32f08 · outbound
AdaMTP: An Adaptive Training Paradigm for Multi-Token Prediction Advances in Neural Information Processing Systems , volume=
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c1e56746-f9a9-4d69-baeb-e4335031de2b · outbound
AdaMTP: An Adaptive Training Paradigm for Multi-Token Prediction EAGLE: Speculative Sampling Requires Rethinking Feature Uncertainty
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 05d3129c-804f-49f8-a138-9d38188f3620 · outbound
AdaMTP: An Adaptive Training Paradigm for Multi-Token Prediction Proceedings of the 29th ACM International Conference on Architectural Support for Programming Languages and Operating Systems, Volume 3 , pages=
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 09811252-cabc-45a1-9504-239f8a8f959f · outbound
AdaMTP: An Adaptive Training Paradigm for Multi-Token Prediction Advances in Neural Information Processing Systems , volume=
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 55ee1ba7-a582-4b59-8970-58ca9abebb03 · outbound
AdaMTP: An Adaptive Training Paradigm for Multi-Token Prediction Large Concept Models: Language Modeling in a Sentence Representation Space
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c365b461-3df8-4e58-8af3-f790fd444351 · outbound
AdaMTP: An Adaptive Training Paradigm for Multi-Token Prediction Measuring Mathematical Problem Solving With the MATH Dataset
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 96ab64fb-ae7a-442a-ab22-c86884bbc001 · outbound
AdaMTP: An Adaptive Training Paradigm for Multi-Token Prediction WizardCoder: Empowering Code Large Language Models with Evol-Instruct
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dab79e4b-f641-400a-8fc0-638e4099e2a3 · outbound
AdaMTP: An Adaptive Training Paradigm for Multi-Token Prediction Unresolved cited work
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eabc3651-6d29-4a61-abc5-e0370d1727db · outbound
AdaMTP: An Adaptive Training Paradigm for Multi-Token Prediction The twelfth international conference on learning representations , year=
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b11c1ef7-37a2-4809-bf71-58caf832463d · outbound
AdaMTP: An Adaptive Training Paradigm for Multi-Token Prediction Training Verifiers to Solve Math Word Problems
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8c450dd1-3485-47e4-9749-6d5e12d3efb1 · outbound
AdaMTP: An Adaptive Training Paradigm for Multi-Token Prediction Program Synthesis with Large Language Models
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 56d7718e-3358-42a4-ad22-98e782a646cf · outbound
AdaMTP: An Adaptive Training Paradigm for Multi-Token Prediction Advances in neural information processing systems , volume=
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 479bba2b-2309-4c94-acf7-449e546e8d59 · outbound
AdaMTP: An Adaptive Training Paradigm for Multi-Token Prediction Evaluating Large Language Models Trained on Code
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b2c7d324-7caa-48f1-aa34-ad9b4fa3b362 · outbound
AdaMTP: An Adaptive Training Paradigm for Multi-Token Prediction Measuring Massive Multitask Language Understanding
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 66c737ae-6767-43ff-986b-b459f46ab715 · outbound
AdaMTP: An Adaptive Training Paradigm for Multi-Token Prediction Instruction-Following Evaluation for Large Language Models
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 172356cb-0b9d-4c88-af09-8d1cb59fc6b5 · outbound
AdaMTP: An Adaptive Training Paradigm for Multi-Token Prediction , author=
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.