Pith. sign in

Paper Citation Record · LEDGER

AdaMTP: An Adaptive Training Paradigm for Multi-Token Prediction

As of 16 August 2026, this Paper Citation Record lists 56 of 56 outbound references and 0 inbound Pith citation observations for arXiv:2608.00434.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.00434 v1

Coverage vector

measured 56 of 56 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T15:25:35.727421Z

measured 56 of 56 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

56 of 56 outbound references displayed

  • verified exact0
  • verified fuzzy4
  • unresolved52
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 6ca8042f-fe68-43f0-95c0-6a4cae5f0e81 · outbound

This paper cites arXiv preprint arXiv:2505.17505 , year=.

AdaMTP: An Adaptive Training Paradigm for Multi-Token Prediction arXiv preprint arXiv:2505.17505 , year=

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-15T15:25:35.488365Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:25:35.488365Z digest=sha256:17ac67a4380dede5050be064912dab89a9c3ce4d521e6f41c546e9709ca37bb2

Observation b15b3e1e-1ce3-4422-a3d1-93d64aaf5815 · outbound

This paper cites Multi-Token Prediction Needs Registers.

AdaMTP: An Adaptive Training Paradigm for Multi-Token Prediction Multi-Token Prediction Needs Registers

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-15T15:25:35.492948Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:25:35.492948Z digest=sha256:267a0270bcb98605fe7f33434be21bedc535b1ebb8f995b6c5a593de0763cc1f

Observation 4af58e76-e53f-4e04-abac-96ec96e0b318 · outbound

This paper cites Faster Language Models with Better Multi-Token Prediction Using Tensor Decomposition.

AdaMTP: An Adaptive Training Paradigm for Multi-Token Prediction Faster Language Models with Better Multi-Token Prediction Using Tensor Decomposition

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T15:25:35.497472Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:25:35.497472Z digest=sha256:7c90b5f3fabba0a8c483472aae36fc346cd1fb8ef2c0fc13854ab448ad9bd55c

Observation fddb3267-0753-401d-a3bc-791dfcf6b3f7 · outbound

This paper cites arXiv preprint arXiv:2510.14751 , year=.

AdaMTP: An Adaptive Training Paradigm for Multi-Token Prediction arXiv preprint arXiv:2510.14751 , year=

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-15T15:25:35.502284Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:25:35.502284Z digest=sha256:5e5260370c0514c448f70de51a5c54b8c9ad47ae84d795aa3b185a3c0becfe7c

Observation ab5214b2-cd6d-42a0-b2d6-dabf286e9e62 · outbound

This paper cites Better & Faster Large Language Models via Multi-token Prediction.

AdaMTP: An Adaptive Training Paradigm for Multi-Token Prediction Better & Faster Large Language Models via Multi-token Prediction

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-15T15:25:35.506809Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:25:35.506809Z digest=sha256:d837b43fd814269cceb589d1db2911849f006c4c474e6295aa8a6a9e87b8c9df

Observation cb42265e-5eeb-47ad-bdfb-2e05c2109b78 · outbound

This paper cites Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=.

AdaMTP: An Adaptive Training Paradigm for Multi-Token Prediction Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T15:25:35.511445Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:25:35.511445Z digest=sha256:12b726426f4764e4f2458991a6fe27c27243ac734738f1e98f03628983e86f83

Observation ff1c6df5-af5c-4786-b8cb-c9c05f5430d5 · outbound

This paper cites DeepSeek-V3 Technical Report.

AdaMTP: An Adaptive Training Paradigm for Multi-Token Prediction DeepSeek-V3 Technical Report

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T15:25:35.515460Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:25:35.515460Z digest=sha256:3811114b255700b01b2386691a39acaeab4e6060cec5c5e24349a34169d9dfb5

Observation 8c78a3ba-b27a-42dc-8fae-fe74713e6b4b · outbound

This paper cites arXiv preprint arXiv:2512.24617 , year=.

AdaMTP: An Adaptive Training Paradigm for Multi-Token Prediction arXiv preprint arXiv:2512.24617 , year=

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T15:25:35.520379Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:25:35.520379Z digest=sha256:cfb0995f7ff56171908081af76288b28b79d77a5601762d54958024c8e5a1cf4

Observation 59b27054-2001-4abb-bfbd-55df76adbf8b · outbound

This paper cites Efficient Joint Prediction of Multiple Future Tokens.

AdaMTP: An Adaptive Training Paradigm for Multi-Token Prediction Efficient Joint Prediction of Multiple Future Tokens

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T15:25:35.525235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:25:35.525235Z digest=sha256:d1595bd2903cd6654bae3fc3c89bde58f59c90e3b362553ecf300c72ff810399

Observation 7901a405-6c62-489d-a70e-1292971f33c1 · outbound

This paper cites arXiv preprint arXiv:2509.18362 , year=.

AdaMTP: An Adaptive Training Paradigm for Multi-Token Prediction arXiv preprint arXiv:2509.18362 , year=

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T15:25:35.529828Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:25:35.529828Z digest=sha256:5a0057d1b02ad42eacb47cda92ca17e40394b6cd3272b6b0860ad85d1edaae9a

Observation 9e6ae26b-e742-4625-9c4d-5d6be734e0cd · outbound

This paper cites Medusa: Simple LLM Inference Acceleration Framework with Multiple Decoding Heads.

AdaMTP: An Adaptive Training Paradigm for Multi-Token Prediction Medusa: Simple LLM Inference Acceleration Framework with Multiple Decoding Heads

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T15:25:35.533490Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:25:35.533490Z digest=sha256:ac3c151b13f41b3530c7af25433148c66c1aacd9bf8a225f3552eebe8c36b3c4

Observation fae43951-a2c7-4818-9d31-503092c9033b · outbound

This paper cites arXiv preprint arXiv:2508.19228 , year=.

AdaMTP: An Adaptive Training Paradigm for Multi-Token Prediction arXiv preprint arXiv:2508.19228 , year=

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-15T15:25:35.540077Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:25:35.540077Z digest=sha256:167366a07b730357d81907a952f9d7e65abf11e182a52e1447ff652222147295

Observation 11058976-3e5b-4c53-aeca-19444d6933a2 · outbound

This paper cites Your LLM Knows the Future: Uncovering Its Multi-Token Prediction Potential.

AdaMTP: An Adaptive Training Paradigm for Multi-Token Prediction Your LLM Knows the Future: Uncovering Its Multi-Token Prediction Potential

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-15T15:25:35.544478Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:25:35.544478Z digest=sha256:58af90dd1553f9e42f0af6604f73ea71bf5686b3b78163da8f000caad5259f88

Observation a0425c5c-cd91-4de9-9fb7-6485cfb8a0b1 · outbound

This paper cites MiMo: Unlocking the Reasoning Potential of Language Model -- From Pretraining to Posttraining.

AdaMTP: An Adaptive Training Paradigm for Multi-Token Prediction MiMo: Unlocking the Reasoning Potential of Language Model -- From Pretraining to Posttraining

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-15T15:25:35.548459Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:25:35.548459Z digest=sha256:58b148c613799aa658d290260aa4ada696f68d1a90e89011761d8c134b2ccdba

Observation 49e34df3-9ec8-4be1-bd75-ee6fc9069bf1 · outbound

This paper cites Findings of the Association for Computational Linguistics: EMNLP 2020 , pages=.

AdaMTP: An Adaptive Training Paradigm for Multi-Token Prediction Findings of the Association for Computational Linguistics: EMNLP 2020 , pages=

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:25:36.836408Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T15:25:35.552740Z digest=sha256:b0fbe828637485c8366d0c3c2b7e7cb8fbf4081246952ca7c548d09adf6030da

Observation 8305692f-b842-403e-ae79-19833b7bd8f4 · outbound

This paper cites journal of machine learning research , volume=.

AdaMTP: An Adaptive Training Paradigm for Multi-Token Prediction journal of machine learning research , volume=

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-15T15:25:35.556471Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:25:35.556471Z digest=sha256:b7b1fb2d2f5387e3473c6e27c35ecb14b6147fa12293489e954e7438363a0f45

Observation 97eec831-accd-4d0d-81d9-b292476d3de8 · outbound

This paper cites SqueezeLLM: Dense-and-Sparse Quantization.

AdaMTP: An Adaptive Training Paradigm for Multi-Token Prediction SqueezeLLM: Dense-and-Sparse Quantization

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-15T15:25:35.561194Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:25:35.561194Z digest=sha256:578fca52571bf7c84d9e0b9fe339d8bc8491c769a0ec335f6dd0db0a172b65ba

Observation b7a82594-e9ac-4fc4-80ea-f36e7ceb11bb · outbound

This paper cites Proceedings of machine learning and systems , volume=.

AdaMTP: An Adaptive Training Paradigm for Multi-Token Prediction Proceedings of machine learning and systems , volume=

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-15T15:25:35.565402Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:25:35.565402Z digest=sha256:28cb0568e91e93bf5386797c63511f0a439138096af5f6c8b0814345959063a1

Observation 7216b7b8-fb1a-45ea-98d3-c39d1f76900a · outbound

This paper cites Advances in neural information processing systems , volume=.

AdaMTP: An Adaptive Training Paradigm for Multi-Token Prediction Advances in neural information processing systems , volume=

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-15T15:25:35.569367Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:25:35.569367Z digest=sha256:2ed919ecb89f4c4f22f038a9817ce500ead0ca8314ef31e64832c5ffeec2a860

Observation cdb2b6dc-7b8a-43fd-968b-b0030e0584f8 · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

AdaMTP: An Adaptive Training Paradigm for Multi-Token Prediction Advances in Neural Information Processing Systems , volume=

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:25:36.805981Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T15:25:35.573313Z digest=sha256:6ff69d1e306ed408185b94a016de542c21e6fe90bde360be9f476d3bb2186856

Observation 9197c0e0-8b5d-4333-9d4e-92b4ba92d94f · outbound

This paper cites International conference on machine learning , pages=.

AdaMTP: An Adaptive Training Paradigm for Multi-Token Prediction International conference on machine learning , pages=

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-15T15:25:35.577111Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:25:35.577111Z digest=sha256:0f96277aa41bb0af06661ece8588d89dfb9b68660ae8ab190dc856af98244066

Observation cf1ec5e0-6419-4e95-9a30-a9d11bec1ba8 · outbound

This paper cites A Simple and Effective Pruning Approach for Large Language Models.

AdaMTP: An Adaptive Training Paradigm for Multi-Token Prediction A Simple and Effective Pruning Approach for Large Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-15T15:25:35.581263Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:25:35.581263Z digest=sha256:873f580953bc34b54b79b36087a308c3fb788557de762b9be5b60ead92ce455c

Observation 0b6e21e7-5dcb-461b-ba01-d2f755056a13 · outbound

This paper cites MiniLLM: On-Policy Distillation of Large Language Models.

AdaMTP: An Adaptive Training Paradigm for Multi-Token Prediction MiniLLM: On-Policy Distillation of Large Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-15T15:25:35.584918Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:25:35.584918Z digest=sha256:17584fbb398b64ba77a52cde79650acbf8440adb336ba5b9ddb8044b57dee9d0

Observation c0b25a40-0d95-4f6e-8990-46ce33cdf93d · outbound

This paper cites Distilling the Knowledge in a Neural Network.

AdaMTP: An Adaptive Training Paradigm for Multi-Token Prediction Distilling the Knowledge in a Neural Network

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-15T15:25:35.588779Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:25:35.588779Z digest=sha256:22b2bd3b0d604ae804a2280a2ca08455981d7f32d1dbf0d31a8f02b0f2e5aedb

Observation 684808a8-7cb9-42a2-99f7-92cb59bca5ca · outbound

This paper cites Instruction Tuning with GPT-4.

AdaMTP: An Adaptive Training Paradigm for Multi-Token Prediction Instruction Tuning with GPT-4

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-15T15:25:35.593350Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:25:35.593350Z digest=sha256:904cb976f5f61e58d16c65b55a3d377e055f0e932919a851321748da5fd965d5

Observation 9352eff0-d483-4feb-b47b-8d8aa54f07d0 · outbound

This paper cites Findings of the Association for Computational Linguistics: ACL 2023 , pages=.

AdaMTP: An Adaptive Training Paradigm for Multi-Token Prediction Findings of the Association for Computational Linguistics: ACL 2023 , pages=

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-15T15:25:35.597216Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:25:35.597216Z digest=sha256:3cfac1e6a49f8109b6ee4a69f7bc6e76fef5dcbc08ab5461deac1774f619da16

Observation 47e84ab0-6e18-4f92-9d22-cee4dac4cd70 · outbound

This paper cites Proceedings of the 61st annual meeting of the association for computational linguistics (volume 1: long papers) , pages=.

AdaMTP: An Adaptive Training Paradigm for Multi-Token Prediction Proceedings of the 61st annual meeting of the association for computational linguistics (volume 1: long papers) , pages=

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-15T15:25:35.601190Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:25:35.601190Z digest=sha256:cd8ad328cadf48aa05da7cd003b2a56a8badedb8929f5a4a207d82a216ac4b8e

Observation 2754cfba-0c70-4c6b-b757-e3c30c72157f · outbound

This paper cites International conference on machine learning , pages=.

AdaMTP: An Adaptive Training Paradigm for Multi-Token Prediction International conference on machine learning , pages=

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-15T15:25:35.604775Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:25:35.604775Z digest=sha256:fd5cd1194818b6cc4c2fbc18cbad61113c153667db08695498ddadfc4919cc9a

Observation 68b7cbdc-223d-4165-9389-3710a91fcd55 · outbound

This paper cites Advances in neural information processing systems , volume=.

AdaMTP: An Adaptive Training Paradigm for Multi-Token Prediction Advances in neural information processing systems , volume=

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-15T15:25:35.608730Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:25:35.608730Z digest=sha256:ef3c195a88a97a33b521e86625b882d32c641338e0b5483796c5d56aa5242c66

Observation 81b89411-4885-4e9d-b25b-3f9bfe3fb54e · outbound

This paper cites Generating Long Sequences with Sparse Transformers.

AdaMTP: An Adaptive Training Paradigm for Multi-Token Prediction Generating Long Sequences with Sparse Transformers

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-15T15:25:35.613055Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:25:35.613055Z digest=sha256:8c78ba7c88a9d5cf78bc62cc80e95d4b8b3abf909cbf539b2cfa3e2d9ae694cd

Observation 07f1da7e-6c0b-4e28-a59a-6fc269ebb59f · outbound

This paper cites MoBA: Mixture of Block Attention for Long-Context LLMs.

AdaMTP: An Adaptive Training Paradigm for Multi-Token Prediction MoBA: Mixture of Block Attention for Long-Context LLMs

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-15T15:25:35.617482Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:25:35.617482Z digest=sha256:1f5723b94ea459a77443252bdab91d584fad79f105a4d81e41ebf7ab3fab7e98

Observation 4341a82d-73e1-40df-a8f2-61b47b70508f · outbound

This paper cites DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model.

AdaMTP: An Adaptive Training Paradigm for Multi-Token Prediction DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-15T15:25:35.621851Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:25:35.621851Z digest=sha256:c76cb26063cd372fc9c2049761130e2e57e698f378d9e3ddb5d2db0bdba452e7

Observation 4578af6f-bcc6-41b3-8bf3-a61977e00d1f · outbound

This paper cites Advances in neural information processing systems , volume=.

AdaMTP: An Adaptive Training Paradigm for Multi-Token Prediction Advances in neural information processing systems , volume=

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-15T15:25:35.626115Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:25:35.626115Z digest=sha256:7aea99e7c84d8334f4feab9c24e520a225a8fdba41c5d5c311bec94571a67a69

Observation 93a6d0a6-88c6-4d2c-a717-34668cd039cd · outbound

This paper cites FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning.

AdaMTP: An Adaptive Training Paradigm for Multi-Token Prediction FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-15T15:25:35.629844Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:25:35.629844Z digest=sha256:db61c7578061c05c8be8283ad2c117adec5f00d876798dee0e4746fd7d835068

Observation c4b3350b-d095-4e51-a2ee-e117f5dcb70f · outbound

This paper cites Proceedings of the 29th symposium on operating systems principles , pages=.

AdaMTP: An Adaptive Training Paradigm for Multi-Token Prediction Proceedings of the 29th symposium on operating systems principles , pages=

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-15T15:25:35.634067Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:25:35.634067Z digest=sha256:b7f72ed3cb2804c1feeddc98d65f22e1f3ec74b543cf48a4ff4e34664ad9f28f

Observation 7f4fc9a8-475d-48af-bd96-b1dbe02f3a2f · outbound

This paper cites Advances in neural information processing systems , volume=.

AdaMTP: An Adaptive Training Paradigm for Multi-Token Prediction Advances in neural information processing systems , volume=

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-15T15:25:35.638626Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:25:35.638626Z digest=sha256:395669fe72d0b62f32b350e7804c7fa597c86b417197846408b197a387428951

Observation 417c9b87-6b3d-4572-9a70-c8128b0923aa · outbound

This paper cites International Conference on Machine Learning , pages=.

AdaMTP: An Adaptive Training Paradigm for Multi-Token Prediction International Conference on Machine Learning , pages=

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-15T15:25:35.642428Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:25:35.642428Z digest=sha256:82588732006731f52fffee985d468b9d9a922614d949d4d654632531709e1401

Observation 9308dce7-57e6-4084-bbbc-044c91fdb25e · outbound

This paper cites Accelerating Large Language Model Decoding with Speculative Sampling.

AdaMTP: An Adaptive Training Paradigm for Multi-Token Prediction Accelerating Large Language Model Decoding with Speculative Sampling

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-15T15:25:35.646913Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:25:35.646913Z digest=sha256:2a3f2c4aa6ce7e52f70345aa3a6d3b87da847745f1fae7a735fc173cc371e8a1

Observation affac131-8332-409d-85cf-12b8e3c40c94 · outbound

This paper cites CCF International Conference on Natural Language Processing and Chinese Computing , pages=.

AdaMTP: An Adaptive Training Paradigm for Multi-Token Prediction CCF International Conference on Natural Language Processing and Chinese Computing , pages=

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:25:36.733156Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T15:25:35.651547Z digest=sha256:581ab300c6d071efb0e372fba075cc51a963e13d278876fc565e6e85c0498f45

Observation 0d9a0d3b-0c2d-4d8f-bcde-c13d645c6218 · outbound

This paper cites DistillSpec: Improving Speculative Decoding via Knowledge Distillation.

AdaMTP: An Adaptive Training Paradigm for Multi-Token Prediction DistillSpec: Improving Speculative Decoding via Knowledge Distillation

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-15T15:25:35.655628Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:25:35.655628Z digest=sha256:23a7097b1c6eb93d82a96965c2768619a1ae3d21f192a7220dc009dbcb999161

Observation eacf1cdc-5280-41d9-8bb3-33c2eec32f08 · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

AdaMTP: An Adaptive Training Paradigm for Multi-Token Prediction Advances in Neural Information Processing Systems , volume=

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-15T15:25:35.660295Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:25:35.660295Z digest=sha256:cc14ccd08c2aa65193ddddc70c7d12a65735f2c63d68cd2aea69bd7a84272633

Observation c1e56746-f9a9-4d69-baeb-e4335031de2b · outbound

This paper cites EAGLE: Speculative Sampling Requires Rethinking Feature Uncertainty.

AdaMTP: An Adaptive Training Paradigm for Multi-Token Prediction EAGLE: Speculative Sampling Requires Rethinking Feature Uncertainty

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-15T15:25:35.664561Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:25:35.664561Z digest=sha256:66fd703d132de98a971b9570f80da13d280b53ec7e1ba8439b0f764b97f190cf

Observation 05d3129c-804f-49f8-a138-9d38188f3620 · outbound

This paper cites Proceedings of the 29th ACM International Conference on Architectural Support for Programming Languages and Operating Systems, Volume 3 , pages=.

AdaMTP: An Adaptive Training Paradigm for Multi-Token Prediction Proceedings of the 29th ACM International Conference on Architectural Support for Programming Languages and Operating Systems, Volume 3 , pages=

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-15T15:25:35.668809Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:25:35.668809Z digest=sha256:e3d5bc5e0da427e8d2266546742ade83160dfb7116b607803446d36827028475

Observation 09811252-cabc-45a1-9504-239f8a8f959f · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

AdaMTP: An Adaptive Training Paradigm for Multi-Token Prediction Advances in Neural Information Processing Systems , volume=

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:25:36.706553Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-15T15:25:35.672899Z digest=sha256:5a2fc0084a1a197e04ca7e5a96f5573cd1391e302747856f7c3b9c8d9b839899

Observation 55ee1ba7-a582-4b59-8970-58ca9abebb03 · outbound

This paper cites Large Concept Models: Language Modeling in a Sentence Representation Space.

AdaMTP: An Adaptive Training Paradigm for Multi-Token Prediction Large Concept Models: Language Modeling in a Sentence Representation Space

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-15T15:25:35.677765Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:25:35.677765Z digest=sha256:3a4f618e17cae50a74611781b85615e444eddaedfaa1843b70844a8499150d2c

Observation c365b461-3df8-4e58-8af3-f790fd444351 · outbound

This paper cites Measuring Mathematical Problem Solving With the MATH Dataset.

AdaMTP: An Adaptive Training Paradigm for Multi-Token Prediction Measuring Mathematical Problem Solving With the MATH Dataset

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-15T15:25:35.683168Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:25:35.683168Z digest=sha256:33ef3434cb18febf804312f978d60a062a6abe3369813551d22e85e49d9b1e19

Observation 96ab64fb-ae7a-442a-ab22-c86884bbc001 · outbound

This paper cites WizardCoder: Empowering Code Large Language Models with Evol-Instruct.

AdaMTP: An Adaptive Training Paradigm for Multi-Token Prediction WizardCoder: Empowering Code Large Language Models with Evol-Instruct

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-15T15:25:35.687143Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:25:35.687143Z digest=sha256:ce8383f7edef46afb9a28bad71d9212e473ea1fe371d66d174d9e7137a774541

Observation dab79e4b-f641-400a-8fc0-638e4099e2a3 · outbound

This paper cites an unresolved cited work.

AdaMTP: An Adaptive Training Paradigm for Multi-Token Prediction Unresolved cited work

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-15T15:25:35.691476Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:25:35.691476Z digest=sha256:18ff80897bc9f0ac2c892e2bab65b0da60f72226a069887bd0cdf36e1bd9796b

Observation eabc3651-6d29-4a61-abc5-e0370d1727db · outbound

This paper cites The twelfth international conference on learning representations , year=.

AdaMTP: An Adaptive Training Paradigm for Multi-Token Prediction The twelfth international conference on learning representations , year=

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-15T15:25:35.696223Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:25:35.696223Z digest=sha256:7f358ca23d9b69f5d9cf40fa137bea0b56d3a04a9c7a0883f032ec6e544f486b

Observation b11c1ef7-37a2-4809-bf71-58caf832463d · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

AdaMTP: An Adaptive Training Paradigm for Multi-Token Prediction Training Verifiers to Solve Math Word Problems

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-15T15:25:35.700300Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:25:35.700300Z digest=sha256:9b82f79e0bc7aefa2e4ff11b9538332fe1f673a437060e8f26c4f268d051faea

Observation 8c450dd1-3485-47e4-9749-6d5e12d3efb1 · outbound

This paper cites Program Synthesis with Large Language Models.

AdaMTP: An Adaptive Training Paradigm for Multi-Token Prediction Program Synthesis with Large Language Models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-15T15:25:35.704273Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:25:35.704273Z digest=sha256:6b8290dd911f5b1e4d374b7dc400cc7d98e13c97ca5ea039d8d644203ec0c505

Observation 56d7718e-3358-42a4-ad22-98e782a646cf · outbound

This paper cites Advances in neural information processing systems , volume=.

AdaMTP: An Adaptive Training Paradigm for Multi-Token Prediction Advances in neural information processing systems , volume=

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-15T15:25:35.708837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:25:35.708837Z digest=sha256:3e872bfdcc7c0caa4171f94d35ebe8de197c0e70075c35a4b7992921b7f6ce81

Observation 479bba2b-2309-4c94-acf7-449e546e8d59 · outbound

This paper cites Evaluating Large Language Models Trained on Code.

AdaMTP: An Adaptive Training Paradigm for Multi-Token Prediction Evaluating Large Language Models Trained on Code

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-15T15:25:35.713859Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:25:35.713859Z digest=sha256:cb7754a97d3e5dfae752d4766a37785fef94fbf2237386712181c9b6329933d9

Observation b2c7d324-7caa-48f1-aa34-ad9b4fa3b362 · outbound

This paper cites Measuring Massive Multitask Language Understanding.

AdaMTP: An Adaptive Training Paradigm for Multi-Token Prediction Measuring Massive Multitask Language Understanding

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-15T15:25:35.718483Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:25:35.718483Z digest=sha256:62bf3542728a9c9517c677a91aa28f447959d0e92f733cca06ae01708d4fecf0

Observation 66c737ae-6767-43ff-986b-b459f46ab715 · outbound

This paper cites Instruction-Following Evaluation for Large Language Models.

AdaMTP: An Adaptive Training Paradigm for Multi-Token Prediction Instruction-Following Evaluation for Large Language Models

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-15T15:25:35.722864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:25:35.722864Z digest=sha256:a9d96c1fca7e6ee5fedb0fa7e8b173f9f89380e4f7da66bf94f310e6ab6be881

Observation 172356cb-0b9d-4c88-af09-8d1cb59fc6b5 · outbound

This paper cites , author=.

AdaMTP: An Adaptive Training Paradigm for Multi-Token Prediction , author=

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-15T15:25:35.727421Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:25:35.727421Z digest=sha256:824061629bdf6d69635d1d6299ab6118e1cd061cf2976f334ee4bf771088b078

Pith citing papers

No inbound Pith citation observations are available.