Pith. sign in

Paper Citation Record · LEDGER

NITP: Next Implicit Token Prediction for LLM Pre-training

As of 15 August 2026, this Paper Citation Record lists 41 of 41 outbound references and 0 inbound Pith citation observations for arXiv:2605.24956.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.24956 v3

Coverage vector

measured 41 of 41 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-14T18:45:28.635910Z

measured 41 of 41 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

41 of 41 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved40
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation f872ea8d-66a4-4c18-858f-a13568cdc3d2 · outbound

This paper cites GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints.

NITP: Next Implicit Token Prediction for LLM Pre-training GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints

Reference 1

Resolution
unresolved
no resolver link, observed 2026-07-14T18:45:28.635910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T18:45:28.635910Z digest=sha256:328e343195d089a38859939406e29eaaacc018afe5522819f80e80f86f665f53

Observation 1c3ab298-d7ac-4ac3-b52e-158f88d19a75 · outbound

This paper cites Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey.

NITP: Next Implicit Token Prediction for LLM Pre-training Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey

Reference 2

Resolution
unresolved
no resolver link, observed 2026-07-14T18:45:28.635910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T18:45:28.635910Z digest=sha256:035ecaa1425e023ee70b6dcb89052f061b05dc402f8594163988581f2792b475

Observation 284b326c-502b-4174-ab73-84a170a6ca51 · outbound

This paper cites Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge.

NITP: Next Implicit Token Prediction for LLM Pre-training Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge

Reference 3

Resolution
unresolved
no resolver link, observed 2026-07-14T18:45:28.635910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T18:45:28.635910Z digest=sha256:e9ccea9f8ee56b9f026c7fcb3927f3b7410e78ec0b4532733a8afa7c02fc35e4

Observation 45c18772-fa57-40d2-b270-10e9b8b39c34 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

NITP: Next Implicit Token Prediction for LLM Pre-training Training Verifiers to Solve Math Word Problems

Reference 4

Resolution
unresolved
no resolver link, observed 2026-07-14T18:45:28.635910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T18:45:28.635910Z digest=sha256:94920dbf007a1a04521556bc831b8b19006a5f5e9d67ee0354562796cdcc8ed5

Observation 45296ac4-6d8c-4cf3-a99d-4716d05b9dfa · outbound

This paper cites How Contextual are Contextualized Word Representations? Comparing the Geometry of BERT, ELMo, and GPT-2 Embeddings.

NITP: Next Implicit Token Prediction for LLM Pre-training How Contextual are Contextualized Word Representations? Comparing the Geometry of BERT, ELMo, and GPT-2 Embeddings

Reference 5

Resolution
unresolved
no resolver link, observed 2026-07-14T18:45:28.635910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T18:45:28.635910Z digest=sha256:4de911d7f46026a25b8c6f1de25392c60fc0f322c4c05348efc6c32c10564163

Observation fe23d654-343f-4ae3-83cb-cb1bc98aa189 · outbound

This paper cites Representation Degeneration Problem in Training Natural Language Generation Models.

NITP: Next Implicit Token Prediction for LLM Pre-training Representation Degeneration Problem in Training Natural Language Generation Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-07-14T18:45:28.635910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T18:45:28.635910Z digest=sha256:6ebbb07aac27d36f69b45e93d06ff4cf7a9d575648a359044ec8b53a6f28b511

Observation bedc496f-98f8-4828-a1bb-346767839919 · outbound

This paper cites The Pile: An 800GB Dataset of Diverse Text for Language Modeling.

NITP: Next Implicit Token Prediction for LLM Pre-training The Pile: An 800GB Dataset of Diverse Text for Language Modeling

Reference 7

Resolution
unresolved
no resolver link, observed 2026-07-14T18:45:28.635910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T18:45:28.635910Z digest=sha256:d0dcf01e8aac4307fdd417ba70aaa6cd6e9566ee74a50e0c59f6c9a7cdef49fc

Observation fee8a5db-d6e9-4976-98eb-aa11357ca289 · outbound

This paper cites Better & Faster Large Language Models via Multi-token Prediction.

NITP: Next Implicit Token Prediction for LLM Pre-training Better & Faster Large Language Models via Multi-token Prediction

Reference 8

Resolution
unresolved
no resolver link, observed 2026-07-14T18:45:28.635910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T18:45:28.635910Z digest=sha256:7a6cddcb8a0cd78e9a8fa3dc3be5e1b8e116146d72e6d492592fa179017a5ebd

Observation fb9c5d06-ff17-4927-bb5f-b367f5d17433 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

NITP: Next Implicit Token Prediction for LLM Pre-training DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 9

Resolution
unresolved
no resolver link, observed 2026-07-14T18:45:28.635910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T18:45:28.635910Z digest=sha256:2f0c156d3cbacb54c7fd2e8a1f0349d8fce65fe18222d6260538162dd50b5763

Observation 20f24f6c-ba20-4f29-8194-7b684153a2dd · outbound

This paper cites Measuring Massive Multitask Language Understanding.

NITP: Next Implicit Token Prediction for LLM Pre-training Measuring Massive Multitask Language Understanding

Reference 10

Resolution
unresolved
no resolver link, observed 2026-07-14T18:45:28.635910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T18:45:28.635910Z digest=sha256:064a4a687f061b8ee2b2e5a336d87d21b8aa963b08ef850662830117e19fe2e1

Observation 8c5556cf-97e5-4241-9fc3-5ca019c6b71f · outbound

This paper cites Training Compute-Optimal Large Language Models.

NITP: Next Implicit Token Prediction for LLM Pre-training Training Compute-Optimal Large Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-07-14T18:45:28.635910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T18:45:28.635910Z digest=sha256:904dfb74897bbe9b8fc86c3c8ab3d32a174a3ffb9e62f0090a60bdc30973ff66

Observation b98441cd-c429-4287-9f12-a99ff671c262 · outbound

This paper cites Tinybert: Distilling bert for natural language understanding.

NITP: Next Implicit Token Prediction for LLM Pre-training Tinybert: Distilling bert for natural language understanding

Reference 12

Resolution
unresolved
no resolver link, observed 2026-07-14T18:45:28.635910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T18:45:28.635910Z digest=sha256:30176e04eb52e5afed54b5b02d376b1d902a24883620dc850328b472ce4d02cb

Observation 3b779c04-d0ac-48bf-bc8f-10e2a7ecb00a · outbound

This paper cites Scaling Laws for Neural Language Models.

NITP: Next Implicit Token Prediction for LLM Pre-training Scaling Laws for Neural Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-07-14T18:45:28.635910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T18:45:28.635910Z digest=sha256:14e330aeb473a5ecb198187301b22ef394ae316981b71076831b5213a47418b4

Observation 23deae2c-26ff-444a-8d49-9ba1b3fce52b · outbound

This paper cites When Choosing Plausible Alternatives, Clever Hans can be Clever.

NITP: Next Implicit Token Prediction for LLM Pre-training When Choosing Plausible Alternatives, Clever Hans can be Clever

Reference 14

Resolution
unresolved
no resolver link, observed 2026-07-14T18:45:28.635910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T18:45:28.635910Z digest=sha256:93c98a6c6cf770be89f48512b0191d2e7359c5506dd9eef7bb4649e587ee3088

Observation 3099c140-0127-490d-8ac2-4375c0ded92b · outbound

This paper cites Self-Distillation for Further Pre-training of Transformers.

NITP: Next Implicit Token Prediction for LLM Pre-training Self-Distillation for Further Pre-training of Transformers

Reference 15

Resolution
unresolved
no resolver link, observed 2026-07-14T18:45:28.635910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T18:45:28.635910Z digest=sha256:212da14615c1e3ed4a34c696c24158706bc822b92f15a753e0ccaf9b8c06dc45

Observation 4e076dbe-51b2-468d-9bd7-c5cea22c1e0a · outbound

This paper cites DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model.

NITP: Next Implicit Token Prediction for LLM Pre-training DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model

Reference 16

Resolution
unresolved
no resolver link, observed 2026-07-14T18:45:28.635910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T18:45:28.635910Z digest=sha256:4ce68304f3a9ba1f9f1a67633b76aa3107caa49da3808659305ba2a5e80c8aa9

Observation e08aa440-a3fe-4438-93b4-fc33196a1850 · outbound

This paper cites Fantastic Semantics and Where to Find Them: Investigating Which Layers of Generative LLMs Reflect Lexical Semantics.

NITP: Next Implicit Token Prediction for LLM Pre-training Fantastic Semantics and Where to Find Them: Investigating Which Layers of Generative LLMs Reflect Lexical Semantics

Reference 17

Resolution
unresolved
no resolver link, observed 2026-07-14T18:45:28.635910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T18:45:28.635910Z digest=sha256:9eaed0aeaca06c0c5ca72beee050182492a7487c2ce7f42ec6110c84cca600b8

Observation bad0bc0a-3a26-41f8-b3fa-c9e66bb9b2ce · outbound

This paper cites Language Models are Few-Shot Learners.

NITP: Next Implicit Token Prediction for LLM Pre-training Language Models are Few-Shot Learners

Reference 18

Resolution
unresolved
no resolver link, observed 2026-07-14T18:45:28.635910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T18:45:28.635910Z digest=sha256:1824aecc646dd2c321cea2bccc78f391ddbff4e51333a7b30e6f290c99e7c4d3

Observation 2099a6f5-6b24-4663-bab6-e0f1aa5c9ee5 · outbound

This paper cites Mteb: Massive text embedding benchmark.

NITP: Next Implicit Token Prediction for LLM Pre-training Mteb: Massive text embedding benchmark

Reference 19

Resolution
unresolved
no resolver link, observed 2026-07-14T18:45:28.635910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T18:45:28.635910Z digest=sha256:dde602101a8a055b05403ed77a5da5e06838225619f0253744484349c130d529

Observation 8ab55965-31d0-4884-9c49-0a38bbd37a0e · outbound

This paper cites Representation Learning with Contrastive Predictive Coding.

NITP: Next Implicit Token Prediction for LLM Pre-training Representation Learning with Contrastive Predictive Coding

Reference 20

Resolution
unresolved
no resolver link, observed 2026-07-14T18:45:28.635910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T18:45:28.635910Z digest=sha256:161de8e379107299d998118e0a7b4f2b19e2b29fc03088a9196b17a63045f763

Observation 4e1aa5cd-9562-4bbb-ac68-0669e7eb8dde · outbound

This paper cites Future Lens: Anticipating Subsequent Tokens from a Single Hidden State.

NITP: Next Implicit Token Prediction for LLM Pre-training Future Lens: Anticipating Subsequent Tokens from a Single Hidden State

Reference 21

Resolution
unresolved
no resolver link, observed 2026-07-14T18:45:28.635910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T18:45:28.635910Z digest=sha256:f5a3906331a81e3b816385894ad7e56a5572b8ede033b586c57ae30a6f89e67c

Observation bd5c26b7-6ce6-4f8e-b151-da3a192d7300 · outbound

This paper cites FitNets: Hints for Thin Deep Nets.

NITP: Next Implicit Token Prediction for LLM Pre-training FitNets: Hints for Thin Deep Nets

Reference 22

Resolution
unresolved
no resolver link, observed 2026-07-14T18:45:28.635910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T18:45:28.635910Z digest=sha256:de211efb0c1e87171d7cdc5785137a9bda0527b8a264e2bb6481cfa0dee96bb5

Observation e912ef23-9fff-4880-a03c-c67b89d2f48e · outbound

This paper cites Your LLM Knows the Future: Uncovering Its Multi-Token Prediction Potential.

NITP: Next Implicit Token Prediction for LLM Pre-training Your LLM Knows the Future: Uncovering Its Multi-Token Prediction Potential

Reference 23

Resolution
unresolved
no resolver link, observed 2026-07-14T18:45:28.635910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T18:45:28.635910Z digest=sha256:25e1addf7a8680baefed00d11946c109ea06fe925c8e849bb6055d61ef010315

Observation abf7d550-48e4-43ae-9c88-02639a54e9e0 · outbound

This paper cites GLU Variants Improve Transformer.

NITP: Next Implicit Token Prediction for LLM Pre-training GLU Variants Improve Transformer

Reference 24

Resolution
unresolved
no resolver link, observed 2026-07-14T18:45:28.635910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T18:45:28.635910Z digest=sha256:702080ad9ed9ae96d3e65dfcc08ae8b3289896c343e5d893dd8597ea2b198a74

Observation 0a267b32-6956-43bc-af60-19367af57619 · outbound

This paper cites Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer.

NITP: Next Implicit Token Prediction for LLM Pre-training Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer

Reference 25

Resolution
unresolved
no resolver link, observed 2026-07-14T18:45:28.635910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T18:45:28.635910Z digest=sha256:615e752f9ff6380fe7b25254ffe5cbd224494001e1e8e8653bc32f20f15405c8

Observation 25fbb8c4-dd70-4ea6-a349-250c75b9e57b · outbound

This paper cites Layer by Layer: Uncovering Hidden Representations in Language Models.

NITP: Next Implicit Token Prediction for LLM Pre-training Layer by Layer: Uncovering Hidden Representations in Language Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-07-14T18:45:28.635910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T18:45:28.635910Z digest=sha256:af11f0123c4c53c77ceacbad01691d77ee017d156cdf209665e17e18109c9161

Observation d7be940a-e20c-4407-add2-74746e405e6a · outbound

This paper cites Patient Knowledge Distillation for BERT Model Compression.

NITP: Next Implicit Token Prediction for LLM Pre-training Patient Knowledge Distillation for BERT Model Compression

Reference 27

Resolution
unresolved
no resolver link, observed 2026-07-14T18:45:28.635910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T18:45:28.635910Z digest=sha256:6f029105df468fdef3e87b1eb0c6ba4095faa24e4b2ec372a81e9f8dc354d61d

Observation 6867df25-a06b-4523-b272-5c9f14dc8d43 · outbound

This paper cites Contrastive distillation on intermediate representations for language model compression.

NITP: Next Implicit Token Prediction for LLM Pre-training Contrastive distillation on intermediate representations for language model compression

Reference 28

Resolution
unresolved
no resolver link, observed 2026-07-14T18:45:28.635910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T18:45:28.635910Z digest=sha256:d359b760d71b758cada373d7164675ad9c50624dacecf3f7a2fa9903da94daf9

Observation 80db826e-7060-4cb0-b06e-9e99428f908b · outbound

This paper cites LLM Pretraining with Continuous Concepts.

NITP: Next Implicit Token Prediction for LLM Pre-training LLM Pretraining with Continuous Concepts

Reference 29

Resolution
unresolved
no resolver link, observed 2026-07-14T18:45:28.635910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T18:45:28.635910Z digest=sha256:3947ca53a22aa45e49c74f3b8a9bc92d953620688c805e9fa47d24c1f2459748

Observation 2bd389a7-40bc-461f-893c-3f7aa8870495 · outbound

This paper cites Com- monsenseqa: A question answering challenge targeting commonsense knowledge.

NITP: Next Implicit Token Prediction for LLM Pre-training Com- monsenseqa: A question answering challenge targeting commonsense knowledge

Reference 30

Resolution
unresolved
no resolver link, observed 2026-07-14T18:45:28.635910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T18:45:28.635910Z digest=sha256:c25592aed90ae444aa3056340b0e586e3235ace6de816c647416e644b4c9d91c

Observation d8b4d4d1-d6fc-4534-b532-4fab4cef1492 · outbound

This paper cites Kimi K2: Open Agentic Intelligence.

NITP: Next Implicit Token Prediction for LLM Pre-training Kimi K2: Open Agentic Intelligence

Reference 31

Resolution
unresolved
no resolver link, observed 2026-07-14T18:45:28.635910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T18:45:28.635910Z digest=sha256:452a9ac79922bc077c5866a1d373f452fe098a6aa7521a227561fafb6532416b

Observation a1f1d358-8aa6-49fa-8bdb-b54f577e9629 · outbound

This paper cites Understanding Warmup-Stable-Decay Learning Rates: A River Valley Loss Landscape Perspective.

NITP: Next Implicit Token Prediction for LLM Pre-training Understanding Warmup-Stable-Decay Learning Rates: A River Valley Loss Landscape Perspective

Reference 32

Resolution
unresolved
no resolver link, observed 2026-07-14T18:45:28.635910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T18:45:28.635910Z digest=sha256:557bd79457b2e30f235e3f74f2030f6ecd8adb4f635d88eec17e6b80630d66c3

Observation 9badcd3a-81d1-49f7-9aa3-0f92107f0e51 · outbound

This paper cites Sheared LLaMA: Accelerating Language Model Pre-training via Structured Pruning.

NITP: Next Implicit Token Prediction for LLM Pre-training Sheared LLaMA: Accelerating Language Model Pre-training via Structured Pruning

Reference 33

Resolution
unresolved
no resolver link, observed 2026-07-14T18:45:28.635910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T18:45:28.635910Z digest=sha256:b558f431436028fb9df3012e462d9853f630329ea37d9cc2c662202bc9ca4a68

Observation 1d4a81b6-8d2c-4bee-b12b-acec38b0ce69 · outbound

This paper cites Qwen3 Technical Report.

NITP: Next Implicit Token Prediction for LLM Pre-training Qwen3 Technical Report

Reference 34

Resolution
unresolved
no resolver link, observed 2026-07-14T18:45:28.635910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T18:45:28.635910Z digest=sha256:51c659c494bb75583c51e070e5f4ca7dd2b16e2b6ef8f7c992cfe2d88302bb5f

Observation 5bf1e222-ee6d-4a1a-8ed2-0a47cad18eef · outbound

This paper cites Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?.

NITP: Next Implicit Token Prediction for LLM Pre-training Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?

Reference 35

Resolution
unresolved
no resolver link, observed 2026-07-14T18:45:28.635910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T18:45:28.635910Z digest=sha256:85c9a4f592c20e990cae1a9875fb3b0875b86c2bcc7dc405c8609b7eef23afed

Observation b4e6b3dd-02c7-4bc0-9151-b766bc10cc32 · outbound

This paper cites Qwen3 Embedding: Advancing Text Embedding and Reranking Through Foundation Models.

NITP: Next Implicit Token Prediction for LLM Pre-training Qwen3 Embedding: Advancing Text Embedding and Reranking Through Foundation Models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-07-14T18:45:28.635910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T18:45:28.635910Z digest=sha256:3651a3e97969338d39d907f078348b6f78de17cfce8bd2a447af6c18d544fb25

Observation 544d1b85-a7b1-43dd-a17b-48903d3d0421 · outbound

This paper cites Repre- sentation degeneration problem in prompt-based models for natural language understanding.

NITP: Next Implicit Token Prediction for LLM Pre-training Repre- sentation degeneration problem in prompt-based models for natural language understanding

Reference 37

Resolution
unresolved
no resolver link, observed 2026-07-14T18:45:28.635910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T18:45:28.635910Z digest=sha256:b60b74d58b4a9a22eccd3f9c6694749b6225cb1d564f4b846a33e9c104002fae

Observation 912ffb0b-1a9b-40b0-802e-9bbb1c9cf8e9 · outbound

This paper cites Agieval: A human- centric benchmark for evaluating foundation models.

NITP: Next Implicit Token Prediction for LLM Pre-training Agieval: A human- centric benchmark for evaluating foundation models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-07-14T18:45:28.635910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T18:45:28.635910Z digest=sha256:213c7383d60f600b9d3efc0fbe91a48da45e9a26f43ee4d5d94b1d9ee6bf958e

Observation 152b2086-2de2-4e05-925e-69cfe7c30b62 · outbound

This paper cites M., Fuadi, E.

NITP: Next Implicit Token Prediction for LLM Pre-training M., Fuadi, E

Reference 39

Resolution
unresolved
no resolver link, observed 2026-07-14T18:45:28.635910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T18:45:28.635910Z digest=sha256:3b49139d4304acfe5b44ba15752ddc148ea704e64cff007f1cd8621a26f1c0ef

Observation 7cbb8c04-2c59-45aa-b79b-526ed788c2ff · outbound

This paper cites The learning rate and global batch size are scaled according to model size, while the context length is fixed to 8192 tokens for all experiments.

NITP: Next Implicit Token Prediction for LLM Pre-training The learning rate and global batch size are scaled according to model size, while the context length is fixed to 8192 tokens for all experiments

Reference 40

Resolution
unresolved
no resolver link, observed 2026-07-14T18:45:28.635910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T18:45:28.635910Z digest=sha256:80f78ecde76c62bfab7dd2d17cae53b649f131b32a1a51a5c3a1c9ff8e70e163

Observation 77a4a9a2-93e2-448e-b89f-02ed21f891e3 · outbound

This paper cites 2.006 (PPL 7.43 vs.

NITP: Next Implicit Token Prediction for LLM Pre-training 2.006 (PPL 7.43 vs

Reference 41

Resolution
malformed identifier
no resolver link, observed 2026-07-14T18:45:28.635910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T18:45:28.635910Z digest=sha256:6f0b0bfdd7c20ba72d8e84d50d738055acd1aed67cd9bb2b62e590b43e49dfd4

Pith citing papers

No inbound Pith citation observations are available.