Pith. sign in

Paper Citation Record · LEDGER

Tracing LLM Behavior to the Training Data with Empirical Next-Token Distributions

As of 10 August 2026, this Paper Citation Record lists 12 of 12 outbound references and 0 inbound Pith citation observations for arXiv:2607.14306.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.14306 v2

Coverage vector

measured 12 of 12 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-04T01:45:19.127391Z

measured 12 of 12 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

12 of 12 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved12
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation aa6f2505-82ca-46f3-a3ac-67c21dfb0bc3 · outbound

This paper cites The Pile: An 800GB Dataset of Diverse Text for Language Modeling.

Tracing LLM Behavior to the Training Data with Empirical Next-Token Distributions The Pile: An 800GB Dataset of Diverse Text for Language Modeling

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-04T01:45:18.540645Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:45:18.540645Z digest=sha256:8e0b31ac53c76e08b568e46f84cb5a05f6d2807022d586088ecf7c4e3bd091b3

Observation 150c4578-eaf8-4766-b1d3-ced5db3b2652 · outbound

This paper cites Why Larger Models Learn More: Effects of Capacity, Interference, and Rare-Task Retention.

Tracing LLM Behavior to the Training Data with Empirical Next-Token Distributions Why Larger Models Learn More: Effects of Capacity, Interference, and Rare-Task Retention

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-04T01:45:18.695606Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:45:18.695606Z digest=sha256:0f411e851b9d7ce21952d809a2e64e2193c9b7298a5c0f7c9117ed51c4aebf07

Observation 780bf20f-282f-46c5-85ef-4a661bf8fd7a · outbound

This paper cites In-context Learning and Induction Heads.

Tracing LLM Behavior to the Training Data with Empirical Next-Token Distributions In-context Learning and Induction Heads

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-04T01:45:18.749882Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:45:18.749882Z digest=sha256:70324913f7492e5348bb1680fd80c29beefc263ac0832ed69c3275d2ee9dfa9b

Observation 0cf14fd0-76dc-4043-90c9-9f4180ce1986 · outbound

This paper cites Backtracking mathematical reasoning of language models to the pretraining data.

Tracing LLM Behavior to the Training Data with Empirical Next-Token Distributions Backtracking mathematical reasoning of language models to the pretraining data

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-04T01:45:18.821579Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:45:18.821579Z digest=sha256:25056cd7955e06ec9c3cf9fa4afbb26cf19481f30e159b37a72c9050fffbe474

Observation 46f28f8b-a6e8-48f7-bf91-c2189e23b6f3 · outbound

This paper cites Polypythias: Stability and outliers across fifty language model pre-training runs.

Tracing LLM Behavior to the Training Data with Empirical Next-Token Distributions Polypythias: Stability and outliers across fifty language model pre-training runs

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-04T01:45:18.971744Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:45:18.971744Z digest=sha256:86e34981a03212fbcfa2520710d5bede7e5ab94b683b59eff65358da2ace3cba

Observation 3676417d-e443-41f8-927c-9cc3d4a0c35f · outbound

This paper cites Jiasheng Ye, Peiju Liu, Tianxiang Sun, Jun Zhan, Yunhua Zhou, and Xipeng Qiu.

Tracing LLM Behavior to the Training Data with Empirical Next-Token Distributions Jiasheng Ye, Peiju Liu, Tianxiang Sun, Jun Zhan, Yunhua Zhou, and Xipeng Qiu

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-04T01:45:19.127391Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:45:19.127391Z digest=sha256:4d15ec8f762cb82e5ee18128ffd359d28ca3657282e81d32c1d8f1582f4d26b8

Observation 5838fa0b-fa07-42dd-8f97-28d6d0efeb9d · outbound

This paper cites Studying Large Language Model Generalization with Influence Functions.

Tracing LLM Behavior to the Training Data with Empirical Next-Token Distributions Studying Large Language Model Generalization with Influence Functions

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-04T01:45:18.619376Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:45:18.619376Z digest=sha256:2f5e9a81460db56255dbb4cd17a295cb2ec88b0fc21477bd3847ea8363e24235

Observation 13dc8bbc-f492-4229-9cb5-4ada2d9b8f04 · outbound

This paper cites Stealing Part of a Production Language Model.

Tracing LLM Behavior to the Training Data with Empirical Next-Token Distributions Stealing Part of a Production Language Model

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-04T01:45:18.338994Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:45:18.338994Z digest=sha256:687c68ab6ad347f207f386ad59ff27760632e29dd5aefa5db34bae7c6a368d5d

Observation 90fcf031-a76e-4a02-b970-b99ba6125a6a · outbound

This paper cites Language contamination helps explains the cross-lingual capabilities of english pretrained models.

Tracing LLM Behavior to the Training Data with Empirical Next-Token Distributions Language contamination helps explains the cross-lingual capabilities of english pretrained models

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-04T01:45:18.298168Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:45:18.298168Z digest=sha256:1e4d1f43fd73b781a2ad88d3080d26cf84aae2334daf42668b807fd67be0c68c

Observation c3aec7f0-926f-49ce-8ecf-62598d29d059 · outbound

This paper cites Pretraining data statistics shape the phases of learning entity comparison in language models.

Tracing LLM Behavior to the Training Data with Empirical Next-Token Distributions Pretraining data statistics shape the phases of learning entity comparison in language models

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-04T01:45:18.393265Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:45:18.393265Z digest=sha256:1057f5bd2432dfd16b1fd76ff017ba4f736b021dc33429e392b88161d163a43e

Observation 08e6c2ee-e8ef-488e-968a-145f1d8e271b · outbound

This paper cites Jason Wei, Dan Garrette, Tal Linzen, and Ellie Pavlick.

Tracing LLM Behavior to the Training Data with Empirical Next-Token Distributions Jason Wei, Dan Garrette, Tal Linzen, and Ellie Pavlick

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-04T01:45:19.044382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:45:19.044382Z digest=sha256:36aeaab4fb84ed1ddd087528485719d435d60e28cc83f73aa90155636580f507

Observation df60d6c8-5a36-4e5f-9d43-e4f19ada691b · outbound

This paper cites Measuring Causal Effects of Data Statistics on Language Model's `Factual' Predictions.

Tracing LLM Behavior to the Training Data with Empirical Next-Token Distributions Measuring Causal Effects of Data Statistics on Language Model's `Factual' Predictions

Reference 2026

Resolution
unresolved
no resolver link, observed 2026-08-04T01:45:18.458866Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:45:18.458866Z digest=sha256:7383bdeadb5bfb7de6ace71a01f889c48d73adcdae520f98d897ebe764d9b853

Pith citing papers

No inbound Pith citation observations are available.