Pith. sign in

Paper Citation Record · LEDGER

OLMoTrace: Tracing Language Model Outputs Back to Trillions of Training Tokens

As of 6 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 8 inbound Pith citation observations for arXiv:2504.07096.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2504.07096 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 8 of 8 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 8 of 8 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T20:44:59.409853Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation e1db5eac-f82c-4a69-995b-cd52412fb470 · inbound

Low-Perplexity LLM-Generated Sequences and Where To Find Them cites this paper.

Low-Perplexity LLM-Generated Sequences and Where To Find Them OLMoTrace: Tracing Language Model Outputs Back to Trillions of Training Tokens

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T20:44:59.409853Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T20:44:59.409853Z digest=sha256:422aee389f7c2e74cb18134c5c1070d7f96a10673c59833b39c35ad10ccabdd4

Observation 819ea96e-a4b0-4ee4-9f9a-5ba1be651a56 · inbound

"Lost-in-the-Later": Framework for Quantifying Contextual Grounding in Large Language Models cites this paper.

"Lost-in-the-Later": Framework for Quantifying Contextual Grounding in Large Language Models OLMoTrace: Tracing Language Model Outputs Back to Trillions of Training Tokens

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T19:32:20.702007Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:32:20.702007Z digest=sha256:1ee8243d0f548bbf3946574e8f57f20bcdd8ee2aa7c9f11eb48a6e4a6ddd4c2b

Observation bf831371-c2d7-4d56-9a17-20d0d8f0e699 · inbound

LLM generation novelty through the lens of semantic similarity cites this paper.

LLM generation novelty through the lens of semantic similarity OLMoTrace: Tracing Language Model Outputs Back to Trillions of Training Tokens

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-04T07:22:06.014745Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T07:22:06.014745Z digest=sha256:79386603b9d436016097e39bd38b3c30362330d2c6e9c34d1c4e4744b386feb4

Observation 469caa50-e1ec-44ba-8913-c999ae21a589 · inbound

Feature Starvation as Geometric Instability in Sparse Autoencoders cites this paper.

Feature Starvation as Geometric Instability in Sparse Autoencoders OLMoTrace: Tracing Language Model Outputs Back to Trillions of Training Tokens

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-08T20:09:10.079342Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-08T16:31:38.468164Z digest=sha256:1eae95df0d23ddefb0de1891b1f2bb6669b1d587576a171023ea3d7bd28fcf6d

Observation d26c2632-171e-4347-96d3-8bcdceb63399 · inbound

DataDignity: Training Data Attribution for Large Language Models cites this paper.

DataDignity: Training Data Attribution for Large Language Models OLMoTrace: Tracing Language Model Outputs Back to Trillions of Training Tokens

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-11T19:26:10.330691Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-08T11:53:19.594779Z digest=sha256:dc01fc315d5250671e3e6a7c95c93b36491b8e36430691fd8c8e3358a2703d1b

Observation bf5dce4d-4239-41ff-97f7-9adebfef0684 · inbound

Efficient and Scalable Provenance Tracking for LLM-Generated Code Snippets cites this paper.

Efficient and Scalable Provenance Tracking for LLM-Generated Code Snippets OLMoTrace: Tracing Language Model Outputs Back to Trillions of Training Tokens

Reference 15

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T10:53:19.746679Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-29T10:46:16.314823Z digest=sha256:c13c1325f61069fbd7a6f7e09c857099103d2a571a30f39cc3b4d7f7b0c52dd1

Observation 2b047b8d-8f1c-4481-954d-a0cfe957cf73 · inbound

LLMs Can Leak Training Data But Do They Want To? A Propensity-Aware Evaluation of Memorization in LLMs cites this paper.

LLMs Can Leak Training Data But Do They Want To? A Propensity-Aware Evaluation of Memorization in LLMs OLMoTrace: Tracing Language Model Outputs Back to Trillions of Training Tokens

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-07-02T12:56:57.470433Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-06-28T01:40:53.284131Z digest=sha256:c545bea96c2def235b473cfbb73a47a829cd8b9dadb484c025b04ea156080bb7

Observation 80756a46-a9e3-4c12-a284-1158ff03aba3 · inbound

RELIANCE: Curating and Evaluating Reproductive Health Information on Social Media cites this paper.

RELIANCE: Curating and Evaluating Reproductive Health Information on Social Media OLMoTrace: Tracing Language Model Outputs Back to Trillions of Training Tokens

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-07-03T13:28:19.276871Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-06-27T07:56:36.600805Z digest=sha256:fa8e01ef97cc6bd17f669b613e6ab341bc02ee777f466da9ba45148c11f7ce06