Pith. sign in

Paper Citation Record · LEDGER

Training on the Benchmark Is Not All You Need

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 3 inbound Pith citation observations for arXiv:2409.01790.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2409.01790 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 3 of 3 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T14:51:15.442357Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-07T10:19:14.868572Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 11a417e7-a81b-403a-b639-4071f017c38e · inbound

Unbiased Evaluation of Large Language Models from a Causal Perspective cites this paper.

Unbiased Evaluation of Large Language Models from a Causal Perspective Training on the Benchmark Is Not All You Need

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-08T14:51:15.442357Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:51:15.442357Z digest=sha256:ca7d11dab725a0ab292d709ae92a02e77240e7fc0a1b7cb1f7648892c533888b

Observation b4c6546f-481d-4c0f-a8a6-1c9b078a59b3 · inbound

Simulating Training Data Leakage in Multiple-Choice Benchmarks for LLM Evaluation cites this paper.

Simulating Training Data Leakage in Multiple-Choice Benchmarks for LLM Evaluation Training on the Benchmark Is Not All You Need

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:27.501910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:27.501910Z digest=sha256:fcef29dd7c027d48dab7f0d4dc10ad896ebc053e5ccf5a2a9fedcd312319af17

Observation 5ee39aaa-cdd3-4af3-9b7b-a25b7cd89e8d · inbound

Loki's Dance of Illusions: A Comprehensive Survey of Hallucination in Large Language Models cites this paper.

Loki's Dance of Illusions: A Comprehensive Survey of Hallucination in Large Language Models Training on the Benchmark Is Not All You Need

Reference 168

Resolution
verified exact
local_arxiv, observed 2026-08-07T10:19:15.045306Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T10:19:07.723477Z digest=sha256:ba10b357d99efb4f20659f0c42f63ec25497bc98425f9cdfbbdddc26ad320e76