Pith. sign in

Paper Citation Record · LEDGER

NLP Evaluation in trouble: On the Need to Measure LLM Data Contamination for each Benchmark

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 20 inbound Pith citation observations for arXiv:2310.18018.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2310.18018 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 20 of 20 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 20 of 20 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T14:51:15.453440Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

8
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation d38930be-bc42-42ce-8362-ea6b0e0573df · inbound

Unbiased Evaluation of Large Language Models from a Causal Perspective cites this paper.

Unbiased Evaluation of Large Language Models from a Causal Perspective NLP Evaluation in trouble: On the Need to Measure LLM Data Contamination for each Benchmark

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-08T14:51:15.453440Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:51:15.453440Z digest=sha256:1748fe3d54fc587245382763de44d897f6583f9cfd0219e964f09ff69704927b

Observation 9171a74f-991e-4ff4-b18d-8ed9ca0db66c · inbound

RewardAnything: Generalizable Principle-Following Reward Models cites this paper.

RewardAnything: Generalizable Principle-Following Reward Models NLP Evaluation in trouble: On the Need to Measure LLM Data Contamination for each Benchmark

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:07.003448Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:04:07.003448Z digest=sha256:12026719930e786a50b63619732de9cc64ed8748e38e4f9ba0a30b1502c5cb50

Observation 6ff839f5-66d4-4b38-947c-22946783829d · inbound

Establishing Trustworthy LLM Evaluation via Shortcut Neuron Analysis cites this paper.

Establishing Trustworthy LLM Evaluation via Shortcut Neuron Analysis NLP Evaluation in trouble: On the Need to Measure LLM Data Contamination for each Benchmark

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T10:54:26.345390Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:54:26.345390Z digest=sha256:ed7dfee2c33d4a81836ef268dbe02faa3dbc224c849cfd02f257b40a3f3d2faa

Observation 67128382-2d9a-44be-83dd-fd934e1ae9bd · inbound

League of LLMs: A Benchmark-Free Paradigm for Mutual Evaluation of Large Language Models cites this paper.

League of LLMs: A Benchmark-Free Paradigm for Mutual Evaluation of Large Language Models NLP Evaluation in trouble: On the Need to Measure LLM Data Contamination for each Benchmark

Reference 36

Resolution
metadata mismatch
arxiv_id, observed 2026-05-19T03:22:01.356975Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-19T03:17:06.457421Z digest=sha256:346b60ba786f77f1bd22dfca4afb35290c7067e96e44a331f569ddce377a34a0

Observation 0f330cad-9e76-467d-95ef-727506a63474 · inbound

On the Fitness Landscape in the $NK$ Model cites this paper.

On the Fitness Landscape in the $NK$ Model NLP Evaluation in trouble: On the Need to Measure LLM Data Contamination for each Benchmark

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-05T19:28:54.967333Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:28:54.967333Z digest=sha256:b0434bf96687ad91fff35361bde6a5513d4a1a1150280626a30d682dc82b0c33

Observation 56c0d6cd-85ae-444a-96c4-7198018eabf1 · inbound

Artificial Phantasia: Emergent Mental Imagery in Large Language Models cites this paper.

Artificial Phantasia: Emergent Mental Imagery in Large Language Models NLP Evaluation in trouble: On the Need to Measure LLM Data Contamination for each Benchmark

Reference 72

Resolution
verified exact
arxiv_id, observed 2026-05-21T21:44:22.869054Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-21T21:41:39.111769Z digest=sha256:b6763da29fc3a1775f110800f625ce5190bceda7c74b68c78224ab7b7bde5786

Observation e5b6228e-fa83-4191-b596-db823287d718 · inbound

Correcting Prompt Dependence in LLM Benchmarks: A Bayesian Hierarchical Model with Embedding-Space Clustering cites this paper.

Correcting Prompt Dependence in LLM Benchmarks: A Bayesian Hierarchical Model with Embedding-Space Clustering NLP Evaluation in trouble: On the Need to Measure LLM Data Contamination for each Benchmark

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-04T11:19:49.706763Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:19:49.706763Z digest=sha256:8a1dac7d18893d344997fb6ebb8d1ba560f8ae8e68da45cf2c2348ceed9bd9e9

Observation 293cbe29-27d6-4e41-91e6-87ae083e33c7 · inbound

DecompSR: A dataset for decomposed analyses of compositional multihop spatial reasoning cites this paper.

DecompSR: A dataset for decomposed analyses of compositional multihop spatial reasoning NLP Evaluation in trouble: On the Need to Measure LLM Data Contamination for each Benchmark

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-18T01:20:34.337660Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-18T01:18:44.523602Z digest=sha256:1593eb5e557750c93c3b87fb8e9e53da473038d943b44fe02e30c0e9242788ee

Observation 5fe6ffde-e7d3-4b18-9162-da15dd263a19 · inbound

DecompSR: A dataset for decomposed analyses of compositional multihop spatial reasoning cites this paper.

DecompSR: A dataset for decomposed analyses of compositional multihop spatial reasoning NLP Evaluation in trouble: On the Need to Measure LLM Data Contamination for each Benchmark

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-04T00:11:51.863397Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T00:11:51.863397Z digest=sha256:b0c75234ddb83475613c9365cc589c839f6585a7187bac752b7f902df27a01c2

Observation d6048266-81af-4ed5-9c79-d42da57ac594 · inbound

When Generic Prompt Improvements Hurt: Evaluation-Driven Iteration for LLM Applications cites this paper.

When Generic Prompt Improvements Hurt: Evaluation-Driven Iteration for LLM Applications NLP Evaluation in trouble: On the Need to Measure LLM Data Contamination for each Benchmark

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-03T06:46:26.337035Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:46:26.337035Z digest=sha256:bae1c5780c9c1cb340fb4a5705be25b1b5ee69569fff0a578af89167a54442aa

Observation 14e051e7-4ac8-4f34-8c88-9452415addd0 · inbound

Assessing Capabilities of Large Language Models in Social Media Analytics: A Multi-task Quest cites this paper.

Assessing Capabilities of Large Language Models in Social Media Analytics: A Multi-task Quest NLP Evaluation in trouble: On the Need to Measure LLM Data Contamination for each Benchmark

Reference 74

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T12:41:01.820540Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-10T03:20:12.777878Z digest=sha256:7b510fee92ec5525722b9e7bd1f599ef9b529d35badaf43b3ffb037ca5fb61c9

Observation 0c62ac00-20a6-47ee-9c7c-338a0ba8f1a1 · inbound

ActuBench: A Multi-Agent LLM Pipeline for Generation and Evaluation of Actuarial Reasoning Tasks cites this paper.

ActuBench: A Multi-Agent LLM Pipeline for Generation and Evaluation of Actuarial Reasoning Tasks NLP Evaluation in trouble: On the Need to Measure LLM Data Contamination for each Benchmark

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:46:05.187194Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T00:35:24.397273Z digest=sha256:63eacb39131d52cef91269648c854d0aac4390f2aadcea20b509d5b86a8a0005

Observation 69286f6b-21dd-4371-9c4e-7f613ba6fab7 · inbound

Towards Evaluation Engineering: An Empirical Study of ML Evaluation Harnesses in the Wild cites this paper.

Towards Evaluation Engineering: An Empirical Study of ML Evaluation Harnesses in the Wild NLP Evaluation in trouble: On the Need to Measure LLM Data Contamination for each Benchmark

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-06-30T14:44:45.135213Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-30T14:41:07.354007Z digest=sha256:c78695e07307858723793d799f5ff17f28cb8e5c1a744947b1b17ea488086c65

Observation 66eb410d-85f7-4e23-b259-15edd22a1667 · inbound

Pretraining Data Exposure in Large Language Models: A Survey of Membership Inference, Data Contamination, and Security Implications cites this paper.

Pretraining Data Exposure in Large Language Models: A Survey of Membership Inference, Data Contamination, and Security Implications NLP Evaluation in trouble: On the Need to Measure LLM Data Contamination for each Benchmark

Reference 46

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T17:24:56.613744Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-30T17:20:16.735285Z digest=sha256:473571004fcbcef2533259197dd2f0bed0989aeb2381fe4db4b7a5366aa2c304

Observation 8bbe3091-2d39-4119-8607-e5f378ccfd67 · inbound

LiveK12Bench: Have Large Multimodal Models Truly Conquered High School-level Examinations? cites this paper.

LiveK12Bench: Have Large Multimodal Models Truly Conquered High School-level Examinations? NLP Evaluation in trouble: On the Need to Measure LLM Data Contamination for each Benchmark

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-06-29T17:33:45.019663Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-29T17:33:03.397468Z digest=sha256:5c68e213cacfd20b69074047d93db6210ad2556f418f380103cb52e547bbfed2

Observation 8ba959b7-1a87-4e52-8441-eaf435884bdf · inbound

The Case for Model Science: Verify, Explore, Steer, Refine cites this paper.

The Case for Model Science: Verify, Explore, Steer, Refine NLP Evaluation in trouble: On the Need to Measure LLM Data Contamination for each Benchmark

Reference 24

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T21:16:13.563798Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-28T17:24:32.311565Z digest=sha256:28bd54688890ad7c8462f160e6706fc0c8f9b17dd9276760189f95fdb720d975

Observation 66ad5f14-c0ad-47f2-80ba-698b5e72f080 · inbound

Dissecting model behavior through agent trajectories cites this paper.

Dissecting model behavior through agent trajectories NLP Evaluation in trouble: On the Need to Measure LLM Data Contamination for each Benchmark

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-07-03T20:18:56.917233Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-27T01:27:39.812496Z digest=sha256:59816fe874561622f43970a3abbe7f153d1125b6343f82a11a74812611746eda

Observation 17835370-b4a7-4b61-bee8-ff4db2e32631 · inbound

Testing Frontier Large Language Models' Physics Literacy in Parallel Physical Worlds cites this paper.

Testing Frontier Large Language Models' Physics Literacy in Parallel Physical Worlds NLP Evaluation in trouble: On the Need to Measure LLM Data Contamination for each Benchmark

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-07-02T19:27:18.679688Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-07-02T19:18:43.558804Z digest=sha256:9a6f3c2b8941e8004066bdfbc683349344e9847d46654ee257fb83c992b0c425

Observation 4482a391-6e89-4f76-9714-9d72df27b199 · inbound

Pre-Flight: A Benchmark for Evaluating Large Language Models on Aviation Operational Knowledge cites this paper.

Pre-Flight: A Benchmark for Evaluating Large Language Models on Aviation Operational Knowledge NLP Evaluation in trouble: On the Need to Measure LLM Data Contamination for each Benchmark

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T13:38:18.445574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-03T13:36:31.189451Z digest=sha256:88ea71315ba23d3a1100f1e07ab7e687b20ce69e6375e4ea7c7ecf16dde4816f

Observation f07d090d-375a-4152-ad26-b306a029ab47 · inbound

Information Discernment in Large Language Models cites this paper.

Information Discernment in Large Language Models NLP Evaluation in trouble: On the Need to Measure LLM Data Contamination for each Benchmark

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-02T13:26:26.237664Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:26:26.237664Z digest=sha256:3e08ca9f04c987faea3bbf91b737d3373eea38966b42e276cad5b95462278eac