Pith. sign in

Paper Citation Record · LEDGER

Rethinking Benchmark and Contamination for Language Models with Rephrased Samples

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 42 inbound Pith citation observations for arXiv:2311.04850.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2311.04850 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 42 of 42 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 42 of 42 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T18:19:46.939354Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

7
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 26dde097-e60d-47b9-ba7b-9ca5f4315064 · inbound

Benchmark Data Contamination of Large Language Models: A Survey cites this paper.

Benchmark Data Contamination of Large Language Models: A Survey Rethinking Benchmark and Contamination for Language Models with Rephrased Samples

Reference 170

Resolution
verified exact
arxiv_id, observed 2026-05-22T23:10:41.134008Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-22T23:10:40.420241Z digest=sha256:27ee69d60e7822143570cdb5aef21073e9bccee202def6dfb8072477049c8bc5

Observation 79ffe552-61e5-4673-bf9b-bc4fde5114c9 · inbound

DataComp-LM: In search of the next generation of training sets for language models cites this paper.

DataComp-LM: In search of the next generation of training sets for language models Rethinking Benchmark and Contamination for Language Models with Rephrased Samples

Reference 208

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T22:58:17.334977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-17T22:58:16.523267Z digest=sha256:82bdb4effbe9b4991fb481a794437d74828744c15d72cbf7a89cb35a16d4897c

Observation d04fd98e-06bd-41e4-af68-e6b519a0ec47 · inbound

ChatGLM: A Family of Large Language Models from GLM-130B to GLM-4 All Tools cites this paper.

ChatGLM: A Family of Large Language Models from GLM-130B to GLM-4 All Tools Rethinking Benchmark and Contamination for Language Models with Rephrased Samples

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-05-11T08:08:09.653347Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-11T08:08:09.444352Z digest=sha256:42d64f70105690fd993f0ecf72c46acd03b0e44457561f5d330f029e5d4445c4

Observation 721c1b03-93f3-4b00-a4dc-c92cdd4ae55f · inbound

Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation cites this paper.

Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation Rethinking Benchmark and Contamination for Language Models with Rephrased Samples

Reference 136

Resolution
unresolved
no resolver link, observed 2026-08-08T15:06:55.384462Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T15:06:55.384462Z digest=sha256:6ab2547b8a0562ce8c9b67ff2969d8bcf1dfbb8b4ab32f7c1003d271bf615324

Observation 246859ee-fbaf-4caa-ba72-24d790e7d46a · inbound

Unbiased Evaluation of Large Language Models from a Causal Perspective cites this paper.

Unbiased Evaluation of Large Language Models from a Causal Perspective Rethinking Benchmark and Contamination for Language Models with Rephrased Samples

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-08T14:51:15.489455Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:51:15.489455Z digest=sha256:c033cdeece96aa5d7458054f04fd802826b0ce7255677ac71d8611b7bc1d1a19

Observation 775566c9-a666-4c3c-a592-902782fcb963 · inbound

CapBencher: Give Your LLM Benchmark a Built-in Alarm for Test-Set Overfitting cites this paper.

CapBencher: Give Your LLM Benchmark a Built-in Alarm for Test-Set Overfitting Rethinking Benchmark and Contamination for Language Models with Rephrased Samples

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T14:44:35.529969Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:44:35.529969Z digest=sha256:bdec4b7fbc1408b5c475a73ba02a469f151778833dc7afda5a472c4aa5b9a44c

Observation fb1ad3e0-5e03-4ab1-aeff-d7c4f5f9f56d · inbound

SwiftEval: Developing a Language-Specific Benchmark for LLM-generated Code Evaluation cites this paper.

SwiftEval: Developing a Language-Specific Benchmark for LLM-generated Code Evaluation Rethinking Benchmark and Contamination for Language Models with Rephrased Samples

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T12:29:18.525978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:29:18.525978Z digest=sha256:45845cdbd096c0f51c412856f84799d97b9b31a386eb4e52e90574cabe8ae93a

Observation 0a0588e0-0c37-4bd2-95fb-f1c7e168d36e · inbound

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation cites this paper.

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation Rethinking Benchmark and Contamination for Language Models with Rephrased Samples

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-07T12:08:17.369531Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:08:17.369531Z digest=sha256:a100d5d481d4c7e71944c2b0048e215e196234a6f0aadace65e23db2e46ca261

Observation b690f495-6cb2-4638-a9e5-a77112b77557 · inbound

OpenCodeReasoning-II: A Simple Test Time Scaling Approach via Self-Critique cites this paper.

OpenCodeReasoning-II: A Simple Test Time Scaling Approach via Self-Critique Rethinking Benchmark and Contamination for Language Models with Rephrased Samples

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-06T18:11:20.143746Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:11:20.143746Z digest=sha256:e19c063c8ebfe17aa637b29ae79f1fdfb8f24993f26c62f9c00cdcfaf7ee2b03

Observation 1bb22ce2-c68b-422d-948d-2f1f77c32e39 · inbound

League of LLMs: A Benchmark-Free Paradigm for Mutual Evaluation of Large Language Models cites this paper.

League of LLMs: A Benchmark-Free Paradigm for Mutual Evaluation of Large Language Models Rethinking Benchmark and Contamination for Language Models with Rephrased Samples

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-05-19T03:22:01.341453Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-19T03:17:06.457421Z digest=sha256:aa48e521ef2e75233b0a1e42b74756a0a0bb4f26870b1546d8ee7f3ab84e960d

Observation 8fca8ffb-2cd4-4adf-b168-b5219a677eb7 · inbound

Llama-3.1-FoundationAI-SecurityLLM-8B-Instruct Technical Report cites this paper.

Llama-3.1-FoundationAI-SecurityLLM-8B-Instruct Technical Report Rethinking Benchmark and Contamination for Language Models with Rephrased Samples

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-06T05:57:29.511258Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:57:29.511258Z digest=sha256:8aefd4bc3ac85309d3b9e03e35a2fee0a93041349f0399633af474576437cb5e

Observation 54c99d7c-0932-4665-88df-4cafaffd7561 · inbound

LogitTrace: Detecting Benchmark Contamination via Layerwise Logit Trajectories cites this paper.

LogitTrace: Detecting Benchmark Contamination via Layerwise Logit Trajectories Rethinking Benchmark and Contamination for Language Models with Rephrased Samples

Reference 26

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T14:36:28.943124Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-18T14:32:57.426996Z digest=sha256:d82d3b79979094a09b360a66ab4a1544df3327b396a71f4be5db377d1c065502

Observation 1cf0bda5-0b6e-4ddb-b0ff-0793c0142615 · inbound

Correcting Prompt Dependence in LLM Benchmarks: A Bayesian Hierarchical Model with Embedding-Space Clustering cites this paper.

Correcting Prompt Dependence in LLM Benchmarks: A Bayesian Hierarchical Model with Embedding-Space Clustering Rethinking Benchmark and Contamination for Language Models with Rephrased Samples

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-04T11:19:51.012398Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:19:51.012398Z digest=sha256:870270dad037d2e5a4a31fd19904a046152e06490965957e212489cda47cacde

Observation d25e5c67-f333-4fda-b471-faf6e2b0bdff · inbound

Compiled AI: Deterministic Code Generation for LLM-Based Workflow Automation cites this paper.

Compiled AI: Deterministic Code Generation for LLM-Based Workflow Automation Rethinking Benchmark and Contamination for Language Models with Rephrased Samples

Reference 22

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T23:40:51.520975Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T18:57:34.386632Z digest=sha256:94e21e755bec43ee98fca65037f8e0a80e46ddef5f2bedf43acfd261a7077c95

Observation 548a2693-0ac5-442b-aac0-89d3bec954ea · inbound

DRBENCHER: Can Your Agent Identify the Entity, Retrieve Its Properties and Do the Math? cites this paper.

DRBENCHER: Can Your Agent Identify the Entity, Retrieve Its Properties and Do the Math? Rethinking Benchmark and Contamination for Language Models with Rephrased Samples

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-11T08:11:04.832167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T16:46:24.780270Z digest=sha256:3c7bba52260c7eae44e82a9bc0ca0bc07c30947796c6bc7eb818c738403e4b8a

Observation 47e2479c-fb1f-496d-827e-eaaf883b2921 · inbound

TPS-CalcBench: A Benchmark and Diagnostic Evaluation Framework for LLM Analytical Calculation Competence in Hypersonic Thermal Protection System Engineering cites this paper.

TPS-CalcBench: A Benchmark and Diagnostic Evaluation Framework for LLM Analytical Calculation Competence in Hypersonic Thermal Protection System Engineering Rethinking Benchmark and Contamination for Language Models with Rephrased Samples

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-05-10T09:23:37.514127Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-10T05:25:59.343884Z digest=sha256:8bf0b01c5933c4bbf0206d1365b347c94a6e185e66c196aec73a2b02d69863be

Observation 22b7cd4d-8090-460d-9758-5f01a02de39b · inbound

When Prompt Under-Specification Improves Code Correctness: An Exploratory Study of Prompt Wording and Structure Effects on LLM-Based Code Generation cites this paper.

When Prompt Under-Specification Improves Code Correctness: An Exploratory Study of Prompt Wording and Structure Effects on LLM-Based Code Generation Rethinking Benchmark and Contamination for Language Models with Rephrased Samples

Reference 44

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T22:17:06.819446Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-08T03:00:26.137401Z digest=sha256:38991412db803d661dfb8dcdd91bd3dabfd1cbb0118906857126626b7c105f7b

Observation 6fd5b37e-9661-4f13-97f0-0af6e7d0b045 · inbound

Measuring AI Reasoning: A Guide for Researchers cites this paper.

Measuring AI Reasoning: A Guide for Researchers Rethinking Benchmark and Contamination for Language Models with Rephrased Samples

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-05-09T06:05:36.959600Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-08T18:53:18.586923Z digest=sha256:4b6ca5df16a5395c0d31d6c2456ca88a47401b88ba7cdb91579ec33853b22361

Observation 1b15f2a6-e35a-4bde-a1ba-9bf71adfead8 · inbound

Agent Island: A Saturation- and Contamination-Resistant Benchmark from Multiagent Games cites this paper.

Agent Island: A Saturation- and Contamination-Resistant Benchmark from Multiagent Games Rethinking Benchmark and Contamination for Language Models with Rephrased Samples

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-11T17:46:17.276350Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-08T17:06:32.814188Z digest=sha256:f2617fc5a16ebe22cb8db1b41c95c36d767e20414d01c5132176b8220839554f

Observation 9d53cd78-ddd4-4c96-9571-3b76c9fa1010 · inbound

Measuring Evaluation-Context Divergence in Open-Weight LLMs: A Paired-Prompt Protocol with Pilot Evidence of Alignment-Pipeline-Specific Heterogeneity cites this paper.

Measuring Evaluation-Context Divergence in Open-Weight LLMs: A Paired-Prompt Protocol with Pilot Evidence of Alignment-Pipeline-Specific Heterogeneity Rethinking Benchmark and Contamination for Language Models with Rephrased Samples

Reference 18

Resolution
metadata mismatch
arxiv_id, observed 2026-05-08T22:04:18.090953Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-08T10:23:02.697982Z digest=sha256:b49e21a6fc8dd9c1eb6847bbb2e6f8962ca060c3d5febf267642668d12b5aed3

Observation 1c5fe167-2a6f-4ffc-b264-cb3a266a33fe · inbound

GSM-SEM: Benchmark and Framework for Generating Semantically Variant Augmentations cites this paper.

GSM-SEM: Benchmark and Framework for Generating Semantically Variant Augmentations Rethinking Benchmark and Contamination for Language Models with Rephrased Samples

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-11T04:55:59.899239Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-11T00:57:55.616036Z digest=sha256:8c80dc3d852bd5b03a9a6cc6012c7db99f4ef2f9a0f7637e1c58e76047a28a25

Observation bcaeda91-673a-474a-8963-2564111db7eb · inbound

GSM-SEM: Benchmark and Framework for Generating Semantically Variant Augmentations cites this paper.

GSM-SEM: Benchmark and Framework for Generating Semantically Variant Augmentations Rethinking Benchmark and Contamination for Language Models with Rephrased Samples

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-06-30T23:45:07.976186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-30T23:42:00.965554Z digest=sha256:9e38d8c568269a6cbf81727365f55edb457fc51f7f21208bdb36b7233cceb511

Observation ec4241e4-e7c0-4b4c-8c06-4660896cd370 · inbound

Can Agent Benchmarks Support Their Scores? Evidence-Supported Bounds for Interactive-Agent Evaluation cites this paper.

Can Agent Benchmarks Support Their Scores? Evidence-Supported Bounds for Interactive-Agent Evaluation Rethinking Benchmark and Contamination for Language Models with Rephrased Samples

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-12T05:41:23.900869Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-12T05:05:55.592359Z digest=sha256:0a161bcaa59cf2fac81c57b25788ec27c53f4abcde0da79affb60ddd5541902b

Observation abfb013e-e714-4037-85ac-acc881089016 · inbound

Decaf: Improving Neural Decompilation with Automatic Feedback and Search cites this paper.

Decaf: Improving Neural Decompilation with Automatic Feedback and Search Rethinking Benchmark and Contamination for Language Models with Rephrased Samples

Reference 32

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T02:07:08.093098Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-13T02:03:19.836506Z digest=sha256:7ae1f9518105d23001e903f20f744a7947a72e5d2fd6916676052f93d7dc0eb0

Observation 481aa828-3a0e-498f-9f0b-f56df90e7622 · inbound

Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack cites this paper.

Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack Rethinking Benchmark and Contamination for Language Models with Rephrased Samples

Reference 60

Resolution
verified exact
arxiv_id, observed 2026-05-14T20:32:56.862291Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-14T20:31:50.043920Z digest=sha256:4dc0ed02bb8f754652f4c823344cf698eb0a1f161ac3c03176f33a4f8e5e314f

Observation a0f175ac-6597-48d7-9bf4-51fc3977f5ec · inbound

LLM Benchmark Datasets Should Be Contamination-Resistant cites this paper.

LLM Benchmark Datasets Should Be Contamination-Resistant Rethinking Benchmark and Contamination for Language Models with Rephrased Samples

Reference 90

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T07:23:07.058655Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-20T07:19:50.354875Z digest=sha256:c9b1dc6e1828e8893d382d4c7727d7c4b27ab43eaa8fc0536054e23f067435b1

Observation ae61037a-4a0c-4149-9764-5047fc27a22e · inbound

Provable Joint Decontamination for Benchmarking Multiple Large Language Models cites this paper.

Provable Joint Decontamination for Benchmarking Multiple Large Language Models Rethinking Benchmark and Contamination for Language Models with Rephrased Samples

Reference 172

Resolution
metadata mismatch
arxiv_id, observed 2026-05-22T00:44:29.489044Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-22T00:40:54.038367Z digest=sha256:b50162510986c5069db2cba8f2fea97d1af4a119009398977fe9f3d7ed476837

Observation e50e38d5-a330-4082-9fc1-b47897ecc449 · inbound

The Illusion of Reasoning: Exposing Evasive Data Contamination in LLMs via Zero-CoT Truncation cites this paper.

The Illusion of Reasoning: Exposing Evasive Data Contamination in LLMs via Zero-CoT Truncation Rethinking Benchmark and Contamination for Language Models with Rephrased Samples

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-05-22T08:06:15.310194Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-22T08:05:42.212459Z digest=sha256:beff191f86e3b61b45f91accfae1237772a74a532b8093541e43e99b794b1df0

Observation 81a8cf3c-06b8-49b8-bbd7-3da728bad548 · inbound

How Hard is it to Rig a Benchmark? A Social Choice Analysis of Leaderboard Robustness cites this paper.

How Hard is it to Rig a Benchmark? A Social Choice Analysis of Leaderboard Robustness Rethinking Benchmark and Contamination for Language Models with Rephrased Samples

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-25T04:40:24.444660Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-25T04:37:01.537128Z digest=sha256:3d2c5d2dd4a8a9675ff347632fc3b820e2d22ae066ce3afae40c77ada37bad5a

Observation e11c3d40-017b-429b-8690-ff2370e45d98 · inbound

TRACER: A Semantic-Aware Framework for Fine-Grained Contamination Detection in Code LLMs cites this paper.

TRACER: A Semantic-Aware Framework for Fine-Grained Contamination Detection in Code LLMs Rethinking Benchmark and Contamination for Language Models with Rephrased Samples

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-06-30T15:34:47.895008Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-30T15:32:05.976335Z digest=sha256:f05d91dee7a88f7077afbb7327323dab5b2c0b65ee5f05acb929882e74ee9387

Observation 04ea9e21-1c59-4575-a510-f812fd2b7aeb · inbound

Towards Evaluation Engineering: An Empirical Study of ML Evaluation Harnesses in the Wild cites this paper.

Towards Evaluation Engineering: An Empirical Study of ML Evaluation Harnesses in the Wild Rethinking Benchmark and Contamination for Language Models with Rephrased Samples

Reference 55

Resolution
verified exact
arxiv_id, observed 2026-06-30T14:44:45.155954Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-30T14:41:07.354007Z digest=sha256:855e9099c812204b90d29d8844d82eef81a1cf75709d669046f4d2955cf6f003

Observation 6c2dc137-5f7d-4787-b49a-71cbffd49ef8 · inbound

Measuring Reasoning Quality in LLMs: A Multi-Dimensional Behavioral Framework cites this paper.

Measuring Reasoning Quality in LLMs: A Multi-Dimensional Behavioral Framework Rethinking Benchmark and Contamination for Language Models with Rephrased Samples

Reference 26

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T13:34:40.588543Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-30T13:27:50.367497Z digest=sha256:093079d6bbb9b8fdfdb925947ef5a630c5a17daefcda99043a956389ab868178

Observation b5a659b1-05df-4a48-8f38-913a67b3cfe8 · inbound

Measuring Reasoning Quality in LLMs: A Multi-Dimensional Behavioral Framework cites this paper.

Measuring Reasoning Quality in LLMs: A Multi-Dimensional Behavioral Framework Rethinking Benchmark and Contamination for Language Models with Rephrased Samples

Reference 26

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T08:05:31.914578Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-01T07:35:51.017797Z digest=sha256:3f5bb1462efef0f4dc30e20d583e2cb8f4a48612516757b2e2c8a678cf020437

Observation 73c90e48-132c-4c28-9ee1-a4949e9206f0 · inbound

Measuring Reasoning Quality in LLMs: A Multi-Dimensional Behavioral Framework cites this paper.

Measuring Reasoning Quality in LLMs: A Multi-Dimensional Behavioral Framework Rethinking Benchmark and Contamination for Language Models with Rephrased Samples

Reference 26

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T23:27:26.731890Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-02T23:22:08.567841Z digest=sha256:687d3ecba2cee12fbd32e1b37ac2ebd87c2cd50e57b1da75d99742ca7b5cf91d

Observation dab2454a-666d-4842-b07a-aed17edf3fd0 · inbound

TSFMAudit: Data Contamination Auditing in Forecasting Time Series Foundation Models cites this paper.

TSFMAudit: Data Contamination Auditing in Forecasting Time Series Foundation Models Rethinking Benchmark and Contamination for Language Models with Rephrased Samples

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-06-30T12:04:38.780577Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-30T12:00:05.807004Z digest=sha256:7be51c87a2021504abf236ba2c5ea502b50606961808d2e29f0c08fa7ec789b1

Observation 429ec9d5-66e8-4def-a403-8144313ea449 · inbound

Embodied-BenchClaw: An Autonomous Multi-Agent System for Embodied Spatial Intelligence Benchmark Construction cites this paper.

Embodied-BenchClaw: An Autonomous Multi-Agent System for Embodied Spatial Intelligence Benchmark Construction Rethinking Benchmark and Contamination for Language Models with Rephrased Samples

Reference 13

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T10:48:02.880656Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-27T09:48:49.786021Z digest=sha256:c4e432bc8b55d25c059ba36aba27cfa34ccb0aaa2602220cb984ced2971e449c

Observation dcbed4ea-62a8-48a1-b727-be0d3cb6004e · inbound

Which Models Are Our Models Built On? Auditing Invisible Dependencies in Modern LLMs cites this paper.

Which Models Are Our Models Built On? Auditing Invisible Dependencies in Modern LLMs Rethinking Benchmark and Contamination for Language Models with Rephrased Samples

Reference 75

Resolution
verified exact
arxiv_id, observed 2026-07-03T10:37:56.593959Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-27T09:57:14.328157Z digest=sha256:307567bc55039660729ae969a8decfbcb93462d1e042e8c227f633e3cc03b6d6

Observation df54faea-0bd2-4570-a54c-c347743948f7 · inbound

SrDetection: A Self-Referential Framework for Data Leakage Detection in Code Large Language Models cites this paper.

SrDetection: A Self-Referential Framework for Data Leakage Detection in Code Large Language Models Rethinking Benchmark and Contamination for Language Models with Rephrased Samples

Reference 12

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T06:24:18.367339Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-30T06:20:02.046020Z digest=sha256:46b45eeba56b88ef5765977a5efe3e3760fa40202d7d2ccceaf157fc7d49eb75

Observation e34d5daa-ff82-4dcf-8c9b-19274d08c5d6 · inbound

Do Agent Benchmarks Measure Capability? Protocol Validity in the Age of Agentic AI cites this paper.

Do Agent Benchmarks Measure Capability? Protocol Validity in the Age of Agentic AI Rethinking Benchmark and Contamination for Language Models with Rephrased Samples

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-01T05:03:42.877641Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T05:03:42.877641Z digest=sha256:d73ca63694a1346de5b59e9887e7d7d3a85512eac999cbdb2fe0c976ed0d51e9

Observation 372bad00-77da-4547-981f-6351c5a94418 · inbound

Temporal Leakage in LLM Backtesting: Measurement, Validation, and Adjusted Scores cites this paper.

Temporal Leakage in LLM Backtesting: Measurement, Validation, and Adjusted Scores Rethinking Benchmark and Contamination for Language Models with Rephrased Samples

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-08T04:31:04.418737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T04:31:04.418737Z digest=sha256:b0dc7c8e0ae0f24bc53991cda7c8356d25a71b7e6833a62bbc712379a4983b44

Observation f312928c-1d49-415a-bc1e-92f5c4a57e6c · inbound

EuroExec: Frontier Language Models Fall Short of Expert Judgment on European Executive Decision Tasks cites this paper.

EuroExec: Frontier Language Models Fall Short of Expert Judgment on European Executive Decision Tasks Rethinking Benchmark and Contamination for Language Models with Rephrased Samples

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-06T22:23:05.614586Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:23:05.614586Z digest=sha256:c7ae6be237b1cef8e7d80942db4e81e739378d30079870a7e80dbfab79e1ec77

Observation 1fd6bd95-accd-4fa4-b360-c9160579218f · inbound

EuroExec: Frontier Language Models Fall Short of Expert Judgment on European Executive Decision Tasks cites this paper.

EuroExec: Frontier Language Models Fall Short of Expert Judgment on European Executive Decision Tasks Rethinking Benchmark and Contamination for Language Models with Rephrased Samples

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-08T18:19:46.939354Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T18:19:46.939354Z digest=sha256:a505e90dd17517623ae3afb6dec7d53984530b92284982b108fa749628d1c16f