Pith. sign in

Paper Citation Record · LEDGER

Rethinking Benchmark and Contamination for Language Models with Rephrased Samples

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 42 inbound Pith citation observations for arXiv:2311.04850.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2311.04850 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 42 of 42 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 42 of 42 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T18:19:46.939354Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

7
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 26dde097-e60d-47b9-ba7b-9ca5f4315064 · inbound

Benchmark Data Contamination of Large Language Models: A Survey cites this paper.

Benchmark Data Contamination of Large Language Models: A Survey Rethinking Benchmark and Contamination for Language Models with Rephrased Samples

Reference 170

Resolution
verified exact
arxiv_id, observed 2026-05-22T23:10:41.134008Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-22T23:10:40.420241Z digest=sha256:ab489aa501aafffd31cf1959f3e573bc17f234188e4e9e59fa962b09b357f9e3

Observation 79ffe552-61e5-4673-bf9b-bc4fde5114c9 · inbound

DataComp-LM: In search of the next generation of training sets for language models cites this paper.

DataComp-LM: In search of the next generation of training sets for language models Rethinking Benchmark and Contamination for Language Models with Rephrased Samples

Reference 208

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T22:58:17.334977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-17T22:58:16.523267Z digest=sha256:979cf082e694acedba3e8d11c029a3798229c71069fb85bd82a9004dd6767283

Observation d04fd98e-06bd-41e4-af68-e6b519a0ec47 · inbound

ChatGLM: A Family of Large Language Models from GLM-130B to GLM-4 All Tools cites this paper.

ChatGLM: A Family of Large Language Models from GLM-130B to GLM-4 All Tools Rethinking Benchmark and Contamination for Language Models with Rephrased Samples

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-05-11T08:08:09.653347Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-11T08:08:09.444352Z digest=sha256:578a505cc8f9088aaf1edc718790865a9640693b8772a3b67db4aa0df4695235

Observation 721c1b03-93f3-4b00-a4dc-c92cdd4ae55f · inbound

Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation cites this paper.

Can We Trust AI Benchmarks? An Interdisciplinary Review of Current Issues in AI Evaluation Rethinking Benchmark and Contamination for Language Models with Rephrased Samples

Reference 136

Resolution
unresolved
no resolver link, observed 2026-08-08T15:06:55.384462Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T15:06:55.384462Z digest=sha256:6ab2547b8a0562ce8c9b67ff2969d8bcf1dfbb8b4ab32f7c1003d271bf615324

Observation 246859ee-fbaf-4caa-ba72-24d790e7d46a · inbound

Unbiased Evaluation of Large Language Models from a Causal Perspective cites this paper.

Unbiased Evaluation of Large Language Models from a Causal Perspective Rethinking Benchmark and Contamination for Language Models with Rephrased Samples

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-08T14:51:15.489455Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:51:15.489455Z digest=sha256:63a8a43706e92ef9c2cf5d613240317fb37f18b3ef9512911589a470243931bb

Observation 775566c9-a666-4c3c-a592-902782fcb963 · inbound

CapBencher: Give Your LLM Benchmark a Built-in Alarm for Test-Set Overfitting cites this paper.

CapBencher: Give Your LLM Benchmark a Built-in Alarm for Test-Set Overfitting Rethinking Benchmark and Contamination for Language Models with Rephrased Samples

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T14:44:35.529969Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:44:35.529969Z digest=sha256:bdec4b7fbc1408b5c475a73ba02a469f151778833dc7afda5a472c4aa5b9a44c

Observation fb1ad3e0-5e03-4ab1-aeff-d7c4f5f9f56d · inbound

SwiftEval: Developing a Language-Specific Benchmark for LLM-generated Code Evaluation cites this paper.

SwiftEval: Developing a Language-Specific Benchmark for LLM-generated Code Evaluation Rethinking Benchmark and Contamination for Language Models with Rephrased Samples

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T12:29:18.525978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:29:18.525978Z digest=sha256:45845cdbd096c0f51c412856f84799d97b9b31a386eb4e52e90574cabe8ae93a

Observation 0a0588e0-0c37-4bd2-95fb-f1c7e168d36e · inbound

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation cites this paper.

BenchHub: A Unified Benchmark Suite for Holistic and Customizable LLM Evaluation Rethinking Benchmark and Contamination for Language Models with Rephrased Samples

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-07T12:08:17.369531Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:08:17.369531Z digest=sha256:a100d5d481d4c7e71944c2b0048e215e196234a6f0aadace65e23db2e46ca261

Observation b690f495-6cb2-4638-a9e5-a77112b77557 · inbound

OpenCodeReasoning-II: A Simple Test Time Scaling Approach via Self-Critique cites this paper.

OpenCodeReasoning-II: A Simple Test Time Scaling Approach via Self-Critique Rethinking Benchmark and Contamination for Language Models with Rephrased Samples

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-06T18:11:20.143746Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:11:20.143746Z digest=sha256:e19c063c8ebfe17aa637b29ae79f1fdfb8f24993f26c62f9c00cdcfaf7ee2b03

Observation 1bb22ce2-c68b-422d-948d-2f1f77c32e39 · inbound

League of LLMs: A Benchmark-Free Paradigm for Mutual Evaluation of Large Language Models cites this paper.

League of LLMs: A Benchmark-Free Paradigm for Mutual Evaluation of Large Language Models Rethinking Benchmark and Contamination for Language Models with Rephrased Samples

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-05-19T03:22:01.341453Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-19T03:17:06.457421Z digest=sha256:3922873e5a9f7e11b52f01ad6397aa02e1175299650a571e3682cde53af2e4ee

Observation 8fca8ffb-2cd4-4adf-b168-b5219a677eb7 · inbound

Llama-3.1-FoundationAI-SecurityLLM-8B-Instruct Technical Report cites this paper.

Llama-3.1-FoundationAI-SecurityLLM-8B-Instruct Technical Report Rethinking Benchmark and Contamination for Language Models with Rephrased Samples

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-06T05:57:29.511258Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:57:29.511258Z digest=sha256:8aefd4bc3ac85309d3b9e03e35a2fee0a93041349f0399633af474576437cb5e

Observation 54c99d7c-0932-4665-88df-4cafaffd7561 · inbound

LogitTrace: Detecting Benchmark Contamination via Layerwise Logit Trajectories cites this paper.

LogitTrace: Detecting Benchmark Contamination via Layerwise Logit Trajectories Rethinking Benchmark and Contamination for Language Models with Rephrased Samples

Reference 26

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T14:36:28.943124Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-18T14:32:57.426996Z digest=sha256:c0cf0771b022366f676bd0e83a4d855b98c8dc0561f2e3f1b403f421eeabe41a

Observation 1cf0bda5-0b6e-4ddb-b0ff-0793c0142615 · inbound

Correcting Prompt Dependence in LLM Benchmarks: A Bayesian Hierarchical Model with Embedding-Space Clustering cites this paper.

Correcting Prompt Dependence in LLM Benchmarks: A Bayesian Hierarchical Model with Embedding-Space Clustering Rethinking Benchmark and Contamination for Language Models with Rephrased Samples

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-04T11:19:51.012398Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:19:51.012398Z digest=sha256:870270dad037d2e5a4a31fd19904a046152e06490965957e212489cda47cacde

Observation d25e5c67-f333-4fda-b471-faf6e2b0bdff · inbound

Compiled AI: Deterministic Code Generation for LLM-Based Workflow Automation cites this paper.

Compiled AI: Deterministic Code Generation for LLM-Based Workflow Automation Rethinking Benchmark and Contamination for Language Models with Rephrased Samples

Reference 22

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T23:40:51.520975Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T18:57:34.386632Z digest=sha256:a54c665ab02dc18cc5dda022221abc0653f9cc8d0aa86a32ea350c4fbaa85faa

Observation 548a2693-0ac5-442b-aac0-89d3bec954ea · inbound

DRBENCHER: Can Your Agent Identify the Entity, Retrieve Its Properties and Do the Math? cites this paper.

DRBENCHER: Can Your Agent Identify the Entity, Retrieve Its Properties and Do the Math? Rethinking Benchmark and Contamination for Language Models with Rephrased Samples

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-11T08:11:04.832167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T16:46:24.780270Z digest=sha256:64d65ea9be08ff9903e6dc991d7e18598ece4db76e4acb9a971da0992feaa75d

Observation 47e2479c-fb1f-496d-827e-eaaf883b2921 · inbound

TPS-CalcBench: A Benchmark and Diagnostic Evaluation Framework for LLM Analytical Calculation Competence in Hypersonic Thermal Protection System Engineering cites this paper.

TPS-CalcBench: A Benchmark and Diagnostic Evaluation Framework for LLM Analytical Calculation Competence in Hypersonic Thermal Protection System Engineering Rethinking Benchmark and Contamination for Language Models with Rephrased Samples

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-05-10T09:23:37.514127Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-10T05:25:59.343884Z digest=sha256:153a5b0958ba37b0c783baf902268cc5a74d6b71a124c4a3793f5bc0fb41b265

Observation 22b7cd4d-8090-460d-9758-5f01a02de39b · inbound

When Prompt Under-Specification Improves Code Correctness: An Exploratory Study of Prompt Wording and Structure Effects on LLM-Based Code Generation cites this paper.

When Prompt Under-Specification Improves Code Correctness: An Exploratory Study of Prompt Wording and Structure Effects on LLM-Based Code Generation Rethinking Benchmark and Contamination for Language Models with Rephrased Samples

Reference 44

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T22:17:06.819446Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-08T03:00:26.137401Z digest=sha256:be9bca9e39f4e363b80fb1584f53a356766dfa48a4137aec4966040da534102c

Observation 6fd5b37e-9661-4f13-97f0-0af6e7d0b045 · inbound

Measuring AI Reasoning: A Guide for Researchers cites this paper.

Measuring AI Reasoning: A Guide for Researchers Rethinking Benchmark and Contamination for Language Models with Rephrased Samples

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-05-09T06:05:36.959600Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-08T18:53:18.586923Z digest=sha256:47ed6dd7b0c33c20d891095680e819e9848673084101402dfa4140b8af7c7079

Observation 1b15f2a6-e35a-4bde-a1ba-9bf71adfead8 · inbound

Agent Island: A Saturation- and Contamination-Resistant Benchmark from Multiagent Games cites this paper.

Agent Island: A Saturation- and Contamination-Resistant Benchmark from Multiagent Games Rethinking Benchmark and Contamination for Language Models with Rephrased Samples

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-11T17:46:17.276350Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-08T17:06:32.814188Z digest=sha256:526b974876176077bf3c2dbcee551e048eafccf545641c9c4602e1ec9588cee7

Observation 9d53cd78-ddd4-4c96-9571-3b76c9fa1010 · inbound

Measuring Evaluation-Context Divergence in Open-Weight LLMs: A Paired-Prompt Protocol with Pilot Evidence of Alignment-Pipeline-Specific Heterogeneity cites this paper.

Measuring Evaluation-Context Divergence in Open-Weight LLMs: A Paired-Prompt Protocol with Pilot Evidence of Alignment-Pipeline-Specific Heterogeneity Rethinking Benchmark and Contamination for Language Models with Rephrased Samples

Reference 18

Resolution
metadata mismatch
arxiv_id, observed 2026-05-08T22:04:18.090953Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-08T10:23:02.697982Z digest=sha256:f80d4502af0ed1922d60be23a74bb508ad11ea81b42b99a9c28f1cb1e10da566

Observation 1c5fe167-2a6f-4ffc-b264-cb3a266a33fe · inbound

GSM-SEM: Benchmark and Framework for Generating Semantically Variant Augmentations cites this paper.

GSM-SEM: Benchmark and Framework for Generating Semantically Variant Augmentations Rethinking Benchmark and Contamination for Language Models with Rephrased Samples

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-11T04:55:59.899239Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-11T00:57:55.616036Z digest=sha256:248ceda930dbf88eeaa3ea2db2b697e0eb438fda07cc3b5d7408ba93f985854d

Observation bcaeda91-673a-474a-8963-2564111db7eb · inbound

GSM-SEM: Benchmark and Framework for Generating Semantically Variant Augmentations cites this paper.

GSM-SEM: Benchmark and Framework for Generating Semantically Variant Augmentations Rethinking Benchmark and Contamination for Language Models with Rephrased Samples

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-06-30T23:45:07.976186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-30T23:42:00.965554Z digest=sha256:c4f16ff979c0e323c885bc4e9925dd167ca2cac6eaf62eae13fb28cb0b20a694

Observation ec4241e4-e7c0-4b4c-8c06-4660896cd370 · inbound

Can Agent Benchmarks Support Their Scores? Evidence-Supported Bounds for Interactive-Agent Evaluation cites this paper.

Can Agent Benchmarks Support Their Scores? Evidence-Supported Bounds for Interactive-Agent Evaluation Rethinking Benchmark and Contamination for Language Models with Rephrased Samples

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-12T05:41:23.900869Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-12T05:05:55.592359Z digest=sha256:dea67df1e2036ad1299fef36b0dc7d21c18e72ded07bf6838479f16adc70ed36

Observation abfb013e-e714-4037-85ac-acc881089016 · inbound

Decaf: Improving Neural Decompilation with Automatic Feedback and Search cites this paper.

Decaf: Improving Neural Decompilation with Automatic Feedback and Search Rethinking Benchmark and Contamination for Language Models with Rephrased Samples

Reference 32

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T02:07:08.093098Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-13T02:03:19.836506Z digest=sha256:9c0e7bc025994d3f0ccdb7120d80d3d8109592817b7e972477d67629a713e81a

Observation 481aa828-3a0e-498f-9f0b-f56df90e7622 · inbound

Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack cites this paper.

Do Androids Dream of Breaking the Game? Systematically Auditing AI Agent Benchmarks with BenchJack Rethinking Benchmark and Contamination for Language Models with Rephrased Samples

Reference 60

Resolution
verified exact
arxiv_id, observed 2026-05-14T20:32:56.862291Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-14T20:31:50.043920Z digest=sha256:892023241eece73ae2befe17d68c0e22f1b8dc806775fc34d51e1c257b018204

Observation a0f175ac-6597-48d7-9bf4-51fc3977f5ec · inbound

LLM Benchmark Datasets Should Be Contamination-Resistant cites this paper.

LLM Benchmark Datasets Should Be Contamination-Resistant Rethinking Benchmark and Contamination for Language Models with Rephrased Samples

Reference 90

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T07:23:07.058655Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-20T07:19:50.354875Z digest=sha256:783a0fbd49606894b58e1ffa04037f84d0a5f13d0a7294de2e54c6e7f398326a

Observation ae61037a-4a0c-4149-9764-5047fc27a22e · inbound

Provable Joint Decontamination for Benchmarking Multiple Large Language Models cites this paper.

Provable Joint Decontamination for Benchmarking Multiple Large Language Models Rethinking Benchmark and Contamination for Language Models with Rephrased Samples

Reference 172

Resolution
metadata mismatch
arxiv_id, observed 2026-05-22T00:44:29.489044Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-22T00:40:54.038367Z digest=sha256:9bde49c50c36ab57117a2eb9b76b36fd0753b264aec5f6fed8147feadc02bcf3

Observation e50e38d5-a330-4082-9fc1-b47897ecc449 · inbound

The Illusion of Reasoning: Exposing Evasive Data Contamination in LLMs via Zero-CoT Truncation cites this paper.

The Illusion of Reasoning: Exposing Evasive Data Contamination in LLMs via Zero-CoT Truncation Rethinking Benchmark and Contamination for Language Models with Rephrased Samples

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-05-22T08:06:15.310194Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-22T08:05:42.212459Z digest=sha256:57e7c9704508a33ff6e95c6d0cb6c38ba642da07962de3628b04aee701c8d55b

Observation 81a8cf3c-06b8-49b8-bbd7-3da728bad548 · inbound

How Hard is it to Rig a Benchmark? A Social Choice Analysis of Leaderboard Robustness cites this paper.

How Hard is it to Rig a Benchmark? A Social Choice Analysis of Leaderboard Robustness Rethinking Benchmark and Contamination for Language Models with Rephrased Samples

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-25T04:40:24.444660Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-25T04:37:01.537128Z digest=sha256:2983bfbfd6e481532f9cfe0c85a90fbcc72694ca6da0b5936b6a07274cea6da5

Observation e11c3d40-017b-429b-8690-ff2370e45d98 · inbound

TRACER: A Semantic-Aware Framework for Fine-Grained Contamination Detection in Code LLMs cites this paper.

TRACER: A Semantic-Aware Framework for Fine-Grained Contamination Detection in Code LLMs Rethinking Benchmark and Contamination for Language Models with Rephrased Samples

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-06-30T15:34:47.895008Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-30T15:32:05.976335Z digest=sha256:578c75f7e90742e62c1984e4d01465ccda61ae493057f5873a0e5bada45a9525

Observation 04ea9e21-1c59-4575-a510-f812fd2b7aeb · inbound

Towards Evaluation Engineering: An Empirical Study of ML Evaluation Harnesses in the Wild cites this paper.

Towards Evaluation Engineering: An Empirical Study of ML Evaluation Harnesses in the Wild Rethinking Benchmark and Contamination for Language Models with Rephrased Samples

Reference 55

Resolution
verified exact
arxiv_id, observed 2026-06-30T14:44:45.155954Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-30T14:41:07.354007Z digest=sha256:b48d99cd887a14c776d61708d5a3745e0dd0478deec8f750cecc6cfbe224c7ec

Observation 6c2dc137-5f7d-4787-b49a-71cbffd49ef8 · inbound

Measuring Reasoning Quality in LLMs: A Multi-Dimensional Behavioral Framework cites this paper.

Measuring Reasoning Quality in LLMs: A Multi-Dimensional Behavioral Framework Rethinking Benchmark and Contamination for Language Models with Rephrased Samples

Reference 26

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T13:34:40.588543Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-30T13:27:50.367497Z digest=sha256:5e98fa674613234badd76d2c86263a77ec3ce70e32b49ea3f5cfe9adc1f4f5db

Observation b5a659b1-05df-4a48-8f38-913a67b3cfe8 · inbound

Measuring Reasoning Quality in LLMs: A Multi-Dimensional Behavioral Framework cites this paper.

Measuring Reasoning Quality in LLMs: A Multi-Dimensional Behavioral Framework Rethinking Benchmark and Contamination for Language Models with Rephrased Samples

Reference 26

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T08:05:31.914578Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-01T07:35:51.017797Z digest=sha256:c31d886f56ef442155e865d001befbfd2bfcc6f11461631148d551419553aa6b

Observation 73c90e48-132c-4c28-9ee1-a4949e9206f0 · inbound

Measuring Reasoning Quality in LLMs: A Multi-Dimensional Behavioral Framework cites this paper.

Measuring Reasoning Quality in LLMs: A Multi-Dimensional Behavioral Framework Rethinking Benchmark and Contamination for Language Models with Rephrased Samples

Reference 26

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T23:27:26.731890Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-02T23:22:08.567841Z digest=sha256:8af7dd22e232c5d14b010bbc90e3e188fc91d8f4c8c396ec059a892b9f87a40e

Observation dab2454a-666d-4842-b07a-aed17edf3fd0 · inbound

TSFMAudit: Data Contamination Auditing in Forecasting Time Series Foundation Models cites this paper.

TSFMAudit: Data Contamination Auditing in Forecasting Time Series Foundation Models Rethinking Benchmark and Contamination for Language Models with Rephrased Samples

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-06-30T12:04:38.780577Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-30T12:00:05.807004Z digest=sha256:c7f44d8f017103d63a9f93b57ccc1d71e20009ce77db53bf29e6520524579fe8

Observation 429ec9d5-66e8-4def-a403-8144313ea449 · inbound

Embodied-BenchClaw: An Autonomous Multi-Agent System for Embodied Spatial Intelligence Benchmark Construction cites this paper.

Embodied-BenchClaw: An Autonomous Multi-Agent System for Embodied Spatial Intelligence Benchmark Construction Rethinking Benchmark and Contamination for Language Models with Rephrased Samples

Reference 13

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T10:48:02.880656Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-27T09:48:49.786021Z digest=sha256:12249e030ffa78e7d7c25bf26fbff6c1b72667073722c02de07d207bcbf6ecb6

Observation dcbed4ea-62a8-48a1-b727-be0d3cb6004e · inbound

Which Models Are Our Models Built On? Auditing Invisible Dependencies in Modern LLMs cites this paper.

Which Models Are Our Models Built On? Auditing Invisible Dependencies in Modern LLMs Rethinking Benchmark and Contamination for Language Models with Rephrased Samples

Reference 75

Resolution
verified exact
arxiv_id, observed 2026-07-03T10:37:56.593959Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-27T09:57:14.328157Z digest=sha256:518deb86f83b0d4b6f21cc05095a751a3745cfbb1bbe186479e1422455a300b5

Observation df54faea-0bd2-4570-a54c-c347743948f7 · inbound

SrDetection: A Self-Referential Framework for Data Leakage Detection in Code Large Language Models cites this paper.

SrDetection: A Self-Referential Framework for Data Leakage Detection in Code Large Language Models Rethinking Benchmark and Contamination for Language Models with Rephrased Samples

Reference 12

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T06:24:18.367339Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-30T06:20:02.046020Z digest=sha256:4499d89ca6eb50ca96a4d4e86a99112d230d78621d321750e6552cb9fae2438b

Observation e34d5daa-ff82-4dcf-8c9b-19274d08c5d6 · inbound

Do Agent Benchmarks Measure Capability? Protocol Validity in the Age of Agentic AI cites this paper.

Do Agent Benchmarks Measure Capability? Protocol Validity in the Age of Agentic AI Rethinking Benchmark and Contamination for Language Models with Rephrased Samples

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-01T05:03:42.877641Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T05:03:42.877641Z digest=sha256:d73ca63694a1346de5b59e9887e7d7d3a85512eac999cbdb2fe0c976ed0d51e9

Observation 372bad00-77da-4547-981f-6351c5a94418 · inbound

Temporal Leakage in LLM Backtesting: Measurement, Validation, and Adjusted Scores cites this paper.

Temporal Leakage in LLM Backtesting: Measurement, Validation, and Adjusted Scores Rethinking Benchmark and Contamination for Language Models with Rephrased Samples

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-08T04:31:04.418737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T04:31:04.418737Z digest=sha256:e20c460f9083d065f5a9c962a71c81cd0a4fc2ddbdf7ab747f3cf150918ae6a1

Observation f312928c-1d49-415a-bc1e-92f5c4a57e6c · inbound

EuroExec: Frontier Language Models Fall Short of Expert Judgment on European Executive Decision Tasks cites this paper.

EuroExec: Frontier Language Models Fall Short of Expert Judgment on European Executive Decision Tasks Rethinking Benchmark and Contamination for Language Models with Rephrased Samples

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-06T22:23:05.614586Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:23:05.614586Z digest=sha256:58eef8d20bbabe5d80e9e1834b4d0e5c1f13d02a1fa00e9b5e9d51047a8eb846

Observation 1fd6bd95-accd-4fa4-b360-c9160579218f · inbound

EuroExec: Frontier Language Models Fall Short of Expert Judgment on European Executive Decision Tasks cites this paper.

EuroExec: Frontier Language Models Fall Short of Expert Judgment on European Executive Decision Tasks Rethinking Benchmark and Contamination for Language Models with Rephrased Samples

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-08T18:19:46.939354Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T18:19:46.939354Z digest=sha256:9d104ae5d296d3010fb09f12a412602fe99240ab991e3e9e775f0412774b5d00