Pith. sign in

Paper Citation Record · LEDGER

Don't Make Your LLM an Evaluation Benchmark Cheater

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 39 inbound Pith citation observations for arXiv:2311.01964.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2311.01964 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 39 of 39 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 39 of 39 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-09T10:26:07.046518Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

18
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 3aaa96ce-fe7c-4b16-930d-13eb3909db90 · inbound

Bridging Language and Items for Retrieval and Recommendation: Benchmarking LLMs as Semantic Encoders cites this paper.

Bridging Language and Items for Retrieval and Recommendation: Benchmarking LLMs as Semantic Encoders Don't Make Your LLM an Evaluation Benchmark Cheater

Reference 62

Resolution
verified exact
arxiv_id, observed 2026-05-24T03:05:56.990257Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-24T03:03:51.053556Z digest=sha256:9b2d25935ea6185f9c5eef835af4079181742120e7f8a51e366965e7d1471bd9

Observation 09a21ebc-9711-4355-941f-f7a01f2d0c76 · inbound

LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code cites this paper.

LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code Don't Make Your LLM an Evaluation Benchmark Cheater

Reference 223

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T17:34:42.839749Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-10T17:34:42.565806Z digest=sha256:3ed987f78d49cfad3fe7218ef56e069ffa10c057e5888f8565f6e569c9019ff6

Observation 01ef215d-f693-4a35-a813-56451877a25c · inbound

Benchmark Data Contamination of Large Language Models: A Survey cites this paper.

Benchmark Data Contamination of Large Language Models: A Survey Don't Make Your LLM an Evaluation Benchmark Cheater

Reference 187

Resolution
verified exact
arxiv_id, observed 2026-05-22T23:10:41.160577Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-22T23:10:40.420241Z digest=sha256:2fb69801174b336278ba03f50f1737b160efbb5bec6aec3a4b6376907c2d9b57

Observation 6e87ec1d-90ca-4af7-9894-a408d1a97434 · inbound

Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs cites this paper.

Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs Don't Make Your LLM an Evaluation Benchmark Cheater

Reference 155

Resolution
verified exact
arxiv_id, observed 2026-05-17T00:05:03.901359Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-17T00:05:03.547664Z digest=sha256:0cf7971943650eca026adc6ac1a4d497d87a7546f91772e58078d7bd2156b414

Observation 11e37873-e86e-4360-8c9d-92532dc29dd4 · inbound

MetaMorph: Multimodal Understanding and Generation via Instruction Tuning cites this paper.

MetaMorph: Multimodal Understanding and Generation via Instruction Tuning Don't Make Your LLM an Evaluation Benchmark Cheater

Reference 197

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T07:51:13.131940Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-17T07:51:12.953777Z digest=sha256:bb27f3c66508627ad7afc9e1a0541b08392f09d1fb358aea0b263ea704dea933

Observation baffcdb0-75e8-4e3c-ba63-4cef2ebc7c87 · inbound

Training an LLM-as-a-Judge Model: Pipeline, Insights, and Practical Lessons cites this paper.

Training an LLM-as-a-Judge Model: Pipeline, Insights, and Practical Lessons Don't Make Your LLM an Evaluation Benchmark Cheater

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-09T10:26:07.046518Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T10:26:07.046518Z digest=sha256:45a0c78dd9887a198cdfb022495795a39ebb456c4ade8da3bf4ceb8fb206f0f5

Observation a45fb792-6052-4f45-a804-274711e4840e · inbound

Unbiased Evaluation of Large Language Models from a Causal Perspective cites this paper.

Unbiased Evaluation of Large Language Models from a Causal Perspective Don't Make Your LLM an Evaluation Benchmark Cheater

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-08T14:51:15.499571Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:51:15.499571Z digest=sha256:b087412d225ef5a5cf98ac5c3314e0c65700297d563888b8038e0c9003192826

Observation da61f099-43c0-4391-b0e5-9c8fdf2484e3 · inbound

Evaluation of Deep Audio Representations for Hearables cites this paper.

Evaluation of Deep Audio Representations for Hearables Don't Make Your LLM an Evaluation Benchmark Cheater

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-08T14:47:07.948894Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:47:07.948894Z digest=sha256:fb858f8bc0f8c1749729ebf2444b9ccd711d646e97c056c061279f4cfcc8a409

Observation 34222d53-6312-45f7-80db-5a4f47696879 · inbound

Large Language Models in Code Co-generation for Safe Autonomous Vehicles cites this paper.

Large Language Models in Code Co-generation for Safe Autonomous Vehicles Don't Make Your LLM an Evaluation Benchmark Cheater

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T14:13:41.882888Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:13:41.882888Z digest=sha256:f3378070b9cc2282d253c911cf1df44c986940ad625c24666f559e2e826dc4ca

Observation 5e2361a0-1672-422d-8945-2dc470f45371 · inbound

Establishing Trustworthy LLM Evaluation via Shortcut Neuron Analysis cites this paper.

Establishing Trustworthy LLM Evaluation via Shortcut Neuron Analysis Don't Make Your LLM an Evaluation Benchmark Cheater

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T10:54:26.413010Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:54:26.413010Z digest=sha256:9413115bb8a1097502212b0c4e615c4c5556871546aa08d960bf831df8c9a501

Observation 3983b9a3-1339-405a-8c93-7ae58305dce3 · inbound

A Conceptual Framework for AI Capability Evaluations cites this paper.

A Conceptual Framework for AI Capability Evaluations Don't Make Your LLM an Evaluation Benchmark Cheater

Reference 103

Resolution
unresolved
no resolver link, observed 2026-08-06T23:25:53.137574Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:25:53.137574Z digest=sha256:d7230bc969ed45f71f8ec8d5714c4203ec3f7662856eb2de3668daa29ce51146

Observation c201a5ad-d818-43c7-b7a5-0128eba0b6ae · inbound

Can Vision Language Models Understand Mimed Actions? cites this paper.

Can Vision Language Models Understand Mimed Actions? Don't Make Your LLM an Evaluation Benchmark Cheater

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-07T00:23:50.793745Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:23:50.793745Z digest=sha256:9d2e6427b4663002b2e30fe27f207104b41176a877da33409d6b33cebcf77830

Observation 11652142-6a23-4b3c-b5d8-bee15e27fe51 · inbound

Establishing Best Practices for Building Rigorous Agentic Benchmarks cites this paper.

Establishing Best Practices for Building Rigorous Agentic Benchmarks Don't Make Your LLM an Evaluation Benchmark Cheater

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-06T20:24:28.438581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:24:28.438581Z digest=sha256:4ac6cee011478fb3d11962d7d47e4731413a933c519b0654a7d23d9263cee8f7

Observation 5a5eeb3d-ad4e-4eb4-807a-a8526dd0d0aa · inbound

Toward Valid Measurement Of (Un)fairness For Generative AI: A Proposal For Systematization Through The Lens Of Fair Equality of Chances cites this paper.

Toward Valid Measurement Of (Un)fairness For Generative AI: A Proposal For Systematization Through The Lens Of Fair Equality of Chances Don't Make Your LLM an Evaluation Benchmark Cheater

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-06T19:47:49.856525Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:47:49.856525Z digest=sha256:9653d7cf00bf3ed66750a6ca26282d5d8edbd363a3fefd152920fac4537e10a3

Observation 5ecc2e53-7f59-435d-afc4-b5bb1194e4e6 · inbound

User Behavior Prediction as a Generic, Robust, Scalable, and Low-Cost Evaluation Strategy for Estimating Generalization in LLMs cites this paper.

User Behavior Prediction as a Generic, Robust, Scalable, and Low-Cost Evaluation Strategy for Estimating Generalization in LLMs Don't Make Your LLM an Evaluation Benchmark Cheater

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-06T21:43:41.764235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:43:41.764235Z digest=sha256:84d7ac2f8fae7ad47e176f3e738a66a3bd397aabd82de7f73933a06e352d1acc

Observation a680c7fb-277d-4b9c-a183-19a2afd2d214 · inbound

Deprecating Benchmarks: Criteria and Framework cites this paper.

Deprecating Benchmarks: Criteria and Framework Don't Make Your LLM an Evaluation Benchmark Cheater

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-06T19:07:40.530184Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:07:40.530184Z digest=sha256:a221674d229faddc4e53f6e95408ac8432a194357cd0585a7ec1cd7c3e9856fd

Observation 97aa3bc0-878f-47d6-8b62-1e86e5b135dd · inbound

ISACL: Internal State Analyzer for Copyrighted Training Data Leakage cites this paper.

ISACL: Internal State Analyzer for Copyrighted Training Data Leakage Don't Make Your LLM an Evaluation Benchmark Cheater

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-05T16:50:28.964070Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T16:50:28.964070Z digest=sha256:8a46fc89e453fed32436e22060cb20cb0413c735c3d2b53c294aed2b4af76f7b

Observation 814f58bc-4f4b-435d-92ee-04ac5bae3038 · inbound

InSQuAD: In-Context Learning for Efficient Retrieval via Submodular Mutual Information to Enforce Quality and Diversity cites this paper.

InSQuAD: In-Context Learning for Efficient Retrieval via Submodular Mutual Information to Enforce Quality and Diversity Don't Make Your LLM an Evaluation Benchmark Cheater

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-05T14:44:28.055272Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:44:28.055272Z digest=sha256:b114053855ab561fbd366daa331c7cfb67293810bf9bb07bf9e28830ca861ce8

Observation ca9ba4eb-c67c-4ba0-94b8-5b5319d0bcc5 · inbound

LogitTrace: Detecting Benchmark Contamination via Layerwise Logit Trajectories cites this paper.

LogitTrace: Detecting Benchmark Contamination via Layerwise Logit Trajectories Don't Make Your LLM an Evaluation Benchmark Cheater

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-18T14:36:28.964642Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-18T14:32:57.426996Z digest=sha256:a5ee7275b536e4e157eb279911effa2e566ddd14a954b3f02799add504577569

Observation 62e80328-de9e-46ef-8346-c51a5dec4830 · inbound

Make a Video Call with LLM: A Measurement Campaign over Six Mainstream Apps cites this paper.

Make a Video Call with LLM: A Measurement Campaign over Six Mainstream Apps Don't Make Your LLM an Evaluation Benchmark Cheater

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-04T13:29:04.649764Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T13:29:04.649764Z digest=sha256:54ec4f877165e5a25227436cafe93fab7dd92d4db03d86a713ed115a16cf42af

Observation 13edee3e-fa11-45a1-a31e-00ca70fb6dd7 · inbound

SciCoQA: Quality Assurance for Scientific Paper--Code Alignment cites this paper.

SciCoQA: Quality Assurance for Scientific Paper--Code Alignment Don't Make Your LLM an Evaluation Benchmark Cheater

Reference 11

Resolution
malformed identifier
arxiv_id, observed 2026-05-16T13:20:57.867847Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-16T13:20:07.741347Z digest=sha256:ff8912cd3155d65298a8b40fded1e06fbf7955f4384cf7aeacfebf822ae28c8b

Observation 0be77a86-62d8-4578-a0e6-0e23617115be · inbound

Benchmark Leakage Trap: Can We Trust LLM-based Recommendation? cites this paper.

Benchmark Leakage Trap: Can We Trust LLM-based Recommendation? Don't Make Your LLM an Evaluation Benchmark Cheater

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-02T23:32:09.809875Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T23:32:09.809875Z digest=sha256:0922145c8d09f7580914e6a276e04c107f43f752da1a5f553998aef93e94dd1e

Observation fb03d784-d5ec-47bb-b80b-073e1a2b49b2 · inbound

Agentic Business Process Management: A Research Manifesto cites this paper.

Agentic Business Process Management: A Research Manifesto Don't Make Your LLM an Evaluation Benchmark Cheater

Reference 84

Resolution
verified exact
arxiv_id, observed 2026-05-15T08:25:18.811678Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-15T08:24:56.763438Z digest=sha256:f6a3e3b6f80220c9797bcd3aa164264b1cf26606bf7b49018cee9e87efe9bf38

Observation 22ed82ce-384d-4802-8fd1-d4c1ae345d65 · inbound

LiveFact: A Dynamic, Time-Aware Benchmark for LLM-Driven Fake News Detection cites this paper.

LiveFact: A Dynamic, Time-Aware Benchmark for LLM-Driven Fake News Detection Don't Make Your LLM an Evaluation Benchmark Cheater

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-10T21:55:49.393337Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T20:30:53.094207Z digest=sha256:7123f45063e1de4621459efe4ffb767938fa853f9beeda5f8db6c1d49d4c8f67

Observation a69ccd46-f3ac-4193-967f-d2b183c8a583 · inbound

Riemann-Bench: A Benchmark for Moonshot Mathematics cites this paper.

Riemann-Bench: A Benchmark for Moonshot Mathematics Don't Make Your LLM an Evaluation Benchmark Cheater

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:36:00.811359Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T18:02:42.607682Z digest=sha256:9792628b9a85c7ab502cb3bc6df4efca272d3381c5666c236510622c9e87607e

Observation 227c3cac-82a2-40f0-be47-0d92e3231ded · inbound

How Independent are Large Language Models? A Statistical Framework for Auditing Behavioral Entanglement and Reweighting Verifier Ensembles cites this paper.

How Independent are Large Language Models? A Statistical Framework for Auditing Behavioral Entanglement and Reweighting Verifier Ensembles Don't Make Your LLM an Evaluation Benchmark Cheater

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-11T07:20:58.932401Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T17:13:04.305435Z digest=sha256:9b951363551e5c509661506bda2ecc5387243cb0c5a4e22a2577bcc49269cc0a

Observation 5d80a5c7-a48b-4c8f-831f-22fbfb20d00a · inbound

ActuBench: A Multi-Agent LLM Pipeline for Generation and Evaluation of Actuarial Reasoning Tasks cites this paper.

ActuBench: A Multi-Agent LLM Pipeline for Generation and Evaluation of Actuarial Reasoning Tasks Don't Make Your LLM an Evaluation Benchmark Cheater

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:46:05.207324Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T00:35:24.397273Z digest=sha256:b6b74412fe799c33657e7436d3c049c58f5b756c37670dfa0e08b174e0120183

Observation d1555bcf-9371-468c-b0c0-4efe0b8c70e2 · inbound

Training a General Purpose Automated Red Teaming Model cites this paper.

Training a General Purpose Automated Red Teaming Model Don't Make Your LLM an Evaluation Benchmark Cheater

Reference 5

Resolution
malformed identifier
arxiv_id, observed 2026-05-11T19:41:08.319868Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-08T11:20:09.224745Z digest=sha256:ba4d2a26445a0e24ff1f2026b8031ae9fd8ef9d071d7223334a153e05eeddff2

Observation dbdfad1f-74b3-4fa4-9dc4-5dd2f0084784 · inbound

STELLAR-E: a Synthetic, Tailored, End-to-end LLM Application Rigorous Evaluator cites this paper.

STELLAR-E: a Synthetic, Tailored, End-to-end LLM Application Rigorous Evaluator Don't Make Your LLM an Evaluation Benchmark Cheater

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-11T22:01:10.945701Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-08T03:39:30.528601Z digest=sha256:0bbf935c2e3010002e687128835666118ec9f1e906280e21c7884fb8032634eb

Observation 2e23b564-b627-4049-8ba4-b2fd10035453 · inbound

Programming with Data: Test-Driven Data Engineering for Self-Improving LLMs from Raw Corpora cites this paper.

Programming with Data: Test-Driven Data Engineering for Self-Improving LLMs from Raw Corpora Don't Make Your LLM an Evaluation Benchmark Cheater

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-05-11T22:11:16.280192Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-08T03:13:30.648190Z digest=sha256:64297becc1f9751081b3e50bf224fd6104e2d30fa8f98f53b7160f514641010c

Observation f8720662-6515-4a69-bcce-5408c93bda67 · inbound

Generating Leakage-Free Benchmarks for Robust RAG Evaluation cites this paper.

Generating Leakage-Free Benchmarks for Robust RAG Evaluation Don't Make Your LLM an Evaluation Benchmark Cheater

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-12T03:06:18.948679Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-12T03:04:39.253429Z digest=sha256:29a816068fdb0f78c49a8bc0eae123d269f32b18814c65dc885354a2a6f0380b

Observation d5bf708a-2ecb-4cc9-8afa-52f0f61be15f · inbound

Provable Joint Decontamination for Benchmarking Multiple Large Language Models cites this paper.

Provable Joint Decontamination for Benchmarking Multiple Large Language Models Don't Make Your LLM an Evaluation Benchmark Cheater

Reference 181

Resolution
verified exact
arxiv_id, observed 2026-05-22T00:44:29.503089Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-22T00:40:54.038367Z digest=sha256:cedf0c19f65afc203503038b6f63b5f176cf287c595a2d9b2441985b0f14ba3d

Observation 6ac7c5d0-9059-4a44-9a04-cdcff63d2890 · inbound

How Hard is it to Rig a Benchmark? A Social Choice Analysis of Leaderboard Robustness cites this paper.

How Hard is it to Rig a Benchmark? A Social Choice Analysis of Leaderboard Robustness Don't Make Your LLM an Evaluation Benchmark Cheater

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-25T04:40:24.435108Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-25T04:37:01.537128Z digest=sha256:3063874e940ee84d2a648bb77bfe6edf4d79c579ab5de34cecd7da5472c73c67

Observation 7feb40be-8f95-4344-97f7-80e1ed66da49 · inbound

Pretraining Data Exposure in Large Language Models: A Survey of Membership Inference, Data Contamination, and Security Implications cites this paper.

Pretraining Data Exposure in Large Language Models: A Survey of Membership Inference, Data Contamination, and Security Implications Don't Make Your LLM an Evaluation Benchmark Cheater

Reference 67

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T17:24:56.540890Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-30T17:20:16.735285Z digest=sha256:437259af7f1e4a4dd6f885f5b449f58255afac52bf0695dc8b8bc2fa15d2166c

Observation 3f65e975-0d9d-4cb3-a563-06e43cae517f · inbound

Welfare, Improvability, and Variance: A Principal-Agent Approach to Optimal Benchmark Item Aggregation cites this paper.

Welfare, Improvability, and Variance: A Principal-Agent Approach to Optimal Benchmark Item Aggregation Don't Make Your LLM an Evaluation Benchmark Cheater

Reference 73

Resolution
metadata mismatch
arxiv_id, observed 2026-06-28T23:52:49.465940Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-28T23:49:57.580051Z digest=sha256:1c36f01b2dde2443b84ea4ac476f1679952894eaf9ae126ff985719ec8a8daf8

Observation 119018cb-5232-4f83-8943-cac3dbc2a450 · inbound

Flaws in the LLM Automation Narrative cites this paper.

Flaws in the LLM Automation Narrative Don't Make Your LLM an Evaluation Benchmark Cheater

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-07-03T08:17:45.290264Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-27T10:52:36.252919Z digest=sha256:ff28ce4015f538fda4fd5df6a11dc31b48fdf85651ce84e1ef69625ae132bdca

Observation c4eb3bdd-b34a-46fb-aca3-ecf7dea39295 · inbound

Defeat Devices in AI Systems cites this paper.

Defeat Devices in AI Systems Don't Make Your LLM an Evaluation Benchmark Cheater

Reference 61

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T08:44:27.875364Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-30T08:34:58.879346Z digest=sha256:a216b61e15283fb0a8265744c09172c38d636b0fe68605535969ca9fa2d19a45

Observation be1c851b-6620-4735-9841-8d4ecfb0c51c · inbound

From Execution to Education: A Bloom-Aligned Framework for Measuring Educational Control in LLMs cites this paper.

From Execution to Education: A Bloom-Aligned Framework for Measuring Educational Control in LLMs Don't Make Your LLM an Evaluation Benchmark Cheater

Reference 208

Resolution
verified exact
local_arxiv, observed 2026-07-10T13:57:06.719623Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-07-10T13:49:17.343893Z digest=sha256:c943815497158504a4f5ea154b34690e7d1910e9757e37f142eef8af0ce8494d

Observation 4bd18683-eceb-42fc-acc7-3c6ea3c4f451 · inbound

Guideline-as-Oracle: Zero-Annotation Training of an Ophthalmic Telephone Triage Agent cites this paper.

Guideline-as-Oracle: Zero-Annotation Training of an Ophthalmic Telephone Triage Agent Don't Make Your LLM an Evaluation Benchmark Cheater

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T17:07:02.934623Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:07:02.934623Z digest=sha256:2f9bd0d255d9db99278b3881ac8d0b4c7a5d0197a1342e6281766c9bccd05336