Pith. sign in

Paper Citation Record · LEDGER

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks

As of 18 August 2026, this Paper Citation Record lists 79 of 79 outbound references and 0 inbound Pith citation observations for arXiv:2507.17747.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.17747 v2

Coverage vector

measured 79 of 79 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T14:44:09.471855Z

measured 79 of 79 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

79 of 79 outbound references displayed

  • verified exact5
  • verified fuzzy30
  • unresolved44
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation a83d6f7c-cc0d-4c07-9845-7e11451947d1 · outbound

This paper cites Claude 3.5 sonnet.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Claude 3.5 sonnet

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:44:11.909879Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T14:44:09.233992Z digest=sha256:bdcf27c7a27d100ee4fff5839de069e001a974bde9a75ff692e2f5d4d2f29337

Observation 0c351428-9b2f-4c72-bcd6-0b4c6cf93302 · outbound

This paper cites Claude 3.5 haiku.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Claude 3.5 haiku

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:44:11.901402Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T14:44:09.237996Z digest=sha256:17bc2293e6c8d0606ed0822f36e9aaa62225dd507e51221cacb9b24ef09b45d8

Observation 00efc9cd-aaad-4095-8e28-dba4fd18b6a8 · outbound

This paper cites ARC-AGI-2 + ARC Prize 2025 is Live! Blog Post, https://arcprize.org/blog/announcing-arc-agi-2-and-arc-prize-2025, March 24 2025.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks ARC-AGI-2 + ARC Prize 2025 is Live! Blog Post, https://arcprize.org/blog/announcing-arc-agi-2-and-arc-prize-2025, March 24 2025

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:44:11.893059Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T14:44:09.241713Z digest=sha256:4aaf19344e6ea303c83de2752eee4b1e68d98443324fa16ff3cabbea002e3bc8

Observation 168020df-faa3-4617-bcb0-3f661ff0428e · outbound

This paper cites Benchmarking Foundation Models with Language-Model-as-an-Examiner.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Benchmarking Foundation Models with Language-Model-as-an-Examiner

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T14:44:09.245058Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:44:09.245058Z digest=sha256:34f585701cd446a2812d257c52ad5a4b3ca6870c360e32e7dbd1a2ed3ff123a9

Observation d2899a6c-dd33-44ef-9b7f-4fa173d23ace · outbound

This paper cites Leak, cheat, repeat: Data contamination and evaluation malpractices in closed-source llms.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Leak, cheat, repeat: Data contamination and evaluation malpractices in closed-source llms

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:44:11.884260Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T14:44:09.248532Z digest=sha256:3de929392e97d09ad68528c64b51c461c96c263404bb384b906e831860c0291c

Observation 8e75f131-7bc5-43fa-b6b7-b20bf3ce0b61 · outbound

This paper cites Adversarial multi-agent evaluation of large language models through iterative debates.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Adversarial multi-agent evaluation of large language models through iterative debates

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T14:44:09.252377Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:44:09.252377Z digest=sha256:67abe1d53c03f93e9c99c26c1e1a8c801d60396ccc520e5f7a49d8399c86993b

Observation 325cfa56-0309-4fc2-bdd2-eb2b2dc9dbb9 · outbound

This paper cites Flageval.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Flageval

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:44:11.875447Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T14:44:09.255787Z digest=sha256:6cd06a5c4eafc8cf57974ef4903787a3d171807d07a6fc816f45ebc25ce98920

Observation 28e252be-bc05-477b-9b7c-6b1eb1892a54 · outbound

This paper cites CoT-Kinetics: A Theoretical Modeling Assessing LRM Reasoning Process.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks CoT-Kinetics: A Theoretical Modeling Assessing LRM Reasoning Process

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T14:44:09.258694Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:44:09.258694Z digest=sha256:c4df8964ca10b57abe3c7cd767f8ccfd49632f8ee26d38b2e696767181488b78

Observation 91128ddb-5cc4-4299-be72-ffd00b990fdd · outbound

This paper cites an unresolved cited work.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Unresolved cited work

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T14:44:09.261955Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:44:09.261955Z digest=sha256:bd842d3e4d99816bca8f30f0ae8cf1a9dadc10d6310760a06f147141fe20c75b

Observation f00b4b03-06ee-4086-9adc-237c2aaeaa5e · outbound

This paper cites Sparks of Artificial General Intelligence: Early experiments with GPT-4.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Sparks of Artificial General Intelligence: Early experiments with GPT-4

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T14:44:09.264731Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:44:09.264731Z digest=sha256:6af54c691a133561227bdf8de8a040be18c059c965aaec6bd1e7718543a13992

Observation d8b5ddd1-7732-46aa-a3de-03b2b08f4d5c · outbound

This paper cites The Role of Deductive and Inductive Reasoning in Large Language Models.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks The Role of Deductive and Inductive Reasoning in Large Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T14:44:09.268106Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:44:09.268106Z digest=sha256:15958b6e95618c36974e998a07239e61c6e15facd95853281583414e95d63cc8

Observation 9dca74ff-0d91-4e32-ae9c-e9c730ce254e · outbound

This paper cites Are we on the right way for evaluating large vision-language models? In A.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Are we on the right way for evaluating large vision-language models? In A

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:44:11.866662Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T14:44:09.271284Z digest=sha256:ef7d05b762e87639880c80a4a6b62d3e72b03ee490940baa8bbd22dabbe65567

Observation 2000b724-3372-452c-975c-a70ce8d73640 · outbound

This paper cites Jordan, Joseph E.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Jordan, Joseph E

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:44:11.858241Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T14:44:09.274046Z digest=sha256:c885064455626223874cdd94c9ee7f5588d3fc0cfb951b281e0cc46ed735fee7

Observation 8877c7e4-7337-4901-9ce7-33441fd4fad0 · outbound

This paper cites ARC Prize 2024: Technical Report.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks ARC Prize 2024: Technical Report

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T14:44:09.277264Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:44:09.277264Z digest=sha256:44241866fbe56adb609fc7433e309658e530fda3f9d6901601ed548a67db56b4

Observation 8e8b6024-00f3-460a-a420-2b2ccca90a02 · outbound

This paper cites Abstraction and reasoning corpus for artificial general intelligence (arc-agi), 2019.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Abstraction and reasoning corpus for artificial general intelligence (arc-agi), 2019

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:44:11.849645Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T14:44:09.280482Z digest=sha256:2f7ef0b34a524c7ab66bcb3e3c6f4eb7fe81738b7ce8c3b56f7c03c70a60427c

Observation 85429f59-fd84-4ab2-a982-cdf409324331 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Training Verifiers to Solve Math Word Problems

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T14:44:09.283458Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:44:09.283458Z digest=sha256:42d0441271c3fe984f106318c0eb81ec365a65262822496efc8dfb9c6bf170d1

Observation 48812bb5-00e1-424c-b0e9-4dafced156f7 · outbound

This paper cites DeepSeek-V3 Technical Report.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks DeepSeek-V3 Technical Report

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T14:44:09.286882Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:44:09.286882Z digest=sha256:ed7f9daaf2b474a340d3c3610429c21b3a72fc03672ce00aaa17cec65e9653d4

Observation a3ba679b-47be-4e9a-b1ff-285680b552b7 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T14:44:09.289691Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:44:09.289691Z digest=sha256:92845e1fd3df05259a5a36b2b70bdcd431a3f5728b5e2419c41de8e0c5e75cb8

Observation ed6de936-dfbd-410d-b224-bb625d1ae89a · outbound

This paper cites Investigating data contamination in modern benchmarks for large language models.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Investigating data contamination in modern benchmarks for large language models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T14:44:09.292859Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:44:09.292859Z digest=sha256:26dd1eedc3eb3d40ee461d530324868c104d5375ee643097ce08f9aad4b3f80b

Observation a740c12f-d3a0-49c5-b8d7-ef683acdcc95 · outbound

This paper cites Stabilizing modality gap & lowering gradient norms improve zero-shot adversarial robustness of vlms.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Stabilizing modality gap & lowering gradient norms improve zero-shot adversarial robustness of vlms

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:44:11.840252Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T14:44:09.295646Z digest=sha256:54b21f35aad0da577cc01582f53ae9f195da1b7c1886f60248a93030128e0c8d

Observation 8f4b08d6-f95c-48ea-9b21-134bb8dd3aba · outbound

This paper cites Improving zero-shot adversarial robustness in vision-language models by closed-form alignment of adversarial path simplices.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Improving zero-shot adversarial robustness in vision-language models by closed-form alignment of adversarial path simplices

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:44:11.830837Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T14:44:09.298363Z digest=sha256:bce1e939d0f5d298a44785bf2c5c2642e419ddff5e2af5cf017e357fefa3112d

Observation 6bb92465-f91d-41de-9fff-5c67a3152038 · outbound

This paper cites Improving Factuality and Reasoning in Language Models through Multiagent Debate.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Improving Factuality and Reasoning in Language Models through Multiagent Debate

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T14:44:09.301684Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:44:09.301684Z digest=sha256:ca8451836584aa175c34c37178116a814387e1787ca6fab9334a372f8089aceb

Observation a810d3f1-b2a7-4cbb-9524-d02ddea554e3 · outbound

This paper cites an unresolved cited work.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Unresolved cited work

Reference 23

Resolution
verified exact
raw_fallback, observed 2026-08-06T14:44:09.824746Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T14:44:09.304648Z digest=sha256:a8754ecf564f9ce1baaf56f92575b3959d04333df6392025a8cf1fd760e10026

Observation 834de928-4ab6-4b41-a962-c2fad6bd8d13 · outbound

This paper cites FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T14:44:09.307985Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:44:09.307985Z digest=sha256:433250aaf7dae218794e2ea98ed22831d06c71d775aba79f30fd80a25c29a39b

Observation 0a8ff0aa-536d-4f47-a10f-ed347fddadcb · outbound

This paper cites Time travel in llms: Tracing data contamination in large language models.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Time travel in llms: Tracing data contamination in large language models

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:44:11.821333Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T14:44:09.311518Z digest=sha256:e6afe3b0e3b76f6b1f812ad19e2b9fe16742a72ca63057cecc46e7bd1e62a573

Observation 29f26e3a-ff53-43ea-8d89-28c698a42493 · outbound

This paper cites The Llama 3 Herd of Models.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks The Llama 3 Herd of Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T14:44:09.314332Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:44:09.314332Z digest=sha256:2fddf4b46bf7edf58c34cf0cfaf498007c84e7e594e6d806290bf4af76400f88

Observation 3f84371c-551c-4272-9c46-4b5923fea23e · outbound

This paper cites A Survey on LLM-as-a-Judge.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks A Survey on LLM-as-a-Judge

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T14:44:09.317974Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:44:09.317974Z digest=sha256:e6076c7a13936f643bee0f9ae5c34375dbdf4ec46ec6a6197254f0cdd0d59b55

Observation 52a345a1-6f49-40e1-810e-a7640175aa05 · outbound

This paper cites Improving Model Evaluation using SMART Filtering of Benchmark Datasets.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Improving Model Evaluation using SMART Filtering of Benchmark Datasets

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-08-06T14:44:09.727362Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T14:44:09.321034Z digest=sha256:c6be3ee353e4a308a1f11b419dbc41170adfc652e2803fbda13d472c8b1072e6

Observation e85e20e6-37bf-4433-a5fb-1c42a9f3b091 · outbound

This paper cites Measuring massive multitask language understanding.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Measuring massive multitask language understanding

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:44:11.812461Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T14:44:09.323991Z digest=sha256:87759b5fae2d8000b01730f1f47b22f66351a337ba95b93498bdacc3bfd7ed95

Observation 1af5e0d3-7f8a-4b50-b586-318c67eb5b8b · outbound

This paper cites Trueskill : A bayesian skill rating system.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Trueskill : A bayesian skill rating system

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:44:11.803349Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T14:44:09.327209Z digest=sha256:5c2187ac5267b36be73ed8d5d80aeaf67b0ca66209303f1b8666daaed8950c2c

Observation 5126eeb1-cf26-4ab6-a34d-14f7327f25db · outbound

This paper cites Lo RA : Low-rank adaptation of large language models.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Lo RA : Low-rank adaptation of large language models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T14:44:09.330561Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:44:09.330561Z digest=sha256:c05b62d15e2f8ae940a46b5ff7e00989abf81cb7786f795a1d1467c161180ee7

Observation 47a8eb31-62ee-40b4-88d8-7030916e14a2 · outbound

This paper cites AI safety via debate.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks AI safety via debate

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T14:44:09.333252Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:44:09.333252Z digest=sha256:dfa708062f3a78a7aaf37a04598f6907b0908873c36772c46dc8d44e18f2f904

Observation c9630232-1386-4a80-9503-c007f5e98ef4 · outbound

This paper cites Mistral 7B.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Mistral 7B

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T14:44:09.336865Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:44:09.336865Z digest=sha256:02636f6492728abd701e2f003078f485ee99ebed3d40e0c0a6c1c4f344fecd10

Observation b86f7a43-fbad-48c8-ab59-5f59d018b4c0 · outbound

This paper cites Mixtral of Experts.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Mixtral of Experts

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T14:44:09.340177Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:44:09.340177Z digest=sha256:ce36b052bde799cf0187653679736002ddfb693f9d89105a894b78656f373886

Observation 87e694a9-d61b-4e1a-be24-bfeeeeac2a80 · outbound

This paper cites Bowman, Tim Rockt \"a schel, and Ethan Perez.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Bowman, Tim Rockt \"a schel, and Ethan Perez

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:44:11.787423Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T14:44:09.343125Z digest=sha256:c20dbec96942a367030204bd7018f608abd4aa1311a691011f1a09deec2ffed8

Observation 325bd2a9-eb0f-4a67-a6c4-a13267cefb10 · outbound

This paper cites Debate Helps Weak-to-Strong Generalization.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Debate Helps Weak-to-Strong Generalization

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T14:44:09.346312Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:44:09.346312Z digest=sha256:27d6161f7929b7eeda67ac4b4e49b98464d1a647d3bbf797ca5483404d64f2eb

Observation dcdbfae7-ea09-477a-93c7-932389ee3b53 · outbound

This paper cites LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T14:44:09.349366Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:44:09.349366Z digest=sha256:ecbe92e1733b013fc282e30f1361e66e2d6b62b00b27cf92aedd9296f5a7e2ec

Observation 88498f70-5335-422b-be70-d39ac2d85e91 · outbound

This paper cites CMMLU: Measuring massive multitask language understanding in Chinese.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks CMMLU: Measuring massive multitask language understanding in Chinese

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T14:44:09.352447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:44:09.352447Z digest=sha256:573871b42174039036318034a0ab9b9c919c7a0f248f98a585b17a7bfdcf6f17

Observation 87452799-2260-4345-bd6e-15bd193c9117 · outbound

This paper cites A Debate-Driven Experiment on LLM Hallucinations and Accuracy.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks A Debate-Driven Experiment on LLM Hallucinations and Accuracy

Reference 39

Resolution
verified exact
local_arxiv, observed 2026-08-06T14:44:09.659401Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T14:44:09.355365Z digest=sha256:6a6c924ae556d5511b8b476e20aaf6361f0b99914c4eaa6876b9f19d55c08a46

Observation ed2720b2-bae9-47a1-b47c-1b33256a0864 · outbound

This paper cites Manning, Christopher R \' e , Diana Acosta - Navas, Drew A.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Manning, Christopher R \' e , Diana Acosta - Navas, Drew A

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:44:11.777855Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T14:44:09.358457Z digest=sha256:e99a6fe6db514fa0579cbd12b144e006f0f25f3f6d7e49d88dc061d49abc2184

Observation b16022a5-2bc7-40ea-b7db-0dbdc8a78cd6 · outbound

This paper cites Encouraging divergent thinking in large language models through multi-agent debate.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Encouraging divergent thinking in large language models through multi-agent debate

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:44:11.754017Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T14:44:09.361156Z digest=sha256:85ce79bda7ff88782e6b3966fc54a26d191c361b5a19fd75aa8521dc4041ec58

Observation bcd0f28c-7357-451f-a019-43dd1ffd8c48 · outbound

This paper cites An empirical analysis on large language models in debate evaluation.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks An empirical analysis on large language models in debate evaluation

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:44:11.699063Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T14:44:09.364110Z digest=sha256:f1b25f27a00c110f8659f136c0260b187c1da36db0431f77d6a7a0dcec75b05e

Observation c70ea09d-581d-4baf-a8ff-a4403bcf8c35 · outbound

This paper cites The Llama 4 herd: The beginning of a new era of natively multimodal AI innovation.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks The Llama 4 herd: The beginning of a new era of natively multimodal AI innovation

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:44:11.638130Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T14:44:09.367183Z digest=sha256:5b8aa752156ddb6078b6e0861e0cf05b3cb3c2f79aba9fb930a87645d80d6f68

Observation 658269b3-bea3-41b0-9ee6-c5226b858a26 · outbound

This paper cites Discursive socratic questioning: Evaluating the faithfulness of language models' understanding of discourse relations.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Discursive socratic questioning: Evaluating the faithfulness of language models' understanding of discourse relations

Reference 44

Resolution
verified exact
doi, observed 2026-08-06T14:44:09.550141Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T14:44:09.370593Z digest=sha256:c1f9221585e79d996f11798a0a6b9afed9945b9273f01dfbb903d6974d384f27

Observation 963a0a2e-6217-4657-988a-4b2724e01846 · outbound

This paper cites Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T14:44:09.373495Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:44:09.373495Z digest=sha256:e970f71369f45ef01c5ae597947024547813fa86f2a76e8b76f5f2c18b9a8648

Observation 471f0037-bffb-4f85-927f-10aaadc82ad1 · outbound

This paper cites Cheaper, better, faster, stronger.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Cheaper, better, faster, stronger

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:44:11.454555Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T14:44:09.376438Z digest=sha256:9ee751433c8d0d831bca9409ad4105da4a8bd203a6703273b968497933263b1a

Observation 48907f32-9cc1-47f7-a2bb-4fc3af06a2d2 · outbound

This paper cites Mistral large.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Mistral large

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:44:11.325353Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T14:44:09.379133Z digest=sha256:cfb682a96c5ccc24856937d13016a7af9d518adbc0c3579cb37cddbe5446623a

Observation e966dac2-cd50-4d15-baaa-94f3454d0b18 · outbound

This paper cites Evaluating the Performance of Large Language Models via Debates.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Evaluating the Performance of Large Language Models via Debates

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-06T14:44:09.381842Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:44:09.381842Z digest=sha256:f4c2f4828c93fac1cdb930c5631cbc614927a09a38341ac44ef3d57b61daaf3c

Observation 0933da57-b03d-4223-9983-e35edb8fc8f0 · outbound

This paper cites GPT-4 Technical Report.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks GPT-4 Technical Report

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-06T14:44:09.384917Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:44:09.384917Z digest=sha256:e26f46fd20d7026137206a1f408fd0942e9d21be3635312b080dde5171b0387e

Observation ec0c7b9f-1c02-404d-b2f5-9f6bb2b5342d · outbound

This paper cites GPT-4o System Card.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks GPT-4o System Card

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T14:44:09.387932Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:44:09.387932Z digest=sha256:56f3bddcd1c49e23193f0a63e2b75ba173f7203cc25118ff254d07d85d936089

Observation f817549b-8f37-4ac2-800d-c183768200b8 · outbound

This paper cites GPT-4o mini: advancing cost-efficient intelligence.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks GPT-4o mini: advancing cost-efficient intelligence

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:44:11.063257Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T14:44:09.391394Z digest=sha256:81b675712476818207ed54f8c2a29f7f4333f5f6d9cdd7f46e9b0e9e4c188ac2

Observation 8c4d6f4f-37f0-4b11-a6fc-4272bd79a562 · outbound

This paper cites OpenAI o1 System Card.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks OpenAI o1 System Card

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-06T14:44:09.394084Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:44:09.394084Z digest=sha256:81ad08591b2ecc403f449f9e2fb761de19620bbef581fd12b1e1261cf0e96da3

Observation 81b8e3f6-eb82-4a21-94f3-71c9f31dfe3d · outbound

This paper cites Chatterji, Faisal Ladhak, and Tatsunori Hashimoto.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Chatterji, Faisal Ladhak, and Tatsunori Hashimoto

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-06T14:44:09.396998Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:44:09.396998Z digest=sha256:7ddccc2be8ada29db20f73dc53d89bea88b03ff410d66d43eb22b6491755f03e

Observation 067938fe-4300-44b8-b2e4-4beba9569ad5 · outbound

This paper cites Mapping global dynamics of benchmark creation and saturation in artificial intelligence.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Mapping global dynamics of benchmark creation and saturation in artificial intelligence

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-06T14:44:09.399824Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:44:09.399824Z digest=sha256:7202d60ec793b36bb478d46e2dda427056d0a29a14f50a2dbd0ac2bb10507570

Observation 7053c87d-d62c-44aa-b800-fd743facd11e · outbound

This paper cites Humanity's Last Exam.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Humanity's Last Exam

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-06T14:44:09.402625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:44:09.402625Z digest=sha256:1910c92acf1a0b56dcee78e55aac854c016dcea1b462a43bb351404836554eb5

Observation f12d1973-fa38-4f70-9902-f1b6e486e20c · outbound

This paper cites Introducing gemini 2.0: our new ai model for the agentic era.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Introducing gemini 2.0: our new ai model for the agentic era

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:44:10.854039Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T14:44:09.405513Z digest=sha256:0d1a733244d043b437d4865399846e6066122f9291294e61f5f7ef3b276be755

Observation 36bb7fd3-ed87-4e09-b768-ccd3d1c8dc91 · outbound

This paper cites Multi-layered evaluation using a fusion of metrics and LLMs as judges in open-domain question answering.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Multi-layered evaluation using a fusion of metrics and LLMs as judges in open-domain question answering

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:44:10.471219Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T14:44:09.408457Z digest=sha256:c324f3e428e17969490158d013494328b013c2580634e943492e282c56b4f108

Observation 07618cbf-e2a8-490c-bec3-fbc58bbe8bf3 · outbound

This paper cites GPQA: A Graduate-Level Google-Proof Q&A Benchmark.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks GPQA: A Graduate-Level Google-Proof Q&A Benchmark

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-06T14:44:09.411020Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:44:09.411020Z digest=sha256:8b412686bcaa93293ca93a4f0668e66432f8fd8d39e42c42c1729b5aab70d9b3

Observation 8bb4c05e-68cc-4951-855a-2f612619ae98 · outbound

This paper cites Nlp evaluation in trouble: On the need to measure llm data contamination for each benchmark.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Nlp evaluation in trouble: On the need to measure llm data contamination for each benchmark

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-06T14:44:09.413880Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:44:09.413880Z digest=sha256:8ba959f252fdb7d10e3456cdc907b181f7b36fab70fa123f1908a37c7ea87b1d

Observation 3cf9e6df-393b-4b02-adde-6800a7912adf · outbound

This paper cites Pretraining on the Test Set Is All You Need.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Pretraining on the Test Set Is All You Need

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-06T14:44:09.416573Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:44:09.416573Z digest=sha256:46db0de4606f0e274e3cf219b583d898f80a2560a70b835eb391570fca626112

Observation d5675478-26e6-415f-b752-bc3c18d2885d · outbound

This paper cites Detecting pretraining data from large language models.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Detecting pretraining data from large language models

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:44:10.319102Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T14:44:09.419654Z digest=sha256:0317ff33c2ecc9fab822c4f2061186a7617841afd6f037f05022acdfb964d02b

Observation db9419df-7567-4932-9c0b-f3dd29fabb58 · outbound

This paper cites Beyond the imitation game: Quantifying and extrapolating the capabilities of language models.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Beyond the imitation game: Quantifying and extrapolating the capabilities of language models

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:44:10.271308Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T14:44:09.422236Z digest=sha256:390e1a931b2e47619e6b9b220944038e7a2161ddee463ca680aed9c2de1220bd

Observation 9717d2a1-315a-4ac1-b51d-f7940e0fb2e5 · outbound

This paper cites MMLU-Pro+: Evaluating Higher-Order Reasoning and Shortcut Learning in LLMs.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks MMLU-Pro+: Evaluating Higher-Order Reasoning and Shortcut Learning in LLMs

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-06T14:44:09.424893Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:44:09.424893Z digest=sha256:3b8e7f06f3f8bab3761f98d371f54668f1e25a5f780deb54c88e76ccbd61c33c

Observation 1647dcfc-81fe-4705-a412-fd81c0ab1680 · outbound

This paper cites an unresolved cited work.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Unresolved cited work

Reference 64

Resolution
unresolved
raw_fallback, observed 2026-08-06T14:44:10.208592Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T14:44:09.427865Z digest=sha256:0422f563821cf299793e456bad211ab009ec6b9b6c816ec2e5733a9f6479c6f0

Observation e7e1d3ac-fd53-430f-8ac1-ee4cd9cb21f5 · outbound

This paper cites an unresolved cited work.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Unresolved cited work

Reference 65

Resolution
unresolved
raw_fallback, observed 2026-08-06T14:44:10.148317Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T14:44:09.430445Z digest=sha256:d262d80c43c64a07544f577a8358261b91645f1ff459a999069d37880913b7f0

Observation 29e7c9eb-c1a0-4b40-a774-be1a12be7040 · outbound

This paper cites Chi, Sharan Narang, Aakanksha Chowdhery, and Denny Zhou.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Chi, Sharan Narang, Aakanksha Chowdhery, and Denny Zhou

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-06T14:44:09.433045Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:44:09.433045Z digest=sha256:19598e8b17a983eb58c6cf883ce8bc664a4d2a400811f02064e6c1a08b137a21

Observation afab2f8b-6539-46ef-83d0-73bada5393a3 · outbound

This paper cites Mmlu-pro: A more robust and challenging multi-task language understanding benchmark.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Mmlu-pro: A more robust and challenging multi-task language understanding benchmark

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:44:10.124690Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T14:44:09.436267Z digest=sha256:b7c2d30d87720f39783fc91c37a92bc54ddfacea677a08f634c530c85dc1dcc3

Observation ff3d821e-b5f0-47d2-aae8-11173c64161c · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Chain-of-thought prompting elicits reasoning in large language models

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:44:10.115093Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T14:44:09.439020Z digest=sha256:17dd3d6f464587a93fcfe2542b0b6fdb9e426fe87551ecc622d9c26072c12fcd

Observation 6c16eb36-7f0b-4507-a872-929d4fc131ee · outbound

This paper cites Livebench: A challenging, contamination-free LLM benchmark.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Livebench: A challenging, contamination-free LLM benchmark

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:44:10.105651Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T14:44:09.441678Z digest=sha256:b8d3ba7f4a3ab43791660c0b292391f31d94d75df7e461f4007f1d5b5db5e54c

Observation ff34c8d0-bf09-4d70-9d71-6d5258bfea53 · outbound

This paper cites QUD eval: The evaluation of questions under discussion discourse parsing.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks QUD eval: The evaluation of questions under discussion discourse parsing

Reference 70

Resolution
verified exact
doi, observed 2026-08-06T14:44:09.505113Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T14:44:09.444655Z digest=sha256:0b622a412bf2a943b117d4ef043c310653c8f423eeadf1a8d985d91165fc5807

Observation f51edf2e-e3ea-4b73-b8f5-05d222f22c20 · outbound

This paper cites Benchmark Data Contamination of Large Language Models: A Survey.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Benchmark Data Contamination of Large Language Models: A Survey

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-06T14:44:09.448188Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:44:09.448188Z digest=sha256:80e7121d79d2208762d2fb6cb24e89586f63624ed35164a69c78c1f34eb80deb

Observation faaed509-9aa0-41be-8f71-1d8c24b38fac · outbound

This paper cites KIEval: A Knowledge-grounded Interactive Evaluation Framework for Large Language Models.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks KIEval: A Knowledge-grounded Interactive Evaluation Framework for Large Language Models

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-06T14:44:09.451179Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:44:09.451179Z digest=sha256:509c37ca02e18b5b14a82f13f6555690501f8e9db94813856278c3084d08f89f

Observation d1754a25-0780-4bf6-8e85-c1d0fa4eb958 · outbound

This paper cites Auto-Arena: Automating LLM Evaluations with Agent Peer Battles and Committee Discussions.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Auto-Arena: Automating LLM Evaluations with Agent Peer Battles and Committee Discussions

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-06T14:44:09.454188Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:44:09.454188Z digest=sha256:c3a5b3413f473795d5e101e3f530096f2ae6fc9f234e4ce993e22e642f676be9

Observation 28e1a61d-5759-49a8-a35d-c52eb1b52a35 · outbound

This paper cites Xing, Hao Zhang, Joseph E.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Xing, Hao Zhang, Joseph E

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:44:10.096336Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T14:44:09.457014Z digest=sha256:f26935aaa4bc42874930e3d119afd010ffdfcfb4ad29b2dfdba39adf66a9994b

Observation aa7be3b1-c45d-452b-a424-1fc0ded6b452 · outbound

This paper cites Dyval: Dynamic evaluation of large language models for reasoning tasks.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Dyval: Dynamic evaluation of large language models for reasoning tasks

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:44:10.086077Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T14:44:09.459891Z digest=sha256:2bae531eb851da19b0bef310dba817971bf8fc1826df42e0eac85c227a7cce56

Observation 7d0e033a-3e88-47c9-bea6-db99e9f6c28f · outbound

This paper cites write newline.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks write newline

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-06T14:44:09.462617Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:44:09.462617Z digest=sha256:f55b34f9965d1be0ad86745347ec22e3aa001cf5403099ec87c77f3ac4325926

Observation 6edf6c22-fd91-48f8-8d29-35f64be39ffd · outbound

This paper cites @esa (Ref.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks @esa (Ref

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-06T14:44:09.465929Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:44:09.465929Z digest=sha256:0f49af575a881e5099bdc01b357127302fdca0dcc326105adbe39a9aa7526f1d

Observation 4339ae32-1247-4087-955b-68eaa972ae45 · outbound

This paper cites an unresolved cited work.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Unresolved cited work

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-06T14:44:09.468994Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:44:09.468994Z digest=sha256:5524f520cafe1681b1ae61edf6c7d6f2b6cbe73c1100ba38a351f83e47c4b46d

Observation 5fde41a7-2a87-47c6-9831-2c3a0ca7d3a2 · outbound

This paper cites an unresolved cited work.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Unresolved cited work

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-06T14:44:09.471855Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:44:09.471855Z digest=sha256:15b9f25f3a38fb6edffd45933cdb93f80e4cf87fbb0378ee6f527bfc24574ebc

Pith citing papers

No inbound Pith citation observations are available.