Pith. sign in

Paper Citation Record · LEDGER

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks

As of 8 August 2026, this Paper Citation Record lists 79 of 79 outbound references and 0 inbound Pith citation observations for arXiv:2507.17747.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.17747 v2

Coverage vector

measured 79 of 79 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T14:44:09.471855Z

measured 79 of 79 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

79 of 79 outbound references displayed

  • verified exact5
  • verified fuzzy30
  • unresolved44
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation a83d6f7c-cc0d-4c07-9845-7e11451947d1 · outbound

This paper cites Claude 3.5 sonnet.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Claude 3.5 sonnet

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:44:11.909879Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T14:44:09.233992Z digest=sha256:db74767cf24ac1feb3cca793b1599b461a790517cfd9671c2be55e9b2019f68e

Observation 0c351428-9b2f-4c72-bcd6-0b4c6cf93302 · outbound

This paper cites Claude 3.5 haiku.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Claude 3.5 haiku

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:44:11.901402Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T14:44:09.237996Z digest=sha256:d5dc10f2fd08b81902ca3bf161c3d59f917a6098d41bc3e9658d338d377fb6c5

Observation 00efc9cd-aaad-4095-8e28-dba4fd18b6a8 · outbound

This paper cites ARC-AGI-2 + ARC Prize 2025 is Live! Blog Post, https://arcprize.org/blog/announcing-arc-agi-2-and-arc-prize-2025, March 24 2025.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks ARC-AGI-2 + ARC Prize 2025 is Live! Blog Post, https://arcprize.org/blog/announcing-arc-agi-2-and-arc-prize-2025, March 24 2025

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:44:11.893059Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T14:44:09.241713Z digest=sha256:1c249dba23cd401f595a29baf3ba031f63f47a7ce85aca4a68eaa9f58c2f2be3

Observation 168020df-faa3-4617-bcb0-3f661ff0428e · outbound

This paper cites Benchmarking Foundation Models with Language-Model-as-an-Examiner.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Benchmarking Foundation Models with Language-Model-as-an-Examiner

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T14:44:09.245058Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:44:09.245058Z digest=sha256:f92c8acbee74c2fe926b9fccbad573a2707d2a8cba94d134195a89530fd66637

Observation d2899a6c-dd33-44ef-9b7f-4fa173d23ace · outbound

This paper cites Leak, cheat, repeat: Data contamination and evaluation malpractices in closed-source llms.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Leak, cheat, repeat: Data contamination and evaluation malpractices in closed-source llms

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:44:11.884260Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T14:44:09.248532Z digest=sha256:7c154501da50c25a8773ed6b58da941c10006fc34faac47dd687d8ce4d18db43

Observation 8e75f131-7bc5-43fa-b6b7-b20bf3ce0b61 · outbound

This paper cites Adversarial multi-agent evaluation of large language models through iterative debates.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Adversarial multi-agent evaluation of large language models through iterative debates

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T14:44:09.252377Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:44:09.252377Z digest=sha256:6be0c2a24186e98cff8af327f05b57efd1e438f80c465a0b06076d72746a8925

Observation 325cfa56-0309-4fc2-bdd2-eb2b2dc9dbb9 · outbound

This paper cites Flageval.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Flageval

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:44:11.875447Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T14:44:09.255787Z digest=sha256:725569cc4cb1a55b55a8c157017132d56019d5d48b1f00fd68ed1ee520a18d9d

Observation 28e252be-bc05-477b-9b7c-6b1eb1892a54 · outbound

This paper cites CoT-Kinetics: A Theoretical Modeling Assessing LRM Reasoning Process.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks CoT-Kinetics: A Theoretical Modeling Assessing LRM Reasoning Process

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T14:44:09.258694Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:44:09.258694Z digest=sha256:226cb92c75d490275e9ca9aff8e1f868b1d52fdaf0efb296bc15e2c9e82c6a29

Observation 91128ddb-5cc4-4299-be72-ffd00b990fdd · outbound

This paper cites an unresolved cited work.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Unresolved cited work

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T14:44:09.261955Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:44:09.261955Z digest=sha256:602872e2552bef8d66455aab6dab893380ec69b721c912f4c5b5a0bad5298b35

Observation f00b4b03-06ee-4086-9adc-237c2aaeaa5e · outbound

This paper cites Sparks of Artificial General Intelligence: Early experiments with GPT-4.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Sparks of Artificial General Intelligence: Early experiments with GPT-4

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T14:44:09.264731Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:44:09.264731Z digest=sha256:40258123152564922cf2ddab460bb5eece5e3446e9bbee899d6d81b36c48bc24

Observation d8b5ddd1-7732-46aa-a3de-03b2b08f4d5c · outbound

This paper cites The Role of Deductive and Inductive Reasoning in Large Language Models.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks The Role of Deductive and Inductive Reasoning in Large Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T14:44:09.268106Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:44:09.268106Z digest=sha256:b8f0c3309f11eb9fd0f9ba2e3979b4db9c35f43efd253fb02af5f46c65f11371

Observation 9dca74ff-0d91-4e32-ae9c-e9c730ce254e · outbound

This paper cites Are we on the right way for evaluating large vision-language models? In A.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Are we on the right way for evaluating large vision-language models? In A

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:44:11.866662Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T14:44:09.271284Z digest=sha256:2b3bf9473e9dc880af2d44d92b5ff2a64d662d19ac0f7f13e2e412c0d63577e5

Observation 2000b724-3372-452c-975c-a70ce8d73640 · outbound

This paper cites Jordan, Joseph E.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Jordan, Joseph E

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:44:11.858241Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T14:44:09.274046Z digest=sha256:e9ec06ca61af58b6a82d9829773580bdcdcb6795e32a3c7c5c3d485ff629c46f

Observation 8877c7e4-7337-4901-9ce7-33441fd4fad0 · outbound

This paper cites ARC Prize 2024: Technical Report.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks ARC Prize 2024: Technical Report

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T14:44:09.277264Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:44:09.277264Z digest=sha256:666c359d6cc4c45c2c64b59f33f69eed670b052dfb720e8db92069ba51e4c271

Observation 8e8b6024-00f3-460a-a420-2b2ccca90a02 · outbound

This paper cites Abstraction and reasoning corpus for artificial general intelligence (arc-agi), 2019.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Abstraction and reasoning corpus for artificial general intelligence (arc-agi), 2019

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:44:11.849645Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T14:44:09.280482Z digest=sha256:d14809d8bf27d61316f1cfd50a4af018cf475f7fdf9c14cc8511fa6ab958b757

Observation 85429f59-fd84-4ab2-a982-cdf409324331 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Training Verifiers to Solve Math Word Problems

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T14:44:09.283458Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:44:09.283458Z digest=sha256:129153a7bd520f46deda3644987caff9c0795ea3e3c340acd6c345371c0edcb7

Observation 48812bb5-00e1-424c-b0e9-4dafced156f7 · outbound

This paper cites DeepSeek-V3 Technical Report.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks DeepSeek-V3 Technical Report

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T14:44:09.286882Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:44:09.286882Z digest=sha256:ea78a7c9b53435a816f790625fa6ab8f4f0cbabe909243f63164edc487605dbf

Observation a3ba679b-47be-4e9a-b1ff-285680b552b7 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T14:44:09.289691Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:44:09.289691Z digest=sha256:58c26d85ba3189d14c4024fee686f8a236598b78e555c9241861b6b5980ea808

Observation ed6de936-dfbd-410d-b224-bb625d1ae89a · outbound

This paper cites Investigating data contamination in modern benchmarks for large language models.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Investigating data contamination in modern benchmarks for large language models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T14:44:09.292859Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:44:09.292859Z digest=sha256:f1a90db1e6dc9052e96aff010012de0478456d515348b21ce7da81671354ed0d

Observation a740c12f-d3a0-49c5-b8d7-ef683acdcc95 · outbound

This paper cites Stabilizing modality gap & lowering gradient norms improve zero-shot adversarial robustness of vlms.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Stabilizing modality gap & lowering gradient norms improve zero-shot adversarial robustness of vlms

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:44:11.840252Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T14:44:09.295646Z digest=sha256:c85d4c80e179b6fe887a5b979e9bdca7b8605df48c5065b13a34b3568159f55d

Observation 8f4b08d6-f95c-48ea-9b21-134bb8dd3aba · outbound

This paper cites Improving zero-shot adversarial robustness in vision-language models by closed-form alignment of adversarial path simplices.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Improving zero-shot adversarial robustness in vision-language models by closed-form alignment of adversarial path simplices

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:44:11.830837Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T14:44:09.298363Z digest=sha256:50b8d138267b8a43646e970652e21b4a50f85c830683a987c8dcc5645758562d

Observation 6bb92465-f91d-41de-9fff-5c67a3152038 · outbound

This paper cites Improving Factuality and Reasoning in Language Models through Multiagent Debate.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Improving Factuality and Reasoning in Language Models through Multiagent Debate

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T14:44:09.301684Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:44:09.301684Z digest=sha256:65733ca45f00bb32caf6c074db9d5c07f0129c8cfa52ed1974322ebbbfce68c1

Observation a810d3f1-b2a7-4cbb-9524-d02ddea554e3 · outbound

This paper cites an unresolved cited work.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Unresolved cited work

Reference 23

Resolution
verified exact
raw_fallback, observed 2026-08-06T14:44:09.824746Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T14:44:09.304648Z digest=sha256:0ba39f0daa33966c59730d7905a5deaeda9cdb4e2e8399c5539cf6ebfe5e5917

Observation 834de928-4ab6-4b41-a962-c2fad6bd8d13 · outbound

This paper cites FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks FrontierMath: A Benchmark for Evaluating Advanced Mathematical Reasoning in AI

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T14:44:09.307985Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:44:09.307985Z digest=sha256:60694dfd4b10977a5a3412cd1a5b0bd2999c2f349db2410903f62828a1bbfebe

Observation 0a8ff0aa-536d-4f47-a10f-ed347fddadcb · outbound

This paper cites Time travel in llms: Tracing data contamination in large language models.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Time travel in llms: Tracing data contamination in large language models

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:44:11.821333Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T14:44:09.311518Z digest=sha256:eb8b2afde1a1bf1c95513be4513c5901c676351cf35a7b41898d86e0af3cf0ac

Observation 29f26e3a-ff53-43ea-8d89-28c698a42493 · outbound

This paper cites The Llama 3 Herd of Models.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks The Llama 3 Herd of Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T14:44:09.314332Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:44:09.314332Z digest=sha256:51415b57150fae2d621e3d3be9761647673d06d11bcf037961ae8d634d7b81fd

Observation 3f84371c-551c-4272-9c46-4b5923fea23e · outbound

This paper cites A Survey on LLM-as-a-Judge.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks A Survey on LLM-as-a-Judge

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T14:44:09.317974Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:44:09.317974Z digest=sha256:5622fecf4957bd666b766d311ed794d357ef35423c21ac6b3d8ea9987bdc85f2

Observation 52a345a1-6f49-40e1-810e-a7640175aa05 · outbound

This paper cites Improving Model Evaluation using SMART Filtering of Benchmark Datasets.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Improving Model Evaluation using SMART Filtering of Benchmark Datasets

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-08-06T14:44:09.727362Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T14:44:09.321034Z digest=sha256:85e36a39c756f62a385c5e68cf513fc0963146d268d2f8794894913490bfe56f

Observation e85e20e6-37bf-4433-a5fb-1c42a9f3b091 · outbound

This paper cites Measuring massive multitask language understanding.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Measuring massive multitask language understanding

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:44:11.812461Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T14:44:09.323991Z digest=sha256:5d2777afd4e0f1c2823f78dfcd0614beb9496a3cfd1f4b4892ac93a6b62ed027

Observation 1af5e0d3-7f8a-4b50-b586-318c67eb5b8b · outbound

This paper cites Trueskill : A bayesian skill rating system.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Trueskill : A bayesian skill rating system

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:44:11.803349Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T14:44:09.327209Z digest=sha256:72fecfb133f6c463621188c9601426ef7daadb02cbf581baf3cd48f6c32f2058

Observation 5126eeb1-cf26-4ab6-a34d-14f7327f25db · outbound

This paper cites Lo RA : Low-rank adaptation of large language models.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Lo RA : Low-rank adaptation of large language models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T14:44:09.330561Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:44:09.330561Z digest=sha256:0b2846505e3519e2e66b9cfb378d126f42e70e759217f7005b53be1908c8ef65

Observation 47a8eb31-62ee-40b4-88d8-7030916e14a2 · outbound

This paper cites AI safety via debate.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks AI safety via debate

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T14:44:09.333252Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:44:09.333252Z digest=sha256:fe9da43f45ff9a2a3c7f3a9f0b4eb889113a28ac19d414bc26b287b01ecf8d32

Observation c9630232-1386-4a80-9503-c007f5e98ef4 · outbound

This paper cites Mistral 7B.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Mistral 7B

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T14:44:09.336865Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:44:09.336865Z digest=sha256:85fff108a55b1b17596397fbdf01b7d540609a7ab08f6e97ea6a9ac9d3993da0

Observation b86f7a43-fbad-48c8-ab59-5f59d018b4c0 · outbound

This paper cites Mixtral of Experts.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Mixtral of Experts

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T14:44:09.340177Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:44:09.340177Z digest=sha256:902e28cc94599fd5f7f0a988a6084b6b25d76a3f18f6dac789d9cacc3c51f258

Observation 87e694a9-d61b-4e1a-be24-bfeeeeac2a80 · outbound

This paper cites Bowman, Tim Rockt \"a schel, and Ethan Perez.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Bowman, Tim Rockt \"a schel, and Ethan Perez

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:44:11.787423Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T14:44:09.343125Z digest=sha256:10a220b729154daaf633738782ce29814730812cb79f13b25b84e659ebb168d6

Observation 325bd2a9-eb0f-4a67-a6c4-a13267cefb10 · outbound

This paper cites Debate Helps Weak-to-Strong Generalization.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Debate Helps Weak-to-Strong Generalization

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T14:44:09.346312Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:44:09.346312Z digest=sha256:fd731c392b7ebc89c777a2289841050d645de0d141479ffa0c0cf3ddfa100876

Observation dcdbfae7-ea09-477a-93c7-932389ee3b53 · outbound

This paper cites LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T14:44:09.349366Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:44:09.349366Z digest=sha256:3c8e7e14222a008d015a891b37746e37228d0bf675ddafd7b3ffc8ee3da5db09

Observation 88498f70-5335-422b-be70-d39ac2d85e91 · outbound

This paper cites CMMLU: Measuring massive multitask language understanding in Chinese.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks CMMLU: Measuring massive multitask language understanding in Chinese

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T14:44:09.352447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:44:09.352447Z digest=sha256:d7fd99df9b63b20d0c8d218c595e6b2b1ddeb2acb92bb8c7e7647cf96923ffe4

Observation 87452799-2260-4345-bd6e-15bd193c9117 · outbound

This paper cites A Debate-Driven Experiment on LLM Hallucinations and Accuracy.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks A Debate-Driven Experiment on LLM Hallucinations and Accuracy

Reference 39

Resolution
verified exact
local_arxiv, observed 2026-08-06T14:44:09.659401Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T14:44:09.355365Z digest=sha256:dbd6e6dc9bbd998ab12c2d7cca958f6ed2deded07be6dab885ac9eb851a9defd

Observation ed2720b2-bae9-47a1-b47c-1b33256a0864 · outbound

This paper cites Manning, Christopher R \' e , Diana Acosta - Navas, Drew A.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Manning, Christopher R \' e , Diana Acosta - Navas, Drew A

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:44:11.777855Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T14:44:09.358457Z digest=sha256:533c57a6d0af1f2d8d76170b6a3a97a2aa0d1fb8d3d0bfdc3a62b8aee0719acf

Observation b16022a5-2bc7-40ea-b7db-0dbdc8a78cd6 · outbound

This paper cites Encouraging divergent thinking in large language models through multi-agent debate.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Encouraging divergent thinking in large language models through multi-agent debate

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:44:11.754017Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T14:44:09.361156Z digest=sha256:ed1f1c1bd9d70066ea8850cb836a2ebbec52bf717ff7d1f86ee9a6848f4e93fa

Observation bcd0f28c-7357-451f-a019-43dd1ffd8c48 · outbound

This paper cites An empirical analysis on large language models in debate evaluation.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks An empirical analysis on large language models in debate evaluation

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:44:11.699063Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T14:44:09.364110Z digest=sha256:0f070f53705a381d6c5d2bafc09575c605b097dd60a0ae8ab895dc6f3d9b9699

Observation c70ea09d-581d-4baf-a8ff-a4403bcf8c35 · outbound

This paper cites The Llama 4 herd: The beginning of a new era of natively multimodal AI innovation.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks The Llama 4 herd: The beginning of a new era of natively multimodal AI innovation

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:44:11.638130Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T14:44:09.367183Z digest=sha256:89f51b3c8cc259191fe252c7a397248a5bd78a6077c82b35cae02989ed4fd293

Observation 658269b3-bea3-41b0-9ee6-c5226b858a26 · outbound

This paper cites Discursive socratic questioning: Evaluating the faithfulness of language models' understanding of discourse relations.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Discursive socratic questioning: Evaluating the faithfulness of language models' understanding of discourse relations

Reference 44

Resolution
verified exact
doi, observed 2026-08-06T14:44:09.550141Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T14:44:09.370593Z digest=sha256:957b3dc47a192129dd276e43c20b7ad2574873eb10dd3394bd7d3e5c90ca34d5

Observation 963a0a2e-6217-4657-988a-4b2724e01846 · outbound

This paper cites Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T14:44:09.373495Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:44:09.373495Z digest=sha256:a3765ddc9528c1ea18ae40fb5b3ac68aea3c4d5738df6c0673fab3323023d8cf

Observation 471f0037-bffb-4f85-927f-10aaadc82ad1 · outbound

This paper cites Cheaper, better, faster, stronger.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Cheaper, better, faster, stronger

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:44:11.454555Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T14:44:09.376438Z digest=sha256:813b2e742e129bc84f754169c451e9b7716eb5a78ec335bff1562ff7b542b331

Observation 48907f32-9cc1-47f7-a2bb-4fc3af06a2d2 · outbound

This paper cites Mistral large.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Mistral large

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:44:11.325353Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T14:44:09.379133Z digest=sha256:876a04e69782c7909ee7499ed3d14e0df81f82455200271f9624194de584e65c

Observation e966dac2-cd50-4d15-baaa-94f3454d0b18 · outbound

This paper cites Evaluating the Performance of Large Language Models via Debates.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Evaluating the Performance of Large Language Models via Debates

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-06T14:44:09.381842Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:44:09.381842Z digest=sha256:ca71cb96ff4d8eb0554f392c62f460d3546df2b9a87b2e4998c8c80f1ba1d4e8

Observation 0933da57-b03d-4223-9983-e35edb8fc8f0 · outbound

This paper cites GPT-4 Technical Report.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks GPT-4 Technical Report

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-06T14:44:09.384917Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:44:09.384917Z digest=sha256:c21c85ac6de63bcf21ad83b2e2c432964fc74674bd69e228c620597340ee81ad

Observation ec0c7b9f-1c02-404d-b2f5-9f6bb2b5342d · outbound

This paper cites GPT-4o System Card.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks GPT-4o System Card

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T14:44:09.387932Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:44:09.387932Z digest=sha256:f524eb737c2fd1d6a42927dcdf8e4b60b9258c2c9fcf105cc715f9b57caa40a3

Observation f817549b-8f37-4ac2-800d-c183768200b8 · outbound

This paper cites GPT-4o mini: advancing cost-efficient intelligence.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks GPT-4o mini: advancing cost-efficient intelligence

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:44:11.063257Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T14:44:09.391394Z digest=sha256:6fa5b206895b4198fa21ada05b2e7620a3aa4e1de83d22bed264f97c3316311c

Observation 8c4d6f4f-37f0-4b11-a6fc-4272bd79a562 · outbound

This paper cites OpenAI o1 System Card.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks OpenAI o1 System Card

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-06T14:44:09.394084Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:44:09.394084Z digest=sha256:6a16398d0b18ef5ae3cf55bc5413bc44e33860e5ba300e384c94d91e541d8059

Observation 81b8e3f6-eb82-4a21-94f3-71c9f31dfe3d · outbound

This paper cites Chatterji, Faisal Ladhak, and Tatsunori Hashimoto.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Chatterji, Faisal Ladhak, and Tatsunori Hashimoto

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-06T14:44:09.396998Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:44:09.396998Z digest=sha256:c182b400750066e54fb2b0fd24ca74e840b21699542fed8b3a2290407de0dcf3

Observation 067938fe-4300-44b8-b2e4-4beba9569ad5 · outbound

This paper cites Mapping global dynamics of benchmark creation and saturation in artificial intelligence.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Mapping global dynamics of benchmark creation and saturation in artificial intelligence

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-06T14:44:09.399824Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:44:09.399824Z digest=sha256:3b101b2ff0dd12a6eb0e13dc2c9c8fb82d6d51a6a4dc971040b1341cb630bab1

Observation 7053c87d-d62c-44aa-b800-fd743facd11e · outbound

This paper cites Humanity's Last Exam.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Humanity's Last Exam

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-06T14:44:09.402625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:44:09.402625Z digest=sha256:df572e60eeacd7d95977ec1c89273eac4f2942bee0fc02d556028480ff3390df

Observation f12d1973-fa38-4f70-9902-f1b6e486e20c · outbound

This paper cites Introducing gemini 2.0: our new ai model for the agentic era.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Introducing gemini 2.0: our new ai model for the agentic era

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:44:10.854039Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T14:44:09.405513Z digest=sha256:3471cc41ea343df12c0b31374b8362d58e6c92039da78186c231327fddeec690

Observation 36bb7fd3-ed87-4e09-b768-ccd3d1c8dc91 · outbound

This paper cites Multi-layered evaluation using a fusion of metrics and LLMs as judges in open-domain question answering.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Multi-layered evaluation using a fusion of metrics and LLMs as judges in open-domain question answering

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:44:10.471219Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T14:44:09.408457Z digest=sha256:aef31862fd14a990e8c1db6ac6160ecd7d2901b34380e44aaa22a112de2fb83c

Observation 07618cbf-e2a8-490c-bec3-fbc58bbe8bf3 · outbound

This paper cites GPQA: A Graduate-Level Google-Proof Q&A Benchmark.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks GPQA: A Graduate-Level Google-Proof Q&A Benchmark

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-06T14:44:09.411020Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:44:09.411020Z digest=sha256:afca59f902fbbb8f3c83f7fffa60a976d36a2a0ab82f5b2203dba400af578a1e

Observation 8bb4c05e-68cc-4951-855a-2f612619ae98 · outbound

This paper cites Nlp evaluation in trouble: On the need to measure llm data contamination for each benchmark.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Nlp evaluation in trouble: On the need to measure llm data contamination for each benchmark

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-06T14:44:09.413880Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:44:09.413880Z digest=sha256:ba1b5cbbd3919b973fe6f35ab74244f33d4536339d29c1db3b0bbd10c9fbef11

Observation 3cf9e6df-393b-4b02-adde-6800a7912adf · outbound

This paper cites Pretraining on the Test Set Is All You Need.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Pretraining on the Test Set Is All You Need

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-06T14:44:09.416573Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:44:09.416573Z digest=sha256:1714e8a642400bd2c489d60c3a343e5160148be4a597ebf24cc803b04b98b2d4

Observation d5675478-26e6-415f-b752-bc3c18d2885d · outbound

This paper cites Detecting pretraining data from large language models.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Detecting pretraining data from large language models

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:44:10.319102Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T14:44:09.419654Z digest=sha256:cd286a152ea8b2df19f34d739363b5ec3627174a60bf4dd3f67a848c1b71cf47

Observation db9419df-7567-4932-9c0b-f3dd29fabb58 · outbound

This paper cites Beyond the imitation game: Quantifying and extrapolating the capabilities of language models.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Beyond the imitation game: Quantifying and extrapolating the capabilities of language models

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:44:10.271308Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T14:44:09.422236Z digest=sha256:0950774adb37853a5c48b9691ed3ec36bed17df7496f0138350008b74bc551d0

Observation 9717d2a1-315a-4ac1-b51d-f7940e0fb2e5 · outbound

This paper cites MMLU-Pro+: Evaluating Higher-Order Reasoning and Shortcut Learning in LLMs.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks MMLU-Pro+: Evaluating Higher-Order Reasoning and Shortcut Learning in LLMs

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-06T14:44:09.424893Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:44:09.424893Z digest=sha256:40c546fb2bbbc24da08d5d262096626ba70e84068112f2abe497c82ff36c3591

Observation 1647dcfc-81fe-4705-a412-fd81c0ab1680 · outbound

This paper cites an unresolved cited work.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Unresolved cited work

Reference 64

Resolution
unresolved
raw_fallback, observed 2026-08-06T14:44:10.208592Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T14:44:09.427865Z digest=sha256:c129c8b5b996ec20e336364f5f12b482c9e1b14d2ea1cdcaeeb8bf3d7ebdfd6c

Observation e7e1d3ac-fd53-430f-8ac1-ee4cd9cb21f5 · outbound

This paper cites an unresolved cited work.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Unresolved cited work

Reference 65

Resolution
unresolved
raw_fallback, observed 2026-08-06T14:44:10.148317Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T14:44:09.430445Z digest=sha256:0cfe74cd833b4472484f727fc57069b3610f18faeb82cc8ca42d0b32de5616b7

Observation 29e7c9eb-c1a0-4b40-a774-be1a12be7040 · outbound

This paper cites Chi, Sharan Narang, Aakanksha Chowdhery, and Denny Zhou.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Chi, Sharan Narang, Aakanksha Chowdhery, and Denny Zhou

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-06T14:44:09.433045Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:44:09.433045Z digest=sha256:e03bbe0e08ff411e453b686ee7fc70d71247b690cf6edd7b29d244f7459cc36d

Observation afab2f8b-6539-46ef-83d0-73bada5393a3 · outbound

This paper cites Mmlu-pro: A more robust and challenging multi-task language understanding benchmark.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Mmlu-pro: A more robust and challenging multi-task language understanding benchmark

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:44:10.124690Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T14:44:09.436267Z digest=sha256:362ab5465b4f823d2f4717d5fd061c96df14e073fe6d926b2d9845007aa69b8e

Observation ff3d821e-b5f0-47d2-aae8-11173c64161c · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Chain-of-thought prompting elicits reasoning in large language models

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:44:10.115093Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T14:44:09.439020Z digest=sha256:b6bf22a9bd2e669e38299a0732ee387532d0184543db23065ed33f492ea35647

Observation 6c16eb36-7f0b-4507-a872-929d4fc131ee · outbound

This paper cites Livebench: A challenging, contamination-free LLM benchmark.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Livebench: A challenging, contamination-free LLM benchmark

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:44:10.105651Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T14:44:09.441678Z digest=sha256:f15ca3da97487b87ae3518c1aa805a38d8bf1e283b7279087a1102d7e2c85076

Observation ff34c8d0-bf09-4d70-9d71-6d5258bfea53 · outbound

This paper cites QUD eval: The evaluation of questions under discussion discourse parsing.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks QUD eval: The evaluation of questions under discussion discourse parsing

Reference 70

Resolution
verified exact
doi, observed 2026-08-06T14:44:09.505113Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T14:44:09.444655Z digest=sha256:62288534d6bf15232ad01f725a6f46bfb048afce8320a17260f0176f9c2a993c

Observation f51edf2e-e3ea-4b73-b8f5-05d222f22c20 · outbound

This paper cites Benchmark Data Contamination of Large Language Models: A Survey.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Benchmark Data Contamination of Large Language Models: A Survey

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-06T14:44:09.448188Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:44:09.448188Z digest=sha256:50fc8974af6d551c3a077372b1bd39eb3a25305a50139f30ea940b8e82293dac

Observation faaed509-9aa0-41be-8f71-1d8c24b38fac · outbound

This paper cites KIEval: A Knowledge-grounded Interactive Evaluation Framework for Large Language Models.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks KIEval: A Knowledge-grounded Interactive Evaluation Framework for Large Language Models

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-06T14:44:09.451179Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:44:09.451179Z digest=sha256:deaadbe599ec94ee6ccc9b27d8a565d2fa48c2510c6fb1fa94322aa70510188d

Observation d1754a25-0780-4bf6-8e85-c1d0fa4eb958 · outbound

This paper cites Auto-Arena: Automating LLM Evaluations with Agent Peer Battles and Committee Discussions.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Auto-Arena: Automating LLM Evaluations with Agent Peer Battles and Committee Discussions

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-06T14:44:09.454188Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:44:09.454188Z digest=sha256:995ae815fb9a8b7a33bbc8874921606636669c84c4c4d3e83e76c935b1b58f02

Observation 28e1a61d-5759-49a8-a35d-c52eb1b52a35 · outbound

This paper cites Xing, Hao Zhang, Joseph E.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Xing, Hao Zhang, Joseph E

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:44:10.096336Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T14:44:09.457014Z digest=sha256:dd2dffb2858fa7e423be5b6c844d4105650daa0419a026b1fdbacb3a195fb957

Observation aa7be3b1-c45d-452b-a424-1fc0ded6b452 · outbound

This paper cites Dyval: Dynamic evaluation of large language models for reasoning tasks.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Dyval: Dynamic evaluation of large language models for reasoning tasks

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:44:10.086077Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T14:44:09.459891Z digest=sha256:99d7a0508514e0cea9812ba7e49ab2db1aec4c2379f7b273df273cb3f675023a

Observation 7d0e033a-3e88-47c9-bea6-db99e9f6c28f · outbound

This paper cites write newline.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks write newline

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-06T14:44:09.462617Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:44:09.462617Z digest=sha256:ed3640440ee1b48d3d3addab2f2a99373939944935ea750079e7b559eb229629

Observation 6edf6c22-fd91-48f8-8d29-35f64be39ffd · outbound

This paper cites @esa (Ref.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks @esa (Ref

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-06T14:44:09.465929Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:44:09.465929Z digest=sha256:b784b944adef4003e7760643b112d6cd437b421b3d81bdaafcca970a09c2fd55

Observation 4339ae32-1247-4087-955b-68eaa972ae45 · outbound

This paper cites an unresolved cited work.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Unresolved cited work

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-06T14:44:09.468994Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:44:09.468994Z digest=sha256:1c3aac01433218b5e6c60f9205a384a78a7b698253e3a4b051d08915d97bc81a

Observation 5fde41a7-2a87-47c6-9831-2c3a0ca7d3a2 · outbound

This paper cites an unresolved cited work.

Pretraining on the Test Set Is No Longer All You Need: A Debate-Driven Approach to QA Benchmarks Unresolved cited work

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-06T14:44:09.471855Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:44:09.471855Z digest=sha256:ddc5345093e68ae4837c141227e3f666da3ea17cc30c3c655455e01342b8520d

Pith citing papers

No inbound Pith citation observations are available.