Pith. sign in

Paper Citation Record · LEDGER

DebateBench: A Challenging Long Context Reasoning Benchmark For Large Language Models

As of 10 August 2026, this Paper Citation Record lists 25 of 25 outbound references and 1 inbound Pith citation observation for arXiv:2502.06279.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.06279 v1

Coverage vector

measured 25 of 25 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-08T16:11:18.608790Z

measured 26 of 26 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:45:15.313592Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-07T12:45:27.776502Z

Reference resolution

25 of 25 outbound references displayed

  • verified exact1
  • verified fuzzy0
  • unresolved24
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation cd8caff9-11fb-4582-8ecc-16047f0ac32b · outbound

This paper cites Have LLMs Advanced Enough? A Challenging Problem Solving Benchmark For Large Language Models.

DebateBench: A Challenging Long Context Reasoning Benchmark For Large Language Models Have LLMs Advanced Enough? A Challenging Problem Solving Benchmark For Large Language Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-08T16:11:18.486154Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T16:11:18.486154Z digest=sha256:14e8764d7cbcccc2830b442eb515796369caad0e50754947b4accd55f600f762

Observation d4031c7c-5b05-4bda-9964-018e3747ea35 · outbound

This paper cites Program Synthesis with Large Language Models.

DebateBench: A Challenging Long Context Reasoning Benchmark For Large Language Models Program Synthesis with Large Language Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-08T16:11:18.491829Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T16:11:18.491829Z digest=sha256:fa9f015bac19ab43816d358c85b0c6c9e2d904dc4241c8f8167e9b0a68793e67

Observation 917d2ff8-be1a-4ed0-9420-bdd2cb478dbd · outbound

This paper cites Sparks of Artificial General Intelligence: Early experiments with GPT-4.

DebateBench: A Challenging Long Context Reasoning Benchmark For Large Language Models Sparks of Artificial General Intelligence: Early experiments with GPT-4

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-08T16:11:18.496913Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T16:11:18.496913Z digest=sha256:5483840e5cc8bee4bd4f8750cb3340ae299fbb93cee574e663fe89dacd4360d2

Observation d77e5042-c4b6-4cdd-89bf-00c456542f15 · outbound

This paper cites Evaluating Large Language Models Trained on Code.

DebateBench: A Challenging Long Context Reasoning Benchmark For Large Language Models Evaluating Large Language Models Trained on Code

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-08T16:11:18.501581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T16:11:18.501581Z digest=sha256:523f58fb73a82869c672a9ff3621e4871cd8ba4830220c8a3603e9e695af55db

Observation 10063763-7df0-49db-882a-aa354370fab7 · outbound

This paper cites Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge.

DebateBench: A Challenging Long Context Reasoning Benchmark For Large Language Models Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-08T16:11:18.506182Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T16:11:18.506182Z digest=sha256:37c1277a9f5d4df0d7c5d453701f040dc691d78110d324b8e9f61c976555abbc

Observation 94fc4e2f-e2dc-43c5-8869-d05148b5e03a · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

DebateBench: A Challenging Long Context Reasoning Benchmark For Large Language Models Training Verifiers to Solve Math Word Problems

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-08T16:11:18.511898Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T16:11:18.511898Z digest=sha256:0bfd8279771235cff3dc2beafd32f92a51f900691635475a2d0888ff6a008158

Observation 9d8a23f7-f692-427d-8402-ec6ee201c926 · outbound

This paper cites A Survey on In-context Learning.

DebateBench: A Challenging Long Context Reasoning Benchmark For Large Language Models A Survey on In-context Learning

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-08T16:11:18.518463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T16:11:18.518463Z digest=sha256:5a1b27c4aaf63e63ba50fe7e8e55f13597a7a235253b76f380830c2373c55c49

Observation 677caec8-45cd-4a68-a371-e34078015f5e · outbound

This paper cites an unresolved cited work.

DebateBench: A Challenging Long Context Reasoning Benchmark For Large Language Models Unresolved cited work

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-08T16:11:18.524453Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T16:11:18.524453Z digest=sha256:cd2d3250b93721eaf583e541adb62a2ba9d9d89ae141399af2cb6b183a3c1481

Observation 2e3f2138-e44b-46d9-a738-fb1e493fde1a · outbound

This paper cites Measuring Massive Multitask Language Understanding.

DebateBench: A Challenging Long Context Reasoning Benchmark For Large Language Models Measuring Massive Multitask Language Understanding

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-08T16:11:18.529248Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T16:11:18.529248Z digest=sha256:d9c2f8ecdc6365852de2943fa894dfdfe727512fd73367464ef492635f983983

Observation 6e8f89c7-074f-4661-af23-24e73cf61f5a · outbound

This paper cites Measuring Mathematical Problem Solving With the MATH Dataset.

DebateBench: A Challenging Long Context Reasoning Benchmark For Large Language Models Measuring Mathematical Problem Solving With the MATH Dataset

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-08T16:11:18.534365Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T16:11:18.534365Z digest=sha256:4194c94e431116063ff8febd0694f1870c341a728f8632cc8ed702c428c7d6b6

Observation 18711a9e-2ab0-4e71-b4c3-9c3d00785754 · outbound

This paper cites OpenAI o1 System Card.

DebateBench: A Challenging Long Context Reasoning Benchmark For Large Language Models OpenAI o1 System Card

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-08T16:11:18.539393Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T16:11:18.539393Z digest=sha256:c81a053807d584307dfa2f17e8e0605af31a09a6d873d1af18626281d100b5fb

Observation 43a8759a-9757-47ea-8e2a-8842b90ae72b · outbound

This paper cites ETHIC: Evaluating Large Language Models on Long-Context Tasks with High Information Coverage.

DebateBench: A Challenging Long Context Reasoning Benchmark For Large Language Models ETHIC: Evaluating Large Language Models on Long-Context Tasks with High Information Coverage

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-08T16:11:18.544009Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T16:11:18.544009Z digest=sha256:1b0ffb201f068d45bf341ed8d5624caa9947d241e0248ed6dfedc68b2f2b274b

Observation d604e143-5e9c-4847-ba4c-12d71d9e42c0 · outbound

This paper cites TruthfulQA: Measuring How Models Mimic Human Falsehoods.

DebateBench: A Challenging Long Context Reasoning Benchmark For Large Language Models TruthfulQA: Measuring How Models Mimic Human Falsehoods

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-08T16:11:18.548724Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T16:11:18.548724Z digest=sha256:07abb4d6a2aac509624ee41f1d0018492b687d3ad4f77ee01471aeb3929b101b

Observation 5e4ad9ce-873a-4e1e-a28b-fec3ba94127f · outbound

This paper cites an unresolved cited work.

DebateBench: A Challenging Long Context Reasoning Benchmark For Large Language Models Unresolved cited work

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-08T16:11:18.553353Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T16:11:18.553353Z digest=sha256:0f07aa2dfe12de5ef3067d16acf4d919cb4121ce952d743eb991bc9e1c8e9224

Observation d7d359cc-271e-48f2-b038-c0de03b0bcdb · outbound

This paper cites GPT-4 Technical Report.

DebateBench: A Challenging Long Context Reasoning Benchmark For Large Language Models GPT-4 Technical Report

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-08T16:11:18.557726Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T16:11:18.557726Z digest=sha256:9603c71f36e9a20edcefdeb683b783bb7165c1a9d7141f3fcea828bd7e95f37c

Observation 7d17e15d-de30-4fcd-9749-cf50fd25a879 · outbound

This paper cites an unresolved cited work.

DebateBench: A Challenging Long Context Reasoning Benchmark For Large Language Models Unresolved cited work

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-08T16:11:18.562510Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T16:11:18.562510Z digest=sha256:edc231149217f2fe57092c8737066605b290b7b2db31fd359a872cb45eab94fc

Observation eb7ae997-d196-417d-982a-be9c65607ffc · outbound

This paper cites VivesDebate-Speech: A Corpus of Spoken Argumentation to Leverage Audio Features for Argument Mining.

DebateBench: A Challenging Long Context Reasoning Benchmark For Large Language Models VivesDebate-Speech: A Corpus of Spoken Argumentation to Leverage Audio Features for Argument Mining

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-08-08T16:11:18.772504Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-08T16:11:18.566920Z digest=sha256:801e09256122ab890e885defa1dd0a925511d18841df744e570249f6b2706d46

Observation ece87034-6b18-40ac-81e5-f3ef7ad81d5a · outbound

This paper cites WinoGrande: An Adversarial Winograd Schema Challenge at Scale.

DebateBench: A Challenging Long Context Reasoning Benchmark For Large Language Models WinoGrande: An Adversarial Winograd Schema Challenge at Scale

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-08T16:11:18.572248Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T16:11:18.572248Z digest=sha256:7d9a00fc4b3ad8a836d04bf8b026ff9b4468b30eb0e2dfb69d93635a07d03cd4

Observation 8e27cc1d-708c-43d0-a8e3-657e5178bde9 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

DebateBench: A Challenging Long Context Reasoning Benchmark For Large Language Models Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-08T16:11:18.577374Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T16:11:18.577374Z digest=sha256:77477a11160ec38a0500ba56ad80591c788f14b631c1bcfdfabe2cd0429cb7df

Observation 72ec0687-a5c1-4e87-a590-60a3940f4dff · outbound

This paper cites Whisper-GPT -- Continuous Discrete Hybrid Representation Language Models For Speech And Music.

DebateBench: A Challenging Long Context Reasoning Benchmark For Large Language Models Whisper-GPT -- Continuous Discrete Hybrid Representation Language Models For Speech And Music

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-08T16:11:18.582452Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T16:11:18.582452Z digest=sha256:a37cd16cbb93d63790346babcd4916d84b66556687395dee0327615b1b015e6b

Observation ea3510d8-9973-4378-8c73-4fed6a197109 · outbound

This paper cites SuperGLUE: A Stickier Benchmark for General-Purpose Language Understanding Systems.

DebateBench: A Challenging Long Context Reasoning Benchmark For Large Language Models SuperGLUE: A Stickier Benchmark for General-Purpose Language Understanding Systems

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-08T16:11:18.587695Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T16:11:18.587695Z digest=sha256:ee5e52e92738ee7cb11ba1570a94046beb99a322874cef319a65cdf1a29b20eb

Observation 9c71265b-4f29-4198-8779-1fd2c153ee48 · outbound

This paper cites HellaSwag: Can a Machine Really Finish Your Sentence?.

DebateBench: A Challenging Long Context Reasoning Benchmark For Large Language Models HellaSwag: Can a Machine Really Finish Your Sentence?

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-08T16:11:18.592741Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T16:11:18.592741Z digest=sha256:528c179cf3642ec484b2e1c9bd7c6e23301df8be7b1e8139adb53dff71ef277d

Observation 73a3bd7c-ae58-4ed2-91b9-a1abe31580eb · outbound

This paper cites Balancing Specialized and General Skills in LLMs: The Impact of Modern Tuning and Data Strategy.

DebateBench: A Challenging Long Context Reasoning Benchmark For Large Language Models Balancing Specialized and General Skills in LLMs: The Impact of Modern Tuning and Data Strategy

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-08T16:11:18.598661Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T16:11:18.598661Z digest=sha256:f0347b6656fb9cec9c3d4feae8bbe2c0b176e393c56fc8d6771846d954487501

Observation 688c9413-e64d-40c8-bd05-822a9ae189b1 · outbound

This paper cites online" 'onlinestring :=.

DebateBench: A Challenging Long Context Reasoning Benchmark For Large Language Models online" 'onlinestring :=

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-08T16:11:18.603687Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T16:11:18.603687Z digest=sha256:8af14d2dcd40138e5e6ccaf928591a100d1bbb9f3056f895da2f5e214d4fd523

Observation a9700a8c-5fee-4d22-afdd-6d5ca9c596d9 · outbound

This paper cites write newline.

DebateBench: A Challenging Long Context Reasoning Benchmark For Large Language Models write newline

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-08T16:11:18.608790Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T16:11:18.608790Z digest=sha256:112e85724dbc5c4fe10b984ded9f77b1c96516f71eb7648c0724afd785c10b9c

Pith citing papers

Observation 5096160d-dc35-4a12-bd7d-d52ca33d0792 · inbound

SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models cites this paper.

SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models DebateBench: A Challenging Long Context Reasoning Benchmark For Large Language Models

Reference 41

Resolution
verified exact
local_arxiv, observed 2026-08-07T12:45:27.847772Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:45:15.313592Z digest=sha256:120ef7dac2bece4bbdbf7f5f6d05d914f32efe3f00ede06f7015528d26c6e69d