Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-08T16:11:18.608790Z
Paper Citation Record · LEDGER
As of 10 August 2026, this Paper Citation Record lists 25 of 25 outbound references and 1 inbound Pith citation observation for arXiv:2502.06279.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-08T16:11:18.608790Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T12:45:15.313592Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-07T12:45:27.776502Z
25 of 25 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation cd8caff9-11fb-4582-8ecc-16047f0ac32b · outbound
DebateBench: A Challenging Long Context Reasoning Benchmark For Large Language Models Have LLMs Advanced Enough? A Challenging Problem Solving Benchmark For Large Language Models
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d4031c7c-5b05-4bda-9964-018e3747ea35 · outbound
DebateBench: A Challenging Long Context Reasoning Benchmark For Large Language Models Program Synthesis with Large Language Models
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 917d2ff8-be1a-4ed0-9420-bdd2cb478dbd · outbound
DebateBench: A Challenging Long Context Reasoning Benchmark For Large Language Models Sparks of Artificial General Intelligence: Early experiments with GPT-4
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d77e5042-c4b6-4cdd-89bf-00c456542f15 · outbound
DebateBench: A Challenging Long Context Reasoning Benchmark For Large Language Models Evaluating Large Language Models Trained on Code
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 10063763-7df0-49db-882a-aa354370fab7 · outbound
DebateBench: A Challenging Long Context Reasoning Benchmark For Large Language Models Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 94fc4e2f-e2dc-43c5-8869-d05148b5e03a · outbound
DebateBench: A Challenging Long Context Reasoning Benchmark For Large Language Models Training Verifiers to Solve Math Word Problems
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9d8a23f7-f692-427d-8402-ec6ee201c926 · outbound
DebateBench: A Challenging Long Context Reasoning Benchmark For Large Language Models A Survey on In-context Learning
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 677caec8-45cd-4a68-a371-e34078015f5e · outbound
DebateBench: A Challenging Long Context Reasoning Benchmark For Large Language Models Unresolved cited work
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2e3f2138-e44b-46d9-a738-fb1e493fde1a · outbound
DebateBench: A Challenging Long Context Reasoning Benchmark For Large Language Models Measuring Massive Multitask Language Understanding
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6e8f89c7-074f-4661-af23-24e73cf61f5a · outbound
DebateBench: A Challenging Long Context Reasoning Benchmark For Large Language Models Measuring Mathematical Problem Solving With the MATH Dataset
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 18711a9e-2ab0-4e71-b4c3-9c3d00785754 · outbound
DebateBench: A Challenging Long Context Reasoning Benchmark For Large Language Models OpenAI o1 System Card
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 43a8759a-9757-47ea-8e2a-8842b90ae72b · outbound
DebateBench: A Challenging Long Context Reasoning Benchmark For Large Language Models ETHIC: Evaluating Large Language Models on Long-Context Tasks with High Information Coverage
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d604e143-5e9c-4847-ba4c-12d71d9e42c0 · outbound
DebateBench: A Challenging Long Context Reasoning Benchmark For Large Language Models TruthfulQA: Measuring How Models Mimic Human Falsehoods
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5e4ad9ce-873a-4e1e-a28b-fec3ba94127f · outbound
DebateBench: A Challenging Long Context Reasoning Benchmark For Large Language Models Unresolved cited work
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d7d359cc-271e-48f2-b038-c0de03b0bcdb · outbound
DebateBench: A Challenging Long Context Reasoning Benchmark For Large Language Models GPT-4 Technical Report
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7d17e15d-de30-4fcd-9749-cf50fd25a879 · outbound
DebateBench: A Challenging Long Context Reasoning Benchmark For Large Language Models Unresolved cited work
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eb7ae997-d196-417d-982a-be9c65607ffc · outbound
DebateBench: A Challenging Long Context Reasoning Benchmark For Large Language Models VivesDebate-Speech: A Corpus of Spoken Argumentation to Leverage Audio Features for Argument Mining
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation ece87034-6b18-40ac-81e5-f3ef7ad81d5a · outbound
DebateBench: A Challenging Long Context Reasoning Benchmark For Large Language Models WinoGrande: An Adversarial Winograd Schema Challenge at Scale
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8e27cc1d-708c-43d0-a8e3-657e5178bde9 · outbound
DebateBench: A Challenging Long Context Reasoning Benchmark For Large Language Models Llama 2: Open Foundation and Fine-Tuned Chat Models
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 72ec0687-a5c1-4e87-a590-60a3940f4dff · outbound
DebateBench: A Challenging Long Context Reasoning Benchmark For Large Language Models Whisper-GPT -- Continuous Discrete Hybrid Representation Language Models For Speech And Music
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ea3510d8-9973-4378-8c73-4fed6a197109 · outbound
DebateBench: A Challenging Long Context Reasoning Benchmark For Large Language Models SuperGLUE: A Stickier Benchmark for General-Purpose Language Understanding Systems
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9c71265b-4f29-4198-8779-1fd2c153ee48 · outbound
DebateBench: A Challenging Long Context Reasoning Benchmark For Large Language Models HellaSwag: Can a Machine Really Finish Your Sentence?
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 73a3bd7c-ae58-4ed2-91b9-a1abe31580eb · outbound
DebateBench: A Challenging Long Context Reasoning Benchmark For Large Language Models Balancing Specialized and General Skills in LLMs: The Impact of Modern Tuning and Data Strategy
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 688c9413-e64d-40c8-bd05-822a9ae189b1 · outbound
DebateBench: A Challenging Long Context Reasoning Benchmark For Large Language Models online" 'onlinestring :=
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a9700a8c-5fee-4d22-afdd-6d5ca9c596d9 · outbound
DebateBench: A Challenging Long Context Reasoning Benchmark For Large Language Models write newline
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5096160d-dc35-4a12-bd7d-d52ca33d0792 · inbound
SocialMaze: A Benchmark for Evaluating Social Reasoning in Large Language Models DebateBench: A Challenging Long Context Reasoning Benchmark For Large Language Models
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.