Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-05T15:51:32.176410Z
Paper Citation Record · LEDGER
As of 9 August 2026, this Paper Citation Record lists 32 of 32 outbound references and 3 inbound Pith citation observations for arXiv:2508.19363.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-05T15:51:32.176410Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-07-14T17:12:24.565155Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-01T17:15:51.605247Z
32 of 32 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 11ef3de4-451f-432f-94a9-e05fefabb761 · outbound
LongReasonArena: A Long Reasoning Benchmark for Large Language Models L-eval: Instituting standardized evaluation for long context language models
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 83b9d2d6-6fee-483c-b6d4-057613265663 · outbound
LongReasonArena: A Long Reasoning Benchmark for Large Language Models Claude 3.7 sonnet system card
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f04d89fb-afaa-4542-8e93-f7255cdc5847 · outbound
LongReasonArena: A Long Reasoning Benchmark for Large Language Models Program synthesis with large language models, 2021
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9e3bdcfd-9c47-4dff-ba7d-6eeeda6934bf · outbound
LongReasonArena: A Long Reasoning Benchmark for Large Language Models Longbench: A bilingual, multitask benchmark for long context understanding
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 37d218c6-2947-44f0-99b2-e246952933e1 · outbound
LongReasonArena: A Long Reasoning Benchmark for Large Language Models Unresolved cited work
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 22942380-c5c6-4ad2-93db-3d973e3a7bd7 · outbound
LongReasonArena: A Long Reasoning Benchmark for Large Language Models LR B ench: Evaluating long-chain reflective reasoning capabilities of large language models via constraint satisfaction problems
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 9086b855-b637-48ae-8fab-424d8e9deb8a · outbound
LongReasonArena: A Long Reasoning Benchmark for Large Language Models Unresolved cited work
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation da4897a4-a170-4645-ba63-efc47f0ed0a0 · outbound
LongReasonArena: A Long Reasoning Benchmark for Large Language Models Zhang, Hanwei Xu, Hao Yang, Haowei Zhang, Honghui Ding, Huajian Xin, Huazuo Gao, Hui Li, Hui Qu, J
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation b56c4e72-9614-4507-9c6a-3c851a8248c7 · outbound
LongReasonArena: A Long Reasoning Benchmark for Large Language Models The Llama 3 Herd of Models
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9d3662b6-92dc-4e8f-bda7-ff9e42432c56 · outbound
LongReasonArena: A Long Reasoning Benchmark for Large Language Models Measuring coding challenge competence with apps
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 007f3dc9-ac80-485e-b61e-b99be5be8f9c · outbound
LongReasonArena: A Long Reasoning Benchmark for Large Language Models Ruler: What’s the real context size of your long-context language models? In First Conference on Language Modeling , 2024
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 8e8f6709-e8a3-46eb-8d97-95febb8296d2 · outbound
LongReasonArena: A Long Reasoning Benchmark for Large Language Models Qwen2.5-Coder Technical Report
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aa31b8e0-2957-483c-b73d-525ff64a29ce · outbound
LongReasonArena: A Long Reasoning Benchmark for Large Language Models Livecodebench: Holistic and contamination free evaluation of large language models for code
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 82128b88-5087-4324-a59e-86ccf3fc497a · outbound
LongReasonArena: A Long Reasoning Benchmark for Large Language Models Needle In A Haystack - pressure testing LLM s
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation cde42c8e-5be1-4a33-ad16-f3fdeee6ab84 · outbound
LongReasonArena: A Long Reasoning Benchmark for Large Language Models Babilong: Testing the limits of llms with long context reasoning-in-a-haystack
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 587cae72-04bb-4b6c-a35d-1a70eee7464c · outbound
LongReasonArena: A Long Reasoning Benchmark for Large Language Models Efficient memory management for large language model serving with pagedattention
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 7369c73c-db9a-48ec-82b6-ce5e909688fc · outbound
LongReasonArena: A Long Reasoning Benchmark for Large Language Models Longgenbench: Long-context generation benchmark
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation d530507a-6eb7-49d2-8df3-d70b33244b23 · outbound
LongReasonArena: A Long Reasoning Benchmark for Large Language Models Codei/o: Condensing reasoning patterns via code input-output prediction, 2025
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation afd5edd6-3041-45d2-a19e-bfe4605685c2 · outbound
LongReasonArena: A Long Reasoning Benchmark for Large Language Models Longreason: A synthetic long-context reasoning benchmark via context expansion
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fb36d692-747d-4768-a391-949e6895bfd3 · outbound
LongReasonArena: A Long Reasoning Benchmark for Large Language Models ALR$^2$: A Retrieve-then-Reason Framework for Long-context Question Answering
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0bbb19dd-f518-456a-895a-dd56813da96e · outbound
LongReasonArena: A Long Reasoning Benchmark for Large Language Models Large Language Models as Code Executors: An Exploratory Study
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 62d65c44-5632-4c15-96dd-8b952bc853cf · outbound
LongReasonArena: A Long Reasoning Benchmark for Large Language Models GPT-4o System Card
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 07dddfad-b035-405d-9d03-1cec5bda189b · outbound
LongReasonArena: A Long Reasoning Benchmark for Large Language Models OpenAI o1 System Card
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 68c2adb3-59fc-4371-99b8-a3b3b0b3d0e8 · outbound
LongReasonArena: A Long Reasoning Benchmark for Large Language Models QwQ-32B : Embracing the power of reinforcement learning, March 2025
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 51b21b92-873c-4ea0-a5cb-cdc7dd413440 · outbound
LongReasonArena: A Long Reasoning Benchmark for Large Language Models Longgenbench: Benchmarking long-form generation in long context llms, 2025
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation ba883100-1b0f-45d7-b249-02cb092fba8e · outbound
LongReasonArena: A Long Reasoning Benchmark for Large Language Models Thoughts are all over the place: On the underthinking of o1-like llms, 2025
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation c6f8f972-9d78-458d-8b70-9aa23eaf3561 · outbound
LongReasonArena: A Long Reasoning Benchmark for Large Language Models Chain-of-thought prompting elicits reasoning in large language models
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 2d3677ff-fa16-463b-99da-667f05f9e955 · outbound
LongReasonArena: A Long Reasoning Benchmark for Large Language Models Tree of thoughts: Deliberate problem solving with large language models
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation ed79a072-bcfb-4a41-bc25-554558b72305 · outbound
LongReasonArena: A Long Reasoning Benchmark for Large Language Models Qwen2.5 Technical Report
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c9be9688-3096-4da2-9b6f-978494eda5b0 · outbound
LongReasonArena: A Long Reasoning Benchmark for Large Language Models bench: Extending long context evaluation beyond 100k tokens
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation c9c1372d-36af-4af5-b030-0f44c37b9f30 · outbound
LongReasonArena: A Long Reasoning Benchmark for Large Language Models DocPuzzle: A Process-Aware Benchmark for Evaluating Realistic Long-Context Reasoning Capabilities
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 2448123f-d8e5-4f46-ada2-a791e302adcd · outbound
LongReasonArena: A Long Reasoning Benchmark for Large Language Models GSM-Infinite: How Do Your LLMs Behave over Infinitely Increasing Context Length and Reasoning Complexity?
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 90c34bf1-150e-42b7-9342-24f0518e0937 · inbound
Efficient Evaluation of LLM Performance with Statistical Guarantees LongReasonArena: A Long Reasoning Benchmark for Large Language Models
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 64e5264d-9d28-4c1e-91f4-950aeb31e9a3 · inbound
Cognitive Episodes in LLM Reasoning Traces Enable Interpretable Human Item Difficulty Prediction LongReasonArena: A Long Reasoning Benchmark for Large Language Models
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation b887b421-a269-4587-a575-485520251693 · inbound
Cognitive Episodes in LLM Reasoning Traces Enable Interpretable Human Item Difficulty Prediction LongReasonArena: A Long Reasoning Benchmark for Large Language Models
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.