Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T04:33:13.512658Z
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 61 of 61 outbound references and 2 inbound Pith citation observations for arXiv:2506.10481.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T04:33:13.512658Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-05T21:18:28.946759Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-05T21:18:29.388938Z
61 of 61 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 677cd943-be5c-4dcf-a8c9-ac7b9834aae1 · outbound
OIBench: Benchmarking Strong Reasoning Models with Olympiad in Informatics Evaluating Large Language Models Trained on Code
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5fb3d385-2c8a-4fb5-9930-86e32248768a · outbound
OIBench: Benchmarking Strong Reasoning Models with Olympiad in Informatics Program Synthesis with Large Language Models
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1eed0699-97f2-441f-8413-d8a27d8de4ee · outbound
OIBench: Benchmarking Strong Reasoning Models with Olympiad in Informatics GSM-Symbolic: Understanding the Limitations of Mathematical Reasoning in Large Language Models
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8971f75e-77ab-4780-8d59-2d75b01c8113 · outbound
OIBench: Benchmarking Strong Reasoning Models with Olympiad in Informatics On Memorization of Large Language Models in Logical Reasoning
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7933ce0e-30ce-49d6-a056-1b2d95da9aec · outbound
OIBench: Benchmarking Strong Reasoning Models with Olympiad in Informatics Competition-Level Code Generation with AlphaCode
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 89daff14-5780-49b5-be69-a6f71cfab640 · outbound
OIBench: Benchmarking Strong Reasoning Models with Olympiad in Informatics CodeElo: Benchmarking Competition-level Code Generation of LLMs with Human-comparable Elo Ratings
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b523538d-5be5-47d3-99ad-e3c9115faf50 · outbound
OIBench: Benchmarking Strong Reasoning Models with Olympiad in Informatics Can Language Models Solve Olympiad Programming?
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e5402a56-4632-4d42-b63f-2a855608d68c · outbound
OIBench: Benchmarking Strong Reasoning Models with Olympiad in Informatics LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0df138f1-78c1-4e82-8fc4-c382798f3814 · outbound
OIBench: Benchmarking Strong Reasoning Models with Olympiad in Informatics Effibench: Bench- marking the efficiency of automatically generated code
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 17bca638-1c3d-4a3d-9784-a2dc78abc840 · outbound
OIBench: Benchmarking Strong Reasoning Models with Olympiad in Informatics A performance study of llm-generated code on leetcode
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 54080233-0ace-4240-9f4a-2f7ce7a3622c · outbound
OIBench: Benchmarking Strong Reasoning Models with Olympiad in Informatics OpenAI o1 System Card
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 642496ed-ad65-45d1-b167-ee0c48c136be · outbound
OIBench: Benchmarking Strong Reasoning Models with Olympiad in Informatics Chain-of-thought prompting elicits reasoning in large language models
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3901de89-a2fe-44eb-8d03-327450d93252 · outbound
OIBench: Benchmarking Strong Reasoning Models with Olympiad in Informatics Towards reasoning in large language models: A survey
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b1450cc4-1139-4609-83d5-ce03f3fd2944 · outbound
OIBench: Benchmarking Strong Reasoning Models with Olympiad in Informatics Towards Large Reasoning Models: A Survey of Reinforced Reasoning with Large Language Models
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 44bfdff2-8214-40e2-9022-85520e23a5fa · outbound
OIBench: Benchmarking Strong Reasoning Models with Olympiad in Informatics Training Verifiers to Solve Math Word Problems
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6b3a8478-b429-413d-927c-2b2ce0f6eba6 · outbound
OIBench: Benchmarking Strong Reasoning Models with Olympiad in Informatics Measuring mathematical problem solving with the MATH dataset
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f94983e6-4a73-4d5e-ab04-21ddc0139989 · outbound
OIBench: Benchmarking Strong Reasoning Models with Olympiad in Informatics Cohen, Ruslan Salakhut- dinov, and Christopher D
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5ed49304-8946-4d3d-97d9-ba86e3bd6312 · outbound
OIBench: Benchmarking Strong Reasoning Models with Olympiad in Informatics Logicbench: Towards systematic evaluation of logical reasoning ability of large language models
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation fb1688cf-6872-4115-987d-2288f1ad5ade · outbound
OIBench: Benchmarking Strong Reasoning Models with Olympiad in Informatics Criticbench: Benchmarking llms for critique-correct reasoning
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e6186f45-f387-4eaa-8c0c-0ba4ae4d219d · outbound
OIBench: Benchmarking Strong Reasoning Models with Olympiad in Informatics Inference-Time Computations for LLM Reasoning and Planning: A Benchmark and Insights
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0b740906-92a7-4fd4-96b5-eacbd9f605c5 · outbound
OIBench: Benchmarking Strong Reasoning Models with Olympiad in Informatics LogicVista: Multimodal LLM Logical Reasoning Benchmark in Visual Contexts
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 257705c7-de3a-4bd4-beb8-34177dabd9b4 · outbound
OIBench: Benchmarking Strong Reasoning Models with Olympiad in Informatics Test of Time: A Benchmark for Evaluating LLMs on Temporal Reasoning
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 76549ae5-f538-4580-a655-95ee05d74ed8 · outbound
OIBench: Benchmarking Strong Reasoning Models with Olympiad in Informatics LogicGame: Benchmarking Rule-Based Reasoning Abilities of Large Language Models
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d3e75ec1-cd0f-449f-b2f6-6dc02a3ad2d9 · outbound
OIBench: Benchmarking Strong Reasoning Models with Olympiad in Informatics Are llms capable of data-based statistical and causal reasoning? benchmarking advanced quantitative reasoning with data
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d92827bc-67a8-4fd7-b70b-55ec6e095b75 · outbound
OIBench: Benchmarking Strong Reasoning Models with Olympiad in Informatics Codereval: A benchmark of pragmatic code generation with generative pre-trained models
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 936321a4-c108-462a-bd8c-590b38ec447c · outbound
OIBench: Benchmarking Strong Reasoning Models with Olympiad in Informatics Beyond Correctness: Benchmarking Multi-dimensional Code Generation for Large Language Models
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0655c6db-019a-414a-90cd-857c259b7560 · outbound
OIBench: Benchmarking Strong Reasoning Models with Olympiad in Informatics CodeScore: Evaluating Code Generation by Learning Code Execution
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 899dd107-2403-4bd5-81cc-ccb94887e1b4 · outbound
OIBench: Benchmarking Strong Reasoning Models with Olympiad in Informatics Cruxeval: A benchmark for code reasoning, understanding and execution
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e2cd57c7-dded-4765-bbc7-79afff3440b4 · outbound
OIBench: Benchmarking Strong Reasoning Models with Olympiad in Informatics Unresolved cited work
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cf56bd4a-4897-40bf-b472-3e3af2586d02 · outbound
OIBench: Benchmarking Strong Reasoning Models with Olympiad in Informatics Isolating Language-Coding from Problem-Solving: Benchmarking LLMs with PseudoEval
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d710289c-aca7-4203-b897-827a8eb75b65 · outbound
OIBench: Benchmarking Strong Reasoning Models with Olympiad in Informatics Is Your Benchmark Still Useful? Dynamic Benchmarking for Code Language Models
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 81670f9e-1a1b-4357-9c81-a8e809fea339 · outbound
OIBench: Benchmarking Strong Reasoning Models with Olympiad in Informatics Dynamic Benchmarking of Reasoning Capabilities in Code Large Language Models Under Data Contamination
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7cbaf7ec-3ed3-4e85-ac20-d6eec217d703 · outbound
OIBench: Benchmarking Strong Reasoning Models with Olympiad in Informatics DICE: Detecting In-distribution Contamination in LLM's Fine-tuning Phase for Math Reasoning
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 86c6e912-7fc3-4588-a43e-04df53aa5026 · outbound
OIBench: Benchmarking Strong Reasoning Models with Olympiad in Informatics Skywork: A More Open Bilingual Foundation Model
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9a361e93-bc4a-4fe3-8ff8-29e3ba0a6029 · outbound
OIBench: Benchmarking Strong Reasoning Models with Olympiad in Informatics Benchmarking Benchmark Leakage in Large Language Models
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 21f92f1a-3f2d-4dfb-aa4f-5b0138120960 · outbound
OIBench: Benchmarking Strong Reasoning Models with Olympiad in Informatics Investigating data contamination in modern benchmarks for large language models
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e1523533-2660-400b-a1a6-ae70109142aa · outbound
OIBench: Benchmarking Strong Reasoning Models with Olympiad in Informatics Benchmark Data Contamination of Large Language Models: A Survey
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 581e0455-150e-48ab-9d57-5697dd66c467 · outbound
OIBench: Benchmarking Strong Reasoning Models with Olympiad in Informatics Docker: lightweight linux containers for consistent development and deployment
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 84463afe-3213-45f6-a19b-96af79e2c121 · outbound
OIBench: Benchmarking Strong Reasoning Models with Olympiad in Informatics DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5c822408-c600-4247-902c-ca8fd4e6e333 · outbound
OIBench: Benchmarking Strong Reasoning Models with Olympiad in Informatics Adding Error Bars to Evals: A Statistical Approach to Language Model Evaluations
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 10f87df3-ad7f-464e-b7ed-19503c52def0 · outbound
OIBench: Benchmarking Strong Reasoning Models with Olympiad in Informatics Nolazco-Flores, Lori Landay, Matthew Thomas Jackson, Paul Röttger, Philip H
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 78bc563c-3cb6-405f-8318-ba1d6dad94cc · outbound
OIBench: Benchmarking Strong Reasoning Models with Olympiad in Informatics DeepSeek-V3 Technical Report
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fe56bd4e-ee86-4a6d-b0d5-d7659287e70a · outbound
OIBench: Benchmarking Strong Reasoning Models with Olympiad in Informatics Qwen2.5-Coder Technical Report
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0f70dad5-8764-4312-90b2-39fd99cf8493 · outbound
OIBench: Benchmarking Strong Reasoning Models with Olympiad in Informatics Qwen2.5 Technical Report
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0529cc92-095b-44cb-8316-c9290a998c21 · outbound
OIBench: Benchmarking Strong Reasoning Models with Olympiad in Informatics The Llama 3 Herd of Models
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3aeae70a-eee1-41d7-9a3a-e4ddcdf67436 · outbound
OIBench: Benchmarking Strong Reasoning Models with Olympiad in Informatics doubao-pro-32k
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation cc5cf503-ffa5-4de0-bbac-16aa997f02ab · outbound
OIBench: Benchmarking Strong Reasoning Models with Olympiad in Informatics GPT-4o System Card
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f8db0d76-cfd0-479b-b6cf-7caf3d627a57 · outbound
OIBench: Benchmarking Strong Reasoning Models with Olympiad in Informatics claude 3.5 sonnet
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 403f9abe-e5ea-4cf3-bc42-4d22cfb671a1 · outbound
OIBench: Benchmarking Strong Reasoning Models with Olympiad in Informatics Qwq-32b: Embracing the power of reinforcement learning, March 2025
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 285ab06e-1f0e-4e19-b97e-d93f3d7f8337 · outbound
OIBench: Benchmarking Strong Reasoning Models with Olympiad in Informatics Unresolved cited work
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 9c620678-4793-4185-978f-c60b35414ddc · outbound
OIBench: Benchmarking Strong Reasoning Models with Olympiad in Informatics Unresolved cited work
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation deed81ee-5c08-42de-9023-ce48a68ffa60 · outbound
OIBench: Benchmarking Strong Reasoning Models with Olympiad in Informatics Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 60544d7f-15a2-489a-8b3c-118c1fc619ac · outbound
OIBench: Benchmarking Strong Reasoning Models with Olympiad in Informatics Unresolved cited work
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 19fc4506-4f00-48ae-b270-3aa91ee352fc · outbound
OIBench: Benchmarking Strong Reasoning Models with Olympiad in Informatics Planning-Driven Programming: A Large Language Model Programming Workflow
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 715587bd-377a-4a17-973e-dbeb65973865 · outbound
OIBench: Benchmarking Strong Reasoning Models with Olympiad in Informatics SWE-RL: Advancing LLM Reasoning via Reinforcement Learning on Open Software Evolution
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 27f0621f-2e87-4b49-b354-a7e7f545caa3 · outbound
OIBench: Benchmarking Strong Reasoning Models with Olympiad in Informatics MdEval: Massively Multilingual Code Debugging
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fbc1b3c7-92aa-45d0-916b-6e48fd573c5a · outbound
OIBench: Benchmarking Strong Reasoning Models with Olympiad in Informatics Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4113a733-6537-474f-8a8e-bc155c22e733 · outbound
OIBench: Benchmarking Strong Reasoning Models with Olympiad in Informatics Codetransocean: A comprehensive multilingual benchmark for code translation
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 946b8748-aacf-4824-b769-b34ddf9071d9 · outbound
OIBench: Benchmarking Strong Reasoning Models with Olympiad in Informatics Unresolved cited work
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation fc7e4390-da56-4d66-955b-7a41fb5b3a58 · outbound
OIBench: Benchmarking Strong Reasoning Models with Olympiad in Informatics Gonzalez, Hao Zhang, and Ion Stoica
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 9991e6ca-2d3c-4e5b-b22d-0a670a1e87ba · outbound
OIBench: Benchmarking Strong Reasoning Models with Olympiad in Informatics Unresolved cited work
Reference 2025
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 13979a02-a595-4766-af05-c8b0caeed694 · inbound
AutoCodeBench: Large Language Models are Automatic Code Benchmark Generators OIBench: Benchmarking Strong Reasoning Models with Olympiad in Informatics
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation dbbbccce-7a7a-4d22-ad08-be68d7ac57e0 · inbound
UniCode: Augmenting Evaluation for Code Reasoning OIBench: Benchmarking Strong Reasoning Models with Olympiad in Informatics
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.