Pith. sign in

Paper Citation Record · LEDGER

OIBench: Benchmarking Strong Reasoning Models with Olympiad in Informatics

As of 13 August 2026, this Paper Citation Record lists 61 of 61 outbound references and 2 inbound Pith citation observations for arXiv:2506.10481.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.10481 v1

Coverage vector

measured 61 of 61 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T04:33:13.512658Z

measured 63 of 63 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T21:18:28.946759Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T21:18:29.388938Z

Reference resolution

61 of 61 outbound references displayed

  • verified exact1
  • verified fuzzy17
  • unresolved41
  • parse uncertain2
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 677cd943-be5c-4dcf-a8c9-ac7b9834aae1 · outbound

This paper cites Evaluating Large Language Models Trained on Code.

OIBench: Benchmarking Strong Reasoning Models with Olympiad in Informatics Evaluating Large Language Models Trained on Code

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T04:33:07.868175Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:33:07.868175Z digest=sha256:610095bc424f519428e3d33aaad07e2b6046982377ece3eafc4b771f9dd72a02

Observation 5fb3d385-2c8a-4fb5-9930-86e32248768a · outbound

This paper cites Program Synthesis with Large Language Models.

OIBench: Benchmarking Strong Reasoning Models with Olympiad in Informatics Program Synthesis with Large Language Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T04:33:07.895300Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:33:07.895300Z digest=sha256:5cfa17fbaf3c0eab258e0c12b4297f19bf9d91d6319a9eb3ff96d5461b09b284

Observation 1eed0699-97f2-441f-8413-d8a27d8de4ee · outbound

This paper cites GSM-Symbolic: Understanding the Limitations of Mathematical Reasoning in Large Language Models.

OIBench: Benchmarking Strong Reasoning Models with Olympiad in Informatics GSM-Symbolic: Understanding the Limitations of Mathematical Reasoning in Large Language Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T04:33:07.938420Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:33:07.938420Z digest=sha256:aa490a1d9a5f8862edece08936364bf57b0465d40756c21125867b198e5d8e78

Observation 8971f75e-77ab-4780-8d59-2d75b01c8113 · outbound

This paper cites On Memorization of Large Language Models in Logical Reasoning.

OIBench: Benchmarking Strong Reasoning Models with Olympiad in Informatics On Memorization of Large Language Models in Logical Reasoning

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T04:33:08.113023Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:33:08.113023Z digest=sha256:d478955ff8e1089a02a8ee014ac1509cca4c16c1f301bd667c070df69d00a197

Observation 7933ce0e-30ce-49d6-a056-1b2d95da9aec · outbound

This paper cites Competition-Level Code Generation with AlphaCode.

OIBench: Benchmarking Strong Reasoning Models with Olympiad in Informatics Competition-Level Code Generation with AlphaCode

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T04:33:08.236487Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:33:08.236487Z digest=sha256:d12dd5b689d117bf0b574b33c2a732023216ee4aa35afe423e22cfe2e0d7adaa

Observation 89daff14-5780-49b5-be69-a6f71cfab640 · outbound

This paper cites CodeElo: Benchmarking Competition-level Code Generation of LLMs with Human-comparable Elo Ratings.

OIBench: Benchmarking Strong Reasoning Models with Olympiad in Informatics CodeElo: Benchmarking Competition-level Code Generation of LLMs with Human-comparable Elo Ratings

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T04:33:08.345343Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:33:08.345343Z digest=sha256:e26b1c89cf04e3d05bb8a99b960645e826eb761fae741cab6516c9f740666a7f

Observation b523538d-5be5-47d3-99ad-e3c9115faf50 · outbound

This paper cites Can Language Models Solve Olympiad Programming?.

OIBench: Benchmarking Strong Reasoning Models with Olympiad in Informatics Can Language Models Solve Olympiad Programming?

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T04:33:08.494278Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:33:08.494278Z digest=sha256:527c3516c6d0b45bda725c863e59c6c3ee96fa846d9d1842039ac61144a00f5d

Observation e5402a56-4632-4d42-b63f-2a855608d68c · outbound

This paper cites LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code.

OIBench: Benchmarking Strong Reasoning Models with Olympiad in Informatics LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T04:33:08.592196Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:33:08.592196Z digest=sha256:ec6d23fe9318a092f5dc898d35c3e0cfb9e8951b295eb61722477a81b8e25a66

Observation 0df138f1-78c1-4e82-8fc4-c382798f3814 · outbound

This paper cites Effibench: Bench- marking the efficiency of automatically generated code.

OIBench: Benchmarking Strong Reasoning Models with Olympiad in Informatics Effibench: Bench- marking the efficiency of automatically generated code

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:33:14.355418Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T04:33:08.686791Z digest=sha256:47637f3932a5bef27db1464421bc6226ee74cf2de310e6da944aa0ead83be60a

Observation 17bca638-1c3d-4a3d-9784-a2dc78abc840 · outbound

This paper cites A performance study of llm-generated code on leetcode.

OIBench: Benchmarking Strong Reasoning Models with Olympiad in Informatics A performance study of llm-generated code on leetcode

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:33:14.347033Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T04:33:08.803139Z digest=sha256:898f64cd108294edca6add30b669d964bac6febbaca0a7e54bf31c347f66ffd1

Observation 54080233-0ace-4240-9f4a-2f7ce7a3622c · outbound

This paper cites OpenAI o1 System Card.

OIBench: Benchmarking Strong Reasoning Models with Olympiad in Informatics OpenAI o1 System Card

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T04:33:09.009199Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:33:09.009199Z digest=sha256:11e1e9f81cbe2bce16b593d125515d008fb5a28075455879e5fcd50248dd67e6

Observation 642496ed-ad65-45d1-b167-ee0c48c136be · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models.

OIBench: Benchmarking Strong Reasoning Models with Olympiad in Informatics Chain-of-thought prompting elicits reasoning in large language models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T04:33:09.150957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:33:09.150957Z digest=sha256:85e33843aa3051cfd532fdb4cbe4d6713dc7e5a9e4b07b27fd6930cc07046509

Observation 3901de89-a2fe-44eb-8d03-327450d93252 · outbound

This paper cites Towards reasoning in large language models: A survey.

OIBench: Benchmarking Strong Reasoning Models with Olympiad in Informatics Towards reasoning in large language models: A survey

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:33:14.334289Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T04:33:09.261275Z digest=sha256:f01419ce7d2d6d52cbb23ce5dc4d4579fc5bf6cde25e742f3b00d359d96d6423

Observation b1450cc4-1139-4609-83d5-ce03f3fd2944 · outbound

This paper cites Towards Large Reasoning Models: A Survey of Reinforced Reasoning with Large Language Models.

OIBench: Benchmarking Strong Reasoning Models with Olympiad in Informatics Towards Large Reasoning Models: A Survey of Reinforced Reasoning with Large Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T04:33:09.361597Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:33:09.361597Z digest=sha256:37796424670858e576d788c6a347fa4cce089b1ca027a91e0c2fb4a5e04a1cdc

Observation 44bfdff2-8214-40e2-9022-85520e23a5fa · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

OIBench: Benchmarking Strong Reasoning Models with Olympiad in Informatics Training Verifiers to Solve Math Word Problems

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T04:33:09.512190Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:33:09.512190Z digest=sha256:4be1f9ef79e09074184452ff9b643988d594a690075078b4493b3ec7bfc817bd

Observation 6b3a8478-b429-413d-927c-2b2ce0f6eba6 · outbound

This paper cites Measuring mathematical problem solving with the MATH dataset.

OIBench: Benchmarking Strong Reasoning Models with Olympiad in Informatics Measuring mathematical problem solving with the MATH dataset

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T04:33:09.609005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:33:09.609005Z digest=sha256:453daa1456f4a4fbadb14bce3050e10ba6ea32e0a53d37e9835d89e9b4af3440

Observation f94983e6-4a73-4d5e-ab04-21ddc0139989 · outbound

This paper cites Cohen, Ruslan Salakhut- dinov, and Christopher D.

OIBench: Benchmarking Strong Reasoning Models with Olympiad in Informatics Cohen, Ruslan Salakhut- dinov, and Christopher D

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:33:14.322073Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T04:33:09.793808Z digest=sha256:549b1545deea43b4cd749e544e748e1778f99f26d086c63ba167a4346c13b963

Observation 5ed49304-8946-4d3d-97d9-ba86e3bd6312 · outbound

This paper cites Logicbench: Towards systematic evaluation of logical reasoning ability of large language models.

OIBench: Benchmarking Strong Reasoning Models with Olympiad in Informatics Logicbench: Towards systematic evaluation of logical reasoning ability of large language models

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:33:14.314247Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T04:33:09.948383Z digest=sha256:a34c3ce74401e7a5cfe9e0a3b80fe3d547a3b3a8e85ad6f9729b13340dfa7000

Observation fb1688cf-6872-4115-987d-2288f1ad5ade · outbound

This paper cites Criticbench: Benchmarking llms for critique-correct reasoning.

OIBench: Benchmarking Strong Reasoning Models with Olympiad in Informatics Criticbench: Benchmarking llms for critique-correct reasoning

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:33:14.306856Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T04:33:10.071881Z digest=sha256:4d954d80e01ca0e28b37cb47cc10751978ea276b2cbd6f0d21f23aa092bdccab

Observation e6186f45-f387-4eaa-8c0c-0ba4ae4d219d · outbound

This paper cites Inference-Time Computations for LLM Reasoning and Planning: A Benchmark and Insights.

OIBench: Benchmarking Strong Reasoning Models with Olympiad in Informatics Inference-Time Computations for LLM Reasoning and Planning: A Benchmark and Insights

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T04:33:10.196140Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:33:10.196140Z digest=sha256:e23dae36f15594ef3b4c8b580392af3209a402388db94eb3d8a847c6c47f37cc

Observation 0b740906-92a7-4fd4-96b5-eacbd9f605c5 · outbound

This paper cites LogicVista: Multimodal LLM Logical Reasoning Benchmark in Visual Contexts.

OIBench: Benchmarking Strong Reasoning Models with Olympiad in Informatics LogicVista: Multimodal LLM Logical Reasoning Benchmark in Visual Contexts

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T04:33:10.271856Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:33:10.271856Z digest=sha256:f7dc5a9ab36f0057066478507c99504ade2c966c40a865eb393d1b929c6a47bb

Observation 257705c7-de3a-4bd4-beb8-34177dabd9b4 · outbound

This paper cites Test of Time: A Benchmark for Evaluating LLMs on Temporal Reasoning.

OIBench: Benchmarking Strong Reasoning Models with Olympiad in Informatics Test of Time: A Benchmark for Evaluating LLMs on Temporal Reasoning

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T04:33:10.361435Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:33:10.361435Z digest=sha256:86be0598894d2e5f68cc329df2fbbc888e134168b0905b4da771f322fe0943d2

Observation 76549ae5-f538-4580-a655-95ee05d74ed8 · outbound

This paper cites LogicGame: Benchmarking Rule-Based Reasoning Abilities of Large Language Models.

OIBench: Benchmarking Strong Reasoning Models with Olympiad in Informatics LogicGame: Benchmarking Rule-Based Reasoning Abilities of Large Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T04:33:10.456529Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:33:10.456529Z digest=sha256:60ec8d6adf4c7192fcf0c16a12f827fcdfc272836bdf697b73488660da2952be

Observation d3e75ec1-cd0f-449f-b2f6-6dc02a3ad2d9 · outbound

This paper cites Are llms capable of data-based statistical and causal reasoning? benchmarking advanced quantitative reasoning with data.

OIBench: Benchmarking Strong Reasoning Models with Olympiad in Informatics Are llms capable of data-based statistical and causal reasoning? benchmarking advanced quantitative reasoning with data

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:33:14.299145Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T04:33:10.553871Z digest=sha256:42179314a0a379812c3638959b672c8fb210a8c0b60db78be2a112cc19f9cc3e

Observation d92827bc-67a8-4fd7-b70b-55ec6e095b75 · outbound

This paper cites Codereval: A benchmark of pragmatic code generation with generative pre-trained models.

OIBench: Benchmarking Strong Reasoning Models with Olympiad in Informatics Codereval: A benchmark of pragmatic code generation with generative pre-trained models

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:33:14.291143Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T04:33:10.636451Z digest=sha256:5521eb5c70be07ed064a159085bfdbb6cd5c2a97a4f762421b2bb6a98e306033

Observation 936321a4-c108-462a-bd8c-590b38ec447c · outbound

This paper cites Beyond Correctness: Benchmarking Multi-dimensional Code Generation for Large Language Models.

OIBench: Benchmarking Strong Reasoning Models with Olympiad in Informatics Beyond Correctness: Benchmarking Multi-dimensional Code Generation for Large Language Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T04:33:10.690842Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:33:10.690842Z digest=sha256:64c1271d23fb24bbf0b09abf8ef105187b7bbd490018c75cd007b79015673a45

Observation 0655c6db-019a-414a-90cd-857c259b7560 · outbound

This paper cites CodeScore: Evaluating Code Generation by Learning Code Execution.

OIBench: Benchmarking Strong Reasoning Models with Olympiad in Informatics CodeScore: Evaluating Code Generation by Learning Code Execution

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T04:33:10.781213Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:33:10.781213Z digest=sha256:11b4951294ce0c5b91fe69d2daed88a3a376885bf33f541e6fb506eb249a753a

Observation 899dd107-2403-4bd5-81cc-ccb94887e1b4 · outbound

This paper cites Cruxeval: A benchmark for code reasoning, understanding and execution.

OIBench: Benchmarking Strong Reasoning Models with Olympiad in Informatics Cruxeval: A benchmark for code reasoning, understanding and execution

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:33:14.283713Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T04:33:10.889582Z digest=sha256:2bf8465b066367a73bd49d1681462764b039f74bf1b6b138a94fd5c93f273a69

Observation e2cd57c7-dded-4765-bbc7-79afff3440b4 · outbound

This paper cites an unresolved cited work.

OIBench: Benchmarking Strong Reasoning Models with Olympiad in Informatics Unresolved cited work

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T04:33:10.980550Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:33:10.980550Z digest=sha256:4b6b966563d2c80f98f3e83aa29cdc63cb7cc0d004c6ee694f35ec15eaa31ac5

Observation cf56bd4a-4897-40bf-b472-3e3af2586d02 · outbound

This paper cites Isolating Language-Coding from Problem-Solving: Benchmarking LLMs with PseudoEval.

OIBench: Benchmarking Strong Reasoning Models with Olympiad in Informatics Isolating Language-Coding from Problem-Solving: Benchmarking LLMs with PseudoEval

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T04:33:11.064635Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:33:11.064635Z digest=sha256:79b36358163e795ba56b14c451feac0add970a114b5a7da3bffcb2c7a6f29761

Observation d710289c-aca7-4203-b897-827a8eb75b65 · outbound

This paper cites Is Your Benchmark Still Useful? Dynamic Benchmarking for Code Language Models.

OIBench: Benchmarking Strong Reasoning Models with Olympiad in Informatics Is Your Benchmark Still Useful? Dynamic Benchmarking for Code Language Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T04:33:11.149875Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:33:11.149875Z digest=sha256:2b60268e7b2e389796878c6ecb76a37d4d8e4b07ca7e29e9d064dcbe0cc62d2a

Observation 81670f9e-1a1b-4357-9c81-a8e809fea339 · outbound

This paper cites Dynamic Benchmarking of Reasoning Capabilities in Code Large Language Models Under Data Contamination.

OIBench: Benchmarking Strong Reasoning Models with Olympiad in Informatics Dynamic Benchmarking of Reasoning Capabilities in Code Large Language Models Under Data Contamination

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T04:33:11.223662Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:33:11.223662Z digest=sha256:3eb48a9dde34e86ab8a924c074779c09acdced6a2fabb28460cdd734830462bb

Observation 7cbaf7ec-3ed3-4e85-ac20-d6eec217d703 · outbound

This paper cites DICE: Detecting In-distribution Contamination in LLM's Fine-tuning Phase for Math Reasoning.

OIBench: Benchmarking Strong Reasoning Models with Olympiad in Informatics DICE: Detecting In-distribution Contamination in LLM's Fine-tuning Phase for Math Reasoning

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T04:33:11.291003Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:33:11.291003Z digest=sha256:017e33eaf79e4e2efa368b0074b57b8083ffa4a7478cd96e46161e55b98c85ef

Observation 86c6e912-7fc3-4588-a43e-04df53aa5026 · outbound

This paper cites Skywork: A More Open Bilingual Foundation Model.

OIBench: Benchmarking Strong Reasoning Models with Olympiad in Informatics Skywork: A More Open Bilingual Foundation Model

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T04:33:11.374239Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:33:11.374239Z digest=sha256:e4f321016162a12dcc18ceafa31815275e8754ab0b63c07f23fd1d98cf6e5dfe

Observation 9a361e93-bc4a-4fe3-8ff8-29e3ba0a6029 · outbound

This paper cites Benchmarking Benchmark Leakage in Large Language Models.

OIBench: Benchmarking Strong Reasoning Models with Olympiad in Informatics Benchmarking Benchmark Leakage in Large Language Models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T04:33:11.477391Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:33:11.477391Z digest=sha256:8bd1238fe74e5cd3bb20685dac59aed43db3b1b55a4119567e5d5f87ff0be78b

Observation 21f92f1a-3f2d-4dfb-aa4f-5b0138120960 · outbound

This paper cites Investigating data contamination in modern benchmarks for large language models.

OIBench: Benchmarking Strong Reasoning Models with Olympiad in Informatics Investigating data contamination in modern benchmarks for large language models

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:33:14.276180Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T04:33:11.551800Z digest=sha256:1df28a80a5d46603d71f469d05c8161df65b10b063bb9aa14dd4ca4ae48b5b6c

Observation e1523533-2660-400b-a1a6-ae70109142aa · outbound

This paper cites Benchmark Data Contamination of Large Language Models: A Survey.

OIBench: Benchmarking Strong Reasoning Models with Olympiad in Informatics Benchmark Data Contamination of Large Language Models: A Survey

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T04:33:11.622079Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:33:11.622079Z digest=sha256:420cd95888f6c62bfb682addb3ff5d387ffb9568d536d8e0b85ded1a41ab186c

Observation 581e0455-150e-48ab-9d57-5697dd66c467 · outbound

This paper cites Docker: lightweight linux containers for consistent development and deployment.

OIBench: Benchmarking Strong Reasoning Models with Olympiad in Informatics Docker: lightweight linux containers for consistent development and deployment

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:33:14.269060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T04:33:11.694824Z digest=sha256:f81adb9fba86706c540ff91452e871d8e4f209d8a7724e02ddfa4f9572971179

Observation 84463afe-3213-45f6-a19b-96af79e2c121 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

OIBench: Benchmarking Strong Reasoning Models with Olympiad in Informatics DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T04:33:11.777217Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:33:11.777217Z digest=sha256:69dbfb6a2ff77d3c7c936e1c5a7479cd08ac6aada4ed1c1ace2e4a64ecd93855

Observation 5c822408-c600-4247-902c-ca8fd4e6e333 · outbound

This paper cites Adding Error Bars to Evals: A Statistical Approach to Language Model Evaluations.

OIBench: Benchmarking Strong Reasoning Models with Olympiad in Informatics Adding Error Bars to Evals: A Statistical Approach to Language Model Evaluations

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T04:33:11.879063Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:33:11.879063Z digest=sha256:de83d1dbd03b6b87db6e7dd831ec55b791ce443b820d015e94b89498d0e47222

Observation 10f87df3-ad7f-464e-b7ed-19503c52def0 · outbound

This paper cites Nolazco-Flores, Lori Landay, Matthew Thomas Jackson, Paul Röttger, Philip H.

OIBench: Benchmarking Strong Reasoning Models with Olympiad in Informatics Nolazco-Flores, Lori Landay, Matthew Thomas Jackson, Paul Röttger, Philip H

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:33:14.261358Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T04:33:11.977474Z digest=sha256:3f527c92e609c361829fed43f3154efcb5001e7b3a8cd327982758ce163ab4a7

Observation 78bc563c-3cb6-405f-8318-ba1d6dad94cc · outbound

This paper cites DeepSeek-V3 Technical Report.

OIBench: Benchmarking Strong Reasoning Models with Olympiad in Informatics DeepSeek-V3 Technical Report

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T04:33:12.079334Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:33:12.079334Z digest=sha256:897f79265631674dcc6c432e98e179a295a25c22b8eefa53d0ed2907045ab957

Observation fe56bd4e-ee86-4a6d-b0d5-d7659287e70a · outbound

This paper cites Qwen2.5-Coder Technical Report.

OIBench: Benchmarking Strong Reasoning Models with Olympiad in Informatics Qwen2.5-Coder Technical Report

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T04:33:12.172149Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:33:12.172149Z digest=sha256:9006d32da13b806f6c5f6dd5ab20f6ed868127cf24b50d5219b89dfce3cd5830

Observation 0f70dad5-8764-4312-90b2-39fd99cf8493 · outbound

This paper cites Qwen2.5 Technical Report.

OIBench: Benchmarking Strong Reasoning Models with Olympiad in Informatics Qwen2.5 Technical Report

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T04:33:12.250698Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:33:12.250698Z digest=sha256:2c05a399210d529b9256c9ed32fa0337e139f5a5970f0fa1da23bad0e30278a9

Observation 0529cc92-095b-44cb-8316-c9290a998c21 · outbound

This paper cites The Llama 3 Herd of Models.

OIBench: Benchmarking Strong Reasoning Models with Olympiad in Informatics The Llama 3 Herd of Models

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T04:33:12.323405Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:33:12.323405Z digest=sha256:16f523999cedc58d3a93cc4ce924128b25f478999210c0a6f07e614c200dffde

Observation 3aeae70a-eee1-41d7-9a3a-e4ddcdf67436 · outbound

This paper cites doubao-pro-32k.

OIBench: Benchmarking Strong Reasoning Models with Olympiad in Informatics doubao-pro-32k

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:33:14.253388Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T04:33:12.399148Z digest=sha256:52ddf2d78d5497aa136f041dd2bda0826f21a3aefa82326133071697c1b466a4

Observation cc5cf503-ffa5-4de0-bbac-16aa997f02ab · outbound

This paper cites GPT-4o System Card.

OIBench: Benchmarking Strong Reasoning Models with Olympiad in Informatics GPT-4o System Card

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T04:33:12.478817Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:33:12.478817Z digest=sha256:b764a28d533ba0d61de98d0fefb805ec6cecd7c2dfd07131e334b0d3b936ebb5

Observation f8db0d76-cfd0-479b-b6cf-7caf3d627a57 · outbound

This paper cites claude 3.5 sonnet.

OIBench: Benchmarking Strong Reasoning Models with Olympiad in Informatics claude 3.5 sonnet

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:33:14.238282Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T04:33:12.507631Z digest=sha256:c4512bf336b3193a87735e90358b71de65b0956bc23844c1535af9ec1fc42d6a

Observation 403f9abe-e5ea-4cf3-bc42-4d22cfb671a1 · outbound

This paper cites Qwq-32b: Embracing the power of reinforcement learning, March 2025.

OIBench: Benchmarking Strong Reasoning Models with Olympiad in Informatics Qwq-32b: Embracing the power of reinforcement learning, March 2025

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:33:14.230706Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T04:33:12.617668Z digest=sha256:765b41595551ac8cb4dfca3f8bbf08357359a5beb87cec57d48749b96c6fb36c

Observation 285ab06e-1f0e-4e19-b97e-d93f3d7f8337 · outbound

This paper cites an unresolved cited work.

OIBench: Benchmarking Strong Reasoning Models with Olympiad in Informatics Unresolved cited work

Reference 51

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:33:14.222489Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T04:33:12.706993Z digest=sha256:fb68730bd8b2899e78d70676dc3eb3dd07517409ae9418fed5c7f34b578f7ef4

Observation 9c620678-4793-4185-978f-c60b35414ddc · outbound

This paper cites an unresolved cited work.

OIBench: Benchmarking Strong Reasoning Models with Olympiad in Informatics Unresolved cited work

Reference 52

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:33:14.214256Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T04:33:12.833113Z digest=sha256:7ee525311bd55d5650cc2b9b1db3b52a0aed7050eff33d38516b0cc0ec16a814

Observation deed81ee-5c08-42de-9023-ce48a68ffa60 · outbound

This paper cites Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters.

OIBench: Benchmarking Strong Reasoning Models with Olympiad in Informatics Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T04:33:13.126641Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:33:13.126641Z digest=sha256:07382f77a40a0403835a687df95bfdbbb19f3b73ddb942eadd0e01a576d98adf

Observation 60544d7f-15a2-489a-8b3c-118c1fc619ac · outbound

This paper cites an unresolved cited work.

OIBench: Benchmarking Strong Reasoning Models with Olympiad in Informatics Unresolved cited work

Reference 54

Resolution
parse uncertain
raw_fallback, observed 2026-08-07T04:33:14.206611Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T04:33:12.976100Z digest=sha256:9d69226c20afcad1e862148d9cb422a599bdc3464ee56a347bc818811ce599d8

Observation 19fc4506-4f00-48ae-b270-3aa91ee352fc · outbound

This paper cites Planning-Driven Programming: A Large Language Model Programming Workflow.

OIBench: Benchmarking Strong Reasoning Models with Olympiad in Informatics Planning-Driven Programming: A Large Language Model Programming Workflow

Reference 55

Resolution
verified exact
local_arxiv, observed 2026-08-07T04:33:13.651606Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T04:33:13.290319Z digest=sha256:46e9f95be6dffc26f017941d62646c8b35c431ec7bc3a1ce454f6125b20e1e0f

Observation 715587bd-377a-4a17-973e-dbeb65973865 · outbound

This paper cites SWE-RL: Advancing LLM Reasoning via Reinforcement Learning on Open Software Evolution.

OIBench: Benchmarking Strong Reasoning Models with Olympiad in Informatics SWE-RL: Advancing LLM Reasoning via Reinforcement Learning on Open Software Evolution

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T04:33:13.242807Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:33:13.242807Z digest=sha256:aad679342da508663995ef663a21323ef71cc4f7eaeb52ae26624f42eda282d0

Observation 27f0621f-2e87-4b49-b354-a7e7f545caa3 · outbound

This paper cites MdEval: Massively Multilingual Code Debugging.

OIBench: Benchmarking Strong Reasoning Models with Olympiad in Informatics MdEval: Massively Multilingual Code Debugging

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T04:33:13.471509Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:33:13.471509Z digest=sha256:283d72bf87103793cb2d6251583f02233a384801e6e422fe27464a8458f18454

Observation fbc1b3c7-92aa-45d0-916b-6e48fd573c5a · outbound

This paper cites Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism.

OIBench: Benchmarking Strong Reasoning Models with Olympiad in Informatics Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T04:33:13.374423Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:33:13.374423Z digest=sha256:e766cf62ae95f85e9d965e0c499b9ef50667683d0f385c084360df18c38de94a

Observation 4113a733-6537-474f-8a8e-bc155c22e733 · outbound

This paper cites Codetransocean: A comprehensive multilingual benchmark for code translation.

OIBench: Benchmarking Strong Reasoning Models with Olympiad in Informatics Codetransocean: A comprehensive multilingual benchmark for code translation

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:33:14.191077Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T04:33:13.509440Z digest=sha256:27317f952e2a6d6d3cc621e9488718ebd41ecb4e5c12f02bd92ebd38e645a48f

Observation 946b8748-aacf-4824-b769-b34ddf9071d9 · outbound

This paper cites an unresolved cited work.

OIBench: Benchmarking Strong Reasoning Models with Olympiad in Informatics Unresolved cited work

Reference 60

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:33:14.199024Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T04:33:13.480183Z digest=sha256:eddcd165f19a8122fa6b1f9e54e766d183c1b7cde8430ddce87e871b30626189

Observation fc7e4390-da56-4d66-955b-7a41fb5b3a58 · outbound

This paper cites Gonzalez, Hao Zhang, and Ion Stoica.

OIBench: Benchmarking Strong Reasoning Models with Olympiad in Informatics Gonzalez, Hao Zhang, and Ion Stoica

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:33:14.174519Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T04:33:13.512658Z digest=sha256:d958cb85e84d1817a15d4b7513befe651f89bf140106f6aa7445a65f131381b1

Observation 9991e6ca-2d3c-4e5b-b22d-0a670a1e87ba · outbound

This paper cites an unresolved cited work.

OIBench: Benchmarking Strong Reasoning Models with Olympiad in Informatics Unresolved cited work

Reference 2025

Resolution
parse uncertain
raw_fallback, observed 2026-08-07T04:33:14.245758Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-07T04:33:12.468028Z digest=sha256:3b55234de2ede314908af9e6ff19002e78366b005aed78868b27724152a2bb5f

Pith citing papers

Observation 13979a02-a595-4766-af05-c8b0caeed694 · inbound

AutoCodeBench: Large Language Models are Automatic Code Benchmark Generators cites this paper.

AutoCodeBench: Large Language Models are Automatic Code Benchmark Generators OIBench: Benchmarking Strong Reasoning Models with Olympiad in Informatics

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-08-05T21:18:29.439278Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-05T21:18:28.946759Z digest=sha256:6882272161f8bd37c89f0853aa487008604d626c57e92693417e700944697a05

Observation dbbbccce-7a7a-4d22-ad08-be68d7ac57e0 · inbound

UniCode: Augmenting Evaluation for Code Reasoning cites this paper.

UniCode: Augmenting Evaluation for Code Reasoning OIBench: Benchmarking Strong Reasoning Models with Olympiad in Informatics

Reference 2023

Resolution
malformed identifier
no resolver link, observed 2026-08-04T09:39:47.742161Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:39:47.742161Z digest=sha256:9fb064f397ce334441df04c2a41bc51834e2c08c13c248e3b5c4cbd55551a854