Pith. sign in

Paper Citation Record · LEDGER

OIBench: Benchmarking Strong Reasoning Models with Olympiad in Informatics

As of 8 August 2026, this Paper Citation Record lists 61 of 61 outbound references and 2 inbound Pith citation observations for arXiv:2506.10481.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.10481 v1

Coverage vector

measured 61 of 61 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T04:33:13.512658Z

measured 63 of 63 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T21:18:28.946759Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T21:18:29.388938Z

Reference resolution

61 of 61 outbound references displayed

  • verified exact1
  • verified fuzzy17
  • unresolved41
  • parse uncertain2
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 677cd943-be5c-4dcf-a8c9-ac7b9834aae1 · outbound

This paper cites Evaluating Large Language Models Trained on Code.

OIBench: Benchmarking Strong Reasoning Models with Olympiad in Informatics Evaluating Large Language Models Trained on Code

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T04:33:07.868175Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:33:07.868175Z digest=sha256:8829e28ba4fb21c9acafbcb12a161c8b42915712e40bd61827c4e7c7fa0591aa

Observation 5fb3d385-2c8a-4fb5-9930-86e32248768a · outbound

This paper cites Program Synthesis with Large Language Models.

OIBench: Benchmarking Strong Reasoning Models with Olympiad in Informatics Program Synthesis with Large Language Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T04:33:07.895300Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:33:07.895300Z digest=sha256:c8c5761de06469a59eb8acd3612f11c9edd56080e8d8021fcd925cdba6f5c4cd

Observation 1eed0699-97f2-441f-8413-d8a27d8de4ee · outbound

This paper cites GSM-Symbolic: Understanding the Limitations of Mathematical Reasoning in Large Language Models.

OIBench: Benchmarking Strong Reasoning Models with Olympiad in Informatics GSM-Symbolic: Understanding the Limitations of Mathematical Reasoning in Large Language Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T04:33:07.938420Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:33:07.938420Z digest=sha256:0da7c1fa2008d51a579aac1cb768aa3440c80a5badac66a4ca7700f7f2c8180b

Observation 8971f75e-77ab-4780-8d59-2d75b01c8113 · outbound

This paper cites On Memorization of Large Language Models in Logical Reasoning.

OIBench: Benchmarking Strong Reasoning Models with Olympiad in Informatics On Memorization of Large Language Models in Logical Reasoning

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T04:33:08.113023Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:33:08.113023Z digest=sha256:2e120d5e9264f068abe5fb8e9b644520ee25f594462e640be8a4a8ec715551a2

Observation 7933ce0e-30ce-49d6-a056-1b2d95da9aec · outbound

This paper cites Competition-Level Code Generation with AlphaCode.

OIBench: Benchmarking Strong Reasoning Models with Olympiad in Informatics Competition-Level Code Generation with AlphaCode

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T04:33:08.236487Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:33:08.236487Z digest=sha256:8e1fe50f554d9de894a32722ccd4be69921c38c541e20dc2dfe7cda188863eed

Observation 89daff14-5780-49b5-be69-a6f71cfab640 · outbound

This paper cites CodeElo: Benchmarking Competition-level Code Generation of LLMs with Human-comparable Elo Ratings.

OIBench: Benchmarking Strong Reasoning Models with Olympiad in Informatics CodeElo: Benchmarking Competition-level Code Generation of LLMs with Human-comparable Elo Ratings

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T04:33:08.345343Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:33:08.345343Z digest=sha256:aad5c3b57f4481effe3625815d1049409a2ad8fc12f74b797e56e8b4c9f4336f

Observation b523538d-5be5-47d3-99ad-e3c9115faf50 · outbound

This paper cites Can Language Models Solve Olympiad Programming?.

OIBench: Benchmarking Strong Reasoning Models with Olympiad in Informatics Can Language Models Solve Olympiad Programming?

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T04:33:08.494278Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:33:08.494278Z digest=sha256:95a71a723f418d02dd48986c8a7027e86c34b326a4c1f96fa1a05c525d251379

Observation e5402a56-4632-4d42-b63f-2a855608d68c · outbound

This paper cites LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code.

OIBench: Benchmarking Strong Reasoning Models with Olympiad in Informatics LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T04:33:08.592196Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:33:08.592196Z digest=sha256:c40e4a1cc0e570a634081f85cdb4a8395ca303664360ca73ea28d3874b8a27a6

Observation 0df138f1-78c1-4e82-8fc4-c382798f3814 · outbound

This paper cites Effibench: Bench- marking the efficiency of automatically generated code.

OIBench: Benchmarking Strong Reasoning Models with Olympiad in Informatics Effibench: Bench- marking the efficiency of automatically generated code

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:33:14.355418Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T04:33:08.686791Z digest=sha256:693fbb109f07606992b74ed2a2858d53ef550946a6f7054166942aca675bddd9

Observation 17bca638-1c3d-4a3d-9784-a2dc78abc840 · outbound

This paper cites A performance study of llm-generated code on leetcode.

OIBench: Benchmarking Strong Reasoning Models with Olympiad in Informatics A performance study of llm-generated code on leetcode

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:33:14.347033Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T04:33:08.803139Z digest=sha256:007651e696f56d7a3723b11d292b649696cbfb299d2037d09187629fb59b9552

Observation 54080233-0ace-4240-9f4a-2f7ce7a3622c · outbound

This paper cites OpenAI o1 System Card.

OIBench: Benchmarking Strong Reasoning Models with Olympiad in Informatics OpenAI o1 System Card

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T04:33:09.009199Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:33:09.009199Z digest=sha256:3af81346e4d78f73899d40b1b4f3e0230855bd7d24a2c8c3b84feafbf60fc813

Observation 642496ed-ad65-45d1-b167-ee0c48c136be · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models.

OIBench: Benchmarking Strong Reasoning Models with Olympiad in Informatics Chain-of-thought prompting elicits reasoning in large language models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T04:33:09.150957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:33:09.150957Z digest=sha256:d75f11a6b9491484f0cd6891030e3cce5d8e17a5d77a15f0558584ba8b1f95ce

Observation 3901de89-a2fe-44eb-8d03-327450d93252 · outbound

This paper cites Towards reasoning in large language models: A survey.

OIBench: Benchmarking Strong Reasoning Models with Olympiad in Informatics Towards reasoning in large language models: A survey

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:33:14.334289Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T04:33:09.261275Z digest=sha256:9db761ecf6c2bcfd2a6ba0829d9681b54b184f3d868c27b6e36c63fbd0a2a43c

Observation b1450cc4-1139-4609-83d5-ce03f3fd2944 · outbound

This paper cites Towards Large Reasoning Models: A Survey of Reinforced Reasoning with Large Language Models.

OIBench: Benchmarking Strong Reasoning Models with Olympiad in Informatics Towards Large Reasoning Models: A Survey of Reinforced Reasoning with Large Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T04:33:09.361597Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:33:09.361597Z digest=sha256:5586cba20b37b55f6f892da11f451b96d0a5730d076a0cd8c5ce3e8896159311

Observation 44bfdff2-8214-40e2-9022-85520e23a5fa · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

OIBench: Benchmarking Strong Reasoning Models with Olympiad in Informatics Training Verifiers to Solve Math Word Problems

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T04:33:09.512190Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:33:09.512190Z digest=sha256:29438706226635df674ff93b85a60df3133bb019990f395b1deda4c53e8bbd9d

Observation 6b3a8478-b429-413d-927c-2b2ce0f6eba6 · outbound

This paper cites Measuring mathematical problem solving with the MATH dataset.

OIBench: Benchmarking Strong Reasoning Models with Olympiad in Informatics Measuring mathematical problem solving with the MATH dataset

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T04:33:09.609005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:33:09.609005Z digest=sha256:a6ae83b2678c2a499f6f78ea5fde81d0d9161e5aae00657b316ffb306199fdf7

Observation f94983e6-4a73-4d5e-ab04-21ddc0139989 · outbound

This paper cites Cohen, Ruslan Salakhut- dinov, and Christopher D.

OIBench: Benchmarking Strong Reasoning Models with Olympiad in Informatics Cohen, Ruslan Salakhut- dinov, and Christopher D

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:33:14.322073Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T04:33:09.793808Z digest=sha256:2cc4a6aa503a43915198b5452b848d700827c21c16d62f7965aeb4f3a60c4212

Observation 5ed49304-8946-4d3d-97d9-ba86e3bd6312 · outbound

This paper cites Logicbench: Towards systematic evaluation of logical reasoning ability of large language models.

OIBench: Benchmarking Strong Reasoning Models with Olympiad in Informatics Logicbench: Towards systematic evaluation of logical reasoning ability of large language models

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:33:14.314247Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T04:33:09.948383Z digest=sha256:591f524e85e05806d7755e996af6d17f0d137d362b3f003d7af5cfc1131fdaa7

Observation fb1688cf-6872-4115-987d-2288f1ad5ade · outbound

This paper cites Criticbench: Benchmarking llms for critique-correct reasoning.

OIBench: Benchmarking Strong Reasoning Models with Olympiad in Informatics Criticbench: Benchmarking llms for critique-correct reasoning

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:33:14.306856Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T04:33:10.071881Z digest=sha256:81b23afceb6608bbc68f158d370f00ff5b95e5359c6ad1d2b1929231f2b6e25e

Observation e6186f45-f387-4eaa-8c0c-0ba4ae4d219d · outbound

This paper cites Inference-Time Computations for LLM Reasoning and Planning: A Benchmark and Insights.

OIBench: Benchmarking Strong Reasoning Models with Olympiad in Informatics Inference-Time Computations for LLM Reasoning and Planning: A Benchmark and Insights

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T04:33:10.196140Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:33:10.196140Z digest=sha256:9b474807253224186d1c2d8f14263346671bb750216e68a331f1afcd7eabb7df

Observation 0b740906-92a7-4fd4-96b5-eacbd9f605c5 · outbound

This paper cites LogicVista: Multimodal LLM Logical Reasoning Benchmark in Visual Contexts.

OIBench: Benchmarking Strong Reasoning Models with Olympiad in Informatics LogicVista: Multimodal LLM Logical Reasoning Benchmark in Visual Contexts

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T04:33:10.271856Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:33:10.271856Z digest=sha256:c3dd443cb173e16e596c10cc2b29c25f749c38280bd8ea18c209f211098d93dc

Observation 257705c7-de3a-4bd4-beb8-34177dabd9b4 · outbound

This paper cites Test of Time: A Benchmark for Evaluating LLMs on Temporal Reasoning.

OIBench: Benchmarking Strong Reasoning Models with Olympiad in Informatics Test of Time: A Benchmark for Evaluating LLMs on Temporal Reasoning

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T04:33:10.361435Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:33:10.361435Z digest=sha256:47cd86c5207a052366b1014d241cef03ed7ebbbf6ec9ad21a5540bde3286743a

Observation 76549ae5-f538-4580-a655-95ee05d74ed8 · outbound

This paper cites LogicGame: Benchmarking Rule-Based Reasoning Abilities of Large Language Models.

OIBench: Benchmarking Strong Reasoning Models with Olympiad in Informatics LogicGame: Benchmarking Rule-Based Reasoning Abilities of Large Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T04:33:10.456529Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:33:10.456529Z digest=sha256:9f0e4ee704731bc2a1ccbb2b44dd878c6e81c9c992deceecf71507d4aff06c33

Observation d3e75ec1-cd0f-449f-b2f6-6dc02a3ad2d9 · outbound

This paper cites Are llms capable of data-based statistical and causal reasoning? benchmarking advanced quantitative reasoning with data.

OIBench: Benchmarking Strong Reasoning Models with Olympiad in Informatics Are llms capable of data-based statistical and causal reasoning? benchmarking advanced quantitative reasoning with data

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:33:14.299145Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T04:33:10.553871Z digest=sha256:8aff2ea59e12adaa5feb15913dedcfa73709866980454559a97d7ecf2a1b221a

Observation d92827bc-67a8-4fd7-b70b-55ec6e095b75 · outbound

This paper cites Codereval: A benchmark of pragmatic code generation with generative pre-trained models.

OIBench: Benchmarking Strong Reasoning Models with Olympiad in Informatics Codereval: A benchmark of pragmatic code generation with generative pre-trained models

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:33:14.291143Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T04:33:10.636451Z digest=sha256:51f2b2607bf86893c868db70c5360684f08a4b7520adcf5781d42257828be1fc

Observation 936321a4-c108-462a-bd8c-590b38ec447c · outbound

This paper cites Beyond Correctness: Benchmarking Multi-dimensional Code Generation for Large Language Models.

OIBench: Benchmarking Strong Reasoning Models with Olympiad in Informatics Beyond Correctness: Benchmarking Multi-dimensional Code Generation for Large Language Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T04:33:10.690842Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:33:10.690842Z digest=sha256:3a522c6fb450b9608df1cdc61b2cdde80832715e202fb876a45c8431207a3c62

Observation 0655c6db-019a-414a-90cd-857c259b7560 · outbound

This paper cites CodeScore: Evaluating Code Generation by Learning Code Execution.

OIBench: Benchmarking Strong Reasoning Models with Olympiad in Informatics CodeScore: Evaluating Code Generation by Learning Code Execution

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T04:33:10.781213Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:33:10.781213Z digest=sha256:3f876e0d855545ee7d682b00d17b054d89d2d649bea3a3ce192c35906351b487

Observation 899dd107-2403-4bd5-81cc-ccb94887e1b4 · outbound

This paper cites Cruxeval: A benchmark for code reasoning, understanding and execution.

OIBench: Benchmarking Strong Reasoning Models with Olympiad in Informatics Cruxeval: A benchmark for code reasoning, understanding and execution

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:33:14.283713Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T04:33:10.889582Z digest=sha256:b23b4842ca0b78041637527a458d0aa0d377d37c5bcd556cfbdf14e9aa9628d2

Observation e2cd57c7-dded-4765-bbc7-79afff3440b4 · outbound

This paper cites an unresolved cited work.

OIBench: Benchmarking Strong Reasoning Models with Olympiad in Informatics Unresolved cited work

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T04:33:10.980550Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:33:10.980550Z digest=sha256:bcfba8bf07f09029e0cde17f49fa4cca47a3ae0af8077933088b5cae8a9284ac

Observation cf56bd4a-4897-40bf-b472-3e3af2586d02 · outbound

This paper cites Isolating Language-Coding from Problem-Solving: Benchmarking LLMs with PseudoEval.

OIBench: Benchmarking Strong Reasoning Models with Olympiad in Informatics Isolating Language-Coding from Problem-Solving: Benchmarking LLMs with PseudoEval

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T04:33:11.064635Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:33:11.064635Z digest=sha256:629f09b305fe7b1715446b49a269040728fd50aa6a5fd42c6d13fb344b96e6ea

Observation d710289c-aca7-4203-b897-827a8eb75b65 · outbound

This paper cites Is Your Benchmark Still Useful? Dynamic Benchmarking for Code Language Models.

OIBench: Benchmarking Strong Reasoning Models with Olympiad in Informatics Is Your Benchmark Still Useful? Dynamic Benchmarking for Code Language Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T04:33:11.149875Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:33:11.149875Z digest=sha256:947269c366c97bf3f9c5f6f79acb117def9c70d3784726411c34b5c536c6adc5

Observation 81670f9e-1a1b-4357-9c81-a8e809fea339 · outbound

This paper cites Dynamic Benchmarking of Reasoning Capabilities in Code Large Language Models Under Data Contamination.

OIBench: Benchmarking Strong Reasoning Models with Olympiad in Informatics Dynamic Benchmarking of Reasoning Capabilities in Code Large Language Models Under Data Contamination

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T04:33:11.223662Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:33:11.223662Z digest=sha256:e9384651825a276120ab7d14af66d8cdc1135a9806e480df01d1e70fc6441a54

Observation 7cbaf7ec-3ed3-4e85-ac20-d6eec217d703 · outbound

This paper cites DICE: Detecting In-distribution Contamination in LLM's Fine-tuning Phase for Math Reasoning.

OIBench: Benchmarking Strong Reasoning Models with Olympiad in Informatics DICE: Detecting In-distribution Contamination in LLM's Fine-tuning Phase for Math Reasoning

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T04:33:11.291003Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:33:11.291003Z digest=sha256:8c6474cd66e7b9ea10af121641ccb3790930dc88ff81ec58cccb10a781e38539

Observation 86c6e912-7fc3-4588-a43e-04df53aa5026 · outbound

This paper cites Skywork: A More Open Bilingual Foundation Model.

OIBench: Benchmarking Strong Reasoning Models with Olympiad in Informatics Skywork: A More Open Bilingual Foundation Model

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T04:33:11.374239Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:33:11.374239Z digest=sha256:053b00c72e5b02c4fe168da845b7c79528f31e2aea69760a4ef53527a86a9262

Observation 9a361e93-bc4a-4fe3-8ff8-29e3ba0a6029 · outbound

This paper cites Benchmarking Benchmark Leakage in Large Language Models.

OIBench: Benchmarking Strong Reasoning Models with Olympiad in Informatics Benchmarking Benchmark Leakage in Large Language Models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T04:33:11.477391Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:33:11.477391Z digest=sha256:deaa406058f1c3273b47b7d2e18ddd3f96de0db2ad29e1e43ffeaafc0ae794fd

Observation 21f92f1a-3f2d-4dfb-aa4f-5b0138120960 · outbound

This paper cites Investigating data contamination in modern benchmarks for large language models.

OIBench: Benchmarking Strong Reasoning Models with Olympiad in Informatics Investigating data contamination in modern benchmarks for large language models

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:33:14.276180Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T04:33:11.551800Z digest=sha256:b8a1bd9724fd032b69ba08d750ced02ec995fce378ee078a7eeb1fd19425f0a1

Observation e1523533-2660-400b-a1a6-ae70109142aa · outbound

This paper cites Benchmark Data Contamination of Large Language Models: A Survey.

OIBench: Benchmarking Strong Reasoning Models with Olympiad in Informatics Benchmark Data Contamination of Large Language Models: A Survey

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T04:33:11.622079Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:33:11.622079Z digest=sha256:e1f150c1ab09d7904eb551c0339c0f4325df523f8d489b7f7cdcf13fb560ddb3

Observation 581e0455-150e-48ab-9d57-5697dd66c467 · outbound

This paper cites Docker: lightweight linux containers for consistent development and deployment.

OIBench: Benchmarking Strong Reasoning Models with Olympiad in Informatics Docker: lightweight linux containers for consistent development and deployment

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:33:14.269060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T04:33:11.694824Z digest=sha256:a0ce72775c93aaeb31b5ca33883b84a30d15bce4f9d9cb420bb17aad9387aa8d

Observation 84463afe-3213-45f6-a19b-96af79e2c121 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

OIBench: Benchmarking Strong Reasoning Models with Olympiad in Informatics DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T04:33:11.777217Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:33:11.777217Z digest=sha256:e113b489eb4de50eb79186df27ef6ff3a91b3a3aaf672649dc0f35c5cc67a9f3

Observation 5c822408-c600-4247-902c-ca8fd4e6e333 · outbound

This paper cites Adding Error Bars to Evals: A Statistical Approach to Language Model Evaluations.

OIBench: Benchmarking Strong Reasoning Models with Olympiad in Informatics Adding Error Bars to Evals: A Statistical Approach to Language Model Evaluations

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T04:33:11.879063Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:33:11.879063Z digest=sha256:4c1156608ed6fac5d8875f35ba36703413e8a2af992bb2540508a548a8ced938

Observation 10f87df3-ad7f-464e-b7ed-19503c52def0 · outbound

This paper cites Nolazco-Flores, Lori Landay, Matthew Thomas Jackson, Paul Röttger, Philip H.

OIBench: Benchmarking Strong Reasoning Models with Olympiad in Informatics Nolazco-Flores, Lori Landay, Matthew Thomas Jackson, Paul Röttger, Philip H

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:33:14.261358Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T04:33:11.977474Z digest=sha256:7686696f0814c504b6cedbd06d15e9a2d84a244c27908098980d0db156188f4a

Observation 78bc563c-3cb6-405f-8318-ba1d6dad94cc · outbound

This paper cites DeepSeek-V3 Technical Report.

OIBench: Benchmarking Strong Reasoning Models with Olympiad in Informatics DeepSeek-V3 Technical Report

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T04:33:12.079334Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:33:12.079334Z digest=sha256:442c5560c7d1380fd7f7577c072d8451e8255afb29d8b03a11787e8274fc361e

Observation fe56bd4e-ee86-4a6d-b0d5-d7659287e70a · outbound

This paper cites Qwen2.5-Coder Technical Report.

OIBench: Benchmarking Strong Reasoning Models with Olympiad in Informatics Qwen2.5-Coder Technical Report

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T04:33:12.172149Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:33:12.172149Z digest=sha256:f4180da7d8d5ca45dfde1c2c2852c6dd9582dfd45f1537e909fa02463493bf84

Observation 0f70dad5-8764-4312-90b2-39fd99cf8493 · outbound

This paper cites Qwen2.5 Technical Report.

OIBench: Benchmarking Strong Reasoning Models with Olympiad in Informatics Qwen2.5 Technical Report

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T04:33:12.250698Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:33:12.250698Z digest=sha256:dd1e0b469657ec3a3ac7063df9f445dbea90952f27fafe709783277f7caa6fc6

Observation 0529cc92-095b-44cb-8316-c9290a998c21 · outbound

This paper cites The Llama 3 Herd of Models.

OIBench: Benchmarking Strong Reasoning Models with Olympiad in Informatics The Llama 3 Herd of Models

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T04:33:12.323405Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:33:12.323405Z digest=sha256:638e0e258ec138ff7ab27aaf5200367398fb14c28242b818bdf6ed6c90cf4b52

Observation 3aeae70a-eee1-41d7-9a3a-e4ddcdf67436 · outbound

This paper cites doubao-pro-32k.

OIBench: Benchmarking Strong Reasoning Models with Olympiad in Informatics doubao-pro-32k

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:33:14.253388Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T04:33:12.399148Z digest=sha256:75c298a854a403902a16d769acc89d591e7d0f0d4adf1101fb1ebf91281812f8

Observation cc5cf503-ffa5-4de0-bbac-16aa997f02ab · outbound

This paper cites GPT-4o System Card.

OIBench: Benchmarking Strong Reasoning Models with Olympiad in Informatics GPT-4o System Card

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T04:33:12.478817Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:33:12.478817Z digest=sha256:464edbd049f9a1b7c1b439292626cd4f61dda401cc1f6d1d517c5d1e5edd7922

Observation f8db0d76-cfd0-479b-b6cf-7caf3d627a57 · outbound

This paper cites claude 3.5 sonnet.

OIBench: Benchmarking Strong Reasoning Models with Olympiad in Informatics claude 3.5 sonnet

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:33:14.238282Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T04:33:12.507631Z digest=sha256:a639c832f9ad12f59a074082f012224a04279cfe4408059aee11dd6777ad37cf

Observation 403f9abe-e5ea-4cf3-bc42-4d22cfb671a1 · outbound

This paper cites Qwq-32b: Embracing the power of reinforcement learning, March 2025.

OIBench: Benchmarking Strong Reasoning Models with Olympiad in Informatics Qwq-32b: Embracing the power of reinforcement learning, March 2025

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:33:14.230706Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T04:33:12.617668Z digest=sha256:6ca73b6d770aa75cc2884883ce57b2aa669d106a88b96af6d283eeb830f6df03

Observation 285ab06e-1f0e-4e19-b97e-d93f3d7f8337 · outbound

This paper cites an unresolved cited work.

OIBench: Benchmarking Strong Reasoning Models with Olympiad in Informatics Unresolved cited work

Reference 51

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:33:14.222489Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T04:33:12.706993Z digest=sha256:fe418157173f85629f9890db662adaeabf8fade25fda85cc50032b8141bbc6a1

Observation 9c620678-4793-4185-978f-c60b35414ddc · outbound

This paper cites an unresolved cited work.

OIBench: Benchmarking Strong Reasoning Models with Olympiad in Informatics Unresolved cited work

Reference 52

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:33:14.214256Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T04:33:12.833113Z digest=sha256:458ca4f6588d2f8cf8018cf329697aabb7ca85e5d32afc2ad51ff3dc02baebaf

Observation deed81ee-5c08-42de-9023-ce48a68ffa60 · outbound

This paper cites Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters.

OIBench: Benchmarking Strong Reasoning Models with Olympiad in Informatics Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T04:33:13.126641Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:33:13.126641Z digest=sha256:5de6a1f5fb0f610a40dc789725f0583604ab38b11b0ae149e6c36c51f8df5a17

Observation 60544d7f-15a2-489a-8b3c-118c1fc619ac · outbound

This paper cites an unresolved cited work.

OIBench: Benchmarking Strong Reasoning Models with Olympiad in Informatics Unresolved cited work

Reference 54

Resolution
parse uncertain
raw_fallback, observed 2026-08-07T04:33:14.206611Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T04:33:12.976100Z digest=sha256:a038a23c9cf14b1fd4b1e754d2118593f1b561f4b5620a62de6fcb5b32a19fd1

Observation 19fc4506-4f00-48ae-b270-3aa91ee352fc · outbound

This paper cites Planning-Driven Programming: A Large Language Model Programming Workflow.

OIBench: Benchmarking Strong Reasoning Models with Olympiad in Informatics Planning-Driven Programming: A Large Language Model Programming Workflow

Reference 55

Resolution
verified exact
local_arxiv, observed 2026-08-07T04:33:13.651606Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T04:33:13.290319Z digest=sha256:2636dc55cb9e8627e392f7a35fb4425592cc1ae4e45f3c41d4d24bda29a2365f

Observation 715587bd-377a-4a17-973e-dbeb65973865 · outbound

This paper cites SWE-RL: Advancing LLM Reasoning via Reinforcement Learning on Open Software Evolution.

OIBench: Benchmarking Strong Reasoning Models with Olympiad in Informatics SWE-RL: Advancing LLM Reasoning via Reinforcement Learning on Open Software Evolution

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T04:33:13.242807Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:33:13.242807Z digest=sha256:e164d4c01eec8aa256206fb067637bfb241434541dcf1f0534e3279f86a7fb1b

Observation 27f0621f-2e87-4b49-b354-a7e7f545caa3 · outbound

This paper cites MdEval: Massively Multilingual Code Debugging.

OIBench: Benchmarking Strong Reasoning Models with Olympiad in Informatics MdEval: Massively Multilingual Code Debugging

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T04:33:13.471509Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:33:13.471509Z digest=sha256:0ec5bd6c6c4e7581912503809e1e95cb9a1a28c8191f7093e3b2e2bbbff41c63

Observation fbc1b3c7-92aa-45d0-916b-6e48fd573c5a · outbound

This paper cites Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism.

OIBench: Benchmarking Strong Reasoning Models with Olympiad in Informatics Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T04:33:13.374423Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:33:13.374423Z digest=sha256:0245879065d4da9e7aa8fb8ab5cb2b51df5d02b59aef6d4af9c82b0c9fedbb4f

Observation 4113a733-6537-474f-8a8e-bc155c22e733 · outbound

This paper cites Codetransocean: A comprehensive multilingual benchmark for code translation.

OIBench: Benchmarking Strong Reasoning Models with Olympiad in Informatics Codetransocean: A comprehensive multilingual benchmark for code translation

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:33:14.191077Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T04:33:13.509440Z digest=sha256:598ec0f8e993d2e6a483c5c9b9993ad6c1e021fd34ea014935f463856b6adb10

Observation 946b8748-aacf-4824-b769-b34ddf9071d9 · outbound

This paper cites an unresolved cited work.

OIBench: Benchmarking Strong Reasoning Models with Olympiad in Informatics Unresolved cited work

Reference 60

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:33:14.199024Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T04:33:13.480183Z digest=sha256:0756a0dbd08f800368d93bef1512b6521d787a8e8bccc31fcd42f8235c23c98f

Observation fc7e4390-da56-4d66-955b-7a41fb5b3a58 · outbound

This paper cites Gonzalez, Hao Zhang, and Ion Stoica.

OIBench: Benchmarking Strong Reasoning Models with Olympiad in Informatics Gonzalez, Hao Zhang, and Ion Stoica

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:33:14.174519Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T04:33:13.512658Z digest=sha256:16764bef7d169b1d5b35d7629815625fba790818d38cf72171dc1e2d5ee3bc68

Observation 9991e6ca-2d3c-4e5b-b22d-0a670a1e87ba · outbound

This paper cites an unresolved cited work.

OIBench: Benchmarking Strong Reasoning Models with Olympiad in Informatics Unresolved cited work

Reference 2025

Resolution
parse uncertain
raw_fallback, observed 2026-08-07T04:33:14.245758Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T04:33:12.468028Z digest=sha256:461341309231a12cfcad2bdaa7004514c213a0d5d908f50925b2ab5685f84517

Pith citing papers

Observation 13979a02-a595-4766-af05-c8b0caeed694 · inbound

AutoCodeBench: Large Language Models are Automatic Code Benchmark Generators cites this paper.

AutoCodeBench: Large Language Models are Automatic Code Benchmark Generators OIBench: Benchmarking Strong Reasoning Models with Olympiad in Informatics

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-08-05T21:18:29.439278Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-05T21:18:28.946759Z digest=sha256:19522a9c0ae40965f846e9192f72988d2acdb71818e13c87f57b611c4fb0b98b

Observation dbbbccce-7a7a-4d22-ad08-be68d7ac57e0 · inbound

UniCode: Augmenting Evaluation for Code Reasoning cites this paper.

UniCode: Augmenting Evaluation for Code Reasoning OIBench: Benchmarking Strong Reasoning Models with Olympiad in Informatics

Reference 2023

Resolution
malformed identifier
no resolver link, observed 2026-08-04T09:39:47.742161Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:39:47.742161Z digest=sha256:2cf488b3b0e90576c01798b1f79aa1e3604ed4bed78fd17f0c844e380fb2c2a9