Pith. sign in

Paper Citation Record · LEDGER

ResearchCodeBench: Benchmarking LLMs on Implementing Novel Machine Learning Research Code

As of 22 August 2026, this Paper Citation Record lists 11 of 11 outbound references and 17 inbound Pith citation observations for arXiv:2506.02314.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.02314 v1

Coverage vector

measured 11 of 11 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:30:04.614811Z

measured 28 of 28 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 17 of 17 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T15:24:46.242048Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-29T12:23:24.671544Z

Reference resolution

11 of 11 outbound references displayed

  • verified exact0
  • verified fuzzy2
  • unresolved8
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 691111d9-f02b-4677-9914-91bc4770ad7d · outbound

This paper cites an unresolved cited work.

ResearchCodeBench: Benchmarking LLMs on Implementing Novel Machine Learning Research Code Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:30:05.291244Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T11:30:04.158768Z digest=sha256:c5985148e2ea3f83ecd9346aed5d3a278d249382ab8d2624113cd97f0361a360

Observation 2ddd7919-e2b9-479b-afd5-50c3228893df · outbound

This paper cites calculate area.

ResearchCodeBench: Benchmarking LLMs on Implementing Novel Machine Learning Research Code calculate area

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:30:05.132937Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T11:30:04.233185Z digest=sha256:d4f61d9fcb246b6f7e06fa5a349291d7b8b921caf9a3b446b0db9cfc27952f7d

Observation 66794533-d14b-4e1b-9c2e-d2ffd3814057 · outbound

This paper cites an unresolved cited work.

ResearchCodeBench: Benchmarking LLMs on Implementing Novel Machine Learning Research Code Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:30:05.012101Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T11:30:04.285489Z digest=sha256:937b14e0a1cbff7da303b7d6dc7e62de3038e22102bd3fa90a09c00b0dfdc4ae

Observation 25375523-4700-4625-b181-7f40447c8193 · outbound

This paper cites 19 Here is the code that you need to complete: {context_code_str + masked_code_str} Please implement the missing code in the TODO blocks.

ResearchCodeBench: Benchmarking LLMs on Implementing Novel Machine Learning Research Code 19 Here is the code that you need to complete: {context_code_str + masked_code_str} Please implement the missing code in the TODO blocks

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:30:04.891715Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T11:30:04.324269Z digest=sha256:6a148c0ea6c95c3c760aeda97280dee5ad80b04afcc43552cd6e33b3c3aacef0

Observation 5953f810-81cf-4204-a698-f1788d9337c7 · outbound

This paper cites an unresolved cited work.

ResearchCodeBench: Benchmarking LLMs on Implementing Novel Machine Learning Research Code Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:30:05.837972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T11:30:04.386780Z digest=sha256:4efe37e23c0a8c5811d4e6182fa67db3398aee9569f07eb722e5be89e220f09e

Observation 89d98f07-0444-41d4-9ab2-b95c1ec61730 · outbound

This paper cites an unresolved cited work.

ResearchCodeBench: Benchmarking LLMs on Implementing Novel Machine Learning Research Code Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:30:05.730723Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T11:30:04.444609Z digest=sha256:80d02ececee9c8ce22d5bc6440cdfc719420a5a1c4d696c248023f3b66ee4d79

Observation c8fcdc46-aa79-4c0b-929c-710df3a816f8 · outbound

This paper cites an unresolved cited work.

ResearchCodeBench: Benchmarking LLMs on Implementing Novel Machine Learning Research Code Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:30:05.570551Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T11:30:04.490287Z digest=sha256:52d706a4468a6a1f4149b4a70721d131d42e04856d0242e7bb5050f863200367

Observation 8074b593-43c7-4a83-a570-8071abf0a365 · outbound

This paper cites an unresolved cited work.

ResearchCodeBench: Benchmarking LLMs on Implementing Novel Machine Learning Research Code Unresolved cited work

Reference 14

Resolution
unresolved
raw_fallback, observed 2026-08-07T11:30:05.437871Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T11:30:04.567217Z digest=sha256:68c5dfb630d76dbcadea617849d046e00c452da069101d9d960a8590bb9254d0

Observation 5825ffe0-a1f3-4f1c-af74-b36e3b47bd51 · outbound

This paper cites an unresolved cited work.

ResearchCodeBench: Benchmarking LLMs on Implementing Novel Machine Learning Research Code Unresolved cited work

Reference 15

Resolution
malformed identifier
raw_fallback, observed 2026-08-07T11:30:04.772792Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T11:30:04.614811Z digest=sha256:8db95399f9d6fee42ee2f09bef254e83103e4924da76c14e440979680f91f269

Observation f75f7814-d361-42b0-a987-39bff45af3f1 · outbound

This paper cites MCTS-Judge: Test-Time Scaling in LLM-as-a-Judge for Code Correctness Evaluation.

ResearchCodeBench: Benchmarking LLMs on Implementing Novel Machine Learning Research Code MCTS-Judge: Test-Time Scaling in LLM-as-a-Judge for Code Correctness Evaluation

Reference 1118

Resolution
unresolved
no resolver link, observed 2026-08-07T11:30:03.875329Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:30:03.875329Z digest=sha256:5ec92d6bf54f5165f5da09379b2ce19e5570652ec3862499b650ee8cb24e16b0

Observation b3a50432-2799-4222-8072-24ab5db5925e · outbound

This paper cites The AI Review Lottery: Widespread AI-Assisted Peer Reviews Boost Paper Scores and Acceptance Rates.

ResearchCodeBench: Benchmarking LLMs on Implementing Novel Machine Learning Research Code The AI Review Lottery: Widespread AI-Assisted Peer Reviews Boost Paper Scores and Acceptance Rates

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-07T11:30:03.829875Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:30:03.829875Z digest=sha256:f2f535b17fef14bc24eb14817b0d8d0fe0986eabe2fd97992b7a2237eafc455d

Pith citing papers

Observation b555cb75-8a55-487b-8429-0ff88055c6cd · inbound

The Ideation-Execution Gap: Execution Outcomes of LLM-Generated versus Human Research Ideas cites this paper.

The Ideation-Execution Gap: Execution Outcomes of LLM-Generated versus Human Research Ideas ResearchCodeBench: Benchmarking LLMs on Implementing Novel Machine Learning Research Code

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T22:46:32.360665Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:46:32.360665Z digest=sha256:edcc1a4e938855770f04064267fcdd09f1a7efb4c1b2a060ae8d4af8b8fb7458

Observation 46c50fdc-078a-42df-9163-041d793d16ef · inbound

Robot builds a robot's brain: AI generated drone command and control station hosted in the sky cites this paper.

Robot builds a robot's brain: AI generated drone command and control station hosted in the sky ResearchCodeBench: Benchmarking LLMs on Implementing Novel Machine Learning Research Code

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T04:51:20.803345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:51:20.803345Z digest=sha256:418e868c653ef5dcbccf5fbf5f92d851bf28d6da1cecd5f2f4e5d5d9d374671b

Observation 3fdb24da-5aa2-4e20-9e2d-5d85a0516e6a · inbound

TASE: Token Awareness and Structured Evaluation for Multilingual Language Models cites this paper.

TASE: Token Awareness and Structured Evaluation for Multilingual Language Models ResearchCodeBench: Benchmarking LLMs on Implementing Novel Machine Learning Research Code

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-05T23:21:22.599676Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:21:22.599676Z digest=sha256:98b34500a1d56f2249e4b6b62ba0c6e5cd545cbf72e2f2b28bbdda0ad8501fdb

Observation 8845b873-a0d6-4b73-a5a1-001b7f67e920 · inbound

CodeDistiller: Automatically Generating Code Libraries for Scientific Coding Agents cites this paper.

CodeDistiller: Automatically Generating Code Libraries for Scientific Coding Agents ResearchCodeBench: Benchmarking LLMs on Implementing Novel Machine Learning Research Code

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-21T18:15:28.008514Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-21T18:15:13.924995Z digest=sha256:2e5752634a51a259563a5f999bf940a6c228976069e0e79778ed7c87ef596f60

Observation 4ad5c576-5f75-422e-89b7-6473640b0e45 · inbound

ReplicatorBench: Benchmarking LLM Agents for Replicability in Social and Behavioral Sciences cites this paper.

ReplicatorBench: Benchmarking LLM Agents for Replicability in Social and Behavioral Sciences ResearchCodeBench: Benchmarking LLMs on Implementing Novel Machine Learning Research Code

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-16T02:02:06.935796Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-16T02:01:07.555124Z digest=sha256:063b3926c925f9ad824c19232099221e031ee5bb65dcdf3baa30faa7bb38c9b9

Observation 7d50b9dd-91ab-491b-b104-8c950183688f · inbound

FactReview: Evidence-Grounded Reviews with Literature Positioning and Execution-Based Claim Verification cites this paper.

FactReview: Evidence-Grounded Reviews with Literature Positioning and Execution-Based Claim Verification ResearchCodeBench: Benchmarking LLMs on Implementing Novel Machine Learning Research Code

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-13T17:33:02.271560Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-13T17:32:06.864338Z digest=sha256:7e7aabd4e2a0f6794db978a409f9deb7f82690e2414e1bc72fd54638c92718cc

Observation 9827d15f-5293-4b57-aa18-9223ba4c8115 · inbound

SciPredict: Can LLMs Predict the Outcomes of Scientific Experiments in Natural Sciences? cites this paper.

SciPredict: Can LLMs Predict the Outcomes of Scientific Experiments in Natural Sciences? ResearchCodeBench: Benchmarking LLMs on Implementing Novel Machine Learning Research Code

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-11T09:36:03.778992Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-10T15:55:34.768853Z digest=sha256:cc5a8019f499135772779f86779d7fa2a28d373236941186eb6989883e08647d

Observation eabd1c0e-4c1b-4689-a89f-5ed82c8c11d7 · inbound

AblateCell: A Reproduce-then-Ablate Agent for Virtual Cell Repositories cites this paper.

AblateCell: A Reproduce-then-Ablate Agent for Virtual Cell Repositories ResearchCodeBench: Benchmarking LLMs on Implementing Novel Machine Learning Research Code

Reference 42

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T13:01:04.244734Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-10T02:30:39.346257Z digest=sha256:27415e0d80c736ca4a62957bb1ff6f15c9f9d43441aea350de3b405fb33d1a55

Observation eddb464f-bb32-48c8-91c2-7cbca15f67af · inbound

Is Agentic AI Ready for Real-World Hardware Engineering? A Deep Dive with Phoenix-bench cites this paper.

Is Agentic AI Ready for Real-World Hardware Engineering? A Deep Dive with Phoenix-bench ResearchCodeBench: Benchmarking LLMs on Implementing Novel Machine Learning Research Code

Reference 33

Resolution
metadata mismatch
arxiv_id, observed 2026-05-19T17:52:43.035333Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-19T17:49:01.198956Z digest=sha256:437e393e0a5d2a4081b90690f8876aea48978889daf3eb19c13b10e0f634f088

Observation f78992f4-4ce3-462a-b6a9-68a132b93bb6 · inbound

MLReplicate: Benchmarking Autonomous Research Systems for Machine Learning Reproducibility cites this paper.

MLReplicate: Benchmarking Autonomous Research Systems for Machine Learning Reproducibility ResearchCodeBench: Benchmarking LLMs on Implementing Novel Machine Learning Research Code

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-20T20:03:44.142766Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-20T19:59:40.519962Z digest=sha256:89508ce27b811c19e9141b2b6eb89684e8a69a87d9b652ee7e02d128c2226387

Observation 78d3d367-84ca-450b-bab8-24e16520e1a2 · inbound

ASPI: Seeking Ambiguity Clarification Amplifies Prompt Injection Vulnerability in LLM Agents cites this paper.

ASPI: Seeking Ambiguity Clarification Amplifies Prompt Injection Vulnerability in LLM Agents ResearchCodeBench: Benchmarking LLMs on Implementing Novel Machine Learning Research Code

Reference 73

Resolution
metadata mismatch
arxiv_id, observed 2026-05-19T23:57:53.363396Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-19T23:54:15.987953Z digest=sha256:78312c063ecbe29492563140b41391603cbac52de1cf1b773bc9bbf7ad76b64c

Observation 87403219-8c32-45e1-bfd6-113e90b5ff56 · inbound

AI for Auto-Research: Roadmap & User Guide cites this paper.

AI for Auto-Research: Roadmap & User Guide ResearchCodeBench: Benchmarking LLMs on Implementing Novel Machine Learning Research Code

Reference 71

Resolution
verified exact
arxiv_id, observed 2026-05-20T10:33:12.684486Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-20T10:30:50.256635Z digest=sha256:51d973088827c267aae016d6a28c27021d4e1361804c8486902fac3843b44348

Observation dfaac76a-3c6a-4e6f-a532-dff57ebd65bf · inbound

AI for Auto-Research: Roadmap & User Guide cites this paper.

AI for Auto-Research: Roadmap & User Guide ResearchCodeBench: Benchmarking LLMs on Implementing Novel Machine Learning Research Code

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-02T13:43:35.697198Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:43:35.697198Z digest=sha256:96a0d53de2f864bbd90235b930bd4fd99d851c53b68b09b5d939fe6401e98e8b

Observation 943dad94-76be-4e79-8e7e-ac03b93f6b2a · inbound

From paper to benchmark: agentic, framework-based reproduction of under-specified methods in machine health intelligence cites this paper.

From paper to benchmark: agentic, framework-based reproduction of under-specified methods in machine health intelligence ResearchCodeBench: Benchmarking LLMs on Implementing Novel Machine Learning Research Code

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-06-29T12:23:24.672894Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-29T12:13:58.111299Z digest=sha256:ff4121ed0f81ba0936ae688fade1833205abe820f3f3811568727c950f0d6c3e

Observation 15cc6769-5a69-4933-b731-70b6c788d4fb · inbound

Can LLMs Build a MaxSAT Solver from Papers? The CoreForge Experience cites this paper.

Can LLMs Build a MaxSAT Solver from Papers? The CoreForge Experience ResearchCodeBench: Benchmarking LLMs on Implementing Novel Machine Learning Research Code

Reference 2009

Resolution
unresolved
no resolver link, observed 2026-08-02T01:00:22.372434Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:00:22.372434Z digest=sha256:2a8e8f7a76f40036817022eecb1292c4c0d84074e96e6cd1e5127d3c318d233b

Observation 88708047-8690-4df7-a06e-1ca8438e43fc · inbound

From Execution to Capability: Scientific Experience Consolidation via Procedural Knowledge Synthesis cites this paper.

From Execution to Capability: Scientific Experience Consolidation via Procedural Knowledge Synthesis ResearchCodeBench: Benchmarking LLMs on Implementing Novel Machine Learning Research Code

Reference 6

Resolution
unresolved
no resolver link, observed 2026-07-31T14:26:52.414276Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T14:26:52.414276Z digest=sha256:706cbedd3f5c2498f11c5905b6292c6e7db2259315ad8a0244833064ad583939

Observation 63a653a7-9bd5-43ab-b6a6-5c508d9d876e · inbound

Revibing Code from Papers: Reimplementing HCI Artifacts cites this paper.

Revibing Code from Papers: Reimplementing HCI Artifacts ResearchCodeBench: Benchmarking LLMs on Implementing Novel Machine Learning Research Code

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-15T15:24:46.242048Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:24:46.242048Z digest=sha256:0f8e6236c5c92a624932b79ddc4b26debe92a901d2135c433eed88855b0581f6