Pith. sign in

Paper Citation Record · LEDGER

RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts

As of 18 August 2026, this Paper Citation Record lists 64 of 64 outbound references and 64 inbound Pith citation observations for arXiv:2411.15114.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.15114 v2

Coverage vector

measured 64 of 64 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T14:33:13.231482Z

measured 128 of 128 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 64 of 64 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T12:40:14.612983Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

64 of 64 outbound references displayed

  • verified exact1
  • verified fuzzy17
  • unresolved44
  • parse uncertain0
  • malformed identifier2
  • metadata mismatch0

External citation measurements

2
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 93e50550-a810-40fc-a4a7-8283917db63c · outbound

This paper cites Evaluating Large Language Models Trained on Code.

RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts Evaluating Large Language Models Trained on Code

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T14:33:13.069528Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:33:13.069528Z digest=sha256:b937ae987d237c13d46a9fba119845e82d6f2c964e06dec5f28a381f38c84c73

Observation 9214f6e2-9035-4c84-9c81-1fd6d4366a98 · outbound

This paper cites Competition-level code generation with AlphaCode.

RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts Competition-level code generation with AlphaCode

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-12T14:33:13.073102Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:33:13.073102Z digest=sha256:8ecbf3d8515d25121a6c8642fbc19d52b05e2bc107a9b2f9ae099129fcf2bb3e

Observation d68dc253-f549-4593-aebf-c77561f0f0ee · outbound

This paper cites Textbooks Are All You Need.

RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts Textbooks Are All You Need

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-12T14:33:13.076189Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:33:13.076189Z digest=sha256:49a4c55c8860c1ff4301a301a51b8c89b6b22be103b91c987f90a3373a2b9a3d

Observation bea0344b-a812-44b7-aef2-106eb00a4911 · outbound

This paper cites The Llama 3 Herd of Models.

RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts The Llama 3 Herd of Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T14:33:13.078973Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:33:13.078973Z digest=sha256:1162a5ad9d29fa5eeb9110ea128438ebfaac15bae098f585f82c6d8bb3188e5a

Observation 419e9304-59a9-477e-88c0-8b9dfed02aa1 · outbound

This paper cites an unresolved cited work.

RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-12T14:33:13.703896Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T14:33:13.081910Z digest=sha256:c3fb7a05498be43d8a2f08f7493bcc80be232e68141ed788b1676a6e1ea5d6c4

Observation 2f4af9d2-ff40-4e0e-9bec-fe388e460293 · outbound

This paper cites OpenHands: An Open Platform for AI Software Developers as Generalist Agents.

RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts OpenHands: An Open Platform for AI Software Developers as Generalist Agents

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-12T14:33:13.084827Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:33:13.084827Z digest=sha256:6b2887b127145c07c5d5fff950b4408949de3e3a6d6ca6975477d7639e3d9add

Observation fb9d432a-3125-41da-b3f2-89fed418c256 · outbound

This paper cites Interviewing AI researchers on automation of AI r&d (2024).

RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts Interviewing AI researchers on automation of AI r&d (2024)

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:33:13.697544Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T14:33:13.088517Z digest=sha256:5f5487d81e7f62cb480a264b8be0fbf0d588660e0fd11221d4ed8ddc4c0d4825

Observation d94f04f1-6f4d-4ba2-b655-c0a96dd5ba44 · outbound

This paper cites an unresolved cited work.

RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts Unresolved cited work

Reference 8

Resolution
verified exact
doi, observed 2026-08-12T14:33:13.252997Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T14:33:13.090814Z digest=sha256:6fbf8b5f845a827c6c3b2d27f08940fb71370f5988cce65c9af094d9f3bedb49

Observation 13b94930-6795-4441-aff6-0483157c500f · outbound

This paper cites Explosive growth from AI automation: A review of the arguments.

RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts Explosive growth from AI automation: A review of the arguments

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-12T14:33:13.093306Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:33:13.093306Z digest=sha256:a20be91f3569cc625323ea8c31f4a12c8a1921f8a4d5344ffa1ba281f62ebf5b

Observation 9108dd92-eba2-470c-b99b-80ec96d37abd · outbound

This paper cites OpenAI preparedness framework (beta).

RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts OpenAI preparedness framework (beta)

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:33:13.691722Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T14:33:13.095724Z digest=sha256:0deeb94af80e45bc432a5ae6b53011a8502918fa5c367e5676f51bbd1a37ebfa

Observation daad44ba-e91c-40f9-8909-3730c0508d4a · outbound

This paper cites Frontier safety framework, version 1.0.

RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts Frontier safety framework, version 1.0

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:33:13.685477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T14:33:13.097831Z digest=sha256:04d8939aa633043a1f32cf9ab9303fb03ca75b231f0a7d02d875780969473bd9

Observation caa42555-43ef-46bc-9dda-6754402a07df · outbound

This paper cites Anthropic responsible scaling policy.

RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts Anthropic responsible scaling policy

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:33:13.678875Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T14:33:13.099932Z digest=sha256:67c967eb6aead05305cc01ca26c0a8a7d1e99e38d8e3f158f846080a6b621574

Observation cfac877e-9688-40de-b039-a5b053e7a108 · outbound

This paper cites Recital 110 of the eu artificial intelligence act (2024).

RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts Recital 110 of the eu artificial intelligence act (2024)

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:33:13.671828Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T14:33:13.101999Z digest=sha256:20a215ce487896dc7116feb57ac78dd1c61d1b73e2d3ae33534afe18f4a39a10

Observation 8825082f-4866-415c-8950-0e5c94ab0ab6 · outbound

This paper cites an unresolved cited work.

RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts Unresolved cited work

Reference 14

Resolution
unresolved
raw_fallback, observed 2026-08-12T14:33:13.663824Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T14:33:13.103890Z digest=sha256:0577f3d2f7dc95d52b449ce228f54b0ce4faaaed47373522fcb066a73d6207a9

Observation a2741764-7fae-47ce-9f96-def74a3d4ac0 · outbound

This paper cites The bletchley declaration on AI safety (2023).

RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts The bletchley declaration on AI safety (2023)

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:33:13.656928Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T14:33:13.105870Z digest=sha256:799ba1a5324aa58bc7c5346f2566b19c825bc8aadb6de0e184d8cb009d931033

Observation 2aebe51c-8d48-42b4-a261-88fb47af02e5 · outbound

This paper cites Frontier AI safety commitments, AI seoul summit 2024.

RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts Frontier AI safety commitments, AI seoul summit 2024

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:33:13.650122Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T14:33:13.107747Z digest=sha256:d2fc73eef88149abe3d90106f6f7cbf329d3f4a796249df122629d0e0a71d86a

Observation 2d114eb8-0e65-4d24-81f6-aee212b8eaf4 · outbound

This paper cites Building an early warning system for llm-aided biological threat creation (2024).

RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts Building an early warning system for llm-aided biological threat creation (2024)

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:33:13.643125Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T14:33:13.109734Z digest=sha256:c3e37003fa7eae960177541e819f50f53a4682271b51485ad9c92b41222eb6d5

Observation bfcb303c-7e77-41d0-9a92-19acb60f329a · outbound

This paper cites CYBERSECEVAL 3: Advancing the Evaluation of Cybersecurity Risks and Capabilities in Large Language Models.

RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts CYBERSECEVAL 3: Advancing the Evaluation of Cybersecurity Risks and Capabilities in Large Language Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-12T14:33:13.112305Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:33:13.112305Z digest=sha256:648a131016857a93121543ac33a8a14a98108db5403b83192d9df5b06178696c

Observation 21553f6f-d87d-4220-8f52-d8c10669d594 · outbound

This paper cites What a compute-centric framework says about takeoff speeds (2023).

RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts What a compute-centric framework says about takeoff speeds (2023)

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:33:13.636145Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T14:33:13.115443Z digest=sha256:0cfb7d8acf514c7a18c0c28168fffb390c8fa552603c012066597ca018096b37

Observation 049cc47b-1736-4ae4-ba5f-dd9eddbf9ed7 · outbound

This paper cites Sabotage Evaluations for Frontier Models.

RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts Sabotage Evaluations for Frontier Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-12T14:33:13.117927Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:33:13.117927Z digest=sha256:9dee798cb0d07e12894bc1e148937cffdc84de796ed837e97d7a389d92627790

Observation 43984ef0-bde5-411f-8d9a-bb6bf903d172 · outbound

This paper cites an unresolved cited work.

RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts Unresolved cited work

Reference 21

Resolution
unresolved
raw_fallback, observed 2026-08-12T14:33:13.628799Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T14:33:13.120597Z digest=sha256:675465dacf74f18881343f1d343bacc3afca47ce8fe74c93558fe236c0338a06

Observation b52f8e02-5694-430a-881c-ebe71f379032 · outbound

This paper cites an unresolved cited work.

RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts Unresolved cited work

Reference 22

Resolution
unresolved
raw_fallback, observed 2026-08-12T14:33:13.622601Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T14:33:13.122748Z digest=sha256:bceb9538d3e63ef0bf6a1ad6934f2a13c740b0eddac883e7b48b61e8c36eabb7

Observation dda5e2a4-9525-41fa-bd72-aba38ab4de70 · outbound

This paper cites Raising the bar on swe-bench verified with claude 3.5 sonnet (2024).

RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts Raising the bar on swe-bench verified with claude 3.5 sonnet (2024)

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:33:13.616633Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T14:33:13.125629Z digest=sha256:55d8d7dec50566bb6eef6b060872d426db24430f00ddab3a6f4c33c08f95884c

Observation 055c5b1d-80d5-4728-9bcd-93629c551826 · outbound

This paper cites Artificial Intelligence Index Report 2024.

RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts Artificial Intelligence Index Report 2024

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-12T14:33:13.127921Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:33:13.127921Z digest=sha256:f642e84c71d6c41f4ac5d60d0c77a299af747bee8ca5e8e084df54299021001e

Observation 26c207d8-7c32-4898-8e65-e5ff9de0db1e · outbound

This paper cites MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering.

RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-12T14:33:13.130371Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:33:13.130371Z digest=sha256:d727999ecd42d66991258fe7fa5a2400d23963bbf1b864d0ad0053cea4c44018

Observation 357ddd75-67dc-4f11-8700-e9dae2fa006f · outbound

This paper cites SciCode: A Research Coding Benchmark Curated by Scientists.

RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts SciCode: A Research Coding Benchmark Curated by Scientists

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-12T14:33:13.132899Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:33:13.132899Z digest=sha256:44755f21a91d6432b25a8c108609c1c7ff8c97e151cd22325564c5c32de39735

Observation 1f305af8-669c-4c5f-bd9c-842163806679 · outbound

This paper cites GPQA: A Graduate-Level Google-Proof Q&A Benchmark.

RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts GPQA: A Graduate-Level Google-Proof Q&A Benchmark

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-12T14:33:13.135584Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:33:13.135584Z digest=sha256:983a7cd43bce67cfb09ac457cf9a1aa1827dce50473aafd6594f38918347a2d9

Observation 1d91d83a-f5ba-4e82-9d6a-ddcc1cad57c0 · outbound

This paper cites GAIA: a benchmark for General AI Assistants.

RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts GAIA: a benchmark for General AI Assistants

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-12T14:33:13.138150Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:33:13.138150Z digest=sha256:fa0feceaf6c11e1c76b9ee2bbbc983316c5c04dc729fd32095e76f4ac30a6994

Observation 5db7ccbf-5e43-407d-99a4-bbdb04f4c4d7 · outbound

This paper cites an unresolved cited work.

RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts Unresolved cited work

Reference 29

Resolution
unresolved
raw_fallback, observed 2026-08-12T14:33:13.610070Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T14:33:13.140723Z digest=sha256:9270e6d9ef652e0f3e6e57475d4aad865a68280521b73dde1139a2584489e2ef

Observation 33f1df9b-aaf8-44c7-b2d6-80125436611e · outbound

This paper cites WebArena: A Realistic Web Environment for Building Autonomous Agents.

RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts WebArena: A Realistic Web Environment for Building Autonomous Agents

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-12T14:33:13.143333Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:33:13.143333Z digest=sha256:66604136a37280a46aabef0b2705bbf2bbe62b38701f21352d466295054c2071

Observation 482e5e49-c928-48db-80dd-957436abcce4 · outbound

This paper cites MLAgentBench: Evaluating Language Agents on Machine Learning Experimentation.

RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts MLAgentBench: Evaluating Language Agents on Machine Learning Experimentation

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-12T14:33:13.146051Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:33:13.146051Z digest=sha256:52b67f054e12aeaf4811aa960da3ece681b90c1446f5d827f605ebc4843a5c02

Observation 70053717-4393-4d76-a4f4-d829c6c4311b · outbound

This paper cites DISCOVERYWORLD: A Virtual Environment for Developing and Evaluating Automated Scientific Discovery Agents.

RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts DISCOVERYWORLD: A Virtual Environment for Developing and Evaluating Automated Scientific Discovery Agents

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-12T14:33:13.148692Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:33:13.148692Z digest=sha256:3b002b7d71ff11ec7d8dbf5bc606d26f3c7edbb38e8c5b908fc67fa970f26a74

Observation ab39c1ef-c811-4eb3-9f1c-5e419341688d · outbound

This paper cites H-ARC: A Robust Estimate of Human Performance on the Abstraction and Reasoning Corpus Benchmark.

RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts H-ARC: A Robust Estimate of Human Performance on the Abstraction and Reasoning Corpus Benchmark

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-12T14:33:13.151881Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:33:13.151881Z digest=sha256:234eb10498e8481a493a1ca3825204e05b8427d6689412ec04302b15a538d341

Observation 7add920a-c7e6-4f45-bd36-39f67fc91252 · outbound

This paper cites Training language models to follow instructions with human feedback.

RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts Training language models to follow instructions with human feedback

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-12T14:33:13.154509Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:33:13.154509Z digest=sha256:652905cfc9fd7bb3c20328b576a5dd909f6a6b59abd31c43c0f03325e74efed1

Observation 9a730257-0cea-427b-94c2-65df035f2290 · outbound

This paper cites Constitutional AI: Harmlessness from AI Feedback.

RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts Constitutional AI: Harmlessness from AI Feedback

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-12T14:33:13.157074Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:33:13.157074Z digest=sha256:148d7a26a8a4aa61cadf81d8a9c74db46543b4ea57af1fb2c9b4a3728da5bdbd

Observation 36689a9d-015e-49bd-955b-1b15cf741994 · outbound

This paper cites Nemotron-4 340B Technical Report.

RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts Nemotron-4 340B Technical Report

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-12T14:33:13.159698Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:33:13.159698Z digest=sha256:e1ce33b73786477e1f83b2023885f9f36b34a86bbe2c40fb0ff6b895db8a203f

Observation be93dd10-ed34-4eaf-a0ed-340d20b151c3 · outbound

This paper cites an unresolved cited work.

RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts Unresolved cited work

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-12T14:33:13.162197Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:33:13.162197Z digest=sha256:381fe57b4c97f5582340eebba5c0f321859c491c20cd5db909bbe08c2666d1a4

Observation 9306505f-d8a0-4dff-aa71-b43aee00c239 · outbound

This paper cites EvoPrompting: Language Models for Code-Level Neural Architecture Search.

RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts EvoPrompting: Language Models for Code-Level Neural Architecture Search

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-12T14:33:13.165447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:33:13.165447Z digest=sha256:de8ab8da2ffa9903d3612c5ca1fb55612935842fc29699067b0ad90042576c90

Observation 4cf5c454-0308-43ae-8b62-b4ad963cb9ed · outbound

This paper cites DS-Agent: Automated Data Science by Empowering Large Language Models with Case-Based Reasoning.

RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts DS-Agent: Automated Data Science by Empowering Large Language Models with Case-Based Reasoning

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-12T14:33:13.167944Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:33:13.167944Z digest=sha256:4eef1c796b3615454639242d8014c9e7d8ca1930b0df56c0b92eb411c5218fe4

Observation 0306d3bd-d6ea-4cfb-9ff7-ef2b4388d32f · outbound

This paper cites SUPER: Evaluating Agents on Setting Up and Executing Tasks from Research Repositories.

RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts SUPER: Evaluating Agents on Setting Up and Executing Tasks from Research Repositories

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-12T14:33:13.170581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:33:13.170581Z digest=sha256:ffa5174fdf78ea2e6ba474b4fa9dafcbe82f006d92ef2ef0e41e878993e2ee3f

Observation 97edb651-32a6-4773-b83e-73d493d1b791 · outbound

This paper cites The AI Scientist: Towards Fully Automated Open-Ended Scientific Discovery.

RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts The AI Scientist: Towards Fully Automated Open-Ended Scientific Discovery

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-12T14:33:13.173140Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:33:13.173140Z digest=sha256:5b544c464064e44b69d398700189e8a402e00eae395c4a9a645a19c8617ce47a

Observation d4886c2d-9919-4e81-a473-72b7b55735e7 · outbound

This paper cites Eureka: Human-Level Reward Design via Coding Large Language Models.

RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts Eureka: Human-Level Reward Design via Coding Large Language Models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-12T14:33:13.176588Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:33:13.176588Z digest=sha256:b0805398acfab605a4c34f5669a823acb24d87ad5149f8147b7d480eecfc8131

Observation 619dfc4e-82ba-4856-ba1e-e93636af51f6 · outbound

This paper cites OMNI-EPIC: Open-endedness via Models of human Notions of Interestingness with Environments Programmed in Code.

RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts OMNI-EPIC: Open-endedness via Models of human Notions of Interestingness with Environments Programmed in Code

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-12T14:33:13.179246Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:33:13.179246Z digest=sha256:afcdd13d118cdb4359c732c840a7a3a2c34616667d72d30b1e249fc720fde5cf

Observation 943925d1-4f81-4b0c-8197-e482e3e8d8bd · outbound

This paper cites Discovering Preference Optimization Algorithms with and for Large Language Models.

RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts Discovering Preference Optimization Algorithms with and for Large Language Models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-12T14:33:13.181972Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:33:13.181972Z digest=sha256:6dd177820ab6a9c9afae65c917321fb98dcd44b9ef9071b13cbc7d7410c602ca

Observation 374b2181-0d01-4879-a464-a963c73c8b91 · outbound

This paper cites SciAgent: Tool-augmented Language Models for Scientific Reasoning.

RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts SciAgent: Tool-augmented Language Models for Scientific Reasoning

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-12T14:33:13.184324Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:33:13.184324Z digest=sha256:8cdd98d24c46a99535167a091b6aa396cac96c23401e12d3d85242de3a736cdb

Observation ce8cc2a6-f7fa-485e-af96-2020ce6cf0ce · outbound

This paper cites ChemCrow: Augmenting large-language models with chemistry tools.

RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts ChemCrow: Augmenting large-language models with chemistry tools

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-12T14:33:13.186866Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:33:13.186866Z digest=sha256:0246904847c16df76e5d094e8fac603217750a9933ba9df86377bbc0f2fe84dd

Observation b46ed719-5073-413b-b47a-0747638c7318 · outbound

This paper cites Scientific Large Language Models: A Survey on Biological & Chemical Domains.

RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts Scientific Large Language Models: A Survey on Biological & Chemical Domains

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-12T14:33:13.189430Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:33:13.189430Z digest=sha256:0272372239a6385ab3ba3092d23e10f2a6f12d8656216de71cdc7eb924628f04

Observation 9a49e1d6-cb3f-4d06-a864-f12545dc0181 · outbound

This paper cites AutoML in the Age of Large Language Models: Current Challenges, Future Opportunities and Risks.

RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts AutoML in the Age of Large Language Models: Current Challenges, Future Opportunities and Risks

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-12T14:33:13.191545Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:33:13.191545Z digest=sha256:6c404fc2d1ca5a84d9732334ad1b6993aa1923a8cf68478b5a38ea6ee75fd6b6

Observation d7a33066-7114-4595-95a8-d5e6ea53ad7a · outbound

This paper cites Chip Placement with Deep Reinforcement Learning.

RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts Chip Placement with Deep Reinforcement Learning

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-12T14:33:13.194082Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:33:13.194082Z digest=sha256:3fc6e40eaff2db93ab99b0730843644866233160897e67816622403c518e80e0

Observation bd8aacd5-e317-4e49-aa2b-d1e8e4d17979 · outbound

This paper cites & Aydos, G.

RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts & Aydos, G

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:33:13.603105Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T14:33:13.196795Z digest=sha256:452558e456c6f291d0792f50966108f40ca52e7de73cca9511780eefa0322726

Observation 27a19e35-bd54-415a-8ae2-b056bd5b6fbc · outbound

This paper cites Vivaria: Open-source platform for agent evaluations (2024).

RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts Vivaria: Open-source platform for agent evaluations (2024)

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:33:13.596263Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T14:33:13.199205Z digest=sha256:f1e00d98d0caed38a47b68ba3fd089b686138241fd6ecca508e73fe180c241ff

Observation 2a51f86b-eb88-4539-b1a0-ff870c907679 · outbound

This paper cites Claude 3.5 Sonnet model card addendum (2024).

RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts Claude 3.5 Sonnet model card addendum (2024)

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:33:13.589134Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T14:33:13.201630Z digest=sha256:40748083e1fb0a69512e29a2282fd02c3beebd1622aeadc4a3946b8b4ea00ae1

Observation b537316d-67d4-46f4-a304-82e7eff5addb · outbound

This paper cites OpenAI o1 system card (2024).

RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts OpenAI o1 system card (2024)

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:33:13.580503Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T14:33:13.204017Z digest=sha256:82210afb639fc1c09dedcfdcf8639b0bffb8a6fe5f07bfcbf11794fecd0a98ad

Observation b7390dc8-00a2-4757-98b3-64235c222ea9 · outbound

This paper cites AIDE: Data science automation technical report (2024).

RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts AIDE: Data science automation technical report (2024)

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:33:13.573410Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T14:33:13.206318Z digest=sha256:97c5d5420d6ccd306a8449c4aa8560e2319009b492e671bcc4e7a110a8a03d06

Observation 7693bafe-b60a-40ed-85b6-4579cff93977 · outbound

This paper cites Details about METR’s preliminary evaluation of OpenAI o1-preview (2024).

RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts Details about METR’s preliminary evaluation of OpenAI o1-preview (2024)

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:33:13.565835Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T14:33:13.208642Z digest=sha256:2b5e847f5ec2d2baecfa415e20dab399d66369dce1171aa1ed5f878d210ba346

Observation a4dbaa28-eb60-414f-96d8-590c9bce99bc · outbound

This paper cites Claude 3.5 Sonnet: Quality, performance & price analysis (2024).

RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts Claude 3.5 Sonnet: Quality, performance & price analysis (2024)

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T14:33:13.558707Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T14:33:13.210938Z digest=sha256:258966a6a091e16d07d23d4a2aa28aefdfe77e1431d022a83ece476e56068ec9

Observation 4008e77c-c644-4c3c-a78f-76a5be9c7e79 · outbound

This paper cites an unresolved cited work.

RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts Unresolved cited work

Reference 57

Resolution
unresolved
raw_fallback, observed 2026-08-12T14:33:13.552270Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T14:33:13.213202Z digest=sha256:ff06ce8cb4e6624ab5c4a76fd3beeed48416fce64fc6b0c2997e272a32079995

Observation 02f7ff6e-504e-4a93-bc58-097c5a6ada3c · outbound

This paper cites Direct Preference Optimization: Your Language Model is Secretly a Reward Model.

RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts Direct Preference Optimization: Your Language Model is Secretly a Reward Model

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-12T14:33:13.215554Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:33:13.215554Z digest=sha256:0e15506b858a2792f0b65ec75222c00077f049bdc41fa1f2835095a3dfa60383

Observation b023ccb1-9588-4ffa-bd1e-97c18c884f76 · outbound

This paper cites Is this score similar to what you predicted or measured yourself, or does it come as a surprise?.

RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts Is this score similar to what you predicted or measured yourself, or does it come as a surprise?

Reference 59

Resolution
malformed identifier
raw_fallback, observed 2026-08-12T14:33:13.545394Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T14:33:13.218367Z digest=sha256:cb409e3625d78f1d79daeaf9403ca0e41f8cef29c414b2f577c4ed8ac71b592a

Observation 46c3d2b7-a08a-4178-868e-fe1ab97cb6d9 · outbound

This paper cites an unresolved cited work.

RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts Unresolved cited work

Reference 60

Resolution
unresolved
raw_fallback, observed 2026-08-12T14:33:13.536710Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T14:33:13.222060Z digest=sha256:494d5716ffe6acf51bdf0162c928f2a0c680bb1521d3e6abb303534ab1ed1149

Observation 0f6e2c1e-2e63-4d88-9548-82a0c955aa1f · outbound

This paper cites an unresolved cited work.

RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts Unresolved cited work

Reference 61

Resolution
unresolved
raw_fallback, observed 2026-08-12T14:33:13.529316Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T14:33:13.224373Z digest=sha256:b08f07584cb42ed4f466fedd56eb6ede182320a945303754b4e80abf86aa931b

Observation 6231b3b9-ee6a-4a25-baf1-34cc040b76dd · outbound

This paper cites an unresolved cited work.

RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts Unresolved cited work

Reference 62

Resolution
unresolved
raw_fallback, observed 2026-08-12T14:33:13.521984Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T14:33:13.226781Z digest=sha256:11479516aae3cfbfee727d8ddeb3669a603fa12ad45463828d32a2a05c5647d0

Observation 9a6f4ce8-f8af-4924-a545-1e699ff2598a · outbound

This paper cites an unresolved cited work.

RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts Unresolved cited work

Reference 63

Resolution
unresolved
raw_fallback, observed 2026-08-12T14:33:13.514722Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T14:33:13.229141Z digest=sha256:ab8c405887b88dd1bddc87d471f5ca335b22922677ee7ad4c96cd888e5b9014c

Observation f6629272-2716-425d-b21e-14db69ae21ed · outbound

This paper cites cuda" "cuda.

RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts cuda" "cuda

Reference 64

Resolution
malformed identifier
raw_fallback, observed 2026-08-12T14:33:13.507078Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T14:33:13.231482Z digest=sha256:950b079047ef7c379487a4522b46607202cee002424392dadd135e6a7c5e9a50

Pith citing papers

Observation 4b89afad-6c15-4cd6-a734-fdf2480c86bc · inbound

Frontier Models are Capable of In-context Scheming cites this paper.

Frontier Models are Capable of In-context Scheming RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-16T14:22:01.681340Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-16T14:22:01.616448Z digest=sha256:f79245e2e6ec818c78490a53770561ca8166cbd4d77080e5350b7570c7ab9baa

Observation d9fa4ac5-63e7-45de-986b-c9ba862334fb · inbound

Humanity's Last Exam cites this paper.

Humanity's Last Exam RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts

Reference 60

Resolution
verified exact
arxiv_id, observed 2026-05-10T18:40:50.326504Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-10T18:40:50.139345Z digest=sha256:ee353b1a8fd70783269df70d680220774aeaf238eaf9bd25a873e4edd48b3d39

Observation 7dddfdc1-9de0-4ea1-b0b6-ecb02f46e6d0 · inbound

The AI Agent Index cites this paper.

The AI Agent Index RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-09T14:48:35.872296Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:48:35.872296Z digest=sha256:3ab4f4eae40887797732f6625de55ac6948a649cb8b593353cb124b4a2f5ae0a

Observation 0db8ffb1-e533-489d-b388-a0a9ea0c6fec · inbound

KernelBench: Can LLMs Write Efficient GPU Kernels? cites this paper.

KernelBench: Can LLMs Write Efficient GPU Kernels? RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-15T16:55:02.090533Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-15T16:55:01.976356Z digest=sha256:b251bedab91cdaf2cbb66973d665490836d81ea607257010fb9d00626224912a

Observation 915b7f2d-f33e-4ae4-9db6-045f0a38f059 · inbound

AI Behind Closed Doors: a Primer on The Governance of Internal Deployment cites this paper.

AI Behind Closed Doors: a Primer on The Governance of Internal Deployment RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-16T12:40:14.612983Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:40:14.612983Z digest=sha256:518e6d32bcc38f8a3a7ea7380649087a6facdd5514807bf6fbc6d7982da87b81

Observation 18990c90-06f6-4e82-967c-1f3b67587aff · inbound

Bare Minimum Mitigations for Autonomous AI Development cites this paper.

Bare Minimum Mitigations for Autonomous AI Development RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-16T11:30:47.179402Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T11:30:47.179402Z digest=sha256:e406b201c306bf5e86c1576affe7f61a6a06d40e20adb55346b8985093e9c258

Observation 7f9a1d21-dc25-471a-8e49-39fd2441209b · inbound

Can AI Agents Design and Implement Drug Discovery Pipelines? cites this paper.

Can AI Agents Design and Implement Drug Discovery Pipelines? RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-16T05:44:03.801916Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:44:03.801916Z digest=sha256:7d979aea436126200b4fb82f14fbb457056873a61a40ac4fd883a838ff3b09f0

Observation 19dcc8fb-763b-47e7-8ec9-771e5de241fe · inbound

AI Governance to Avoid Extinction: The Strategic Landscape and Actionable Research Questions cites this paper.

AI Governance to Avoid Extinction: The Strategic Landscape and Actionable Research Questions RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts

Reference 195

Resolution
unresolved
no resolver link, observed 2026-08-15T23:27:30.618243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:27:30.618243Z digest=sha256:1b0158100a2ddd2703bb1b7f230a6f8eb46d42ad127891175b3eba857fe97484

Observation 35d19ee8-5d53-4ede-b2c2-bd9c4fc69316 · inbound

LLMs Outperform Experts on Challenging Biology Benchmarks cites this paper.

LLMs Outperform Experts on Challenging Biology Benchmarks RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T22:50:49.704789Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:50:49.704789Z digest=sha256:b3ff411e2d25f6ead479b3f7113bd604e090d159247f789a8c56808c155f001d

Observation 79ec15b4-8781-468b-8d3e-beeea998a9e1 · inbound

TimeSeriesGym: A Scalable Benchmark for (Time Series) Machine Learning Engineering Agents cites this paper.

TimeSeriesGym: A Scalable Benchmark for (Time Series) Machine Learning Engineering Agents RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-15T20:18:59.964481Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:18:59.964481Z digest=sha256:f52a3e1164ba3651f71c2f74d859db4f98e09ca76712cdba9215c01746b51bf1

Observation e5b67da1-0357-4675-8eb8-5547e5883c99 · inbound

Agent-X: Evaluating Deep Multimodal Reasoning in Vision-Centric Agentic Tasks cites this paper.

Agent-X: Evaluating Deep Multimodal Reasoning in Vision-Centric Agentic Tasks RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-07T12:18:55.881071Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:18:55.881071Z digest=sha256:998a0b38fe50ff167dbd0d706e2eb4e04b7cc153a0e87ed93e13a27aeba11986

Observation 603c9afb-9026-40fe-9f5a-dc6febb693fb · inbound

TextAtari: 100K Frames Game Playing with Language Agents cites this paper.

TextAtari: 100K Frames Game Playing with Language Agents RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-07T10:51:58.052862Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:51:58.052862Z digest=sha256:ea0f53e8097aff79c738d2abf5466eaf253cb4a5de07bda3d7cb9acb7385a98c

Observation f650702d-9842-4d40-99ca-3e58f8cf9c10 · inbound

Deep Research Agents: A Systematic Examination And Roadmap cites this paper.

Deep Research Agents: A Systematic Examination And Roadmap RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts

Reference 118

Resolution
unresolved
no resolver link, observed 2026-08-06T23:26:56.952422Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:26:56.952422Z digest=sha256:f9dd06ecc215d37107a12abb1caec4b9da552f217b8e038cfef346afc44f71a7

Observation 36100b82-0a6a-453d-8fd8-7252d4bb2915 · inbound

Establishing Best Practices for Building Rigorous Agentic Benchmarks cites this paper.

Establishing Best Practices for Building Rigorous Agentic Benchmarks RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-06T20:24:27.421486Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:24:27.421486Z digest=sha256:0836b2167a45975227309c928cf5461eaa545099c45ae5edf717c3afb87e56bc

Observation b3799e82-49f1-4e93-b917-0b42741574ce · inbound

Evaluating the Critical Risks of Amazon's Nova Premier under the Frontier Model Safety Framework cites this paper.

Evaluating the Critical Risks of Amazon's Nova Premier under the Frontier Model Safety Framework RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T19:39:59.279843Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:39:59.279843Z digest=sha256:be7d02542125b24cba13ae833c8178249737998d7875623d1c8d38a19047086e

Observation a58694cb-4e26-4f00-876c-ee10cd3eb27b · inbound

Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities cites this paper.

Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts

Reference 89

Resolution
verified exact
arxiv_id, observed 2026-05-19T05:52:07.806483Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-19T05:48:02.828938Z digest=sha256:e0a0d4c8fd62a288890b991ff922c78338c2a0ac73304812db042c94afefd18b

Observation 9e58febf-03fe-4d5f-8648-635514eec14a · inbound

AblationBench: Evaluating Automated Planning of Ablations in Empirical AI Research cites this paper.

AblationBench: Evaluating Automated Planning of Ablations in Empirical AI Research RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-06T19:00:08.585286Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:00:08.585286Z digest=sha256:a01f14d84bc565edbfbc6fe054b7c5e2cf7be00afee00f657f9e1c5b3924b546

Observation 09cd1878-e82b-4063-bde5-4ff58152aeb2 · inbound

Exploring Design of Multi-Agent LLM Dialogues for Research Ideation cites this paper.

Exploring Design of Multi-Agent LLM Dialogues for Research Ideation RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-06T18:24:26.153839Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:24:26.153839Z digest=sha256:bfd566b9a535e6384bf8386d8e2714efee214d314a084eaba83ebf477ff114e1

Observation 203e8758-23fb-496e-bb2c-9753f4b9bdb4 · inbound

Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity cites this paper.

Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T18:12:29.939412Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:12:29.939412Z digest=sha256:1204087be24a07a17285fe13174a21a21419c9548c76adc9935a5d93beaebeed

Observation 7eca4a44-634e-4288-828e-ad2fcd68d077 · inbound

Chain of Thought Monitorability: A New and Fragile Opportunity for AI Safety cites this paper.

Chain of Thought Monitorability: A New and Fragile Opportunity for AI Safety RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts

Reference 29

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T14:19:44.776446Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-20T14:19:44.695462Z digest=sha256:96e48ea86c4bbd85cee3b10881840770c70f566c54094fdef7151c48e235127c

Observation f89e799f-9104-420c-9c93-7bfaa6b22443 · inbound

Scheming Ability in LLM-to-LLM Strategic Interactions cites this paper.

Scheming Ability in LLM-to-LLM Strategic Interactions RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-18T07:51:03.835431Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-18T07:50:30.597108Z digest=sha256:6475f700c1f20895b5e1c2d404eb1d64013dd4f1c3f7cc1ef0cbe85cac73f3a4

Observation 76b13409-8d2b-4366-afea-6a06c93a1a20 · inbound

EcoGym: Evaluating LLMs for Long-Horizon Plan-and-Execute in Interactive Economies cites this paper.

EcoGym: Evaluating LLMs for Long-Horizon Plan-and-Execute in Interactive Economies RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-16T05:57:24.639150Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-16T05:53:29.860037Z digest=sha256:b050310fddea250d1f817a2b2546f10f7db2b4b8dbc9437c0b24af503e38e58f

Observation 9126949c-9665-45a0-b476-7436e91deb07 · inbound

SPM-Bench: Benchmarking Large Language Models for Scanning Probe Microscopy cites this paper.

SPM-Bench: Benchmarking Large Language Models for Scanning Probe Microscopy RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-02T20:36:06.369914Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T20:36:06.369914Z digest=sha256:3f70c4cc2a96fc10b42f5f195d0b64ba71d7893abe453de433f9a9b7b03d46a6

Observation b030444d-0020-438e-a874-0edb4351c374 · inbound

Frontier-Eng: Benchmarking Self-Evolving Agents on Real-World Engineering Tasks with Generative Optimization cites this paper.

Frontier-Eng: Benchmarking Self-Evolving Agents on Real-World Engineering Tasks with Generative Optimization RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-11T09:05:57.739286Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-10T16:17:32.290531Z digest=sha256:750af6b2c02192d2ad5c8bbcb32cbd4cdf25b5741531060bc1a245bea2267399

Observation 8c7a22b9-3c70-4efa-97fd-886a14846ce8 · inbound

TREX: Automating LLM Fine-tuning via Agent-Driven Tree-based Exploration cites this paper.

TREX: Automating LLM Fine-tuning via Agent-Driven Tree-based Exploration RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-11T11:46:34.653244Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-10T12:34:29.808503Z digest=sha256:874f9a3d66fd27c0fa19dd96447f821ac7ae18a5dd5b8670a3f70e7ef55c53de

Observation c7d436cc-01f1-4907-92f6-2883a4edf091 · inbound

Risk Reporting for Developers' Internal AI Model Use cites this paper.

Risk Reporting for Developers' Internal AI Model Use RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-11T23:16:16.444053Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-07T17:47:21.321820Z digest=sha256:248da5df0e1ec0482b8aa17cf6cded0a7571cbdfc39658da1b96a2265154d498

Observation 43f82950-7e26-4aa2-994c-4d56fa16925b · inbound

Principles and Guidelines for Randomized Controlled Trials in AI Evaluation cites this paper.

Principles and Guidelines for Randomized Controlled Trials in AI Evaluation RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts

Reference 34

Resolution
metadata mismatch
arxiv_id, observed 2026-05-09T06:05:36.486105Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-08T18:55:02.645026Z digest=sha256:c2421547cc212837eee9be63d75e91613a623e134a18404d483344aa3cf9be2b

Observation b4e7854c-6681-4446-ae70-14c4cba12e9a · inbound

Collider-Bench: Benchmarking AI Agents with Particle Physics Analysis Reproduction cites this paper.

Collider-Bench: Benchmarking AI Agents with Particle Physics Analysis Reproduction RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts

Reference 21

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T06:05:06.762989Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-15T06:04:03.605898Z digest=sha256:b3843f1a922832a04038860916d4a2a7b2d2284477049ddd1f5ef75fa9023a5d

Observation 002c6092-2b3e-4304-a6b9-38cc4ca0c284 · inbound

FrontierSmith: Synthesizing Open-Ended Coding Problems at Scale cites this paper.

FrontierSmith: Synthesizing Open-Ended Coding Problems at Scale RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-15T02:03:29.062735Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-15T02:02:25.597640Z digest=sha256:07bdfb32ccfc95974e7b1ac012012f6113b0281c9564ac070c2f7134832c7e33

Observation f16d3fd5-9469-4b23-8d75-671c0be93582 · inbound

FML-bench: A Controlled Study of AI Research Agent Strategies from the Perspective of Search Dynamics cites this paper.

FML-bench: A Controlled Study of AI Research Agent Strategies from the Perspective of Search Dynamics RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts

Reference 57

Resolution
verified exact
arxiv_id, observed 2026-05-20T14:28:21.448976Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-20T14:25:15.565386Z digest=sha256:018e9f351f6519ac9ebfee6100f4415e3227c60ff19ab85e58a05be6928b9e27

Observation 98f353b7-e769-4be9-910b-36a33c63fd71 · inbound

FML-bench: A Controlled Study of AI Research Agent Strategies from the Perspective of Search Dynamics cites this paper.

FML-bench: A Controlled Study of AI Research Agent Strategies from the Perspective of Search Dynamics RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts

Reference 57

Resolution
verified exact
arxiv_id, observed 2026-06-30T19:05:00.924661Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-30T19:00:30.961402Z digest=sha256:d1103c668cf1f9e1dafeb4e4a301a60281ab5b836c7f2acded309ffccde19552

Observation f9630578-942a-41d2-9260-a9449fdd468e · inbound

AI for Auto-Research: Roadmap & User Guide cites this paper.

AI for Auto-Research: Roadmap & User Guide RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts

Reference 222

Resolution
verified exact
arxiv_id, observed 2026-05-20T10:33:12.557130Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-20T10:30:50.256635Z digest=sha256:6d0eabbadc545d1dc506a621299397ce1e483a3efb73c6188669b62295bbba3d

Observation 5328db6c-78ad-45aa-86de-e3618d706e3f · inbound

AI for Auto-Research: Roadmap & User Guide cites this paper.

AI for Auto-Research: Roadmap & User Guide RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts

Reference 221

Resolution
unresolved
no resolver link, observed 2026-08-02T13:43:49.951321Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T13:43:49.951321Z digest=sha256:c1a0816d13b2963910603eafb33ba3e2166c69bf5fc9293fbcae77f4530df5cd

Observation 321ddc94-9675-4172-88d5-3f37ac4c7741 · inbound

How Far Are We From True Auto-Research? cites this paper.

How Far Are We From True Auto-Research? RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts

Reference 26

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T09:58:11.242622Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-20T09:56:16.160551Z digest=sha256:24b77408c680d2405e8870616ca796e7f9862e9cf97e7bdcfd13e4e8ad54c3bb

Observation 95e6e6dd-7a45-4c86-a98d-45990b1871f9 · inbound

Hack-Verifiable Environments: Towards Evaluating Reward Hacking at Scale cites this paper.

Hack-Verifiable Environments: Towards Evaluating Reward Hacking at Scale RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-21T06:59:45.559572Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-21T06:56:27.532299Z digest=sha256:43f3801195e8e8a1a5cf299cb64a75e155fc4c231d791a66f8ef4141c99f4d97

Observation 23a23acf-0525-40c6-9917-a28501eda992 · inbound

Benchmarks are Not Enough: RAMP for Runtime Assessing of Agentic Models in Production Systems cites this paper.

Benchmarks are Not Enough: RAMP for Runtime Assessing of Agentic Models in Production Systems RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-06-29T15:33:32.798892Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-29T15:32:21.737028Z digest=sha256:d1948e9cea67703717c28bc35c71e8d29d7bc5cd8ccee9f40440ae2e22308fef

Observation abf9f5f6-6c96-483c-8388-2bd6d7e356b7 · inbound

The Meta-Agent Challenge: Are Current Agents Capable of Autonomous Agent Development? cites this paper.

The Meta-Agent Challenge: Are Current Agents Capable of Autonomous Agent Development? RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-07-02T07:56:47.690735Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-28T06:29:42.665765Z digest=sha256:2e15241d1e8d935c48cd6497d59b6492dd4cf549653d6c83a8f1c09028c00944

Observation 52945ae3-1b6f-4313-b036-adb3f275750d · inbound

AutoLab: Can Frontier Models Solve Long-Horizon Auto Research and Engineering Tasks? cites this paper.

AutoLab: Can Frontier Models Solve Long-Horizon Auto Research and Engineering Tasks? RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts

Reference 42

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T07:46:46.619488Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-06-28T06:37:03.670571Z digest=sha256:c92fe263fcd97990f99b8803cae062291bc192beb123f665d0dcbe0f5c3a27ad

Observation 9668af3d-f1e5-45f1-98eb-3e0245013815 · inbound

InquiTree: Evaluating AI Agents in the Scientific Inquiry Loop with Paper-Derived Research Trees cites this paper.

InquiTree: Evaluating AI Agents in the Scientific Inquiry Loop with Paper-Derived Research Trees RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts

Reference 9

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T04:07:37.255049Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-06-27T14:06:59.471772Z digest=sha256:02900aec85bf6e947911c46b9bf74a8535493423200e63b2bcfe24915e2ac629

Observation efde55ab-60f9-4aae-8e1d-c230c3c7495a · inbound

Toward Generalist Autonomous Research via Hypothesis-Tree Refinement cites this paper.

Toward Generalist Autonomous Research via Hypothesis-Tree Refinement RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts

Reference 160

Resolution
verified exact
arxiv_id, observed 2026-06-27T09:40:47.033546Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-06-27T09:34:41.800309Z digest=sha256:b38f51f4da30b5b203bf42b7251879fd9ede857ccfd4435add3f7b3f9916f1c6

Observation 4cc5987e-a720-497d-9c48-74c0d0c44ba6 · inbound

Position: Coding Benchmarks Are Misaligned with Agentic Software Engineering cites this paper.

Position: Coding Benchmarks Are Misaligned with Agentic Software Engineering RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-07-03T22:08:58.899735Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-26T23:48:58.497927Z digest=sha256:95bd2ae392c351bf36aac655632734d0d13151348688b37c1e190a101b2784bf

Observation 5d39a32f-a09d-4ee2-a9dc-45b587457b01 · inbound

Position: Coding Benchmarks Are Misaligned with Agentic Software Engineering cites this paper.

Position: Coding Benchmarks Are Misaligned with Agentic Software Engineering RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-02T11:06:03.847203Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T11:06:03.847203Z digest=sha256:6443cfded5d10bd389dd119423b42145a693f9226296a276aad2b57288a71cc0

Observation 4806bec1-2f3f-447a-b62b-400a84e06fa3 · inbound

Learning the ARTS of Search for Automated Discovery cites this paper.

Learning the ARTS of Search for Automated Discovery RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-07-04T08:09:41.833473Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-26T12:04:13.117307Z digest=sha256:4019ce6c904df7697dbda25ed193da6d993fe2c72a526327d7cce0d8ce00b8f2

Observation 1ac7c5ff-0cf2-4e78-b901-f62eaf2d6854 · inbound

Discovering Crystal Structure Prediction Algorithms with an AI Co-Scientist cites this paper.

Discovering Crystal Structure Prediction Algorithms with an AI Co-Scientist RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-07-04T10:19:46.987416Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-26T09:03:00.675456Z digest=sha256:7838c653bff0371e53eef230184f3018dbdf329ffcd9f283cd7fa2062d4307af

Observation a0aaa883-453e-4868-ac34-d39a298bd0e0 · inbound

MirrorCode: AI can rebuild entire programs from behavior alone cites this paper.

MirrorCode: AI can rebuild entire programs from behavior alone RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-06-30T06:44:19.703971Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-30T06:35:32.667865Z digest=sha256:f77fabd5b65dcbe34403fad75e3c4a6015c0502bf89c3c73b0383cdd9230ef65

Observation c974db13-d3e9-4d21-8d11-6794418d9f38 · inbound

MirrorCode: AI can rebuild entire programs from behavior alone cites this paper.

MirrorCode: AI can rebuild entire programs from behavior alone RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-02T09:35:02.886229Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T09:35:02.886229Z digest=sha256:347dfdded85dc3447406893c417426bcd999ad0e1e2c16ebaa3fdd698a46ca7f

Observation 9c361421-912f-4c8e-8c9b-46ec8d5cb72a · inbound

Two AI Metrics Diverged: Will it Make All the Difference? cites this paper.

Two AI Metrics Diverged: Will it Make All the Difference? RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts

Reference 68

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T12:36:56.156858Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-07-02T12:29:24.439779Z digest=sha256:8f03ca41f1d436d6ce74d567569b446287c78f2ba41c42ae160a923f764d5855

Observation 4031574a-1bc3-4e7d-84a4-046c5a8f602e · inbound

RuBench: A Repository-Level Agentic Coding Benchmark with Natively Authored Russian Task Specifications cites this paper.

RuBench: A Repository-Level Agentic Coding Benchmark with Natively Authored Russian Task Specifications RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-07-08T07:04:44.520871Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-07-08T06:55:40.830922Z digest=sha256:19303f7aba64b050b6456507177497c455a4a7338bb18098cbff49ba220178f4

Observation 58fcc192-7e31-4da4-944f-af8b18859fb3 · inbound

RuBench: A Repository-Level Agentic Coding Benchmark with Natively Authored Russian Task Specifications cites this paper.

RuBench: A Repository-Level Agentic Coding Benchmark with Natively Authored Russian Task Specifications RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-02T08:21:50.147998Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T08:21:50.147998Z digest=sha256:4649f4d91cdacaa00f99b8d9dcc3e5df4ca2eaa2f8d9405c571bd85b8a9511da

Observation 1296820b-a97b-4a1e-ab5e-dfdc4571568b · inbound

Long-Horizon-Terminal-Bench: Testing the Limits of Agents on Long-Horizon Terminal Tasks with Dense Reward-Based Grading cites this paper.

Long-Horizon-Terminal-Bench: Testing the Limits of Agents on Long-Horizon Terminal Tasks with Dense Reward-Based Grading RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts

Reference 44

Resolution
unresolved
no resolver link, observed 2026-07-13T05:28:45.311405Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T05:28:45.311405Z digest=sha256:cceb873daa8cdc955fc0d26aed7976928d3b69e65f01d19bef9108634a2932ce

Observation c4e9e2a6-d05e-4c08-b8cf-89f4e4b97842 · inbound

Long-Horizon-Terminal-Bench: Testing the Limits of Agents on Long-Horizon Terminal Tasks with Dense Reward-Based Grading cites this paper.

Long-Horizon-Terminal-Bench: Testing the Limits of Agents on Long-Horizon Terminal Tasks with Dense Reward-Based Grading RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts

Reference 44

Resolution
unresolved
no resolver link, observed 2026-07-14T15:24:58.589243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T15:24:58.589243Z digest=sha256:c70563468fd18bdcb420138639509f18e1250c2880c9e4d22cd79e90f2e193f1

Observation 4300af11-4056-4bf9-af7a-3b20f2f358d5 · inbound

When Does Restricting a Coding Agent to execute_code Help? A Regime $\times$ Agent-Design Ablation cites this paper.

When Does Restricting a Coding Agent to execute_code Help? A Regime $\times$ Agent-Design Ablation RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts

Reference 37

Resolution
unresolved
no resolver link, observed 2026-07-14T10:46:01.272433Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T10:46:01.272433Z digest=sha256:0f4be3b2c32cec459ba7e8cfa0f4d0ed4cb3c9d3d9a7fb9caf70f6d1b5b08296

Observation 45bc81b3-20ed-4ca4-a3a7-bf39ecc86496 · inbound

ResearchArena: Evaluating Sabotage and Monitoring in Automated AI R&D cites this paper.

ResearchArena: Evaluating Sabotage and Monitoring in Automated AI R&D RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-01T12:50:00.490138Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T12:50:00.490138Z digest=sha256:8cc8e5151bf05d09b15ecd01c77b7d76427c992e83630d79418dcadda89e7e3c

Observation 4082e15f-b500-40b7-96f1-5feb288df4fa · inbound

Efficiency Matters in Autonomous Research cites this paper.

Efficiency Matters in Autonomous Research RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts

Reference 18

Resolution
unresolved
no resolver link, observed 2026-07-31T09:14:06.207902Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T09:14:06.207902Z digest=sha256:18cf6807109354c575e8e892fd664f87da03e2d9d305b1b1055fe4915d44ab9d

Observation d6f79981-5a6c-4191-b434-dd3ce80357e7 · inbound

Messier: A High-Resolution Corpus for Cross-Benchmark Agent Evaluation cites this paper.

Messier: A High-Resolution Corpus for Cross-Benchmark Agent Evaluation RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-01T01:16:43.437896Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T01:16:43.437896Z digest=sha256:3c759e2f5c25217297525ff98fe50b2369ccbecb48b2a370b7bc9afdef6829b1

Observation d0bf2f4c-2256-4be0-95f0-6a8f25daeea7 · inbound

Safeguards Based on Copyable Context Cannot Provide Reliable Safety for LLMs cites this paper.

Safeguards Based on Copyable Context Cannot Provide Reliable Safety for LLMs RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts

Reference 64

Resolution
unresolved
no resolver link, observed 2026-07-31T22:25:22.349139Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T22:25:22.349139Z digest=sha256:470a87e5b3a508966db6266ddbf2946b3239a0183fee14e73c082fa879bec671

Observation c77e66a2-c16c-433a-b39f-7b6101b3ccfc · inbound

To Add Is Machine, To Delete Is Human: Measuring and Mitigating Deletion Avoidance in LLM Code Editing cites this paper.

To Add Is Machine, To Delete Is Human: Measuring and Mitigating Deletion Avoidance in LLM Code Editing RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-03T01:30:17.866119Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T01:30:17.866119Z digest=sha256:3a190f5903f150ff33446eed7d6f5cf1411363c26297ded3bb4e208d2079a693

Observation a0cc8a4e-0e7d-46b8-a1fd-d9b77f400289 · inbound

Argus: A General-Purpose Agentic Runtime for Long-Horizon Reasoning cites this paper.

Argus: A General-Purpose Agentic Runtime for Long-Horizon Reasoning RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T04:16:45.749872Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T04:16:45.749872Z digest=sha256:8040317ab76630f11f29b7ed425f3410702af45e86818a57455318ede614f69e

Observation fd24008f-7d94-46d1-b45c-723447228ba8 · inbound

Predicting Task Difficulty Without Rollouts cites this paper.

Predicting Task Difficulty Without Rollouts RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts

Reference 1978

Resolution
unresolved
no resolver link, observed 2026-08-07T23:25:20.492888Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T23:25:20.492888Z digest=sha256:699d4bd86bce87390c871b7e5b662208fd1b22258c7938054cd19acb69004c63

Observation c3ee986a-611a-463a-941e-2e856cb77cad · inbound

HarnessOpt-Bench: Evaluating LLMs at Harness Optimization cites this paper.

HarnessOpt-Bench: Evaluating LLMs at Harness Optimization RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T06:02:25.759327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:02:25.759327Z digest=sha256:e60bd70a31267ed4cf20e9321a29b7458ba8f368c366f6f6d54da4b8eaf92121

Observation b1c4f2d5-3f9b-4e58-a2ea-052cc44da010 · inbound

QuantumMind: Constraint-Grounded Agentic Reasoning for Speedup Analysis in Quantum Computing cites this paper.

QuantumMind: Constraint-Grounded Agentic Reasoning for Speedup Analysis in Quantum Computing RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T00:24:44.676649Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T00:24:44.676649Z digest=sha256:6e9fdfcff532a5bb8fa2d1aa26965c2577eb92554837f265e8a200e5066c3f6b

Observation 06f8b810-3b52-4410-a6bc-6f129843f1ff · inbound

Evo-Bench: Can Language Models Improve Agent Harness? cites this paper.

Evo-Bench: Can Language Models Improve Agent Harness? RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-11T23:44:24.561432Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:44:24.561432Z digest=sha256:95372d9125aeb85ea093b4a593ebf4fee99c6c312862b6caf38ff67af00ad3a2

Observation 40910431-7f02-4f95-ab99-34b730e701de · inbound

Evo-Bench: Can Language Models Improve Agent Harness? cites this paper.

Evo-Bench: Can Language Models Improve Agent Harness? RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-14T04:22:26.843491Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:22:26.843491Z digest=sha256:45f605f8fa7751d546013a3367e6f2a3e287464652c96bf0a47f673e49fb73a8

Observation b631cdce-8a01-4068-a59f-eb8bc494689e · inbound

VALG: An Agentic System for ML Theory Research cites this paper.

VALG: An Agentic System for ML Theory Research RE-Bench: Evaluating frontier AI R&D capabilities of language model agents against human experts

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-15T17:44:14.918108Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:44:14.918108Z digest=sha256:c62d25c25c0e2390b6830be3f240ecec1504b7f9dbe64168372f613545f5c196