Pith. sign in

Paper Citation Record · LEDGER

ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery

As of 16 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 67 inbound Pith citation observations for arXiv:2410.05080.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2410.05080 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 67 of 67 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 67 of 67 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T12:22:36.386473Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

6
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation ea070a9b-bd37-49f5-9c32-6d0352d86f8a · inbound

AstroMLab 3: Achieving GPT-4o Level Performance in Astronomy with a Specialized 8B-Parameter Large Language Model cites this paper.

AstroMLab 3: Achieving GPT-4o Level Performance in Astronomy with a Specialized 8B-Parameter Large Language Model ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-12T21:14:04.994413Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:14:04.994413Z digest=sha256:319a6ca69324aac4241f984b5232376724fbb20c62503fd0551faad4611e958d

Observation 29434ba2-8431-4776-9788-b08b6d2a65f3 · inbound

AIGS: Generating Science from AI-Powered Automated Falsification cites this paper.

AIGS: Generating Science from AI-Powered Automated Falsification ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-12T19:02:19.083819Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:02:19.083819Z digest=sha256:bddcd9bf888732a705de7bde48fcf1567da9eb1a9600121f31f22723856e120a

Observation cfb50cf9-b706-4218-be20-4918cdd4f817 · inbound

LLM4SR: A Survey on Large Language Models for Scientific Research cites this paper.

LLM4SR: A Survey on Large Language Models for Scientific Research ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-10T21:39:25.157389Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:39:25.157389Z digest=sha256:50681b1418c979386b3fd7d2a2062bafc9d91ed88b586d59e504eaef8402b2c3

Observation 56403468-0208-471c-8d4d-23555eee82b9 · inbound

Position: Multimodal Large Language Models Can Significantly Advance Scientific Reasoning cites this paper.

Position: Multimodal Large Language Models Can Significantly Advance Scientific Reasoning ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-23T04:32:32.948145Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-05-23T04:30:38.804702Z digest=sha256:2ffc0815afbaa360f99dd8c145de26835571417f42be0069b0a71f4836a742ac

Observation 54e66ece-9b82-4ad2-94d5-dfa5d0a6c5ed · inbound

Sparks of Science: Hypothesis Generation Using Structured Paper Data cites this paper.

Sparks of Science: Hypothesis Generation Using Structured Paper Data ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-16T12:22:36.386473Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T12:22:36.386473Z digest=sha256:4e41a37a48d54673b7c63af74946c56505d466c145ee057671989ec85d0cdce8

Observation 35d435ba-a4ea-4159-9de9-391ecf6e661d · inbound

From LLM Reasoning to Autonomous AI Agents: A Comprehensive Review cites this paper.

From LLM Reasoning to Autonomous AI Agents: A Comprehensive Review ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery

Reference 98

Resolution
verified exact
arxiv_id, observed 2026-05-15T02:57:38.399364Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-15T02:57:37.873567Z digest=sha256:6059f8e3745ddcfb4d7b2362b5cb8c958b7989f7844652c87762a0f1b8bc1c2b

Observation ef1c04c4-3289-40a4-9644-8b12fc8bbd36 · inbound

Can AI Agents Design and Implement Drug Discovery Pipelines? cites this paper.

Can AI Agents Design and Implement Drug Discovery Pipelines? ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-16T05:44:03.815973Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T05:44:03.815973Z digest=sha256:84f693879b33387bb25939e58a70d4cb4b4be7caf9c59038bf89e03e1f3d6705

Observation 817f76da-d195-4d65-8fae-2f5f5bd6b4e5 · inbound

ResearchCodeAgent: An LLM Multi-Agent System for Automated Codification of Research Methodologies cites this paper.

ResearchCodeAgent: An LLM Multi-Agent System for Automated Codification of Research Methodologies ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-16T05:53:26.078433Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T05:53:26.078433Z digest=sha256:b6258e9fe2126c6c953ad19b882c0b59f41b1941a24a38522e9ad6727d0b19c8

Observation 57b8442d-9b0d-43ef-8a8a-b526983ee376 · inbound

BioDSA-1K: Benchmarking Data Science Agents for Biomedical Research cites this paper.

BioDSA-1K: Benchmarking Data Science Agents for Biomedical Research ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T15:09:32.267308Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:09:32.267308Z digest=sha256:18b3feb122f3dc0642b8f8d84e1ce82204b2c7b400ad2263be9b93b3595a1bf7

Observation 88508b9d-2c41-4609-a8b3-28574d48c1da · inbound

EXP-Bench: Can AI Conduct AI Research Experiments? cites this paper.

EXP-Bench: Can AI Conduct AI Research Experiments? ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T12:20:42.230426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:20:42.230426Z digest=sha256:6ff1d6bdadce1705e774c99fae791fc4b5cf56cfae0be483d436d333f8503e01

Observation b604650e-6e04-4d51-8e11-39f3f68716b7 · inbound

Agent-X: Evaluating Deep Multimodal Reasoning in Vision-Centric Agentic Tasks cites this paper.

Agent-X: Evaluating Deep Multimodal Reasoning in Vision-Centric Agentic Tasks ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T12:18:50.860148Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:18:50.860148Z digest=sha256:daf7752e2d205437f0d01bce896216ae6c97ef894ae83e5fcebeb654ebb41ba1

Observation d9e984a3-57ee-4d7a-bf01-87859956317b · inbound

TextAtari: 100K Frames Game Playing with Language Agents cites this paper.

TextAtari: 100K Frames Game Playing with Language Agents ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-07T10:51:57.940291Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:51:57.940291Z digest=sha256:35d99826ec1aa21d9b143f32202212bdc22143f05f73e377011870fd57fcabb6

Observation 6a271e68-1308-4909-8500-8452e86d43b9 · inbound

Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey cites this paper.

Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery

Reference 98

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:19.467715Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:19.467715Z digest=sha256:36e173e5a555643f54f82ecf35398262969f09eb1a9863f88c2b6c27d26211a3

Observation 8c408575-b0e8-4906-a807-f5ffc28cd7db · inbound

DeepResearch Bench: A Comprehensive Benchmark for Deep Research Agents cites this paper.

DeepResearch Bench: A Comprehensive Benchmark for Deep Research Agents ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T08:07:39.431114Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-16T08:07:39.384613Z digest=sha256:bbef771b26da858911e713ebe1ee47b9fa941ebe798ee1ecadc1d5c3772bf073

Observation 0b835d3e-c055-4cc5-a265-10a43a846660 · inbound

Perspective on Utilizing Foundation Models for Laboratory Automation in Materials Research cites this paper.

Perspective on Utilizing Foundation Models for Laboratory Automation in Materials Research ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T00:55:16.881460Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:55:16.881460Z digest=sha256:ff50860cbbf8e139619ad8b2f7840b710d2540c973a3499612b27777ff5e5e3c

Observation 5d1f6fcf-7734-4252-98b0-c18c8d5ba657 · inbound

SHADE-Arena: Evaluating Sabotage and Monitoring in LLM Agents cites this paper.

SHADE-Arena: Evaluating Sabotage and Monitoring in LLM Agents ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-15T19:54:52.985789Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:54:52.985789Z digest=sha256:d54725534f7d89e2fd163f7f303c870bc82a38dc4f51fc1f06b51bb3abf2df5d

Observation 3e509b4e-623e-4f89-aee5-3cf5d378c250 · inbound

Deep Research Agents: A Systematic Examination And Roadmap cites this paper.

Deep Research Agents: A Systematic Examination And Roadmap ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T23:26:56.564539Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:26:56.564539Z digest=sha256:74b892ee5c908ffee3f9667e8429b105cc73cf3e65d87557b99e46b15280e0ff

Observation 43a63161-5ccf-4255-a97f-5f0e215d83d2 · inbound

Doc2Agent: Scalable Generation of Tool-Using Agents from API Documentation cites this paper.

Doc2Agent: Scalable Generation of Tool-Using Agents from API Documentation ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T23:05:06.557517Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:05:06.557517Z digest=sha256:f3c7b48385a7e930a567effdd768f9b9867b372f165542280171eb6abed1bc76

Observation cb1f5943-ef43-4962-834f-af9a07824c62 · inbound

RExBench: Can coding agents autonomously implement AI research extensions? cites this paper.

RExBench: Can coding agents autonomously implement AI research extensions? ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-19T07:37:08.883664Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-19T07:33:39.675929Z digest=sha256:14903998d5f0e17c4cc158902dfb24d7e4e2737ef5dd428af1f81c5ac856d3a2

Observation 23eb0a43-0b73-46a5-a1dd-d05265413501 · inbound

AI4Research: A Survey of Artificial Intelligence for Scientific Research cites this paper.

AI4Research: A Survey of Artificial Intelligence for Scientific Research ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery

Reference 127

Resolution
unresolved
no resolver link, observed 2026-08-06T20:45:12.313531Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:45:12.313531Z digest=sha256:43eeb4ebaf3f8d8640b80c7e5b11da7cd44c852b2f96d3c5fc98a3f5736afa36

Observation 41d00269-018b-4580-a70c-f1251c57e29f · inbound

A Survey of Self-Evolving Agents: What, When, How, and Where to Evolve on the Path to Artificial Super Intelligence cites this paper.

A Survey of Self-Evolving Agents: What, When, How, and Where to Evolve on the Path to Artificial Super Intelligence ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery

Reference 74

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T22:23:15.664362Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-05-14T22:23:14.621091Z digest=sha256:b3835c83301e61700444ef09aba6ca2e3ca201e7df793382d557182ddd38e86e

Observation 2094d729-a247-4398-969c-1bfd1da0a94a · inbound

Evaluation and Benchmarking of LLM Agents: A Survey cites this paper.

Evaluation and Benchmarking of LLM Agents: A Survey ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T12:44:21.520560Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:44:21.520560Z digest=sha256:4fc040c986401e072150ac399a69343f6d2cdf42f757147120de43c123059692

Observation 91b010aa-dd33-4aed-8976-093534736d28 · inbound

How Far Are AI Scientists from Changing the World? cites this paper.

How Far Are AI Scientists from Changing the World? ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T10:55:14.623103Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T10:55:14.623103Z digest=sha256:8589d237c427b0d9104998340f4ce4fd87667ed3c435b6c0e47fc5c554fba91c

Observation 4ebaa39f-9807-4cd0-99e4-a865f665e472 · inbound

GeoAnalystBench: A GeoAI benchmark for assessing large language models for spatial analysis workflow and code generation cites this paper.

GeoAnalystBench: A GeoAI benchmark for assessing large language models for spatial analysis workflow and code generation ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T16:24:25.055333Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T16:24:25.055333Z digest=sha256:fbe4e2ceec59c661bcf2fd6411a005ce2ca0aeaf00ddf649791fb0dd26d0643b

Observation 56d6a0f3-0ff2-47df-9368-50f9ee986c56 · inbound

CFDLLMBench: A Benchmark Suite for Evaluating Large Language Models in Computational Fluid Dynamics cites this paper.

CFDLLMBench: A Benchmark Suite for Evaluating Large Language Models in Computational Fluid Dynamics ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-18T15:06:32.002005Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-18T15:05:37.519850Z digest=sha256:8fb6c3cfa8cde8b0d4e060032c3fb6fde3a86be1d422b89213fbe24833979ce0

Observation 6d279899-50c3-4d4f-9689-283aec17daf3 · inbound

How can we assess human-agent interactions? Case studies in software agent design cites this paper.

How can we assess human-agent interactions? Case studies in software agent design ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-04T10:36:36.771781Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T10:36:36.771781Z digest=sha256:83f6efde9f9010a7087b5154779cde14ffa71dd6e4e5c70a71af44c9f9a2818e

Observation 808a1a95-e296-4a59-b58f-27eb79084f56 · inbound

Gated KalmaNet: A Fading Memory Layer Through Test-Time Ridge Regression cites this paper.

Gated KalmaNet: A Fading Memory Layer Through Test-Time Ridge Regression ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-21T18:00:27.120470Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-21T17:59:23.826110Z digest=sha256:ca4dbb74e48bb7dfd96d8eb3449830722a516e08e46fcb49d687c08f6399a683

Observation 23341957-a990-4a2c-91e3-80cffd3963ab · inbound

ThunderAgent: A Simple, Fast and Program-Aware Agentic Inference System cites this paper.

ThunderAgent: A Simple, Fast and Program-Aware Agentic Inference System ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-02T23:32:11.594053Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:32:11.594053Z digest=sha256:b16ef565559af45fe5776f6e34502696a2bb7cba7d7972ff92ec3575ef044208

Observation 3bf352c4-3f33-44a0-9224-6a15da628d5a · inbound

OpenHospital: A Thing-in-itself Arena for Evolving and Benchmarking LLM-based Collective Intelligence cites this paper.

OpenHospital: A Thing-in-itself Arena for Evolving and Benchmarking LLM-based Collective Intelligence ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery

Reference 2

Resolution
unresolved
no resolver link, observed 2026-07-14T21:00:05.911327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T21:00:05.911327Z digest=sha256:c7cdb2de577c9ce70f541697282d16fea8f31241cc36a9ae6fa1a75f4689c875

Observation 88755ec6-6b79-4c94-94c8-2e215a67053b · inbound

TREX: Automating LLM Fine-tuning via Agent-Driven Tree-based Exploration cites this paper.

TREX: Automating LLM Fine-tuning via Agent-Driven Tree-based Exploration ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-11T11:46:34.990899Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-10T12:34:29.808503Z digest=sha256:9d28de89d19f21fc93e66119a5ded064860101d9738fdfbdc32f2816f9eba9fe

Observation 512bc90b-06e6-4d78-9603-fd06ed6e08fb · inbound

EvoMaster: A Foundational Evolving Agent Framework for Agentic Science at Scale cites this paper.

EvoMaster: A Foundational Evolving Agent Framework for Agentic Science at Scale ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-10T06:01:13.389705Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-10T05:59:01.010437Z digest=sha256:1873b4e44a111fbaf64394eb3249f3dcbee6692890eb52dd9b4051481bb2175a

Observation eb0aac0c-7f63-4c95-9f21-f7fb03cd1434 · inbound

EvoMaster: A Foundational Evolving Agent Framework for Agentic Science at Scale cites this paper.

EvoMaster: A Foundational Evolving Agent Framework for Agentic Science at Scale ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-07-05T17:51:14.762793Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-07-05T17:45:55.631459Z digest=sha256:6bddaacd7f7c53f0c4ff03d733a0b88c71f589970d5406f1fe6d3c6821f84145

Observation e3d50f7d-c52f-4211-b7dd-ab70da14f2fd · inbound

AI scientists produce results without reasoning scientifically cites this paper.

AI scientists produce results without reasoning scientifically ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-11T12:16:07.053196Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-10T03:56:34.581133Z digest=sha256:d3378605f3003c82d932c9ca3bef5eaafa4947be5424b408bdf979dd36c709a6

Observation 008b456a-fafd-4dc0-bca1-ff4fbf8f2be5 · inbound

Agentic-imodels: Evolving agentic interpretability tools via autoresearch cites this paper.

Agentic-imodels: Evolving agentic interpretability tools via autoresearch ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-11T23:36:36.314009Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-07T16:37:43.371592Z digest=sha256:920b8c3a1c2fad878bf9b49b55d35551da9e50f72ec8a1e86570c16c8faaf2f8

Observation 7db4a2ba-ff3c-4566-88de-01bd7160842b · inbound

Experiment-as-Code Labs: A Declarative Stack for AI-Driven Scientific Discovery cites this paper.

Experiment-as-Code Labs: A Declarative Stack for AI-Driven Scientific Discovery ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T17:31:07.935615Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-05-08T17:28:41.217810Z digest=sha256:ee558f3164c7b27af6971b69e7876d5b7c62a3a0731d2928ca9558b4455491d3

Observation d1d1eb3b-a264-41a7-9589-301092f1dc7f · inbound

Experiment-as-Code Labs: A Declarative Stack for AI-Driven Scientific Discovery cites this paper.

Experiment-as-Code Labs: A Declarative Stack for AI-Driven Scientific Discovery ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-05-21T00:13:52.856073Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-05-21T00:13:07.546472Z digest=sha256:bc5fccd069ea856a39760456ce2b524c16d956952aa41ebbb2ef2522d0a2bd64

Observation bd5541f9-db46-4f0e-9f26-4ed70b4ec7d4 · inbound

gwBenchmarks: Stress-Testing LLM Agents on High-Precision Gravitational Wave Astronomy cites this paper.

gwBenchmarks: Stress-Testing LLM Agents on High-Precision Gravitational Wave Astronomy ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-13T02:17:06.419263Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-13T02:14:48.091245Z digest=sha256:319ff1f816e2e1dd43c00e37f696a2cf188267059923d9f24b7f1a4dd246261a

Observation 7c24fc5b-fbc7-401a-8e85-214ee5d8e4b9 · inbound

Neurodata Without Boredom: Benchmarking Agentic AI for Data Reuse cites this paper.

Neurodata Without Boredom: Benchmarking Agentic AI for Data Reuse ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery

Reference 20

Resolution
metadata mismatch
arxiv_id, observed 2026-05-14T20:42:59.091537Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-14T20:23:34.514071Z digest=sha256:9b031c93fc68a9ce0c614efbf5829c00c58e16e98c07366991141f9d9fe12ee5

Observation f9d3a9a8-70bf-4e06-8679-25a5b4d2b280 · inbound

Neurodata Without Boredom: Benchmarking Agentic AI for Data Reuse cites this paper.

Neurodata Without Boredom: Benchmarking Agentic AI for Data Reuse ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery

Reference 20

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T04:55:04.185995Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-15T04:51:17.519200Z digest=sha256:44350faed1c8341c0d405d009314044c9e540f881b4951fd1f8f24c6beb2779f

Observation d9150b61-a096-4425-8ef3-c022481c9690 · inbound

Sheaf-Theoretic Transport and Obstruction for Detecting Scientific Theory Shift in AI Agents cites this paper.

Sheaf-Theoretic Transport and Obstruction for Detecting Scientific Theory Shift in AI Agents ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-15T05:49:48.155452Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-05-15T05:45:13.652112Z digest=sha256:9e4442cf24e005cae36cc5b01626f99feb2647558c70c4a7d30191827d1b9b5e

Observation 9ae8b600-47c5-42ad-9d29-ecefac5a3aa4 · inbound

MLReplicate: Benchmarking Autonomous Research Systems for Machine Learning Reproducibility cites this paper.

MLReplicate: Benchmarking Autonomous Research Systems for Machine Learning Reproducibility ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T20:03:44.130932Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-20T19:59:40.519962Z digest=sha256:a461fdf21c6fa3b459a72f38ffd362c2fc3c7d13dfbb573a35a096cf84f40dda

Observation 4d64616a-0ed4-4864-8077-eb79813f0a3a · inbound

FML-bench: A Controlled Study of AI Research Agent Strategies from the Perspective of Search Dynamics cites this paper.

FML-bench: A Controlled Study of AI Research Agent Strategies from the Perspective of Search Dynamics ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-20T14:28:21.416859Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-20T14:25:15.565386Z digest=sha256:1b6a7eb1845710763ea11e3c09a3b3fdc7d99b0f187cdd602ac9a636f279828c

Observation 6c4e6d4f-f2e8-47be-8365-17f4c87e9d6a · inbound

FML-bench: A Controlled Study of AI Research Agent Strategies from the Perspective of Search Dynamics cites this paper.

FML-bench: A Controlled Study of AI Research Agent Strategies from the Perspective of Search Dynamics ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-06-30T19:05:00.919428Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-06-30T19:00:30.961402Z digest=sha256:8786970ea39860c5f39922aa9d3b50bf2e914d9fd199821b19eaf2b23db7ceb2

Observation b816edfd-d501-41f4-a7d6-3f0852dfaea2 · inbound

SCICONVBENCH: Benchmarking LLMs on Multi-Turn Clarification for Task Formulation in Computational Science cites this paper.

SCICONVBENCH: Benchmarking LLMs on Multi-Turn Clarification for Task Formulation in Computational Science ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-20T10:38:12.174291Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-20T10:36:09.724234Z digest=sha256:4c2b69cc75cdf2d50655c5f866025f559d3006f674f8bb8f1f6ea44a9bb5af69

Observation f05f4a13-68ad-4357-8163-5a251d1de504 · inbound

How Far Are We From True Auto-Research? cites this paper.

How Far Are We From True Auto-Research? ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery

Reference 22

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T09:58:11.275989Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-05-20T09:56:16.160551Z digest=sha256:18700d2bab6cf246ac5833550151f9af9eee5ea164b39ad87dac0053a800435e

Observation 64537fce-d2f5-4113-82fb-612ab3852546 · inbound

Matter to Mechanism: A Benchmark for AI Co-Scientists in Materials and Battery Research cites this paper.

Matter to Mechanism: A Benchmark for AI Co-Scientists in Materials and Battery Research ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery

Reference 43

Resolution
metadata mismatch
arxiv_id, observed 2026-06-28T12:12:07.893078Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-06-28T12:08:10.552789Z digest=sha256:b28aaef74dac65e466205a3b0fa995eb38d7d9c2d7c9896124627458af3f255f

Observation f4bb8ced-03cd-4618-af61-ea1daf4023b1 · inbound

InquiTree: Evaluating AI Agents in the Scientific Inquiry Loop with Paper-Derived Research Trees cites this paper.

InquiTree: Evaluating AI Agents in the Scientific Inquiry Loop with Paper-Derived Research Trees ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T04:07:37.268768Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-06-27T14:06:59.471772Z digest=sha256:b581046667f8ca7bd09ebb32e80d448e65cc20b22e7d52648332880f40f96a8a

Observation eec0945f-3e54-43de-a0ad-010904ff15c6 · inbound

Auto-Configuring Scientific Simulators with Lightweight Coding-Agent Adapters cites this paper.

Auto-Configuring Scientific Simulators with Lightweight Coding-Agent Adapters ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-07-03T01:47:32.158536Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-06-27T16:14:26.278017Z digest=sha256:6131c933872129c5447eb95080c38c5f81801776a3d0f527f5f7bced09c5b3c7

Observation 2584bc5b-0ca1-4b18-8b8b-281f88b05362 · inbound

Auto-Configuring Scientific Simulators with Lightweight Coding-Agent Adapters cites this paper.

Auto-Configuring Scientific Simulators with Lightweight Coding-Agent Adapters ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-06-29T15:23:33.009993Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-06-29T05:25:29.078764Z digest=sha256:57d599253783deb776d3cc68a6736a094b8eeb65722c4ee29f78eb11742cff3a

Observation 56fd1037-0225-4f58-8cd5-8a890818a602 · inbound

Data Journalist Agent: Transforming Data into Verifiable Multimodal Stories cites this paper.

Data Journalist Agent: Transforming Data into Verifiable Multimodal Stories ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-07-03T04:37:37.774873Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-06-27T13:43:02.919248Z digest=sha256:b830685a4b97c09cd20ead13b6f45ef270da8219b0be6854a6d3fad92fabba61

Observation de9f9559-3f59-4240-9449-302bfad1419c · inbound

Toward Generalist Autonomous Research via Hypothesis-Tree Refinement cites this paper.

Toward Generalist Autonomous Research via Hypothesis-Tree Refinement ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery

Reference 118

Resolution
metadata mismatch
arxiv_id, observed 2026-06-27T09:40:47.024579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-06-27T09:34:41.800309Z digest=sha256:e25c8008cd276fbfc1bbee256cc49ba55ee51a2a0355a25d00c411466804722d

Observation 5e743964-8983-4ece-a96c-7cb1ab00d7b8 · inbound

Beyond NL2Code: A Structured Survey of Multimodal Code Intelligence cites this paper.

Beyond NL2Code: A Structured Survey of Multimodal Code Intelligence ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery

Reference 10

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T17:38:44.063091Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-06-27T03:57:19.028507Z digest=sha256:b090b72869ab1b1bfd18328a74e4e9f62a9d13a1bd7eb36918aa6c05b81e0bd6

Observation bd4861ec-7298-450a-83f5-50c4d597d14e · inbound

MetaSyn: A Benchmark for LLM Agents on Meta-Analysis Articles from Nature Portfolio cites this paper.

MetaSyn: A Benchmark for LLM Agents on Meta-Analysis Articles from Nature Portfolio ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-07-03T17:48:45.699996Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-06-27T03:48:51.497061Z digest=sha256:c8467c4db1e2eaebe493ffa8db507475ae73bab6ec95b9ea6768932f9b5bf2ec

Observation 7a4b9f53-6523-4f29-94fa-10146fbcdeeb · inbound

MetaSyn: A Benchmark for LLM Agents on Meta-Analysis Articles from Nature Portfolio cites this paper.

MetaSyn: A Benchmark for LLM Agents on Meta-Analysis Articles from Nature Portfolio ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-06-30T11:54:38.905083Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-06-30T10:23:03.764108Z digest=sha256:97c655e1f7221b7ad47e943e63f5ad8b22cdd0adaab0c6a7b69f27f07f2d2be2

Observation cc54a9da-76af-46ac-abc7-ab3a7a9f9119 · inbound

MetaSyn: A Benchmark for LLM Agents on Meta-Analysis Articles from Nature Portfolio cites this paper.

MetaSyn: A Benchmark for LLM Agents on Meta-Analysis Articles from Nature Portfolio ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-07-02T22:07:25.722606Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-07-02T22:05:37.948336Z digest=sha256:698b93180797abd244b0e6a14ceb62daa45e4e6776b0fed109208190ad809241

Observation 20aa0762-abaa-4548-a728-44d9794ff469 · inbound

MetaSyn: A Benchmark for LLM Agents on Meta-Analysis Articles from Nature Portfolio cites this paper.

MetaSyn: A Benchmark for LLM Agents on Meta-Analysis Articles from Nature Portfolio ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-02T11:08:45.071671Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T11:08:45.071671Z digest=sha256:b4baf293232e86a3498add4a86832c2e6403c178e97d97014cd7d401abc74934

Observation ffaa95bd-4fa8-4215-a572-169e113bc2a2 · inbound

OPD-Evolver: Cultivating Holistic Agent Evolver via On-Policy Distillation cites this paper.

OPD-Evolver: Cultivating Holistic Agent Evolver via On-Policy Distillation ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery

Reference 93

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T20:48:56.358749Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-06-27T01:07:49.603969Z digest=sha256:f2bc6edfcff9f17fb8268a7e17c6d849d74433cf3b2acb69daf93b8d65d6d4b8

Observation 97ae98a8-5aca-4e41-8377-b824b5aae258 · inbound

SafeClawBench: Separating Semantic, Audit-Evidence, and Sandbox Harm in Tool-Using LLM Agents cites this paper.

SafeClawBench: Separating Semantic, Audit-Evidence, and Sandbox Harm in Tool-Using LLM Agents ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-07-03T22:18:59.717212Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-06-26T23:40:53.439549Z digest=sha256:33ab922269aed7449a5cc82950927f60cb7f164c76b609123eda8c581b871ccf

Observation 4ebafdfb-d718-4a22-906d-22d78ff414fc · inbound

Deep Research in Physical Sciences: A Multi-Agent Framework and Comprehensive Benchmark cites this paper.

Deep Research in Physical Sciences: A Multi-Agent Framework and Comprehensive Benchmark ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-07-04T02:49:24.272487Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-06-26T19:17:17.373463Z digest=sha256:79d7435ad10b91aa990d753b70d126b0710389a1b55b0e2baba58c6a932334c6

Observation 44d78091-2566-498b-b749-93971a525d1a · inbound

Marginal Advantage Accumulation for Memory-Driven Agent Self-Evolution cites this paper.

Marginal Advantage Accumulation for Memory-Driven Agent Self-Evolution ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery

Reference 39

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T03:39:30.801028Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-06-26T17:45:49.272070Z digest=sha256:3f29197eb18fccbef5d4cb6c0ae1b9a4c1e069d547af41a92e64d49505a1f895

Observation 39f1b902-6d5b-4f37-bece-6c0a835556e3 · inbound

How Do Tool-Augmented LLM Agents Perform on Real-World Energy Analytics Tasks? cites this paper.

How Do Tool-Augmented LLM Agents Perform on Real-World Energy Analytics Tasks? ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-07-04T15:29:56.683593Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-06-26T01:34:10.103638Z digest=sha256:d73ad2c82e6381289a46e3c665fa554ac4c79d2ffcfa454835c0d9d6c2cc5d7e

Observation 6aca5516-39b5-4cb5-a91b-53c26b4307a6 · inbound

Recursive Self-Improvement in AI: From Bounded Self-Refinement to Autonomous Research Loops cites this paper.

Recursive Self-Improvement in AI: From Bounded Self-Refinement to Autonomous Research Loops ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery

Reference 188

Resolution
metadata mismatch
local_arxiv, observed 2026-07-09T03:45:55.719982Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-07-09T03:36:57.168246Z digest=sha256:6bffa61bf38ec00c059f04528ec9131693c968b08d7f1365fd49932dc192ba04

Observation ebd57547-9138-42c3-9bcf-b982a4fac180 · inbound

Evolutionary Intelligence for Scientific Discovery: From Evolutionary Computation to Cumulative Discovery Systems cites this paper.

Evolutionary Intelligence for Scientific Discovery: From Evolutionary Computation to Cumulative Discovery Systems ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery

Reference 48

Resolution
unresolved
no resolver link, observed 2026-07-13T00:55:17.107245Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T00:55:17.107245Z digest=sha256:f1a3ddb3eb3e2c358ea2f0c405f00cf04943af885766dd97ff3f004196e8b21e

Observation f3b9591a-9c4c-4707-b2a8-8863478fcf5e · inbound

Imaging-101: Benchmarking LLM Coding Agents on Scientific Computational Imaging cites this paper.

Imaging-101: Benchmarking LLM Coding Agents on Scientific Computational Imaging ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery

Reference 22

Resolution
unresolved
no resolver link, observed 2026-07-14T09:17:00.467000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T09:17:00.467000Z digest=sha256:806ec7ca921107a410f8d5a21d8d85ea39775f975bb371047ebc5c7ded8d50b7

Observation 8c12ae4c-2cc0-45c9-bd76-b4e32ef82e35 · inbound

SciCodePile: A 128GB Corpus and Executable Benchmark for Challenging Scientific Code Generation cites this paper.

SciCodePile: A 128GB Corpus and Executable Benchmark for Challenging Scientific Code Generation ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-01T13:28:32.903465Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T13:28:32.903465Z digest=sha256:b3aef5372f05c94d40a2b58d8d78fdb43b89de8d0411bce30afe77661d507d5b

Observation 487e51cc-4ac0-47cd-a22e-49062cd04b05 · inbound

SciDataSailor: Deep Scientific Data Exploring cites this paper.

SciDataSailor: Deep Scientific Data Exploring ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-01T15:02:37.896688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:02:37.896688Z digest=sha256:0b212a878e43482ed3a18cb7fdcff65bccc7b5be4196131a691fa8ea40961f3b

Observation a3ab6b94-0a2d-4abd-b2ed-7c5c83063fe8 · inbound

Adversarial Fast-Moving Real-World Domains as Test Beds for Benchmarking AI Scientist Capabilities cites this paper.

Adversarial Fast-Moving Real-World Domains as Test Beds for Benchmarking AI Scientist Capabilities ScienceAgentBench: Toward Rigorous Assessment of Language Agents for Data-Driven Scientific Discovery

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-05T16:49:30.068458Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:49:30.068458Z digest=sha256:f4e76f6ba8c8879aaeabbbe104078f2add614e7a1babc9d02e890238736772d4