Pith. sign in

Paper Citation Record · LEDGER

ML-Dev-Bench: Comparative Analysis of AI Agents on ML development workflows

As of 12 August 2026, this Paper Citation Record lists 9 of 9 outbound references and 6 inbound Pith citation observations for arXiv:2502.00964.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.00964 v3

Coverage vector

measured 9 of 9 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-09T17:06:26.773812Z

measured 15 of 15 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 6 of 6 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:49:04.506588Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-30T19:05:00.939699Z

Reference resolution

9 of 9 outbound references displayed

  • verified exact0
  • verified fuzzy5
  • unresolved4
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 3c4753bd-b585-4cfe-9898-b1b2965f5585 · outbound

This paper cites write newline.

ML-Dev-Bench: Comparative Analysis of AI Agents on ML development workflows write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-09T17:06:26.722864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T17:06:26.722864Z digest=sha256:393ca35dc2ed026ca9e457684a705ac054f7d63ac5c1c0f98934122b664cd604

Observation 8217769e-db29-482b-af97-40573b29935b · outbound

This paper cites Program synthesis with large language models, 2021.

ML-Dev-Bench: Comparative Analysis of AI Agents on ML development workflows Program synthesis with large language models, 2021

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-09T17:06:26.729693Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T17:06:26.729693Z digest=sha256:dc70cbad2adf1efc59562cfed57c53db9c309b9c4eeea9b94a4382bddf42e468

Observation 1c1ef3d7-93bd-4946-92f2-8b176be502f5 · outbound

This paper cites Mle-bench: Evaluating machine learning agents on machine learning engineering, 2024.

ML-Dev-Bench: Comparative Analysis of AI Agents on ML development workflows Mle-bench: Evaluating machine learning agents on machine learning engineering, 2024

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T17:06:26.941455Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-09T17:06:26.735267Z digest=sha256:4f68c6430796e32773153c9a41053771598845946017f10e98de435c760a9f63

Observation 0127cd28-a6e0-4457-a703-82413003d552 · outbound

This paper cites Imagenette.

ML-Dev-Bench: Comparative Analysis of AI Agents on ML development workflows Imagenette

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T17:06:26.917512Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-09T17:06:26.740953Z digest=sha256:4a0de9a5b1e43d06a2b2ad065c1f42258cfbe496608eea06bcbfd18881017348

Observation 0a65d93f-cf79-49e7-b564-ca45ae337d83 · outbound

This paper cites Mlagentbench: Evaluating language agents on machine learning experimentation, 2024.

ML-Dev-Bench: Comparative Analysis of AI Agents on ML development workflows Mlagentbench: Evaluating language agents on machine learning experimentation, 2024

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T17:06:26.895909Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-09T17:06:26.749204Z digest=sha256:adf8777b7d3bca3f471ade7fa484d1ee5e772ecd2b9b6d790f03c60154464441

Observation 94bac6e4-7d07-4e63-adc0-06ed9f379220 · outbound

This paper cites Jimenez, John Yang, Alexander Wettig, Shunyu Yao, Kexin Pei, Ofir Press, and Karthik Narasimhan.

ML-Dev-Bench: Comparative Analysis of AI Agents on ML development workflows Jimenez, John Yang, Alexander Wettig, Shunyu Yao, Kexin Pei, Ofir Press, and Karthik Narasimhan

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-09T17:06:26.755541Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T17:06:26.755541Z digest=sha256:8662f0e2d099348641188e9641afbb7d3fe41e0bd5749af9ca2004e4a79ef428

Observation 2e5efd7f-945b-4efd-8e0d-286914883f82 · outbound

This paper cites Ml-bench: Evaluating large language models and agents for machine learning tasks on repository-level code, 2024.

ML-Dev-Bench: Comparative Analysis of AI Agents on ML development workflows Ml-bench: Evaluating large language models and agents for machine learning tasks on repository-level code, 2024

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T17:06:26.866974Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-09T17:06:26.761383Z digest=sha256:b10fb96de6efac5771799418f00b2aa5fa21a8fa3ef9d37523a34f6bf5b733a4

Observation b56fe7ed-a73f-47d0-90b7-004f6c047411 · outbound

This paper cites Xu, Xiangru Tang, Mingchen Zhuge, Jiayi Pan, Yueqi Song, Bowen Li, Jaskirat Singh, Hoang H.

ML-Dev-Bench: Comparative Analysis of AI Agents on ML development workflows Xu, Xiangru Tang, Mingchen Zhuge, Jiayi Pan, Yueqi Song, Bowen Li, Jaskirat Singh, Hoang H

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T17:06:26.845831Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-09T17:06:26.767245Z digest=sha256:9aa2d03137b8f7ea8afd1a3705cc55e39a181ddd0024d32e8fa4683144aca894

Observation c6ad4745-af62-466d-b8e3-a1126e3b4913 · outbound

This paper cites React: Synergizing reasoning and acting in language models, 2023.

ML-Dev-Bench: Comparative Analysis of AI Agents on ML development workflows React: Synergizing reasoning and acting in language models, 2023

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-09T17:06:26.773812Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T17:06:26.773812Z digest=sha256:12232144a9a094e4531ef26329d65d44e6aa0cf335337b62f56c9add93e85e51

Pith citing papers

Observation 0beef106-8e32-460f-ac4f-fb8be80c23bf · inbound

AI Scientists Fail Without Strong Implementation Capability cites this paper.

AI Scientists Fail Without Strong Implementation Capability ML-Dev-Bench: Comparative Analysis of AI Agents on ML development workflows

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T11:49:04.506588Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:49:04.506588Z digest=sha256:c7101b4080a83a715b562a2bb7b7b320d7d2d431d5baa142596bee526d080c15

Observation 560361e4-aec9-4c8a-8535-08f6237a660f · inbound

RExBench: Can coding agents autonomously implement AI research extensions? cites this paper.

RExBench: Can coding agents autonomously implement AI research extensions? ML-Dev-Bench: Comparative Analysis of AI Agents on ML development workflows

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-19T07:37:08.916542Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-19T07:33:39.675929Z digest=sha256:a450be223e84000bcb5df50603c41b1ebfc281dc2a79e1b951d3d0284e80cf11

Observation d2935d8f-ecb8-4424-b53e-823fb388506f · inbound

How Far Are AI Scientists from Changing the World? cites this paper.

How Far Are AI Scientists from Changing the World? ML-Dev-Bench: Comparative Analysis of AI Agents on ML development workflows

Reference 122

Resolution
unresolved
no resolver link, observed 2026-08-06T10:55:14.931704Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T10:55:14.931704Z digest=sha256:009f5dfaa202867930fd7a50c8a8a7a3e15d1f6c804ff472233ada71e8be129c

Observation 090899ee-1e72-423a-ab1a-cdd33192a375 · inbound

FML-bench: A Controlled Study of AI Research Agent Strategies from the Perspective of Search Dynamics cites this paper.

FML-bench: A Controlled Study of AI Research Agent Strategies from the Perspective of Search Dynamics ML-Dev-Bench: Comparative Analysis of AI Agents on ML development workflows

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-20T14:28:21.397811Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-20T14:25:15.565386Z digest=sha256:4e1b4b05ded7eec46a9b350b065590a82d2d2d1912dd44d4104f3579bfa69370

Observation cc480f07-cbee-42dd-a438-fc8875d30cdb · inbound

FML-bench: A Controlled Study of AI Research Agent Strategies from the Perspective of Search Dynamics cites this paper.

FML-bench: A Controlled Study of AI Research Agent Strategies from the Perspective of Search Dynamics ML-Dev-Bench: Comparative Analysis of AI Agents on ML development workflows

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-06-30T19:05:00.941301Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-30T19:00:30.961402Z digest=sha256:a9877c63abf4ce07a8a9b40d615b8f8ab89b672657241bd0c8f305c82971dbf8

Observation 0f4065f3-5a16-4462-9a48-8fe445ebb60c · inbound

Rethinking AI Cloud Infrastructure for Agentic Serving Systems with the Aries Experimentation Framework cites this paper.

Rethinking AI Cloud Infrastructure for Agentic Serving Systems with the Aries Experimentation Framework ML-Dev-Bench: Comparative Analysis of AI Agents on ML development workflows

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-03T14:14:12.016405Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T14:14:12.016405Z digest=sha256:7bf3670eaf29e8479d35c899cd3ba4208c9e61cf210e7fe4c83fde43bb64f425