Pith. sign in

Paper Citation Record · LEDGER

ML-Bench: Evaluating Large Language Models and Agents for Machine Learning Tasks on Repository-Level Code

As of 22 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 20 inbound Pith citation observations for arXiv:2311.09835.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2311.09835 v5

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 20 of 20 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 20 of 20 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T20:41:08.909862Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-01T14:05:46.736275Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation ebf24730-2b79-4bb7-ac86-c752e9e58ab9 · inbound

MetaGPT: Meta Programming for A Multi-Agent Collaborative Framework cites this paper.

MetaGPT: Meta Programming for A Multi-Agent Collaborative Framework ML-Bench: Evaluating Large Language Models and Agents for Machine Learning Tasks on Repository-Level Code

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-11T03:43:18.933691Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-05-11T03:43:18.632292Z digest=sha256:4d90f33737765490f3c06d2564877103adeaa5131a01fccfdac020b5c7b5287a

Observation 230b2e79-42ba-4507-bf2f-bb3b8e302bd2 · inbound

MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering cites this paper.

MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering ML-Bench: Evaluating Large Language Models and Agents for Machine Learning Tasks on Repository-Level Code

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-23T19:13:21.535616Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-05-23T19:11:20.600633Z digest=sha256:7be8cc15d0ba423cc918b20da07ff03e22b48e56f4bfa03b959d917f62d1d637

Observation a4df26b6-6cc6-4622-9203-23ed6fd64527 · inbound

AIGS: Generating Science from AI-Powered Automated Falsification cites this paper.

AIGS: Generating Science from AI-Powered Automated Falsification ML-Bench: Evaluating Large Language Models and Agents for Machine Learning Tasks on Repository-Level Code

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-12T19:02:19.266981Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:02:19.266981Z digest=sha256:a81e73e1ba1cfdafac3707a4b32319a1ae50cf5583e85474d8cb537e7b2781ce

Observation 19ebd076-1d70-427d-912f-77193fece6a3 · inbound

OSS-Bench: Benchmark Generator for Coding LLMs cites this paper.

OSS-Bench: Benchmark Generator for Coding LLMs ML-Bench: Evaluating Large Language Models and Agents for Machine Learning Tasks on Repository-Level Code

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T20:41:08.909862Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:41:08.909862Z digest=sha256:723ff996bfb6f40e90cd5c5fab44f0d584fd56a4bf687d4f3bd947927876adaf

Observation ebfff6ab-3eb5-467d-97f9-cefbfd68fe81 · inbound

TimeSeriesGym: A Scalable Benchmark for (Time Series) Machine Learning Engineering Agents cites this paper.

TimeSeriesGym: A Scalable Benchmark for (Time Series) Machine Learning Engineering Agents ML-Bench: Evaluating Large Language Models and Agents for Machine Learning Tasks on Repository-Level Code

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-15T20:18:59.955968Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:18:59.955968Z digest=sha256:fb50ce6232be8103153e1a9883c4d19ef76871e036e8d2a4e720b49f9b7c199e

Observation 3308b1be-fe72-42e4-a9a4-f700438c81a8 · inbound

RepoMaster: Autonomous Exploration and Understanding of GitHub Repositories for Complex Task Solving cites this paper.

RepoMaster: Autonomous Exploration and Understanding of GitHub Repositories for Complex Task Solving ML-Bench: Evaluating Large Language Models and Agents for Machine Learning Tasks on Repository-Level Code

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T13:53:16.268485Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:53:16.268485Z digest=sha256:93890e1baaf9410b82c9b68e6ee97b805e447c8a89d429a871d7fa3389c5ce77

Observation 0a3c29a7-66ad-42f9-8b46-abb6bc0485fa · inbound

From Standalone LLMs to Integrated Intelligence: A Survey of Compound Al Systems cites this paper.

From Standalone LLMs to Integrated Intelligence: A Survey of Compound Al Systems ML-Bench: Evaluating Large Language Models and Agents for Machine Learning Tasks on Repository-Level Code

Reference 173

Resolution
verified exact
arxiv_id, observed 2026-05-19T11:52:16.430624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-19T11:49:36.574471Z digest=sha256:f97c3c222e9edd817d8e31b48127fd712dc9581c494eeac4aeeb29e3e0cedb69

Observation d59ecd38-95a2-4267-810e-79a9d494d2f9 · inbound

Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey cites this paper.

Evolutionary Perspectives on the Evaluation of LLM-Based AI Agents: A Comprehensive Survey ML-Bench: Evaluating Large Language Models and Agents for Machine Learning Tasks on Repository-Level Code

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T06:00:18.492754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:00:18.492754Z digest=sha256:c06f448bc1783c88529d0ac739abc93e9300e8b5b80d08368a740666827c454e

Observation 883b6106-d5c8-4d33-95d9-44f1c38e3cb9 · inbound

LMR-BENCH: Evaluating LLM Agent's Ability on Reproducing Language Modeling Research cites this paper.

LMR-BENCH: Evaluating LLM Agent's Ability on Reproducing Language Modeling Research ML-Bench: Evaluating Large Language Models and Agents for Machine Learning Tasks on Repository-Level Code

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T23:51:19.113797Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:51:19.113797Z digest=sha256:d6b2a9315117e86751794228d15366ee33e6181c076b17d57441521791fef747

Observation 9c8a6073-8551-4f3e-9228-a2b53f7f9e2a · inbound

Deep Research Agents: A Systematic Examination And Roadmap cites this paper.

Deep Research Agents: A Systematic Examination And Roadmap ML-Bench: Evaluating Large Language Models and Agents for Machine Learning Tasks on Repository-Level Code

Reference 105

Resolution
unresolved
no resolver link, observed 2026-08-06T23:26:56.893012Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:26:56.893012Z digest=sha256:89f983e31e9fefce3b51d971f57ae6cebb3afc597c7958aaf353978db1ab2dc6

Observation c6f21ed6-1294-43ac-9006-709465b70578 · inbound

AbGen: Evaluating Large Language Models in Ablation Study Design and Evaluation for Scientific Research cites this paper.

AbGen: Evaluating Large Language Models in Ablation Study Design and Evaluation for Scientific Research ML-Bench: Evaluating Large Language Models and Agents for Machine Learning Tasks on Repository-Level Code

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T16:30:31.039397Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:30:31.039397Z digest=sha256:a86f2c06134d3605f90b291974e9001b220976e6b5806bb6e1bdbb89e238e268

Observation 744c909d-1d45-4ef6-aa4d-29e09ab1a074 · inbound

Reinforcement Learning for Machine Learning Engineering Agents cites this paper.

Reinforcement Learning for Machine Learning Engineering Agents ML-Bench: Evaluating Large Language Models and Agents for Machine Learning Tasks on Repository-Level Code

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-05T12:24:02.277658Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:24:02.277658Z digest=sha256:90530ecf5ebfd5b309f252ef3991dda9f3b57fac0eeaa4eb71e7266158a56df7

Observation 174c2508-feb8-4a56-86a8-86a342144568 · inbound

Compiling Large Multi-Modal Requirement Documents into Runnable Software Systems: From an Agentic Test-Driven Perspective cites this paper.

Compiling Large Multi-Modal Requirement Documents into Runnable Software Systems: From an Agentic Test-Driven Perspective ML-Bench: Evaluating Large Language Models and Agents for Machine Learning Tasks on Repository-Level Code

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-02T23:31:16.475615Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:31:16.475615Z digest=sha256:39769bf384416a950a8a667241474108abdb845c02d8c2490ab2593455f31182

Observation 18816a07-66c5-4ba8-a943-1a7a84a17885 · inbound

MLS-Bench: A Holistic and Rigorous Assessment of AI Systems on Building Better AI cites this paper.

MLS-Bench: A Holistic and Rigorous Assessment of AI Systems on Building Better AI ML-Bench: Evaluating Large Language Models and Agents for Machine Learning Tasks on Repository-Level Code

Reference 93

Resolution
verified exact
arxiv_id, observed 2026-05-12T08:21:25.084754Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-12T01:13:35.990078Z digest=sha256:f642eaff9b5c43106359829e565563fe9e471708f0f7fe86bda0709b41423143

Observation cecd6380-a041-4259-aa06-cedd57d3e95b · inbound

MLS-Bench: A Holistic and Rigorous Assessment of AI Systems on Building Better AI cites this paper.

MLS-Bench: A Holistic and Rigorous Assessment of AI Systems on Building Better AI ML-Bench: Evaluating Large Language Models and Agents for Machine Learning Tasks on Repository-Level Code

Reference 95

Resolution
verified exact
arxiv_id, observed 2026-07-01T13:25:46.005528Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-30T23:12:57.154537Z digest=sha256:222625962fc9d87c71c39e3273106c7d235c625e841a51c4109f96a8d7bf5f63

Observation ef75ebca-0ea7-44ac-8c98-c58ed522f8cb · inbound

MLS-Bench: A Holistic and Rigorous Assessment of AI Systems on Building Better AI cites this paper.

MLS-Bench: A Holistic and Rigorous Assessment of AI Systems on Building Better AI ML-Bench: Evaluating Large Language Models and Agents for Machine Learning Tasks on Repository-Level Code

Reference 94

Resolution
unresolved
no resolver link, observed 2026-07-12T17:14:49.310598Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T17:14:49.310598Z digest=sha256:36a0d6af7112d50d05e5973fc85e183a85be7c23d902a6d5e15d0e50e088a21e

Observation 637a3010-188f-4e81-8dbc-ff3578f8ba1d · inbound

ProfiliTable: Profiling-Driven Tabular Data Processing via Agentic Workflows cites this paper.

ProfiliTable: Profiling-Driven Tabular Data Processing via Agentic Workflows ML-Bench: Evaluating Large Language Models and Agents for Machine Learning Tasks on Repository-Level Code

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-13T05:57:22.501265Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-13T05:56:36.312877Z digest=sha256:20479523513b3eeb6c2427fd32505868aa3473bb63d242644acb8198bbb732c9

Observation 26cc257b-1017-4935-887b-1bbe3b2806a1 · inbound

ProfiliTable: Profiling-Driven Tabular Data Processing via Agentic Workflows cites this paper.

ProfiliTable: Profiling-Driven Tabular Data Processing via Agentic Workflows ML-Bench: Evaluating Large Language Models and Agents for Machine Learning Tasks on Repository-Level Code

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-07-01T14:05:46.737909Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-30T22:18:45.189576Z digest=sha256:3471f90aac34c4ba37a3754dc3cf9fe9ea88109e7fbf92197cc7738a79a1326d

Observation 3e1dffac-109e-4453-9e1c-48b967604fc4 · inbound

Matryoshka Agent: Unfolding Sub-Agents for Long-Horizon Machine Learning Engineering cites this paper.

Matryoshka Agent: Unfolding Sub-Agents for Long-Horizon Machine Learning Engineering ML-Bench: Evaluating Large Language Models and Agents for Machine Learning Tasks on Repository-Level Code

Reference 46

Resolution
unresolved
no resolver link, observed 2026-07-31T01:39:48.754483Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T01:39:48.754483Z digest=sha256:377f15ea8cb8ca8f9c68605063e7551fcb010bb69c54615ecfd47ccad9d83ef3

Observation b8946c5c-24df-47d6-8697-164497746df3 · inbound

AgentHPOBench: A Benchmark For Evaluating LLM Agents as Sequential Hyperparameter Optimizers cites this paper.

AgentHPOBench: A Benchmark For Evaluating LLM Agents as Sequential Hyperparameter Optimizers ML-Bench: Evaluating Large Language Models and Agents for Machine Learning Tasks on Repository-Level Code

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-03T03:15:14.160911Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:15:14.160911Z digest=sha256:8bca84fbd85d5605178ead3804ee05e01fb0e58dcd5da4fd6b1e70fb44f2d815