Pith. sign in

Paper Citation Record · LEDGER

ML2B: Benchmarking LLMs on Cross-Lingual ML Pipeline Generation

As of 18 August 2026, this Paper Citation Record lists 45 of 45 outbound references and 0 inbound Pith citation observations for arXiv:2509.22768.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2509.22768 v3

Coverage vector

measured 45 of 45 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-04T14:52:50.638681Z

measured 45 of 45 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

45 of 45 outbound references displayed

  • verified exact1
  • verified fuzzy0
  • unresolved43
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e420ec97-8354-432d-a81f-1547a576a8ea · outbound

This paper cites MEGA : Multilingual evaluation of generative AI.

ML2B: Benchmarking LLMs on Cross-Lingual ML Pipeline Generation MEGA : Multilingual evaluation of generative AI

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-04T14:52:45.456078Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:52:45.456078Z digest=sha256:51f83a289655d10f2dfad4d3e8884d3b3e190bf646eb9026c2c85acb29978bb9

Observation c4338902-d974-4a93-876b-1961c01736a7 · outbound

This paper cites Don’t push the button! exploring data leakage risks in machine learning and transfer learning.

ML2B: Benchmarking LLMs on Cross-Lingual ML Pipeline Generation Don’t push the button! exploring data leakage risks in machine learning and transfer learning

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-04T14:52:45.542088Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:52:45.542088Z digest=sha256:370c0e4c78ba9f07eb5365c098267662d6ff7a47c37b453ecf5b3cfd4699d0a1

Observation fda118ab-8731-4786-9027-99de33468a22 · outbound

This paper cites Do, Yan Xu, and Pascale Fung.

ML2B: Benchmarking LLMs on Cross-Lingual ML Pipeline Generation Do, Yan Xu, and Pascale Fung

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-04T14:52:45.681840Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:52:45.681840Z digest=sha256:e4aee06f634f701e4d59ce7a2c2b3fa1f78346cd26822d58c7999e135a38a176

Observation dca7e741-24e2-4292-9f0e-f13a89578fa3 · outbound

This paper cites MLE -bench: Evaluating machine learning agents on machine learning engineering.

ML2B: Benchmarking LLMs on Cross-Lingual ML Pipeline Generation MLE -bench: Evaluating machine learning agents on machine learning engineering

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-04T14:52:45.852056Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:52:45.852056Z digest=sha256:d3958a068348055b4db58f1d055ebe686c89543eb9c2c9cf2f17ba46754e46bb

Observation 10b37b62-bc47-488b-9a0e-fce59ebf2109 · outbound

This paper cites Evaluating Large Language Models Trained on Code.

ML2B: Benchmarking LLMs on Cross-Lingual ML Pipeline Generation Evaluating Large Language Models Trained on Code

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-04T14:52:45.983121Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:52:45.983121Z digest=sha256:4d2b5b32e77b745b2ce7cc51a605c1a02204a9a21d85e8324aa64b50b0ceacef

Observation a45a150c-1add-4c85-82db-133695def190 · outbound

This paper cites R o C ode: A dataset for measuring code intelligence from problem definitions in R omanian.

ML2B: Benchmarking LLMs on Cross-Lingual ML Pipeline Generation R o C ode: A dataset for measuring code intelligence from problem definitions in R omanian

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-04T14:52:46.093206Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:52:46.093206Z digest=sha256:0977374e2dcc292e86c59313aa9c4aa8788463a9dafbe463b3c617e102a0428a

Observation 425917cb-f6cb-41aa-8a6e-6738841ed173 · outbound

This paper cites Abstract Interpretation-Based Data Leakage Static Analysis.

ML2B: Benchmarking LLMs on Cross-Lingual ML Pipeline Generation Abstract Interpretation-Based Data Leakage Static Analysis

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-04T14:52:46.250742Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:52:46.250742Z digest=sha256:7d799af986e8471de97b8420ab375859beb84c82d23aa31d565de5a4444409eb

Observation 39787f77-bd8c-43c9-9bcf-20c0f0673a44 · outbound

This paper cites Code4ml: a large-scale dataset of annotated machine learning code.

ML2B: Benchmarking LLMs on Cross-Lingual ML Pipeline Generation Code4ml: a large-scale dataset of annotated machine learning code

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-04T14:52:46.373728Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:52:46.373728Z digest=sha256:eab84210fb900f9bd4e5d92ea85c692db12827398d5858dacbbb5f2ecb2ca6d3

Observation f4603158-7b91-40af-9387-07492a75bb3d · outbound

This paper cites Neural architecture search: A survey.

ML2B: Benchmarking LLMs on Cross-Lingual ML Pipeline Generation Neural architecture search: A survey

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-04T14:52:46.445465Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:52:46.445465Z digest=sha256:b1e0cfa4883ee4c607cd19baf658529c0e2e326f8b9982a2764cc1b123208def

Observation 06e7f6d6-bf49-4634-9cae-cf4c12d5ebb3 · outbound

This paper cites Efficient and robust automated machine learning.

ML2B: Benchmarking LLMs on Cross-Lingual ML Pipeline Generation Efficient and robust automated machine learning

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-04T14:52:46.648892Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:52:46.648892Z digest=sha256:5d36b432f93aba564407704f2edec5b72185122ede8e98bdd04aa6bb5b8b7cc8

Observation cdd74d53-5673-4387-b47e-707c26cb6546 · outbound

This paper cites How Good Are GPT Models at Machine Translation? A Comprehensive Evaluation.

ML2B: Benchmarking LLMs on Cross-Lingual ML Pipeline Generation How Good Are GPT Models at Machine Translation? A Comprehensive Evaluation

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-04T14:52:46.813328Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:52:46.813328Z digest=sha256:63e94433ecf2c945b7afce376d09d4a98dc9de6f0bdd67596b98b95aef2dc928

Observation 64d76b61-f079-48b1-a95a-190502e4a1bc · outbound

This paper cites DA -code: Agent data science code generation benchmark for large language models.

ML2B: Benchmarking LLMs on Cross-Lingual ML Pipeline Generation DA -code: Agent data science code generation benchmark for large language models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-04T14:52:46.924451Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:52:46.924451Z digest=sha256:ff3afa030b26301ceb76f6243e112db81e6e2412d211a12e360b905f82395d56

Observation a5f330a1-39c7-4150-82bd-85682835644d · outbound

This paper cites CodeSearchNet Challenge: Evaluating the State of Semantic Code Search.

ML2B: Benchmarking LLMs on Cross-Lingual ML Pipeline Generation CodeSearchNet Challenge: Evaluating the State of Semantic Code Search

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-04T14:52:47.012365Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:52:47.012365Z digest=sha256:d271f595942aea0f4cbbe05cd6818f69470bf7dc6bf226fb88f70d33b525d654

Observation d35eff5a-5cfa-4009-bc98-c4dad0339ada · outbound

This paper cites AIDE: AI-Driven Exploration in the Space of Code.

ML2B: Benchmarking LLMs on Cross-Lingual ML Pipeline Generation AIDE: AI-Driven Exploration in the Space of Code

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-04T14:52:47.132960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:52:47.132960Z digest=sha256:268eef15c628200f29587d64b4a3aab611303b73199193c39e18af42743b5755

Observation 67fa34a3-6f01-4847-9cac-5ac1b1c4e261 · outbound

This paper cites Is ChatGPT A Good Translator? Yes With GPT-4 As The Engine.

ML2B: Benchmarking LLMs on Cross-Lingual ML Pipeline Generation Is ChatGPT A Good Translator? Yes With GPT-4 As The Engine

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-04T14:52:47.365507Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:52:47.365507Z digest=sha256:a132ef29169586513b43a3df9982cc7d1bde40aabc7566f568fdb739e66a5204

Observation 52378ff4-98e4-4e66-95d1-24e301beb581 · outbound

This paper cites Better to ask in english: Cross-lingual evaluation of large language models for healthcare queries.

ML2B: Benchmarking LLMs on Cross-Lingual ML Pipeline Generation Better to ask in english: Cross-lingual evaluation of large language models for healthcare queries

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-04T14:52:47.579659Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:52:47.579659Z digest=sha256:6adf3883e8a3e96a99f726a04aa800df5893bda5bb2314320ebaeb78bff5c317

Observation ad604c15-4f00-481d-a0ba-c1830dfb1ad1 · outbound

This paper cites Leakage and the reproducibility crisis in machine-learning-based science.

ML2B: Benchmarking LLMs on Cross-Lingual ML Pipeline Generation Leakage and the reproducibility crisis in machine-learning-based science

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-04T14:52:47.761260Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:52:47.761260Z digest=sha256:52044a88693744066b65cfb9dde748581a61da8cb5cbad1fa335417585901dd7

Observation 91e5a2f4-be41-4d93-9f6c-8bad564ac27e · outbound

This paper cites Ds-1000: a natural and reliable benchmark for data science code generation.

ML2B: Benchmarking LLMs on Cross-Lingual ML Pipeline Generation Ds-1000: a natural and reliable benchmark for data science code generation

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-04T14:52:47.897537Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:52:47.897537Z digest=sha256:6d389e5d4c6a0f53189965af0e96bddf9a25f128524f30b2516362094b6fe040

Observation a4eb0139-5aa9-4a31-b3ab-889cb471791b · outbound

This paper cites H2o automl: Scalable automatic machine learning.

ML2B: Benchmarking LLMs on Cross-Lingual ML Pipeline Generation H2o automl: Scalable automatic machine learning

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-04T14:52:47.995256Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:52:47.995256Z digest=sha256:d6a5d5079f4ce3d5b168f1445c30798d338d837476e4e7bd14f051b249181752

Observation a4d37812-aefe-45e9-98a7-5c376415acb8 · outbound

This paper cites Bridging the Language Gap: Enhancing Multilingual Prompt-Based Code Generation in LLMs via Zero-Shot Cross-Lingual Transfer.

ML2B: Benchmarking LLMs on Cross-Lingual ML Pipeline Generation Bridging the Language Gap: Enhancing Multilingual Prompt-Based Code Generation in LLMs via Zero-Shot Cross-Lingual Transfer

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-04T14:52:48.068875Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:52:48.068875Z digest=sha256:4a32b42b4ba26fc1e56409909a4c4489d4e5aec29dae5aa35702bd4768b6014d

Observation d4a1793e-a95c-4941-a3a7-2bb8a7f54953 · outbound

This paper cites StarCoder: may the source be with you!.

ML2B: Benchmarking LLMs on Cross-Lingual ML Pipeline Generation StarCoder: may the source be with you!

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-04T14:52:48.179470Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:52:48.179470Z digest=sha256:493c440e854d545e4db2da194e47bc4337971e493f5eacbe410bbd7c8d7e27f8

Observation 6bfeca5d-667a-4960-8f82-747f5eb2a5a2 · outbound

This paper cites DARTS : Differentiable architecture search.

ML2B: Benchmarking LLMs on Cross-Lingual ML Pipeline Generation DARTS : Differentiable architecture search

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-04T14:52:48.326957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:52:48.326957Z digest=sha256:c0bf99600fc18e48dbe3a6756c613bd869037cda2dfb65e7a7028ab18dc2756a

Observation aee95852-047f-4726-848b-90c4e94a2711 · outbound

This paper cites On Leakage of Code Generation Evaluation Datasets.

ML2B: Benchmarking LLMs on Cross-Lingual ML Pipeline Generation On Leakage of Code Generation Evaluation Datasets

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-04T14:52:48.436986Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:52:48.436986Z digest=sha256:cdd98c62ab2c638ec35a8e6ca3c6fc4ab9340f300583b1c66f34cb4bcd1e9c7c

Observation bc99bfad-adb8-401d-99c3-edd33fc5aef6 · outbound

This paper cites Evaluating programming language confusion.

ML2B: Benchmarking LLMs on Cross-Lingual ML Pipeline Generation Evaluating programming language confusion

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-04T14:52:48.607551Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:52:48.607551Z digest=sha256:69f16f9fa3889edc9a400c26b60b670ceb68bf0736a040218a1cc8e850c3b5d8

Observation 9f7c6e13-9f1b-40f0-bdf6-d2d5f498f3a1 · outbound

This paper cites Crosslingual generalization through multitask finetuning.

ML2B: Benchmarking LLMs on Cross-Lingual ML Pipeline Generation Crosslingual generalization through multitask finetuning

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-04T14:52:48.698716Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:52:48.698716Z digest=sha256:b7d9f8db753701f3fac16f757089641b332618994af744129cf0d245f2660747

Observation f0c9f926-3437-4da0-a311-d138b1341510 · outbound

This paper cites Olson, Nathan Bartley, Ryan J.

ML2B: Benchmarking LLMs on Cross-Lingual ML Pipeline Generation Olson, Nathan Bartley, Ryan J

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-04T14:52:48.870432Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:52:48.870432Z digest=sha256:2a7ae90d9874468ecbdae6dc3020c46531ec1a568c11aea41328b34acb989480

Observation 243d93b9-8f3d-4874-9e8c-97848239dccd · outbound

This paper cites Dscodebench: A realistic benchmark for data science code generation.

ML2B: Benchmarking LLMs on Cross-Lingual ML Pipeline Generation Dscodebench: A realistic benchmark for data science code generation

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-04T14:52:49.054155Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:52:49.054155Z digest=sha256:606112830cbf1d55d7a0a75dcc5eea61a2bec3bd3b44092de9a6a8bb7adbc2d5

Observation 59e74f38-68e8-484c-b5ff-0a7d8079fcc3 · outbound

This paper cites Efficient neural architecture search via parameters sharing.

ML2B: Benchmarking LLMs on Cross-Lingual ML Pipeline Generation Efficient neural architecture search via parameters sharing

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-04T14:52:49.202574Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:52:49.202574Z digest=sha256:cfd702c3708b8716647be4b1da2272c08a4723a00882a884d65e74612e6ee8bd

Observation 936bb467-1a9a-4eeb-851c-286e5ddfbb1a · outbound

This paper cites Meta kaggle code, 2023.

ML2B: Benchmarking LLMs on Cross-Lingual ML Pipeline Generation Meta kaggle code, 2023

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-04T14:52:49.284097Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:52:49.284097Z digest=sha256:a6956a96b834567cb85202aee2cfc2bfe4936027dc831238610725c80b6d6fd0

Observation 4a33406c-9442-49cb-8bf8-8d48daa9312e · outbound

This paper cites m H uman E val - a multilingual benchmark to evaluate large language models for code generation.

ML2B: Benchmarking LLMs on Cross-Lingual ML Pipeline Generation m H uman E val - a multilingual benchmark to evaluate large language models for code generation

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-04T14:52:49.398954Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:52:49.398954Z digest=sha256:23a7eb07559dae9b062c7997f1e7ad3a18f827ab03928b9598cfbfe339580403

Observation 3f05b8c5-4f13-46ec-b7c1-fe72e521d0d3 · outbound

This paper cites Do GPTs Produce Less Literal Translations?.

ML2B: Benchmarking LLMs on Cross-Lingual ML Pipeline Generation Do GPTs Produce Less Literal Translations?

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-04T14:52:49.483019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:52:49.483019Z digest=sha256:9ed0bcc66e8c032329e6dbfcb120d82d63882f843d8d6f4a8c0da179547cc3c6

Observation 65082097-3b06-4b20-8026-2dbca25f9768 · outbound

This paper cites Code Llama: Open Foundation Models for Code.

ML2B: Benchmarking LLMs on Cross-Lingual ML Pipeline Generation Code Llama: Open Foundation Models for Code

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-04T14:52:49.569202Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:52:49.569202Z digest=sha256:f3293fa57e6262c51a726ad3daba0bc68881da03c1397a1a4f42a84ef9df3367

Observation c6db1d1c-91fa-49b1-8832-53d5913799c8 · outbound

This paper cites Sasse, E.

ML2B: Benchmarking LLMs on Cross-Lingual ML Pipeline Generation Sasse, E

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-04T14:52:49.671243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:52:49.671243Z digest=sha256:466fadc73a564e66c869bf192cefc16624275e292c3e444880710b3d27cbd6d1

Observation d57d06e2-b08b-40d7-a4fd-2a0050e7339b · outbound

This paper cites Biocoder: a benchmark for bioinformatics code generation with large language models.

ML2B: Benchmarking LLMs on Cross-Lingual ML Pipeline Generation Biocoder: a benchmark for bioinformatics code generation with large language models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-04T14:52:49.760663Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:52:49.760663Z digest=sha256:63c4e24591682034d99455f9c50ea8968f22c75957d265ef6c4ab9ec59af6e79

Observation c6c3e951-9187-446f-8974-1a3d1d2f5592 · outbound

This paper cites SciCode: A Research Coding Benchmark Curated by Scientists.

ML2B: Benchmarking LLMs on Cross-Lingual ML Pipeline Generation SciCode: A Research Coding Benchmark Curated by Scientists

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-04T14:52:49.846361Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:52:49.846361Z digest=sha256:a41b3346b1adb34eec1659a90cf0a4aa20c7a566c86633643e45dc9293766e43

Observation da2f3360-d57d-4ea2-b2bf-ee6767351e45 · outbound

This paper cites LightAutoML: AutoML Solution for a Large Financial Services Ecosystem.

ML2B: Benchmarking LLMs on Cross-Lingual ML Pipeline Generation LightAutoML: AutoML Solution for a Large Financial Services Ecosystem

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-04T14:52:49.910664Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:52:49.910664Z digest=sha256:3259be339748d296f5dfab19caf096e3e450e2e837bb018f49482c983ca7d800

Observation 326cfc55-677f-4256-a3a3-b34b38dc4559 · outbound

This paper cites OpenHands: An Open Platform for AI Software Developers as Generalist Agents.

ML2B: Benchmarking LLMs on Cross-Lingual ML Pipeline Generation OpenHands: An Open Platform for AI Software Developers as Generalist Agents

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-04T14:52:49.994121Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:52:49.994121Z digest=sha256:244ccf81ee1180f004ef4bc34dcaca0c087e0d17c8c7ca2cea2e6711e68cb982

Observation 0373be4a-3c79-484c-9b03-7ede74aba817 · outbound

This paper cites MCoNaLa: A Benchmark for Code Generation from Multiple Natural Languages.

ML2B: Benchmarking LLMs on Cross-Lingual ML Pipeline Generation MCoNaLa: A Benchmark for Code Generation from Multiple Natural Languages

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-04T14:52:50.080171Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:52:50.080171Z digest=sha256:1af4bb96a014a46e9681882cadfb19bbfc6f28cb9c9aa0ba65112422f1a3b580

Observation 3606950e-8203-4733-bc79-7f5e90003c6f · outbound

This paper cites Data Leakage in Notebooks: Static Detection and Better Processes.

ML2B: Benchmarking LLMs on Cross-Lingual ML Pipeline Generation Data Leakage in Notebooks: Static Detection and Better Processes

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-04T14:52:50.160601Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:52:50.160601Z digest=sha256:8df20878bd0eae70f21eed7587aa96c8d7a1718b1b80a4d976d7cfdf75a3e173

Observation 454558da-8267-4e9d-ada2-7797949bcc52 · outbound

This paper cites LessLeak-Bench: A First Investigation of Data Leakage in LLMs Across 83 Software Engineering Benchmarks.

ML2B: Benchmarking LLMs on Cross-Lingual ML Pipeline Generation LessLeak-Bench: A First Investigation of Data Leakage in LLMs Across 83 Software Engineering Benchmarks

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-04T14:52:50.211290Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:52:50.211290Z digest=sha256:31ccfda23495c5bef660412b1c4b3d6a31b03f200dc8923e6ec14df527157a2e

Observation 47f62696-d9f4-4954-9b87-f220d173af03 · outbound

This paper cites an unresolved cited work.

ML2B: Benchmarking LLMs on Cross-Lingual ML Pipeline Generation Unresolved cited work

Reference 41

Resolution
verified exact
doi, observed 2026-08-04T14:53:21.126268Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-08-04T14:52:50.314164Z digest=sha256:d508848fb77657977e353fda00940c71714c3e76a34b60bcd4c4b6e1f2bf6169

Observation f4f7e8e6-5ff7-4d32-ba10-0a1256f3afdd · outbound

This paper cites write newline.

ML2B: Benchmarking LLMs on Cross-Lingual ML Pipeline Generation write newline

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-04T14:52:50.383648Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:52:50.383648Z digest=sha256:8c1d5f16b33c9429e731f180713e3474bd3aa75bda57f418996936fcade8d4f6

Observation fcd9988c-2c0a-4aba-9cbc-28b12ce13d36 · outbound

This paper cites @esa (Ref.

ML2B: Benchmarking LLMs on Cross-Lingual ML Pipeline Generation @esa (Ref

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-04T14:52:50.475733Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:52:50.475733Z digest=sha256:cd84a673d1067b9bec12c639e5a7497d312f0491cabf5ecdd07425b88a010570

Observation 3977f7a0-e450-451e-aab6-3c1c6756bf8f · outbound

This paper cites an unresolved cited work.

ML2B: Benchmarking LLMs on Cross-Lingual ML Pipeline Generation Unresolved cited work

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-04T14:52:50.558496Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:52:50.558496Z digest=sha256:7278b83bb39e12566ad87d3a55df50b6fa2221290603c81ba95aa3f82f54752e

Observation 1570824f-2e9a-4fab-9e0c-7eb165238048 · outbound

This paper cites 9rD= <8rrr, 5jTXyyy5V ꬦM`0HΞ=[X5Z ꬀[t ߿XܹSԮ];i qsv Ÿu9 >|Xo =z(00PΝ n: UVzǝ]2Z 4 ڸqV 6c4j(] rR g ǎiJHH=--M + P u5k, iӦ . egg+::Z /VLL ] 3Fwy P uӧ-` (WWW+.

ML2B: Benchmarking LLMs on Cross-Lingual ML Pipeline Generation 9rD= <8rrr, 5jTXyyy5V ꬦM`0HΞ=[X5Z ꬀[t ߿XܹSԮ];i qsv Ÿu9 >|Xo =z(00PΝ n: UVzǝ]2Z 4 ڸqV 6c4j(] rR g ǎiJHH=--M + P u5k, iӦ . egg+::Z /VLL ] 3Fwy P uӧ-` (WWW+

Reference 45

Resolution
malformed identifier
no resolver link, observed 2026-08-04T14:52:50.638681Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:52:50.638681Z digest=sha256:af62a261a8ce4694ee284698dd377eca4e2805efebb6c1b3f715aefa1523959c

Pith citing papers

No inbound Pith citation observations are available.