Pith. sign in

Paper Citation Record · LEDGER

ML2B: Benchmarking LLMs on Cross-Lingual ML Pipeline Generation

As of 8 August 2026, this Paper Citation Record lists 45 of 45 outbound references and 0 inbound Pith citation observations for arXiv:2509.22768.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2509.22768 v3

Coverage vector

measured 45 of 45 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-04T14:52:50.638681Z

measured 45 of 45 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

45 of 45 outbound references displayed

  • verified exact1
  • verified fuzzy0
  • unresolved43
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e420ec97-8354-432d-a81f-1547a576a8ea · outbound

This paper cites MEGA : Multilingual evaluation of generative AI.

ML2B: Benchmarking LLMs on Cross-Lingual ML Pipeline Generation MEGA : Multilingual evaluation of generative AI

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-04T14:52:45.456078Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:52:45.456078Z digest=sha256:e9af1d4781c4ec418fb4568a39771ddf2c4992a7be7f547e9752dcd88d459f82

Observation c4338902-d974-4a93-876b-1961c01736a7 · outbound

This paper cites Don’t push the button! exploring data leakage risks in machine learning and transfer learning.

ML2B: Benchmarking LLMs on Cross-Lingual ML Pipeline Generation Don’t push the button! exploring data leakage risks in machine learning and transfer learning

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-04T14:52:45.542088Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:52:45.542088Z digest=sha256:6c56050607b3eaadefe153889a17af097ba04597337407a8c5ce16afd5309c10

Observation fda118ab-8731-4786-9027-99de33468a22 · outbound

This paper cites Do, Yan Xu, and Pascale Fung.

ML2B: Benchmarking LLMs on Cross-Lingual ML Pipeline Generation Do, Yan Xu, and Pascale Fung

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-04T14:52:45.681840Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:52:45.681840Z digest=sha256:6e65ee5567af886872c26977e8e1b43b78814f1f6a3e2c56b043b75e6ff7f19a

Observation dca7e741-24e2-4292-9f0e-f13a89578fa3 · outbound

This paper cites MLE -bench: Evaluating machine learning agents on machine learning engineering.

ML2B: Benchmarking LLMs on Cross-Lingual ML Pipeline Generation MLE -bench: Evaluating machine learning agents on machine learning engineering

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-04T14:52:45.852056Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:52:45.852056Z digest=sha256:a3f5431e5951aa89a87463b850e10fc44871fab3ac15fc494cdcfefd92839658

Observation 10b37b62-bc47-488b-9a0e-fce59ebf2109 · outbound

This paper cites Evaluating Large Language Models Trained on Code.

ML2B: Benchmarking LLMs on Cross-Lingual ML Pipeline Generation Evaluating Large Language Models Trained on Code

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-04T14:52:45.983121Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:52:45.983121Z digest=sha256:669f69b8412d7c9be827ae196fbf5ab62d203df993a6461ed27d98beebc65e14

Observation a45a150c-1add-4c85-82db-133695def190 · outbound

This paper cites R o C ode: A dataset for measuring code intelligence from problem definitions in R omanian.

ML2B: Benchmarking LLMs on Cross-Lingual ML Pipeline Generation R o C ode: A dataset for measuring code intelligence from problem definitions in R omanian

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-04T14:52:46.093206Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:52:46.093206Z digest=sha256:4cda53958cad327c97f4f2c0f9e0799a7cb139de3a0f14e4250d658c7f2f8438

Observation 425917cb-f6cb-41aa-8a6e-6738841ed173 · outbound

This paper cites Abstract Interpretation-Based Data Leakage Static Analysis.

ML2B: Benchmarking LLMs on Cross-Lingual ML Pipeline Generation Abstract Interpretation-Based Data Leakage Static Analysis

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-04T14:52:46.250742Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:52:46.250742Z digest=sha256:669362fecebd4986e48c4ac137ecc66a664f294fb6746ca704b16979a6790556

Observation 39787f77-bd8c-43c9-9bcf-20c0f0673a44 · outbound

This paper cites Code4ml: a large-scale dataset of annotated machine learning code.

ML2B: Benchmarking LLMs on Cross-Lingual ML Pipeline Generation Code4ml: a large-scale dataset of annotated machine learning code

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-04T14:52:46.373728Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:52:46.373728Z digest=sha256:e518d97b3dfffe8623a690ea657b4d650ab5278363db5a6c3e7499d5f0f1e9eb

Observation f4603158-7b91-40af-9387-07492a75bb3d · outbound

This paper cites Neural architecture search: A survey.

ML2B: Benchmarking LLMs on Cross-Lingual ML Pipeline Generation Neural architecture search: A survey

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-04T14:52:46.445465Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:52:46.445465Z digest=sha256:2cdc50f971af37ddaf4b55ad31ad62f7aea8448a81b7c0bec5403ed084afbf95

Observation 06e7f6d6-bf49-4634-9cae-cf4c12d5ebb3 · outbound

This paper cites Efficient and robust automated machine learning.

ML2B: Benchmarking LLMs on Cross-Lingual ML Pipeline Generation Efficient and robust automated machine learning

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-04T14:52:46.648892Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:52:46.648892Z digest=sha256:5411a74c63f41d11d0db9027e8ec78a0833d1ed498ff0698bbcf4b26fd7440a6

Observation cdd74d53-5673-4387-b47e-707c26cb6546 · outbound

This paper cites How Good Are GPT Models at Machine Translation? A Comprehensive Evaluation.

ML2B: Benchmarking LLMs on Cross-Lingual ML Pipeline Generation How Good Are GPT Models at Machine Translation? A Comprehensive Evaluation

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-04T14:52:46.813328Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:52:46.813328Z digest=sha256:5af11bed977ee10e8c6c70ac9eb05600407a88f49c93e46525152b32a72afda3

Observation 64d76b61-f079-48b1-a95a-190502e4a1bc · outbound

This paper cites DA -code: Agent data science code generation benchmark for large language models.

ML2B: Benchmarking LLMs on Cross-Lingual ML Pipeline Generation DA -code: Agent data science code generation benchmark for large language models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-04T14:52:46.924451Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:52:46.924451Z digest=sha256:cfc2e0b2e8ad5559a7b2248955014411bb56f9afed0543a6202aa992a5f5721f

Observation a5f330a1-39c7-4150-82bd-85682835644d · outbound

This paper cites CodeSearchNet Challenge: Evaluating the State of Semantic Code Search.

ML2B: Benchmarking LLMs on Cross-Lingual ML Pipeline Generation CodeSearchNet Challenge: Evaluating the State of Semantic Code Search

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-04T14:52:47.012365Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:52:47.012365Z digest=sha256:0c8b8817870427f0f754edc92a2e04efabd08ecac82a147dbc8b2312698a3663

Observation d35eff5a-5cfa-4009-bc98-c4dad0339ada · outbound

This paper cites AIDE: AI-Driven Exploration in the Space of Code.

ML2B: Benchmarking LLMs on Cross-Lingual ML Pipeline Generation AIDE: AI-Driven Exploration in the Space of Code

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-04T14:52:47.132960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:52:47.132960Z digest=sha256:88cd8d15fce2a3ade514cf02f24aa028dfaca2e23e877e71e7f8b9802e9c8923

Observation 67fa34a3-6f01-4847-9cac-5ac1b1c4e261 · outbound

This paper cites Is ChatGPT A Good Translator? Yes With GPT-4 As The Engine.

ML2B: Benchmarking LLMs on Cross-Lingual ML Pipeline Generation Is ChatGPT A Good Translator? Yes With GPT-4 As The Engine

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-04T14:52:47.365507Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:52:47.365507Z digest=sha256:46a3f9e1785f311208a7e98ecb28109b3c5667d67cfbb352fcbbe68ee5046f3b

Observation 52378ff4-98e4-4e66-95d1-24e301beb581 · outbound

This paper cites Better to ask in english: Cross-lingual evaluation of large language models for healthcare queries.

ML2B: Benchmarking LLMs on Cross-Lingual ML Pipeline Generation Better to ask in english: Cross-lingual evaluation of large language models for healthcare queries

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-04T14:52:47.579659Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:52:47.579659Z digest=sha256:4c5d23e1a4339e9bb2deb5d2967e8d4b041953a6bb6022d01b666aaf6bf29515

Observation ad604c15-4f00-481d-a0ba-c1830dfb1ad1 · outbound

This paper cites Leakage and the reproducibility crisis in machine-learning-based science.

ML2B: Benchmarking LLMs on Cross-Lingual ML Pipeline Generation Leakage and the reproducibility crisis in machine-learning-based science

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-04T14:52:47.761260Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:52:47.761260Z digest=sha256:bbbf1cbdef3bc0843ab77862eed92c6b8052dca844a62ae3470ec29a67188c5d

Observation 91e5a2f4-be41-4d93-9f6c-8bad564ac27e · outbound

This paper cites Ds-1000: a natural and reliable benchmark for data science code generation.

ML2B: Benchmarking LLMs on Cross-Lingual ML Pipeline Generation Ds-1000: a natural and reliable benchmark for data science code generation

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-04T14:52:47.897537Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:52:47.897537Z digest=sha256:952c6368578c8fa12402dd973a51c85356a28b43d36b342d38b84bf8d12a3072

Observation a4eb0139-5aa9-4a31-b3ab-889cb471791b · outbound

This paper cites H2o automl: Scalable automatic machine learning.

ML2B: Benchmarking LLMs on Cross-Lingual ML Pipeline Generation H2o automl: Scalable automatic machine learning

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-04T14:52:47.995256Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:52:47.995256Z digest=sha256:79e4e1c3c7730eec390fd80f90744b848b8c197a55d23fa6db337b44ae6649c3

Observation a4d37812-aefe-45e9-98a7-5c376415acb8 · outbound

This paper cites Bridging the Language Gap: Enhancing Multilingual Prompt-Based Code Generation in LLMs via Zero-Shot Cross-Lingual Transfer.

ML2B: Benchmarking LLMs on Cross-Lingual ML Pipeline Generation Bridging the Language Gap: Enhancing Multilingual Prompt-Based Code Generation in LLMs via Zero-Shot Cross-Lingual Transfer

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-04T14:52:48.068875Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:52:48.068875Z digest=sha256:19019ee8edf13fae8034b246180ec1f7e3ed3779dd80da8db1e37de6eee2a9d6

Observation d4a1793e-a95c-4941-a3a7-2bb8a7f54953 · outbound

This paper cites StarCoder: may the source be with you!.

ML2B: Benchmarking LLMs on Cross-Lingual ML Pipeline Generation StarCoder: may the source be with you!

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-04T14:52:48.179470Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:52:48.179470Z digest=sha256:9457aea1ca72130d727dc05f782f63466614b4051c844edfcca104d633dbd0d0

Observation 6bfeca5d-667a-4960-8f82-747f5eb2a5a2 · outbound

This paper cites DARTS : Differentiable architecture search.

ML2B: Benchmarking LLMs on Cross-Lingual ML Pipeline Generation DARTS : Differentiable architecture search

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-04T14:52:48.326957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:52:48.326957Z digest=sha256:b7ca9622663ddf643a997eb9bb48cd689d33eee061ac9de800333d5ed01cdb68

Observation aee95852-047f-4726-848b-90c4e94a2711 · outbound

This paper cites On Leakage of Code Generation Evaluation Datasets.

ML2B: Benchmarking LLMs on Cross-Lingual ML Pipeline Generation On Leakage of Code Generation Evaluation Datasets

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-04T14:52:48.436986Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:52:48.436986Z digest=sha256:a47d5d72af91cbaa126309cf54c3170fef1506d68375e077c1f4789163459a7e

Observation bc99bfad-adb8-401d-99c3-edd33fc5aef6 · outbound

This paper cites Evaluating programming language confusion.

ML2B: Benchmarking LLMs on Cross-Lingual ML Pipeline Generation Evaluating programming language confusion

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-04T14:52:48.607551Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:52:48.607551Z digest=sha256:8d931b78358fd4186012b6e6531057d5c3f8819b7ceff418e5c2b4896864c523

Observation 9f7c6e13-9f1b-40f0-bdf6-d2d5f498f3a1 · outbound

This paper cites Crosslingual generalization through multitask finetuning.

ML2B: Benchmarking LLMs on Cross-Lingual ML Pipeline Generation Crosslingual generalization through multitask finetuning

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-04T14:52:48.698716Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:52:48.698716Z digest=sha256:8ffaa025926529f766939e35b0c1642ebf6b4c1df574ebd86ed6b558e0b1d78c

Observation f0c9f926-3437-4da0-a311-d138b1341510 · outbound

This paper cites Olson, Nathan Bartley, Ryan J.

ML2B: Benchmarking LLMs on Cross-Lingual ML Pipeline Generation Olson, Nathan Bartley, Ryan J

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-04T14:52:48.870432Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:52:48.870432Z digest=sha256:22318a2b1142d9af7bd89bdf3edad40557b35ab5f66ae8ebdf5880fcab730795

Observation 243d93b9-8f3d-4874-9e8c-97848239dccd · outbound

This paper cites Dscodebench: A realistic benchmark for data science code generation.

ML2B: Benchmarking LLMs on Cross-Lingual ML Pipeline Generation Dscodebench: A realistic benchmark for data science code generation

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-04T14:52:49.054155Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:52:49.054155Z digest=sha256:8e51b5fe24a52b3ad99f444a65d907a1be860df271641d4fd188ee91e902d0b8

Observation 59e74f38-68e8-484c-b5ff-0a7d8079fcc3 · outbound

This paper cites Efficient neural architecture search via parameters sharing.

ML2B: Benchmarking LLMs on Cross-Lingual ML Pipeline Generation Efficient neural architecture search via parameters sharing

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-04T14:52:49.202574Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:52:49.202574Z digest=sha256:c86f813c3b4c922cb909506fe3af2ef95b0f3f16567300e8a735ded3f3502d8c

Observation 936bb467-1a9a-4eeb-851c-286e5ddfbb1a · outbound

This paper cites Meta kaggle code, 2023.

ML2B: Benchmarking LLMs on Cross-Lingual ML Pipeline Generation Meta kaggle code, 2023

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-04T14:52:49.284097Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:52:49.284097Z digest=sha256:b53a9560105e1a897351308579f73808ca5716630982a2c26d5ece64737b7e49

Observation 4a33406c-9442-49cb-8bf8-8d48daa9312e · outbound

This paper cites m H uman E val - a multilingual benchmark to evaluate large language models for code generation.

ML2B: Benchmarking LLMs on Cross-Lingual ML Pipeline Generation m H uman E val - a multilingual benchmark to evaluate large language models for code generation

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-04T14:52:49.398954Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:52:49.398954Z digest=sha256:d8406fb35b06f15eb5f08d1141ae703c0aa424f6ec8a19be6aa43854ccf67cbf

Observation 3f05b8c5-4f13-46ec-b7c1-fe72e521d0d3 · outbound

This paper cites Do GPTs Produce Less Literal Translations?.

ML2B: Benchmarking LLMs on Cross-Lingual ML Pipeline Generation Do GPTs Produce Less Literal Translations?

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-04T14:52:49.483019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:52:49.483019Z digest=sha256:5601b4cd086816ae6f278499343312586b1396aaf1ad93a379d033047343d2f2

Observation 65082097-3b06-4b20-8026-2dbca25f9768 · outbound

This paper cites Code Llama: Open Foundation Models for Code.

ML2B: Benchmarking LLMs on Cross-Lingual ML Pipeline Generation Code Llama: Open Foundation Models for Code

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-04T14:52:49.569202Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:52:49.569202Z digest=sha256:b38e457d8d916d8b119e2e7b0106b68b123d98dfb43bece64d9d65ecb7d65d7b

Observation c6db1d1c-91fa-49b1-8832-53d5913799c8 · outbound

This paper cites Sasse, E.

ML2B: Benchmarking LLMs on Cross-Lingual ML Pipeline Generation Sasse, E

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-04T14:52:49.671243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:52:49.671243Z digest=sha256:614a102561b4fdf79b57560f2ff572ab26335604e74ee0ab2b47de7f762598e3

Observation d57d06e2-b08b-40d7-a4fd-2a0050e7339b · outbound

This paper cites Biocoder: a benchmark for bioinformatics code generation with large language models.

ML2B: Benchmarking LLMs on Cross-Lingual ML Pipeline Generation Biocoder: a benchmark for bioinformatics code generation with large language models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-04T14:52:49.760663Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:52:49.760663Z digest=sha256:aa2598a27505d4da42f71bfe69f1a5ae25326050b007633a70c0ab45e7ed593c

Observation c6c3e951-9187-446f-8974-1a3d1d2f5592 · outbound

This paper cites SciCode: A Research Coding Benchmark Curated by Scientists.

ML2B: Benchmarking LLMs on Cross-Lingual ML Pipeline Generation SciCode: A Research Coding Benchmark Curated by Scientists

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-04T14:52:49.846361Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:52:49.846361Z digest=sha256:99ac6870f400f78d991abbaf988181a80eba41f3480d3a4c6ece13599e9c5741

Observation da2f3360-d57d-4ea2-b2bf-ee6767351e45 · outbound

This paper cites LightAutoML: AutoML Solution for a Large Financial Services Ecosystem.

ML2B: Benchmarking LLMs on Cross-Lingual ML Pipeline Generation LightAutoML: AutoML Solution for a Large Financial Services Ecosystem

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-04T14:52:49.910664Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:52:49.910664Z digest=sha256:c3974ecf6a2bf29c2e852702c6393013df58028c7b242afdd74daa55d1385146

Observation 326cfc55-677f-4256-a3a3-b34b38dc4559 · outbound

This paper cites OpenHands: An Open Platform for AI Software Developers as Generalist Agents.

ML2B: Benchmarking LLMs on Cross-Lingual ML Pipeline Generation OpenHands: An Open Platform for AI Software Developers as Generalist Agents

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-04T14:52:49.994121Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:52:49.994121Z digest=sha256:4e0efb6bb16c75d610963dbb947318ff6fde1c37afbb32060dda75f0c9109094

Observation 0373be4a-3c79-484c-9b03-7ede74aba817 · outbound

This paper cites MCoNaLa: A Benchmark for Code Generation from Multiple Natural Languages.

ML2B: Benchmarking LLMs on Cross-Lingual ML Pipeline Generation MCoNaLa: A Benchmark for Code Generation from Multiple Natural Languages

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-04T14:52:50.080171Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:52:50.080171Z digest=sha256:fb76ad6ce966a7b5cba097d1e43a190c6d1995fa836c017e8f8164834048a10a

Observation 3606950e-8203-4733-bc79-7f5e90003c6f · outbound

This paper cites Data Leakage in Notebooks: Static Detection and Better Processes.

ML2B: Benchmarking LLMs on Cross-Lingual ML Pipeline Generation Data Leakage in Notebooks: Static Detection and Better Processes

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-04T14:52:50.160601Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:52:50.160601Z digest=sha256:bac943d9a11145f984ec4c9c3fbd68bd62197ce7d2e6c10f0be86f7685c3664a

Observation 454558da-8267-4e9d-ada2-7797949bcc52 · outbound

This paper cites LessLeak-Bench: A First Investigation of Data Leakage in LLMs Across 83 Software Engineering Benchmarks.

ML2B: Benchmarking LLMs on Cross-Lingual ML Pipeline Generation LessLeak-Bench: A First Investigation of Data Leakage in LLMs Across 83 Software Engineering Benchmarks

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-04T14:52:50.211290Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:52:50.211290Z digest=sha256:da5583a7ad57173cb7eceabdd982b8b5a496080b6594451030f421656108188d

Observation 47f62696-d9f4-4954-9b87-f220d173af03 · outbound

This paper cites an unresolved cited work.

ML2B: Benchmarking LLMs on Cross-Lingual ML Pipeline Generation Unresolved cited work

Reference 41

Resolution
verified exact
doi, observed 2026-08-04T14:53:21.126268Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-04T14:52:50.314164Z digest=sha256:83e327fd2daa0a4d837b699601ebae8712b6765ce02462528987c3edf5c36f2b

Observation f4f7e8e6-5ff7-4d32-ba10-0a1256f3afdd · outbound

This paper cites write newline.

ML2B: Benchmarking LLMs on Cross-Lingual ML Pipeline Generation write newline

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-04T14:52:50.383648Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:52:50.383648Z digest=sha256:8e3f31a7daba15bbe98ca1e813048841f0b572fd9a8f0756489dec185c38b026

Observation fcd9988c-2c0a-4aba-9cbc-28b12ce13d36 · outbound

This paper cites @esa (Ref.

ML2B: Benchmarking LLMs on Cross-Lingual ML Pipeline Generation @esa (Ref

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-04T14:52:50.475733Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:52:50.475733Z digest=sha256:d637b8891641f302ee43a8fc799e3f34d8b95c38252dcecad8513694600a2bd7

Observation 3977f7a0-e450-451e-aab6-3c1c6756bf8f · outbound

This paper cites an unresolved cited work.

ML2B: Benchmarking LLMs on Cross-Lingual ML Pipeline Generation Unresolved cited work

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-04T14:52:50.558496Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:52:50.558496Z digest=sha256:df25b20442f561e11bf397dcab8236e8f2067217d7207b1446ddfa08ca13ed4e

Observation 1570824f-2e9a-4fab-9e0c-7eb165238048 · outbound

This paper cites 9rD= <8rrr, 5jTXyyy5V ꬦM`0HΞ=[X5Z ꬀[t ߿XܹSԮ];i qsv Ÿu9 >|Xo =z(00PΝ n: UVzǝ]2Z 4 ڸqV 6c4j(] rR g ǎiJHH=--M + P u5k, iӦ . egg+::Z /VLL ] 3Fwy P uӧ-` (WWW+.

ML2B: Benchmarking LLMs on Cross-Lingual ML Pipeline Generation 9rD= <8rrr, 5jTXyyy5V ꬦM`0HΞ=[X5Z ꬀[t ߿XܹSԮ];i qsv Ÿu9 >|Xo =z(00PΝ n: UVzǝ]2Z 4 ڸqV 6c4j(] rR g ǎiJHH=--M + P u5k, iӦ . egg+::Z /VLL ] 3Fwy P uӧ-` (WWW+

Reference 45

Resolution
malformed identifier
no resolver link, observed 2026-08-04T14:52:50.638681Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:52:50.638681Z digest=sha256:77f297d0c66a41cc8acff3b5d7574f65a298c828df7d852254ec673738b1e4a9

Pith citing papers

No inbound Pith citation observations are available.