Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-05T12:24:02.385113Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 62 of 62 outbound references and 7 inbound Pith citation observations for arXiv:2509.01684.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-05T12:24:02.385113Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-07-31T01:39:49.218941Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-01T13:25:45.948001Z
62 of 62 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 4bf46cd9-754d-4002-aa2b-02ba6e56517d · outbound
Reinforcement Learning for Machine Learning Engineering Agents SWE-bench: Can Language Models Resolve Real-World GitHub Issues?
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 144bf1fe-35c3-4062-83d3-ac95553d6215 · outbound
Reinforcement Learning for Machine Learning Engineering Agents Swe-agent: Agent-computer interfaces enable automated software engineering
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 594ff46b-6c87-4ac0-8d3a-9f4b2bdd3cc8 · outbound
Reinforcement Learning for Machine Learning Engineering Agents ResearchAgent: Iterative Research Idea Generation over Scientific Literature with Large Language Models
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 47afa8af-5f77-41d0-8d91-daac13edd68a · outbound
Reinforcement Learning for Machine Learning Engineering Agents MLAgentBench: Evaluating Language Agents on Machine Learning Experimentation
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9966b5d1-224c-4fea-9a94-78f61bc54f1f · outbound
Reinforcement Learning for Machine Learning Engineering Agents MLE-bench: Evaluating Machine Learning Agents on Machine Learning Engineering
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9576ef01-37db-4382-b092-150327a1b092 · outbound
Reinforcement Learning for Machine Learning Engineering Agents Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c020786e-7f1e-4d7b-8ad6-9f2aee623858 · outbound
Reinforcement Learning for Machine Learning Engineering Agents Inference Scaling Laws: An Empirical Analysis of Compute-Optimal Inference for Problem-Solving with Language Models
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1f2ffa2e-60cf-4c26-9e4a-77610ae19182 · outbound
Reinforcement Learning for Machine Learning Engineering Agents Reinforcement learning: An introduction, volume 1
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0e0cd02a-24e6-4518-85a8-972d9b855e2a · outbound
Reinforcement Learning for Machine Learning Engineering Agents Qwen2.5 Technical Report
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 901e5f44-53a8-4c0b-96f4-3e9d5675e960 · outbound
Reinforcement Learning for Machine Learning Engineering Agents Markov decision processes: discrete stochastic dynamic programming
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9383b62b-6c1f-4e93-858a-e3c44dd157c0 · outbound
Reinforcement Learning for Machine Learning Engineering Agents Simple statistical gradient-following algorithms for connectionist reinforce- ment learning
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d2c05991-2083-4e0e-b769-9c67a13b0e83 · outbound
Reinforcement Learning for Machine Learning Engineering Agents Proximal Policy Optimization Algorithms
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ce8144c7-d782-4342-ba92-85ec442ee0c6 · outbound
Reinforcement Learning for Machine Learning Engineering Agents Efficient exploration in reinforcement learning
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation df3c1527-bd08-4cde-9319-751b1c54317c · outbound
Reinforcement Learning for Machine Learning Engineering Agents On the sample complexity of reinforcement learning
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 23b6d3c8-25cc-4ef7-bd32-6497a846884b · outbound
Reinforcement Learning for Machine Learning Engineering Agents Rllib: Abstractions for distributed reinforcement learning
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 3e505bb6-a99a-4cf9-9558-1c793b2d0777 · outbound
Reinforcement Learning for Machine Learning Engineering Agents Acme: A Research Framework for Distributed Reinforcement Learning
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ebc7e94f-e431-4133-bf0c-d2d9da6c37a8 · outbound
Reinforcement Learning for Machine Learning Engineering Agents HybridFlow: A Flexible and Efficient RLHF Framework
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c69dd7fe-66c9-4769-9702-7e26521eac41 · outbound
Reinforcement Learning for Machine Learning Engineering Agents Openhands: An open platform for ai software developers as generalist agents
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 87eb226d-023f-4db7-a7d6-41b1764864f8 · outbound
Reinforcement Learning for Machine Learning Engineering Agents LangChain, October 2022
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f032b7bd-e628-47bf-871d-e3551d297aa3 · outbound
Reinforcement Learning for Machine Learning Engineering Agents Unresolved cited work
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 744c909d-1d45-4ef6-aa4d-29e09ab1a074 · outbound
Reinforcement Learning for Machine Learning Engineering Agents ML-Bench: Evaluating Large Language Models and Agents for Machine Learning Tasks on Repository-Level Code
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bbc19c93-e2ff-42b8-b25e-cd50ec355bb9 · outbound
Reinforcement Learning for Machine Learning Engineering Agents AutoKaggle: A Multi-Agent Framework for Autonomous Data Science Competitions
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 61707d3e-b884-480a-81fa-d9330985d6e4 · outbound
Reinforcement Learning for Machine Learning Engineering Agents Mlrc-bench: Can language agents solve machine learning research challenges? arXiv preprint arXiv:2504.09702, 2025
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 526feaf1-92f5-44d8-b182-03db1f964a93 · outbound
Reinforcement Learning for Machine Learning Engineering Agents Large language models orchestrating structured reasoning achieve kaggle grandmaster level
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b49b0a67-8f9d-4605-8bc0-3af9bf940daa · outbound
Reinforcement Learning for Machine Learning Engineering Agents Exploring LLM Agents for Cleaning Tabular Machine Learning Datasets
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4ad69974-7d04-4d41-9bf4-127080bd50a7 · outbound
Reinforcement Learning for Machine Learning Engineering Agents HardML: A Benchmark For Evaluating Data Science And Machine Learning knowledge and reasoning in AI
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e5ce4061-220c-472e-a887-5c0ffc568eca · outbound
Reinforcement Learning for Machine Learning Engineering Agents AutoML-GPT: Automatic Machine Learning with GPT
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f59c9e87-c56e-446c-ab10-a4a91451733a · outbound
Reinforcement Learning for Machine Learning Engineering Agents Large Language Model Agent for Hyper-Parameter Optimization
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 33638af2-a3fa-4291-bb47-147f2d248376 · outbound
Reinforcement Learning for Machine Learning Engineering Agents Large Language Models for Constructing and Optimizing Machine Learning Workflows: A Survey
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1bf3b52d-1d92-4dc9-99c0-deaf8c7e7ba6 · outbound
Reinforcement Learning for Machine Learning Engineering Agents AIDE: AI-Driven Exploration in the Space of Code
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 32b4121b-5f6d-406c-a8f9-763cddf1b368 · outbound
Reinforcement Learning for Machine Learning Engineering Agents I-mcts: Enhancing agentic automl via introspective monte carlo tree search
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f69bf019-b73c-4d00-bfab-c8cbadbbab01 · outbound
Reinforcement Learning for Machine Learning Engineering Agents Large Language Models Cannot Self-Correct Reasoning Yet
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f040e31f-271f-434e-905f-a6f0fe38d6ab · outbound
Reinforcement Learning for Machine Learning Engineering Agents What Did I Do Wrong? Quantifying LLMs' Sensitivity and Consistency to Prompt Engineering
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 47208247-8f1c-4afa-a06e-06c1ced3c7b3 · outbound
Reinforcement Learning for Machine Learning Engineering Agents Policy gradient meth- ods for reinforcement learning with function approximation
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b4c2fccd-badd-4374-bad3-4d06a283d41d · outbound
Reinforcement Learning for Machine Learning Engineering Agents A natural policy gradient
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2ce322a9-b301-4171-9f00-edee91f7b9eb · outbound
Reinforcement Learning for Machine Learning Engineering Agents Trust region policy optimization
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9c921657-35d1-4b7f-9254-1fb19008fabf · outbound
Reinforcement Learning for Machine Learning Engineering Agents Q-learning
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 01289653-dd89-4bdf-b7b1-b015b37a0994 · outbound
Reinforcement Learning for Machine Learning Engineering Agents Mujoco: A physics engine for model-based control
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c862a379-db9e-4500-a98d-c075fb028046 · outbound
Reinforcement Learning for Machine Learning Engineering Agents The arcade learning environment: An evaluation platform for general agents
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5ff97f97-2147-4292-8299-ebd823436c31 · outbound
Reinforcement Learning for Machine Learning Engineering Agents TensorFlow Agents: Efficient Batched Reinforcement Learning in TensorFlow
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0789f35c-5da4-4f0e-98f8-939f45184155 · outbound
Reinforcement Learning for Machine Learning Engineering Agents Training language models to follow instructions with human feedback
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e821198e-723f-440c-a807-9c844497fc6a · outbound
Reinforcement Learning for Machine Learning Engineering Agents Direct preference optimization: Your language model is secretly a reward model
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fba52266-aafe-477b-87f4-e98595bc8d66 · outbound
Reinforcement Learning for Machine Learning Engineering Agents Brown, Miljan Martic, Shane Legg, and Dario Amodei
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0c9982e0-0b5c-44a5-846c-e75b2aa4f9eb · outbound
Reinforcement Learning for Machine Learning Engineering Agents Ziegler, Nisan Stiennon, Jeffrey Wu, Tom B
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 791f7e20-5df9-4eac-bc45-a78e5bbeb58e · outbound
Reinforcement Learning for Machine Learning Engineering Agents Rlaif: Scaling reinforcement learning from human feedback with ai feedback
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 110fa51b-0c9e-426f-aed5-b3fd2a2f26c8 · outbound
Reinforcement Learning for Machine Learning Engineering Agents Reinforcement learning for reasoning in small llms: What works and what doesn’t
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a9be106b-4916-4424-b1d1-011921a36fee · outbound
Reinforcement Learning for Machine Learning Engineering Agents SWE-RL: Advancing LLM Reasoning via Reinforcement Learning on Open Software Evolution
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8d4775dd-103e-4381-b55a-ae26a0d00da0 · outbound
Reinforcement Learning for Machine Learning Engineering Agents DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 054e1061-806e-475b-af2f-b24adf327aa1 · outbound
Reinforcement Learning for Machine Learning Engineering Agents Agentbench: Evaluating llms as agents, 2023
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f725624d-fd98-4f03-9ba1-766c3fe32d60 · outbound
Reinforcement Learning for Machine Learning Engineering Agents Context-Aware Language Modeling for Goal-Oriented Dialogue Systems
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 6d631304-f8d8-4a31-bb62-83143c552e6e · outbound
Reinforcement Learning for Machine Learning Engineering Agents Digirl: Training in-the-wild device-control agents with autonomous reinforcement learning
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c4b31ec5-ca97-4e35-a30f-4380ee3cba57 · outbound
Reinforcement Learning for Machine Learning Engineering Agents CHAI: A CHatbot AI for Task-Oriented Dialogue with Offline Reinforcement Learning
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e560e437-4891-4104-bbe9-cfd7031438cb · outbound
Reinforcement Learning for Machine Learning Engineering Agents Xu, Hao Zhu, Xuhui Zhou, Robert Lo, Abishek Sridhar, Xianyi Cheng, Tianyue Ou, Yonatan Bisk, Daniel Fried and Uri Alon, and Graham Neubig
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a9c7e422-c2ae-4624-8714-edd191db63da · outbound
Reinforcement Learning for Machine Learning Engineering Agents Language Understanding for Text-based Games Using Deep Reinforcement Learning
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation bd9f13a9-47f1-4a00-b732-6889e6f75555 · outbound
Reinforcement Learning for Machine Learning Engineering Agents Offline RL for Natural Language Generation with Implicit Language Q Learning
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ef0ee11a-e246-440f-8cbe-5c3556f6446b · outbound
Reinforcement Learning for Machine Learning Engineering Agents Process Reward Models for LLM Agents: Practical Framework and Directions
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 86d3afb9-4361-46b8-b3d1-ff99a13bd789 · outbound
Reinforcement Learning for Machine Learning Engineering Agents Generative Verifiers: Reward Modeling as Next-Token Prediction
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation afebba1c-3ec7-4a55-b95c-cf857a34b4f3 · outbound
Reinforcement Learning for Machine Learning Engineering Agents Generative Reward Models
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b270066d-cec2-4cf4-9af0-0983af8ac565 · outbound
Reinforcement Learning for Machine Learning Engineering Agents Autonomous Evaluation and Refinement of Digital Agents
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5ef0f259-b734-480b-83f0-32aa46ffb8de · outbound
Reinforcement Learning for Machine Learning Engineering Agents Code as Reward: Empowering Reinforcement Learning with VLMs
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a27a564b-280b-4a49-950f-36f9344456a2 · outbound
Reinforcement Learning for Machine Learning Engineering Agents LLM-as-a-Judge & Reward Model: What They Can and Cannot Do
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d6fa3e54-ec90-476f-bda8-0103cde99a26 · outbound
Reinforcement Learning for Machine Learning Engineering Agents ‘ re- quest_id,requester_received_pizza t3_i8iy4,0 t3_1mfqi0,0 etc “‘ • Data snippet: -> /workdir/random-acts-of-pizza/prepared/public/test.json: [
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b3137228-5778-4160-88a4-893fd1c75603 · inbound
PACEvolve++: Improving Test-time Learning for Evolutionary Search Agents Reinforcement Learning for Machine Learning Engineering Agents
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5630521f-8c37-4617-a14a-a9b49de7a90b · inbound
MLS-Bench: A Holistic and Rigorous Assessment of AI Systems on Building Better AI Reinforcement Learning for Machine Learning Engineering Agents
Reference 114
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 57a77419-c52c-4371-a291-6d2b39a723ff · inbound
MLS-Bench: A Holistic and Rigorous Assessment of AI Systems on Building Better AI Reinforcement Learning for Machine Learning Engineering Agents
Reference 116
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 86ff746c-ce33-464e-926e-0bee3a1a5c7f · inbound
MLS-Bench: A Holistic and Rigorous Assessment of AI Systems on Building Better AI Reinforcement Learning for Machine Learning Engineering Agents
Reference 115
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 83d2cc52-cd90-48ae-a519-dba65306dc3c · inbound
Revisiting DAgger in the Era of LLM-Agents Reinforcement Learning for Machine Learning Engineering Agents
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 082a7f23-a38a-4140-b7dd-43e98d831d0b · inbound
Reinforcement Learning with Verifiable Physics: Post-training LLMs with Continuous Rewards Reinforcement Learning for Machine Learning Engineering Agents
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b223f18e-0fb4-47d0-86b2-c12b00bf2c00 · inbound
Matryoshka Agent: Unfolding Sub-Agents for Long-Horizon Machine Learning Engineering Reinforcement Learning for Machine Learning Engineering Agents
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.