Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-05T16:22:51.681755Z
Paper Citation Record · LEDGER
As of 9 August 2026, this Paper Citation Record lists 47 of 47 outbound references and 10 inbound Pith citation observations for arXiv:2508.18669.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-05T16:22:51.681755Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-03T18:03:39.498611Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-07-08T15:35:07.151222Z
47 of 47 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation cd035e9a-a855-4838-9a5a-097e5365af56 · outbound
MUA-RL: Multi-turn User-interacting Agent Reinforcement Learning for agentic tool use Kimi K2: Open Agentic Intelligence
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 05603cab-74cd-4ddd-a45a-4f6d17706760 · outbound
MUA-RL: Multi-turn User-interacting Agent Reinforcement Learning for agentic tool use DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a0fb646b-3a1e-45cb-adb4-6f1ff32937aa · outbound
MUA-RL: Multi-turn User-interacting Agent Reinforcement Learning for agentic tool use ReTool: Reinforcement Learning for Strategic Tool Use in LLMs
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8fe26866-1bab-4b30-bb80-e618129c16b4 · outbound
MUA-RL: Multi-turn User-interacting Agent Reinforcement Learning for agentic tool use Gonzalez, and Ion Stoica
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 20ea23a1-92a8-4a08-8f9b-c2c8a16bfe37 · outbound
MUA-RL: Multi-turn User-interacting Agent Reinforcement Learning for agentic tool use RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3e7b5f80-c70a-47ab-acb1-122be6293cfe · outbound
MUA-RL: Multi-turn User-interacting Agent Reinforcement Learning for agentic tool use Proximal Policy Optimization Algorithms
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a1dd30f9-0879-46db-82d8-0df364935d88 · outbound
MUA-RL: Multi-turn User-interacting Agent Reinforcement Learning for agentic tool use Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 631af179-9c0a-4ff3-8b35-4918fdcf5670 · outbound
MUA-RL: Multi-turn User-interacting Agent Reinforcement Learning for agentic tool use Buy 4 reinforce samples, get a baseline for free! Learn- ing,Learning, Mar 2019
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 724f0b98-6405-4417-b4f9-e4fd58dc0f24 · outbound
MUA-RL: Multi-turn User-interacting Agent Reinforcement Learning for agentic tool use DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e26e9302-4be4-4c86-8859-9b443d212786 · outbound
MUA-RL: Multi-turn User-interacting Agent Reinforcement Learning for agentic tool use Star: Bootstrapping reasoning with reasoning
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 0e42f05a-7af7-4101-b757-a994d78ff299 · outbound
MUA-RL: Multi-turn User-interacting Agent Reinforcement Learning for agentic tool use Reasoning with Language Model is Planning with World Model
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8eceef6f-2047-4f61-82d9-b510836faff8 · outbound
MUA-RL: Multi-turn User-interacting Agent Reinforcement Learning for agentic tool use Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a0bc6eb0-5b3c-41a5-bc9e-111085d807b4 · outbound
MUA-RL: Multi-turn User-interacting Agent Reinforcement Learning for agentic tool use Code-r1: Reproducing r1 for code with reliable rewards
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 57ac40bf-5d47-480f-9deb-01ad51d2eccb · outbound
MUA-RL: Multi-turn User-interacting Agent Reinforcement Learning for agentic tool use VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ea85cc3e-a2a6-40ac-a17d-32a8c242341c · outbound
MUA-RL: Multi-turn User-interacting Agent Reinforcement Learning for agentic tool use Program of Thoughts Prompting: Disentangling Computation from Reasoning for Numerical Reasoning Tasks
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cb121056-e228-4f72-9957-c62926009f79 · outbound
MUA-RL: Multi-turn User-interacting Agent Reinforcement Learning for agentic tool use Instructerc: Reforming emotion recognition in conversation with a retrieval multi-task llms framework
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 955b1022-2129-4f30-99d0-34d7346f240c · outbound
MUA-RL: Multi-turn User-interacting Agent Reinforcement Learning for agentic tool use Hammer: Robust function-calling for on-device language models via function masking
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1ee74172-ff78-4f40-9175-cbb1a1b76af0 · outbound
MUA-RL: Multi-turn User-interacting Agent Reinforcement Learning for agentic tool use xlam: A family of large action models to empower ai agent systems
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 3d4b5ff1-95d3-4335-9607-07b90c82f338 · outbound
MUA-RL: Multi-turn User-interacting Agent Reinforcement Learning for agentic tool use Can a Single Model Master Both Multi-turn Conversations and Tool Use? CoALM: A Unified Conversational Agentic Language Model
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 115a9a25-065f-4794-89d2-efa50f9e47bf · outbound
MUA-RL: Multi-turn User-interacting Agent Reinforcement Learning for agentic tool use R1-Searcher: Incentivizing the Search Capability in LLMs via Reinforcement Learning
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4aa972d9-522e-40aa-9b7d-71a8afa19c0e · outbound
MUA-RL: Multi-turn User-interacting Agent Reinforcement Learning for agentic tool use ZeroSearch: Incentivize the Search Capability of LLMs without Searching
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a7d18c52-045b-4390-a996-eb97c0cad2ef · outbound
MUA-RL: Multi-turn User-interacting Agent Reinforcement Learning for agentic tool use ToRL: Scaling Tool-Integrated RL
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e52003de-437d-4409-ae87-f10b1522336c · outbound
MUA-RL: Multi-turn User-interacting Agent Reinforcement Learning for agentic tool use Agentic Reasoning and Tool Integration for LLMs via Reinforcement Learning
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6028b4a8-4a88-4342-a4a3-59af3a053aaa · outbound
MUA-RL: Multi-turn User-interacting Agent Reinforcement Learning for agentic tool use AgentInstruct: Toward Generative Teaching with Agentic Flows
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a3d7c793-3654-490f-927a-e3eda35011e3 · outbound
MUA-RL: Multi-turn User-interacting Agent Reinforcement Learning for agentic tool use StableToolBench: Towards Stable Large-Scale Benchmarking on Tool Learning of Large Language Models
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 116e2fdc-2134-4e72-b837-be93e9df192a · outbound
MUA-RL: Multi-turn User-interacting Agent Reinforcement Learning for agentic tool use Model Context Protocol (MCP): Landscape, Security Threats, and Future Research Directions
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3fa98f77-b13e-4083-8f07-8d24fc1c566e · outbound
MUA-RL: Multi-turn User-interacting Agent Reinforcement Learning for agentic tool use DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cd6e296d-3ea8-473e-b4a2-77d08c328fd4 · outbound
MUA-RL: Multi-turn User-interacting Agent Reinforcement Learning for agentic tool use Unresolved cited work
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 84ac58dc-b084-4e81-822e-dab037fe247c · outbound
MUA-RL: Multi-turn User-interacting Agent Reinforcement Learning for agentic tool use ReSearch: Learning to Reason with Search for LLMs via Reinforcement Learning
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2a752a56-5318-43b5-8de1-6689bd32a4e0 · outbound
MUA-RL: Multi-turn User-interacting Agent Reinforcement Learning for agentic tool use WebThinker: Empowering Large Reasoning Models with Deep Research Capability
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 16a9385a-cff2-4ccf-b27f-660bae9b7ddb · outbound
MUA-RL: Multi-turn User-interacting Agent Reinforcement Learning for agentic tool use Plm-based world models for text-based games
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation d255c769-1a32-4ece-a74a-6ca7d801855c · outbound
MUA-RL: Multi-turn User-interacting Agent Reinforcement Learning for agentic tool use Generative agents: Interactive simulacra of human behavior
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 07ad9f53-0fc0-4992-8509-da22c3a4d494 · outbound
MUA-RL: Multi-turn User-interacting Agent Reinforcement Learning for agentic tool use ToolRL: Reward is All Tool Learning Needs
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 159b4990-e5c7-4400-b315-9e5dbff73feb · outbound
MUA-RL: Multi-turn User-interacting Agent Reinforcement Learning for agentic tool use Nemotron-Research-Tool-N1: Exploring Tool-Using Language Models with Reinforced Reasoning
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ed009504-4a98-4e4c-bb84-d7d1fe05d32d · outbound
MUA-RL: Multi-turn User-interacting Agent Reinforcement Learning for agentic tool use Reinforcing multi-turn reasoning in llm agents via turn-level credit assignment
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 28acd782-d161-40c5-9e51-5a45b0b240a9 · outbound
MUA-RL: Multi-turn User-interacting Agent Reinforcement Learning for agentic tool use Qwen3 Technical Report
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 29c0c48d-3474-413f-bc63-c3baff453695 · outbound
MUA-RL: Multi-turn User-interacting Agent Reinforcement Learning for agentic tool use Decoupled Weight Decay Regularization
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 99c154c9-34e3-4a41-9e52-733796921874 · outbound
MUA-RL: Multi-turn User-interacting Agent Reinforcement Learning for agentic tool use HybridFlow: A Flexible and Efficient RLHF Framework
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c329e446-20e4-4e70-a34b-45e5ac0a92eb · outbound
MUA-RL: Multi-turn User-interacting Agent Reinforcement Learning for agentic tool use $\tau$-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 86f669ba-ffd7-4308-b669-38a751da45e4 · outbound
MUA-RL: Multi-turn User-interacting Agent Reinforcement Learning for agentic tool use GPT-4o System Card
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 260d555d-7799-4ba3-93ef-e9f0e6d4d55d · outbound
MUA-RL: Multi-turn User-interacting Agent Reinforcement Learning for agentic tool use $\tau^2$-Bench: Evaluating Conversational Agents in a Dual-Control Environment
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 89b5a9ae-d96d-434f-aaf2-dc4ff909cb15 · outbound
MUA-RL: Multi-turn User-interacting Agent Reinforcement Learning for agentic tool use The berkeley function calling leaderboard (bfcl): From tool use to agentic evaluation of large language models
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation b4a34744-4040-472f-a8f5-24a6df2c4a35 · outbound
MUA-RL: Multi-turn User-interacting Agent Reinforcement Learning for agentic tool use Acebench: Who wins the match point in tool usage? arXiv preprint arXiv:2501.12851, 2025
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d533bfd7-f531-4160-bdf8-d683f7845a16 · outbound
MUA-RL: Multi-turn User-interacting Agent Reinforcement Learning for agentic tool use DeepSeek-V3 Technical Report
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 14f7e64c-b556-46ef-bbd3-11a642a4470e · outbound
MUA-RL: Multi-turn User-interacting Agent Reinforcement Learning for agentic tool use GLM-4.5: Agentic, Reasoning, and Coding (ARC) Foundation Models
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7cb328aa-1ec2-4dc1-84cd-094d76a1f226 · outbound
MUA-RL: Multi-turn User-interacting Agent Reinforcement Learning for agentic tool use term":"Sakura
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 9b796e61-8568-4fab-b7ea-ccc1bd211a66 · outbound
MUA-RL: Multi-turn User-interacting Agent Reinforcement Learning for agentic tool use 02:30:00
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 6d4fd521-4b00-4928-b9ea-5952643f3ce4 · inbound
Experience-Evolving Multi-Turn Tool-Use Agent with Hybrid Episodic-Procedural Memory MUA-RL: Multi-turn User-interacting Agent Reinforcement Learning for agentic tool use
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f10900cb-86fe-4da9-8034-b54a86907c95 · inbound
Agent-World: Scaling Real-World Environment Synthesis for Evolving General Agent Intelligence MUA-RL: Multi-turn User-interacting Agent Reinforcement Learning for agentic tool use
Reference 134
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation ee64902f-ac2d-4f45-bee9-a041bc1bbb7f · inbound
Generate, Filter, Control, Replay: A Comprehensive Survey of Rollout Strategies for LLM Reinforcement Learning MUA-RL: Multi-turn User-interacting Agent Reinforcement Learning for agentic tool use
Reference 181
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation a31372e8-1eb0-4ce9-b5fa-7c954d5b0eb0 · inbound
CustomerSim: Benchmarking and Aligning Multimodal Language Models as Retail User Simulators MUA-RL: Multi-turn User-interacting Agent Reinforcement Learning for agentic tool use
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 3bf48290-6246-46b3-98c8-504dd83bb253 · inbound
CustomerSim: Benchmarking and Aligning Multimodal Language Models as Retail User Simulators MUA-RL: Multi-turn User-interacting Agent Reinforcement Learning for agentic tool use
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4812cde0-8357-4799-be70-1e70dcc89e76 · inbound
When Simulation Lies: A Sim-to-Real Benchmark and Domain-Randomized RL Recipe for Tool-Use Agents MUA-RL: Multi-turn User-interacting Agent Reinforcement Learning for agentic tool use
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 633cafca-e809-453b-a331-369110412474 · inbound
Trust Region On-Policy Distillation MUA-RL: Multi-turn User-interacting Agent Reinforcement Learning for agentic tool use
Reference 141
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 72b083e3-7392-4955-898f-74a650db6a31 · inbound
Synthesize and Reward -- Reinforcement Learning for Multi-Step Tool Use in Live Environments MUA-RL: Multi-turn User-interacting Agent Reinforcement Learning for agentic tool use
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 73dd061a-8945-445a-86b7-ba01e5e90857 · inbound
CurateEvo: Data-Curation Evolving for Agentic Post-Training MUA-RL: Multi-turn User-interacting Agent Reinforcement Learning for agentic tool use
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation b1bb2712-7032-43aa-84fb-fa6502ce6bdc · inbound
AgentOmnia: Scaling Agentic Models for Full-Scenario Applications MUA-RL: Multi-turn User-interacting Agent Reinforcement Learning for agentic tool use
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.