Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 34 inbound Pith citation observations for arXiv:2503.15478.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T15:26:56.255061Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-06-30T17:34:57.712764Z
0 of 0 outbound references displayed
External citation measurements
No source-named external measurement is stored.
No outbound reference observations are available for this paper version.
Observation 14a9e272-3755-4f82-8ebb-4c5d0b603de9 · inbound
lmgame-Bench: How Good are LLMs at Playing Games? SWEET-RL: Training Multi-Turn LLM Agents on Collaborative Reasoning Tasks
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b169f929-fb11-4c83-8be7-a284473d8032 · inbound
ARPO:End-to-End Policy Optimization for GUI Agents with Experience Replay SWEET-RL: Training Multi-Turn LLM Agents on Collaborative Reasoning Tasks
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 73ce5631-804b-42e7-a0dc-7ac9d5f4e945 · inbound
Get Experience from Practice: LLM Agents with Record & Replay SWEET-RL: Training Multi-Turn LLM Agents on Collaborative Reasoning Tasks
Reference 93
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b6dca655-ce83-4790-b963-216ffd7ff111 · inbound
SPA-RL: Reinforcing LLM Agents via Stepwise Progress Attribution SWEET-RL: Training Multi-Turn LLM Agents on Collaborative Reasoning Tasks
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1454204d-6f50-4207-9c68-4aac41fc2c92 · inbound
OWL: Optimized Workforce Learning for General Multi-Agent Assistance in Real-World Task Automation SWEET-RL: Training Multi-Turn LLM Agents on Collaborative Reasoning Tasks
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4f807b02-017d-4bed-a1be-c32a4eab9262 · inbound
TimeHC-RL: Temporal-aware Hierarchical Cognitive Reinforcement Learning for Enhancing LLMs' Social Intelligence SWEET-RL: Training Multi-Turn LLM Agents on Collaborative Reasoning Tasks
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 125d7087-241d-4a98-9ab3-b3ac2c9313c8 · inbound
ARIA: Training Language Agents with Intention-Driven Reward Aggregation SWEET-RL: Training Multi-Turn LLM Agents on Collaborative Reasoning Tasks
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 17e33e8c-9b01-4cc9-a34c-5da8a28ac835 · inbound
Self-Challenging Language Model Agents SWEET-RL: Training Multi-Turn LLM Agents on Collaborative Reasoning Tasks
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3d37a5d8-7487-4031-94a2-c63c0b08b245 · inbound
Response-Level Rewards Are All You Need for Online Reinforcement Learning in LLMs: A Mathematical Perspective SWEET-RL: Training Multi-Turn LLM Agents on Collaborative Reasoning Tasks
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7625a69a-a765-439d-8c35-df5ee4ddd1c3 · inbound
PAG: Multi-Turn Reinforced LLM Self-Correction with Policy as Generative Verifier SWEET-RL: Training Multi-Turn LLM Agents on Collaborative Reasoning Tasks
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1f0dc4f4-f5e2-4c7e-91bb-808ff5e823e3 · inbound
Conversation Forests: The Key to Fine Tuning Large Language Models for Multi-Turn Medical Conversations is Branching SWEET-RL: Training Multi-Turn LLM Agents on Collaborative Reasoning Tasks
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 633bb555-fc2c-4003-9188-8f313ed67ce6 · inbound
How to Train a Leader: Hierarchical Reasoning in Multi-Agent LLMs SWEET-RL: Training Multi-Turn LLM Agents on Collaborative Reasoning Tasks
Reference 71
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dd58bca8-26d4-4d0c-b6ab-e0f0866f15b4 · inbound
General Modular Harness for LLM Agents in Multi-Turn Gaming Environments SWEET-RL: Training Multi-Turn LLM Agents on Collaborative Reasoning Tasks
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aa65f488-02e3-49d7-8feb-539207c4b1eb · inbound
The Landscape of Agentic Reinforcement Learning for LLMs: A Survey SWEET-RL: Training Multi-Turn LLM Agents on Collaborative Reasoning Tasks
Reference 169
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation df998f68-4a2b-44f7-8386-e11527426948 · inbound
RLBoost: Harvesting Preemptible Resources for Cost-Efficient Reinforcement Learning on LLMs SWEET-RL: Training Multi-Turn LLM Agents on Collaborative Reasoning Tasks
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 43c7aad6-49c0-49e6-906c-976e2360c91e · inbound
Tree Training: Accelerating Agentic LLMs Training via Shared Prefix Reuse SWEET-RL: Training Multi-Turn LLM Agents on Collaborative Reasoning Tasks
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 1a6179a0-be0e-4fec-a489-4edb3cfde94d · inbound
ABBEL: Learning Natural-Language Belief States for Memory-Efficient Interaction SWEET-RL: Training Multi-Turn LLM Agents on Collaborative Reasoning Tasks
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d137f4c8-ea46-414d-83e9-9e000525dd33 · inbound
Agentic Reasoning for Large Language Models SWEET-RL: Training Multi-Turn LLM Agents on Collaborative Reasoning Tasks
Reference 235
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 1a126558-2aa2-41fa-822f-bb3b714ebddb · inbound
Towards Generalizable Reasoning: Group Causal Counterfactual Policy Optimization for LLM Reasoning SWEET-RL: Training Multi-Turn LLM Agents on Collaborative Reasoning Tasks
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 04ab4af9-3fd6-40cc-b505-75a93ffb9a05 · inbound
Multi-Turn Reinforcement Learning for Tool-Calling Agents with Iterative Reward Calibration SWEET-RL: Training Multi-Turn LLM Agents on Collaborative Reasoning Tasks
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation f3d42dab-7dce-4859-ac9e-c1ed570faaeb · inbound
ActivityEditor: Learning to Synthesize Physically Valid Human Mobility SWEET-RL: Training Multi-Turn LLM Agents on Collaborative Reasoning Tasks
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 26c3d4ed-4455-4ae1-9410-e2bf0424da5c · inbound
PRIME: Training Free Proactive Reasoning via Iterative Memory Evolution for User-Centric Agent SWEET-RL: Training Multi-Turn LLM Agents on Collaborative Reasoning Tasks
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation a0fdffef-ee8e-439a-865e-3d2bcf9dfcf6 · inbound
GUI Agents with Reinforcement Learning: Toward Digital Inhabitants SWEET-RL: Training Multi-Turn LLM Agents on Collaborative Reasoning Tasks
Reference 88
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 11058383-d232-4920-a4e6-e116d4043f25 · inbound
CustomerSim: Benchmarking and Aligning Multimodal Language Models as Retail User Simulators SWEET-RL: Training Multi-Turn LLM Agents on Collaborative Reasoning Tasks
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 3390848c-dd81-46ca-98fc-9a18b501fb71 · inbound
CustomerSim: Benchmarking and Aligning Multimodal Language Models as Retail User Simulators SWEET-RL: Training Multi-Turn LLM Agents on Collaborative Reasoning Tasks
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b9ad27e3-ca85-4d65-96da-dcefaabc01cb · inbound
Step Rejection Fine-Tuning: A Practical Distillation Recipe SWEET-RL: Training Multi-Turn LLM Agents on Collaborative Reasoning Tasks
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 754fb7ba-0268-41af-9679-6dd5bcb6f288 · inbound
PAIR: Prefix-Aware Internal Reward Model for Multi-Turn Agent Optimization SWEET-RL: Training Multi-Turn LLM Agents on Collaborative Reasoning Tasks
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 7e2d00a5-471e-438c-94c9-204a55df7b01 · inbound
PAIR: Prefix-Aware Internal Reward Model for Multi-Turn Agent Optimization SWEET-RL: Training Multi-Turn LLM Agents on Collaborative Reasoning Tasks
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f5193e78-721e-4d63-849f-01c0f30ae694 · inbound
Unlocking Proactivity in Task-Oriented Dialogue SWEET-RL: Training Multi-Turn LLM Agents on Collaborative Reasoning Tasks
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 8d6e6666-2b71-4cb6-b8a8-6c70dc737ad5 · inbound
Unlocking Proactivity in Task-Oriented Dialogue SWEET-RL: Training Multi-Turn LLM Agents on Collaborative Reasoning Tasks
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 2e6c0734-d37c-44a7-ab34-e99e658384e5 · inbound
AppWorld-UL: Benchmarking Diverse Agent-User Interactions for Tool-Use SWEET-RL: Training Multi-Turn LLM Agents on Collaborative Reasoning Tasks
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fb439d78-5096-4625-ab82-618defb3f50a · inbound
Teaching LLMs to Self-Evolve: Cultivating Core Meta-Skills with Reinforcement Learning SWEET-RL: Training Multi-Turn LLM Agents on Collaborative Reasoning Tasks
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0188d5a1-8c76-42b0-9bc7-0a7d4b590917 · inbound
ODYSSE: Episode-wise Policy Optimization for Personalized Agentic Reasoning SWEET-RL: Training Multi-Turn LLM Agents on Collaborative Reasoning Tasks
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 29aec9b2-098f-4024-ada1-5ec02895ecfb · inbound
TCPO: Turn-Level Credit Policy Optimization SWEET-RL: Training Multi-Turn LLM Agents on Collaborative Reasoning Tasks
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.