Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T15:07:10.796768Z
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 29 of 29 outbound references and 16 inbound Pith citation observations for arXiv:2505.16282.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T15:07:10.796768Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-06T19:25:14.276099Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-04T20:00:08.408321Z
29 of 29 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation a97a4621-6c2e-4838-8b44-aac68c14df56 · outbound
ARPO:End-to-End Policy Optimization for GUI Agents with Experience Replay Qwen2.5-VL Technical Report
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a5544cf9-7703-4840-aa66-c648472154de · outbound
ARPO:End-to-End Policy Optimization for GUI Agents with Experience Replay SeeClick: Harnessing GUI Grounding for Advanced Visual GUI Agents
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 06a674f1-8a7e-48ab-8c65-6f05654887e2 · outbound
ARPO:End-to-End Policy Optimization for GUI Agents with Experience Replay KTO: Model Alignment as Prospect Theoretic Optimization
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a204acfe-dfc4-4120-a7ae-98cec497d88e · outbound
ARPO:End-to-End Policy Optimization for GUI Agents with Experience Replay Navigating the Digital World as Humans Do: Universal Visual Grounding for GUI Agents
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 96a71639-ea78-4543-8235-922a76b611cc · outbound
ARPO:End-to-End Policy Optimization for GUI Agents with Experience Replay DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f2923ec7-65ca-4e66-a110-9ab98fa6804c · outbound
ARPO:End-to-End Policy Optimization for GUI Agents with Experience Replay Cogagent: A visual language model for gui agents
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 08626187-7624-473a-848e-7420015aef08 · outbound
ARPO:End-to-End Policy Optimization for GUI Agents with Experience Replay Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language Models
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 48d13b54-9277-4d20-9f3b-7d8563c4eae4 · outbound
ARPO:End-to-End Policy Optimization for GUI Agents with Experience Replay OpenAI o1 System Card
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4b5481a0-74d7-401d-881c-9dac23cee2e4 · outbound
ARPO:End-to-End Policy Optimization for GUI Agents with Experience Replay OmniACT: A Dataset and Benchmark for Enabling Multimodal Generalist Autonomous Agents for Desktop and Web
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e0a72d0c-1f2b-439e-b76a-0815a1f77501 · outbound
ARPO:End-to-End Policy Optimization for GUI Agents with Experience Replay Efficient memory management for large language model serving with pagedattention
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 723a7e63-be07-4902-9360-6ab11888e498 · outbound
ARPO:End-to-End Policy Optimization for GUI Agents with Experience Replay Code-r1: Reproducing r1 for code with reliable rewards
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8ea43a1d-8027-47f5-b596-6acbc00c0eb8 · outbound
ARPO:End-to-End Policy Optimization for GUI Agents with Experience Replay Seg-Zero: Reasoning-Chain Guided Segmentation via Cognitive Reinforcement
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bc7d8539-4150-48cb-832c-6b4548dd2e82 · outbound
ARPO:End-to-End Policy Optimization for GUI Agents with Experience Replay Decoupled Weight Decay Regularization
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2a897a41-6265-4c2c-abd7-a4289024e73f · outbound
ARPO:End-to-End Policy Optimization for GUI Agents with Experience Replay ScreenAgent: A Vision Language Model-driven Computer Control Agent
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 53b21184-ee19-4e15-a7de-b8d3d4fa34d7 · outbound
ARPO:End-to-End Policy Optimization for GUI Agents with Experience Replay ToolRL: Reward is All Tool Learning Needs
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 202e5032-e30c-4814-905a-6ed48af0b141 · outbound
ARPO:End-to-End Policy Optimization for GUI Agents with Experience Replay UI-TARS: Pioneering Automated GUI Interaction with Native Agents
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 15b11a46-da31-4f47-aea5-c4c68fc8cadd · outbound
ARPO:End-to-End Policy Optimization for GUI Agents with Experience Replay Direct preference optimization: Your language model is secretly a reward model
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 1f8adde8-ef64-4dfe-a2b9-59b519a037ef · outbound
ARPO:End-to-End Policy Optimization for GUI Agents with Experience Replay Proximal Policy Optimization Algorithms
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1d28fa66-0dad-4959-9ec7-efc2a223b9c2 · outbound
ARPO:End-to-End Policy Optimization for GUI Agents with Experience Replay DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 440ba8ab-3650-4f73-bf52-0b10af21267c · outbound
ARPO:End-to-End Policy Optimization for GUI Agents with Experience Replay HybridFlow: A Flexible and Efficient RLHF Framework
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6ababca5-5736-47a2-924b-fd00f9d5b887 · outbound
ARPO:End-to-End Policy Optimization for GUI Agents with Experience Replay RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eda3ee73-fe85-4711-a4cc-bec481c35515 · outbound
ARPO:End-to-End Policy Optimization for GUI Agents with Experience Replay Chain-of-thought prompting elicits reasoning in large language models
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 8f180272-3fd7-41f7-af45-ed5731c23050 · outbound
ARPO:End-to-End Policy Optimization for GUI Agents with Experience Replay OS-ATLAS: A Foundation Action Model for Generalist GUI Agents
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 58f4df80-56d4-438d-8f34-5af8fc91a67a · outbound
ARPO:End-to-End Policy Optimization for GUI Agents with Experience Replay GUI-R1 : A Generalist R1-Style Vision-Language Action Model For GUI Agents
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f4d6f9ea-76b2-4537-a06c-9ecfa4747bd3 · outbound
ARPO:End-to-End Policy Optimization for GUI Agents with Experience Replay Osworld: Benchmarking multimodal agents for open-ended tasks in real computer environments
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation ca674936-63b4-4926-aa11-f4c6420fc5b6 · outbound
ARPO:End-to-End Policy Optimization for GUI Agents with Experience Replay Aguvis: Unified Pure Vision Agents for Autonomous GUI Interaction
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a363f308-e927-4e0c-9fbf-e202d39c717e · outbound
ARPO:End-to-End Policy Optimization for GUI Agents with Experience Replay Aria-UI: Visual Grounding for GUI Instructions
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f9423a99-3bc6-4f52-a8cb-ee85a07db547 · outbound
ARPO:End-to-End Policy Optimization for GUI Agents with Experience Replay DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b169f929-fb11-4c83-8be7-a284473d8032 · outbound
ARPO:End-to-End Policy Optimization for GUI Agents with Experience Replay SWEET-RL: Training Multi-Turn LLM Agents on Collaborative Reasoning Tasks
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c7df2476-8c0a-48d9-8a9f-2cf60b09acd7 · inbound
MobileGUI-RL: Advancing Mobile GUI Agent through Reinforcement Learning in Online Environment ARPO:End-to-End Policy Optimization for GUI Agents with Experience Replay
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation efa7606d-90a1-4c89-a24d-f755e76324f5 · inbound
UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning ARPO:End-to-End Policy Optimization for GUI Agents with Experience Replay
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 0acd32bd-456b-4de0-87eb-b50696880523 · inbound
Plan Then Action:High-Level Planning Guidance Reinforcement Learning for LLM Reasoning ARPO:End-to-End Policy Optimization for GUI Agents with Experience Replay
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 61c88db6-3fd5-4294-852b-ee3cc6395ca1 · inbound
AutoTool: Dynamic Tool Selection and Integration for Agentic Reasoning ARPO:End-to-End Policy Optimization for GUI Agents with Experience Replay
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f7a6fa83-59c1-42f8-bd82-ecc76dcd26e5 · inbound
Agentic Reasoning for Large Language Models ARPO:End-to-End Policy Optimization for GUI Agents with Experience Replay
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation a00a8d54-94c6-4eff-8a4b-afedab8048e3 · inbound
Turing Test on Screen: A Benchmark for Mobile GUI Agent Humanization ARPO:End-to-End Policy Optimization for GUI Agents with Experience Replay
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 32ddfcb8-f0f4-4f19-ad06-b0543b4f3018 · inbound
Faithful Mobile GUI Agents with Guided Advantage Estimator ARPO:End-to-End Policy Optimization for GUI Agents with Experience Replay
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 3ab69e72-872e-4417-b40f-66d6b72ad9eb · inbound
On Training Large Language Models for Long-Horizon Tasks: An Empirical Study of Horizon Length ARPO:End-to-End Policy Optimization for GUI Agents with Experience Replay
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 61a0ea97-88b7-4880-83e2-222fca09fa1a · inbound
Learning to Explore: Scaling Agentic Reasoning via Exploration-Aware Policy Optimization ARPO:End-to-End Policy Optimization for GUI Agents with Experience Replay
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation e43a52a5-801e-44c0-a4e7-1f00e7d6cdf0 · inbound
Learning to Explore: Scaling Agentic Reasoning via Exploration-Aware Policy Optimization ARPO:End-to-End Policy Optimization for GUI Agents with Experience Replay
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 494b1bdb-c798-4aa8-b384-8e948dd3bd4f · inbound
ToolCUA: Towards Optimal GUI-Tool Path Orchestration for Computer Use Agents ARPO:End-to-End Policy Optimization for GUI Agents with Experience Replay
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation ad7335ec-cd17-4d1b-85cc-f1e71fcfc475 · inbound
RLVR without Ineffective Samples: Group Prioritized Off-Policy Optimization for LLM Reasoning ARPO:End-to-End Policy Optimization for GUI Agents with Experience Replay
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 44daa594-25cc-420c-a599-f689908c2cdb · inbound
StainFlow: Entity-Stain Tracking and Evidence Linking for Process Rewards in GUI Agents ARPO:End-to-End Policy Optimization for GUI Agents with Experience Replay
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 46e7f937-ef14-4065-9ae5-ae9c9c4aa368 · inbound
Modularized Reinforcement Learning on LLMs: From MDP Creation to Exploration and Learning ARPO:End-to-End Policy Optimization for GUI Agents with Experience Replay
Reference 125
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 3450fa5e-24d1-467e-99a6-dbf7f0c44049 · inbound
Learning with a Single Rollout via Monte Carlo Pass@k Critic ARPO:End-to-End Policy Optimization for GUI Agents with Experience Replay
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 490492e7-74a4-411e-afa8-b66bcd8c2a8d · inbound
DocArena: Turning Raw Documents into Controllable Training Environments for Document Search Agents ARPO:End-to-End Policy Optimization for GUI Agents with Experience Replay
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.