Pith. sign in

Paper Citation Record · LEDGER

ARPO:End-to-End Policy Optimization for GUI Agents with Experience Replay

As of 16 August 2026, this Paper Citation Record lists 29 of 29 outbound references and 17 inbound Pith citation observations for arXiv:2505.16282.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.16282 v1

Coverage vector

measured 29 of 29 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:07:10.796768Z

measured 46 of 46 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 17 of 17 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-11T20:19:15.434119Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T20:00:08.408321Z

Reference resolution

29 of 29 outbound references displayed

  • verified exact0
  • verified fuzzy5
  • unresolved24
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation a97a4621-6c2e-4838-8b44-aac68c14df56 · outbound

This paper cites Qwen2.5-VL Technical Report.

ARPO:End-to-End Policy Optimization for GUI Agents with Experience Replay Qwen2.5-VL Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T15:07:08.732949Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:07:08.732949Z digest=sha256:ac09f36d19a262995c7899cea98d48ec67a089981ce840d36a6b7ab3f8b90ec3

Observation a5544cf9-7703-4840-aa66-c648472154de · outbound

This paper cites SeeClick: Harnessing GUI Grounding for Advanced Visual GUI Agents.

ARPO:End-to-End Policy Optimization for GUI Agents with Experience Replay SeeClick: Harnessing GUI Grounding for Advanced Visual GUI Agents

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T15:07:08.818285Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:07:08.818285Z digest=sha256:bd6f3e61a7ebc98f39d556ef3310d2ce70549870431ea90c9384cf6a5f56930e

Observation 06a674f1-8a7e-48ab-8c65-6f05654887e2 · outbound

This paper cites KTO: Model Alignment as Prospect Theoretic Optimization.

ARPO:End-to-End Policy Optimization for GUI Agents with Experience Replay KTO: Model Alignment as Prospect Theoretic Optimization

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T15:07:08.943210Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:07:08.943210Z digest=sha256:cbfe4c3d58f4ecfe0b7e5c8bd79bf4a271fc1290d5bee6ac906098211d40fa8e

Observation a204acfe-dfc4-4120-a7ae-98cec497d88e · outbound

This paper cites Navigating the Digital World as Humans Do: Universal Visual Grounding for GUI Agents.

ARPO:End-to-End Policy Optimization for GUI Agents with Experience Replay Navigating the Digital World as Humans Do: Universal Visual Grounding for GUI Agents

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T15:07:09.054423Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:07:09.054423Z digest=sha256:71a3c1862ad52ca6fa379a2f54a54fd274f0c60546afd0ead178e52dbf51e03f

Observation 96a71639-ea78-4543-8235-922a76b611cc · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

ARPO:End-to-End Policy Optimization for GUI Agents with Experience Replay DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T15:07:09.148499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:07:09.148499Z digest=sha256:ce7b2811cde26ed9b904e4600e64edba63f41e59ea06130eff32d7798cee758a

Observation f2923ec7-65ca-4e66-a110-9ab98fa6804c · outbound

This paper cites Cogagent: A visual language model for gui agents.

ARPO:End-to-End Policy Optimization for GUI Agents with Experience Replay Cogagent: A visual language model for gui agents

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:07:11.440915Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T15:07:09.257984Z digest=sha256:8ee52e99c47536641d8bff1a5003e89f21ca3dd10eba762266064e6956ab3781

Observation 08626187-7624-473a-848e-7420015aef08 · outbound

This paper cites Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language Models.

ARPO:End-to-End Policy Optimization for GUI Agents with Experience Replay Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T15:07:09.358618Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:07:09.358618Z digest=sha256:6cbd3abe405c23d710a65fdf2a69775bfa39fb1bbd116e4adcdc15aca98f3af8

Observation 48d13b54-9277-4d20-9f3b-7d8563c4eae4 · outbound

This paper cites OpenAI o1 System Card.

ARPO:End-to-End Policy Optimization for GUI Agents with Experience Replay OpenAI o1 System Card

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T15:07:09.452790Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:07:09.452790Z digest=sha256:89ef27ca83c7ca23d387c4b8d54d6bc56babfe5cefc4ba560a0917282a702efc

Observation 4b5481a0-74d7-401d-881c-9dac23cee2e4 · outbound

This paper cites OmniACT: A Dataset and Benchmark for Enabling Multimodal Generalist Autonomous Agents for Desktop and Web.

ARPO:End-to-End Policy Optimization for GUI Agents with Experience Replay OmniACT: A Dataset and Benchmark for Enabling Multimodal Generalist Autonomous Agents for Desktop and Web

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T15:07:09.526745Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:07:09.526745Z digest=sha256:4ae895f58d0702da0b6bbeeae9e0d5cb023e4c91b870848cf224a64d0bc5c8b5

Observation e0a72d0c-1f2b-439e-b76a-0815a1f77501 · outbound

This paper cites Efficient memory management for large language model serving with pagedattention.

ARPO:End-to-End Policy Optimization for GUI Agents with Experience Replay Efficient memory management for large language model serving with pagedattention

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:07:11.374189Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T15:07:09.645013Z digest=sha256:f7b85b3861f0e9936dff9d344a308333508c7bd8efcb43d95322ebd783b881ca

Observation 723a7e63-be07-4902-9360-6ab11888e498 · outbound

This paper cites Code-r1: Reproducing r1 for code with reliable rewards.

ARPO:End-to-End Policy Optimization for GUI Agents with Experience Replay Code-r1: Reproducing r1 for code with reliable rewards

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T15:07:09.737013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:07:09.737013Z digest=sha256:5ecf738a097ddb063b867665bd036c634358aacf0198dad463cf4056692cd130

Observation 8ea43a1d-8027-47f5-b596-6acbc00c0eb8 · outbound

This paper cites Seg-Zero: Reasoning-Chain Guided Segmentation via Cognitive Reinforcement.

ARPO:End-to-End Policy Optimization for GUI Agents with Experience Replay Seg-Zero: Reasoning-Chain Guided Segmentation via Cognitive Reinforcement

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T15:07:09.853262Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:07:09.853262Z digest=sha256:6f6ab956e0b20706403065aebe9df31bf13d2569b7695197f4574b29b7ad6b6a

Observation bc7d8539-4150-48cb-832c-6b4548dd2e82 · outbound

This paper cites Decoupled Weight Decay Regularization.

ARPO:End-to-End Policy Optimization for GUI Agents with Experience Replay Decoupled Weight Decay Regularization

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T15:07:09.981894Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:07:09.981894Z digest=sha256:fa9de63b68ab59e714295ecabf0464a8885397c67519100a1ef258c22f13cf82

Observation 2a897a41-6265-4c2c-abd7-a4289024e73f · outbound

This paper cites ScreenAgent: A Vision Language Model-driven Computer Control Agent.

ARPO:End-to-End Policy Optimization for GUI Agents with Experience Replay ScreenAgent: A Vision Language Model-driven Computer Control Agent

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T15:07:10.028219Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:07:10.028219Z digest=sha256:12fb33c4893ee93234553eda464829b6d097fe9010c0736f32ad7bf2a53f5e40

Observation 53b21184-ee19-4e15-a7de-b8d3d4fa34d7 · outbound

This paper cites ToolRL: Reward is All Tool Learning Needs.

ARPO:End-to-End Policy Optimization for GUI Agents with Experience Replay ToolRL: Reward is All Tool Learning Needs

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T15:07:10.067180Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:07:10.067180Z digest=sha256:7de14e2ced0fcd9d710fd5f7606c506224aacd1c4dd8c8661536bc9759cf0f79

Observation 202e5032-e30c-4814-905a-6ed48af0b141 · outbound

This paper cites UI-TARS: Pioneering Automated GUI Interaction with Native Agents.

ARPO:End-to-End Policy Optimization for GUI Agents with Experience Replay UI-TARS: Pioneering Automated GUI Interaction with Native Agents

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T15:07:10.120332Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:07:10.120332Z digest=sha256:571bbead902e8594e72326d8a1d12a93df4ffe4507f0e536ffc298c682fd58a8

Observation 15b11a46-da31-4f47-aea5-c4c68fc8cadd · outbound

This paper cites Direct preference optimization: Your language model is secretly a reward model.

ARPO:End-to-End Policy Optimization for GUI Agents with Experience Replay Direct preference optimization: Your language model is secretly a reward model

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:07:11.288367Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T15:07:10.174192Z digest=sha256:e27f848d45f51d491ee1174414c5f6f1707c29675f164b129ed4633d4fee7057

Observation 1f8adde8-ef64-4dfe-a2b9-59b519a037ef · outbound

This paper cites Proximal Policy Optimization Algorithms.

ARPO:End-to-End Policy Optimization for GUI Agents with Experience Replay Proximal Policy Optimization Algorithms

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T15:07:10.211575Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:07:10.211575Z digest=sha256:a36cdee65c49c5f082073625c8334499ba1870b84564981dd2f7ee060da4fb9f

Observation 1d28fa66-0dad-4959-9ec7-efc2a223b9c2 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

ARPO:End-to-End Policy Optimization for GUI Agents with Experience Replay DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T15:07:10.279353Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:07:10.279353Z digest=sha256:56348d482079a25611a924ac671390cf428013e3d041a66d5c8c0c37e1161dcc

Observation 440ba8ab-3650-4f73-bf52-0b10af21267c · outbound

This paper cites HybridFlow: A Flexible and Efficient RLHF Framework.

ARPO:End-to-End Policy Optimization for GUI Agents with Experience Replay HybridFlow: A Flexible and Efficient RLHF Framework

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T15:07:10.332778Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:07:10.332778Z digest=sha256:a0ab5aea64f1512b67d2b7e95d4ab0e5ba3ac82a53d17834e136d4fa36bb86d1

Observation 6ababca5-5736-47a2-924b-fd00f9d5b887 · outbound

This paper cites RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning.

ARPO:End-to-End Policy Optimization for GUI Agents with Experience Replay RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T15:07:10.363216Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:07:10.363216Z digest=sha256:353e8a593239e1922e07b0ad3ac964c01686292f924db2768e235e61ed87c7ad

Observation eda3ee73-fe85-4711-a4cc-bec481c35515 · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models.

ARPO:End-to-End Policy Optimization for GUI Agents with Experience Replay Chain-of-thought prompting elicits reasoning in large language models

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:07:11.194077Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T15:07:10.403817Z digest=sha256:6846a1ff366a40735f2d2a08ac37e25858b689bc1048832a94f6c01dc95a36f7

Observation 8f180272-3fd7-41f7-af45-ed5731c23050 · outbound

This paper cites OS-ATLAS: A Foundation Action Model for Generalist GUI Agents.

ARPO:End-to-End Policy Optimization for GUI Agents with Experience Replay OS-ATLAS: A Foundation Action Model for Generalist GUI Agents

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T15:07:10.444576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:07:10.444576Z digest=sha256:081d0c1dd7361432164f2660924a3c237d614315cda8bd0f594f9d8df4d61fc2

Observation 58f4df80-56d4-438d-8f34-5af8fc91a67a · outbound

This paper cites GUI-R1 : A Generalist R1-Style Vision-Language Action Model For GUI Agents.

ARPO:End-to-End Policy Optimization for GUI Agents with Experience Replay GUI-R1 : A Generalist R1-Style Vision-Language Action Model For GUI Agents

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T15:07:10.491309Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:07:10.491309Z digest=sha256:236fb935e3e419d1c5b5f7b08bf71799e65c76528c8ceea5deabda6fc08ba1ef

Observation f4d6f9ea-76b2-4537-a06c-9ecfa4747bd3 · outbound

This paper cites Osworld: Benchmarking multimodal agents for open-ended tasks in real computer environments.

ARPO:End-to-End Policy Optimization for GUI Agents with Experience Replay Osworld: Benchmarking multimodal agents for open-ended tasks in real computer environments

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:07:11.128405Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-07T15:07:10.537390Z digest=sha256:3753618f17674fb38570d36c1c9bc486156ee83dd6c469e9aef3621c7e28b9aa

Observation ca674936-63b4-4926-aa11-f4c6420fc5b6 · outbound

This paper cites Aguvis: Unified Pure Vision Agents for Autonomous GUI Interaction.

ARPO:End-to-End Policy Optimization for GUI Agents with Experience Replay Aguvis: Unified Pure Vision Agents for Autonomous GUI Interaction

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T15:07:10.588099Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:07:10.588099Z digest=sha256:694b3748aefdbaa85b5b1ea91c4cfc2baeb32c1c3b6f102dcb7b74cadf4580b1

Observation a363f308-e927-4e0c-9fbf-e202d39c717e · outbound

This paper cites Aria-UI: Visual Grounding for GUI Instructions.

ARPO:End-to-End Policy Optimization for GUI Agents with Experience Replay Aria-UI: Visual Grounding for GUI Instructions

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T15:07:10.649086Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:07:10.649086Z digest=sha256:1b88239442904988923d5cb8e2d9a1765e52440cad2c0150bbaa745fec3ffd54

Observation f9423a99-3bc6-4f52-a8cb-ee85a07db547 · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

ARPO:End-to-End Policy Optimization for GUI Agents with Experience Replay DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T15:07:10.758735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:07:10.758735Z digest=sha256:8d2fa65655a73eb447b2958bad587945c732880ef8db773ab5c97eb807892a77

Observation b169f929-fb11-4c83-8be7-a284473d8032 · outbound

This paper cites SWEET-RL: Training Multi-Turn LLM Agents on Collaborative Reasoning Tasks.

ARPO:End-to-End Policy Optimization for GUI Agents with Experience Replay SWEET-RL: Training Multi-Turn LLM Agents on Collaborative Reasoning Tasks

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T15:07:10.796768Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:07:10.796768Z digest=sha256:cc8b3b32cc1b4bf66e463e414d7b4013fdb7f983a711099388dec835fe27ea3c

Pith citing papers

Observation c7df2476-8c0a-48d9-8a9f-2cf60b09acd7 · inbound

MobileGUI-RL: Advancing Mobile GUI Agent through Reinforcement Learning in Online Environment cites this paper.

MobileGUI-RL: Advancing Mobile GUI Agent through Reinforcement Learning in Online Environment ARPO:End-to-End Policy Optimization for GUI Agents with Experience Replay

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T19:25:14.276099Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:25:14.276099Z digest=sha256:d126d343a975b508aa9d573eadfca5c167ba418fb14303c824f568902a2fa871

Observation efa7606d-90a1-4c89-a24d-f755e76324f5 · inbound

UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning cites this paper.

UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning ARPO:End-to-End Policy Optimization for GUI Agents with Experience Replay

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-13T10:13:59.047730Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-13T10:13:58.774968Z digest=sha256:4ec8b455978aaf2bac4798cd25cadde101ddc555c9a2ba15202354e79cb4255b

Observation 0acd32bd-456b-4de0-87eb-b50696880523 · inbound

Plan Then Action:High-Level Planning Guidance Reinforcement Learning for LLM Reasoning cites this paper.

Plan Then Action:High-Level Planning Guidance Reinforcement Learning for LLM Reasoning ARPO:End-to-End Policy Optimization for GUI Agents with Experience Replay

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-04T12:52:27.271355Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:52:27.271355Z digest=sha256:23bf0ec1ca14c2bcd6734fb9a37198567469334399d04a67b20a9f7df9ef477e

Observation 61c88db6-3fd5-4294-852b-ee3cc6395ca1 · inbound

AutoTool: Dynamic Tool Selection and Integration for Agentic Reasoning cites this paper.

AutoTool: Dynamic Tool Selection and Integration for Agentic Reasoning ARPO:End-to-End Policy Optimization for GUI Agents with Experience Replay

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-03T16:30:39.322149Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T16:30:39.322149Z digest=sha256:855b77e86af89dbdba4c87c252021a6a7a580f4ad3dc525645786c59bbdca9b6

Observation f7a6fa83-59c1-42f8-bd82-ecc76dcd26e5 · inbound

Agentic Reasoning for Large Language Models cites this paper.

Agentic Reasoning for Large Language Models ARPO:End-to-End Policy Optimization for GUI Agents with Experience Replay

Reference 64

Resolution
verified exact
arxiv_id, observed 2026-05-17T15:14:25.866371Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-17T15:14:25.558878Z digest=sha256:16c1beb73ae4ab5a027f4ed44350399cb5ee1eddb491dc35be9b49dad4846d7f

Observation a00a8d54-94c6-4eff-8a4b-afedab8048e3 · inbound

Turing Test on Screen: A Benchmark for Mobile GUI Agent Humanization cites this paper.

Turing Test on Screen: A Benchmark for Mobile GUI Agent Humanization ARPO:End-to-End Policy Optimization for GUI Agents with Experience Replay

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-15T20:36:35.212903Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-15T20:35:17.075890Z digest=sha256:749473407cc6f6344b99d29d5f7d4a13d1efc9ce1a0e886bd549c8bd600a1b2d

Observation 32ddfcb8-f0f4-4f19-ad06-b0543b4f3018 · inbound

Faithful Mobile GUI Agents with Guided Advantage Estimator cites this paper.

Faithful Mobile GUI Agents with Guided Advantage Estimator ARPO:End-to-End Policy Optimization for GUI Agents with Experience Replay

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-11T16:41:24.147115Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-09T15:17:05.477466Z digest=sha256:e0f97766c0a03936ba8cbcfe0e5fee13df4c71e2441c51fa4b1c49be0c5a0d10

Observation 3ab69e72-872e-4417-b40f-66d6b72ad9eb · inbound

On Training Large Language Models for Long-Horizon Tasks: An Empirical Study of Horizon Length cites this paper.

On Training Large Language Models for Long-Horizon Tasks: An Empirical Study of Horizon Length ARPO:End-to-End Policy Optimization for GUI Agents with Experience Replay

Reference 66

Resolution
metadata mismatch
arxiv_id, observed 2026-05-09T06:40:40.641959Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-05-08T18:13:25.735085Z digest=sha256:f1a766df5f555d8bbb99da9c36d02c4ad4773edee7eeaef24d5137e018f5b3f2

Observation 61a0ea97-88b7-4880-83e2-222fca09fa1a · inbound

Learning to Explore: Scaling Agentic Reasoning via Exploration-Aware Policy Optimization cites this paper.

Learning to Explore: Scaling Agentic Reasoning via Exploration-Aware Policy Optimization ARPO:End-to-End Policy Optimization for GUI Agents with Experience Replay

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-12T03:01:17.873771Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-12T03:00:01.549385Z digest=sha256:6807870b07dfa729012b7d72be8ac6437165734580154ee82f90261e10849c7e

Observation e43a52a5-801e-44c0-a4e7-1f00e7d6cdf0 · inbound

Learning to Explore: Scaling Agentic Reasoning via Exploration-Aware Policy Optimization cites this paper.

Learning to Explore: Scaling Agentic Reasoning via Exploration-Aware Policy Optimization ARPO:End-to-End Policy Optimization for GUI Agents with Experience Replay

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-13T02:17:07.683646Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-13T00:58:25.205386Z digest=sha256:57cebcdad21e8cee0b0323277697f585956bd6d8ca7bb3becd8c375679c9d82f

Observation 494b1bdb-c798-4aa8-b384-8e948dd3bd4f · inbound

ToolCUA: Towards Optimal GUI-Tool Path Orchestration for Computer Use Agents cites this paper.

ToolCUA: Towards Optimal GUI-Tool Path Orchestration for Computer Use Agents ARPO:End-to-End Policy Optimization for GUI Agents with Experience Replay

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-13T03:52:12.365420Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-13T03:51:54.401200Z digest=sha256:966977cf34ec79a4bb07bf6f5683ffc561f74e00198b874aba189b0c81c72b13

Observation ad7335ec-cd17-4d1b-85cc-f1e71fcfc475 · inbound

RLVR without Ineffective Samples: Group Prioritized Off-Policy Optimization for LLM Reasoning cites this paper.

RLVR without Ineffective Samples: Group Prioritized Off-Policy Optimization for LLM Reasoning ARPO:End-to-End Policy Optimization for GUI Agents with Experience Replay

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-07-01T21:16:13.369302Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-06-28T17:25:50.758630Z digest=sha256:901147a7640af09e579ddf8d6a50a4ec0177871ddb37bdf1335052418dda9a06

Observation 44daa594-25cc-420c-a599-f689908c2cdb · inbound

StainFlow: Entity-Stain Tracking and Evidence Linking for Process Rewards in GUI Agents cites this paper.

StainFlow: Entity-Stain Tracking and Evidence Linking for Process Rewards in GUI Agents ARPO:End-to-End Policy Optimization for GUI Agents with Experience Replay

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-07-02T16:47:09.594291Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-06-27T22:24:41.353303Z digest=sha256:a9f197bb332c254a8b42d8394b478f897ce47c36d839fc6e855caf4f03446cb3

Observation 46e7f937-ef14-4065-9ae5-ae9c9c4aa368 · inbound

Modularized Reinforcement Learning on LLMs: From MDP Creation to Exploration and Learning cites this paper.

Modularized Reinforcement Learning on LLMs: From MDP Creation to Exploration and Learning ARPO:End-to-End Policy Optimization for GUI Agents with Experience Replay

Reference 125

Resolution
verified exact
arxiv_id, observed 2026-07-04T08:09:40.731617Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-06-26T12:15:08.304150Z digest=sha256:1625ec90adfc922bf73fa8139292d963aaf019cb7de64932263fc495c841bb18

Observation 3450fa5e-24d1-467e-99a6-dbf7f0c44049 · inbound

Learning with a Single Rollout via Monte Carlo Pass@k Critic cites this paper.

Learning with a Single Rollout via Monte Carlo Pass@k Critic ARPO:End-to-End Policy Optimization for GUI Agents with Experience Replay

Reference 31

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T20:00:08.409789Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-06-25T20:53:10.016447Z digest=sha256:2b7c00cf7517160f2a0f76891093ff7e99cd91d579bc307cf422594d7e1f6d2e

Observation 490492e7-74a4-411e-afa8-b66bcd8c2a8d · inbound

DocArena: Turning Raw Documents into Controllable Training Environments for Document Search Agents cites this paper.

DocArena: Turning Raw Documents into Controllable Training Environments for Document Search Agents ARPO:End-to-End Policy Optimization for GUI Agents with Experience Replay

Reference 32

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T12:53:26.532661Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-06-29T12:50:16.625077Z digest=sha256:4b820d619b2973fe196113a2e0e01a5ed6ef083d644f11ca35d3c44fcd432f70

Observation 6632b7f6-e440-45f9-bf76-c7f0b764d1da · inbound

Software Engineering for and with GUI Agent cites this paper.

Software Engineering for and with GUI Agent ARPO:End-to-End Policy Optimization for GUI Agents with Experience Replay

Reference 137

Resolution
unresolved
no resolver link, observed 2026-08-11T20:19:15.434119Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:19:15.434119Z digest=sha256:3ec031788534bfea9ab68857336df9ecfe46392cc8d78a42673c766201a6ab83