Pith. sign in

Paper Citation Record · LEDGER

ARPO:End-to-End Policy Optimization for GUI Agents with Experience Replay

As of 8 August 2026, this Paper Citation Record lists 29 of 29 outbound references and 16 inbound Pith citation observations for arXiv:2505.16282.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.16282 v1

Coverage vector

measured 29 of 29 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:07:10.796768Z

measured 45 of 45 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 16 of 16 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T19:25:14.276099Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T20:00:08.408321Z

Reference resolution

29 of 29 outbound references displayed

  • verified exact0
  • verified fuzzy5
  • unresolved24
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation a97a4621-6c2e-4838-8b44-aac68c14df56 · outbound

This paper cites Qwen2.5-VL Technical Report.

ARPO:End-to-End Policy Optimization for GUI Agents with Experience Replay Qwen2.5-VL Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T15:07:08.732949Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:07:08.732949Z digest=sha256:a4119363a3dd553e3d98f34436d33a87bf77305761b06563cfa2780767e34da9

Observation a5544cf9-7703-4840-aa66-c648472154de · outbound

This paper cites SeeClick: Harnessing GUI Grounding for Advanced Visual GUI Agents.

ARPO:End-to-End Policy Optimization for GUI Agents with Experience Replay SeeClick: Harnessing GUI Grounding for Advanced Visual GUI Agents

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T15:07:08.818285Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:07:08.818285Z digest=sha256:e55f173ce27a4d225444074d234ddd406195bbf86623852fb4f89354fdc02183

Observation 06a674f1-8a7e-48ab-8c65-6f05654887e2 · outbound

This paper cites KTO: Model Alignment as Prospect Theoretic Optimization.

ARPO:End-to-End Policy Optimization for GUI Agents with Experience Replay KTO: Model Alignment as Prospect Theoretic Optimization

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T15:07:08.943210Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:07:08.943210Z digest=sha256:2ad2ca77307521af50470f6aeefa7cf54b76d03b4d26f55b00f007c791954d15

Observation a204acfe-dfc4-4120-a7ae-98cec497d88e · outbound

This paper cites Navigating the Digital World as Humans Do: Universal Visual Grounding for GUI Agents.

ARPO:End-to-End Policy Optimization for GUI Agents with Experience Replay Navigating the Digital World as Humans Do: Universal Visual Grounding for GUI Agents

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T15:07:09.054423Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:07:09.054423Z digest=sha256:2c6fe8cd62d8b715d5fd2b362cce2c53374ccab81596b1dded5a5309e2e274c5

Observation 96a71639-ea78-4543-8235-922a76b611cc · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

ARPO:End-to-End Policy Optimization for GUI Agents with Experience Replay DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T15:07:09.148499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:07:09.148499Z digest=sha256:23d284df7c0430a9e3a5d8c9c0caff7c183bd7579bc258a55be5fb06ae1e7768

Observation f2923ec7-65ca-4e66-a110-9ab98fa6804c · outbound

This paper cites Cogagent: A visual language model for gui agents.

ARPO:End-to-End Policy Optimization for GUI Agents with Experience Replay Cogagent: A visual language model for gui agents

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:07:11.440915Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:07:09.257984Z digest=sha256:03057f6eb745b65474d2d52002dd659a886ec35a2a5c386654b0cde6685e8010

Observation 08626187-7624-473a-848e-7420015aef08 · outbound

This paper cites Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language Models.

ARPO:End-to-End Policy Optimization for GUI Agents with Experience Replay Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T15:07:09.358618Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:07:09.358618Z digest=sha256:8e33f1d26642baddf44cce2b9199160e9274d189222140de30f1b2696ddedad5

Observation 48d13b54-9277-4d20-9f3b-7d8563c4eae4 · outbound

This paper cites OpenAI o1 System Card.

ARPO:End-to-End Policy Optimization for GUI Agents with Experience Replay OpenAI o1 System Card

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T15:07:09.452790Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:07:09.452790Z digest=sha256:37c2c5cefbf655ea472d39cd116651fbc2e43456c5beef789b610048e4f9d4c0

Observation 4b5481a0-74d7-401d-881c-9dac23cee2e4 · outbound

This paper cites OmniACT: A Dataset and Benchmark for Enabling Multimodal Generalist Autonomous Agents for Desktop and Web.

ARPO:End-to-End Policy Optimization for GUI Agents with Experience Replay OmniACT: A Dataset and Benchmark for Enabling Multimodal Generalist Autonomous Agents for Desktop and Web

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T15:07:09.526745Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:07:09.526745Z digest=sha256:69acf94dace5a1c2da3c3f09c2a71877bd75c7e74b13dabf2faf3367929e8b6c

Observation e0a72d0c-1f2b-439e-b76a-0815a1f77501 · outbound

This paper cites Efficient memory management for large language model serving with pagedattention.

ARPO:End-to-End Policy Optimization for GUI Agents with Experience Replay Efficient memory management for large language model serving with pagedattention

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:07:11.374189Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:07:09.645013Z digest=sha256:f73de481631c5fa5833b55fe5cec62baafee5d49c48dbfd5afd7b0acfad6bdce

Observation 723a7e63-be07-4902-9360-6ab11888e498 · outbound

This paper cites Code-r1: Reproducing r1 for code with reliable rewards.

ARPO:End-to-End Policy Optimization for GUI Agents with Experience Replay Code-r1: Reproducing r1 for code with reliable rewards

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T15:07:09.737013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:07:09.737013Z digest=sha256:57dc42d8982a300e0c750c7de711dd0e0598cd385f299e1e610b3826dffaad3c

Observation 8ea43a1d-8027-47f5-b596-6acbc00c0eb8 · outbound

This paper cites Seg-Zero: Reasoning-Chain Guided Segmentation via Cognitive Reinforcement.

ARPO:End-to-End Policy Optimization for GUI Agents with Experience Replay Seg-Zero: Reasoning-Chain Guided Segmentation via Cognitive Reinforcement

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T15:07:09.853262Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:07:09.853262Z digest=sha256:47bf893079caac4bb87de7a5ecd8b51ded333518ab742e888975f271c922da2e

Observation bc7d8539-4150-48cb-832c-6b4548dd2e82 · outbound

This paper cites Decoupled Weight Decay Regularization.

ARPO:End-to-End Policy Optimization for GUI Agents with Experience Replay Decoupled Weight Decay Regularization

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T15:07:09.981894Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:07:09.981894Z digest=sha256:9051e4487067fd05183d7e43398081864b69482ef8b71155fbd8ae40a9a67332

Observation 2a897a41-6265-4c2c-abd7-a4289024e73f · outbound

This paper cites ScreenAgent: A Vision Language Model-driven Computer Control Agent.

ARPO:End-to-End Policy Optimization for GUI Agents with Experience Replay ScreenAgent: A Vision Language Model-driven Computer Control Agent

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T15:07:10.028219Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:07:10.028219Z digest=sha256:0815897b675399cc1cae10658d66bcab7102a136b471c2f795b42a78231c2928

Observation 53b21184-ee19-4e15-a7de-b8d3d4fa34d7 · outbound

This paper cites ToolRL: Reward is All Tool Learning Needs.

ARPO:End-to-End Policy Optimization for GUI Agents with Experience Replay ToolRL: Reward is All Tool Learning Needs

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T15:07:10.067180Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:07:10.067180Z digest=sha256:0ba4a23f63e0aeb20105066dbacb5ffe97eab5f6de81ab9ed86cd1ccc86ca356

Observation 202e5032-e30c-4814-905a-6ed48af0b141 · outbound

This paper cites UI-TARS: Pioneering Automated GUI Interaction with Native Agents.

ARPO:End-to-End Policy Optimization for GUI Agents with Experience Replay UI-TARS: Pioneering Automated GUI Interaction with Native Agents

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T15:07:10.120332Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:07:10.120332Z digest=sha256:0b4fa9b39df5769a7f9f288ac75e614c612a7a85d8d06684fdd08cf4bd018fb0

Observation 15b11a46-da31-4f47-aea5-c4c68fc8cadd · outbound

This paper cites Direct preference optimization: Your language model is secretly a reward model.

ARPO:End-to-End Policy Optimization for GUI Agents with Experience Replay Direct preference optimization: Your language model is secretly a reward model

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:07:11.288367Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:07:10.174192Z digest=sha256:c2572027461edc786841a461942e2d6adc3e32e9e991fddb7e3ff6ed43a18cc5

Observation 1f8adde8-ef64-4dfe-a2b9-59b519a037ef · outbound

This paper cites Proximal Policy Optimization Algorithms.

ARPO:End-to-End Policy Optimization for GUI Agents with Experience Replay Proximal Policy Optimization Algorithms

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T15:07:10.211575Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:07:10.211575Z digest=sha256:ab6a3bea11133074c729790dc08dbdc91ea1b4c75dffacb11c94000d15ad6490

Observation 1d28fa66-0dad-4959-9ec7-efc2a223b9c2 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

ARPO:End-to-End Policy Optimization for GUI Agents with Experience Replay DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T15:07:10.279353Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:07:10.279353Z digest=sha256:ac7e6ac61d54c6b2364b54de0831c8650e488c95c2acad933d99b71b6f13d1e4

Observation 440ba8ab-3650-4f73-bf52-0b10af21267c · outbound

This paper cites HybridFlow: A Flexible and Efficient RLHF Framework.

ARPO:End-to-End Policy Optimization for GUI Agents with Experience Replay HybridFlow: A Flexible and Efficient RLHF Framework

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T15:07:10.332778Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:07:10.332778Z digest=sha256:6639b6af91810f301775752e9914b9cf6ee8a94d8065d67f69104260e6ccfa4d

Observation 6ababca5-5736-47a2-924b-fd00f9d5b887 · outbound

This paper cites RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning.

ARPO:End-to-End Policy Optimization for GUI Agents with Experience Replay RAGEN: Understanding Self-Evolution in LLM Agents via Multi-Turn Reinforcement Learning

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T15:07:10.363216Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:07:10.363216Z digest=sha256:4ccdb4292ce822acbebb15b8aa1f711e10adb10507debe2899f125c88aacdba1

Observation eda3ee73-fe85-4711-a4cc-bec481c35515 · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models.

ARPO:End-to-End Policy Optimization for GUI Agents with Experience Replay Chain-of-thought prompting elicits reasoning in large language models

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:07:11.194077Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:07:10.403817Z digest=sha256:70318039500d09f27d59e35acb75999d8ad0b3f31251f678274c40930c996dbd

Observation 8f180272-3fd7-41f7-af45-ed5731c23050 · outbound

This paper cites OS-ATLAS: A Foundation Action Model for Generalist GUI Agents.

ARPO:End-to-End Policy Optimization for GUI Agents with Experience Replay OS-ATLAS: A Foundation Action Model for Generalist GUI Agents

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T15:07:10.444576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:07:10.444576Z digest=sha256:639ad51e94cfd96837562ddd58d10c5d92c29fdc54254b9aba2aaf5acdd8fd89

Observation 58f4df80-56d4-438d-8f34-5af8fc91a67a · outbound

This paper cites GUI-R1 : A Generalist R1-Style Vision-Language Action Model For GUI Agents.

ARPO:End-to-End Policy Optimization for GUI Agents with Experience Replay GUI-R1 : A Generalist R1-Style Vision-Language Action Model For GUI Agents

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T15:07:10.491309Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:07:10.491309Z digest=sha256:c4d230aa6095ff25083b87b6ab40906b3d6cdd6aefb067077aa651543126d719

Observation f4d6f9ea-76b2-4537-a06c-9ecfa4747bd3 · outbound

This paper cites Osworld: Benchmarking multimodal agents for open-ended tasks in real computer environments.

ARPO:End-to-End Policy Optimization for GUI Agents with Experience Replay Osworld: Benchmarking multimodal agents for open-ended tasks in real computer environments

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:07:11.128405Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T15:07:10.537390Z digest=sha256:abf3cf3b210f1f64828faba677166f99687dd060751eb225b1bf8635d42b6a95

Observation ca674936-63b4-4926-aa11-f4c6420fc5b6 · outbound

This paper cites Aguvis: Unified Pure Vision Agents for Autonomous GUI Interaction.

ARPO:End-to-End Policy Optimization for GUI Agents with Experience Replay Aguvis: Unified Pure Vision Agents for Autonomous GUI Interaction

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T15:07:10.588099Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:07:10.588099Z digest=sha256:c3540e1a9c7eccc67eed4f85826cfd792a23dc492ecdc2f94b9ab2834fcf41a3

Observation a363f308-e927-4e0c-9fbf-e202d39c717e · outbound

This paper cites Aria-UI: Visual Grounding for GUI Instructions.

ARPO:End-to-End Policy Optimization for GUI Agents with Experience Replay Aria-UI: Visual Grounding for GUI Instructions

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T15:07:10.649086Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:07:10.649086Z digest=sha256:3516a04a3f1154b3d9ef55a0809dc0ad92aeda3d39ad4274c8f627c93df0bbce

Observation f9423a99-3bc6-4f52-a8cb-ee85a07db547 · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

ARPO:End-to-End Policy Optimization for GUI Agents with Experience Replay DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T15:07:10.758735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:07:10.758735Z digest=sha256:aeb79a1b23ba16dfc629e44cc9606a9caeca8e8e54ebf364ffe26b791793c58a

Observation b169f929-fb11-4c83-8be7-a284473d8032 · outbound

This paper cites SWEET-RL: Training Multi-Turn LLM Agents on Collaborative Reasoning Tasks.

ARPO:End-to-End Policy Optimization for GUI Agents with Experience Replay SWEET-RL: Training Multi-Turn LLM Agents on Collaborative Reasoning Tasks

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T15:07:10.796768Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:07:10.796768Z digest=sha256:2d3e561c8b3cee5d0afb88af5f25d83d8bd76fe4d6acfa5ff568fe0b73cb6b19

Pith citing papers

Observation c7df2476-8c0a-48d9-8a9f-2cf60b09acd7 · inbound

MobileGUI-RL: Advancing Mobile GUI Agent through Reinforcement Learning in Online Environment cites this paper.

MobileGUI-RL: Advancing Mobile GUI Agent through Reinforcement Learning in Online Environment ARPO:End-to-End Policy Optimization for GUI Agents with Experience Replay

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T19:25:14.276099Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:25:14.276099Z digest=sha256:0939f738ccf42ca8d913f76bfadb020b8e7898b96844345dc8c8e2a385bb7f14

Observation efa7606d-90a1-4c89-a24d-f755e76324f5 · inbound

UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning cites this paper.

UI-TARS-2 Technical Report: Advancing GUI Agent with Multi-Turn Reinforcement Learning ARPO:End-to-End Policy Optimization for GUI Agents with Experience Replay

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-13T10:13:59.047730Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-13T10:13:58.774968Z digest=sha256:12e62b15cef3a5a1ceefc2e512bcd825319995a171acb8cbba7766ac9a8ffaba

Observation 0acd32bd-456b-4de0-87eb-b50696880523 · inbound

Plan Then Action:High-Level Planning Guidance Reinforcement Learning for LLM Reasoning cites this paper.

Plan Then Action:High-Level Planning Guidance Reinforcement Learning for LLM Reasoning ARPO:End-to-End Policy Optimization for GUI Agents with Experience Replay

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-04T12:52:27.271355Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:52:27.271355Z digest=sha256:937e646c4c28797bb3d7e85852caae1036451ae5376e0483a059f341d7dd3cf2

Observation 61c88db6-3fd5-4294-852b-ee3cc6395ca1 · inbound

AutoTool: Dynamic Tool Selection and Integration for Agentic Reasoning cites this paper.

AutoTool: Dynamic Tool Selection and Integration for Agentic Reasoning ARPO:End-to-End Policy Optimization for GUI Agents with Experience Replay

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-03T16:30:39.322149Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T16:30:39.322149Z digest=sha256:02b414efdd5a7f2727cd5badb0856f5baeb414ae31c83c2f23e79175fadfb1f3

Observation f7a6fa83-59c1-42f8-bd82-ecc76dcd26e5 · inbound

Agentic Reasoning for Large Language Models cites this paper.

Agentic Reasoning for Large Language Models ARPO:End-to-End Policy Optimization for GUI Agents with Experience Replay

Reference 64

Resolution
verified exact
arxiv_id, observed 2026-05-17T15:14:25.866371Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-17T15:14:25.558878Z digest=sha256:1e4e4d9c0c0410ee7dfebe06cdfdd04133fad21e51551d09bc465fc89b0a4258

Observation a00a8d54-94c6-4eff-8a4b-afedab8048e3 · inbound

Turing Test on Screen: A Benchmark for Mobile GUI Agent Humanization cites this paper.

Turing Test on Screen: A Benchmark for Mobile GUI Agent Humanization ARPO:End-to-End Policy Optimization for GUI Agents with Experience Replay

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-15T20:36:35.212903Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-15T20:35:17.075890Z digest=sha256:a2c415d632e3a0c486ce10dfdcd39c8c3622b11722407d04e0d6022f490a698e

Observation 32ddfcb8-f0f4-4f19-ad06-b0543b4f3018 · inbound

Faithful Mobile GUI Agents with Guided Advantage Estimator cites this paper.

Faithful Mobile GUI Agents with Guided Advantage Estimator ARPO:End-to-End Policy Optimization for GUI Agents with Experience Replay

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-11T16:41:24.147115Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-09T15:17:05.477466Z digest=sha256:7c77a94921ef6c32d8a6a582833ac58e7e7e7a9dda18194b6c73154eed98185f

Observation 3ab69e72-872e-4417-b40f-66d6b72ad9eb · inbound

On Training Large Language Models for Long-Horizon Tasks: An Empirical Study of Horizon Length cites this paper.

On Training Large Language Models for Long-Horizon Tasks: An Empirical Study of Horizon Length ARPO:End-to-End Policy Optimization for GUI Agents with Experience Replay

Reference 66

Resolution
metadata mismatch
arxiv_id, observed 2026-05-09T06:40:40.641959Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-08T18:13:25.735085Z digest=sha256:6f7e31d453bc3f8501391c97970cb4e85108a41735aac27f27e814ee887a8919

Observation 61a0ea97-88b7-4880-83e2-222fca09fa1a · inbound

Learning to Explore: Scaling Agentic Reasoning via Exploration-Aware Policy Optimization cites this paper.

Learning to Explore: Scaling Agentic Reasoning via Exploration-Aware Policy Optimization ARPO:End-to-End Policy Optimization for GUI Agents with Experience Replay

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-12T03:01:17.873771Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-12T03:00:01.549385Z digest=sha256:9eddc935a250d05a38362b754a92af112cdb146551032859f3a73b6f728708f5

Observation e43a52a5-801e-44c0-a4e7-1f00e7d6cdf0 · inbound

Learning to Explore: Scaling Agentic Reasoning via Exploration-Aware Policy Optimization cites this paper.

Learning to Explore: Scaling Agentic Reasoning via Exploration-Aware Policy Optimization ARPO:End-to-End Policy Optimization for GUI Agents with Experience Replay

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-13T02:17:07.683646Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-13T00:58:25.205386Z digest=sha256:8c40de9301ee952cec1f7abd8770c561350228c194707748600b71307c149ee0

Observation 494b1bdb-c798-4aa8-b384-8e948dd3bd4f · inbound

ToolCUA: Towards Optimal GUI-Tool Path Orchestration for Computer Use Agents cites this paper.

ToolCUA: Towards Optimal GUI-Tool Path Orchestration for Computer Use Agents ARPO:End-to-End Policy Optimization for GUI Agents with Experience Replay

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-13T03:52:12.365420Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-13T03:51:54.401200Z digest=sha256:9a81fc4a5a118709a03cdb9c7a3d842313d76038dac45bf94b52ae8fe90604bd

Observation ad7335ec-cd17-4d1b-85cc-f1e71fcfc475 · inbound

RLVR without Ineffective Samples: Group Prioritized Off-Policy Optimization for LLM Reasoning cites this paper.

RLVR without Ineffective Samples: Group Prioritized Off-Policy Optimization for LLM Reasoning ARPO:End-to-End Policy Optimization for GUI Agents with Experience Replay

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-07-01T21:16:13.369302Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-28T17:25:50.758630Z digest=sha256:91cf06539fb4e9c0b7189845031dc7807536f8b9d0c1bd1346fce0d6f3827c4a

Observation 44daa594-25cc-420c-a599-f689908c2cdb · inbound

StainFlow: Entity-Stain Tracking and Evidence Linking for Process Rewards in GUI Agents cites this paper.

StainFlow: Entity-Stain Tracking and Evidence Linking for Process Rewards in GUI Agents ARPO:End-to-End Policy Optimization for GUI Agents with Experience Replay

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-07-02T16:47:09.594291Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-27T22:24:41.353303Z digest=sha256:974c12e98de41199142ad056f6105c116eb7ee42547f01f7814da1c9b0492cc1

Observation 46e7f937-ef14-4065-9ae5-ae9c9c4aa368 · inbound

Modularized Reinforcement Learning on LLMs: From MDP Creation to Exploration and Learning cites this paper.

Modularized Reinforcement Learning on LLMs: From MDP Creation to Exploration and Learning ARPO:End-to-End Policy Optimization for GUI Agents with Experience Replay

Reference 125

Resolution
verified exact
arxiv_id, observed 2026-07-04T08:09:40.731617Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-26T12:15:08.304150Z digest=sha256:4f6b1f80c2add2079fe69f022e338abfbcd5bb2d0f35e28477172e1e94e5e24b

Observation 3450fa5e-24d1-467e-99a6-dbf7f0c44049 · inbound

Learning with a Single Rollout via Monte Carlo Pass@k Critic cites this paper.

Learning with a Single Rollout via Monte Carlo Pass@k Critic ARPO:End-to-End Policy Optimization for GUI Agents with Experience Replay

Reference 31

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T20:00:08.409789Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-25T20:53:10.016447Z digest=sha256:72ebbef4542475fb66f3650a556aaef8d2d2f765b2351505cb6550afa9a61a36

Observation 490492e7-74a4-411e-afa8-b66bcd8c2a8d · inbound

DocArena: Turning Raw Documents into Controllable Training Environments for Document Search Agents cites this paper.

DocArena: Turning Raw Documents into Controllable Training Environments for Document Search Agents ARPO:End-to-End Policy Optimization for GUI Agents with Experience Replay

Reference 32

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T12:53:26.532661Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-29T12:50:16.625077Z digest=sha256:c4eb1f33b0b47be659b7dfc9b97ce1f39956b8ad9c4a3a496f11fb404d831f61