Pith. sign in

Paper Citation Record · LEDGER

ProgRM: Build Better GUI Agents with Progress Rewards

As of 15 August 2026, this Paper Citation Record lists 45 of 45 outbound references and 5 inbound Pith citation observations for arXiv:2505.18121.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.18121 v1

Coverage vector

measured 45 of 45 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:37:49.267563Z

measured 50 of 50 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 5 of 5 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-28T19:26:37.281563Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T16:47:09.687885Z

Reference resolution

45 of 45 outbound references displayed

  • verified exact1
  • verified fuzzy11
  • unresolved33
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation aec7cd57-76d8-4c11-a7ff-73119782bc71 · outbound

This paper cites Digi-q: Learning vlm q-value functions for training device-control agents.

ProgRM: Build Better GUI Agents with Progress Rewards Digi-q: Learning vlm q-value functions for training device-control agents

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:37:51.949072Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:37:45.129305Z digest=sha256:f511dd703f3efd77589419c6aaf41b38f01a72d4266202b181168c69517d089e

Observation f7254977-e58d-4cd0-abab-2582825f887b · outbound

This paper cites Digirl: Training in-the-wild device-control agents with autonomous re- inforcement learning.

ProgRM: Build Better GUI Agents with Progress Rewards Digirl: Training in-the-wild device-control agents with autonomous re- inforcement learning

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:37:51.762133Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:37:45.198300Z digest=sha256:7605cdee0b5be2e102b3417dfd2bf74d84852cd7bd272414123325cb6371276b

Observation a5895103-8609-41bb-847b-5452b5004258 · outbound

This paper cites Learning about progress from experts.

ProgRM: Build Better GUI Agents with Progress Rewards Learning about progress from experts

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:37:51.581616Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:37:45.251487Z digest=sha256:4dd98534417d38b786c05778cf477d0d7322666e5ecafb66e1c82821161c6039

Observation 879f22e9-00ab-40d8-9a23-badacf93344a · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

ProgRM: Build Better GUI Agents with Progress Rewards Training Verifiers to Solve Math Word Problems

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:45.337260Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:45.337260Z digest=sha256:da60d5423bf50ba5d5fb38f6a9f18ba3f2f7e4bb3ec25b7f76a7d79ae3d934cf

Observation 53605823-cbb5-4d82-80f2-6a76092eb65e · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

ProgRM: Build Better GUI Agents with Progress Rewards DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:45.419221Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:45.419221Z digest=sha256:28c8b737ae3bb4a261d4189aba438acc0fb47098255d3e0e9bced1c53b926e71

Observation 9bb8ec98-2eac-4dc7-8c12-73a5d9dd3a60 · outbound

This paper cites Dungeons and data: A large-scale nethack dataset.

ProgRM: Build Better GUI Agents with Progress Rewards Dungeons and data: A large-scale nethack dataset

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:37:51.400251Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:37:45.508048Z digest=sha256:f11ffbd6f0305a93b5905ff1dcc313925b0ae209cb3588657f13eeafa9ff28c3

Observation e7dc26f3-39d4-42cb-91c6-261d8a91c241 · outbound

This paper cites REINFORCE++: Stabilizing Critic-Free Policy Optimization with Global Advantage Normalization.

ProgRM: Build Better GUI Agents with Progress Rewards REINFORCE++: Stabilizing Critic-Free Policy Optimization with Global Advantage Normalization

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:45.597283Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:45.597283Z digest=sha256:0e3ff30d75dcb79e10f4bf64b6bc44dc7622342c01fceeebf9092244fe6e8091

Observation 58e481f7-a907-4cbb-9051-e260b61b71d4 · outbound

This paper cites Exploring Expert Failures Improves LLM Agent Tuning.

ProgRM: Build Better GUI Agents with Progress Rewards Exploring Expert Failures Improves LLM Agent Tuning

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-08-07T14:37:50.068075Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:37:45.644348Z digest=sha256:fe199450b8388984f387804edc4f2843790a5f687200f4be94178953d92ad246

Observation a4ae53f1-e8c2-4bcc-ac12-def7e240e590 · outbound

This paper cites Making language models better reasoners with step-aware verifier.

ProgRM: Build Better GUI Agents with Progress Rewards Making language models better reasoners with step-aware verifier

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:45.740939Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:45.740939Z digest=sha256:6cbac5a1b391fd1ba95d78c0f1bb894b50b74759cc2df2d1b0e1d2f69b22d487

Observation 8b426dbb-a6d4-449f-b9d5-68468ac898a7 · outbound

This paper cites Let’s verify step by step.

ProgRM: Build Better GUI Agents with Progress Rewards Let’s verify step by step

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:45.786546Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:45.786546Z digest=sha256:c4543de146eba0d72d3fa8d94c466f7371905e673f6fad953238cbdc397cc8ab

Observation 4a6e6515-a8a4-4cac-a2c8-9cca2eecd7b2 · outbound

This paper cites UI-R1: Enhancing Efficient Action Prediction of GUI Agents by Reinforcement Learning.

ProgRM: Build Better GUI Agents with Progress Rewards UI-R1: Enhancing Efficient Action Prediction of GUI Agents by Reinforcement Learning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:45.862230Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:45.862230Z digest=sha256:92d0dbf1ba16c5d3eadaaa8aa3350ce8c5ce2581fc7743e8961f0b2e9bbe5682

Observation 50f6f860-4632-4c42-8826-d133a6404fdb · outbound

This paper cites NNetNav: Unsupervised Learning of Browser Agents Through Environment Interaction in the Wild.

ProgRM: Build Better GUI Agents with Progress Rewards NNetNav: Unsupervised Learning of Browser Agents Through Environment Interaction in the Wild

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:45.943894Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:45.943894Z digest=sha256:76b9efb7a9d789c62e93f8a500a9493ab7fcbb83a0722562f2e6bf0361168149

Observation 47146464-9f0a-4fd5-8d7f-bafd72db849c · outbound

This paper cites Xu, Aman Madaan, Jiarui Liu, Robert Lo, Abishek Sridhar, Sudipta Sengupta, Dan Roth, Graham Neubig, and Shuyan Zhou.

ProgRM: Build Better GUI Agents with Progress Rewards Xu, Aman Madaan, Jiarui Liu, Robert Lo, Abishek Sridhar, Sudipta Sengupta, Dan Roth, Graham Neubig, and Shuyan Zhou

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:37:51.202532Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:37:46.013705Z digest=sha256:75c1ceb3188a02face5e0fc45f4f2bc7d29dd0b718a17add1e4aea0600f3b806

Observation f3f932b9-fe28-43b6-9684-dfa389c944ac · outbound

This paper cites Explorer: Scaling Exploration-driven Web Trajectory Synthesis for Multimodal Web Agents.

ProgRM: Build Better GUI Agents with Progress Rewards Explorer: Scaling Exploration-driven Web Trajectory Synthesis for Multimodal Web Agents

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:46.119179Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:46.119179Z digest=sha256:c858264c5c05e372c2a5ef2adf0e656a84082bb06e76cd916bf43f3d8cdf022e

Observation b5b3b549-039f-4284-99eb-97810c3fe91c · outbound

This paper cites Autonomous Evaluation and Refinement of Digital Agents.

ProgRM: Build Better GUI Agents with Progress Rewards Autonomous Evaluation and Refinement of Digital Agents

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:46.228152Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:46.228152Z digest=sha256:bac4ff4c485759dd29deeecad8dc1c7152c9c4277c4414242f0e155817f75501

Observation 4efd62e6-4159-4d0d-b049-b694982913da · outbound

This paper cites WebRL: Training LLM Web Agents via Self-Evolving Online Curriculum Reinforcement Learning.

ProgRM: Build Better GUI Agents with Progress Rewards WebRL: Training LLM Web Agents via Self-Evolving Online Curriculum Reinforcement Learning

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:46.329762Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:46.329762Z digest=sha256:e2ee1403e219bc2771984f1d4b1f5cd8343cf3d7796c8fe066c5ef8aab4f46a8

Observation 8e89bd61-c05a-4608-9851-7de29189cb15 · outbound

This paper cites UI-TARS: Pioneering Automated GUI Interaction with Native Agents.

ProgRM: Build Better GUI Agents with Progress Rewards UI-TARS: Pioneering Automated GUI Interaction with Native Agents

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:46.431629Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:46.431629Z digest=sha256:44fa367f782a63df03bf208ba2bfcf05b47f17dad803b391e2d665d2d1091290

Observation 18d80b30-3761-4e5e-b55e-c5e36827d8ce · outbound

This paper cites Optimizing Test-Time Compute via Meta Reinforcement Fine-Tuning.

ProgRM: Build Better GUI Agents with Progress Rewards Optimizing Test-Time Compute via Meta Reinforcement Fine-Tuning

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:46.502834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:46.502834Z digest=sha256:9f563dd19c221602e174fd3160af7533c833b7dbe56ac874d7daea590a2dd5fc

Observation d4bdfdd7-bfd0-451a-a2a5-ffbd21fb45d1 · outbound

This paper cites Sentence-bert: Sentence embeddings using siamese bert- networks.

ProgRM: Build Better GUI Agents with Progress Rewards Sentence-bert: Sentence embeddings using siamese bert- networks

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:37:51.001358Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:37:46.642165Z digest=sha256:4c7f4df01d98aa639bd90b6b094d9e0b6a9ebffd7413ee7e3c4ed5c52f5aa3b8

Observation c464d765-6783-4c58-ba4f-a10440cdcbe1 · outbound

This paper cites Step: Stacked llm policies for web actions.

ProgRM: Build Better GUI Agents with Progress Rewards Step: Stacked llm policies for web actions

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:37:50.788218Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:37:46.921764Z digest=sha256:7c913ff26c492a441820c4aa556718a1cc12143746d78c5a2d7cae6ef2ceaf9d

Observation c9cf3a05-12da-493e-98f2-dce10c1ab90d · outbound

This paper cites R1-Searcher: Incentivizing the Search Capability in LLMs via Reinforcement Learning.

ProgRM: Build Better GUI Agents with Progress Rewards R1-Searcher: Incentivizing the Search Capability in LLMs via Reinforcement Learning

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:47.033881Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:47.033881Z digest=sha256:e49533c50450c8e66f526d8ab215c30035689c438506d1de20800e117c731709

Observation 3eb32102-56d0-490f-a9e5-92e377d9949e · outbound

This paper cites Learn-by-interact: A Data-Centric Framework for Self-Adaptive Agents in Realistic Environments.

ProgRM: Build Better GUI Agents with Progress Rewards Learn-by-interact: A Data-Centric Framework for Self-Adaptive Agents in Realistic Environments

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:47.108403Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:47.108403Z digest=sha256:1beddcbbafa8084d6866fbef1bcca08a077da1472ae20b2dfaa85abf956550e4

Observation 486f8fcf-9f3c-4375-be1f-f7656a1a18bf · outbound

This paper cites OS-Genesis: Automating GUI Agent Trajectory Construction via Reverse Task Synthesis.

ProgRM: Build Better GUI Agents with Progress Rewards OS-Genesis: Automating GUI Agent Trajectory Construction via Reverse Task Synthesis

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:47.185783Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:47.185783Z digest=sha256:392c48f123bef9c2ba1c5150ea9897421e59db40651ca6ed1f5e0d66d4f59498

Observation c6421f46-7dfa-4f7a-8464-a8c8d1c8b002 · outbound

This paper cites Mobile-Agent: Autonomous Multi-Modal Mobile Device Agent with Visual Perception.

ProgRM: Build Better GUI Agents with Progress Rewards Mobile-Agent: Autonomous Multi-Modal Mobile Device Agent with Visual Perception

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:47.238146Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:47.238146Z digest=sha256:8222aa73145a603a3a759f753cdbea9f3fe85f36e46ced723caed4e0f41cd679

Observation 19a62709-4c66-4c0c-b607-d22e1aa6c55a · outbound

This paper cites Math-shepherd: Verify and reinforce llms step-by-step without human annotations.

ProgRM: Build Better GUI Agents with Progress Rewards Math-shepherd: Verify and reinforce llms step-by-step without human annotations

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:47.380315Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:47.380315Z digest=sha256:666449c7b407ad892ec4e28a6deb42e519a18bf2e4fa9f228acf81b4ecd8236e

Observation 053bc850-e181-465d-8960-9813e0b5f08f · outbound

This paper cites DistRL: An Asynchronous Distributed Reinforcement Learning Framework for On-Device Control Agents.

ProgRM: Build Better GUI Agents with Progress Rewards DistRL: An Asynchronous Distributed Reinforcement Learning Framework for On-Device Control Agents

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:47.485240Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:47.485240Z digest=sha256:2a1a5afeaf3ffd105e95d717f7dc194ce24d9a2db2b20bb8f3835a0a68e1fbf7

Observation 8029fb17-4eaa-44da-a0b9-37a8b23ba8c9 · outbound

This paper cites Reinforcing Language Agents via Policy Optimization with Action Decomposition.

ProgRM: Build Better GUI Agents with Progress Rewards Reinforcing Language Agents via Policy Optimization with Action Decomposition

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:47.590688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:47.590688Z digest=sha256:6889615a66f3b7ee7eb55960865688181f2546da93a0485bc2c136e5a35976c1

Observation a3e49663-2092-4537-8873-2260e3312495 · outbound

This paper cites GUI-R1 : A Generalist R1-Style Vision-Language Action Model For GUI Agents.

ProgRM: Build Better GUI Agents with Progress Rewards GUI-R1 : A Generalist R1-Style Vision-Language Action Model For GUI Agents

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:47.700415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:47.700415Z digest=sha256:feebc882a6e18e7d2e31ef474b2c59707f0ec47be4ae279d9aee47ca0d4cdcd4

Observation 19c05f48-6249-4353-acce-6092cd98e096 · outbound

This paper cites Logic-RL: Unleashing LLM Reasoning with Rule-Based Reinforcement Learning.

ProgRM: Build Better GUI Agents with Progress Rewards Logic-RL: Unleashing LLM Reasoning with Rule-Based Reinforcement Learning

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:47.804784Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:47.804784Z digest=sha256:c38c67edf64905ad352f54dc1ea43d101354555e3c1d38c041c56012a0346baf

Observation 5a6a36ef-b5cf-4007-84c6-d943abba95e4 · outbound

This paper cites Aguvis: Unified Pure Vision Agents for Autonomous GUI Interaction.

ProgRM: Build Better GUI Agents with Progress Rewards Aguvis: Unified Pure Vision Agents for Autonomous GUI Interaction

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:47.937341Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:47.937341Z digest=sha256:26e66544c41a7d134baad15e1be9d0eae444e8c9c15396c3cc178c25af7760f0

Observation 835c7d55-3d95-4219-9486-da9488b80267 · outbound

This paper cites Agenttrek: Agent trajectory synthesis via guiding replay with web tutorials.

ProgRM: Build Better GUI Agents with Progress Rewards Agenttrek: Agent trajectory synthesis via guiding replay with web tutorials

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:37:50.672358Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:37:48.039904Z digest=sha256:eb211f9621fd7534ccc83205aae2afe21bd415d18c0a40c19edb0955f9ba0d1c

Observation 3450faf6-7e13-4dc7-8f4e-c5518ff85d63 · outbound

This paper cites Qwen2.5 Technical Report.

ProgRM: Build Better GUI Agents with Progress Rewards Qwen2.5 Technical Report

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:48.119240Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:48.119240Z digest=sha256:a407e6a69f08078d260634510563a1950baa078e5142c32fc4b0f5f5e58ecca2

Observation bd0674c3-a3c1-4f74-af92-f26678f2827c · outbound

This paper cites OVM, Outcome-supervised Value Models for Planning in Mathematical Reasoning.

ProgRM: Build Better GUI Agents with Progress Rewards OVM, Outcome-supervised Value Models for Planning in Mathematical Reasoning

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:48.213972Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:48.213972Z digest=sha256:c80eaefb58b192dcfdb656773c9412b24c9627d6a8bb21cc5faaf681bec5221e

Observation 750b3d33-83c5-4606-8ed4-199bcbfa0d59 · outbound

This paper cites Free Process Rewards without Process Labels.

ProgRM: Build Better GUI Agents with Progress Rewards Free Process Rewards without Process Labels

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:48.320653Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:48.320653Z digest=sha256:a41fbbe7a683d59ee4ab36033fcae5a009d96c9f26b5314cd160c666fec54a88

Observation 5dc867a9-8a39-4ef1-ae06-9424c5cdf5d8 · outbound

This paper cites UFO: A UI-Focused Agent for Windows OS Interaction.

ProgRM: Build Better GUI Agents with Progress Rewards UFO: A UI-Focused Agent for Windows OS Interaction

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:48.529789Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:48.529789Z digest=sha256:01f7abea20f3988c181ce84d752f94c0ff6d9a218b09ab49c56175e0be7d9f60

Observation 8db64b3d-63b0-47ef-b6ef-d92909160a9a · outbound

This paper cites Appagent: Multimodal agents as smartphone users.

ProgRM: Build Better GUI Agents with Progress Rewards Appagent: Multimodal agents as smartphone users

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:48.664282Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:48.664282Z digest=sha256:2d5dd9b36dfacd5d491c6199db1e26efac32c9acbd9828ce7b09cdaeaf6e3a40

Observation d601e41c-ad3e-4be6-92fe-9e5ad8d03326 · outbound

This paper cites Large language models are semi-parametric reinforcement learning agents.

ProgRM: Build Better GUI Agents with Progress Rewards Large language models are semi-parametric reinforcement learning agents

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:37:50.549721Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:37:48.771877Z digest=sha256:b3c88d43e45754574c2cae8327a07234866f48e493e937fe69eef2a64676cbda

Observation 32ad9c9e-7ec1-40c1-99e7-bdf899b9763d · outbound

This paper cites Mobile-Env: Building Qualified Evaluation Benchmarks for LLM-GUI Interaction.

ProgRM: Build Better GUI Agents with Progress Rewards Mobile-Env: Building Qualified Evaluation Benchmarks for LLM-GUI Interaction

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:48.872226Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:48.872226Z digest=sha256:e0ff13524c8c6e2c5db69d9cd08f62226e858891a2edbdc818b53fb78f65103b

Observation 129c17e0-bb24-4cf3-ad7e-bbdcd6ef5525 · outbound

This paper cites The Lessons of Developing Process Reward Models in Mathematical Reasoning.

ProgRM: Build Better GUI Agents with Progress Rewards The Lessons of Developing Process Reward Models in Mathematical Reasoning

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:48.959454Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:48.959454Z digest=sha256:cd4938360ed9e71dfdbd5f90235865197939a34b9cd0a39f495a6ea213901e83

Observation a0dee5fd-9a2b-44d8-a18e-b0cb28957dc1 · outbound

This paper cites Gpt-4v(ision) is a generalist web agent, if grounded.

ProgRM: Build Better GUI Agents with Progress Rewards Gpt-4v(ision) is a generalist web agent, if grounded

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:37:50.386446Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:37:49.041108Z digest=sha256:e47dafd032d6587650edd178c7cf4584fa81a91ba8ae3fe38e270bd80dafaf51

Observation 54802348-9e11-49bd-a93e-166b08e21376 · outbound

This paper cites VEM: Environment-Free Exploration for Training GUI Agent with Value Environment Model.

ProgRM: Build Better GUI Agents with Progress Rewards VEM: Environment-Free Exploration for Training GUI Agent with Value Environment Model

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:49.160588Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:49.160588Z digest=sha256:d524cbe4a8faf0e20dc70278535444e9b88a11ffe8e532533263e3a7a4f58147

Observation 901c9048-5ede-4c28-ae38-b51dac1b22fd · outbound

This paper cites Proposer-Agent-Evaluator(PAE): Autonomous Skill Discovery For Foundation Model Internet Agents.

ProgRM: Build Better GUI Agents with Progress Rewards Proposer-Agent-Evaluator(PAE): Autonomous Skill Discovery For Foundation Model Internet Agents

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:49.207033Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:49.207033Z digest=sha256:8c03e61c987e8c62c1f2f9a2b616155ba8c38b92f7b6a7ea29b1c8e741594f6a

Observation b897fa9b-2afd-4675-9485-e29701257712 · outbound

This paper cites prototype.

ProgRM: Build Better GUI Agents with Progress Rewards prototype

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:37:50.229956Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T14:37:49.267563Z digest=sha256:f5ac9b715c74cef999045092d36c845cb52ab70bb2f9c3ab65b515446bf9c3eb

Observation 853b27f3-2c39-475c-a86a-ae09ebfd1b8f · outbound

This paper cites URL https://doi.org/10.18653/v1/D19-1410.

ProgRM: Build Better GUI Agents with Progress Rewards URL https://doi.org/10.18653/v1/D19-1410

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:46.769331Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:46.769331Z digest=sha256:ee5d7169185d953b8f9c7e0fb0cb100833fbe7516b00c2c3dfc6534e912709f9

Observation 9d04d547-2c34-45ac-a980-11f7e7716d76 · outbound

This paper cites Free Process Rewards without Process Labels.

ProgRM: Build Better GUI Agents with Progress Rewards Free Process Rewards without Process Labels

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T14:37:48.423228Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:37:48.423228Z digest=sha256:b296ebf7a5a3d9318c2c373a64eb5bb32aac14852e8fe45b43d41827307eed26

Pith citing papers

Observation 32164952-9169-4ee7-af71-03c15a4fad7f · inbound

UI-Oceanus: Scaling GUI Agents with Synthetic Environmental Dynamics cites this paper.

UI-Oceanus: Scaling GUI Agents with Synthetic Environmental Dynamics ProgRM: Build Better GUI Agents with Progress Rewards

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-05-16T03:00:31.666131Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-05-16T02:59:27.807789Z digest=sha256:9e7d1723e3b94c0d9121131a7270278a3d97131d2098e8611d396a30cbe5038b

Observation 81223177-eb8d-47d6-92ed-1976515667bf · inbound

GUI Agents with Reinforcement Learning: Toward Digital Inhabitants cites this paper.

GUI Agents with Reinforcement Learning: Toward Digital Inhabitants ProgRM: Build Better GUI Agents with Progress Rewards

Reference 83

Resolution
verified exact
arxiv_id, observed 2026-05-12T10:31:29.615649Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-07T05:48:00.486572Z digest=sha256:bee69908940c29309b667d8bc0ddca26f2abf64a5567178caa8a225789ef0ac7

Observation 331eab4f-af11-4384-8355-7588e38095c9 · inbound

Securing Computer-Use Agents: A Unified Architecture-Lifecycle Framework for Deployment-Grounded Reliability cites this paper.

Securing Computer-Use Agents: A Unified Architecture-Lifecycle Framework for Deployment-Grounded Reliability ProgRM: Build Better GUI Agents with Progress Rewards

Reference 114

Resolution
verified exact
arxiv_id, observed 2026-05-11T04:35:56.866360Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-11T01:15:19.239355Z digest=sha256:546c00f81a43764693fb23081b9091e1575576750e36f9c46d90c37d1371afe0

Observation 106cf729-2ff8-4143-ace8-1a0164978cba · inbound

Decomposed On-Policy Distillation for Vision-Language Reasoning: Steering Gradients for Visual Grounding cites this paper.

Decomposed On-Policy Distillation for Vision-Language Reasoning: Steering Gradients for Visual Grounding ProgRM: Build Better GUI Agents with Progress Rewards

Reference 44

Resolution
metadata mismatch
arxiv_id, observed 2026-06-28T19:32:35.010485Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-06-28T19:26:37.281563Z digest=sha256:f1f26d6db07ee7aae6cfc489f77360ece69c429b2ee09e683382d0323408f96e

Observation 6c8feefd-d674-4c20-a654-fa417d286531 · inbound

StainFlow: Entity-Stain Tracking and Evidence Linking for Process Rewards in GUI Agents cites this paper.

StainFlow: Entity-Stain Tracking and Evidence Linking for Process Rewards in GUI Agents ProgRM: Build Better GUI Agents with Progress Rewards

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-07-02T16:47:09.689614Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-27T22:24:41.353303Z digest=sha256:19c67dc71dcb6acbb943eeb411875bc147e5a63ceca278f43703b509e215d6bf