CUARewardBench: A benchmark for evaluating reward models on computer-using agent

Let’s verify step by step · 2024 · arXiv 2510.18596

4 Pith papers cite this work. Polarity classification is still indexing.

4 Pith papers citing it

read on arXiv browse 4 citing papers

citation-role summary

dataset 2

citation-polarity summary

background 1 use dataset 1

representative citing papers

PRO-CUA: Process-Reward Optimization for Computer Use Agents

cs.AI · 2026-05-27 · unverdicted · novelty 7.0

PRO-CUA trains CUAs via decoupled on-policy rollouts and PRM-guided step-level optimization to enable dense credit assignment without expert trajectories or golden answers.

Security Considerations for Multi-agent Systems

cs.CR · 2026-03-09 · unverdicted · novelty 6.0

No existing AI security framework covers a majority of the 193 identified multi-agent system threats in any category, with OWASP Agentic Security Initiative achieving the highest overall coverage at 65.3%.

Securing Computer-Use Agents: A Unified Architecture-Lifecycle Framework for Deployment-Grounded Reliability

cs.CL · 2026-05-08 · unverdicted · novelty 4.0

The paper develops a unified framework that organizes computer-use agent reliability around perception-decision-execution layers and creation-deployment-operation-maintenance stages to map security and alignment interventions.

Reinforcement Learning from Human Feedback

cs.LG · 2025-04-16

citing papers explorer

Showing 1 of 1 citing paper after filters.

PRO-CUA: Process-Reward Optimization for Computer Use Agents cs.AI · 2026-05-27 · unverdicted · none · ref 2
PRO-CUA trains CUAs via decoupled on-policy rollouts and PRM-guided step-level optimization to enable dense credit assignment without expert trajectories or golden answers.

CUARewardBench: A benchmark for evaluating reward models on computer-using agent

citation-role summary

citation-polarity summary

fields

years

verdicts

roles

polarities

representative citing papers

citing papers explorer