Pith. sign in

REVIEW 3 major objections 6 minor

OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward Models

T0 review · 3 major / 6 minor · reviewed 2026-07-31 · grok-4.5

Pith's one-line read Vision-language judges of computer-using agents systematically accept failed runs as successes, and only expensive frontier models hold up—until open reward models trained on a new human-gold benchmark match them at a fraction of the cost.

desk verdict Solid human-gold judge benchmark plus usable open reward models; the main soft spot is ensemble-distilled training labels, not the measurement claims. read the letter →

arxiv 2607.28609 v2 pith:GU67LITU submitted 2026-07-30 cs.AI cs.CLcs.CV

classification cs.AIcs.CLcs.CV
keywords computer-usingagentsVLM-as-judgerewardmodelstrajectoryevaluationleniencybiascross-platformGUIOSRewardOS-Shepherd
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Computer-using agents leave long trajectories of screenshots, actions, and thoughts, and the field has started treating vision-language models as automatic judges of whether those runs actually completed the task. This paper shows that those judges are not yet trustworthy: even the strongest ones share a leniency bias, especially on hard false successes where the agent claims done but the environment never reached the goal. On a fresh cross-platform benchmark with human gold labels, frontier accuracy collapses from near 90% on the full set to below 70% on the hard subset, while cheap open models fare worse. The authors release that benchmark, a 100K reasoning-annotated training corpus, and open 9B/35B reward models that recover commercial-level judging at roughly 30–60× lower cost than the frontier and keep the de-biasing on held-out agent benchmarks. The practical stake is a reward signal that evaluation, data curation, and reinforcement learning can actually afford to run at scale.

What carries the argument

OSReward: a cross-platform benchmark of 1019 human-gold CUA trajectories (with Hard and Multi subsets) that measures judges themselves, plus OS-Shepherd-100K and the two-stage SFT+RL OS-Shepherd reward models that target false successes directly.

What would settle it

Hold OS-Shepherd and the frontier judges to a new batch of long-horizon desktop trajectories with independent multi-annotator human gold, and check whether fail-recall on false successes stays near the paper’s ~60% hard-set level or collapses back toward the untuned base’s near-zero catch rate.

Watch

Extended reading notes

Core claim

VLM judges of CUA trajectories fall short of an ideal judge and share one dominant failure mode—over-accepting incomplete tasks as successes—driven more by the agent’s text history than by the screenshots; the few judges accurate enough to trust cost too much for training-scale use, while open OS-Shepherd models trained on a new agreement-filtered corpus close most of that gap at 30–60× lower cost and transfer leniency resistance out of distribution.

Load-bearing premise

That high-agreement labels from strong vision-language judges—after dropping ambiguous cases—are reliable enough training targets even though those judges herd and share the same leniency bias the models are meant to unlearn.

Editorial extensions

If this is right

  • Training-time CUA reward no longer has to be bought only at frontier API prices; a small self-hostable judge can sit near the cost–accuracy frontier.
  • CUA evaluation and data filtering should treat false successes as the primary error mode and measure fail-recall, not only binary accuracy.
  • Judging prompts and reward pipelines should keep full action/thought text; screenshot count and click markers move aggregate accuracy little.
  • A single learned judge can begin to replace per-task human-written verifiers across mobile, web, and desktop once de-biasing transfers.
  • Fine-grained alignment and efficiency grading remain much weaker than binary outcome judging and need separate calibration work.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If text narrative dominates the verdict, agents that learn to write confident closing claims could systematically game reward models unless screen-verification is forced in the objective.
  • Agreement filtering may quietly discard the hardest genuine failures, so the next corpus gains may come from adversarial mining of judge disagreement rather than more unanimous easy labels.
  • The platform gap (desktop hardest, mobile easiest) suggests reward-model progress will stall on long GUI+CLI desktop runs unless those are overweighted in training.
  • Soft, confidence-weighted labels look more promising than majority vote, since judges herd on the same hard cases.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 6 minor

Summary. The paper introduces OSReward, a cross-platform benchmark of 1019 human-gold CUA trajectories for evaluating VLM judges, plus OSReward-Hard and OSReward-Multi. Across 27 judges it reports a shared leniency bias (over-accepting incomplete tasks), a sharp accuracy drop on Hard (best ~70%), and a cost–accuracy frontier on which only expensive frontier models remain usable. It then releases OS-Shepherd-100K (ensemble-labeled, agreement-filtered) and trains open 9B/35B reward models that approach mid-tier commercial accuracy at much lower cost and transfer fail-recall gains to OSWorld, WebArena, and AndroidWorld. Primary measurement claims rest on multi-stage human labels; the open models rest on distilled VLM-ensemble supervision plus a short GRPO stage targeting false successes.

Significance. If the human-gold measurements hold, the paper supplies the first standardized, multi-platform stress test of CUA trajectory judges and documents a concrete, shared failure mode (narrative-driven false successes) that matters for evaluation, curation, and RL. The released benchmark, reasoning-annotated corpus, and self-hostable 9B/35B checkpoints are immediately usable artifacts; the cost–accuracy framing and held-out transfer results strengthen the case that reliable CUA reward need not be frontier-priced. The human pipeline (peer-screened instructions, diverse agent backbones, triple annotation plus meta-review, Hard re-verification, reported agreement) is a genuine strength relative to reused-trajectory judge studies.

major comments (3)
  1. [§6.1–6.2] §6.1–6.2 and Fig. 10/14: OS-Shepherd-100K labels are produced by the same VLM-judge class shown in §4.3 and §5.3 to herd and share leniency, with agreement filtering and a short GRPO pass on mined false successes as the main mitigations. This does not undermine the human-gold measurement claims on OSReward, but it is load-bearing for any claim that the open models are independently reliable rather than distilled ensemble imitators. The manuscript should state this limitation more explicitly (e.g., in §6.3 or the conclusion), quantify residual disagreement between the ensemble label and a human audit on a held-out training subsample if feasible, and avoid language that equates ensemble agreement with ground truth.
  2. [§7, Fig. 11] §7 / Fig. 11: Transfer is reported as agreement with each benchmark’s human-written verifier, which the paper correctly notes are imperfect (citing Xie et al., 2025). The text should more clearly separate “agreement with verifier” from “accuracy,” and, where possible, break out false-positive vs. false-negative shifts so readers can see that the transferred gain is specifically fail-catching rather than overall verifier mimicry. This is especially important on OSWorld, where §E.5 already shows FPs dominate judge errors.
  3. [Fig. 2, §5.4] Fig. 2, §5.4, §E.1: Cost figures mix official API list prices with “May 2026 market rates” for open-weight models and report 30–60× savings vs. frontier. The comparison is directionally convincing for OS-Shepherd-9B vs. Opus/GPT-5.5, but the paper should fix the pricing date/source table, state whether self-hosted GPU amortization is included in the “API-equivalent” $1.36 figure, and ensure the abstract/body multiplier is computed on a single, documented basket (full-set vs. Hard, same token assumptions).
minor comments (6)
  1. [Abstract] Abstract vs. body: the provided abstract text elsewhere says “30–60% lower cost” while the manuscript body and Fig. 2 claim “30–60×”; unify the multiplier everywhere.
  2. [§3.1] §3.1 / §A.4: Mobile environment description is duplicated (“The mobile environment is hosted in an Android emulator…” appears twice in close succession).
  3. [§4.5] §4.5 / Table 2: OSReward-Multi alignment is effectively two-level after removing the single 0-scored run (§B.4); state this in the main Multi discussion so macro-recall is not over-interpreted as three-class grading.
  4. [Table 1] Fig. 1 caption and Table 1: several model marketing names (GPT-5.5, Claude-Opus-4-8, etc.) will date quickly; keep API identifiers from Table 9 adjacent in the main results for reproducibility.
  5. [§5.2] §5.2 takeaway claims actions carry “roughly more signal” than CoT; the −1.8pp vs −7.2pp ablations support text>vision more cleanly than actions>CoT—soften the wording.
  6. [§3.4] Minor typos/style: “with detailed definitions in with detailed definitions in §B.4” (§3.4); “envel⌢pe” artifacts in the author block; ensure Hard-set N=284 platform counts in Fig. 8 match the released metadata.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: primary claims rest on held-out multi-stage human-gold labels and external verifiers, not on quantities defined by the training ensemble.

full rationale

OSReward is an empirical measurement and systems paper, not a first-principles derivation. The load-bearing reliability claims (frontier judges share leniency; accuracy collapses on OSReward-Hard; OS-Shepherd matches commercial judges at much lower cost and transfers de-biasing) are scored against 1019 trajectories with multi-annotator human gold plus meta-review, fully disjoint from OS-Shepherd-100K. Training labels are agreement-filtered VLM ensemble judgments—a standard distillation setup with a known shared-bias risk—but the paper does not redefine success by those labels, does not fit a parameter on the evaluation set and call it a prediction, and does not import a self-authored uniqueness theorem to force the result. Held-out agreement with OSWorld/WebArena/AndroidWorld human-written verifiers and the reported base→SFT→RL fail-recall shift on human-gold Hard further separate measured performance from training-label construction. Self-citations (OS-Genesis, OpenMobile, infrastructure papers) supply data-collection methods and reused open trajectories, not the gold verdicts or the accuracy numbers. No step reduces a claimed prediction to its inputs by construction.

Assumptions & free parameters 4 free parameters · 4 assumptions · 2 invented entities

Empirical systems paper. Load-bearing premises are methodological: human multi-annotator verdicts as gold; fixed judging protocol (last-N screenshots + text); ensemble agreement as scalable proxy labels; list/market API prices for cost claims. No physical constants or fitted scientific laws. Invented items are named datasets/models, not unobserved theoretical entities.

free parameters (4)
  • N trailing screenshots (default 5) = 5
    Main protocol fixes N=5; ablations show aggregate accuracy is flat in N but individual verdicts flip (§4.1, §5.1, Fig. 19).
  • Ensemble agreement filter threshold = ~85% of judged trajectories retained
    Trajectories enter OS-Shepherd-100K only under high cross-judge agreement; exact keep/drop rule shapes the training distribution (§6.1).
  • GRPO RL hyperparameters = lr=1e-6, KL=0.001, ~150 steps
    Learning rate 1e-6, KL 0.001, 8 rollouts, ~150 steps, reward 1.0/0.1/0.0 scheme relocate the operating point on false successes (§D.2).
  • API/market unit prices for cost frontier = e.g. OS-Shepherd-9B ~$1.36 per full-set pass
    Cost–accuracy claims (30–60×) depend on May-2026 list/market prices and average trajectory token mix (Fig. 2, §E.1).
assumptions (4)
  • domain assumption Multi-stage human annotation (3-way + meta-review) yields ground-truth success/fail for CUA trajectories.
    Entire benchmark validity rests on this; stated in §3.3–3.4 and §B.
  • domain assumption A trajectory is FAIL if the agent did not obtain/verify the answer through the environment, even if the answer is factually correct.
    Explicit judging rule shared by annotators and prompt (§3.3, §C.2 grounding rule).
  • ad hoc to paper High-agreement votes among strong VLM judges are sufficiently clean training labels after dropping the ambiguous middle.
    Core scalable-supervision choice in §6.1; motivated by herding analysis in §5.3.
  • domain assumption Standard supervised fine-tuning plus GRPO improves judge calibration without needing new human labels.
    Training recipe §6.2 / §D.2; common RLHF/GRPO practice applied to this setting.
invented entities (2)
  • OSReward / OSReward-Hard / OSReward-Multi independent evidence
    purpose: Standardized human-gold evaluation of CUA trajectory judges across platforms and difficulty.
    Named benchmark artifacts constructed by the authors; falsifiable via released labels and re-annotation.
  • OS-Shepherd-100K and OS-Shepherd 9B/35B independent evidence
    purpose: Open training corpus and self-hostable reward models targeting false-success bias at low cost.
    Released models/data; performance claims are testable on OSReward and external benchmarks.

how reviews work

0 comments
Cite this review

Pith. "Pith review of OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward Models." pith.science (2026). https://pith.science/paper/GU67LITU

@misc{pith2026260728609,
  author       = {Pith},
  title        = {Pith review of: OSReward: Instituting Standardized Evaluation for Cross-Platform Computer-Use Reward Models},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/GU67LITU}},
  note         = {Machine review of arXiv:2607.28609}
}
read the original abstract

Computer-using agents (CUAs) are advancing rapidly across the digital world. A CUA trajectory records the agent's actions, states, and reasoning. Verifying whether it fulfilled the task instruction is central to CUA evaluation, data curation, and reinforcement learning. Neither human-written verifiers nor human annotators can provide such verification at scale, so the field increasingly turns to vision-language models (VLMs) as judges of CUA trajectories. But a fundamental question has long gone unexamined: are these VLM judges reliable enough? To study it systematically, we introduce OSReward, a realistic, high-quality benchmark that evaluates VLM judges on CUA trajectories. The trajectories come from diverse agent backbones executing human-verified instructions across platforms, and are then rigorously labeled with ground-truth verdicts through multi-stage human annotation. Building on it, we derive OSReward-Hard, a challenge set concentrating genuinely hard cases, and OSReward-Multi for fine-grained efficiency and alignment scoring. The most comprehensive evaluation of VLM judges to date finds even state-of-the-art models fall short of an ideal judge, sharing a systematic leniency bias that mislabels failed runs as successes. The few reliable enough to trust are too expensive to run at scale, while affordable open models trail far behind. To close this gap, we construct and release OS-Shepherd-100K, an open corpus of reasoning-annotated trajectory judgments for the CUA community. On it, we train OS-Shepherd (9B and 35B), open reward models that supply low-cost, stable, and reliable reward signals, matching commercial judges at 30-60x lower cost than the frontier. Extensive analyses further inform the design of reliable CUA reward at scale. Our code, benchmark, dataset, and model checkpoints are available at https://os-copilot.github.io/OSReward-Home/.

Figures

Figures reproduced from arXiv: 2607.28609 by the authors.

Figure 1
Figure 1. How well VLM judges score CUA trajectories (a), and the strict–lenient bias they share (b). Contact author(s): qiushisun@connect.hku.hk, chengkz@smail.nju.edu.cn, zjb@nju.edu.cn * Equal contribution † Project co-lead # Corresponding author arXiv:2607.28609v1 [cs.AI] 30 Jul 2026 [PITH_FULL_IMAGE:figures/full_fig_p001_1.png] view at source ↗
Figure 2
Figure 2. Cost against binary accuracy on OSReward-Hard: reliable judges are expensive, and the OS-Shepherd models come nearest their accuracy at a fraction of the cost. collection pipeline, process and screen its output, and curate it together with filtered public corpora into OS-Shepherd-100K, an open corpus of 100K reasoning-annotated trajectory judgments that spans more trajectories, agent backbones, scenarios, and task t… view at source ↗
Figure 3
Figure 3. From realistic environments to raw trajectories: annotators prepare the environments and write grounded instructions on them, agents from four model families execute the instructions. 3. OSReward Our goal is a corpus of realistic, cross-platform CUA trajectories paired with gold verdicts trustworthy enough to measure the judges themselves. Reusing existing benchmarks’ rollouts cannot supply it: their runs carry qual… view at source ↗
Figures from the paper (25 more)
Figure 4
Figure 4. Figure 4: The annotation pipeline. Each pre-filtered trajectory is labeled by three independent annotators; disagreements go to meta review, and the verified gold set is read as three views instruction is run by one to three of them, spanning the Claude, Gemini, Kimi, and Qwen f…
Figure 5
Figure 5. Figure 5: OSReward at a glance: outcome composition of the full and Hard sets, platform mix, and trajectory length; failed runs are markedly longer. latter two are nested subsets of the full set. The Full Set (OSReward). The full OSReward set holds 1019 trajectories spanning all…
Figure 6
Figure 6. Figure 6: Judge bias on OSReward (left) and OSReward-Hard (right). Judges below the diagonal skew lenient; most of the field sits there, and the skew widens on the hard set. 4.4. OSReward-Hard: The Challenge Set The Collapse. The aggregate ∼90% is optimistic. OSReward-Hard, the …
Figure 7
Figure 7. Figure 7: Per-judge error composition, seven representative judges: over-accepting an incomplete task (warm colors) dominates every family. 0 10 20 30 40 50 60 Mean judge binary acc (%) windows (n=29) web (n=60) ubuntu (n=132) mobile (n=63) 42.4% 51.9% 52.1% 58.3% OSReward-Hard:…
Figure 8
Figure 8. Figure 8: Mean per-judge binary accuracy on OSReward-Hard, by platform and by failure type. Failure types are multi-label, so their counts sum to more than the number of failed trajectories. Observations from the Hard Set. OSReward-Hard concentrates on the cases the shared lenie…
Figure 9
Figure 9. Figure 9: Overview of input ablations: Δ binary accuracy per (setting × model) vs. the main setting. This is the mechanism behind the leniency bias of §4.3. A judge that leans on the agent’s own narrative is exactly a judge that a confident closing claim can fool. This suggests …
Figure 10
Figure 10. Figure 10: The OS-Shepherd-100K pipeline: self-collected and open-source trajectories are ensemble￾judged and distilled into the training set. Band widths ∝ trajectory counts. kept trajectory is scored by an ensemble of strong VLM-as-a-Judge runs under the protocol of § 4.1, var…
Figure 11
Figure 11. Figure 11: Judges on three existing CUA benchmarks against each benchmark’s human-written verifier (matched subsets): accuracy and failure recall; means are computed before rounding. up to Qwen3.5-397B-A17B (∼44× the 9B’s size) on all three ( [PITH_FULL_IMAGE:figures/full_fig_p…
Figure 12
Figure 12. Figure 12: The twenty most frequent task types among the roughly 32K web instructions sampled for collection, together about 65% of the pool. Fact-finding lookups dominate, followed by article, academic￾paper, and shopping tasks. A.2. Windows Environment. Our Windows environment…
Figure 13
Figure 13. Figure 13: Failure-type profile over OSReward’s fail trajectories (multi-label shares; the catch-all others tag is excluded). A single run can carry several tags. benchmark was checked for personally identifiable information, and none enters the release. B.3. Failure-Type Taxono…
Figure 14
Figure 14. Figure 14: The de-biasing trajectory: base→ SFT→SFT+RL moves OS-Shepherd-9B from the lenient corner toward the balanced diagonal, with RL supplying the largest hard-set step. SFT RL (both sizes) 9B 35B-A3B Base model Qwen3.5-9B / Qwen3.5-35B-A3B 9B SFT ckpt 35B SFT ckpt Samples …
Figure 15
Figure 15. Figure 15: gives the cost-accuracy view on the full set. Costs are official API list prices where available; open-weight models with no official pricing are charged at May-2026 market rates for models of similar size [PITH_FULL_IMAGE:figures/full_fig_p042_15.png]
Figure 16
Figure 16. Figure 16: Open-weight vs. closed-source binary accuracy (group means over the 27-judge field). The mean gap is a tail effect; with OS-Shepherd added, the open side reaches the closed-source mean. Macro-recall is the balanced accuracy of the emitted levels: threshold-dependent, …
Figure 17
Figure 17. Figure 17: Pairwise agreement. (a) binary-verdict 𝜅; (b) 3-class alignment 𝜅, both hierarchically clustered on 1 − 𝜅. E.4. Inter-Judge Agreement [PITH_FULL_IMAGE:figures/full_fig_p044_17.png]
Figure 18
Figure 18. Figure 18: OSWorld deep dive (external benchmark), pooled over eight judges. Left: accuracy by application domain. Right: accuracy and false-positive rate vs. trajectory length [PITH_FULL_IMAGE:figures/full_fig_p045_18.png]
Figure 19
Figure 19. Figure 19: Binary accuracy vs. the number of trailing screenshots 𝑁. No judge trends with 𝑁; the shaded band marks the 𝑁 = 5–9 regime the main setting draws from. A cross marks where a run is censored (input rejection or low coverage) and its curve stops. 45 [PITH_FULL_IMAGE:fi…
Figure 20
Figure 20. Figure 20: Self-consistency at 𝑇=0.7 over five trials. (a) aggregate accuracy vs. greedy; (b) per-trajectory flip rate (lower = steadier labels). 46 [PITH_FULL_IMAGE:figures/full_fig_p046_20.png]
Figure 21
Figure 21. Figure 21: A recency-grounding failure: the selected documentary visibly dates to 2017, but the agent presents it as recently popular without supporting evidence. 47 [PITH_FULL_IMAGE:figures/full_fig_p047_21.png]
Figure 22
Figure 22. Figure 22: A false-success case in which self-narration overrides terminal evidence: the required git status command never appears and produces no visible output. TRAJECTORY EVIDENCE INSTRUCTION DESKTOP Navigate to "finance.yahoo.com" using the address bar, search for "NVDA" to …
Figure 23
Figure 23. Figure 23: A fine-grained perception failure on Yahoo Finance: after reaching NVDA’s Statistics page, the agent misreads both the Beta window and its displayed value. 48 [PITH_FULL_IMAGE:figures/full_fig_p048_23.png]
Figure 24
Figure 24. Figure 24: A fine-grained visual miss in Audacity: the agent follows convincing action and export cues without independently verifying whether the waveform satisfies the requested first-ten-second Fade Out. TRAJECTORY EVIDENCE INSTRUCTION DESKTOP In Krita, use the Crop Tool to r…
Figure 25
Figure 25. Figure 25: A long-horizon Ubuntu planning failure: repeated image-insertion errors are obscured by a successful final save, causing a visually incorrect LibreCAD document to be mistaken for a completed artifact. 49 [PITH_FULL_IMAGE:figures/full_fig_p049_25.png]
Figure 26
Figure 26. Figure 26: A successful long-horizon desktop case requiring sustained state tracking and verification of the final spreadsheet. 50 [PITH_FULL_IMAGE:figures/full_fig_p050_26.png]
Figure 27
Figure 27. Figure 27: A mobile hard case requiring the judge to verify the nearest open hospital and a fuel stop in the final route. TRAJECTORY EVIDENCE INSTRUCTION WEB Find a hotel in Rome under 180 euros per night for next weekend. Then search for Italian restaurants within walking dista…
Figure 28
Figure 28. Figure 28: A successful Web hard case requiring cross-site constraint tracking and grounded recovery from anti-bot blocks. 51 [PITH_FULL_IMAGE:figures/full_fig_p051_28.png]

Discussion (0). Continue with ORCID to comment.

Pith tools

Reviewed July 31, 2026 · model on record in the stance chip above.