{"id":"5aebc576-be9f-47d9-86dc-a0b07273df61","arxiv_id":"2607.28609","paper_version":1,"verdict":"ACCEPT","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"low","formal_verification":"none","parameter_count":4,"one_line_summary":"VLM judges of CUA trajectories are systematically lenient; OSReward measures this with human gold, and OS-Shepherd open models close most of the reliability gap cheaply.","lead":"Vision-language models used as judges of computer-using agent trajectories share a systematic leniency bias and collapse on hard cases. The authors release a human-gold cross-platform benchmark and open 9B/35B reward models that match commercial judges at a fraction of the cost.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.5","headline":"No significant objection identified beyond the reader's already-flagged ensemble-supervision soft spot.","rationale":"The central empirical claim is supported by construction that avoids the usual confounds: fresh cross-platform trajectories, multi-annotator human gold with meta-review, a hard set enriched for annotator disagreement, and a broad 27-model judge sweep that isolates a shared over-accept mode. Cost and open-model claims are comparative on the same protocol and are backed by released artifacts. The ensemble-as-supervision loop is the genuine soft underbelly of the OS-Shepherd half of the story, and the reader already named it precisely; external verifier agreement and the staged fail-recall movement on human gold keep correctness risk in the low–medium band the reader assigned rather than elevating it. I would not move the verdict off ACCEPT or lower confidence on what was measured. A targeted error-split audit against the ensemble would be the cleanest way to quantify how much of the open models' hard-set balance is true de-biasing versus high-agreement distillation.","tokens_in":45717,"tokens_out":572,"duration_ms":8680,"concrete_test":"On a stratified 100–150 trajectory subset of OSReward-Hard (and a matched full-set slice), compare OS-Shepherd-9B verdicts to (i) the human gold and (ii) the training-ensemble majority; report the fraction of 9B corrections of ensemble errors vs. retained ensemble errors, plus fail-recall stratified by whether the ensemble was unanimous. If most hard-set gains are retained-ensemble agreements rather than corrections, the de-biasing claim weakens toward distillation.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The reader's weakest assumption is the right load-bearing point: OS-Shepherd-100K labels come from the same VLM-judge class shown to herd and share leniency (§4.3, §5.3), with agreement filtering and a short GRPO pass on mined false successes as the main mitigations (§6.1–6.2, Fig. 14). That is a real methodological soft spot for any claim that the open models are independently reliable rather than distilled ensemble imitators. It does not, however, undercut the paper's primary measurement claim—that frontier VLM judges are systematically lenient and collapse on human-gold hard cases—because that claim rests on the 1019 multi-stage human labels, not on the training corpus. Held-out transfer to OSWorld/WebArena/AndroidWorld verifiers and the large base→SFT→RL fail-recall shift on OSReward-Hard further limit how far residual ensemble bias can explain the reported gains. No stronger internal inconsistency or measurement flaw is evident.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.5","summary":"The paper introduces OSReward, a cross-platform benchmark of 1019 human-gold CUA trajectories for evaluating VLM judges, plus OSReward-Hard and OSReward-Multi. Across 27 judges it reports a shared leniency bias (over-accepting incomplete tasks), a sharp accuracy drop on Hard (best ~70%), and a cost–accuracy frontier on which only expensive frontier models remain usable. It then releases OS-Shepherd-100K (ensemble-labeled, agreement-filtered) and trains open 9B/35B reward models that approach mid-tier commercial accuracy at much lower cost and transfer fail-recall gains to OSWorld, WebArena, and AndroidWorld. Primary measurement claims rest on multi-stage human labels; the open models rest on distilled VLM-ensemble supervision plus a short GRPO stage targeting false successes.","tokens_in":46006,"tokens_out":1243,"duration_ms":26889,"significance":"If the human-gold measurements hold, the paper supplies the first standardized, multi-platform stress test of CUA trajectory judges and documents a concrete, shared failure mode (narrative-driven false successes) that matters for evaluation, curation, and RL. The released benchmark, reasoning-annotated corpus, and self-hostable 9B/35B checkpoints are immediately usable artifacts; the cost–accuracy framing and held-out transfer results strengthen the case that reliable CUA reward need not be frontier-priced. The human pipeline (peer-screened instructions, diverse agent backbones, triple annotation plus meta-review, Hard re-verification, reported agreement) is a genuine strength relative to reused-trajectory judge studies.","major_comments":[{"comment":"§6.1–6.2 and Fig. 10/14: OS-Shepherd-100K labels are produced by the same VLM-judge class shown in §4.3 and §5.3 to herd and share leniency, with agreement filtering and a short GRPO pass on mined false successes as the main mitigations. This does not undermine the human-gold measurement claims on OSReward, but it is load-bearing for any claim that the open models are independently reliable rather than distilled ensemble imitators. The manuscript should state this limitation more explicitly (e.g., in §6.3 or the conclusion), quantify residual disagreement between the ensemble label and a human audit on a held-out training subsample if feasible, and avoid language that equates ensemble agreement with ground truth.","section":"§6.1–6.2"},{"comment":"§7 / Fig. 11: Transfer is reported as agreement with each benchmark’s human-written verifier, which the paper correctly notes are imperfect (citing Xie et al., 2025). The text should more clearly separate “agreement with verifier” from “accuracy,” and, where possible, break out false-positive vs. false-negative shifts so readers can see that the transferred gain is specifically fail-catching rather than overall verifier mimicry. This is especially important on OSWorld, where §E.5 already shows FPs dominate judge errors.","section":"§7, Fig. 11"},{"comment":"Fig. 2, §5.4, §E.1: Cost figures mix official API list prices with “May 2026 market rates” for open-weight models and report 30–60× savings vs. frontier. The comparison is directionally convincing for OS-Shepherd-9B vs. Opus/GPT-5.5, but the paper should fix the pricing date/source table, state whether self-hosted GPU amortization is included in the “API-equivalent” $1.36 figure, and ensure the abstract/body multiplier is computed on a single, documented basket (full-set vs. Hard, same token assumptions).","section":"Fig. 2, §5.4"}],"minor_comments":[{"comment":"Abstract vs. body: the provided abstract text elsewhere says “30–60% lower cost” while the manuscript body and Fig. 2 claim “30–60×”; unify the multiplier everywhere.","section":"Abstract"},{"comment":"§3.1 / §A.4: Mobile environment description is duplicated (“The mobile environment is hosted in an Android emulator…” appears twice in close succession).","section":"§3.1"},{"comment":"§4.5 / Table 2: OSReward-Multi alignment is effectively two-level after removing the single 0-scored run (§B.4); state this in the main Multi discussion so macro-recall is not over-interpreted as three-class grading.","section":"§4.5"},{"comment":"Fig. 1 caption and Table 1: several model marketing names (GPT-5.5, Claude-Opus-4-8, etc.) will date quickly; keep API identifiers from Table 9 adjacent in the main results for reproducibility.","section":"Table 1"},{"comment":"§5.2 takeaway claims actions carry “roughly more signal” than CoT; the −1.8pp vs −7.2pp ablations support text>vision more cleanly than actions>CoT—soften the wording.","section":"§5.2"},{"comment":"Minor typos/style: “with detailed definitions in with detailed definitions in §B.4” (§3.4); “envel⌢pe” artifacts in the author block; ensure Hard-set N=284 platform counts in Fig. 8 match the released metadata.","section":"§3.4"}],"recommendation":"minor_revision","confidential_remarks":"Primary contribution is a carefully labeled benchmark plus a clear empirical failure mode; that alone is above bar for a solid accept at many AI venues. The open RM is valuable engineering but methodologically the weaker half; I would not let ensemble-distillation concerns block the paper if the authors add an explicit limitation and tighten cost reporting. No integrity red flags; human-gold construction is the credible core."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"Punchline: this is the first careful cross-platform measurement of whether VLM judges of CUA trajectories are actually reliable, and the answer is no—they share a clear leniency bias and collapse on hard false successes. The authors then ship open 9B/35B reward models that sit near commercial accuracy at a fraction of the cost.\n\nWhat is new is not “use a VLM as judge” (that is already practice). It is end-to-end fresh trajectories across web/mobile/Ubuntu/Windows, multi-stage human gold (~800 hours, three annotators, meta-review, Krippendorff α≈0.8), a Hard split built from annotator disagreements, and a broad 27-model eval with ablations. The text-history ablation is the useful mechanism result: dropping thought/action text hurts far more than changing screenshots, which explains why confident closing claims fool judges. The cost–accuracy frontier and held-out transfer to OSWorld/WebArena/AndroidWorld are practical, not decorative. Releasing the 100K reasoning-annotated corpus and checkpoints is real community value.\n\nSoft spot, in proportion: OS-Shepherd-100K is labeled by the same VLM-judge class the paper shows herds and is lenient. Agreement filtering plus a short GRPO pass on mined false successes is a reasonable mitigation, and the base→SFT→RL fail-recall shift on human-gold Hard plus external transfer limit how much you can dismiss the gains as pure imitation. Still, any claim that the open models are independently reliable rather than distilled ensemble imitators should be read carefully. Multi-axis grading is weak across the board; they treat it as secondary, correctly. Cost numbers depend on list/market prices—fine for ranking, not gospel.\n\nMath is light (empirical systems paper); data pipeline and citation pattern look honest; related single-platform reward benches are acknowledged. This is for people building or training CUAs who need a trajectory-level reward signal. I would bring it to reading group, cite the benchmark and the leniency finding, and send it to peer review without hesitation.","headline":"Solid human-gold judge benchmark plus usable open reward models; the main soft spot is ensemble-distilled training labels, not the measurement claims.","tokens_in":46717,"tokens_out":528,"would_cite":true,"duration_ms":11067,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"Vision-language judges of computer-using agents systematically accept failed runs as successes, and only expensive frontier models hold up—until open reward models trained on a new human-gold benchmark match them at a fraction of the cost.","keywords":["computer-using agents","VLM-as-judge","reward models","trajectory evaluation","leniency bias","cross-platform GUI agents","OSReward","OS-Shepherd"],"falsifier":"Hold OS-Shepherd and the frontier judges to a new batch of long-horizon desktop trajectories with independent multi-annotator human gold, and check whether fail-recall on false successes stays near the paper’s ~60% hard-set level or collapses back toward the untuned base’s near-zero catch rate.","tokens_in":46604,"feed_emoji":"🖥️","tokens_out":991,"duration_ms":27380,"temperature":0.7,"pith_summary":"Computer-using agents leave long trajectories of screenshots, actions, and thoughts, and the field has started treating vision-language models as automatic judges of whether those runs actually completed the task. This paper shows that those judges are not yet trustworthy: even the strongest ones share a leniency bias, especially on hard false successes where the agent claims done but the environment never reached the goal. On a fresh cross-platform benchmark with human gold labels, frontier accuracy collapses from near 90% on the full set to below 70% on the hard subset, while cheap open models fare worse. The authors release that benchmark, a 100K reasoning-annotated training corpus, and open 9B/35B reward models that recover commercial-level judging at roughly 30–60× lower cost than the frontier and keep the de-biasing on held-out agent benchmarks. The practical stake is a reward signal that evaluation, data curation, and reinforcement learning can actually afford to run at scale.","feed_headline":"Agent judges call failed runs successes—open models fix the bill","feed_subtitle":"A human-gold benchmark exposes shared leniency; 9B/35B reward models match commercial judges far cheaper.","key_machinery":"OSReward: a cross-platform benchmark of 1019 human-gold CUA trajectories (with Hard and Multi subsets) that measures judges themselves, plus OS-Shepherd-100K and the two-stage SFT+RL OS-Shepherd reward models that target false successes directly.","core_discovery":"VLM judges of CUA trajectories fall short of an ideal judge and share one dominant failure mode—over-accepting incomplete tasks as successes—driven more by the agent’s text history than by the screenshots; the few judges accurate enough to trust cost too much for training-scale use, while open OS-Shepherd models trained on a new agreement-filtered corpus close most of that gap at 30–60× lower cost and transfer leniency resistance out of distribution.","pith_inferences":["If text narrative dominates the verdict, agents that learn to write confident closing claims could systematically game reward models unless screen-verification is forced in the objective.","Agreement filtering may quietly discard the hardest genuine failures, so the next corpus gains may come from adversarial mining of judge disagreement rather than more unanimous easy labels.","The platform gap (desktop hardest, mobile easiest) suggests reward-model progress will stall on long GUI+CLI desktop runs unless those are overweighted in training.","Soft, confidence-weighted labels look more promising than majority vote, since judges herd on the same hard cases."],"forward_implications":["Training-time CUA reward no longer has to be bought only at frontier API prices; a small self-hostable judge can sit near the cost–accuracy frontier.","CUA evaluation and data filtering should treat false successes as the primary error mode and measure fail-recall, not only binary accuracy.","Judging prompts and reward pipelines should keep full action/thought text; screenshot count and click markers move aggregate accuracy little.","A single learned judge can begin to replace per-task human-written verifiers across mobile, web, and desktop once de-biasing transfers.","Fine-grained alignment and efficiency grading remain much weaker than binary outcome judging and need separate calibration work."],"fun_headline_variants":["VLM judges mark failed CUA runs successes—open models close the gap cheaper","CUA trajectory judges share leniency bias; OS-Shepherd matches at far lower cost","Agent judges over-accept incomplete tasks; 9B/35B models fix it 30-60x cheaper","OSReward shows VLM judges fail on hard CUA cases—open shepherds supply reliable reward","Leniency dominates VLM CUA judges; agreement-filtered corpus trains low-cost fixes"],"cache_read_input_tokens":128,"weakest_assumption_plain":"That high-agreement labels from strong vision-language judges—after dropping ambiguous cases—are reliable enough training targets even though those judges herd and share the same leniency bias the models are meant to unlearn.","fun_headline_variants_meta":{"raw":{"variants":["VLM judges mark failed CUA runs successes—open models close the gap cheaper","CUA trajectory judges share leniency bias; OS-Shepherd matches at far lower cost","Agent judges over-accept incomplete tasks; 9B/35B models fix it 30-60x cheaper","OSReward shows VLM judges fail on hard CUA cases—open shepherds supply reliable reward","Leniency dominates VLM CUA judges; agreement-filtered corpus trains low-cost fixes"]},"model":"grok-4.5","effort":"low","cost_usd":0.004953,"raw_usage":{"total_tokens":1520,"prompt_tokens":933,"num_sources_used":0,"completion_tokens":102,"cost_in_usd_ticks":49528000,"prompt_tokens_details":{"text_tokens":933,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":485,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":933,"tokens_out":102,"duration_ms":7883,"temperature":1.0,"reasoning_tokens":485,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-31T02:16:21.366544+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"Hold OS-Shepherd and the frontier judges to a new batch of long-horizon desktop trajectories with independent multi-annotator human gold, and check whether fail-recall on false successes stays near the paper’s ~60% hard-set level or collapses back toward the untuned base’s near-zero catch rate.","supporting_citations":[],"review_version":1}