{"id":"d52aef52-dd60-4b13-bf4b-fdb931ef605c","arxiv_id":"2608.08070","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"SurgWMBench introduces 1,637 annotated suturing motion segments with 20-frame 2D instrument trajectories and shows that visual generation quality does not imply motion planning accuracy in current surgical world models.","lead":"SurgWMBench is a new benchmark that tests whether AI models can predict a surgical instrument's next positions from surgery videos and recent motion, instead of only generating realistic-looking footage. It evaluates four existing surgical world models and finds that models with good video quality often produce poor instrument trajectories.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Unvalidated, single-anchor ground truth threatens the central mismatch claim: without inter-annotator agreement or validation against instrument masks/kinematics, the ADE gap could reflect annotation ambiguity rather than a true limitation of surgical world models.","rationale":"The reader's weakest assumption—that annotation noise could invalidate the ADE/FDE comparison—is exactly the most load-bearing concern. The paper's central empirical claim is a mismatch between visual generation quality and trajectory accuracy across four baselines. That claim depends on the ground-truth anchor being a valid, consistent proxy for instrument motion. Section 3.3 describes a subjective anchor-selection protocol with no reliability statistics, and the paper itself notes the exact location 'may vary depending on the surgical scene and instrument visibility.' Without inter-annotator agreement or external validation (e.g., against instrument masks or kinematics), the large ADE differences could stem from different models latching onto different plausible anchor interpretations, not from genuine differences in planning-oriented dynamics. This concern is not speculative in a way that would warrant rejection: it is addressable with a modest annotation-reliability study. The internal inconsistency in Appendix Table A.6 (SurgSora and HieraSurg rows appear swapped relative to Table 1) is a real but separate correctness issue; it does not change the verdict because it is fixable and does not by itself undermine the benchmark's central claim. Since the reader already conditioned acceptance on quantifying annotation reliability, the appropriate verdict remains CONDITIONAL, and my read does not move it.","tokens_in":12270,"tokens_out":4406,"duration_ms":46734,"concrete_test":"Re-annotate a random subset of 100 motion segments with two independent annotators and compute per-frame Euclidean distances between their anchor points; also compare each anchor to the SAR-RARP50 instrument segmentation masks by computing distance to the nearest instrument pixel. If the median inter-annotator distance is comparable to the reported ADE gaps (e.g., >20–50 px) or anchors frequently fall outside the instrument mask, the ADE/FDE rankings are not reliable evidence for the headline mismatch.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The benchmark's core supervision is a single manually placed 2D anchor point per frame, defined as 'task-relevant' and 'trackable' (Section 3.3). The paper reports no inter-annotator agreement, no comparison with SAR-RARP50's available instrument segmentation masks, and no validation against robot kinematics. Because the anchor's exact location is allowed to vary by scene and annotator judgment, the ground-truth trajectory may not correspond to a consistent physical point on the instrument. Models are trained to predict this idiosyncratic anchor; a model that generates visually excellent video but locates a different, still-defensible 'task-relevant' point (e.g., needle center instead of tip) would incur large ADE/FDE even if its motion dynamics are correct. The observed mismatch—HieraSurg with 176 px ADE versus iVideoGPT with 51 px ADE (Table 1)—could therefore be an artifact of label ambiguity rather than evidence that visual generation quality does not ensure planning accuracy. The same concern applies to the rollout and perturbation results, which inherit the ground-truth definition. This is more than a noise floor issue: it is a construct-validity threat to the benchmark's claim to measure instrument motion planning.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"SurgWMBench introduces a vision-based benchmark for evaluating surgical world models on short-horizon instrument motion planning. Built from 50 SAR-RARP50 suturing videos, it segments effective needle insertion/extraction motions, uniformly samples 20 frames per segment, and annotates one task-relevant 2D anchor point per frame to form a 20-point instrument trajectory. The benchmark proposes three tasks—short-horizon prediction (Local Transition Fidelity), autoregressive rollout (Closed-loop Rollout Stability), and perturbed-input recovery (Robustness under Perturbation and Shift)—with ADE/FDE and rollout ADE as primary metrics. The authors evaluate VideoGPT, iVideoGPT, HieraSurg, and SurgSora under two settings (pure video generation and joint generation plus trajectory prediction). The central claim is that visual generation quality does not ensure planning-oriented motion accuracy: HieraSurg achieves the best image-level metrics but the worst trajectory error, while iVideoGPT has the best ADE/FDE but not the best visual metrics. Additional perturbation and rollout results indicate limited robustness and stability of current baselines.","tokens_in":12499,"tokens_out":7211,"duration_ms":60986,"significance":"If the benchmark's annotations are valid, SurgWMBench addresses a genuine gap in surgical world-model evaluation: existing metrics such as FVD and CD-FVD measure visual distributional similarity but not trajectory-level accuracy relevant to motion planning. The paper's hierarchical protocol (short-horizon accuracy, closed-loop rollout stability, perturbation robustness) is a well-motivated extension, and the use of a public dataset (SAR-RARP50) with video-level splits is a strength. The reported mismatch between visual generation quality and trajectory accuracy is an interesting and falsifiable finding that would justify the benchmark's existence. However, the benchmark's validity hinges on the reliability of the manually placed 2D anchor points, and the paper currently provides no inter-annotator agreement or external validation against instrument segmentation or kinematics. The paper also omits key experimental details (notably the history length K) and contains an apparent row-swap error in Table 3. With these issues addressed, this could be a useful contribution to the surgical robotics and medical imaging communities.","major_comments":[{"comment":"The benchmark's ground truth is a single manually placed 2D anchor point per frame, with the protocol explicitly allowing the anchor's exact location to vary by surgical scene and annotator judgment. The paper reports no inter-annotator agreement and no validation of the anchors against SAR-RARP50's available instrument segmentation masks or robot kinematics. Because ADE and FDE are computed against these anchors, the observed mismatch between visual generation quality and trajectory accuracy (e.g., HieraSurg ADE 176.20 vs. iVideoGPT 51.32 in Table 1) could partly reflect annotation ambiguity rather than a genuine limitation of the models. Please report inter-annotator distance/agreement statistics, describe how disagreements were resolved, and demonstrate that the anchors correspond to a consistent physical point (e.g., by measuring agreement with the instrument segmentation centroid or against kinematic ground truth).","section":"Section 3.3"},{"comment":"The history length K in Eq. (1) is never specified in the experimental setup or in the tables. All ADE/FDE and rollout values depend on K; for example, Rollout ADE@H in Eq. (7) is computed over frames K+1 through K+H. Without K, the reported numbers cannot be reproduced or compared across baselines. Please state the exact K used in all experiments and, ideally, report results for at least two K values (e.g., K=5 and K=10) to show the sensitivity of the conclusions.","section":"Section 4.1 / Appendix A.2"},{"comment":"There appears to be a row-swap error in Table 3: the row labeled HieraSurg reports ADE@5=64.1 and FDE@5=81.5, while Table 1 gives HieraSurg clean ADE=176.20 and FDE=174.70; the row labeled SurgSora reports ADE@5=176.2 and FDE@5=174.7, while Table 1 gives SurgSora clean ADE=64.12 and FDE=81.54. The @5 values in Table 3 appear to be clean teacher-forcing errors from the swapped models. This inverts the qualitative statements in Appendix A.6 about HieraSurg and SurgSora. Please correct the rows and re-evaluate the rollout analysis and the associated conclusions.","section":"Table 3 / Appendix A.6"},{"comment":"The interpolation strategy in A.4 generates dense pseudo-labels for non-anchor frames through linear, pchip, akima, or cubic spline interpolation, but the experimental setup does not state whether the baselines were trained on the original 20-frame human-annotated trajectories or on the densified sliding-window samples. Since the evaluation ground truth is the sparse 20-point annotation, training on interpolated trajectories may introduce a train/evaluation mismatch. Please clarify which data were used for training and report results for both training settings if both were considered.","section":"Appendix A.4"}],"minor_comments":[{"comment":"The text says that the authors report the Performance Drop Rate (PDR) defined in Eq. (8), but Table 2 only lists ADE and FDE under clean, Gaussian, and Mask conditions without a PDR column. Please either add PDR values or revise the text to state that the table provides the underlying errors from which PDR can be computed.","section":"Section 4.2.3 / Table 2"},{"comment":"The paper states that quality labels are used for difficulty-stratified performance analysis, but no such analysis appears in the main text or the Appendix. Please add the stratified results or remove the claim.","section":"Appendix A.2"},{"comment":"The sentence 'The training, validation, and test splits will follow a unified protocol in the final released version' indicates that the exact split used for the reported experiments is not yet fixed. For a benchmark paper, please provide the exact video-level split (e.g., video IDs or per-split counts) to ensure reproducibility.","section":"Appendix A.2"},{"comment":"The caption merges two settings (pure generation and joint generation plus trajectory prediction) into a single row-wise comparison, but it is not explicit that group (a) and group (b) use the same backbone and training data except for the trajectory head. Please clarify this in the caption or text.","section":"Table 1 caption"},{"comment":"Minor typos and formatting issues: in Eq. (5) the summation subscript appears as '20X' instead of a proper sum symbol, and the label 'Nioottodn' in Figure 1 appears garbled. Please proofread the manuscript.","section":"General"}],"recommendation":"major_revision","confidential_remarks":"The paper is suitable for a robotics or medical image computing venue, and the core idea addresses a real evaluation gap. However, the central mismatch claim depends critically on the validity of the single-anchor annotations; the absence of inter-annotator agreement and any external validation is a serious concern that must be resolved before the benchmark can be trusted. The row-swap in Table 3 also suggests the appendix was not carefully checked. I would ask the authors to add an annotation-validation study and to fix the experimental reporting (K value, split details) in the revision."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The benchmark is a real contribution: it gives the surgical world model community a motion-centric evaluation suite that did not exist, and the headline result—that visual generation quality does not track trajectory accuracy—looks real and worth taking seriously. The construction is thoughtful: 1,637 motion segments from public SAR-RARP50 videos, 20-point 2D trajectories, video-level splits, no mask or action supervision, and metrics computed in original pixel space. The main table is striking: HieraSurg wins on PSNR/LPIPS but has 176 px ADE, while iVideoGPT has the lowest ADE/FDE but middling visual metrics. That pattern is consistent across the perturbation and rollout analyses, which supports the paper's motivation.\n\nThe soft spots are real but mostly addressable. The biggest one is the ground-truth definition: a single 'task-relevant' anchor point whose exact location is allowed to vary by scene, with no inter-annotator agreement and no validation against SAR-RARP50's instrument masks or robot kinematics. That is a genuine construct-validity threat—a model that tracks a different but still defensible point (say, needle center instead of tip) would be penalized even if its motion dynamics are correct. I would not call the central claim dead, but the authors need to report annotation agreement and ideally show that the anchor tracks a consistent physical point. Second, the history length K is never stated in the experimental setup, which makes the numbers hard to reproduce. Third, Table 3 appears to swap the HieraSurg and SurgSora rows relative to Table 1—a factual error that should have been caught. Fourth, the paper repeatedly says splits 'will be reported in the final released version,' which suggests the benchmark is not yet public; for a benchmark paper, releasing data and code is close to a precondition for acceptance. Minor: adding a non-generative trajectory baseline (e.g., an LSTM) would tell whether the task is learnable at all.\n\nNone of these are load-bearing flaws in the sense of refuting the main claim; they are omissions and inconsistencies that a careful revision can fix. The paper is for researchers working on surgical world models or vision-based motion planning, and it deserves a serious referee. My recommendation: send it out, but with a clear request for annotation validation, the missing K, the table fix, and a commitment to release the data and code.","headline":"A genuinely useful benchmark with a plausible central claim, but the unvalidated anchor-point ground truth and missing experimental details keep it from being accepted as-is.","tokens_in":13042,"tokens_out":2198,"would_cite":false,"duration_ms":23226,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Surgical world models can look realistic while failing at instrument motion planning, and a new benchmark measures that gap directly.","keywords":["surgical world models","motion planning benchmark","instrument trajectory prediction","video generation evaluation","robot-assisted surgery","autoregressive rollout","perturbation robustness","ADE FDE metrics"],"falsifier":"Randomly sample a few dozen SurgWMBench segments, have the two annotators re-annotate them independently, and compute per-point inter-annotator distance; if the median disagreement approaches or exceeds the roughly 125-pixel ADE gap that separates iVideoGPT from HieraSurg, the headline mismatch may be an artifact of label noise. A second check: compare the annotated anchors with SAR-RARP50 instrument segmentation masks or kinematic ground truth; if anchors consistently miss the visible instrument tip, the trajectories are not actually instrument motion.","tokens_in":12086,"feed_emoji":"🤖","tokens_out":5398,"duration_ms":47809,"temperature":0.7,"pith_summary":"This paper sets out to show that a surgical world model can generate visually convincing future frames while still being unreliable at the thing that matters for planning: predicting where the instrument will go. To test this, the authors build SurgWMBench from 1,637 real robot-assisted suturing motion segments, each reduced to a 20-frame clip with a hand-annotated 2D instrument anchor trajectory, and evaluate four world-model baselines under three protocols: short-horizon prediction, autoregressive rollout, and perturbed inputs. Their headline result is a mismatch: HieraSurg produces the best-looking frames but the worst trajectory predictions, while iVideoGPT predicts motion best without leading on visual metrics. The paper argues that generation-oriented metrics such as FVD and CD-FVD therefore cannot stand in for planning-oriented evaluation, and that the field needs trajectory-level accuracy, rollout stability, and perturbation recovery as first-class benchmark axes.","feed_headline":"Visual quality does not guarantee surgical motion accuracy","feed_subtitle":"A new benchmark scores four surgical world models on trajectory error, rollout stability, and perturbation recovery.","key_machinery":"The load-bearing object is SurgWMBench itself: a benchmark built from SAR-RARP50 suturing videos, in which each valid needle insertion or extraction is uniformly sampled to 20 frames and annotated with one task-relevant, trackable 2D anchor point per frame, yielding a 20-point instrument trajectory as ground truth. Around this trajectory the benchmark wraps a hierarchical protocol: Local Transition Fidelity (LTF) scores short-horizon teacher-forced prediction with ADE and FDE; Closed-loop Rollout Stability (CRS) scores autoregressive prediction with Rollout ADE@H; and Robustness under Perturbation and Shift (RPS) scores performance drop rate (PDR) under Gaussian noise and random masking. This protocol is what lets the paper separate visual generation quality from motion-planning accuracy.","core_discovery":"The central claim is that visual generation quality does not ensure planning-oriented motion accuracy. In the joint video-generation plus trajectory-prediction setting, HieraSurg achieves the highest PSNR and lowest LPIPS while producing the largest ADE and FDE (176.20 and 174.70 pixels), whereas iVideoGPT has the lowest ADE (51.32) and FDE (65.11) despite weaker visual metrics. Under Gaussian noise and random masking, the better trajectory predictors degrade sharply (e.g., iVideoGPT ADE rises from 51.32 to 131.34 under Gaussian noise), and in closed-loop rollout all models accumulate error with horizon, confirming that local teacher-forced accuracy does not guarantee stable continuous planning. The paper concludes that current surgical world models remain limited in trajectory accuracy, robustness, and planning-oriented dynamics, and that SurgWMBench supplies the missing evaluation layer.","pith_inferences":["The single-anchor, 2D-pixel ground truth likely underweights instrument orientation, needle-tissue interaction, and depth; a natural extension is to evaluate whether anchor-based ADE/FDE rankings survive when predictions are scored against instrument segmentation masks or kinematic recordings.","If the visual-vs-motion mismatch generalizes, one testable corollary is that adding a trajectory head to a video generator will improve ADE/FDE only when the head is trained with motion-focused losses; otherwise the visual backbone may dominate and leave trajectory error high.","Because perturbed inputs cause large PDR for the best clean predictors, a concrete design target for future surgical world models is to condition on trajectory history in a way that is robust to dropout or noise, e.g., via explicit uncertainty estimation or masked-training objectives.","The 20-frame uniform sampling compresses variable-duration motions into a fixed length, so cross-segment comparisons of pixel error assume temporal alignment that the benchmark imposes; extending to variable-horizon or frame-rate-aware evaluation would test whether the reported rankings persist."],"forward_implications":["If the central claim is right, generation-oriented metrics (FVD, CD-FVD, SSIM, PSNR, LPIPS) should not be used as a proxy for surgical world-model quality; trajectory-level ADE/FDE and rollout metrics must be reported for any model intended for planning.","Explicit trajectory supervision does not by itself close the gap: HieraSurg, despite strong visual outputs, remains the worst trajectory predictor, suggesting that visual fidelity and motion dynamics are learned somewhat independently.","Closed-loop stability is a separate axis: every baseline's error grows from horizon @5 to @15, so a model that passes short-horizon tests can still drift when its own predictions are fed back.","Perturbation robustness matters and is currently poor: the stronger clean-input predictors (VideoGPT, iVideoGPT) lose much of their accuracy when the historical trajectory is noisy or masked, while HieraSurg's insensitivity reflects its already-high clean error rather than recovery skill.","SurgWMBench offers a unified, reproducible platform on which future surgical world models can be compared for planning-oriented dynamics rather than only appearance."],"supporting_citations":[{"why":"Baseline that achieves the best ADE/FDE in the joint setting, used to establish the visual-quality/motion-accuracy mismatch.","marker":"[27]"},{"why":"Baseline with the best visual metrics (PSNR/LPIPS) and the worst trajectory error, the key evidence for the central claim.","marker":"[3]"},{"why":"Baseline world model whose trajectory error and rollout behavior are compared alongside the others.","marker":"[5]"},{"why":"Baseline providing the second-best clean trajectory prediction and showing degradation under perturbation.","marker":"[28]"},{"why":"Public source dataset of 50 robot-assisted RARP suturing videos from which all motion segments and frames are derived.","marker":"[17]"},{"why":"Source of FVD, the generation-oriented metric the paper argues is misaligned with motion planning.","marker":"[23]"},{"why":"Source of CD-FVD, another content-biased generation metric the paper argues cannot measure trajectory accuracy.","marker":"[8]"}],"fun_headline_variants":["Surgical world models excel at video, fail at motion planning","SurgWMBench: visual quality does not equal trajectory accuracy","Surgical AI: pretty pictures, poor predictions","New benchmark reveals surgical world models' motion blind spot"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The benchmark's ground truth is a single manually placed 2D anchor point per sampled frame, chosen to be task-relevant and trackable by two annotators, and the paper reports no inter-annotator agreement or cross-check against kinematic or segmentation data, so pixel-level ADE/FDE values could contain annotation noise that drives the reported model ranking.","fun_headline_variants_meta":{"raw":{"variants":["Surgical world models excel at video, fail at motion planning","SurgWMBench: visual quality does not equal trajectory accuracy","Surgical AI: pretty pictures, poor predictions","New benchmark reveals surgical world models' motion blind spot"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000883,"raw_usage":{"total_tokens":3818,"prompt_tokens":950,"completion_tokens":2868,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":566,"completion_tokens_details":{"reasoning_tokens":2801}},"tokens_in":566,"tokens_out":2868,"duration_ms":24167,"temperature":1.0,"reasoning_tokens":2801,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T00:27:16.244200+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Randomly sample a few dozen SurgWMBench segments, have the two annotators re-annotate them independently, and compute per-point inter-annotator distance; if the median disagreement approaches or exceeds the roughly 125-pixel ADE gap that separates iVideoGPT from HieraSurg, the headline mismatch may be an artifact of label noise. A second check: compare the annotated anchors with SAR-RARP50 instrument segmentation masks or kinematic ground truth; if anchors consistently miss the visible instrument tip, the trajectories are not actually instrument motion.","supporting_citations":[],"review_version":1}