{"id":"cac7a18d-65bb-443f-945f-cc5e6bfa13ce","arxiv_id":"2608.06770","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"A controllable surgical world model that fuses a first-frame and hierarchical-mask anchor with edge, depth, and flow expert increments, and outperforms existing baselines on a new Cholec80-based benchmark.","lead":"Surg-UniWorld is a video-generation model that makes fake surgical videos obey user-supplied edge, depth, and motion controls while keeping the scene's anatomy and instruments consistent. It is a candidate building block for surgical training simulators and for generating rare instrument-tissue interactions without collecting new real footage.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Table II compares Surg-UniWorld with a richer condition set (future hierarchical masks plus edge, depth, and flow) against baselines receiving a single control, so the claimed consistent outperformance is not established by matched evidence.","rationale":"I read the paper in good faith. The architecture is novel and internally coherent: the Hierarchical Surgical Anchor (Section IV-B) and Anchor-Relative Modality Experts (Section IV-C) are described with enough detail to be reproducible, and the auxiliary losses (Section IV-E) are reasonable. The load-bearing weakness is empirical rather than conceptual. The central claim of consistent, architecture-driven superiority rests on Table II, whose columns do not hold the input condition set fixed. Surg-UniWorld's always-on input includes the full temporal mask sequence, which encodes the future semantic layout, and its full-control row adds edge, depth, and flow. Baselines receive only one condition type. Therefore the 2.751 dB PSNR improvement and the FVD/FID reductions are confounded by information budget. This is more fundamental than the reader's mask-quality concern: even if the masks are perfectly accurate, their exclusive availability to the proposed method would still invalidate the comparison. The proposed concrete test—rerunning the strongest baselines with the same multi-condition inputs—would settle the issue. If the gap persists under matched inputs, the paper's claims are substantiated; if not, the comparisons must be revised. The reader's conditional verdict remains appropriate, with this concern added to the fixable gaps. The mask-quality gap (no inter-rater agreement reported) is real but secondary to the information asymmetry identified here.","tokens_in":15093,"tokens_out":11575,"duration_ms":101821,"concrete_test":"Run VACE-Wan2.2-5B and ControlNet-Wan2.2-5B with the identical multi-condition input used for Surg-UniWorld's 'All' configuration: the same temporally aligned hierarchical mask sequence plus the same edge, depth, and optical-flow sequences from Cholec80-SurgWAM. If their PSNR/FVD/FID approach or exceed 21.250/92.981/7.359, the headline gains are explained by information asymmetry. If those interfaces cannot accept all four conditions, the paper should report a matched-information comparison (e.g., mask-only versus mask-only at equal mask granularity) and constrain the claim of consistent outperformance accordingly.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Every Surg-UniWorld row in Table II is conditioned on the full temporal hierarchical mask sequence M={M_t}_{t=0}^{T-1}, which is defined in Eq. (1) and used throughout the anchor construction (Eqs. 4-8). M includes semantic masks for all future frames, not just the first frame. On top of M, the 'All' row receives edge, depth, and optical-flow sequences. The baselines are evaluated with a single control: the strongest PSNR baseline is ControlNet-Wan2.2 with edge only (18.499 dB), and the strongest FVD/FID baselines are VACE-Wan2.2 with mask only (104.461 / 11.312). Surg-UniWorld's 21.250 dB, 92.981 FVD, and 7.359 FID therefore reflect a strictly larger information budget: the future semantic layout and four condition streams, versus one stream for the baselines. No baseline is run with the same hierarchical mask sequence plus edge, depth, and flow, so the advantage cannot be attributed to the proposed anchor and expert design rather than to the additional conditioning signals. The inconsistency is visible in the reported table: in the mask-only configuration, Surg-UniWorld has LPIPS 0.297 and FVD 270.996, which are worse than VACE-Wan2.2 mask-only (LPIPS 0.274, FVD 104.461). This directly contradicts the abstract's 'consistently outperforms' claim.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes Surg-UniWorld, a controllable surgical video generation framework built on a frozen Wan2.2 video diffusion backbone. The method introduces a Hierarchical Surgical Anchor constructed from the first frame and a temporally aligned sequence of hierarchical semantic masks (instrument, foreground tissue, background tissue), followed by Anchor-Relative Modality Experts that interpret edge, depth, and optical-flow controls relative to this anchor. A Multimodal Control Expert composes the anchor hint with modality-specific increments in a stage-wise, contribution-preserving manner and injects the result into selected DiT blocks. The authors also construct Cholec80-SurgWAM, a benchmark derived from Cholec80 with hierarchical masks, text descriptions, and aligned edge, depth, and flow controls. Experiments on this benchmark compare Surg-UniWorld with several video generation and controllable generation baselines across quality, temporal consistency, and control-adherence metrics, with additional ablations on architectural choices, losses, and composition strategies. The central claim is that the full-control configuration consistently outperforms all baselines in generation quality, temporal consistency, and multimodal controllability.","tokens_in":1909,"tokens_out":1982,"duration_ms":45390,"significance":"If the empirical claims are supported, Surg-UniWorld would be a useful contribution to surgical world modeling and controllable video generation: it explicitly separates persistent scene anchors from optional modality evidence, introduces a region-aware composition mechanism, and provides a new benchmark with curated hierarchical masks and multimodal controls. The paper is also transparent about many design choices and includes detailed ablations. However, the current evidence does not fully support the headline claim. The main comparison is not matched in conditioning, the control-adherence metrics are computed with the same estimators that generated the input conditions, and all reported numbers appear to come from single runs without error bars or significance tests. These issues are central to the paper's claims and should be addressed before the work can be accepted.","major_comments":[{"comment":"The headline comparison is not matched in conditioning. Surg-UniWorld's Mask row uses the full temporal sequence of hierarchical masks M={M_t}_{t=0}^{T-1} together with the first frame and text, while the All row additionally receives edge, depth, and optical-flow sequences; each baseline in Table II is evaluated with a single control condition. The improvements of the full configuration (PSNR 21.250 vs. 18.499, FVD 92.981 vs. 104.461, FID 7.359 vs. 11.312) therefore reflect a strictly larger information budget and cannot be attributed specifically to the proposed anchor and expert design. Moreover, the Surg-UniWorld mask-only row reports LPIPS 0.297 and FVD 270.996, which are worse than the VACE-Wan2.2 mask-only row (LPIPS 0.274, FVD 104.461), directly contradicting the abstract's statement that Surg-UniWorld consistently outperforms existing controllable video generation methods and surgical world-model baselines. The authors should either run baselines under the same hierarchical mask sequence and the same optional modality set, or qualify the claim to the matched settings.","section":"V-B, Table II"},{"comment":"The control-adherence metrics (Edge F1, Depth si-RMSE, Flow EPE) are computed by re-extracting conditions from the generated video with a fixed modality estimator. In this paper, those estimators are the same HED, Depth Anything 3, and WAFT models that were used in Section III to produce the input edge, depth, and flow annotations. This creates a systematic bias: a model that reproduces estimator-specific artifacts rather than true geometric or motion properties can achieve artificially high adherence scores. Since the multimodal controllability claim rests on these metrics, the authors should validate at least a subset of the control-adherence results with independent estimators or human evaluation, and should discuss the potential circularity explicitly.","section":"V-A2 and Section III"},{"comment":"All quantitative results appear to be from single training runs without error bars, confidence intervals, or significance tests. This is load-bearing because several reported differences are small relative to what one would expect from stochastic training (for example, LPIPS 0.186 vs. 0.190 and FVD 92.981 vs. 98.716 in the ablations of Table III). Without multiple seeds or statistical testing, the ranking of configurations and the claims of consistent improvement are not established. The authors should report variance over at least three runs for the main comparison and for the key ablations, or provide a justified significance analysis.","section":"V-B, Tables II-V"},{"comment":"The hierarchical masks are produced by SAM2 and then manually reviewed and refined, but no inter-rater agreement, mask-quality metrics, or error analysis are reported. These masks are not merely an input: region weights W_i, shape tokens S_i, the anchor bank B_i, the region-aware flow-matching loss in Eq. (16), and the region-wise evaluations in Fig. 7 all depend on them. If the masks contain systematic errors in instrument/tissue boundaries, those errors propagate into every control stage and into the evaluation itself. The authors should report mask quality on a held-out subset and, ideally, a sensitivity analysis with perturbed masks to show that the conclusions are robust.","section":"III and IV-B"}],"minor_comments":[{"comment":"The method reuses the VACE patch embedding and context blocks with shared parameters, while VACE-Wan2.2 is also used as a baseline. The paper should state explicitly how much of the VACE pipeline is reused, whether this reuse is considered part of the proposed adapter or an external component, and what this implies for the comparison with the VACE baseline.","section":"IV-C, Eq. (10)"},{"comment":"The sentence 'Each condition is re-extracted from the generated video using a fixed modality estimator' should clarify that the input control is itself an estimator output, and that the estimator family is the same one used in dataset construction; this point is related to the major concern about circularity.","section":"V-A2"},{"comment":"In the All row, the Flow EPE value appears as '0.0951.327' and 'All21.250' with missing spacing; these formatting errors should be corrected.","section":"Table II"},{"comment":"The notation for the weighted norm is unusual: it is defined as a ratio of weighted sums, but the same notation is later used for a squared weighted norm. The authors should define the weighted squared norm explicitly to avoid confusion.","section":"IV-E, Eq. (15)"},{"comment":"The claim that the consistent improvements across pixel-level, perceptual, and spatiotemporal metrics indicate that the gains extend beyond frame reconstruction to overall temporal realism is too strong given the absence of error bars and the unmatched conditioning noted above; please temper the wording.","section":"V-B.1"}],"recommendation":"major_revision","confidential_remarks":"The manuscript has a clearly described architecture and a substantial new benchmark, but the empirical case is currently overstated. The most serious issue is the unmatched comparison: the full Surg-UniWorld configuration is given future semantic masks plus edge, depth, and flow, while baselines receive only one control. The abstract's consistently outperforms claim is also contradicted by the mask-only row of Table II, where VACE-Wan2.2 achieves lower LPIPS and much lower FVD. I believe these issues are fixable with matched baselines, error bars, and a more careful claim, so I recommend major revision rather than rejection. There is also a novelty-boundary question regarding the reuse of VACE that the authors should address explicitly."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The paper's design is the real news. The Hierarchical Surgical Anchor, anchor-relative modality experts, and contribution-preserving stage-wise composition are not just a restickered ControlNet or VACE. They make a coherent architectural claim: persistent scene identity and semantic layout should be treated as an anchor, while edge, depth, and flow are optional deltas interpreted relative to that anchor. The Cholec80-SurgWAM benchmark, with 6,001 clips and 573,721 curated masks, is a practical resource. The ablations are internally consistent and each component appears to earn its place. The auxiliary losses (temporal structure, control benefit, marginal consistency) are thoughtfully motivated and the counterfactual-loss idea is a nice touch.\n\nWhere it gets soft: the central comparison in Table II is not matched. The full Surg-UniWorld row gets future hierarchical masks plus edge, depth, and flow; the baselines get one control stream apiece. So the 2.75 dB PSNR gain and the 11% FVD reduction may reflect the extra conditioning information, not the anchor-expert design. The mask-only row is the one matched comparison, and there Surg-UniWorld is worse than VACE-Wan2.2 mask on LPIPS (0.297 vs 0.274) and much worse on FVD (270.996 vs 104.461). That directly contradicts the abstract's 'consistently outperforms' claim. The issue is fixable: either run baselines with the same multimodal conditioning, or frame the comparison as a within-model ablation and drop the global superiority claim.\n\nTwo more soft spots. First, all metrics are single-run, no error bars or significance tests; for a paper with margins like these, that is a real gap. Second, the control-adherence metrics (Edge F1, Depth si-RMSE, Flow EPE) re-use the same estimators that produced the input conditions (Depth Anything 3, WAFT, HED). That puts a thumb on the scale for any model that reconstructs the input statistics. The dependence on SAM2 masks with 'manual review' but no inter-rater agreement is worth flagging too, since errors in M propagate into every stage.\n\nThe architecture and benchmark deserve serious referee time. But the evaluation needs matched baselines, variance estimates, and a toned-down claim. Who gets value: people working on surgical video generation and world models, especially anyone who wants a baseline dataset with hierarchical masks and aligned modalities. I would not cite the numbers as evidence of state-of-the-art, but I would cite the dataset and the anchor-relative design if I worked in that subfield.\n\nRecommendation: send it to peer review, but only after the authors redo the comparison on equal footing. The core idea is good enough that this should be a revision, not a rejection.","headline":"A genuinely new surgical control adapter with a smart anchor-relative design, but the headline comparison gives Surg-UniWorld a richer condition set than the baselines, so the 'consistently outperforms' claim does not hold as stated.","tokens_in":15986,"tokens_out":2222,"would_cite":true,"duration_ms":21547,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper claims that controllable surgical video generation works best when persistent scene identity is anchored separately from optional edge, depth, and optical-flow evidence, and demonstrates this with a hierarchical anchor plus…","keywords":["surgical world model","controllable video generation","multimodal control","hierarchical semantic masks","video diffusion","instrument-tissue interaction","flow matching","laparoscopic video generation"],"falsifier":"Take the trained model and feed it masks with known, controlled corruption, for instance swapping instrument and tissue labels in a random 10% of frames or blurring region boundaries, and measure PSNR, FVD, and control adherence. If performance barely changes, the mask-accuracy premise is not load-bearing; if it collapses or reproduces the corruption, the reported advantage over baselines may be partly an artifact of unusually clean masks rather than of the anchor architecture.","tokens_in":14944,"feed_emoji":"🎥","tokens_out":7670,"duration_ms":64537,"temperature":0.7,"pith_summary":"The paper sets out to show that surgical video generation is best controlled by separating persistent scene identity from transient visual evidence. It proposes Surg-UniWorld, which builds a Hierarchical Surgical Anchor from the first frame and hierarchical instrument, foreground-tissue, and background-tissue masks, then treats edge, depth, and optical flow as optional cues interpreted relative to that anchor. On a new benchmark built from laparoscopic clips, the full-control model reports consistent gains over general and surgical baselines in pixel fidelity, temporal realism, and control adherence. The sympathetic reading of the result is that \"what the scene is\" and \"what is moving in it\" need different roles in a generative world model.","feed_headline":"Anchor-based controls lift surgical video quality by 2.75 dB","feed_subtitle":"It ties edge, depth, and flow cues to a fixed first-frame anchor, beating prior generators.","key_machinery":"The load-bearing object is Surg-ARCA, the Surgical Anchor-Relative Control Adapter. Surg-ARCA builds a Hierarchical Surgical Anchor by organizing first-frame appearance tokens into a region-aware key-value memory using masked attention over instrument, foreground-tissue, and background-tissue masks; the memory yields a dense anchor hint and a compact region anchor bank. Modality-specific experts then read edge, depth, or flow evidence against the anchor bank with region-grounded attention and anchor modulation, producing per-modality increments. A Multimodal Control Expert combines the dense anchor hint with these increments through learnable stage-wise scaling, and the resulting hints are added to selected blocks of the frozen video diffusion backbone, so removing a modality subtracts only its own contribution.","core_discovery":"The paper's central claim is that anchoring generation to the first frame and hierarchical semantic masks, then expressing each optional control as an increment relative to that anchor, prevents anatomical distortion, instrument appearance drift, and temporal inconsistency that arise when heterogeneous conditions are fused directly. In the full-control configuration, Surg-UniWorld reports PSNR 21.250, SSIM 0.722, LPIPS 0.186, FVD 92.981, and FID 7.359, improving PSNR by 2.751 dB over the strongest baseline and reducing FVD and FID by about 11.0% and 34.9%. The authors would state the discovery as: a shared anchor with contribution-preserving composition makes arbitrary subsets of edge, depth, and flow control both usable and beneficial.","pith_inferences":["If the anchor mechanism is as general as the paper implies, the same architecture could transfer to other video domains with persistent object identity, such as manipulation or driving, wherever hierarchical masks are available; this is a natural but untested extension.","The manual refinement step in the benchmark construction suggests that mask quality, rather than the diffusion backbone, may be the practical bottleneck; an automatic mask-refinement or uncertainty-weighted anchor could be tested against the current pipeline.","The contribution-preserving composition might enable incremental editing at inference time, for example adding a depth constraint to an already generated anchor-only video without regenerating from scratch; the paper does not test this.","The region-aware flow-matching weighting implies that model performance should be evaluated per region, and the paper does report such breakdowns; future comparisons should do the same to avoid masking instrument-region failures."],"forward_implications":["Generation remains possible with no optional modalities at all, because the anchor-only setting is a trained configuration, so surgical video can be produced from text, a first frame, and masks alone.","Adding each control improves its matched property: edge raises boundary adherence, depth lowers geometric error, and flow lowers motion error, while combinations give more balanced quality across instrument, tissue, and background regions.","Each auxiliary loss contributes: removing the temporal structure, control benefit, or marginal consistency objective degrades FVD or control adherence, meaning the composition behavior is trained rather than emergent.","Because modality subsets are sampled during training and composed additively, the same frozen backbone can serve multiple downstream uses, from anchor-only data augmentation to fully controlled simulation."],"supporting_citations":[{"why":"Supplies the pretrained video diffusion backbone that Surg-ARCA augments with stage-wise control hints.","marker":"[13]"},{"why":"Supplies the video-conditioning interface and context blocks that the modality experts reuse.","marker":"[15]"},{"why":"Defines the control-injection baseline that the paper extends from flat concatenation to anchor-relative experts.","marker":"[14]"},{"why":"Provides the surgical world-model family used as baselines in the prediction and transfer comparisons.","marker":"[4]"},{"why":"Provides a motion-controllable surgical generation baseline.","marker":"[8]"},{"why":"Produces the initial masks that are manually reviewed into the hierarchical instrument and tissue annotations.","marker":"[20]"},{"why":"Provides optical-flow estimates used both for clip selection and as the flow control modality.","marker":"[19]"},{"why":"Supplies the depth estimates used as the depth control condition.","marker":"[21]"},{"why":"Supplies the edge estimates used as the edge control condition.","marker":"[22]"}],"fun_headline_variants":["Anchor-relative control lifts surgical video quality by 2.75 dB","Surg-UniWorld: anchor-based controls cut surgical video FVD by 11%","One anchor to align multimodal controls for realistic surgical videos","Anchor-relative experts keep edge, depth, and flow cues from distorting anatomy","Surg-UniWorld: multimodal control experts with anchor-relative composition"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The hierarchical semantic masks that define the anchor are assumed to be accurate, complete, and temporally aligned; the paper reports manual review and refinement but no mask-quality metric such as inter-annotator agreement or error rates.","fun_headline_variants_meta":{"raw":{"variants":["Anchor-relative control lifts surgical video quality by 2.75 dB","Surg-UniWorld: anchor-based controls cut surgical video FVD by 11%","One anchor to align multimodal controls for realistic surgical videos","Anchor-relative experts keep edge, depth, and flow cues from distorting anatomy","Surg-UniWorld: multimodal control experts with anchor-relative composition"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00138,"raw_usage":{"total_tokens":5578,"prompt_tokens":926,"completion_tokens":4652,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":542,"completion_tokens_details":{"reasoning_tokens":4554}},"tokens_in":542,"tokens_out":4652,"duration_ms":28808,"temperature":1.0,"reasoning_tokens":4554,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T14:29:21.195359+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take the trained model and feed it masks with known, controlled corruption, for instance swapping instrument and tissue labels in a random 10% of frames or blurring region boundaries, and measure PSNR, FVD, and control adherence. If performance barely changes, the mask-accuracy premise is not load-bearing; if it collapses or reproduces the corruption, the reported advantage over baselines may be partly an artifact of unusually clean masks rather than of the anchor architecture.","supporting_citations":[{"cited_title":"Vace: All- in-one video creation and editing,","cited_arxiv_id":null,"evidence_quote":"Supplies the video-conditioning interface and context blocks that the modality experts reuse."},{"cited_title":"Cosmos-h-surgical: Learning surgical robot policies from videos via world modeling,","cited_arxiv_id":null,"evidence_quote":"Provides the surgical world-model family used as baselines in the prediction and transfer comparisons."},{"cited_title":"Surgsora: Object-aware diffusion model for controllable surgical video genera- tion,","cited_arxiv_id":null,"evidence_quote":"Provides a motion-controllable surgical generation baseline."},{"cited_title":"Sam 2: Segment anything in images and videos,","cited_arxiv_id":null,"evidence_quote":"Produces the initial masks that are manually reviewed into the hierarchical instrument and tissue annotations."},{"cited_title":"Holistically-nested edge detection,","cited_arxiv_id":null,"evidence_quote":"Supplies the edge estimates used as the edge control condition."}],"review_version":2}