{"id":"bb57bf00-8934-4dc1-9964-fd4331a93362","arxiv_id":"2608.09298","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"WorldSimProbe is a five-suite benchmark showing that six action-conditioned world models systematically degrade in action-to-motion fidelity and interaction grounding across RoboTwin, ManiSkill, and LIBERO.","lead":"This paper introduces a diagnostic benchmark that tests whether robot world models really turn supplied action commands into matching motion and physically grounded interactions, instead of just generating plausible videos. Across six open-source models and three simulators, it finds systematic failures in action realization, interaction grounding, and interaction dynamics.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Task 5's VLM judge is validated only on simulator references, but the main text claims 93–95% human validation on model generations, which the supplement does not document; the interaction-dynamics findings may therefore be evaluator artifacts.","rationale":"The reader's stated weakest assumption is the reliance on simulator replay as the physical ground truth. That is a legitimate scope caveat, but it does not undercut the paper's actual claim, which is about fidelity 'in simulation': within a chosen simulator, simulator replay is the correct reference for whether the model reproduces that simulator's transitions. The more immediately decisive gap is an internal inconsistency in the Task 5 evidence chain: the main text asserts human validation on model generations, but the supplement only documents human validation on simulator references. Since Task 5 is one of the two novel interaction-focused suites and directly supports the 'structured failures in interaction dynamics' claim, the missing validation is load-bearing. The proposed test would settle whether the VLM artifact-risk is real: if the judge agrees with humans on generated rollouts, the concern is resolved; if not, the primitive-level rankings and the near-zero shake result would need revision. This concern is consistent with the reader's overall CONDITIONAL verdict, so I do not recommend moving the verdict; it sharpens the condition under which the paper's strongest interaction-dynamics claims should be accepted.","tokens_in":21095,"tokens_out":4536,"duration_ms":46529,"concrete_test":"Sample roughly 300 generated Task 5 rollouts stratified by model, simulator, and primitive; obtain three independent human primitive labels per rollout using the same forced-choice protocol as Section B.5; compare the VLM's accepted predictions against the human majority and report per-primitive agreement. If agreement on generated rollouts is comparable to the 94–97% reference-level agreement and 'shake' remains near zero, the RQ3 conclusions survive this check. If agreement is materially lower or concentrated in specific primitives, recompute Task 5 scores and Figure 6 using the human labels on the sampled rollouts (or after correcting/recalibrating the judge) and verify whether the primitive ranking and the 'shake absent' result persist.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The interaction-dynamics conclusions (RQ3, Fig. 6, and the T5 column of Table 2) are computed by Qwen3-VL-8B classifying each generated rollout into one of eight primitives. Section 5.4 states that 'Human validation on model generations yields 93–95% agreement, supporting these findings,' yet the only human validation described in Section B.5 is on 1,864 simulator-reference Task 5 rollouts, where VLM agreement with majority human labels is 94–97%. No human-annotation protocol, dataset, or agreement statistic for generated model rollouts appears anywhere in the supplement. This gap is load-bearing because reference rollouts are clean simulator renders with correct, physically grounded dynamics, whereas the generated rollouts being scored contain exactly the blur, deformation, missing contacts, static frames, and motion artifacts this benchmark is designed to expose. Reference-level VLM accuracy need not transfer to these degraded inputs. In particular, the near-zero 'shake' scores, the primitive-level ordering (tap > drag/drop/push > pull/rotate/knock-over), and the headline claim of 'structured failures in interaction dynamics' would be substantially weakened if the VLM systematically mislabels degraded model rollouts, for instance by classifying a failed shake attempt as static or as a tap. This is a missing validation step on the exact inputs used to compute Task 5 scores.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"WorldSimProbe proposes the Observable Simulator Contract, which states that a world model must map supplied actions to corresponding agent motion and then to physically grounded environment responses. The paper operationalizes this contract with five suites—local action calibration, global trajectory coverage, action-source behavior preservation, interaction grounding, and interaction dynamics—and evaluates six open-source action-conditioned world models on 18,608 filtered instances across RoboTwin, ManiSkill, and LIBERO. Suite-level metrics are calibrated against simulator references only, with thresholds fixed before model inference. The results show attenuation and compression of action responses under local, global, and source variation, false-positive interaction grounding, and primitive-dependent dynamics failures, plus human-judgment and downstream synthetic-data checks that are broadly consistent with the benchmark scores.","tokens_in":21515,"tokens_out":6629,"duration_ms":70649,"significance":"The benchmark is a timely and well-structured contribution: it moves ACWM evaluation from visual- and task-level scores toward capability-level diagnostics, and it is careful about reference-side validation, including simulator-replay filtering, fixed thresholds, and a separate downstream study. The large filtered manifest, the oracle-relative Task 1 construction, and the masked-flow alignment for Tasks 2–3 are sensible and reduce circularity. If the results hold, the central claim—that current ACWMs do not satisfy the Observable Simulator Contract even in simulation—is meaningful for the field. The main reservation is that the Task 5 interaction-dynamics findings rely on a VLM judge whose agreement was demonstrated only on clean simulator references, not on the degraded model generations it actually scores; I therefore treat RQ3 as not yet fully supported.","major_comments":[{"comment":"The main text states that 'Human validation on model generations yields 93–95% agreement, supporting these findings' (§5.4), but the only human annotation described in §B.5 is on 1,864 simulator-reference Task 5 rollouts, for which VLM–majority-human agreement is 94–97%. No protocol, dataset, or agreement statistic for generated model rollouts appears anywhere in the supplement. This distinction is load-bearing: reference rollouts are clean simulator renders with physically grounded dynamics, whereas the rollouts scored for Figure 6 and the T5 column of Table 2 contain the blur, deformation, missing contacts, and static frames the benchmark is designed to expose. Reference-level VLM accuracy does not establish accuracy on those degraded inputs; the near-zero 'shake' scores and the primitive ordering (tap > drag/drop/push > pull/rotate/knock-over) are exactly the findings that could change if the judge systematically mislabels degraded rollouts. Please add human labels on generated rollouts, report the resulting agreement, or visibly restrict the RQ3 claims to what the reference-level validation supports.","section":"§5.4, Supplement §B.5, Eq. (S12)"},{"comment":"Table 2 reports no confidence intervals, per-seed standard deviations, or paired comparisons for the headline cross-model and cross-suite scores, yet §5.2 makes precise ordering claims such as 'LingBot-VA leads RoboTwin and ManiSkill' and 'Ctrl-World leads Interaction Grounding on RoboTwin and LIBERO.' Several adjacent entries differ by only a few points (for example, the Task 5 RoboTwin scores of IRASim and BWM are 20.5 and 20.6, respectively), so without uncertainty estimates or a paired test these fine-grained rankings are not assessable. Please add confidence intervals or significance tests, or soften the ranking language to descriptive comparison.","section":"Table 2, §5.2"}],"minor_comments":[{"comment":"The main text describes a motion gate that checks robot-arm centroid displacement across three frames, but Eq. (S11) uses a 60-pixel threshold in a 256×256 coordinate system; please state the frame-selection procedure and the threshold units in the main text for reproducibility.","section":"§4, Task 4 and Eq. (S11)"},{"comment":"The caption says values are percentages and shading encodes magnitude, but no legend connects shading to value; print the exact values or add a color bar.","section":"Figure 6"},{"comment":"Section A.5 gives the exact final test-manifest total of 18,608, while §5.1 says 'approximately'; use the exact count in the main text.","section":"§5.1 and §A.5"},{"comment":"The Task 5 evaluator prompt collects artifact_flags and physical-plausibility scores, but the paper does not report how often artifacts were flagged or whether the VLM's primitive prediction changed when artifacts were present; this information would help interpret the near-zero shake scores.","section":"Figure S7"},{"comment":"The related-work comparison would benefit from a footnote defining each column check, since 'Causal Probe' and 'Interaction Decomp.' are explained only in the caption text.","section":"Table 1"}],"recommendation":"major_revision","confidential_remarks":"To the editor: the paper is within scope and the experimental design is mostly sound. The one load-bearing gap is the missing validation of the Task 5 VLM on generated rollouts, which can be fixed by adding a human-annotation study on model outputs or by restricting RQ3. I do not see circularity in the benchmark design; the simulator-reference anchoring is appropriate for a simulator-faithfulness benchmark. My recommendation is major_revision rather than reject because the issue is local and fixable."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The thing to know: this paper gives the field a genuinely useful diagnostic tool. Instead of another task-success or video-quality benchmark, it formalizes an Observable Simulator Contract and builds five suites that trace the chain from supplied action to agent motion to interaction response. The RMFA flow-alignment evaluator, the false-interaction grounding protocols, and the primitive-level dynamics breakdown are real contributions. The experimental design is mostly careful: reference filtering uses simulator-side data only, thresholds are fixed from reference data, the Task 1 oracle-relative score is sensible, and the downstream OOD study is a nice validity check. I think the central argument holds up: current ACWMs do degrade systematically under control variation and do hallucinate unsupported interactions.\n\nThe soft spot is exactly where the stress test points. Section 5.4 claims \"Human validation on model generations yields 93–95% agreement,\" but the supplement only documents human annotation on 1,864 simulator-reference Task 5 rollouts. There is no human protocol or agreement statistic for generated rollouts anywhere. This is load-bearing because the generated videos contain the very artifacts (blur, deformation, missing contacts, static frames) that Task 5 is meant to expose. Reference-level VLM accuracy does not imply accuracy on degraded inputs, so the near-zero shake scores and the primitive ordering could be evaluator artifacts. The authors need to either collect human labels on model generations or substantially soften the claim. Minor issues: Table 2 has no confidence intervals, and Task 4 only tests false-positive grounding, though the 50-case audit justifies that choice. The simulator-as-ground-truth limitation is acknowledged only as future work; it is a scope limit, not a fatal flaw.\n\nThis paper deserves a serious referee, but not a clean accept. The fix is concrete and well-scoped: validate the VLM judge on the actual model outputs, add variance reporting, and be upfront that all scores are simulator-relative. I would cite this for the benchmark design once the Task 5 gap is closed.","headline":"A well-designed diagnostic benchmark for action-conditioned world models, held back by one load-bearing validation gap: the Task 5 VLM judge is checked on simulator references, not on the degraded model rollouts it actually scores.","tokens_in":21967,"tokens_out":1304,"would_cite":true,"duration_ms":15273,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"WorldSimProbe establishes that six open-source action-conditioned world models systematically violate a two-link simulator contract: supplied actions fail to induce calibrated motion, and environment responses are often not grounded in…","keywords":["action-conditioned world models","simulator faithfulness","embodied manipulation","world model benchmark","action-to-motion correspondence","interaction grounding","contact hallucination","robotic manipulation evaluation"],"falsifier":"Run the same five WorldSimProbe suites on real-robot rollouts with motion capture and contact sensors instead of simulator replay; if models that score poorly in simulation achieve high action-to-motion and grounding fidelity on hardware, the contract violations would be artifacts of the simulator reference rather than properties of the models.","tokens_in":20909,"feed_emoji":"🤖","tokens_out":9662,"duration_ms":78297,"temperature":0.7,"pith_summary":"The paper argues that action-conditioned world models should be judged as physical simulators, not just video generators: a supplied action stream must induce the matching robot motion, and any environment change must be caused by that motion through real contact. It formalizes this requirement as the Observable Simulator Contract and builds WorldSimProbe, five controlled test suites that intervene on action magnitude, global trajectory variation, control source, contact validity, and interaction primitives. Across six open-source models and more than 18,000 simulator-validated instances on RoboTwin, ManiSkill, and LIBERO, the benchmark finds systematic action-realization degradation under control variation, unsupported interactions, and weak primitive-level dynamics. If the results hold, current action-conditioned world models do not yet satisfy the contract even inside simulation, and standard task-outcome or visual-quality scores hide where the simulator chain breaks.","feed_headline":"Robot world models fail the action-to-motion link","feed_subtitle":"A five-suite benchmark shows six open-source robot world models drop action fidelity under variation.","key_machinery":"The organizing object is the Observable Simulator Contract, which decomposes faithfulness into two consistency conditions: $\\hat{r}_{t:t+H} \\approx \\Phi_R(r_t, e_t, a_{t:t+H})$ for action-realization and $\\hat{e}_{t:t+H} \\approx \\Phi_E(e_t, \\hat{r}_{t:t+H})$ for interaction-response, where $\\Phi_R$ and $\\Phi_E$ are ideal but generally unobserved operators. The benchmark machinery is the five-suite probe chain that tests observable consequences of those operators without accessing them directly: perturbation-response ratios for local calibration, cross-task receiver–donor replay for global trajectory coverage, source-diverse human and policy trajectories for behavior preservation, no-contact interventions with tracked object displacement for interaction grounding, and eight scripted interaction primitives judged by a vision-language model for dynamics. Each suite reports its own score so that a failure can be localized to action realization, grounding, or dynamics rather than being absorbed into one aggregate rollout score.","core_discovery":"The paper's central claim is that simulator faithfulness for action-conditioned world models reduces to two observable links: the supplied action must be realized as corresponding agent motion, and the environment response must be physically supported by that realized motion and its contact events. WorldSimProbe operationalizes this claim with five suites, each constructing simulator-executable interventions whose references come from simulator replay rather than from task success. Evaluated on six open-source models over more than 18,000 instances across three platforms, the benchmark shows that action-realization fidelity degrades as control variation increases, that interaction-grounding failures are dominated by contact hallucination, and that interaction-dynamics fidelity is low and primitive-specific. The benchmark's scores agree with human action-following judgments and with downstream policy success under out-of-distribution controls, supporting the diagnosis that the observed failures are real properties of current models rather than artifacts of a single metric.","pith_inferences":["The two-link contract suggests a deployment-time audit procedure: inject labeled action perturbations into any world model and check whether realized motion and contact responses co-vary, even without simulator ground truth, using human labels or inverse models as references.","The systematic degradation under source-diverse actions implies that training distributions matter more than architecture for simulator faithfulness; a testable extension would train a model on counterfactual and multi-source trajectories and measure whether Tasks 2 and 3 scores rise.","If contact hallucination transfers to real-world rollouts, it would predict a specific planning failure mode: models would act as if grasps and pushes occurred that did not, making them unsafe as policy evaluators without external contact verification.","A natural extension is to convert the five-suite decomposition into a training objective, calibrating action-response ratios, enforcing cross-task coverage, and grounding responses in contact, rather than using the suites only for evaluation."],"forward_implications":["Task-outcome and visual-quality scores overstate action-conditioned world model reliability, because a rollout can look action-aware while its realized motion is miscalibrated relative to the simulator reference.","Action realization degrades systematically as control variation increases: fidelity falls with receiver–donor motion mismatch (mean Spearman $\\rho = -0.433$) and all six models score higher on late-policy than early-policy checkpoints.","Interaction-grounding failures are dominated by false positives rather than missed interactions, with appearance-induced false-contact cases being the hardest condition for every model family.","Interaction-dynamics fidelity is primitive-specific and weak overall, with shake near zero for all models and pull, rotate, and knock-over clearly weaker than tap, drag, drop, and push.","The diagnosis carries downstream consequences: in the controlled synthetic-data study, policies trained on higher-fidelity generated rollouts separate clearly under OOD controls (up to 53% success versus 21%) while remaining similar under standard controls."],"supporting_citations":[{"why":"Provides the RoboTwin simulator and the 50 dual-arm manipulation tasks that supply the reference rollouts and most Task 1–3 instances.","marker":"Mu et al. 2025"},{"why":"Provides the ManiSkill simulator and single-arm Panda reference trajectories used across all five suites.","marker":"Tao et al. 2024"},{"why":"Provides the LIBERO task suite and reference trajectories that anchor the third platform in the benchmark.","marker":"Liu et al. 2023"},{"why":"Supplies IRASim, one of the six action-conditioned world model baselines whose failure modes the benchmark diagnoses.","marker":"Zhu et al. 2024"},{"why":"Supplies Ctrl-World, a controllable generative baseline evaluated in all suites and the best cross-platform overall performer on the macro score.","marker":"Guo et al. 2025"},{"why":"Supplies BWM, one of the six baselines evaluated across the full benchmark.","marker":"Boundless Large Model 2026"},{"why":"Supplies DreamDojo, a generalist human-video-trained baseline evaluated across the five suites.","marker":"Gao et al. 2026"},{"why":"Supplies LingBot-VA, the unified action–video baseline that leads RoboTwin and ManiSkill overall.","marker":"Li et al. 2026c"},{"why":"Supplies Cosmos-3-Nano, the unified action–video baseline used as a low-fidelity comparator in the downstream study.","marker":"Agarwal et al. 2026"}],"fun_headline_variants":["World models lose the link between action and motion","Five-suite probe shows world models drop action fidelity","Robot world models fail the action-to-motion test","WorldSimProbe reveals contact hallucination in world models","Six open-source world models flunk simulator-faithfulness check"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The benchmark treats simulator replay in RoboTwin, ManiSkill, and LIBERO as the ground-truth physical reference for every validity check, threshold, and score, so if those simulators are not faithful to real-world manipulation, the diagnosed failure modes and model rankings inherit their error.","fun_headline_variants_meta":{"raw":{"variants":["World models lose the link between action and motion","Five-suite probe shows world models drop action fidelity","Robot world models fail the action-to-motion test","WorldSimProbe reveals contact hallucination in world models","Six open-source world models flunk simulator-faithfulness check"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000296,"raw_usage":{"total_tokens":1738,"prompt_tokens":987,"completion_tokens":751,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":603,"completion_tokens_details":{"reasoning_tokens":673}},"tokens_in":603,"tokens_out":751,"duration_ms":7171,"temperature":1.0,"reasoning_tokens":673,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T19:53:22.602352+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same five WorldSimProbe suites on real-robot rollouts with motion capture and contact sensors instead of simulator replay; if models that score poorly in simulation achieve high action-to-motion and grounding fidelity on hardware, the contract violations would be artifacts of the simulator reference rather than properties of the models.","supporting_citations":[],"review_version":1}