{"id":"27228c3e-70ee-44ea-84fe-c1526100b0b8","arxiv_id":"2605.27589","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"What-If World is a new paired-prompt benchmark showing that nine state-of-the-art video generation models achieve at most 52% on causal intervention tests and cluster near 28% for open-source systems.","lead":"The paper introduces What-If World, a benchmark of 319 prompt pairs from real driving and manipulation videos that tests whether video generation models produce physically correct differences when one causal detail is changed. A smart generalist should read it because current world models for robotics and self-driving may generate plausible individual videos yet fail basic tests of causal sensitivity needed for planning.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.3","headline":"Benchmark performance tracks visual prominence (14.2% subtle vs 40.4% pronounced), so low scores may reflect change detection limits rather than causal physics failures.","rationale":"The concern matches the reader's weakest_assumption exactly; the abstract itself supplies the supporting observation about visual prominence. This is the single most load-bearing risk to interpreting the headline numbers as evidence of causal-understanding deficits rather than perceptual ones. No other internal inconsistency appears in the reported results.","tokens_in":1818,"tokens_out":370,"duration_ms":20862,"concrete_test":"Bin the 319 pairs by a quantitative visual-salience proxy (mean absolute pixel difference between the two source frames or estimated optical-flow magnitude of the described intervention); recompute mean paired APEO Outcome score per bin. If low-salience bin averages <20% while high-salience >35% and explains >50% of score variance, the causal-interpretation claim requires qualification.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central claim interprets the paired APEO scores (max 52%, open-source ~28%) as evidence that models fail on causal interventions across the six-variable taxonomy. However, the abstract states that scores track visual prominence of the intervention rather than physics tractability. This creates an alternative explanation: models may simply under-detect low-salience visual differences between the paired outputs, even when the underlying physics is modeled correctly. The APEO Outcome component requires videos to \"end in the correct difference,\" but without an independent measure separating perceptual salience from causal inference (e.g., via controlled visual-only baselines or metadata-derived intervention magnitude from nuScenes/DROID), the rubric cannot isolate the intended causal construct. The prompt-pair construction (small wording change, large physical difference) does not rule out this confound.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.3","summary":"The paper introduces What-If World, a benchmark of 319 prompt pairs constructed from real frames in nuScenes and DROID, organized by a six-variable taxonomy of physical variables relevant to driving and manipulation. It proposes the APEO four-part rubric (Adherence, Physics, Environment, Outcome) to score whether video generation models produce outputs that correctly diverge under small prompt changes encoding causal interventions. Evaluation across nine state-of-the-art models reports a maximum paired score of 52% (open-source models cluster near 28%), with every model failing on a substantial fraction of interventions; the abstract notes that scores track visual prominence of the change (14.2% subtle vs. 40.4% pronounced) rather than physics tractability.","tokens_in":1993,"tokens_out":671,"duration_ms":41761,"significance":"If the APEO rubric and paired design can be shown to isolate causal inference from visual salience and prompt artifacts, the benchmark would offer a useful addition to existing single-video evaluation protocols for world models in embodied settings. The use of real-world source frames and the explicit taxonomy provide concrete grounding; the reported performance ceiling and prominence correlation are empirical observations that could guide future model development if the causal interpretation holds after addressing potential confounds.","major_comments":[{"comment":"Abstract: The paper states that 'performance appears to track the visual prominence of the intervention rather than the tractability of its underlying physics' and reports 14.2% on subtle vs. 40.4% on pronounced interventions, yet interprets the overall low paired scores (max 52%) as evidence that 'every model tested fails on a large fraction of causal interventions.' This leaves open the alternative that the Outcome component primarily measures change detection rather than causal modeling; without an independent visual-salience baseline or intervention-magnitude metadata from the source datasets, the central claim that the benchmark isolates causal understanding is not fully supported.","section":"Abstract"},{"comment":"Evaluation protocol: The abstract describes the APEO rubric and aggregate scores but provides no information on inter-rater reliability, how the 319 prompt pairs were validated for physical correctness, or statistical significance testing of the reported differences and the 52% ceiling. These omissions are load-bearing for the reliability of the benchmark results and the claim that open-source models cluster near 28%.","section":"Abstract and Results"}],"minor_comments":[{"comment":"The specific nine models evaluated are referenced only as 'state-of-the-art' in the abstract; naming them and providing per-model breakdowns would improve transparency and allow readers to assess whether the 52% ceiling is driven by particular architectures.","section":"Abstract"},{"comment":"The six-variable taxonomy is motivated but would benefit from a table listing one concrete prompt-pair example per variable to illustrate the 'small wording change, large physical difference' construction.","section":"Introduction"}],"recommendation":"major_revision","confidential_remarks":"The manuscript's fit to a methods-oriented vision journal is reasonable, but the evaluation protocol details (reliability, validation) are unusually sparse even for a benchmark paper; this may warrant requesting supplementary material on rater agreement and pair construction before acceptance."},"author_rebuttal":{"model":"grok-4.3","summary":"We thank the referee for the constructive comments. We address each major point below and commit to revisions that improve the manuscript's rigor without altering its core claims.","responses":[{"response":"The manuscript already reports the prominence correlation as an empirical finding and frames the low paired scores as evidence of failure on causal interventions because the Outcome criterion requires the specific physics-predicted divergence, not merely the presence of change. We nevertheless agree that the current design does not fully exclude a change-detection account. In revision we will (1) add a limitations paragraph explicitly discussing this confound and (2) incorporate any available scene-change metadata from nuScenes and DROID. The abstract will be rephrased to reflect this nuance while retaining the reported numbers.","revision_made":"yes","referee_comment":"[Abstract] Abstract: The paper states that 'performance appears to track the visual prominence of the intervention rather than the tractability of its underlying physics' and reports 14.2% on subtle vs. 40.4% on pronounced interventions, yet interprets the overall low paired scores (max 52%) as evidence that 'every model tested fails on a large fraction of causal interventions.' This leaves open the alternative that the Outcome component primarily measures change detection rather than causal modeling; without an independent visual-salience baseline or intervention-magnitude metadata from the source datasets, the central claim that the benchmark isolates causal understanding is not fully supported."},{"response":"We agree these details are necessary. The prompt pairs were constructed and cross-checked by the authors against the source frames and the six-variable taxonomy; we will expand the methods section with a precise description of this validation procedure. We will also report inter-rater agreement statistics for the APEO scoring and add statistical tests (bootstrap confidence intervals or appropriate significance tests) for the subtle-versus-pronounced difference and for the model-performance comparisons.","revision_made":"yes","referee_comment":"[Abstract and Results] Evaluation protocol: The abstract describes the APEO rubric and aggregate scores but provides no information on inter-rater reliability, how the 319 prompt pairs were validated for physical correctness, or statistical significance testing of the reported differences and the 52% ceiling. These omissions are load-bearing for the reliability of the benchmark results and the claim that open-source models cluster near 28%."}],"tokens_in":1583,"tokens_out":500,"duration_ms":38896,"standing_objections":[]},"desk_editor":{"model":"grok-4.3","letter":"The main takeaway is that this paper introduces paired prompts to test whether video models change their output when a single physical variable is altered, and it shows no model clears 52% on the combined score. That paired structure with a shared six-variable taxonomy across driving and manipulation scenes is the genuinely new element. Prior single-video benchmarks cannot catch the case where both outputs look individually plausible yet fail to reflect the intended difference, so the APEO rubric (Adherence, Physics, Environment, Outcome) directly targets that gap. Applying it to 319 pairs drawn from nuScenes and DROID gives a concrete, cross-domain comparison of nine models that is useful for anyone thinking about world models for planning.\n\nThe work does a solid job documenting the scale of the problem: open-source models sit near 28% and every system fails on a substantial fraction of interventions. The taxonomy and the explicit four-part scoring make the evaluation more structured than ad-hoc checks.\n\nThe soft spot is the visual-prominence result. The abstract states that subtle interventions score 14.2% while pronounced ones reach 40.4%, and that performance tracks salience rather than physics tractability. This supplies a plausible alternative account: the models may simply be poor at registering small visual differences between the paired videos, even when the underlying dynamics are modeled correctly. Without an independent control that separates perceptual detection from causal inference, the claim that the low scores demonstrate causal misunderstanding does not land cleanly. Construction and validation details for the prompt pairs and the rubric are also thin in the provided text, which leaves the evaluation harder to reproduce or stress-test.\n\nThe paper is aimed at researchers building or benchmarking generative world models for robotics and driving. Anyone evaluating causal sensitivity in video simulators will find the design worth examining, even if the current numbers need tighter controls before they can be read as pure evidence of causal failure. It deserves peer review because the core idea of testing differential response is worth refining and the empirical sweep across models is worth referee scrutiny.","headline":"The paired benchmark design is a clear step forward but the causal-failure interpretation is weakened by the visual-prominence pattern the paper itself reports.","tokens_in":2532,"tokens_out":474,"would_cite":false,"duration_ms":36962,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.3","headline":"Video generation models fail to adjust outputs correctly when prompts differ by one physical variable, with none exceeding 52% paired accuracy.","keywords":["causal benchmark","world models","video generation","embodied scenarios","physical variables","prompt pairs","paired scoring","APEO rubric"],"falsifier":"A model that achieves paired scores above 70% across interventions of varying visual prominence while maintaining high individual video quality would challenge the reported performance limits and failure rates.","tokens_in":2720,"feed_emoji":"🎥","tokens_out":718,"duration_ms":43927,"temperature":0.7,"pith_summary":"The paper establishes a benchmark to evaluate whether video generation models function as causal world simulators in embodied settings by creating pairs of prompts that differ in only one physical detail and checking if the generated videos reflect the expected physical difference. A sympathetic reader would care because reliable world models are needed for planning in driving and robotics, yet current benchmarks score videos individually and cannot detect when models produce plausible but causally incorrect videos. The evaluation uses a taxonomy of six physical variables and a four-part scoring rubric. Results indicate that all tested models struggle, with performance appearing to track visual prominence rather than the tractability of the underlying physics.","feed_headline":"Video models miss causal changes in what-if prompts","feed_subtitle":"Benchmark of 319 pairs shows top models at 52% accuracy, failing when prompts vary one physical detail.","key_machinery":"The What-If World benchmark of 319 prompt pairs differing by one physical variable, evaluated via the APEO rubric that detects causal response failures missed by single-video scoring.","core_discovery":"We introduce What-If World, 319 prompt pairs built on real frames, organized by a taxonomy of six physical variables shared across driving and manipulation. Each pair is scored with APEO, a four-part rubric checking whether each video follows its prompt (Adherence), is physically consistent (Physics), preserves the shared scene (Environment), and ends in the correct difference (Outcome). Across nine state-of-the-art models, no system exceeds 52% on the paired score, and open-source models cluster near 28%. Every model tested fails on a large fraction of causal interventions, indicating substantial room before these models can reliably support action-conditioned simulation or model-based plan","pith_inferences":["The benchmark could be extended to test whether models handle sequences of multiple interventions or longer time horizons.","Training approaches that explicitly reward correct causal divergence might address the gap between visual plausibility and physical accuracy.","Success on this paired evaluation would be a necessary condition for using such models in closed-loop control tasks."],"forward_implications":["No tested model can yet reliably support action-conditioned simulation or model-based planning.","Every model fails on a large fraction of causal interventions.","Performance tracks visual prominence of the intervention rather than tractability of the physics.","Visually subtle interventions score as low as 14.2% while pronounced ones reach 40.4%."],"fun_headline_variants":["Video models fail on causal what-if prompt pairs","No model tops 52% in What-If World benchmark","Open-source models average 28% on physical change tests","What-If World benchmark finds gaps in video simulation","Models miss physics in single variable prompt variations"],"cache_read_input_tokens":64,"weakest_assumption_plain":"The assumption that the APEO rubric and the 319 prompt pairs provide an unbiased measure of causal understanding rather than being driven by visual salience or prompt artifacts.","fun_headline_variants_meta":{"raw":{"variants":["Video models fail on causal what-if prompt pairs","No model tops 52% in What-If World benchmark","Open-source models average 28% on physical change tests","What-If World benchmark finds gaps in video simulation","Models miss physics in single variable prompt variations"]},"model":"grok-4.3","cost_usd":0.009495,"raw_usage":{"total_tokens":4305,"prompt_tokens":799,"num_sources_used":0,"completion_tokens":74,"cost_in_usd_ticks":94949500,"prompt_tokens_details":{"text_tokens":799,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":3432,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":799,"tokens_out":74,"duration_ms":44787,"temperature":1.0,"reasoning_tokens":3432,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-06-29T18:19:26.650244+00:00","model_set":{"reader":"grok-4.3"},"falsifier":"A model that achieves paired scores above 70% across interventions of varying visual prominence while maintaining high individual video quality would challenge the reported performance limits and failure rates.","supporting_citations":[],"review_version":1}