{"id":"04786494-7ff4-4e55-bd3e-7ac7b4eb737d","arxiv_id":"2605.29339","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":7.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":0,"one_line_summary":"Introduces DMC-CF benchmark from real videos for multimodal causal counterfactual QA, with DGI dynamic evaluation, showing current MLLMs require substantial improvement in real-world causal reasoning.","lead":"The paper introduces DMC-CF, a benchmark built from real-world videos to test multimodal AI models on causal counterfactual reasoning, including a dynamic version using causal graphs to reduce data contamination. A smart generalist might read it to see how far current models are from understanding real cause-and-effect relationships needed for reliable AI.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.3","headline":"No significant objection identified","rationale":"Reader's weakest assumption aligns with the only plausible external validity risk, but the abstract itself supplies no evidence of internal failure in that assumption. Full-text details on graph construction or DGI validation are absent here, yet nothing in the given text creates a load-bearing flaw in the argument as stated.","tokens_in":1693,"tokens_out":255,"duration_ms":14903,"concrete_test":"Re-run the reported MLLM evaluations on a 10% random subsample of DMC-CF-Dynamic questions after confirming (via independent annotation) that each causal graph matches the source video events; if average accuracy shifts by >15% the headline claim on model limitations would require re-examination.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim rests on experimental results from DMC-CF showing MLLMs need improvement in real-world causal reasoning. The abstract describes collection of real-world videos, construction of causal graphs, and use of DGI to create a dynamic benchmark that mitigates contamination. No internal inconsistency, unstated assumption, or methodological gap is detectable from the provided abstract that would undermine the reported conclusion; the benchmark construction is presented as directly addressing prior limitations.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.3","summary":"The paper introduces DMC-CF, a benchmark for multimodal counterfactual QA on causal reasoning. DMC-CF-Static is constructed from real-world videos represented via causal graphs; DMC-CF-Dynamic is derived from it via the Dynamic Graph Intervention (DGI) framework to enable dynamic evaluation that mitigates data contamination. Experiments on the combined benchmark conclude that current MLLMs still require substantial improvement in real-world multimodal causal reasoning.","tokens_in":1766,"tokens_out":481,"duration_ms":18989,"significance":"If the benchmark construction, validation, and experimental results hold, the work supplies a large-scale, real-world resource that improves on prior synthetic or small-scale causal-reasoning datasets and introduces a dynamic intervention mechanism to reduce contamination. This could become a useful standard for diagnosing limitations of statistical learning in MLLMs.","major_comments":[{"comment":"The central claim rests on experimental results showing MLLM limitations, yet the manuscript provides no quantitative tables, model names, accuracy numbers, or error analysis in the abstract or visible sections; without these, the strength of the conclusion cannot be evaluated.","section":"Abstract / Results"},{"comment":"The claim that DGI effectively addresses data contamination (weakest assumption) is load-bearing for DMC-CF-Dynamic; the description does not include concrete validation such as contamination-rate measurements before/after intervention or comparison against static baselines.","section":"DGI framework description"},{"comment":"The assertion that real-world videos yield an unbiased measure of causal reasoning requires supporting statistics (e.g., diversity of scenes, causal-graph complexity distribution, or inter-annotator agreement on graph construction); these are not referenced.","section":"Benchmark construction"}],"minor_comments":[{"comment":"Notation for causal graphs and intervention operators should be defined explicitly with an example figure early in the paper.","section":"Notation / §2"},{"comment":"The abstract states the benchmark is 'large-scale' but supplies no size statistics (number of videos, QA pairs, graph nodes); these should appear in a table in the main text.","section":"Abstract / Dataset statistics"}],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":{"model":"grok-4.3","summary":"We thank the referee for the constructive comments, which identify opportunities to strengthen the presentation of quantitative results and the empirical validation of our methods. We respond to each major comment below.","responses":[{"response":"We agree that the abstract would benefit from explicit quantitative highlights to allow immediate evaluation of the central claims. Detailed results, including the specific MLLMs evaluated, accuracy numbers, and error analysis, appear in Section 4 and the accompanying tables. We will revise the abstract to include a concise summary of key performance metrics.","revision_made":"yes","referee_comment":"[Abstract / Results] The central claim rests on experimental results showing MLLM limitations, yet the manuscript provides no quantitative tables, model names, accuracy numbers, or error analysis in the abstract or visible sections; without these, the strength of the conclusion cannot be evaluated."},{"response":"Section 3.3 describes the DGI mechanism for generating dynamic counterfactual questions via graph interventions. We acknowledge that explicit before-and-after contamination measurements and static-baseline comparisons would provide stronger support. We will add a dedicated validation subsection with these quantitative analyses in the revised manuscript.","revision_made":"yes","referee_comment":"[DGI framework description] The claim that DGI effectively addresses data contamination (weakest assumption) is load-bearing for DMC-CF-Dynamic; the description does not include concrete validation such as contamination-rate measurements before/after intervention or comparison against static baselines."},{"response":"Section 3.1 outlines the real-world video collection and causal-graph construction process. To directly address the request for supporting statistics, we will add a summary table and explicit references to scene diversity, graph-complexity distributions, and inter-annotator agreement in the main text.","revision_made":"yes","referee_comment":"[Benchmark construction] The assertion that real-world videos yield an unbiased measure of causal reasoning requires supporting statistics (e.g., diversity of scenes, causal-graph complexity distribution, or inter-annotator agreement on graph construction); these are not referenced."}],"tokens_in":1331,"tokens_out":454,"duration_ms":24854,"standing_objections":[]},"desk_editor":{"model":"grok-4.3","letter":"The paper's main contribution is DMC-CF, a large-scale benchmark drawn from real-world videos for testing multimodal causal counterfactual reasoning, along with the Dynamic Graph Intervention framework for a dynamic evaluation set.\n\nThis moves past the synthetic or small-scale datasets that have been common. Representing events with causal graphs and intervening on them for dynamic tests is a distinct way to handle contamination.\n\nThe paper does a solid job of motivating the need for more realistic data and proposing a framework to address evaluation issues.\n\nThe central claim is that current MLLMs still require substantial improvement in real-world causal reasoning, based on results from both the static and dynamic parts of the benchmark.\n\nWhere it is thinner is in the details we can see. The abstract mentions the collection and construction but provides no information on methodology for building the graphs, ensuring quality, or the actual model results and error analysis. That makes it hard to judge whether the benchmark truly measures what it claims or if the reported limitations in MLLMs hold up. Without seeing how they collected the videos or verified the causal graphs, it's difficult to know if the benchmark avoids its own biases or if the DGI framework actually works as intended in practice.\n\nThe assumption that real videos plus graph interventions give an unbiased test also needs scrutiny once the full construction is available.\n\nThis paper is for researchers working on multimodal large language models and causal reasoning benchmarks. Anyone building evaluations in this space could find the dynamic approach worth examining.\n\nIt should go through peer review so the construction and results can be properly checked.","headline":"Real-world video benchmark for MLLM causal reasoning with dynamic graph intervention, but thin on construction and result details.","tokens_in":2254,"tokens_out":386,"would_cite":false,"duration_ms":29385,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.3","headline":"A benchmark of real-world videos shows current multimodal large language models still lack causal counterfactual reasoning.","keywords":["multimodal large language models","causal reasoning","counterfactual questions","benchmark construction","dynamic evaluation","real-world videos","data contamination","graph intervention"],"falsifier":"A multimodal model that scores above 80 percent accuracy on the dynamic benchmark after training only on the static videos without exposure to the graph interventions would indicate that the benchmark does not fully test causal reasoning.","tokens_in":2607,"feed_emoji":"📹","tokens_out":636,"duration_ms":21642,"temperature":0.7,"pith_summary":"The paper builds DMC-CF-Static by collecting real-world videos and turning causal events into counterfactual questions to test whether multimodal models grasp cause and effect rather than surface patterns. It then introduces the Dynamic Graph Intervention framework, which encodes events as causal graphs and generates fresh question variants on the fly, creating DMC-CF-Dynamic to reduce the risk that models simply memorize training data. Results across both parts of the benchmark indicate that existing models perform poorly, suggesting statistical training alone does not produce genuine understanding of real-world causality. A sympathetic reader would care because reliable decision-making in changing environments requires models to handle interventions they have never seen before.","feed_headline":"Real-world video benchmark exposes MLLM causal reasoning shortfalls","feed_subtitle":"Models fail to answer counterfactual questions about cause and effect even when events are drawn from authentic footage.","key_machinery":"The Dynamic Graph Intervention (DGI) framework, which encodes causal events as graphs and generates dynamic question variants to create contamination-resistant evaluation sets from the static benchmark.","core_discovery":"The authors construct DMC-CF-Static from real-world videos as a large-scale benchmark for multimodal causal counterfactual reasoning and apply the Dynamic Graph Intervention framework to derive DMC-CF-Dynamic from causal graphs; experiments on the combined benchmark demonstrate that the multimodal causal reasoning capabilities of current multimodal large language models in real-world scenarios still require substantial improvement.","pith_inferences":["If the DGI method successfully blocks contamination, similar graph-based dynamic generation could be applied to other video or image reasoning tasks.","Persistent failure on the benchmark raises the possibility that scaling model size alone will not close the gap without changes to training objectives.","The benchmark could be extended by adding new video domains or by measuring how well models generalize to unseen causal graphs."],"forward_implications":["Existing multimodal models trained by statistical learning do not reliably capture underlying causal relationships in videos.","Synthetic or cartoon-based causal datasets are insufficient proxies for real-world performance.","Future model development must incorporate mechanisms that handle dynamic interventions rather than static pattern matching.","The DMC-CF benchmark supplies a concrete testbed for measuring progress toward causal understanding."],"fun_headline_variants":["Real videos create benchmark for MLLM causal counterfactuals","DMC-CF tests MLLMs using dynamic causal graph interventions","Static and dynamic evaluations reveal MLLM causal gaps","Real-world counterfactual QA benchmark for multimodal models","MLLMs lag in understanding real causal relationships multimodally"],"cache_read_input_tokens":2112,"weakest_assumption_plain":"Real-world videos collected and represented via causal graphs provide an unbiased and contamination-free measure of causal reasoning capabilities.","fun_headline_variants_meta":{"raw":{"variants":["Real videos create benchmark for MLLM causal counterfactuals","DMC-CF tests MLLMs using dynamic causal graph interventions","Static and dynamic evaluations reveal MLLM causal gaps","Real-world counterfactual QA benchmark for multimodal models","MLLMs lag in understanding real causal relationships multimodally"]},"model":"grok-4.3","cost_usd":0.003931,"raw_usage":{"total_tokens":1992,"prompt_tokens":624,"num_sources_used":0,"completion_tokens":70,"cost_in_usd_ticks":39312000,"prompt_tokens_details":{"text_tokens":624,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":1298,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":624,"tokens_out":70,"duration_ms":9776,"temperature":1.0,"reasoning_tokens":1298,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-06-29T08:34:14.502921+00:00","model_set":{"reader":"grok-4.3"},"falsifier":"A multimodal model that scores above 80 percent accuracy on the dynamic benchmark after training only on the static videos without exposure to the graph interventions would indicate that the benchmark does not fully test causal reasoning.","supporting_citations":[],"review_version":1}