{"id":"97a58776-75a9-48bd-85d4-0a62202bdb21","arxiv_id":"2608.09873","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A 1,253-example benchmark for knowledge- and reasoning-intensive science video generation, with expert rubrics, shows large gaps between perceptual realism and scientific correctness in 16 frontier models.","lead":"Sci-VBench is a new benchmark that tests whether text-to-video AI models can generate scientifically correct videos across 60 subjects, using expert-written prompts and rubrics. Across 16 models, visual quality scores cluster tightly while scientific reasoning scores vary widely, with closed-source models ahead on mechanism.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Rubric ground truth for Scientific and Causal Correctness is validated only by internal expert agreement; without an external curriculum or physical-standard check, the central 'mechanism, not appearance' claim partly measures adherence to the authors' rubrics rather than scientific truth.","rationale":"The reader's weakest_assumption and my concern converge: the benchmark's validity as a measure of scientific truth rests on expert rubrics, and the manuscript validates them only through internal consistency. I agree this is the load-bearing point because every downstream quantity—expert labels, non-expert-with-spec correlations, MLLM-judge correlations, and model rankings on PG/SCC—is defined relative to those rubrics. The internal κ values are genuinely useful and rule out the strongest form of the concern (idiosyncratic individual rubrics), which is why I do not recommend REJECT. But they cannot rule out a shared interpretation among the recruited expert pool, especially since prompt construction and rubric review happen inside the same pipeline and five annotators are authors. The textbook-based selection is an argument for external validity, but with no cited curriculum and no released artifacts it remains an assertion. The proposed independent reconstruction test would directly settle whether the released rubrics correspond to an outside standard; until then CONDITIONAL is the right verdict. I set verdict_should_be to UNCHANGED because the reader already made acceptance conditional on this type of validation plus artifact release and statistical reporting; no new adjustment is needed. I did not elevate the VT/human-LPF discrepancy or missing confidence intervals to the primary concern: those affect precision and one sub-claim, whereas rubric external validity underpins the entire benchmark.","tokens_in":22765,"tokens_out":8873,"duration_ms":86697,"concrete_test":"Select a stratified random sample of 60 Sci-VBench prompts (15 per discipline). Recruit 10–15 subject-matter experts who were not involved in Sci-VBench and have no access to the released reference guides or rubrics; ask each to write, from standard curricula and answer keys, the expected phase-based storyline and the key causal transitions for each prompt. Blind the released SCC rubrics and have independent raters compare the two sets of storyline specifications (e.g., quadratic weighted κ on the SCC dimension). If agreement with the released rubrics is below ~0.70—or if the independent experts flag material ambiguity on more than 10% of prompts—the SCC scores partly measure adherence to the authors' interpretive choices, and the central claim should be re-qualified.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim—that current systems fail at mechanism while matching on appearance—depends on Sci-VBench's Scientific and Causal Correctness scores being true measures of scientific correctness. The paper's only support for that premise is internal consistency: §3.5 audits 200 examples by having a second annotator write an independent specification and two scorers grade under both, yielding κ=0.75 for SCC; §4.3 reports expert re-rating stability κ=0.842. These numbers show the rubrics are reliably applied and partly independent of the original author, but they do not show the rubrics match an external ground truth. The annotation pipeline (§3.3) selects concepts from 'canonical textbooks and course materials,' but no textbook/curriculum list or external answer key is cited, and five of the 61 annotators are authors. If the target concepts or expected phase-based storylines embed a particular interpretation (e.g., which outcome the underspecified inclined-plane prompt should show), then expert ratings, non-expert-with-spec scores, and rubric-conditioned MLLM judgments all measure agreement with that interpretation, and the reported proprietary–open-source SCC gap may overstate or misstate genuine scientific failure. The prompt-rewriting result in §5.4—SCC improves 23–52% with more explicit wording—shows a substantial part of the measured gap is attributable to prompt under-specification, making the boundary between 'mechanism failure' and 'intended-inference failure' the central empirical quantity. This does not invalidate the benchmark, but it makes the headline claim conditional on an external validation that is currently absent.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"Sci-VBench is a new benchmark for evaluating knowledge- and reasoning-intensive text-to-video generation in scientific domains. It contains 1,253 expert-authored prompts spanning 60 subjects across Natural Science, Healthcare, Humanities & Social Sciences, and Engineering, each paired with a per-example reference guide and 1-5 scoring rubrics. The paper validates a rubric-based evaluation protocol: expert ratings serve as reference labels, non-expert raters improve agreement when given the evaluation specification, and rubric-conditioned MLLM judges correlate with expert ratings more closely than prior automatic evaluators. Sixteen proprietary and open-source models are benchmarked. The headline findings are that automatic perceptual-quality scores are nearly flat across models, while Prompt Grounding and Scientific and Causal Correctness vary substantially with a proprietary-open-source gap, leading the authors to conclude that what separates current systems is mechanism, not appearance. A prompt-rewriting ablation shows that more explicit prompts improve SCC substantially but do not close the gap.","tokens_in":22989,"tokens_out":6881,"duration_ms":69580,"significance":"If the benchmark's SCC labels are accepted as ground truth, this is an important and useful contribution: it is the first benchmark of its scope to release reusable per-example evaluation specifications, it includes a controlled human study showing that non-experts can be brought close to expert agreement, and it reports strong expert re-rating stability (Cohen's kappa = 0.842). The observed pattern, where perceptual quality is saturated but scientific/causal correctness still separates models, is a falsifiable and practically relevant statement about current video generators. The main reservation is that the reference labels are validated only through internal consistency audits; the central 'mechanism, not appearance' claim is therefore partly a claim about expert-rubric adherence rather than about an external independent standard of scientific truth.","major_comments":[{"comment":"The Scientific and Causal Correctness labels are validated only by internal consistency: the 200-example audit in §3.5 shows quadratic weighted kappa = 0.75 for SCC, and the 300-video re-rating in §4.3 shows Cohen's kappa = 0.842. These numbers demonstrate that the rubrics are applied reliably and are not idiosyncratic to one annotator, but they do not demonstrate that the rubrics correspond to an external scientific standard. Because §3.3 says target concepts are selected from 'canonical textbooks and course materials' but no textbook list or answer key is provided, and because the reference guide fixes the expected phase-based storyline, expert scores, non-expert-with-spec scores, and rubric-conditioned MLLM scores all measure agreement with the authors' operationalization. Given the abstract's claim about 'reliable modeling of scientific and causal dynamics,' the paper should either add an external audit (e.g., an independent expert panel verifying a sample of reference guides against named curriculum sources) or explicitly reframe the contribution as measuring rubric-verified scientific correctness rather than unqualified scientific correctness. Without this, the proprietary-open-source SCC gap in Table 4 could overstate differences in genuine scientific fidelity.","section":"§3.3-3.5, §4.3"},{"comment":"The prompt-rewriting experiment quantifies how much SCC depends on prompt explicitness: +23.3% for Wan2.2-5B and +51.7% for HunyuanVideo-1.5. This is a large fraction of the measured shortfall and shows that much of the SCC gap is attributable to the benchmark's deliberate design choice of omitting outcomes (§3.3) rather than purely to the generators' inability to model mechanisms. Because the rewritten condition was applied only to two open-source models, the proprietary-open-source SCC gap in Table 4 is potentially confounded by differential sensitivity to prompt under-specification. I ask for either (a) rewritten-prompt runs on at least the top proprietary models, or (b) a conditional analysis that reports the SCC gap on a subset of examples where prompt explicitness is controlled. The sentence 'it cannot supply the mechanistic fidelity the generator lacks' is not yet fully supported unless the magnitude of this confound is bounded.","section":"§5.4, Figure 4"},{"comment":"Instance-level Pearson correlations for the best rubric-conditioned MLLM judge are r = 0.545 for Spatiotemporal Consistency and r = 0.535 for Low-level Perceptual Fidelity. These are moderate correlations, not 'relatively high agreement.' Since the automatic scoring in Table 4 uses these judge scores for the SC dimension, the current automatic protocol is not yet reliable for fine-grained model ranking on Spatiotemporal Consistency. The abstract's claim that MLLM-as-Judge systems 'can achieve relatively high agreement with expert judgments, supporting reproducible evaluation at scale' is too broad if it is taken to cover these dimensions. Please report confidence intervals and the number of instances underlying Table 3, specify a threshold for 'high agreement,' and either soften the claim or restrict the scalable-evaluation claim to Prompt Grounding and Scientific and Causal Correctness, where the correlations are substantially higher.","section":"Table 3, §4.3"}],"minor_comments":[{"comment":"Table 3 reports Pearson correlations multiplied by 100; please also state the raw r values and the number of videos in the caption or text for clarity.","section":"Table 3"},{"comment":"The 200-example audit uses quadratic weighted kappa while the 300-video re-rating is reported as Cohen's kappa; please clarify whether the latter is also weighted and, if so, which weighting, since this affects comparability.","section":"§3.5 and §4.3"},{"comment":"For Gemini-Omni-Flash and Seedance-2.0, provider filters rejected some prompts and the averages use successful generations; please report the number of successful videos per model so readers can assess the impact of the missing examples.","section":"Table 4"},{"comment":"Because five of the 61 annotators are authors, a sentence in the main text stating whether authors authored examples, validated examples, or both would improve transparency beyond the appendix biographies.","section":"§3.2 and Appendix A.4"},{"comment":"The knee-jerk reflex rubric example in Figure 6 is very helpful; consider moving one complete rubric example into the main text so readers can see the protocol without consulting the appendix.","section":"Appendix B.3 and Figure 6"}],"recommendation":"major_revision","confidential_remarks":"This is a strong benchmark contribution and likely a good fit for COLM. I would not reject on the external-validity concern alone, since expert-judgment benchmarks routinely define their own ground truth, but the authors should either add an external audit or soften the truth-claim wording. The prompt-rewriting confound in §5.4 is the more serious issue because it directly bears on the headline 'mechanism, not appearance' claim; additional experiments with proprietary models under rewritten prompts would substantially strengthen the paper. No citation or novelty concerns beyond the usual need to differentiate from VideoScience-Bench."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Solid, well-built benchmark contribution with a credible main finding, but the headline claim is a bit ahead of the evidence. What's genuinely new: 1,253 expert-authored prompts across 60 subjects, with per-example reference guides and 1–5 rubrics released alongside. That's a real step beyond VideoScience-Bench and the physics commonsense suites. The rubric-based evaluation protocol is carefully constructed: non-experts with the spec reach 82.5% average correlation with experts, and the strongest MLLM judge gets 63.4% — better than the prior auto-eval methods they compare. The internal consistency audit (independent spec re-annotation, crossed scorer–rubric design) is the right kind of check, and the κ=0.75–0.79 range on reasoning dimensions is reassuring.\n\nThe main finding — VT is nearly flat across models while SCC separates them — is credible and useful. The proprietary–open-source gap is large and consistent across both human and automatic evaluation, and the prompt-rewriting ablation (SCC improves 23–52%) is an honest acknowledgment that part of the gap is prompt under-specification. They don't overclaim; they note the gap narrows but doesn't close.\n\nSoft spots, in order of real weight. First, no dataset or code link appears in the text, so 'released' is unverifiable from the paper. Second, no confidence intervals or significance tests on model rankings; with 150 testmini examples, some mid-table differences may be noise. Third, the external-validity point: rubrics are validated only by internal agreement, not against an independent curriculum or physical standard. That's a legitimate caveat, but I don't think it's fatal. The knee-jerk reflex example shows how concrete and grounded the anchors are; these are not vibes-based rubrics. The concern matters most for underspecified prompts where the expected storyline is a choice, and the prompt-rewriting result shows that's exactly where some measured gap lives. So the paper should either add an external validation or soften the 'mechanism, not appearance' framing to 'mechanism as specified by our rubrics.'\n\nWho this is for: anyone building or evaluating text-to-video models, and anyone working on MLLM-as-judge reliability. It deserves a serious referee; the issues are addressable. I'd engage.","headline":"Solid, well-built benchmark contribution with a credible main finding, but the headline 'mechanism, not appearance' claim rests on rubrics validated only internally and no released artifacts yet.","tokens_in":23598,"tokens_out":3548,"would_cite":true,"duration_ms":32207,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper claims that current text-to-video systems separate on mechanism rather than appearance: scientific and causal correctness scores vary widely across models while perceptual-quality proxies stay nearly flat.","keywords":["text-to-video generation","scientific reasoning","causal correctness","video evaluation benchmark","expert rubric","multidisciplinary","mechanistic fidelity","MLLM-as-judge"],"falsifier":"Have independent subject-matter experts who never saw the released rubrics write their own rubrics for the same prompts and re-score a fixed set of generated videos; substantial divergence in model rankings or per-dimension scores between the two rubric sets would show that the benchmark tracks what rubric authors valued rather than the underlying science.","tokens_in":22521,"feed_emoji":"🧪","tokens_out":8781,"duration_ms":74398,"temperature":0.7,"pith_summary":"Sci-VBench is a benchmark for testing whether text-to-video models can render expert-level scientific mechanisms rather than merely plausible-looking scenes. It contains 1,253 expert-authored prompts spanning 60 subjects in natural science, healthcare, humanities and social sciences, and engineering, each paired with a reference guide and a 1–5 rubric that defines what scientific correctness means for that example. The paper's central claim is that, under this rubric-based protocol, models separate sharply on Prompt Grounding and Scientific and Causal Correctness while automatic perceptual-quality scores stay almost flat, so today's systems differ in mechanism, not appearance. If the evaluation is valid, video generation research should be judged on whether outputs preserve the underlying causal and scientific dynamics, and the finding that even leading models produce systematic mechanistic errors means visual realism has outrun scientific validity.","feed_headline":"What separates AI video models is mechanism, not looks","feed_subtitle":"A new benchmark with expert rubrics finds models still fail on scientific and causal correctness.","key_machinery":"The load-bearing mechanism is the per-example evaluation specification: for each prompt, a domain expert authors a high-level reference guide identifying the target concept, the minimal mechanism, and the expected phase-based storyline, plus a 1–5 anchored rubric for Prompt Grounding, Scientific and Causal Correctness, and Spatiotemporal Consistency. This externalizes the expert knowledge a video judge needs, so non-experts and MLLM judges can apply the same standard without recruiting domain experts; the anchors tie each score to observable evidence, such as whether the struck leg extends at the knee immediately or whether the hammer strikes the tendon region. The automated protocol combines established video-quality vision tools for low-level perceptual fidelity with a rubric-conditioned MLLM-as-judge (a multimodal large language model that watches the video and scores it against the rubric), run three independent times per video and dimension and averaged, which is what allows the paper to separate appearance from mechanism at scale.","core_discovery":"The paper claims that advances in visual realism have not translated into reliable modeling of scientific and causal dynamics, and supports this with a benchmark whose prompts are deliberately minimal: each states only the setup and explicit intervention, omitting the expected outcome, so a successful video must infer and render the mechanism. On the 150-prompt testmini split, expert human ratings and automatic scores both show a wide spread on Scientific and Causal Correctness (roughly 1.1 to 3.4 on a 1–5 scale), while perceptual-quality proxies stay in a narrow band; the proprietary–open-source gap is concentrated on the reasoning dimensions, and no model is strongest across all disciplines. The paper further claims that its rubric-based protocol makes expert evaluation portable: non-experts who receive the reference guide and rubric agree with experts far more closely than non-experts who see only the prompt, and a rubric-conditioned MLLM judge aligns with expert ratings better than prior automatic video scoring methods on the reasoning-centric dimensions.","pith_inferences":["If the rubric-based protocol genuinely tracks scientific correctness, model rankings on Sci-VBench should predict success on downstream tasks that require applying the mechanism, such as generating step-by-step procedure demonstrations or verifying whether a procedure was followed; that prediction is testable but not made in the paper.","Because prompt rewriting improves scientific correctness without fixing temporal consistency, post-training objectives that reward mechanism-level correctness may deliver larger gains than better prompting alone.","The per-discipline performance profiles suggest that domain-specific fine-tuning or retrieval of scientific priors could close part of the open-source gap on the mechanisms those models currently get wrong.","Training a model to optimize the rubric-conditioned judge score directly would reveal whether the benchmark measures scientific ground truth or only surrogate judgments; if such training improved real mechanism fidelity, the benchmark's portability claim would be confirmed."],"forward_implications":["Perceptual-quality proxies alone overstate parity among models, since near-flat automatic quality scores hide large differences in whether generated dynamics obey the underlying mechanism.","Video-generation evaluation should include a distinct scientific-and-causal-correctness dimension with reusable expert rubrics rather than relying on prompt alignment and visual fidelity alone.","Prompt rewriting recovers only part of the mechanistic gap: improvements on scientific correctness and prompt grounding are much larger than improvements on spatiotemporal consistency, indicating generator limitations rather than underspecified instructions.","The proprietary–open-source gap in current systems sits on reasoning-centric dimensions, whereas open-source models can match or beat proprietary ones on spatiotemporal consistency.","No model is uniformly strong across disciplines, so aggregate rankings hide which domain mechanisms a system preserves or violates."],"supporting_citations":[{"why":"Supplies the VBench Video Quality metrics used as the automatic proxy for Low-level Perceptual Fidelity.","marker":"Huang et al., 2024"},{"why":"The closest concurrent benchmark (VideoScience-Bench) that Sci-VBench extends to 60 subjects and portable rubrics.","marker":"Hu et al., 2025b"},{"why":"VideoScore, a prior automatic evaluator compared against the rubric-conditioned judge on agreement with experts.","marker":"He et al., 2024"},{"why":"VideoScore2, a think-before-scoring evaluator used as a comparison in the reliability analysis.","marker":"He et al., 2025"},{"why":"VideoReward, a VLM-based reward model whose dimension scores are compared with expert ratings.","marker":"Liu et al., 2025"},{"why":"ETVA, a question-answering alignment scorer used as a comparison baseline.","marker":"Guan et al., 2025"},{"why":"VideoPhy, a representative physical-commonsense benchmark whose generic scope motivates the need for science-domain evaluation.","marker":"Bansal et al., 2025a"}],"fun_headline_variants":["AI video models: looks good, but can't explain why","New benchmark: AI videos fail at scientific reasoning","Visual realism is easy; causal thinking is the real test","Video AI: pretty pictures, poor physics","Sci-VBench: the gap between video realism and scientific truth"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the expert-authored rubrics give a complete and unbiased specification of scientifically correct behavior for each prompt, so that agreement with expert ratings measures true scientific correctness rather than adherence to the rubric authors' interpretive choices.","fun_headline_variants_meta":{"raw":{"variants":["AI video models: looks good, but can't explain why","New benchmark: AI videos fail at scientific reasoning","Visual realism is easy; causal thinking is the real test","Video AI: pretty pictures, poor physics","Sci-VBench: the gap between video realism and scientific truth"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000179,"raw_usage":{"total_tokens":1284,"prompt_tokens":912,"completion_tokens":372,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":528,"completion_tokens_details":{"reasoning_tokens":308}},"tokens_in":528,"tokens_out":372,"duration_ms":3991,"temperature":1.0,"reasoning_tokens":308,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T05:07:46.705000+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Have independent subject-matter experts who never saw the released rubrics write their own rubrics for the same prompts and re-score a fixed set of generated videos; substantial divergence in model rankings or per-dimension scores between the two rubric sets would show that the benchmark tracks what rubric authors valued rather than the underlying science.","supporting_citations":[],"review_version":1}