Pith. sign in

REVIEW 3 major objections 5 minor 48 references

Sci-VBench: Evaluating Knowledge- and Reasoning-Intensive Video Generation in Science Domains

T0 review · 3 major / 5 minor · reviewed 2026-08-11 · deepseek-v4-flash

Pith's one-line read This paper claims that current text-to-video systems separate on mechanism rather than appearance: scientific and causal correctness scores vary widely across models while perceptual-quality proxies stay nearly flat.

desk verdict Solid, well-built benchmark contribution with a credible main finding, but the headline 'mechanism, not appearance' claim rests on rubrics validated only internally and no released artifacts yet. read the letter →

arxiv 2608.09873 v1 pith:P64UKS6A submitted 2026-08-10 cs.CV cs.AI

classification cs.CVcs.AI
keywords text-to-videogenerationscientificreasoningcausalcorrectnessvideoevaluationbenchmarkexpertrubricmultidisciplinarymechanisticfidelityMLLM-as-judge
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Sci-VBench is a benchmark for testing whether text-to-video models can render expert-level scientific mechanisms rather than merely plausible-looking scenes. It contains 1,253 expert-authored prompts spanning 60 subjects in natural science, healthcare, humanities and social sciences, and engineering, each paired with a reference guide and a 1–5 rubric that defines what scientific correctness means for that example. The paper's central claim is that, under this rubric-based protocol, models separate sharply on Prompt Grounding and Scientific and Causal Correctness while automatic perceptual-quality scores stay almost flat, so today's systems differ in mechanism, not appearance. If the evaluation is valid, video generation research should be judged on whether outputs preserve the underlying causal and scientific dynamics, and the finding that even leading models produce systematic mechanistic errors means visual realism has outrun scientific validity.

What carries the argument

The load-bearing mechanism is the per-example evaluation specification: for each prompt, a domain expert authors a high-level reference guide identifying the target concept, the minimal mechanism, and the expected phase-based storyline, plus a 1–5 anchored rubric for Prompt Grounding, Scientific and Causal Correctness, and Spatiotemporal Consistency. This externalizes the expert knowledge a video judge needs, so non-experts and MLLM judges can apply the same standard without recruiting domain experts; the anchors tie each score to observable evidence, such as whether the struck leg extends at the knee immediately or whether the hammer strikes the tendon region. The automated protocol combines established video-quality vision tools for low-level perceptual fidelity with a rubric-conditioned MLLM-as-judge (a multimodal large language model that watches the video and scores it against the rubric), run three independent times per video and dimension and averaged, which is what allows the paper to separate appearance from mechanism at scale.

What would settle it

Have independent subject-matter experts who never saw the released rubrics write their own rubrics for the same prompts and re-score a fixed set of generated videos; substantial divergence in model rankings or per-dimension scores between the two rubric sets would show that the benchmark tracks what rubric authors valued rather than the underlying science.

Watch

Extended reading notes

Core claim

The paper claims that advances in visual realism have not translated into reliable modeling of scientific and causal dynamics, and supports this with a benchmark whose prompts are deliberately minimal: each states only the setup and explicit intervention, omitting the expected outcome, so a successful video must infer and render the mechanism. On the 150-prompt testmini split, expert human ratings and automatic scores both show a wide spread on Scientific and Causal Correctness (roughly 1.1 to 3.4 on a 1–5 scale), while perceptual-quality proxies stay in a narrow band; the proprietary–open-source gap is concentrated on the reasoning dimensions, and no model is strongest across all disciplines. The paper further claims that its rubric-based protocol makes expert evaluation portable: non-experts who receive the reference guide and rubric agree with experts far more closely than non-experts who see only the prompt, and a rubric-conditioned MLLM judge aligns with expert ratings better than prior automatic video scoring methods on the reasoning-centric dimensions.

Load-bearing premise

The load-bearing premise is that the expert-authored rubrics give a complete and unbiased specification of scientifically correct behavior for each prompt, so that agreement with expert ratings measures true scientific correctness rather than adherence to the rubric authors' interpretive choices.

Editorial extensions

If this is right

  • Perceptual-quality proxies alone overstate parity among models, since near-flat automatic quality scores hide large differences in whether generated dynamics obey the underlying mechanism.
  • Video-generation evaluation should include a distinct scientific-and-causal-correctness dimension with reusable expert rubrics rather than relying on prompt alignment and visual fidelity alone.
  • Prompt rewriting recovers only part of the mechanistic gap: improvements on scientific correctness and prompt grounding are much larger than improvements on spatiotemporal consistency, indicating generator limitations rather than underspecified instructions.
  • The proprietary–open-source gap in current systems sits on reasoning-centric dimensions, whereas open-source models can match or beat proprietary ones on spatiotemporal consistency.
  • No model is uniformly strong across disciplines, so aggregate rankings hide which domain mechanisms a system preserves or violates.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If the rubric-based protocol genuinely tracks scientific correctness, model rankings on Sci-VBench should predict success on downstream tasks that require applying the mechanism, such as generating step-by-step procedure demonstrations or verifying whether a procedure was followed; that prediction is testable but not made in the paper.
  • Because prompt rewriting improves scientific correctness without fixing temporal consistency, post-training objectives that reward mechanism-level correctness may deliver larger gains than better prompting alone.
  • The per-discipline performance profiles suggest that domain-specific fine-tuning or retrieval of scientific priors could close part of the open-source gap on the mechanisms those models currently get wrong.
  • Training a model to optimize the rubric-conditioned judge score directly would reveal whether the benchmark measures scientific ground truth or only surrogate judgments; if such training improved real mechanism fidelity, the benchmark's portability claim would be confirmed.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. Sci-VBench is a new benchmark for evaluating knowledge- and reasoning-intensive text-to-video generation in scientific domains. It contains 1,253 expert-authored prompts spanning 60 subjects across Natural Science, Healthcare, Humanities & Social Sciences, and Engineering, each paired with a per-example reference guide and 1-5 scoring rubrics. The paper validates a rubric-based evaluation protocol: expert ratings serve as reference labels, non-expert raters improve agreement when given the evaluation specification, and rubric-conditioned MLLM judges correlate with expert ratings more closely than prior automatic evaluators. Sixteen proprietary and open-source models are benchmarked. The headline findings are that automatic perceptual-quality scores are nearly flat across models, while Prompt Grounding and Scientific and Causal Correctness vary substantially with a proprietary-open-source gap, leading the authors to conclude that what separates current systems is mechanism, not appearance. A prompt-rewriting ablation shows that more explicit prompts improve SCC substantially but do not close the gap.

Significance. If the benchmark's SCC labels are accepted as ground truth, this is an important and useful contribution: it is the first benchmark of its scope to release reusable per-example evaluation specifications, it includes a controlled human study showing that non-experts can be brought close to expert agreement, and it reports strong expert re-rating stability (Cohen's kappa = 0.842). The observed pattern, where perceptual quality is saturated but scientific/causal correctness still separates models, is a falsifiable and practically relevant statement about current video generators. The main reservation is that the reference labels are validated only through internal consistency audits; the central 'mechanism, not appearance' claim is therefore partly a claim about expert-rubric adherence rather than about an external independent standard of scientific truth.

major comments (3)
  1. [§3.3-3.5, §4.3] The Scientific and Causal Correctness labels are validated only by internal consistency: the 200-example audit in §3.5 shows quadratic weighted kappa = 0.75 for SCC, and the 300-video re-rating in §4.3 shows Cohen's kappa = 0.842. These numbers demonstrate that the rubrics are applied reliably and are not idiosyncratic to one annotator, but they do not demonstrate that the rubrics correspond to an external scientific standard. Because §3.3 says target concepts are selected from 'canonical textbooks and course materials' but no textbook list or answer key is provided, and because the reference guide fixes the expected phase-based storyline, expert scores, non-expert-with-spec scores, and rubric-conditioned MLLM scores all measure agreement with the authors' operationalization. Given the abstract's claim about 'reliable modeling of scientific and causal dynamics,' the paper should either add an external audit (e.g., an independent expert panel verifying a sample of reference guides against named curriculum sources) or explicitly reframe the contribution as measuring rubric-verified scientific correctness rather than unqualified scientific correctness. Without this, the proprietary-open-source SCC gap in Table 4 could overstate differences in genuine scientific fidelity.
  2. [§5.4, Figure 4] The prompt-rewriting experiment quantifies how much SCC depends on prompt explicitness: +23.3% for Wan2.2-5B and +51.7% for HunyuanVideo-1.5. This is a large fraction of the measured shortfall and shows that much of the SCC gap is attributable to the benchmark's deliberate design choice of omitting outcomes (§3.3) rather than purely to the generators' inability to model mechanisms. Because the rewritten condition was applied only to two open-source models, the proprietary-open-source SCC gap in Table 4 is potentially confounded by differential sensitivity to prompt under-specification. I ask for either (a) rewritten-prompt runs on at least the top proprietary models, or (b) a conditional analysis that reports the SCC gap on a subset of examples where prompt explicitness is controlled. The sentence 'it cannot supply the mechanistic fidelity the generator lacks' is not yet fully supported unless the magnitude of this confound is bounded.
  3. [Table 3, §4.3] Instance-level Pearson correlations for the best rubric-conditioned MLLM judge are r = 0.545 for Spatiotemporal Consistency and r = 0.535 for Low-level Perceptual Fidelity. These are moderate correlations, not 'relatively high agreement.' Since the automatic scoring in Table 4 uses these judge scores for the SC dimension, the current automatic protocol is not yet reliable for fine-grained model ranking on Spatiotemporal Consistency. The abstract's claim that MLLM-as-Judge systems 'can achieve relatively high agreement with expert judgments, supporting reproducible evaluation at scale' is too broad if it is taken to cover these dimensions. Please report confidence intervals and the number of instances underlying Table 3, specify a threshold for 'high agreement,' and either soften the claim or restrict the scalable-evaluation claim to Prompt Grounding and Scientific and Causal Correctness, where the correlations are substantially higher.
minor comments (5)
  1. [Table 3] Table 3 reports Pearson correlations multiplied by 100; please also state the raw r values and the number of videos in the caption or text for clarity.
  2. [§3.5 and §4.3] The 200-example audit uses quadratic weighted kappa while the 300-video re-rating is reported as Cohen's kappa; please clarify whether the latter is also weighted and, if so, which weighting, since this affects comparability.
  3. [Table 4] For Gemini-Omni-Flash and Seedance-2.0, provider filters rejected some prompts and the averages use successful generations; please report the number of successful videos per model so readers can assess the impact of the missing examples.
  4. [§3.2 and Appendix A.4] Because five of the 61 annotators are authors, a sentence in the main text stating whether authors authored examples, validated examples, or both would improve transparency beyond the appendix biographies.
  5. [Appendix B.3 and Figure 6] The knee-jerk reflex rubric example in Figure 6 is very helpful; consider moving one complete rubric example into the main text so readers can see the protocol without consulting the appendix.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: Sci-VBench's findings are empirical measurements against an explicitly rubric-defined, expert-anchored standard, and the main claims are corroborated by independent human expert ratings.

full rationale

Sci-VBench's central claims are empirical measurements under an operational definition rather than derivations from their own inputs. Section 3.1 defines Scientific and Causal Correctness as consistency with the per-example reference guide, and the guide and rubric are released as the evaluation standard; this is benchmark construction, not a result reduced to its inputs. The main finding that systems separate on mechanism rather than appearance is supported by both the rubric-conditioned MLLM judge and by independent expert human ratings (Table 4), with expert ratings serving as the reference labels (Section 4.3). The improved agreement of non-experts and MLLM judges when given the evaluation specification (Table 3) is the paper's intended demonstration that the specification is portable; the specification is the independent variable, not a fitted parameter renamed as a prediction. No fitted parameters are used to produce model rankings, and no load-bearing uniqueness theorem or ansatz is imported from the authors' prior work; self-citations appear only in related-work and motivation contexts. The remaining limitation--that rubric correctness is validated by internal expert consistency (kappa = 0.75 and 0.842) rather than an external physical answer key--is a construct-validity question, not circularity: the paper explicitly operationalizes scientific correctness through expert-authored rubrics and does not claim those rubrics are derived from first principles. Under the rule that circularity must be exhibited by a quoted reduction, no circular step can be identified.

Assumptions & free parameters 0 free parameters · 3 assumptions · 0 invented entities

Sci-VBench introduces no free parameters, fitted constants, or invented physical entities. The central claim rests on domain assumptions about visual verifiability of scientific phenomena, the validity of expert-authored rubrics as ground truth, and the reliability of MLLM judges when given rubric anchors.

assumptions (3)
  • domain assumption Visual evidence alone is sufficient to verify the scientific correctness of the selected concepts.
    Section 3.3 explicitly excludes concepts whose correctness cannot be verified from video evidence alone; this makes rubric-based scoring possible but restricts the benchmark to visually observable mechanisms and assumes that the selected phenomena are fully captured on video.
  • domain assumption Expert labels from the benchmark's own annotator pool are an unbiased ground truth.
    Section 4.3 uses expert ratings as reference labels and reports inter-expert agreement of kappa = 0.842, but there is no external validation of the rubric content against an independent standard, so the ground truth is self-referential to the annotator pool.
  • domain assumption MLLM judges can faithfully map rubric anchors to video evidence when conditioned on a single dimension at a time.
    Section 4.2 relies on rubric-conditioned MLLM-as-judge scoring; the paper measures correlation with expert labels but does not independently verify that the MLLM's stated justifications are correct, so the assumption is that rubric conditioning is sufficient for reliable scoring.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Sci-VBench: Evaluating Knowledge- and Reasoning-Intensive Video Generation in Science Domains." pith.science (2026). https://pith.science/paper/P64UKS6A

@misc{pith2026260809873,
  author       = {Pith},
  title        = {Pith review of: Sci-VBench: Evaluating Knowledge- and Reasoning-Intensive Video Generation in Science Domains},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/P64UKS6A}},
  note         = {Machine review of arXiv:2608.09873}
}
read the original abstract

We introduce Sci-VBench, a comprehensive benchmark for evaluating knowledge- and reasoning-intensive video generation across scientific domains. It contains 1,253 expert-annotated examples spanning 60 subjects across four core disciplines: Natural Science, Healthcare, Humanities & Social Sciences, and Engineering. Each example requires models to generate temporally rich videos that demand scientific reasoning and knowledge-grounded synthesis, going beyond surface-level visual plausibility. We further establish a rubric-based evaluation protocol. Our analysis shows that, under this protocol, both non-expert human evaluators and MLLM-as-Judge systems can achieve relatively high agreement with expert judgments, supporting reproducible evaluation at scale. We benchmark 16 frontier proprietary and open-source models and find that, while automatic perceptual-quality scores cluster tightly across systems, performance on Prompt Grounding and Scientific and Causal Correctness varies substantially, with a pronounced proprietary-open-source gap. These findings show that advances in visual realism have not yet translated into reliable modeling of scientific and causal dynamics.

Figures

Figures reproduced from arXiv: 2608.09873 by the authors.

Figure 1
Figure 1. Overview of the Sci-VBench dataset. Top left: distribution of the 1,253 expert￾annotated examples over 60 subjects across four core disciplines. Remaining panels: representative prompt–video examples from each discipline, generated by Gemini-Omni￾Flash, where faithful generation requires grounding in the underlying scientific mechanism. ∗Equal contributions. Correspondence to: Tingyu Song (songtingyu23@mails.ucas.ac… view at source ↗
Figure 2
Figure 2. Overview of the Sci-VBench benchmark construction process. [PITH_FULL_IMAGE:figures/full_fig_p004_2.png] view at source ↗
Figure 3
Figure 3. Per-discipline mean of the SC/PG/SCC judge scores on testmini [PITH_FULL_IMAGE:figures/full_fig_p008_3.png] view at source ↗
Figures from the paper (8 more)
Figure 4
Figure 4. Figure 4: Performance enhancement via prompt rewriting on testmini. Our main results use verbatim prompts (Sec￾tion 5.1); as an ablation, we ask how much of the gap more explicit prompting recovers. We rewrite each testmini prompt with Gemini-3-Flash, instructing it to restate t…
Figure 5
Figure 5. Figure 5: Prompt template for rubric-conditioned MLLM-as-judge scoring on a single [PITH_FULL_IMAGE:figures/full_fig_p021_5.png]
Figure 6
Figure 6. Figure 6: The evaluation specification of the knee-jerk reflex example, reproduced verbatim [PITH_FULL_IMAGE:figures/full_fig_p022_6.png]
Figure 7
Figure 7. Figure 7: Category (1), poor adherence to instructions: the video is inconsistent with the [PITH_FULL_IMAGE:figures/full_fig_p024_7.png]
Figure 8
Figure 8. Figure 8: Category (2), inaccurate simulation of scientific principles: the required mechanism [PITH_FULL_IMAGE:figures/full_fig_p024_8.png]
Figure 9
Figure 9. Figure 9: Category (3), deficiencies in temporal coherence: objects lose permanence over [PITH_FULL_IMAGE:figures/full_fig_p024_9.png]
Figure 10
Figure 10. Figure 10: Category (3), deficiencies in visual quality: the visual style shifts discontinuously. [PITH_FULL_IMAGE:figures/full_fig_p024_10.png]
Figure 11
Figure 11. Figure 11: Further category (3) defects: coarse textures and weak causal links between [PITH_FULL_IMAGE:figures/full_fig_p025_11.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

48 extracted references · 23 canonical work pages

  1. [1]

    CoRR , volume =

    Jingtong Yue and Ziqi Huang and Zhaoxi Chen and Xintao Wang and Pengfei Wan and Ziwei Liu , title =. CoRR , volume =. 2025 , url =. doi:10.48550/ARXIV.2511.08585 , eprinttype =. 2511.08585 , timestamp =

  2. [2]

    CoRR , volume =

    Yue Ma and Kunyu Feng and Zhongyuan Hu and Xinyu Wang and Yucheng Wang and Mingzhe Zheng and Bingyuan Wang and Qinghe Wang and Xuanhua He and Hongfa Wang and Chenyang Zhu and Hongyu Liu and Yingqing He and Zeyu Wang and Zhifeng Li and Xiu Li and Sirui Han and Yike Guo and Wei Liu and Dan Xu and Linfeng Zhang and Qifeng Chen , title =. CoRR , volume =. 202...

  3. [3]

    2024 , url =

    Ziqi Huang and Yinan He and Jiashuo Yu and Fan Zhang and Chenyang Si and Yuming Jiang and Yuanhan Zhang and Tianxing Wu and Qingyang Jin and Nattapol Chanpaisit and Yaohui Wang and Xinyuan Chen and Limin Wang and Dahua Lin and Yu Qiao and Ziwei Liu , title =. 2024 , url =. doi:10.1109/CVPR52733.2024.02060 , timestamp =

  4. [4]

    VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness , journal =

    Dian Zheng and Ziqi Huang and Hongbo Liu and Kai Zou and Yinan He and Fan Zhang and Lulu Gu and Yuanhan Zhang and Jingwen He and Wei. VBench-2.0: Advancing Video Generation Benchmark Suite for Intrinsic Faithfulness , journal =. 2025 , url =. doi:10.48550/ARXIV.2503.21755 , eprinttype =. 2503.21755 , timestamp =

  5. [6]

    CoRR , volume =

    Baiqi Li and Zhiqiu Lin and Deepak Pathak and Jiayao Li and Yixin Fei and Kewen Wu and Tiffany Ling and Xide Xia and Pengchuan Zhang and Graham Neubig and Deva Ramanan , title =. CoRR , volume =. 2024 , url =. doi:10.48550/ARXIV.2406.13743 , eprinttype =. 2406.13743 , timestamp =

  6. [7]

    VideoPhy: Evaluating Physical Commonsense for Video Generation , booktitle =

    Hritik Bansal and Zongyu Lin and Tianyi Xie and Zeshun Zong and Michal Yarom and Yonatan Bitton and Chenfanfu Jiang and Yizhou Sun and Kai. VideoPhy: Evaluating Physical Commonsense for Video Generation , booktitle =. 2025 , url =

  7. [8]

    Gonzalez and Ion Stoica and Song Han and Yao Lu , title =

    Dacheng Li and Yunhao Fang and Yukang Chen and Shuo Yang and Shiyi Cao and Justin Wong and Michael Luo and Xiaolong Wang and Hongxu Yin and Joseph E. Gonzalez and Ion Stoica and Song Han and Yao Lu , title =. CoRR , volume =. 2025 , url =. doi:10.48550/ARXIV.2502.20694 , eprinttype =. 2502.20694 , timestamp =

  8. [9]

    VideoPhy-2:

    Hritik Bansal and Clark Peng and Yonatan Bitton and Roman Goldenberg and Aditya Grover and Kai. VideoPhy-2:. CoRR , volume =. 2025 , url =. doi:10.48550/ARXIV.2503.06800 , eprinttype =. 2503.06800 , timestamp =

Show all 48 references
  1. [10]

    VBench++: Comprehensive and Versatile Benchmark Suite for Video Generative Models , journal =

    Ziqi Huang and Fan Zhang and Xiaojie Xu and Yinan He and Jiashuo Yu and Ziyue Dong and Qianli Ma and Nattapol Chanpaisit and Chenyang Si and Yuming Jiang and Yaohui Wang and Xinyuan Chen and Ying. VBench++: Comprehensive and Versatile Benchmark Suite for Video Generative Model...

  2. [11]

    CoRR , volume =

    Lanxiang Hu and Abhilash Shankarampeta and Yixin Huang and Zilin Dai and Haoyang Yu and Yujie Zhao and Haoqiang Kang and Daniel Zhao and Tajana Rosing and Hao Zhang , title =. CoRR , volume =. 2025 , url =. doi:10.48550/ARXIV.2512.02942 , eprinttype =. 2512.02942 , timestamp =

  3. [12]

    Towards World Simulator: Crafting Physical Commonsense-Based Benchmark for Video Generation , booktitle =

    Fanqing Meng and Jiaqi Liao and Xinyu Tan and Quanfeng Lu and Wenqi Shao and Kaipeng Zhang and Yu Cheng and Dianqi Li and Ping Luo , editor =. Towards World Simulator: Crafting Physical Commonsense-Based Benchmark for Video Generation , booktitle =. 2025 , url =

  4. [13]

    2025 , month = sep, day =

  5. [14]

    2024 , month = feb, day =

    Tim Brooks and Bill Peebles and Connor Holmes and Will DePue and Yufei Guo and Li Jing and David Schnurr and Joe Taylor and Troy Luhman and Eric Luhman and Clarence Ng and Ricky Wang and Aditya Ramesh , title =. 2024 , month = feb, day =

  6. [15]

    2025 , howpublished =

  7. [16]

    The Thirteenth International Conference on Learning Representations,

    Zhuoyi Yang and Jiayan Teng and Wendi Zheng and Ming Ding and Shiyu Huang and Jiazheng Xu and Yuanming Yang and Wenyi Hong and Xiaohan Zhang and Guanyu Feng and Da Yin and Yuxuan Zhang and Weihan Wang and Yean Cheng and Bin Xu and Xiaotao Gu and Yuxiao Dong and Jie Tang , titl...

  8. [17]

    2025 , url =

    Wan: Open and Advanced Large-Scale Video Generative Models , journal =. 2025 , url =. doi:10.48550/ARXIV.2503.20314 , eprinttype =. 2503.20314 , timestamp =

  9. [18]

    2025 , url =

    HunyuanVideo 1.5 Technical Report , journal =. 2025 , url =. doi:10.48550/ARXIV.2511.18870 , eprinttype =. 2511.18870 , timestamp =

  10. [19]

    CoRR , volume =

    Yoav HaCohen and Benny Brazowski and Nisan Chiprut and Yaki Bitterman and Andrew Kvochko and Avishai Berkowitz and Daniel Shalem and Daphna Lifschitz and Dudu Moshe and Eitan Porat and Eitan Richardson and Guy Shiran and Itay Chachy and Jonathan Chetboun and Michael Finkelson ...

  11. [20]

    2025 , url =

    LongCat-Video Technical Report , journal =. 2025 , url =. doi:10.48550/ARXIV.2510.22200 , eprinttype =. 2510.22200 , timestamp =

  12. [21]

    2025 , month = dec, day =

    Kling. 2025 , month = dec, day =

  13. [22]

    Introducing Wan 2.6 , year =

  14. [23]

    The Thirteenth International Conference on Learning Representations,

    Jiachen Li and Qian Long and Jian Zheng and Xiaofeng Gao and Robinson Piramuthu and Wenhu Chen and William Yang Wang , title =. The Thirteenth International Conference on Learning Representations,. 2025 , url =

  15. [24]

    CoRR , volume =

    Roberto Henschel and Levon Khachatryan and Daniil Hayrapetyan and Hayk Poghosyan and Vahram Tadevosyan and Zhangyang Wang and Shant Navasardyan and Humphrey Shi , title =. CoRR , volume =. 2024 , url =. doi:10.48550/ARXIV.2403.14773 , eprinttype =. 2403.14773 , timestamp =

  16. [25]

    CoRR , volume =

    Teng Hu and Zhentao Yu and Zhengguang Zhou and Sen Liang and Yuan Zhou and Qin Lin and Qinglin Lu , title =. CoRR , volume =. 2025 , url =. doi:10.48550/ARXIV.2505.04512 , eprinttype =. 2505.04512 , timestamp =

  17. [26]

    Force Prompting: Video Generation Models Can Learn And Generalize Physics-based Control Signals , booktitle =

    Nate Gillman and Charles Herrmann and Michael Freeman and Daksh Aggarwal and Evan Luo and Deqing Sun and Chen Sun , editor =. Force Prompting: Video Generation Models Can Learn And Generalize Physics-based Control Signals , booktitle =. 2025 , url =. doi:10.52202/085713-3448 ,...

  18. [28]

    VideoScore: Building Automatic Metrics to Simulate Fine-grained Human Feedback for Video Generation , booktitle =

    Xuan He and Dongfu Jiang and Ge Zhang and Max Ku and Achint Soni and Sherman Siu and Haonan Chen and Abhranil Chandra and Ziyan Jiang and Aaran Arulraj and Kai Wang and Quy Duc Do and Yuansheng Ni and Bohan Lyu and Yaswanth Narsupalli and Rongqi Fan and Zhiheng Lyu and Bill Yu...

  19. [29]

    CoRR , volume =

    Xuan He and Dongfu Jiang and Ping Nie and Minghao Liu and Zhengxuan Jiang and Mingyi Su and Wentao Ma and Junru Lin and Chun Ye and Yi Lu and Keming Wu and Benjamin Schneider and Quy Duc Do and Zhuofeng Li and Yiming Jia and Yuxuan Zhang and Guo Cheng and Haozhe Wang and Wangc...

  20. [30]

    Improving Video Generation with Human Feedback , booktitle =

    Jie Liu and Gongye Liu and Jiajun Liang and Ziyang Yuan and Xiaokun Liu and Mingwu Zheng and Xiele Wu and Qiulin Wang and Menghan Xia and Xintao Wang and Xiaohong Liu and Fei Yang and Pengfei Wan and Di Zhang and Kun Gai and Yujiu Yang and Wanli Ouyang , editor =. Improving Vi...

  21. [31]

    2025 , url =

    Kaisi Guan and Zhengfeng Lai and Yuchong Sun and Peng Zhang and Wei Liu and Kieran Liu and Meng Cao and Ruihua Song , title =. 2025 , url =. doi:10.1109/ICCV51701.2025.01978 , timestamp =

  22. [32]

    CoRR , volume =

    NVIDIA , title =. CoRR , volume =. 2025 , url =. doi:10.48550/ARXIV.2511.00062 , eprinttype =. 2511.00062 , timestamp =

  23. [33]

    2025 , url =

    Yilun Zhao and Haowei Zhang and Lujing Xie and Tongyan Hu and Guo Gan and Yitao Long and Zhiyuan Hu and Weiyuan Chen and Chuhan Li and Zhijian Xu and Chengye Wang and Ziyao Shangguan and Zhenwen Liang and Yixin Liu and Chen Zhao and Arman Cohan , title =. 2025 , url =. doi:10....

  24. [34]

    CoRR , volume =

    Lin Fu and Zheyuan Yang and Yang Wang and Tingyu Song and Arman Cohan and Yilun Zhao , title =. CoRR , volume =. 2026 , url =. doi:10.48550/ARXIV.2606.05259 , eprinttype =. 2606.05259 , timestamp =

  25. [35]

    VF-Eval: Evaluating Multimodal LLMs for Generating Feedback on

    Tingyu Song and Tongyan Hu and Guo Gan and Yilun Zhao , editor =. VF-Eval: Evaluating Multimodal LLMs for Generating Feedback on. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers),. 2025 , url =. doi:10.18653/V1/202...

  26. [36]

    CoRR , volume =

    Kairui Hu and Penghao Wu and Fanyi Pu and Wang Xiao and Yuanhan Zhang and Xiang Yue and Bo Li and Ziwei Liu , title =. CoRR , volume =. 2025 , url =. doi:10.48550/ARXIV.2501.13826 , eprinttype =. 2501.13826 , timestamp =

  27. [37]

    SciVideoBench: Benchmarking Scientific Video Reasoning in Large Multimodal Models , journal =

    Andong Deng and Taojiannan Yang and Shoubin Yu and Lincoln Spencer and Mohit Bansal and Chen Chen and Serena Yeung. SciVideoBench: Benchmarking Scientific Video Reasoning in Large Multimodal Models , journal =. 2025 , url =. doi:10.48550/ARXIV.2510.08559 , eprinttype =. 2510.0...

  28. [38]

    2026 , url =

    Seedance 2.0: Advancing Video Generation for World Complexity , journal =. 2026 , url =. doi:10.48550/ARXIV.2604.14148 , eprinttype =. 2604.14148 , timestamp =

  29. [39]

    2026 , url =

    Alibaba Cloud Model Studio: HappyHorse text-to-video API reference , author =. 2026 , url =

  30. [40]

    2026 , url =

    Generate and edit videos with Gemini Omni Flash , author =. 2026 , url =

  31. [41]

    2026 , url =

    Cosmos 3: Omnimodal World Models for Physical AI , author =. 2026 , url =

  32. [42]

    2026 , url =

    MiniMax H3: An Open Model Breaking the Boundaries Between Tasks and Modalities , author =. 2026 , url =

  33. [43]

    The Thirteenth International Conference on Learning Representations,

    Ziyao Shangguan and Chuhan Li and Yuxuan Ding and Yanan Zheng and Yilun Zhao and Tesca Fitzgerald and Arman Cohan , title =. The Thirteenth International Conference on Learning Representations,. 2025 , url =

  34. [44]

    AbGen: Evaluating Large Language Models in Ablation Study Design and Evaluation for Scientific Research , booktitle =

    Yilun Zhao and Weiyuan Chen and Zhijian Xu and Manasi Patwardhan and Chengye Wang and Yixin Liu and Lovekesh Vig and Arman Cohan , editor =. AbGen: Evaluating Large Language Models in Ablation Study Design and Evaluation for Scientific Research , booktitle =. 2025 , url =. doi...

  35. [45]

    SciArena: An Open Evaluation Platform for Non-Verifiable Scientific Literature-Grounded Tasks , booktitle =

    Yilun Zhao and Kaiyan Zhang and Tiansheng Hu and Sihong Wu and Ronan Le Bras and Yixin Liu and Robert Tang and Joseph Chee Chang and Jesse Dodge and Jonathan Bragg and Chen Zhao and Hanna Hajishirzi and Doug Downey and Arman Cohan , editor =. SciArena: An Open Evaluation Platf...

  36. [46]

    SciSketch: An Open-source Framework for Automated Schematic Diagram Generation in Scientific Papers , booktitle =

    Zihang Wang and Yilun Zhao and Kaiyan Zhang and Chen Zhao and Manasi Patwardhan and Arman Cohan , editor =. SciSketch: An Open-source Framework for Automated Schematic Diagram Generation in Scientific Papers , booktitle =. 2025 , url =. doi:10.18653/V1/2025.EMNLP-DEMOS.28 , ti...

  37. [47]

    SciVer: Evaluating Foundation Models for Multimodal Scientific Claim Verification , booktitle =

    Chengye Wang and Yifei Shen and Zexi Kuang and Arman Cohan and Yilun Zhao , editor =. SciVer: Evaluating Foundation Models for Multimodal Scientific Claim Verification , booktitle =. 2025 , url =. doi:10.18653/V1/2025.ACL-LONG.420 , timestamp =

  38. [48]

    Han and Manasi Patwardhan and Arman Cohan , editor =

    Ziyu Chen and Yilun Zhao and Chengye Wang and Rilyn R. Han and Manasi Patwardhan and Arman Cohan , editor =. SciMDR: Advancing Scientific Multimodal Document Reasoning , booktitle =. 2026 , url =. doi:10.18653/V1/2026.ACL-LONG.2070 , timestamp =

  39. [49]

    Can Multimodal Foundation Models Understand Schematic Diagrams? An Empirical Study on Information-Seeking

    Yilun Zhao and Chengye Wang and Chuhan Li and Arman Cohan , editor =. Can Multimodal Foundation Models Understand Schematic Diagrams? An Empirical Study on Information-Seeking. Findings of the Association for Computational Linguistics,. 2025 , url =. doi:10.18653/V1/2025.FINDI...

  40. [50]

    Attention is All you Need , url =

    Vaswani, Ashish and Shazeer, Noam and Parmar, Niki and Uszkoreit, Jakob and Jones, Llion and Gomez, Aidan N and Kaiser, ukasz and Polosukhin, Illia , booktitle =. Attention is All you Need , url =

Pith tools

Reviewed August 11, 2026 · model on record in the stance chip above.