{"paper":{"title":"Do Joint Audio-Video Generation Models Understand Physics?","license":"http://arxiv.org/licenses/nonexclusive-distrib/1.0/","headline":"Joint audio-video generation models generate plausible content but fail to maintain real-world physical consistency across transitions.","cross_cats":["cs.AI","cs.CV","cs.MM"],"primary_cat":"cs.SD","authors_text":"Chenming Ge, Feiyu Du, Hao Fang, Jiageng Liu, Mingwei Xu, Shijian Deng, Weiguo Pian, Xiulong Liu, Yapeng Tian, Zexin Xu, Zijun Cui","submitted_at":"2026-05-08T00:14:07Z","abstract_excerpt":"Joint audio-video generation models are rapidly approaching professional production quality, raising a central question: do they understand audio-visual physics, or merely generate plausible sounds and frames that violate real-world consistency? We introduce AV-Phys Bench, a benchmark for evaluating physical commonsense in joint audio-video generation. AV-Phys Bench tests models across three scene categories: Steady State, Event Transition, and Environment Transition. It covers physics-grounded subcategories drawn from real-world scenes, plus Anti-AV-Physics prompts that deliberately request p"},"claims":{"count":4,"items":[{"kind":"strongest_claim","text":"Across three proprietary and four open-source models, we find that Seedance 2.0 performs best overall, but all models remain far from robust physical understanding. Performance drops sharply on event-driven and environment-driven transitions, and even strong proprietary systems collapse on Anti-AV-Physics prompts.","source":"verdict.strongest_claim","status":"machine_extracted","claim_id":"C1","attestation":"unclaimed"},{"kind":"weakest_assumption","text":"That the five evaluation dimensions (visual/audio semantic adherence, visual/audio physical commonsense, cross-modal physical commonsense) together constitute a sufficient and unbiased measure of physical understanding without requiring additional human validation or external physics simulators.","source":"verdict.weakest_assumption","status":"machine_extracted","claim_id":"C2","attestation":"unclaimed"},{"kind":"one_line_summary","text":"Current joint audio-video generation models lack robust physical commonsense, especially during transitions and when prompted for impossible behaviors.","source":"verdict.one_line_summary","status":"machine_extracted","claim_id":"C3","attestation":"unclaimed"},{"kind":"headline","text":"Joint audio-video generation models generate plausible content but fail to maintain real-world physical consistency across transitions.","source":"verdict.pith_extraction.headline","status":"machine_extracted","claim_id":"C4","attestation":"unclaimed"}],"snapshot_sha256":"ebe658fec31ff2c781ebf50f368fecf2e6b4c2d60806ef4cc6f38a5f4d4e24af"},"source":{"id":"2605.07061","kind":"arxiv","version":2},"verdict":{"id":"d561db9a-627d-436a-b72e-fabf49a89dbc","model_set":{"reader":"grok-4.3"},"created_at":"2026-05-11T02:08:27.608016Z","strongest_claim":"Across three proprietary and four open-source models, we find that Seedance 2.0 performs best overall, but all models remain far from robust physical understanding. Performance drops sharply on event-driven and environment-driven transitions, and even strong proprietary systems collapse on Anti-AV-Physics prompts.","one_line_summary":"Current joint audio-video generation models lack robust physical commonsense, especially during transitions and when prompted for impossible behaviors.","pipeline_version":"pith-pipeline@v0.9.0","weakest_assumption":"That the five evaluation dimensions (visual/audio semantic adherence, visual/audio physical commonsense, cross-modal physical commonsense) together constitute a sufficient and unbiased measure of physical understanding without requiring additional human validation or external physics simulators.","pith_extraction_headline":"Joint audio-video generation models generate plausible content but fail to maintain real-world physical consistency across transitions."},"integrity":{"clean":true,"summary":{"advisory":0,"critical":0,"by_detector":{},"informational":0},"endpoint":"/pith/2605.07061/integrity.json","findings":[],"available":true,"detectors_run":[{"name":"claim_evidence","ran_at":"2026-05-20T11:42:03.394008Z","status":"completed","version":"1.0.0","findings_count":0},{"name":"ai_meta_artifact","ran_at":"2026-05-20T06:37:10.217339Z","status":"completed","version":"1.0.0","findings_count":0},{"name":"doi_title_agreement","ran_at":"2026-05-19T17:31:18.744647Z","status":"completed","version":"1.0.0","findings_count":0},{"name":"doi_compliance","ran_at":"2026-05-19T12:09:37.526702Z","status":"completed","version":"1.0.0","findings_count":0}],"snapshot_sha256":"26745d2109c6111bbf6a833bb8e0c1fb2184db7465cb7d2301cff55b0503fe36"},"references":{"count":0,"sample":[],"resolved_work":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57","internal_anchors":0},"formal_canon":{"evidence_count":2,"snapshot_sha256":"4ee6fccf7e2543a11a4bd827a3535e436fdc46489d7d94694c2a61f634f88688"},"author_claims":{"count":0,"strong_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"builder_version":"pith-number-builder-2026-05-17-v1"}