Pith. sign in

Paper Citation Record · LEDGER

Do Joint Audio-Video Generation Models Understand Physics?

As of 22 August 2026, this Paper Citation Record lists 53 of 53 outbound references and 0 inbound Pith citation observations for arXiv:2605.07061.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.07061 v2

Coverage vector

measured 53 of 53 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-06-30T23:39:22.070629Z

measured 53 of 53 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

53 of 53 outbound references displayed

  • verified exact27
  • verified fuzzy23
  • unresolved1
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 5a8a2620-d35a-4f03-b086-7b814fd71ea4 · outbound

This paper cites Cosmos World Foundation Model Platform for Physical AI.

Do Joint Audio-Video Generation Models Understand Physics? Cosmos World Foundation Model Platform for Physical AI

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-06-30T23:45:08.260783Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-30T23:39:22.070629Z digest=sha256:ce1084bb3b204be66b765a45d716068bd8d598faa7aa70dc95ec2e6b5f557799

Observation c0795d41-239c-494c-919d-5e8ec55ae934 · outbound

This paper cites VideoPhy: Evaluating Physical Commonsense for Video Generation.

Do Joint Audio-Video Generation Models Understand Physics? VideoPhy: Evaluating Physical Commonsense for Video Generation

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-06-30T23:45:08.231096Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-30T23:39:22.070629Z digest=sha256:ff531c521567c979da0d5b8190848067258c91ef981c78639c26d6fe1f5ef85d

Observation 68ca9a9f-28a4-45f6-97c7-000e535ba148 · outbound

This paper cites VideoPhy-2: A Challenging Action-Centric Physical Commonsense Evaluation in Video Generation.

Do Joint Audio-Video Generation Models Understand Physics? VideoPhy-2: A Challenging Action-Centric Physical Commonsense Evaluation in Video Generation

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-06-30T23:45:08.203083Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-30T23:39:22.070629Z digest=sha256:45f376f7ef980302fe903d603ae3d99ab2a08fd78dcb090199adf9ce01f89311

Observation 70f089c7-d8d9-45fc-871e-9cc4bfdee1b2 · outbound

This paper cites Video generation models as world simulators.

Do Joint Audio-Video Generation Models Understand Physics? Video generation models as world simulators

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T10:33:40.268019Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-30T23:39:22.070629Z digest=sha256:948c84adaa8201a12ed1e7977312f234d6bfa86b137ca0395a521a75d83db242

Observation 76e56aab-8d82-414a-b7e8-908829fa1ca6 · outbound

This paper cites Genie: Generative interactive environments.

Do Joint Audio-Video Generation Models Understand Physics? Genie: Generative interactive environments

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T10:33:40.269905Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-30T23:39:22.070629Z digest=sha256:9134e2d2991e5bd8d4154714e422e1c743b8937cd68a4532b795513a5dffa058

Observation a69173e7-7a2c-468a-9dcb-74d5437933cb · outbound

This paper cites T2AV-Compass: Towards Unified Evaluation for Text-to-Audio-Video Generation.

Do Joint Audio-Video Generation Models Understand Physics? T2AV-Compass: Towards Unified Evaluation for Text-to-Audio-Video Generation

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-06-30T23:45:08.236205Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-30T23:39:22.070629Z digest=sha256:63904d23837fe0514a75b1bfb9b2caa8a74cf63731b545598b369f0030754dc3

Observation b2617e39-eb0d-4709-b816-ddbac5719cac · outbound

This paper cites Soundspaces 2.0: A simulation platform for visual-acoustic learning.Advances in Neural Information Processing Systems, 35:8896– 8911.

Do Joint Audio-Video Generation Models Understand Physics? Soundspaces 2.0: A simulation platform for visual-acoustic learning.Advances in Neural Information Processing Systems, 35:8896– 8911

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T10:33:40.271754Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-30T23:39:22.070629Z digest=sha256:c4995914eb93b505cd4906bdc113bd448a067438cd729e1cef6717828c59931c

Observation 162511d4-657a-4d71-8e94-51887078ef76 · outbound

This paper cites SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing.

Do Joint Audio-Video Generation Models Understand Physics? SAVVY: Spatial Awareness via Audio-Visual LLMs through Seeing and Hearing

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-06-30T23:45:08.223574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-30T23:39:22.070629Z digest=sha256:efa7eb46e885a368a2aacb0206d0e6796cd4e6b8415c37044823bd995764be97

Observation 0eb73f90-8817-4f51-b983-38ff60e71dbb · outbound

This paper cites Speed by simplicity: A single-stream architecture for fast audio-video generative foundation model.

Do Joint Audio-Video Generation Models Understand Physics? Speed by simplicity: A single-stream architecture for fast audio-video generative foundation model

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-06-30T23:45:08.241509Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-30T23:39:22.070629Z digest=sha256:0ec67fabad83ce4a1ddc86ae3c8d78852617fec5ff9ae40f9ec84e228ff2e1f5

Observation 75847fe0-e924-4df2-ac3d-91fb842a887a · outbound

This paper cites Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities.

Do Joint Audio-Video Generation Models Understand Physics? Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-06-30T23:45:08.215209Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-30T23:39:22.070629Z digest=sha256:42720aa02941da9dd02a16dc51cab2ad3760d0fd77503caf47ee55e53d2c817d

Observation 385760c1-0697-4437-8535-80c3b003ef5c · outbound

This paper cites Introducing Veo 3.1 and advanced capa- bilities in Flow.

Do Joint Audio-Video Generation Models Understand Physics? Introducing Veo 3.1 and advanced capa- bilities in Flow

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T10:33:40.266378Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-30T23:39:22.070629Z digest=sha256:4a108ec9106603a4885e39a3b980a2d29e2e479492afc53fea513d2d73cf7846

Observation 19520228-7172-4c56-929d-2f27853d27b7 · outbound

This paper cites Technical details inherited from the Veo 3 Tech Re- port,https://storage.googleapis.com/deepmind-media/veo/Veo-3-Tech-Report.

Do Joint Audio-Video Generation Models Understand Physics? Technical details inherited from the Veo 3 Tech Re- port,https://storage.googleapis.com/deepmind-media/veo/Veo-3-Tech-Report

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T10:33:40.273514Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-30T23:39:22.070629Z digest=sha256:6d65a18245015d3c3abc76859c415d52098551f8b12536a91aba332e2605b4b2

Observation dddbdb33-a9e4-4c17-bce5-d2c8fef8bbde · outbound

This paper cites Look, listen, and act: Towards audio-visual embodied navigation.

Do Joint Audio-Video Generation Models Understand Physics? Look, listen, and act: Towards audio-visual embodied navigation

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T10:33:40.259561Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-30T23:39:22.070629Z digest=sha256:37ac7112d0c32e3a6b4552cff8a2242f54ca994f757406d90091076ba2aba7d0

Observation 1ca743af-1dfe-440b-bb86-c2ecf398f61f · outbound

This paper cites "PhyWorldBench": A Comprehensive Evaluation of Physical Realism in Text-to-Video Models.

Do Joint Audio-Video Generation Models Understand Physics? "PhyWorldBench": A Comprehensive Evaluation of Physical Realism in Text-to-Video Models

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-06-30T23:45:08.238608Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-30T23:39:22.070629Z digest=sha256:add9dac9b062beeb3ba06b0f89e36070a7d72fc697f870c699406aa782f55d40

Observation 9cbc7b9f-9d62-433f-81d8-56ab5f6b919b · outbound

This paper cites T2VPhysBench: A First-Principles Benchmark for Physical Consistency in Text-to-Video Generation.

Do Joint Audio-Video Generation Models Understand Physics? T2VPhysBench: A First-Principles Benchmark for Physical Consistency in Text-to-Video Generation

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-06-30T23:45:08.258303Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-30T23:39:22.070629Z digest=sha256:3e8f4fa424b5ade53a44fd7dd670cba6d5037d8c959654d3050742b050c4e6f3

Observation 5975024c-9ed4-454a-b844-4ca3c6bf49ab · outbound

This paper cites LTX-2: Efficient Joint Audio-Visual Foundation Model.

Do Joint Audio-Video Generation Models Understand Physics? LTX-2: Efficient Joint Audio-Visual Foundation Model

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-06-30T23:45:08.209925Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-30T23:39:22.070629Z digest=sha256:05fb693732561a526a4d03ed5d1faa12bb6838e8d7c1a27c0f54de6e97e824f8

Observation 03d9c9c7-4aa9-46ad-8eea-6c77da3d07b2 · outbound

This paper cites Videoscore: Building automatic metrics to simulate fine-grained human feedback for video generation.

Do Joint Audio-Video Generation Models Understand Physics? Videoscore: Building automatic metrics to simulate fine-grained human feedback for video generation

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T10:33:40.257251Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-30T23:39:22.070629Z digest=sha256:f9973f7fdd3ef24d964b15d6c4d50b2d381a41fb199205fe003c00d92fce624b

Observation d4e9e043-088a-44f1-af6c-7cb6fa3e6914 · outbound

This paper cites Clipscore: A reference-free evaluation metric for image captioning.

Do Joint Audio-Video Generation Models Understand Physics? Clipscore: A reference-free evaluation metric for image captioning

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T10:33:40.261363Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-30T23:39:22.070629Z digest=sha256:8dac97312f3584df76e1d0a22a063eada760cc47d64077609d56d68d8ff7d5b3

Observation 3d3576bf-83f4-4c9a-9439-5d73cfdcf179 · outbound

This paper cites VABench: A Comprehensive Benchmark for Audio-Video Generation.

Do Joint Audio-Video Generation Models Understand Physics? VABench: A Comprehensive Benchmark for Audio-Video Generation

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-06-30T23:45:08.228667Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-30T23:39:22.070629Z digest=sha256:1cf9f435c6b0da0a71519309ea4f3a03dbda186bb13329c3d4baa38e8fe61045

Observation ad0976e0-f54c-4ff2-9944-6ad393062896 · outbound

This paper cites Vbench: Comprehensive benchmark suite for video generative models.

Do Joint Audio-Video Generation Models Understand Physics? Vbench: Comprehensive benchmark suite for video generative models

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T10:33:40.264640Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-30T23:39:22.070629Z digest=sha256:34e26a5202da56c95ae2305be33ae69349da656eb27923f7b23065984ee604cb

Observation f3b0d21a-3025-45eb-a22c-c9dee7fb4e05 · outbound

This paper cites A reference-free metric for evaluating music enhancement algorithms.

Do Joint Audio-Video Generation Models Understand Physics? A reference-free metric for evaluating music enhancement algorithms

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T10:33:40.251522Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-30T23:39:22.070629Z digest=sha256:8b13beb2625dd84d3769e26f9d5eff5b6d7a75aa2db0fe23ba23e16f01daa62c

Observation 7085bb9c-47fe-4f79-8bca-cc182a94ab33 · outbound

This paper cites The measurement of observer agreement for categorical data.biometrics, pages 159–174.

Do Joint Audio-Video Generation Models Understand Physics? The measurement of observer agreement for categorical data.biometrics, pages 159–174

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T10:33:40.242940Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-30T23:39:22.070629Z digest=sha256:61a027397053bb72f945e917b7446e9333ff1d1cc464cd210d9d300931036d73

Observation ad03681c-d52a-49b6-a5ef-e09c1e0c8229 · outbound

This paper cites Video generation models: A survey of post-training and alignment.

Do Joint Audio-Video Generation Models Understand Physics? Video generation models: A survey of post-training and alignment

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T10:33:40.240842Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-30T23:39:22.070629Z digest=sha256:bd22c2b915531d90d9da696ed998a02bfb4eba44fc1c62bb0c87d5ef411f21d7

Observation 0317efd0-ebfc-4283-9024-76ef8e8627be · outbound

This paper cites JavisDiT: Joint audio-video diffusion transformer with hierarchical spatio-temporal prior synchronization.

Do Joint Audio-Video Generation Models Understand Physics? JavisDiT: Joint audio-video diffusion transformer with hierarchical spatio-temporal prior synchronization

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-06-30T23:45:08.195881Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-30T23:39:22.070629Z digest=sha256:a0ede67a8dfd1e21db7218cb3a4dd8e52479d115c05d21949de95157739d6886

Observation 0770c935-9b86-4a56-92d2-4906dbd1ab50 · outbound

This paper cites Ilya Loshchilov and Frank Hutter.

Do Joint Audio-Video Generation Models Understand Physics? Ilya Loshchilov and Frank Hutter

Reference 25

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T23:45:08.249848Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-30T23:39:22.070629Z digest=sha256:42badf11c2275ae2763a28a72a1040093ee0d08313f9d3f92b44ad569b477003

Observation 748f8c98-a274-41a2-8867-6dff29af8db8 · outbound

This paper cites Caven: An embodied conversational agent for efficient audio-visual navigation in noisy environments.

Do Joint Audio-Video Generation Models Understand Physics? Caven: An embodied conversational agent for efficient audio-visual navigation in noisy environments

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T10:33:40.247169Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-30T23:39:22.070629Z digest=sha256:b87e2fe47f8fbc8ebe9e830b7d357cad958f35233d6db8541becfa9aca550dee

Observation 9eb70423-d81d-4dac-a462-08c57376539d · outbound

This paper cites Tell what you hear from what you see-video to audio generation through text.Advances in Neural Information Processing Systems, 37:101337– 101366, 2024.

Do Joint Audio-Video Generation Models Understand Physics? Tell what you hear from what you see-video to audio generation through text.Advances in Neural Information Processing Systems, 37:101337– 101366, 2024

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T10:33:40.249135Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-30T23:39:22.070629Z digest=sha256:04c83ddc43675939ab3eb1f28b86373527938446e2610d86ff43b8a9725cf745

Observation 32cb4ad9-c698-4217-9117-a4f494c3acf7 · outbound

This paper cites Ovi: Twin Backbone Cross-Modal Fusion for Audio-Video Generation.

Do Joint Audio-Video Generation Models Understand Physics? Ovi: Twin Backbone Cross-Modal Fusion for Audio-Video Generation

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-06-30T23:45:08.269827Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-30T23:39:22.070629Z digest=sha256:d7d9b661a883d88a93a40e9e16fd606b4505bb5bc9e5419c5c5c5eb2ee312ff3

Observation 5bd31d78-326f-4c67-bd83-a5f8bc71f341 · outbound

This paper cites Tavgbench: Benchmarking text to audible-video generation.

Do Joint Audio-Video Generation Models Understand Physics? Tavgbench: Benchmarking text to audible-video generation

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T10:33:40.239063Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-30T23:39:22.070629Z digest=sha256:e959d5055cf6324aeaf422a1ab7b4370797ce86646780ef36d03bf890feda093

Observation 5a7ff603-bb63-429c-94d6-992a18facb2a · outbound

This paper cites Towards World Simulator: Crafting Physical Commonsense-Based Benchmark for Video Generation.

Do Joint Audio-Video Generation Models Understand Physics? Towards World Simulator: Crafting Physical Commonsense-Based Benchmark for Video Generation

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-06-30T23:45:08.217598Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-30T23:39:22.070629Z digest=sha256:570550f75974bf544f53fd80f3c799ab3ea232a74fedc9bcdd0061bddc0a4152

Observation 3dfafc7a-5b28-4f36-b88c-4f1d3cc6e3f8 · outbound

This paper cites Do gener- ative video models understand physical principles? InProceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, pages 948–958.

Do Joint Audio-Video Generation Models Understand Physics? Do gener- ative video models understand physical principles? InProceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, pages 948–958

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T10:33:40.237267Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-30T23:39:22.070629Z digest=sha256:b32fdaa3bdebb37a0a78a1b7d37d1801361c9bd56b42ce61fddc633a3d5ad777

Observation 566884d4-dc78-4da6-b2f2-fe67cda4520c · outbound

This paper cites Sora 2, 2025.

Do Joint Audio-Video Generation Models Understand Physics? Sora 2, 2025

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T10:33:40.233544Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-30T23:39:22.070629Z digest=sha256:5bc4f7ad453251595869895b45f941e0fe812e374eecae2499f8ac8938a5c9b3

Observation f703365e-a69d-494b-bbf3-5a12f52a8deb · outbound

This paper cites OmniSonic: Towards Universal and Holistic Audio Generation from Video and Text.

Do Joint Audio-Video Generation Models Understand Physics? OmniSonic: Towards Universal and Holistic Audio Generation from Video and Text

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-06-30T23:45:08.233900Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-30T23:39:22.070629Z digest=sha256:7c2d7b8a20fc9e1106b7b13ad54a55b6b5f86c49314bdb5efcdf656489d3fa4e

Observation d376540a-2246-45e3-be74-25853a0dd631 · outbound

This paper cites Seedance 2.0: Advancing Video Generation for World Complexity.

Do Joint Audio-Video Generation Models Understand Physics? Seedance 2.0: Advancing Video Generation for World Complexity

Reference 34

Resolution
verified exact
local_arxiv, observed 2026-06-30T23:45:08.275677Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-30T23:39:22.070629Z digest=sha256:0fe6c846b1f66c9d1e2d3f0c74271f884f958d18fb8e6212e548a2234d87c77e

Observation 084b07ca-7a41-4ddc-9b31-a0b1e4913807 · outbound

This paper cites Savgbench: Benchmarking spatially aligned audio-video generation.

Do Joint Audio-Video Generation Models Understand Physics? Savgbench: Benchmarking spatially aligned audio-video generation

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T10:33:40.231900Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-30T23:39:22.070629Z digest=sha256:478558f66f8770173d4b14ee7ef2656dc095df12d2356c91e5bbb2caf23f4438

Observation 6f2e1459-40fc-48d7-b3f0-8ab806ece0f3 · outbound

This paper cites OpenAI GPT-5 System Card.

Do Joint Audio-Video Generation Models Understand Physics? OpenAI GPT-5 System Card

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-06-30T23:45:08.207476Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-30T23:39:22.070629Z digest=sha256:dbc99e393baf75da8cc698d7d2dd2410ae5ca45e3e3e5f597f6d340b924c1a3a

Observation 4c469386-d564-4fbb-88f1-64ad2f9fc04e · outbound

This paper cites From Vision to Audio and Beyond: A Unified Model for Audio-Visual Representation and Generation.

Do Joint Audio-Video Generation Models Understand Physics? From Vision to Audio and Beyond: A Unified Model for Audio-Visual Representation and Generation

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-06-30T23:45:08.234269Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-30T23:39:22.070629Z digest=sha256:7987d4ae170fc79c7e61c302f572168ad95d089955f5afd58242b8e3d5209618

Observation 8b2f2e9b-780d-47b1-9ca5-5bb0fc51fc2d · outbound

This paper cites T2v- compbench: A comprehensive benchmark for compositional text-to-video generation.

Do Joint Audio-Video Generation Models Understand Physics? T2v- compbench: A comprehensive benchmark for compositional text-to-video generation

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T10:33:40.235367Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-30T23:39:22.070629Z digest=sha256:a6ef4bd12bba34430badb0dec019f80164f605d3c9ad7a366f463a9b2d92e369

Observation f75f7dd6-302c-441e-990d-695c2314aa8a · outbound

This paper cites Sonicbench: Dissecting the physical perception bottleneck in large audio language models.

Do Joint Audio-Video Generation Models Understand Physics? Sonicbench: Dissecting the physical perception bottleneck in large audio language models

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-06-30T23:45:08.208655Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-30T23:39:22.070629Z digest=sha256:8fcb0b36b3688f43829d6cb9982cf9c42d3a11892df38fff10f10791ee895153

Observation b99f848b-3f92-44c0-b1c3-d1f9628e34e9 · outbound

This paper cites Kling-Omni Technical Report.

Do Joint Audio-Video Generation Models Understand Physics? Kling-Omni Technical Report

Reference 40

Resolution
verified exact
local_arxiv, observed 2026-06-30T23:45:08.219778Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-30T23:39:22.070629Z digest=sha256:7053e13e9076d9a99877c6b05ee27de32205aba9d2a71a68029fe171d45f3428

Observation 50d22eea-2a04-447b-aa7b-793f65bc697f · outbound

This paper cites Qwen3.5-Omni Technical Report.

Do Joint Audio-Video Generation Models Understand Physics? Qwen3.5-Omni Technical Report

Reference 41

Resolution
verified exact
local_arxiv, observed 2026-06-30T23:45:08.255291Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-30T23:39:22.070629Z digest=sha256:2e1796d74e0c85def2676b6c9343ca27fbb62bac9cd252af0c9cffcbcf9886bd

Observation cd596f24-076e-4356-a092-3b74d53b140e · outbound

This paper cites Towards Accurate Generative Models of Video: A New Metric & Challenges.

Do Joint Audio-Video Generation Models Understand Physics? Towards Accurate Generative Models of Video: A New Metric & Challenges

Reference 42

Resolution
verified exact
local_arxiv, observed 2026-06-30T23:45:08.226099Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-30T23:39:22.070629Z digest=sha256:9534a87c6e9ee8a9a6a5c9923e53653d0c4db51b7c593851a688ffcec830a63b

Observation 158c2ec1-1e00-4bfa-9c73-b4ea28a3f037 · outbound

This paper cites UniVerse-1: Unified Audio-Video Generation via Stitching of Experts.

Do Joint Audio-Video Generation Models Understand Physics? UniVerse-1: Unified Audio-Video Generation via Stitching of Experts

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-06-30T23:45:08.252808Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-30T23:39:22.070629Z digest=sha256:baae6b35c4463d0372a562dfafd9aebf722c7227f6f02b440fd7926a80fbca27

Observation 6b087763-ae3e-4f2a-bce9-60e6a324b030 · outbound

This paper cites PhyAVBench: A Challenging Audio Physics-Sensitivity Benchmark for Physically Grounded Text-to-Audio-Video Generation.

Do Joint Audio-Video Generation Models Understand Physics? PhyAVBench: A Challenging Audio Physics-Sensitivity Benchmark for Physically Grounded Text-to-Audio-Video Generation

Reference 44

Resolution
verified exact
local_arxiv, observed 2026-06-30T23:45:08.237129Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-30T23:39:22.070629Z digest=sha256:f86b85727f5dc33ce2de845a6cc5d7c0df6494e17833c73b369b7f0d640d99eb

Observation 5ebebb9b-8765-4ce0-aac4-710fb43f94cd · outbound

This paper cites A survey on video diffusion models.ACM Computing Surveys, 57(2):1–42.

Do Joint Audio-Video Generation Models Understand Physics? A survey on video diffusion models.ACM Computing Surveys, 57(2):1–42

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T10:33:40.253226Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-30T23:39:22.070629Z digest=sha256:9ccfe98db82926382e41e09290a0bce7424fb0767fe46c1b16e7d96103c56e78

Observation f7c47df5-db7c-48c4-af0f-dd15dd4d8eca · outbound

This paper cites A Systematic Post-Train Framework for Video Generation.

Do Joint Audio-Video Generation Models Understand Physics? A Systematic Post-Train Framework for Video Generation

Reference 46

Resolution
verified exact
local_arxiv, observed 2026-06-30T23:45:08.213661Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-30T23:39:22.070629Z digest=sha256:1efb6aeef036669f9561e3626cbb32a2a59a7ef09b8928d7eb773ab6d3d9ba65

Observation 3221fd1e-30a0-4801-a237-d9b5617ed036 · outbound

This paper cites ReAct: Synergizing Reasoning and Acting in Language Models.

Do Joint Audio-Video Generation Models Understand Physics? ReAct: Synergizing Reasoning and Acting in Language Models

Reference 47

Resolution
verified exact
local_arxiv, observed 2026-06-30T23:45:08.266662Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-30T23:39:22.070629Z digest=sha256:d4c0059617ed54f1485a37409a8316148492e22a078b16e720742e4c1833acee

Observation eebe27e9-d3a5-46bd-80e6-a1cc79960dd5 · outbound

This paper cites Diverse and aligned audio-to-video generation via text-to-video model adaptation.

Do Joint Audio-Video Generation Models Understand Physics? Diverse and aligned audio-to-video generation via text-to-video model adaptation

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T10:33:40.255030Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-30T23:39:22.070629Z digest=sha256:ab284c4b8b7314321958e537bdd1e580aedb3ed9c4e1abc7ff89e8c93c7cb7df

Observation da48796b-6c81-455b-8b34-3c8f4ba82f53 · outbound

This paper cites Virbo: Multimodal Multilingual Avatar Video Generation in Digital Marketing.

Do Joint Audio-Video Generation Models Understand Physics? Virbo: Multimodal Multilingual Avatar Video Generation in Digital Marketing

Reference 49

Resolution
malformed identifier
arxiv_id, observed 2026-06-30T23:45:08.273174Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-30T23:39:22.070629Z digest=sha256:62fbf04c7fb96d99b425ef7029492a3d16a6301dfab859de8a96bee4021bbf0a

Observation fb640097-828b-45fa-a630-04466d5e7927 · outbound

This paper cites an unresolved cited work.

Do Joint Audio-Video Generation Models Understand Physics? Unresolved cited work

Reference 50

Resolution
unresolved
raw_fallback, observed 2026-07-07T10:33:40.263014Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-30T23:39:22.070629Z digest=sha256:f14ecc5593cb21aac0ef4f6f3ad29c418d5a233139da8d1b09d0c8febf781b65

Observation b2ccf433-5f3d-40cc-a41d-fa46ed0d7a46 · outbound

This paper cites {video.event}.

Do Joint Audio-Video Generation Models Understand Physics? {video.event}

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T10:33:40.275271Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-30T23:39:22.070629Z digest=sha256:983e373823119c22ad3cbd1453aaad9ab69e1c85ca17ba6184d38fab9570833c

Observation 8b5a116c-98b3-4886-827c-d5f100057b03 · outbound

This paper cites would normally be audible if real-world physics held; answer Yes if they are appropriately represented as such (typically silent here).

Do Joint Audio-Video Generation Models Understand Physics? would normally be audible if real-world physics held; answer Yes if they are appropriately represented as such (typically silent here)

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-07-07T10:33:40.230063Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-30T23:39:22.070629Z digest=sha256:aafe1430b66a11f6b027f6e7bfb2e3dadf9f698f244d048679b54b064fa08b49

Observation ea3732ce-2cc2-4940-ac13-347e774178af · outbound

This paper cites the clip is expected to be silent during the depicted event; answer Yes if it is appropriately silent throughout with no audible leak-through.

Do Joint Audio-Video Generation Models Understand Physics? the clip is expected to be silent during the depicted event; answer Yes if it is appropriately silent throughout with no audible leak-through

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-06-30T23:45:08.263931Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-30T23:39:22.070629Z digest=sha256:ee7a058e06e52ecde2fa1fda9895129d090a25ef919b400f7489df45011dc562

Pith citing papers

No inbound Pith citation observations are available.