Pith. sign in

Paper Citation Record · LEDGER

Taming Text-to-Sounding Video Generation via Advanced Modality Condition and Interaction

As of 13 August 2026, this Paper Citation Record lists 34 of 34 outbound references and 4 inbound Pith citation observations for arXiv:2510.03117.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2510.03117 v2

Coverage vector

measured 34 of 34 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-04T12:39:00.193571Z

measured 38 of 38 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-28T15:37:32.767405Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-01T22:16:16.088895Z

Reference resolution

34 of 34 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved34
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation f6d1085d-4c0f-4aaa-92ba-2decbfb146aa · outbound

This paper cites Qwen2.5-VL Technical Report.

Taming Text-to-Sounding Video Generation via Advanced Modality Condition and Interaction Qwen2.5-VL Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-04T12:38:56.059462Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T12:38:56.059462Z digest=sha256:aa703d30304094e6a9b593bca15055194e29183928da8b750cb8056532afded9

Observation cff80d68-0d8b-4dae-bd22-293af6b01781 · outbound

This paper cites Clap learning audio concepts from natural language supervision.

Taming Text-to-Sounding Video Generation via Advanced Modality Condition and Interaction Clap learning audio concepts from natural language supervision

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-04T12:38:56.528792Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T12:38:56.528792Z digest=sha256:e02a06895e4e98f74900cb9425c64e682d38c42d3b8aa1b884d9785445b6a308

Observation 914d8b6e-014e-40d0-84c3-b0e5e2701f6f · outbound

This paper cites Stable audio open.

Taming Text-to-Sounding Video Generation via Advanced Modality Condition and Interaction Stable audio open

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-04T12:38:56.683788Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T12:38:56.683788Z digest=sha256:384f94a0a58568adb1bf75af65d72baa4adf1760b57725f7b70c04735dcd156b

Observation 26c57423-8f68-49ee-9012-7e0b38030f66 · outbound

This paper cites ETVA: Evaluation of Text-to-Video Alignment via Fine-grained Question Generation and Answering.

Taming Text-to-Sounding Video Generation via Advanced Modality Condition and Interaction ETVA: Evaluation of Text-to-Video Alignment via Fine-grained Question Generation and Answering

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-04T12:38:56.851936Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T12:38:56.851936Z digest=sha256:f8292655e55cb80952487d79b220c7b223ec80e3e5aaeb0d3b58c66024adfefe

Observation 295887da-3bab-4ee3-bbf4-dd8932dc092b · outbound

This paper cites A Simple but Strong Baseline for Sounding Video Generation: Effective Adaptation of Audio and Video Diffusion Models for Joint Generation.

Taming Text-to-Sounding Video Generation via Advanced Modality Condition and Interaction A Simple but Strong Baseline for Sounding Video Generation: Effective Adaptation of Audio and Video Diffusion Models for Joint Generation

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-04T12:38:57.293402Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T12:38:57.293402Z digest=sha256:2cbb202c2a51cd873f5143113aa1a737e7fa0ca5e8292e55106d76e4da2fa139

Observation 4dca9b91-a497-4647-95c1-47c63dbb5973 · outbound

This paper cites Auto-Encoding Variational Bayes.

Taming Text-to-Sounding Video Generation via Advanced Modality Condition and Interaction Auto-Encoding Variational Bayes

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-04T12:38:57.563801Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T12:38:57.563801Z digest=sha256:431eaaa9fbd6e8f744d23c800cee3e567fac5f96ec33ded71b3fdf27bce5d2de

Observation c92327c4-f020-4527-9237-4970f7d13999 · outbound

This paper cites Understanding Sounds, Missing the Questions: The Challenge of Object Hallucination in Large Audio-Language Models.

Taming Text-to-Sounding Video Generation via Advanced Modality Condition and Interaction Understanding Sounds, Missing the Questions: The Challenge of Object Hallucination in Large Audio-Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-04T12:38:57.654823Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T12:38:57.654823Z digest=sha256:868d3591358d81cc303367c2fd9bbb758a6fea3b22a401fa20447b749dc661ef

Observation 34bee314-c44f-4b11-9eb2-b14abdf8bebe · outbound

This paper cites Sound-guided semantic video generation.

Taming Text-to-Sounding Video Generation via Advanced Modality Condition and Interaction Sound-guided semantic video generation

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-04T12:38:57.785794Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T12:38:57.785794Z digest=sha256:8dfd221da53e7ed44111a828c5e5e4add1d40df72289b4be4362cf9daa2e75a8

Observation c586ab46-a947-48f8-b992-f7cbc7d348af · outbound

This paper cites Open-Sora Plan: Open-Source Large Video Generation Model.

Taming Text-to-Sounding Video Generation via Advanced Modality Condition and Interaction Open-Sora Plan: Open-Source Large Video Generation Model

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-04T12:38:57.889197Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T12:38:57.889197Z digest=sha256:3dc041c65fb66e19203984445c86d6afbed2f04be8ef7f63c49d765719d8088f

Observation 007728a7-707a-4f8d-9c87-3944ee18eaef · outbound

This paper cites Flow Matching for Generative Modeling.

Taming Text-to-Sounding Video Generation via Advanced Modality Condition and Interaction Flow Matching for Generative Modeling

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-04T12:38:58.063672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T12:38:58.063672Z digest=sha256:6afb5912991f55de9375e3b679fe7a00e7f6a3595662960e3814ba9d9a40be11

Observation cf197551-189e-4413-bfbb-75800a2a5d56 · outbound

This paper cites AudioLDM: Text-to-Audio Generation with Latent Diffusion Models.

Taming Text-to-Sounding Video Generation via Advanced Modality Condition and Interaction AudioLDM: Text-to-Audio Generation with Latent Diffusion Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-04T12:38:58.198477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T12:38:58.198477Z digest=sha256:3ba7cc5ab8843d9f2a753e517ca1de49631752e074ed1e99c64b7dad6b51074f

Observation 42b6baa9-358b-4bba-97f5-d90724595e5d · outbound

This paper cites SyncFlow: Toward Temporally Aligned Joint Audio-Video Generation from Text.

Taming Text-to-Sounding Video Generation via Advanced Modality Condition and Interaction SyncFlow: Toward Temporally Aligned Joint Audio-Video Generation from Text

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-04T12:38:58.383496Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T12:38:58.383496Z digest=sha256:e98c569b02752ca12ae96558e4840aa9ff7fbd49eaa5fe97c25eff4b919240ff

Observation aa6800e8-532a-4740-b0e4-4bc2e4687a68 · outbound

This paper cites On the Audio Hallucinations in Large Audio-Video Language Models.

Taming Text-to-Sounding Video Generation via Advanced Modality Condition and Interaction On the Audio Hallucinations in Large Audio-Video Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-04T12:38:58.546588Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T12:38:58.546588Z digest=sha256:44317291651752b2d978618bf145318926dc59843c7367bd62672e6a2b70f2ba

Observation 55358733-e763-4e3a-98b1-45ca7a1da5a8 · outbound

This paper cites Scalable Diffusion Models with Transformers.

Taming Text-to-Sounding Video Generation via Advanced Modality Condition and Interaction Scalable Diffusion Models with Transformers

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-04T12:38:58.741505Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T12:38:58.741505Z digest=sha256:ffa419783c747e77bba20a324fea0095507a84254c01ab0464cfc8ede8a37cfa

Observation 596cd1fa-c47f-4aef-92c1-d4af1235c8dd · outbound

This paper cites Progressive Distillation for Fast Sampling of Diffusion Models.

Taming Text-to-Sounding Video Generation via Advanced Modality Condition and Interaction Progressive Distillation for Fast Sampling of Diffusion Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-04T12:38:58.851352Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T12:38:58.851352Z digest=sha256:38e40e2f66cecc0ff41add49e69489aa45b76125edf658e7cdb66eab5eed22fe

Observation 89fd54c6-55f7-4915-982e-afac26dd7711 · outbound

This paper cites AVHBench: A Cross-Modal Hallucination Benchmark for Audio-Visual Large Language Models.

Taming Text-to-Sounding Video Generation via Advanced Modality Condition and Interaction AVHBench: A Cross-Modal Hallucination Benchmark for Audio-Visual Large Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-04T12:38:58.958481Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T12:38:58.958481Z digest=sha256:304f8756bc9eaac009b8c79afe1a03b829dab146d24247635f5864873735bffc

Observation 88a031f4-e492-4d90-9bbb-c414ab844e1b · outbound

This paper cites Atom of thoughts for markov llm test-time scaling.arXiv preprint arXiv:2502.12018,.

Taming Text-to-Sounding Video Generation via Advanced Modality Condition and Interaction Atom of thoughts for markov llm test-time scaling.arXiv preprint arXiv:2502.12018,

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-04T12:38:59.139436Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T12:38:59.139436Z digest=sha256:f56522bbc75f7c1e00ace853c83f4100ec524ee719236b0079860685c3699da6

Observation 2c3c091a-23c9-4a18-8790-7f1f7a60d5a1 · outbound

This paper cites Towards Accurate Generative Models of Video: A New Metric & Challenges.

Taming Text-to-Sounding Video Generation via Advanced Modality Condition and Interaction Towards Accurate Generative Models of Video: A New Metric & Challenges

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-04T12:38:59.220106Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T12:38:59.220106Z digest=sha256:23f86b086e26165ea4cbcd21742801f1eb1653277ec5e44941bdf73700793d7f

Observation 4125dad3-9ae4-4b57-a919-f3c7a845552b · outbound

This paper cites an unresolved cited work.

Taming Text-to-Sounding Video Generation via Advanced Modality Condition and Interaction Unresolved cited work

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-04T12:38:59.384196Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T12:38:59.384196Z digest=sha256:e9e3aada8c2cc1af9d88015950360f47b6d6b0dfc980ef39a0952df6f6ceceb9

Observation d0ec8403-afbd-4bd5-8376-38d561ef2c3f · outbound

This paper cites HunyuanVideo: A Systematic Framework For Large Video Generative Models.

Taming Text-to-Sounding Video Generation via Advanced Modality Condition and Interaction HunyuanVideo: A Systematic Framework For Large Video Generative Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-04T12:38:59.460371Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T12:38:59.460371Z digest=sha256:34a9d7d7dafa6e3f51c3ae375f84f1651b070cade6e0566c43924b250bd5a49d

Observation 6af38f93-5e40-4170-9be7-204739b9be4c · outbound

This paper cites Qwen-Image Technical Report.

Taming Text-to-Sounding Video Generation via Advanced Modality Condition and Interaction Qwen-Image Technical Report

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-04T12:38:59.540892Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T12:38:59.540892Z digest=sha256:9a7795b81eb2e71a7b39e6b41f6540aae9b59852cfde71100cb42de025b8c6b7

Observation f022d557-bf23-4f7e-953e-ee71902b74ad · outbound

This paper cites Qwen2.5-Omni Technical Report.

Taming Text-to-Sounding Video Generation via Advanced Modality Condition and Interaction Qwen2.5-Omni Technical Report

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-04T12:38:59.637474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T12:38:59.637474Z digest=sha256:15b8ccb01ddb6e7424e207532df3f3539c6f0062b7b8a49274f108f196fa26f9

Observation 2ebabfad-8135-401c-ad68-f327bf9b834e · outbound

This paper cites Qwen2.5 Technical Report.

Taming Text-to-Sounding Video Generation via Advanced Modality Condition and Interaction Qwen2.5 Technical Report

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-04T12:38:59.738442Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T12:38:59.738442Z digest=sha256:d1ec94ef146b14ad29240ead3830bbb16a156e1a49db307073c8463c36dec3fe

Observation fe916510-e314-4523-bf4a-49e286c8d953 · outbound

This paper cites UniForm: A Unified Multi-Task Diffusion Transformer for Audio-Video Generation.

Taming Text-to-Sounding Video Generation via Advanced Modality Condition and Interaction UniForm: A Unified Multi-Task Diffusion Transformer for Audio-Video Generation

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-04T12:38:59.844041Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T12:38:59.844041Z digest=sha256:d5e929a7dd4305c29a347605dded283f405302c94c35a3d99aae0cf4170cd896

Observation 04060fd3-7145-471d-96fb-d3432a5d9c37 · outbound

This paper cites Open-Sora: Democratizing Efficient Video Production for All.

Taming Text-to-Sounding Video Generation via Advanced Modality Condition and Interaction Open-Sora: Democratizing Efficient Video Production for All

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-04T12:38:59.929675Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T12:38:59.929675Z digest=sha256:fcba564733608137cac9ecfb310116a24f9e13c23b0f536be8b474a1e24e49a1

Observation 6b6a6274-63ef-4a1d-bc98-7fe359387da2 · outbound

This paper cites This technique steers the generation pro- cess towards a desired conditionc(e.g., a text prompt) without needing an external classifier.

Taming Text-to-Sounding Video Generation via Advanced Modality Condition and Interaction This technique steers the generation pro- cess towards a desired conditionc(e.g., a text prompt) without needing an external classifier

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-04T12:39:00.193571Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T12:39:00.193571Z digest=sha256:f9cfb566462620b45446455ff91fb169882dbfccc72ebde9fe626d3367c5b11e

Observation d3105885-199f-46f0-ad1f-7ea8fa4ef1b7 · outbound

This paper cites 3D U-Net: Learning Dense Volumetric Segmentation from Sparse Annotation.

Taming Text-to-Sounding Video Generation via Advanced Modality Condition and Interaction 3D U-Net: Learning Dense Volumetric Segmentation from Sparse Annotation

Reference 2016

Resolution
unresolved
no resolver link, observed 2026-08-04T12:39:00.038859Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T12:39:00.038859Z digest=sha256:cef41d80fbda00da51dc27f5252faed2be7bc99926a7f987b34a7ac80362e713

Observation d25f7a8e-4b31-46dd-bab6-4a479a67d7ec · outbound

This paper cites Wan: Open and Advanced Large-Scale Video Generative Models.

Taming Text-to-Sounding Video Generation via Advanced Modality Condition and Interaction Wan: Open and Advanced Large-Scale Video Generative Models

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-04T12:38:59.279588Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T12:38:59.279588Z digest=sha256:1892275aeace2b139ab67040bad924cc69137e0dc2c106a69d064dd711a7bcaa

Observation b2e16779-9699-405d-8958-8f66d33a3c9f · outbound

This paper cites Denoising Diffusion Probabilistic Models.

Taming Text-to-Sounding Video Generation via Advanced Modality Condition and Interaction Denoising Diffusion Probabilistic Models

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-04T12:38:57.138823Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T12:38:57.138823Z digest=sha256:1530d28590591383df236e1c49168e44ed1f3feceb9b882bbb5b41548a0d4dcb

Observation 23e273f7-3da5-4cf4-9ee3-4e4cefa5f5fa · outbound

This paper cites Qwen2-Audio Technical Report.

Taming Text-to-Sounding Video Generation via Advanced Modality Condition and Interaction Qwen2-Audio Technical Report

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-04T12:38:56.384224Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T12:38:56.384224Z digest=sha256:e829b6bb49ca6aa97b5d033550e536ee4f63a9bd78e9fe4b1e0755146221dfaa

Observation e8144d74-5a43-4bb6-85dd-216833630784 · outbound

This paper cites Fr\'echet Audio Distance: A Metric for Evaluating Music Enhancement Algorithms.

Taming Text-to-Sounding Video Generation via Advanced Modality Condition and Interaction Fr\'echet Audio Distance: A Metric for Evaluating Music Enhancement Algorithms

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-04T12:38:57.438703Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T12:38:57.438703Z digest=sha256:22c9b7da97dc1da40e190af74a945e02f783b42b287ec535a29e4b9dbc7f4d69

Observation a00d6439-4197-49e2-a6c5-bfd1049acff1 · outbound

This paper cites Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets.

Taming Text-to-Sounding Video Generation via Advanced Modality Condition and Interaction Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-04T12:38:56.194933Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T12:38:56.194933Z digest=sha256:f56f7d88f1352e2e38aefccee6e7a140b99c8af4a0f885f0823277bd2182212f

Observation 11f089b3-86bc-4e12-8093-03ad70f608d8 · outbound

This paper cites TEMPO: Prompt-based Generative Pre-trained Transformer for Time Series Forecasting.

Taming Text-to-Sounding Video Generation via Advanced Modality Condition and Interaction TEMPO: Prompt-based Generative Pre-trained Transformer for Time Series Forecasting

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-04T12:38:56.312366Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T12:38:56.312366Z digest=sha256:75eba001df45f73dd7766cdb76807889e0bb553d218949397d8837dc008b3b72

Observation 4fc2be78-903a-4988-b3c2-2a40cfc66194 · outbound

This paper cites Classifier-Free Diffusion Guidance.

Taming Text-to-Sounding Video Generation via Advanced Modality Condition and Interaction Classifier-Free Diffusion Guidance

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-04T12:38:56.998200Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T12:38:56.998200Z digest=sha256:a7604d5c3227a980a2783421aebd3aeca7316333f478dde3b6d02c7cd811c339

Pith citing papers

Observation ba269173-b442-4f69-8027-50f31d2a4c67 · inbound

LTX-2: Efficient Joint Audio-Visual Foundation Model cites this paper.

LTX-2: Efficient Joint Audio-Visual Foundation Model Taming Text-to-Sounding Video Generation via Advanced Modality Condition and Interaction

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-06-29T02:14:22.202186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-13T07:06:20.470686Z digest=sha256:3e90a8cd64da96e8012cb7d39282d85ab1023e1b2a513b1ca4194525e1f0aae6

Observation 82aa4fc2-5520-4e25-8caa-efde5e885ae3 · inbound

Script-a-Video: Deep Structured Audio-visual Captions via Factorized Streams and Relational Grounding cites this paper.

Script-a-Video: Deep Structured Audio-visual Captions via Factorized Streams and Relational Grounding Taming Text-to-Sounding Video Generation via Advanced Modality Condition and Interaction

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T02:14:22.202186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T15:07:45.595260Z digest=sha256:94d0a4cdacc1283c37d492f57cbe4498aa6d74dbc8549e61621f20d9a7d39ea9

Observation 8d553e74-c230-462f-91a2-63a9399b64b8 · inbound

SyncDPO: Enhancing Temporal Synchronization in Video-Audio Joint Generation via Preference Learning cites this paper.

SyncDPO: Enhancing Temporal Synchronization in Video-Audio Joint Generation via Preference Learning Taming Text-to-Sounding Video Generation via Advanced Modality Condition and Interaction

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-06-29T02:14:22.202186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-13T07:01:40.333448Z digest=sha256:3c84ab4f31cd728b58748b612ff7b7864bfb75211402f38a4d612c7ef94765cf

Observation 4f678117-5b37-4ccd-8b31-428359713c1e · inbound

Planar Symmetric Pattern Generation cites this paper.

Planar Symmetric Pattern Generation Taming Text-to-Sounding Video Generation via Advanced Modality Condition and Interaction

Reference 73

Resolution
verified exact
local_arxiv, observed 2026-07-01T22:16:16.090144Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-06-28T15:37:32.767405Z digest=sha256:e7ea4b9000f4d53545c2b67705effd6d1ba5f6529f0de77c8a5bc84f2a81f206