Pith. sign in

Paper Citation Record · LEDGER

Taming Text-to-Sounding Video Generation via Advanced Modality Condition and Interaction

As of 9 August 2026, this Paper Citation Record lists 34 of 34 outbound references and 4 inbound Pith citation observations for arXiv:2510.03117.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2510.03117 v2

Coverage vector

measured 34 of 34 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-04T12:39:00.193571Z

measured 38 of 38 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-28T15:37:32.767405Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-01T22:16:16.088895Z

Reference resolution

34 of 34 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved34
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation f6d1085d-4c0f-4aaa-92ba-2decbfb146aa · outbound

This paper cites Qwen2.5-VL Technical Report.

Taming Text-to-Sounding Video Generation via Advanced Modality Condition and Interaction Qwen2.5-VL Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-04T12:38:56.059462Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T12:38:56.059462Z digest=sha256:90390755d8b73578ce1325d7785d91fe33999f1596134c706515b0eade9486be

Observation cff80d68-0d8b-4dae-bd22-293af6b01781 · outbound

This paper cites Clap learning audio concepts from natural language supervision.

Taming Text-to-Sounding Video Generation via Advanced Modality Condition and Interaction Clap learning audio concepts from natural language supervision

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-04T12:38:56.528792Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T12:38:56.528792Z digest=sha256:3486cfed73db01d80ad6e055a2c9e1bbb1050bbce759bda5f10a32c4a7fdcb39

Observation 914d8b6e-014e-40d0-84c3-b0e5e2701f6f · outbound

This paper cites Stable audio open.

Taming Text-to-Sounding Video Generation via Advanced Modality Condition and Interaction Stable audio open

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-04T12:38:56.683788Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T12:38:56.683788Z digest=sha256:f9aa707a3a53dac601a6b209421f0f3dcd7be3074e7d3d398009f1cd887597d0

Observation 26c57423-8f68-49ee-9012-7e0b38030f66 · outbound

This paper cites ETVA: Evaluation of Text-to-Video Alignment via Fine-grained Question Generation and Answering.

Taming Text-to-Sounding Video Generation via Advanced Modality Condition and Interaction ETVA: Evaluation of Text-to-Video Alignment via Fine-grained Question Generation and Answering

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-04T12:38:56.851936Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T12:38:56.851936Z digest=sha256:7491ffcd47a1d3369dd28b1e75ca719c46075addf77b1a94ac5a75cc0a402651

Observation 295887da-3bab-4ee3-bbf4-dd8932dc092b · outbound

This paper cites A Simple but Strong Baseline for Sounding Video Generation: Effective Adaptation of Audio and Video Diffusion Models for Joint Generation.

Taming Text-to-Sounding Video Generation via Advanced Modality Condition and Interaction A Simple but Strong Baseline for Sounding Video Generation: Effective Adaptation of Audio and Video Diffusion Models for Joint Generation

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-04T12:38:57.293402Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T12:38:57.293402Z digest=sha256:69ada80297b73629849a48b6279b7c6802c9a4f3c63b7a5b2fcec3577c8de406

Observation 4dca9b91-a497-4647-95c1-47c63dbb5973 · outbound

This paper cites Auto-Encoding Variational Bayes.

Taming Text-to-Sounding Video Generation via Advanced Modality Condition and Interaction Auto-Encoding Variational Bayes

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-04T12:38:57.563801Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T12:38:57.563801Z digest=sha256:301c480a8751f611b17a2ea7023315a74285aeb07033c9dc3a7353a667bfb27c

Observation c92327c4-f020-4527-9237-4970f7d13999 · outbound

This paper cites Understanding Sounds, Missing the Questions: The Challenge of Object Hallucination in Large Audio-Language Models.

Taming Text-to-Sounding Video Generation via Advanced Modality Condition and Interaction Understanding Sounds, Missing the Questions: The Challenge of Object Hallucination in Large Audio-Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-04T12:38:57.654823Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T12:38:57.654823Z digest=sha256:aa2237a9113295a2ca9f2295a193a57cabf4d8020d979e7503dc85ae44f16c0a

Observation 34bee314-c44f-4b11-9eb2-b14abdf8bebe · outbound

This paper cites Sound-guided semantic video generation.

Taming Text-to-Sounding Video Generation via Advanced Modality Condition and Interaction Sound-guided semantic video generation

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-04T12:38:57.785794Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T12:38:57.785794Z digest=sha256:7f9b460ce583a8a5b98709c198795f9b58777b763b145b3f7d95e7b3f94031b3

Observation c586ab46-a947-48f8-b992-f7cbc7d348af · outbound

This paper cites Open-Sora Plan: Open-Source Large Video Generation Model.

Taming Text-to-Sounding Video Generation via Advanced Modality Condition and Interaction Open-Sora Plan: Open-Source Large Video Generation Model

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-04T12:38:57.889197Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T12:38:57.889197Z digest=sha256:de945660c54b09b50303f7ab6f4aa33347feedb20f74e96396e9f0da95e6b3d3

Observation 007728a7-707a-4f8d-9c87-3944ee18eaef · outbound

This paper cites Flow Matching for Generative Modeling.

Taming Text-to-Sounding Video Generation via Advanced Modality Condition and Interaction Flow Matching for Generative Modeling

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-04T12:38:58.063672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T12:38:58.063672Z digest=sha256:fddaf0b291758e21a1bdd74381fa3a264ac5409f7650298cd4c30e0cb4c80dd5

Observation cf197551-189e-4413-bfbb-75800a2a5d56 · outbound

This paper cites AudioLDM: Text-to-Audio Generation with Latent Diffusion Models.

Taming Text-to-Sounding Video Generation via Advanced Modality Condition and Interaction AudioLDM: Text-to-Audio Generation with Latent Diffusion Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-04T12:38:58.198477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T12:38:58.198477Z digest=sha256:f05399bd0a62602a46e72551080b5d7df2fdb7543d5586d2e28de7e8524de3fd

Observation 42b6baa9-358b-4bba-97f5-d90724595e5d · outbound

This paper cites SyncFlow: Toward Temporally Aligned Joint Audio-Video Generation from Text.

Taming Text-to-Sounding Video Generation via Advanced Modality Condition and Interaction SyncFlow: Toward Temporally Aligned Joint Audio-Video Generation from Text

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-04T12:38:58.383496Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T12:38:58.383496Z digest=sha256:4c173a3cf788d5025bca1778046c57706ca25a39d396b670d89f1a476f8177bf

Observation aa6800e8-532a-4740-b0e4-4bc2e4687a68 · outbound

This paper cites On the Audio Hallucinations in Large Audio-Video Language Models.

Taming Text-to-Sounding Video Generation via Advanced Modality Condition and Interaction On the Audio Hallucinations in Large Audio-Video Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-04T12:38:58.546588Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T12:38:58.546588Z digest=sha256:a1363e26e02efec26cfa54d33569567ce97a02ddae3b60e55359a2e00984a04c

Observation 55358733-e763-4e3a-98b1-45ca7a1da5a8 · outbound

This paper cites Scalable Diffusion Models with Transformers.

Taming Text-to-Sounding Video Generation via Advanced Modality Condition and Interaction Scalable Diffusion Models with Transformers

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-04T12:38:58.741505Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T12:38:58.741505Z digest=sha256:db2f5c71895e23f243f6edb94db686fa9b332c7168fc8819964780c1462e0909

Observation 596cd1fa-c47f-4aef-92c1-d4af1235c8dd · outbound

This paper cites Progressive Distillation for Fast Sampling of Diffusion Models.

Taming Text-to-Sounding Video Generation via Advanced Modality Condition and Interaction Progressive Distillation for Fast Sampling of Diffusion Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-04T12:38:58.851352Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T12:38:58.851352Z digest=sha256:3ed7c4e67af6b7c04d0f1b23b86d02b0bddaab8b4efe5f5e69ca92d857cd134a

Observation 89fd54c6-55f7-4915-982e-afac26dd7711 · outbound

This paper cites AVHBench: A Cross-Modal Hallucination Benchmark for Audio-Visual Large Language Models.

Taming Text-to-Sounding Video Generation via Advanced Modality Condition and Interaction AVHBench: A Cross-Modal Hallucination Benchmark for Audio-Visual Large Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-04T12:38:58.958481Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T12:38:58.958481Z digest=sha256:48de95ebc7347469f8be2daa909fb1762ec4847815a428a338477e393950938b

Observation 88a031f4-e492-4d90-9bbb-c414ab844e1b · outbound

This paper cites Atom of thoughts for markov llm test-time scaling.arXiv preprint arXiv:2502.12018,.

Taming Text-to-Sounding Video Generation via Advanced Modality Condition and Interaction Atom of thoughts for markov llm test-time scaling.arXiv preprint arXiv:2502.12018,

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-04T12:38:59.139436Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T12:38:59.139436Z digest=sha256:5d1a1b1957f675a108455e79132049875d24d0b0a8dedf5efcceedae2601fba1

Observation 2c3c091a-23c9-4a18-8790-7f1f7a60d5a1 · outbound

This paper cites Towards Accurate Generative Models of Video: A New Metric & Challenges.

Taming Text-to-Sounding Video Generation via Advanced Modality Condition and Interaction Towards Accurate Generative Models of Video: A New Metric & Challenges

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-04T12:38:59.220106Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T12:38:59.220106Z digest=sha256:3b294ff972203474dee0b0c1edbd9cb701824e04096a7e6ea0da5cf8d40c90c8

Observation 4125dad3-9ae4-4b57-a919-f3c7a845552b · outbound

This paper cites an unresolved cited work.

Taming Text-to-Sounding Video Generation via Advanced Modality Condition and Interaction Unresolved cited work

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-04T12:38:59.384196Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T12:38:59.384196Z digest=sha256:4252f8aedf954df36a27cc06585431def45a9aa6dc8eb76062b67a8a2c3c2faf

Observation d0ec8403-afbd-4bd5-8376-38d561ef2c3f · outbound

This paper cites HunyuanVideo: A Systematic Framework For Large Video Generative Models.

Taming Text-to-Sounding Video Generation via Advanced Modality Condition and Interaction HunyuanVideo: A Systematic Framework For Large Video Generative Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-04T12:38:59.460371Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T12:38:59.460371Z digest=sha256:5120cf7de50cc58ea63d0a43c5d5669ce8249626dc9e6c1a50bf6f53301a2596

Observation 6af38f93-5e40-4170-9be7-204739b9be4c · outbound

This paper cites Qwen-Image Technical Report.

Taming Text-to-Sounding Video Generation via Advanced Modality Condition and Interaction Qwen-Image Technical Report

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-04T12:38:59.540892Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T12:38:59.540892Z digest=sha256:bff8d15f0d59ddce1456d2f3c545453b890ca4edc060da95af989e6f92aeacb7

Observation f022d557-bf23-4f7e-953e-ee71902b74ad · outbound

This paper cites Qwen2.5-Omni Technical Report.

Taming Text-to-Sounding Video Generation via Advanced Modality Condition and Interaction Qwen2.5-Omni Technical Report

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-04T12:38:59.637474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T12:38:59.637474Z digest=sha256:79233f340ea82682e9a423f10e7e11f7c0edb9a0c7a44f6d18eeccc9e45b7645

Observation 2ebabfad-8135-401c-ad68-f327bf9b834e · outbound

This paper cites Qwen2.5 Technical Report.

Taming Text-to-Sounding Video Generation via Advanced Modality Condition and Interaction Qwen2.5 Technical Report

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-04T12:38:59.738442Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T12:38:59.738442Z digest=sha256:3e3fa5c819c28b8afdbe3ded9905c0cc55ec4f8849f3d97cde855d2c227d8d22

Observation fe916510-e314-4523-bf4a-49e286c8d953 · outbound

This paper cites UniForm: A Unified Multi-Task Diffusion Transformer for Audio-Video Generation.

Taming Text-to-Sounding Video Generation via Advanced Modality Condition and Interaction UniForm: A Unified Multi-Task Diffusion Transformer for Audio-Video Generation

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-04T12:38:59.844041Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T12:38:59.844041Z digest=sha256:1829b9b3e83467825128d953fb5fe27a6d98e8813b3c7ad697831669380efeed

Observation 04060fd3-7145-471d-96fb-d3432a5d9c37 · outbound

This paper cites Open-Sora: Democratizing Efficient Video Production for All.

Taming Text-to-Sounding Video Generation via Advanced Modality Condition and Interaction Open-Sora: Democratizing Efficient Video Production for All

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-04T12:38:59.929675Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T12:38:59.929675Z digest=sha256:bb5f0ef81b1faacfe67467baeca5c2babe8306e7c5ddd1b764f93ba9792ef8c6

Observation 6b6a6274-63ef-4a1d-bc98-7fe359387da2 · outbound

This paper cites This technique steers the generation pro- cess towards a desired conditionc(e.g., a text prompt) without needing an external classifier.

Taming Text-to-Sounding Video Generation via Advanced Modality Condition and Interaction This technique steers the generation pro- cess towards a desired conditionc(e.g., a text prompt) without needing an external classifier

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-04T12:39:00.193571Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T12:39:00.193571Z digest=sha256:5dee524ed7af4f6f08bc66fc56d157507101dd8765ba08ffde222311e12817dc

Observation d3105885-199f-46f0-ad1f-7ea8fa4ef1b7 · outbound

This paper cites 3D U-Net: Learning Dense Volumetric Segmentation from Sparse Annotation.

Taming Text-to-Sounding Video Generation via Advanced Modality Condition and Interaction 3D U-Net: Learning Dense Volumetric Segmentation from Sparse Annotation

Reference 2016

Resolution
unresolved
no resolver link, observed 2026-08-04T12:39:00.038859Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T12:39:00.038859Z digest=sha256:b3d087fe40748ce6b39dfe4be0158ac7b72d42766bfceb27d47cff7f4341509e

Observation d25f7a8e-4b31-46dd-bab6-4a479a67d7ec · outbound

This paper cites Wan: Open and Advanced Large-Scale Video Generative Models.

Taming Text-to-Sounding Video Generation via Advanced Modality Condition and Interaction Wan: Open and Advanced Large-Scale Video Generative Models

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-04T12:38:59.279588Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T12:38:59.279588Z digest=sha256:711fee84f0560c501ea11a81f293b011a17d2b0fe688a198ff24bbd31efa1c0c

Observation b2e16779-9699-405d-8958-8f66d33a3c9f · outbound

This paper cites Denoising Diffusion Probabilistic Models.

Taming Text-to-Sounding Video Generation via Advanced Modality Condition and Interaction Denoising Diffusion Probabilistic Models

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-04T12:38:57.138823Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T12:38:57.138823Z digest=sha256:4e84a8e87e9ae9c07db4e96a6d3119e21344ae465044dc898fd9d1be9efc1e1a

Observation 23e273f7-3da5-4cf4-9ee3-4e4cefa5f5fa · outbound

This paper cites Qwen2-Audio Technical Report.

Taming Text-to-Sounding Video Generation via Advanced Modality Condition and Interaction Qwen2-Audio Technical Report

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-04T12:38:56.384224Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T12:38:56.384224Z digest=sha256:4e4426473377a48db0a65afdb9715b382de7857bdcecad00453e499ef91c5dd9

Observation e8144d74-5a43-4bb6-85dd-216833630784 · outbound

This paper cites Fr\'echet Audio Distance: A Metric for Evaluating Music Enhancement Algorithms.

Taming Text-to-Sounding Video Generation via Advanced Modality Condition and Interaction Fr\'echet Audio Distance: A Metric for Evaluating Music Enhancement Algorithms

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-04T12:38:57.438703Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T12:38:57.438703Z digest=sha256:0730ff66f33cbc8ce8d1d83faf8d2e1c6f39afd109eb1fd6c7da95f3c192bcc6

Observation a00d6439-4197-49e2-a6c5-bfd1049acff1 · outbound

This paper cites Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets.

Taming Text-to-Sounding Video Generation via Advanced Modality Condition and Interaction Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-04T12:38:56.194933Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T12:38:56.194933Z digest=sha256:f8a08b135b5d9edb936cc3fe0f6925cb295f7a447ceeaf1f58432a9f98adf616

Observation 11f089b3-86bc-4e12-8093-03ad70f608d8 · outbound

This paper cites TEMPO: Prompt-based Generative Pre-trained Transformer for Time Series Forecasting.

Taming Text-to-Sounding Video Generation via Advanced Modality Condition and Interaction TEMPO: Prompt-based Generative Pre-trained Transformer for Time Series Forecasting

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-04T12:38:56.312366Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T12:38:56.312366Z digest=sha256:57c25906d121e77f940bebe390654f3d6ade3ae8e136d2603f5b78fecafd85b9

Observation 4fc2be78-903a-4988-b3c2-2a40cfc66194 · outbound

This paper cites Classifier-Free Diffusion Guidance.

Taming Text-to-Sounding Video Generation via Advanced Modality Condition and Interaction Classifier-Free Diffusion Guidance

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-04T12:38:56.998200Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T12:38:56.998200Z digest=sha256:143d2284a510f802bdcf708d1e8f267d82dd05d370b2b516b5a2c77eb0ad2304

Pith citing papers

Observation ba269173-b442-4f69-8027-50f31d2a4c67 · inbound

LTX-2: Efficient Joint Audio-Visual Foundation Model cites this paper.

LTX-2: Efficient Joint Audio-Visual Foundation Model Taming Text-to-Sounding Video Generation via Advanced Modality Condition and Interaction

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-06-29T02:14:22.202186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-13T07:06:20.470686Z digest=sha256:0d02169476183329e2589cb30f8b7b0c462701b0fbf39b190e1229726ee0fb1c

Observation 82aa4fc2-5520-4e25-8caa-efde5e885ae3 · inbound

Script-a-Video: Deep Structured Audio-visual Captions via Factorized Streams and Relational Grounding cites this paper.

Script-a-Video: Deep Structured Audio-visual Captions via Factorized Streams and Relational Grounding Taming Text-to-Sounding Video Generation via Advanced Modality Condition and Interaction

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T02:14:22.202186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T15:07:45.595260Z digest=sha256:33c505eda81226655deaea65091230e9fa0cf3f7793cc5acbd2ff9424ed1c92f

Observation 8d553e74-c230-462f-91a2-63a9399b64b8 · inbound

SyncDPO: Enhancing Temporal Synchronization in Video-Audio Joint Generation via Preference Learning cites this paper.

SyncDPO: Enhancing Temporal Synchronization in Video-Audio Joint Generation via Preference Learning Taming Text-to-Sounding Video Generation via Advanced Modality Condition and Interaction

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-06-29T02:14:22.202186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-13T07:01:40.333448Z digest=sha256:86b335c3ec9e26ec5a15e9aca560a87208ae5f054a5c39f87927f5449558dac4

Observation 4f678117-5b37-4ccd-8b31-428359713c1e · inbound

Planar Symmetric Pattern Generation cites this paper.

Planar Symmetric Pattern Generation Taming Text-to-Sounding Video Generation via Advanced Modality Condition and Interaction

Reference 73

Resolution
verified exact
local_arxiv, observed 2026-07-01T22:16:16.090144Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-28T15:37:32.767405Z digest=sha256:8fe0c3235a24f73b6b8e47532788702b3b8bfb982b0244673916df93f495f8a4