Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T06:04:29.971825Z
Paper Citation Record · LEDGER
As of 10 August 2026, this Paper Citation Record lists 36 of 36 outbound references and 6 inbound Pith citation observations for arXiv:2508.00733.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T06:04:29.971825Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-03T16:25:58.872482Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-03T03:27:35.612458Z
36 of 36 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 2b310c81-4c27-40cb-8c4f-8af5af29d927 · outbound
AudioGen-Omni: A Unified Multimodal Diffusion Transformer for Video-Synchronized Audio, Speech, and Song Generation VoxSim: A perceptual voice similarity dataset
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 5f9e9fc9-d300-46c9-b166-d826846c8561 · outbound
AudioGen-Omni: A Unified Multimodal Diffusion Transformer for Video-Synchronized Audio, Speech, and Song Generation Beyond Efficiency: A Systematic Survey of Resource-Efficient Large Language Models
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7e15cb2d-cf4c-47ea-9263-17ef14bee244 · outbound
AudioGen-Omni: A Unified Multimodal Diffusion Transformer for Video-Synchronized Audio, Speech, and Song Generation Vggsound: A large-scale audio-visual dataset
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 866d9cf6-06df-4edb-b9e8-78ed43b07394 · outbound
AudioGen-Omni: A Unified Multimodal Diffusion Transformer for Video-Synchronized Audio, Speech, and Song Generation Clotho: An audio captioning dataset
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 437d8bde-4f5d-4bd7-941d-3210bf9447a2 · outbound
AudioGen-Omni: A Unified Multimodal Diffusion Transformer for Video-Synchronized Audio, Speech, and Song Generation CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 01a1f4a8-eec3-40a9-97e4-402710127371 · outbound
AudioGen-Omni: A Unified Multimodal Diffusion Transformer for Video-Synchronized Audio, Speech, and Song Generation Vid2speech: speech reconstruction from silent video
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation ac1f3415-c111-4879-a6d0-e4066fe59202 · outbound
AudioGen-Omni: A Unified Multimodal Diffusion Transformer for Video-Synchronized Audio, Speech, and Song Generation FunASR: A Fundamental End-to-End Speech Recognition Toolkit
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c7bb091b-edf2-421f-808d-1cf2109f21c1 · outbound
AudioGen-Omni: A Unified Multimodal Diffusion Transformer for Video-Synchronized Audio, Speech, and Song Generation ACE-Step: A Step Towards Music Generation Foundation Model
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b5e1a81a-9666-4ad7-80a5-dabf59f440e6 · outbound
AudioGen-Omni: A Unified Multimodal Diffusion Transformer for Video-Synchronized Audio, Speech, and Song Generation Effect of clustering on the me- chanical properties of sic particulate-reinforced aluminum alloy 2024 metal matrix composites
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation edaf99ee-7cd4-4aae-bc7d-8bb2d0200b9d · outbound
AudioGen-Omni: A Unified Multimodal Diffusion Transformer for Video-Synchronized Audio, Speech, and Song Generation Audiocaps: Generating captions for audios in the wild
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation a8231362-d76c-4f73-a1dc-76c8d930a7af · outbound
AudioGen-Omni: A Unified Multimodal Diffusion Transformer for Video-Synchronized Audio, Speech, and Song Generation Efficient Training of Audio Transformers with Patchout
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fff1aa02-8cbb-415d-b915-9fc0222fc5f8 · outbound
AudioGen-Omni: A Unified Multimodal Diffusion Transformer for Video-Synchronized Audio, Speech, and Song Generation Ntire 2023 challenge on efficient super-resolution: Methods and results
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 7f62678e-05c4-4db8-b0a7-0fd6f90625a6 · outbound
AudioGen-Omni: A Unified Multimodal Diffusion Transformer for Video-Synchronized Audio, Speech, and Song Generation Flow Matching for Generative Modeling
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1d75771a-4415-4798-b175-2d463d71af6d · outbound
AudioGen-Omni: A Unified Multimodal Diffusion Transformer for Video-Synchronized Audio, Speech, and Song Generation AudioLDM: Text-to-Audio Generation with Latent Diffusion Models
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a90c9f78-cc57-45cb-a67f-e79f1a80ff1e · outbound
AudioGen-Omni: A Unified Multimodal Diffusion Transformer for Video-Synchronized Audio, Speech, and Song Generation SVTS: Scalable Video-to-Speech Synthesis
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f586a7f7-b5e1-4e4f-99eb-96e5cb855338 · outbound
AudioGen-Omni: A Unified Multimodal Diffusion Transformer for Video-Synchronized Audio, Speech, and Song Generation Effect of large cold deformation on characteristics of age-strengthening of 2024 aluminum alloys
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation e55a50ed-15c8-4340-baba-00de4d55dffa · outbound
AudioGen-Omni: A Unified Multimodal Diffusion Transformer for Video-Synchronized Audio, Speech, and Song Generation Egosonics: Generating synchronized audio for silent egocentric videos
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation f01b5d7c-4379-4464-bb4f-9c504a24c72e · outbound
AudioGen-Omni: A Unified Multimodal Diffusion Transformer for Video-Synchronized Audio, Speech, and Song Generation Dnsmos: A non-intrusive perceptual objective speech quality metric to evaluate noise suppressors
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 3d927eec-728b-4229-b022-3f745738ec2e · outbound
AudioGen-Omni: A Unified Multimodal Diffusion Transformer for Video-Synchronized Audio, Speech, and Song Generation UTMOS: UTokyo-SaruLab System for VoiceMOS Challenge 2022
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9f73405d-b2d7-4205-bea6-b06ea22e0311 · outbound
AudioGen-Omni: A Unified Multimodal Diffusion Transformer for Video-Synchronized Audio, Speech, and Song Generation DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 07986a81-9e19-4ab0-b216-8451bc314589 · outbound
AudioGen-Omni: A Unified Multimodal Diffusion Transformer for Video-Synchronized Audio, Speech, and Song Generation Generalized end-to-end loss for speaker verification
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 85d6e32a-0eb4-48bd-b14e-ec1ec0be758d · outbound
AudioGen-Omni: A Unified Multimodal Diffusion Transformer for Video-Synchronized Audio, Speech, and Song Generation InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 479b8915-b319-4b6c-9bc6-c948f19904f9 · outbound
AudioGen-Omni: A Unified Multimodal Diffusion Transformer for Video-Synchronized Audio, Speech, and Song Generation Qwen2.5-Omni Technical Report
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3d9ffcc9-19c2-4fb1-9bcb-88c4668eab94 · outbound
AudioGen-Omni: A Unified Multimodal Diffusion Transformer for Video-Synchronized Audio, Speech, and Song Generation Forecasting china’s regional energy demand by 2030: A bayesian approach
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation bc06c7c8-a161-4847-8105-a415e8d1c1ba · outbound
AudioGen-Omni: A Unified Multimodal Diffusion Transformer for Video-Synchronized Audio, Speech, and Song Generation LibriTTS: A Corpus Derived from LibriSpeech for Text-to-Speech
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7e2735b9-945d-4da9-bf83-a8d4ac21bf7a · outbound
AudioGen-Omni: A Unified Multimodal Diffusion Transformer for Video-Synchronized Audio, Speech, and Song Generation FoleyCrafter: Bring Silent Videos to Life with Lifelike and Synchronized Sounds
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0d639cc2-a4a9-48bf-9a87-a1675c0146ce · outbound
AudioGen-Omni: A Unified Multimodal Diffusion Transformer for Video-Synchronized Audio, Speech, and Song Generation Synchformer: Efficient synchronization from sparse cues
Reference 2003
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 1dd81864-a54b-460f-8af5-3e6306a72611 · outbound
AudioGen-Omni: A Unified Multimodal Diffusion Transformer for Video-Synchronized Audio, Speech, and Song Generation Temporally aligned audio for video with autoregression
Reference 2017
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation a8811ba4-f8e9-4233-bc25-30129c684a9b · outbound
AudioGen-Omni: A Unified Multimodal Diffusion Transformer for Video-Synchronized Audio, Speech, and Song Generation Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers
Reference 2018
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cf95400b-5535-4a27-8787-6e8f5fe34cb5 · outbound
AudioGen-Omni: A Unified Multimodal Diffusion Transformer for Video-Synchronized Audio, Speech, and Song Generation DeepAudio-V1:Towards Multi-Modal Multi-Stage End-to-End Video to Speech and Audio Generation
Reference 2019
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a9f25310-eeaf-42b5-a2f9-e18080661d87 · outbound
AudioGen-Omni: A Unified Multimodal Diffusion Transformer for Video-Synchronized Audio, Speech, and Song Generation F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching
Reference 2020
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 30c34803-a6cb-4e46-873f-06920348492f · outbound
AudioGen-Omni: A Unified Multimodal Diffusion Transformer for Video-Synchronized Audio, Speech, and Song Generation AudioGen: Textually Guided Audio Generation
Reference 2021
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 46465fa8-6932-426c-8b48-9bfda381a29b · outbound
AudioGen-Omni: A Unified Multimodal Diffusion Transformer for Video-Synchronized Audio, Speech, and Song Generation FLUX.1 Kontext: Flow Matching for In-Context Image Generation and Editing in Latent Space
Reference 2022
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e37e05cc-f7ac-4ffc-84c7-b3678e1ae25c · outbound
AudioGen-Omni: A Unified Multimodal Diffusion Transformer for Video-Synchronized Audio, Speech, and Song Generation Jukebox: A Generative Model for Music
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c770a0db-f760-470e-a2bc-9da33e913861 · outbound
AudioGen-Omni: A Unified Multimodal Diffusion Transformer for Video-Synchronized Audio, Speech, and Song Generation Seed-TTS: A Family of High-Quality Versatile Speech Generation Models
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cc18a5d3-7b85-4190-9dbd-4c6d356f4faa · outbound
AudioGen-Omni: A Unified Multimodal Diffusion Transformer for Video-Synchronized Audio, Speech, and Song Generation Intelligible Lip-to-Speech Synthesis with Speech Units
Reference 2025
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 5db16c02-bee4-4b43-8764-32f82764cc54 · inbound
JoVA: Unified Multimodal Learning for Joint Video-Audio Generation and Editing AudioGen-Omni: A Unified Multimodal Diffusion Transformer for Video-Synchronized Audio, Speech, and Song Generation
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 92cd37fc-20d2-409e-97b6-218e161ecbb2 · inbound
Omni2Sound: Towards Unified Video-Text-to-Audio Generation AudioGen-Omni: A Unified Multimodal Diffusion Transformer for Video-Synchronized Audio, Speech, and Song Generation
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 3cdf2ce2-884a-4cdd-8752-e50636cef16e · inbound
VidAudio-Bench: Benchmarking V2A and VT2A Generation across Four Audio Categories AudioGen-Omni: A Unified Multimodal Diffusion Transformer for Video-Synchronized Audio, Speech, and Song Generation
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation d4725871-7853-4a8c-b72c-0e3b8675d5eb · inbound
ControlFoley: Unified and Controllable Video-to-Audio Generation with Cross-Modal Conflict Handling AudioGen-Omni: A Unified Multimodal Diffusion Transformer for Video-Synchronized Audio, Speech, and Song Generation
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 3baa02ea-0a50-43c0-be32-347d30bc946a · inbound
Towards Streaming Synchronized Spatial Audio Generation via Autoregressive Diffusion Transformer AudioGen-Omni: A Unified Multimodal Diffusion Transformer for Video-Synchronized Audio, Speech, and Song Generation
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 1ae5565b-83bd-4351-8c3a-fa0453af18b8 · inbound
HoliDubber: Holistic Video Dubbing for Complex Acoustic Scenes via Text-Guided Audio Synthesis AudioGen-Omni: A Unified Multimodal Diffusion Transformer for Video-Synchronized Audio, Speech, and Song Generation
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.