Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T23:05:27.453564Z
Paper Citation Record · LEDGER
As of 16 August 2026, this Paper Citation Record lists 19 of 19 outbound references and 10 inbound Pith citation observations for arXiv:2412.15220.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T23:05:27.453564Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-15T15:07:18.202024Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-02T05:56:40.931162Z
19 of 19 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation d715897c-cb03-4450-84b5-333207ec8dfa · outbound
SyncFlow: Toward Temporally Aligned Joint Audio-Video Generation from Text MusicLM: Generating Music From Text
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation abb54f07-cb9b-4d0f-aa35-d6e128898036 · outbound
SyncFlow: Toward Temporally Aligned Joint Audio-Video Generation from Text Imagen Video: High Definition Video Generation with Diffusion Models
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f5540ced-a70e-4c15-89a0-b2b79ce3f1d9 · outbound
SyncFlow: Toward Temporally Aligned Joint Audio-Video Generation from Text Make-An-Audio 2: Temporal-Enhanced Text-to-Audio Generation
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a294bb5e-6dbd-4118-9daa-0d89e0471be6 · outbound
SyncFlow: Toward Temporally Aligned Joint Audio-Video Generation from Text FoleyGen: Visually-Guided Audio Generation
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0d3da6c7-1d27-4941-8bd5-a2b01f48b431 · outbound
SyncFlow: Toward Temporally Aligned Joint Audio-Video Generation from Text Text-to-Audio Generation Synchronized with Videos
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e45ea5c1-09fa-4992-a19d-e2928f636e11 · outbound
SyncFlow: Toward Temporally Aligned Joint Audio-Video Generation from Text Hierarchical Text-Conditional Image Generation with CLIP Latents
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ca96df05-847d-4ac9-b96a-5c33e1c50ce4 · outbound
SyncFlow: Toward Temporally Aligned Joint Audio-Video Generation from Text Make-A-Video: Text-to-Video Generation without Text-Video Data
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 513b7a70-aa2e-4386-9097-a4c11b164d9a · outbound
SyncFlow: Toward Temporally Aligned Joint Audio-Video Generation from Text NaturalSpeech: End-to-End Text to Speech Synthesis with Human-Level Quality
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 187f45ca-e2c5-4f3d-ad1f-92ef1785f6a7 · outbound
SyncFlow: Toward Temporally Aligned Joint Audio-Video Generation from Text Diffusion Models Are Real-Time Game Engines
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 16951b98-ed79-4e0b-884f-4bed32300366 · outbound
SyncFlow: Toward Temporally Aligned Joint Audio-Video Generation from Text Audiobox: Unified Audio Generation with Natural Language Prompts
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1398e0fc-7d6e-480d-9766-a9ac6e4bbea7 · outbound
SyncFlow: Toward Temporally Aligned Joint Audio-Video Generation from Text AV-DiT: Efficient Audio-Visual Diffusion Transformer for Joint Audio and Video Generation
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1b4b0a64-75a7-4fc7-b916-7aa38610716d · outbound
SyncFlow: Toward Temporally Aligned Joint Audio-Video Generation from Text CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation de8052fc-ceeb-4a62-a3fc-39411cc372b1 · outbound
SyncFlow: Toward Temporally Aligned Joint Audio-Video Generation from Text FlowSep: Language-Queried Sound Separation with Rectified Flow Matching
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation faad95cb-ed2f-4b8d-8d8c-e712d49709e3 · outbound
SyncFlow: Toward Temporally Aligned Joint Audio-Video Generation from Text Semantically consistent Video-to-Audio Generation using Multimodal Language Large Model
Reference 2017
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0f8d7c8f-0779-4c12-b2e9-e5a4d6623cd4 · outbound
SyncFlow: Toward Temporally Aligned Joint Audio-Video Generation from Text VideoOFA: Two-Stage Pre-Training for Video-to-Text Generation
Reference 2020
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 350d20ca-4483-415d-a886-5c9fe983c1ac · outbound
SyncFlow: Toward Temporally Aligned Joint Audio-Video Generation from Text Synctalkface: Talking face generation with precise lip-syncing via audio-lip memory
Reference 2021
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation a22d6344-4e16-45de-85ab-90bfdaa21930 · outbound
SyncFlow: Toward Temporally Aligned Joint Audio-Video Generation from Text High Fidelity Text-Guided Music Editing via Single-Stage Flow Matching
Reference 2022
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2defa89d-753a-45e7-a028-b0f7c1ae5177 · outbound
SyncFlow: Toward Temporally Aligned Joint Audio-Video Generation from Text Simple and Controllable Music Generation
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6cdbb547-8170-43c3-9e1e-9e21ee381a27 · outbound
SyncFlow: Toward Temporally Aligned Joint Audio-Video Generation from Text MMDisCo: Multi-Modal Discriminator-Guided Cooperative Diffusion for Joint Audio and Video Generation
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dea83d22-1d9d-49da-ac40-e2eabc4cca0e · inbound
UniVerse-1: Unified Audio-Video Generation via Stitching of Experts SyncFlow: Toward Temporally Aligned Joint Audio-Video Generation from Text
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 42b6baa9-358b-4bba-97f5-d90724595e5d · inbound
Taming Text-to-Sounding Video Generation via Advanced Modality Condition and Interaction SyncFlow: Toward Temporally Aligned Joint Audio-Video Generation from Text
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d7a23a1c-972f-4725-8ad4-3706843219ac · inbound
VideoASMR-Bench: Can AI-Generated ASMR Videos Fool VLMs and Humans? SyncFlow: Toward Temporally Aligned Joint Audio-Video Generation from Text
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation cfb1d7bd-7733-41fe-aa49-4ab177adeeae · inbound
PhyAVBench: A Challenging Audio Physics-Sensitivity Benchmark for Physically Grounded Text-to-Audio-Video Generation SyncFlow: Toward Temporally Aligned Joint Audio-Video Generation from Text
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 2d635f72-789b-474c-9e74-2db868550dc6 · inbound
PhyAVBench: A Challenging Audio Physics-Sensitivity Benchmark for Physically Grounded Text-to-Audio-Video Generation SyncFlow: Toward Temporally Aligned Joint Audio-Video Generation from Text
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation e0f11733-31b9-4529-adee-38832a8e15ee · inbound
Talker-T2AV: Joint Talking Audio-Video Generation with Autoregressive Diffusion Modeling SyncFlow: Toward Temporally Aligned Joint Audio-Video Generation from Text
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 1574171a-d19c-4fb8-a268-7320017c7ec7 · inbound
Unison: Harmonizing Motion, Speech, and Sound for Human-Centric Audio-Video Generation SyncFlow: Toward Temporally Aligned Joint Audio-Video Generation from Text
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation f4d24b73-9292-4f2d-87f7-8cf8abc57f52 · inbound
Unison: Harmonizing Motion, Speech, and Sound for Human-Centric Audio-Video Generation SyncFlow: Toward Temporally Aligned Joint Audio-Video Generation from Text
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 01c6ea4a-f099-458b-8219-85773421ecd2 · inbound
Inference-Time Scaling for Joint Audio-Video Generation SyncFlow: Toward Temporally Aligned Joint Audio-Video Generation from Text
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 00a524b9-2b2f-4888-88d2-c28c0e276ec6 · inbound
AcoustiTrace: When Plausible Sound Violates Physics SyncFlow: Toward Temporally Aligned Joint Audio-Video Generation from Text
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.