Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T05:24:30.637501Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 71 of 71 outbound references and 1 inbound Pith citation observation for arXiv:2506.08003.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T05:24:30.637501Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-05-20T11:34:32.558440Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-20T11:38:14.727310Z
71 of 71 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 59ac458d-94e1-44f3-8a0d-867e9db24c5e · outbound
Audio-Sync Video Generation with Multi-Stream Temporal Control Diverse and aligned audio-to- video generation via text-to-video model adaptation,
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 6e1fd37e-ec24-4972-92c5-1cf29c734613 · outbound
Audio-Sync Video Generation with Multi-Stream Temporal Control Long video generation with time-agnostic vqgan and time-sensitive transformer,
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 6aa8241b-dfda-4759-b83b-3a53ba12a598 · outbound
Audio-Sync Video Generation with Multi-Stream Temporal Control Seeing and hearing: Open-domain visual-audio generation with diffusion latent aligners,
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 173c70ac-33a1-4244-a17b-a8251bd92aaa · outbound
Audio-Sync Video Generation with Multi-Stream Temporal Control Sound-guided semantic video generation,
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 104fdeee-48a2-4e1b-8463-d7fcf09cdd8f · outbound
Audio-Sync Video Generation with Multi-Stream Temporal Control MM-Diffusion: Learning multi-modal diffusion models for joint audio and video generation,
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation eb9bff16-e028-4a01-bc6a-7395298b8204 · outbound
Audio-Sync Video Generation with Multi-Stream Temporal Control TA2V: Text-audio guided video generation,
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 02f5e58b-3e74-41af-9b4c-8f011ee5de61 · outbound
Audio-Sync Video Generation with Multi-Stream Temporal Control Hallo3: Highly dynamic and realistic portrait image animation with video diffusion transformer,
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5b0f35db-ba58-4d9d-ba5f-624f0791daca · outbound
Audio-Sync Video Generation with Multi-Stream Temporal Control Speech drives templates: Co-speech gesture synthesis with learned templates,
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 9db135e9-97a3-4b6e-9623-a31610f879cf · outbound
Audio-Sync Video Generation with Multi-Stream Temporal Control Visualize music using generative arts,
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b6fa7cb1-0213-4c01-84e5-2fb459e1f6a5 · outbound
Audio-Sync Video Generation with Multi-Stream Temporal Control CogVideox: Text-to-video diffusion models with an expert transformer,
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 53b57413-2458-448f-9c31-6ddfcba84fd7 · outbound
Audio-Sync Video Generation with Multi-Stream Temporal Control Structure and content- guided video synthesis with diffusion models,
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 44cb9eee-3cfe-4393-bca8-532eafaafac4 · outbound
Audio-Sync Video Generation with Multi-Stream Temporal Control Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 90957916-2290-4432-99d0-5096cb5e939f · outbound
Audio-Sync Video Generation with Multi-Stream Temporal Control Latent Video Diffusion Models for High-Fidelity Long Video Generation
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e1db0a67-136b-46e0-98a4-1d8f3ddb1577 · outbound
Audio-Sync Video Generation with Multi-Stream Temporal Control High-resolution image synthesis with latent diffusion models,
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 130049ee-dfac-44ae-9cf8-72b93248efbc · outbound
Audio-Sync Video Generation with Multi-Stream Temporal Control Interpretable 3D human action analysis with temporal convolutional networks,
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 9fcda4e2-2151-4ea0-bb10-0a8fbb1ac67c · outbound
Audio-Sync Video Generation with Multi-Stream Temporal Control Is space-time attention all you need for video understanding?,
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5086d922-64e6-440e-be42-6440bf04d4c8 · outbound
Audio-Sync Video Generation with Multi-Stream Temporal Control U-Net: Convolutional networks for biomedical image segmentation,
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation af590379-9528-40cd-a146-62c41fa36b46 · outbound
Audio-Sync Video Generation with Multi-Stream Temporal Control Sora: A Review on Background, Technology, Limitations, and Opportunities of Large Vision Models
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4cb69949-eeaa-4b83-976e-1315e82e3419 · outbound
Audio-Sync Video Generation with Multi-Stream Temporal Control Scalable diffusion models with transformers,
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f185d2ab-9d88-4072-bcac-028ccd651926 · outbound
Audio-Sync Video Generation with Multi-Stream Temporal Control Language model beats diffusion – tokenizer is key to visual generation,
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 74cacca0-99d8-497b-a584-3fae50b1c88e · outbound
Audio-Sync Video Generation with Multi-Stream Temporal Control Wan: Open and Advanced Large-Scale Video Generative Models
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c90e47a0-d7d3-4a13-b202-f45f984da477 · outbound
Audio-Sync Video Generation with Multi-Stream Temporal Control HunyuanVideo: A Systematic Framework For Large Video Generative Models
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1c9c5c64-36ce-4a6d-9ade-88715a693c0a · outbound
Audio-Sync Video Generation with Multi-Stream Temporal Control Step-Video-T2V Technical Report: The Practice, Challenges, and Future of Video Foundation Model
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9386c427-f05d-4b6c-9596-cf0abfce3c4f · outbound
Audio-Sync Video Generation with Multi-Stream Temporal Control Sound2Sight: Generating visual dynamics from sound and context,
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 482fc39c-24fe-4b38-896d-706d3e01cbd0 · outbound
Audio-Sync Video Generation with Multi-Stream Temporal Control CCVS: Context-aware controllable video synthesis,
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 31c2da8d-43a4-4bf0-a56a-2187c03388e3 · outbound
Audio-Sync Video Generation with Multi-Stream Temporal Control The power of sound (TPoS): Audio reactive video generation with stable diffusion,
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0da452e7-443d-444a-b16d-2020746276fd · outbound
Audio-Sync Video Generation with Multi-Stream Temporal Control Audio-synchronized visual animation,
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d141dff9-2d35-4719-a13c-b94753f9f633 · outbound
Audio-Sync Video Generation with Multi-Stream Temporal Control Loopy: Taming audio-driven portrait avatar with long-term motion dependency,
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c3a1929b-d64c-4a7e-b2a7-33342cbf61a2 · outbound
Audio-Sync Video Generation with Multi-Stream Temporal Control AniPortrait: Audio-Driven Synthesis of Photorealistic Portrait Animation
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7fe47a68-6801-4dd5-bcc5-7f8468080216 · outbound
Audio-Sync Video Generation with Multi-Stream Temporal Control CyberHost: A one-stage diffusion framework for audio-driven talking body generation,
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 03919a94-3a27-4efb-87cc-cd6bf72d9b33 · outbound
Audio-Sync Video Generation with Multi-Stream Temporal Control OmniHuman-1: Rethinking the Scaling-Up of One-Stage Conditioned Human Animation Models
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 97c70973-2b49-4825-8c6e-fa7e72169ac0 · outbound
Audio-Sync Video Generation with Multi-Stream Temporal Control Dance any beat: Blending beats with visuals in dance video generation,
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c4d9c21c-b91b-4105-94ef-c6ca6d463b0f · outbound
Audio-Sync Video Generation with Multi-Stream Temporal Control X-Dancer: Expressive Music to Human Dance Video Generation
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7a10a7d8-b6b4-48a6-9fbc-5f009fdd31da · outbound
Audio-Sync Video Generation with Multi-Stream Temporal Control Taming transformers for high-resolution image synthesis,
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation dd210d7b-cf20-478e-bc0d-459bbdea9649 · outbound
Audio-Sync Video Generation with Multi-Stream Temporal Control A style-based generator architecture for generative adversarial networks,
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0292bced-28bf-4495-94c8-8a5709d65e6f · outbound
Audio-Sync Video Generation with Multi-Stream Temporal Control Tr\"aumerAI: Dreaming Music with StyleGAN
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0a69a0f0-b5f1-4d9e-8971-7ff74ac49f2d · outbound
Audio-Sync Video Generation with Multi-Stream Temporal Control UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c80b6ac1-61f6-41c0-b13c-5c9b8137ce49 · outbound
Audio-Sync Video Generation with Multi-Stream Temporal Control Content-Based Video-Music Retrieval Using Soft Intra-Modal Structure Constraint
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a8e9d9d0-2234-436d-b21d-f7333d88753a · outbound
Audio-Sync Video Generation with Multi-Stream Temporal Control AudioSet: An ontology and human-labeled dataset for audio events,
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d1a3ecbc-0034-46f2-bdcb-4a9efbaef075 · outbound
Audio-Sync Video Generation with Multi-Stream Temporal Control VoxCeleb2: Deep Speaker Recognition
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7b68e6ec-ae47-4859-815e-e0e7d579afca · outbound
Audio-Sync Video Generation with Multi-Stream Temporal Control VggSound: A large-scale audio-visual dataset,
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 6bfa8e3d-ce63-4a43-983c-962631ef68ff · outbound
Audio-Sync Video Generation with Multi-Stream Temporal Control Frozen in time: A joint video and image encoder for end-to-end retrieval,
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 32fdc992-0e22-406b-ab5b-a58d40724d58 · outbound
Audio-Sync Video Generation with Multi-Stream Temporal Control InternVid: A large-scale video-text dataset for multimodal understanding and generation,
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a7f64bdf-cd4f-43d4-848f-6e33548caf24 · outbound
Audio-Sync Video Generation with Multi-Stream Temporal Control CelebV-HQ: A large-scale video facial attributes dataset,
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 02f33fb4-2f3b-4c8c-8b03-06a0ea14503c · outbound
Audio-Sync Video Generation with Multi-Stream Temporal Control MovieBench: A Hierarchical Movie Level Dataset for Long Video Generation
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 87cd5216-68e7-4058-b722-039abb8eaf52 · outbound
Audio-Sync Video Generation with Multi-Stream Temporal Control Condensed movies: Story based retrieval with contextual embeddings,
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 8e73c803-64ed-40bb-a691-b58d78a20ef3 · outbound
Audio-Sync Video Generation with Multi-Stream Temporal Control Long Story Short: Story-level Video Understanding from 20K Short Films
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dcc8e9f4-e5a2-46f9-8abb-7876849b57ea · outbound
Audio-Sync Video Generation with Multi-Stream Temporal Control VideoCrafter2: Over- coming data limitations for high-quality video diffusion models,
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d62e1bd4-120d-4c6d-9a71-f055e92dacbb · outbound
Audio-Sync Video Generation with Multi-Stream Temporal Control Video cut detection and analysis tool
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b8fdad6b-7583-4436-a9d0-30ad09f57a16 · outbound
Audio-Sync Video Generation with Multi-Stream Temporal Control Meta Audiobox Aesthetics: Unified Automatic Quality Assessment for Speech, Music, and Sound
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5a25c88f-a587-46d3-8cd0-b7f0467bacc4 · outbound
Audio-Sync Video Generation with Multi-Stream Temporal Control LLaVA-Video: Video Instruction Tuning With Synthetic Data
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2fea16d8-7016-45c3-b221-8c19d508a5a2 · outbound
Audio-Sync Video Generation with Multi-Stream Temporal Control Cinematic sound demixing
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d017bcb6-ff74-45fb-8a8a-c1960fa909e5 · outbound
Audio-Sync Video Generation with Multi-Stream Temporal Control Spleeter: a fast and efficient music source separation tool with pre-trained models,
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 11c5ce5e-efa2-41ea-8d26-e62ee7767c38 · outbound
Audio-Sync Video Generation with Multi-Stream Temporal Control Ultralytics YOLO
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation fe301372-756c-4a2f-905c-8cc24c4c07ce · outbound
Audio-Sync Video Generation with Multi-Stream Temporal Control Meet scribe
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c24d23f4-aca3-4d3f-a2ef-5d519b718316 · outbound
Audio-Sync Video Generation with Multi-Stream Temporal Control TalkNet 2: Non-Autoregressive Depth-Wise Separable Convolutional Model for Speech Synthesis with Explicit Pitch and Duration Prediction
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0505b90b-5a1f-4c55-9af1-34f06ae77fe3 · outbound
Audio-Sync Video Generation with Multi-Stream Temporal Control wav2vec 2.0: a framework for self-supervised learning of speech representations,
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation de1ebc8c-a82f-4075-9e3f-21ffa472914a · outbound
Audio-Sync Video Generation with Multi-Stream Temporal Control Adam: A Method for Stochastic Optimization
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7d5800e6-a8a9-4933-8e39-85d99e5040c8 · outbound
Audio-Sync Video Generation with Multi-Stream Temporal Control Towards Accurate Generative Models of Video: A New Metric & Challenges
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 861757ed-d16a-4659-8865-6d48a251e24f · outbound
Audio-Sync Video Generation with Multi-Stream Temporal Control Learning transferable visual models from natural language supervi- sion,
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation ccb49950-47e0-4796-a667-e079589500d6 · outbound
Audio-Sync Video Generation with Multi-Stream Temporal Control VideoCLIP-XL: Advancing Long Description Understanding for Video CLIP Models
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fbd91222-ad01-4fd6-9b32-3c82239582fe · outbound
Audio-Sync Video Generation with Multi-Stream Temporal Control ImageBind: One embedding space to bind them all,
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 646ecfe8-29ef-4aba-af15-a3ad89219233 · outbound
Audio-Sync Video Generation with Multi-Stream Temporal Control Out of time: automated lip sync in the wild,
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b9b7134a-a243-48a1-95e5-66c7ac1cf334 · outbound
Audio-Sync Video Generation with Multi-Stream Temporal Control DiTCtrl: Exploring Attention Control in Multi-Modal Diffusion Transformer for Tuning-Free Multi-Prompt Longer Video Generation
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2f578fe6-ae65-46dc-8359-c8063755c4de · outbound
Audio-Sync Video Generation with Multi-Stream Temporal Control TA VGBench: Benchmarking text to audible-video generation,
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 1e644474-6768-4be7-b191-5aef89235a09 · outbound
Audio-Sync Video Generation with Multi-Stream Temporal Control MMDisco: Multi-modal discriminator- guided cooperative diffusion for joint audio and video generation,
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d3aa9b78-8bee-4c6f-a667-ef4b871127dc · outbound
Audio-Sync Video Generation with Multi-Stream Temporal Control AV-DiT: Efficient Audio-Visual Diffusion Transformer for Joint Audio and Video Generation
Reference 68
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c59b8c40-6eb5-4122-870f-29532280bded · outbound
Audio-Sync Video Generation with Multi-Stream Temporal Control Hallo2: Long-duration and high-resolution audio-driven portrait image animation,
Reference 69
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 21c9b3e6-2ec4-4247-8fa4-8220abff7b6f · outbound
Audio-Sync Video Generation with Multi-Stream Temporal Control SadTalker: Learning realistic 3d motion coefficients for stylized audio-driven single image talking face animation,
Reference 70
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e04b7e44-6576-4e27-a303-72d88c1d9e66 · outbound
Audio-Sync Video Generation with Multi-Stream Temporal Control Maximum filter vibrato suppression for onset detection,
Reference 71
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 533130be-c7b5-4938-a734-f57842d3dd34 · outbound
Audio-Sync Video Generation with Multi-Stream Temporal Control Determining optical flow,
Reference 72
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation dea592b8-9aa9-4846-b81b-926220787e57 · inbound
InstructAV2AV: Instruction-Guided Audio-Video Joint Editing Audio-Sync Video Generation with Multi-Stream Temporal Control
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.