Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T13:07:35.362688Z
Paper Citation Record · LEDGER
As of 9 August 2026, this Paper Citation Record lists 46 of 46 outbound references and 16 inbound Pith citation observations for arXiv:2505.22647.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T13:07:35.362688Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-06T16:39:35.477467Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-04T10:09:44.572116Z
46 of 46 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation e4bf8d19-192d-4813-9711-d75132ca8e69 · outbound
Let Them Talk: Audio-Driven Multi-Person Conversational Video Generation Emo: Emote portrait alive generating expressive portrait videos with audio2video diffusion model under weak conditions
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7ab23656-362f-4d99-bfd5-0ca9b01b2dd1 · outbound
Let Them Talk: Audio-Driven Multi-Person Conversational Video Generation Hallo: Hierarchical Audio-Driven Visual Synthesis for Portrait Image Animation
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 126822ad-cdb2-4732-b1c0-b087d9131350 · outbound
Let Them Talk: Audio-Driven Multi-Person Conversational Video Generation Hallo3: Highly Dynamic and Realistic Portrait Image Animation with Video Diffusion Transformer
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 83457919-cecf-463f-8ee5-730e3cc5059c · outbound
Let Them Talk: Audio-Driven Multi-Person Conversational Video Generation Echomimic: Lifelike audio-driven portrait animations through editable landmark conditions
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 8f6b2128-93cc-4554-990d-7c199ea931a3 · outbound
Let Them Talk: Audio-Driven Multi-Person Conversational Video Generation LatentSync: Taming Audio-Conditioned Latent Diffusion Models for Lip Sync with SyncNet Supervision
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 54b61ee0-15ce-4696-8d2b-805a45979bb4 · outbound
Let Them Talk: Audio-Driven Multi-Person Conversational Video Generation Loopy: Taming Audio-Driven Portrait Avatar with Long-Term Motion Dependency
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 603278bc-f543-4d50-9d0a-4cb48ffa9b45 · outbound
Let Them Talk: Audio-Driven Multi-Person Conversational Video Generation Cyberhost: A one-stage diffusion framework for audio-driven talking body generation
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 146516e4-b405-484e-8be2-5eff66dcbb08 · outbound
Let Them Talk: Audio-Driven Multi-Person Conversational Video Generation OmniHuman-1: Rethinking the Scaling-Up of One-Stage Conditioned Human Animation Models
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f2aa37e3-d4ee-4e5b-8ad2-d4f0164a6dd9 · outbound
Let Them Talk: Audio-Driven Multi-Person Conversational Video Generation EMO2: End-Effector Guided Audio-Driven Avatar Video Generation
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7ed5f4a0-ce47-4f08-8461-6d8df225a658 · outbound
Let Them Talk: Audio-Driven Multi-Person Conversational Video Generation Echomimicv2: Towards striking, simplified, and semi-body human animation.arXiv preprint arXiv:2411.10061, 2024
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 259b03b1-ff70-43b9-b91d-4b75e3601cab · outbound
Let Them Talk: Audio-Driven Multi-Person Conversational Video Generation FantasyTalking: Realistic Talking Portrait Generation via Coherent Motion Synthesis
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1261ea0a-3beb-4122-8082-c182da4cbd1f · outbound
Let Them Talk: Audio-Driven Multi-Person Conversational Video Generation AnimateDiff: Animate Your Personalized Text-to-Image Diffusion Models without Specific Tuning
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e308acb5-2e92-4ad9-938c-4c6047b69b48 · outbound
Let Them Talk: Audio-Driven Multi-Person Conversational Video Generation Diffusion adversarial post-training for one-step video generation.arXiv preprint arXiv:2501.08316, 2025
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7f95a48d-b2d2-458a-a096-8f3ffc45e4d6 · outbound
Let Them Talk: Audio-Driven Multi-Person Conversational Video Generation Wan: Open and Advanced Large-Scale Video Generative Models
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation be6adcbb-8649-49f2-9bb9-d6df34a74a93 · outbound
Let Them Talk: Audio-Driven Multi-Person Conversational Video Generation Stylesync: High-fidelity generalized and personalized lip sync in style-based generator
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation b3e515dd-9aa9-4a8b-af68-e07dbfa7e31a · outbound
Let Them Talk: Audio-Driven Multi-Person Conversational Video Generation Sadtalker: Learning realistic 3d motion coefficients for stylized audio-driven single image talking face animation
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 6276d75d-38de-4d24-9024-a13fab61bf94 · outbound
Let Them Talk: Audio-Driven Multi-Person Conversational Video Generation Videoretalking: Audio-based lip synchronization for talking head video editing in the wild
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 207c8e16-14de-4c8b-a087-5f80423b576c · outbound
Let Them Talk: Audio-Driven Multi-Person Conversational Video Generation Dpe: Disentanglement of pose and expression for general video portrait editing
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation c5908bc9-7adf-42bb-8e7e-a33af17923d8 · outbound
Let Them Talk: Audio-Driven Multi-Person Conversational Video Generation Styleheat: One-shot high-resolution editable talking face generation via pre-trained stylegan
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation f61d0710-f06b-4136-8457-72c43342f32a · outbound
Let Them Talk: Audio-Driven Multi-Person Conversational Video Generation Toontalker: Cross-domain face reenactment
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 74043573-0dc9-4a75-ae0e-87a21f8a5153 · outbound
Let Them Talk: Audio-Driven Multi-Person Conversational Video Generation V-Express: Conditional Dropout for Progressive Training of Portrait Video Generation
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 495fb50a-5248-408a-a27b-c6bf9729b9ae · outbound
Let Them Talk: Audio-Driven Multi-Person Conversational Video Generation Nonlinear 3d face morphable model
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation d9bf26c2-7b55-4e30-b37a-59afa8d46960 · outbound
Let Them Talk: Audio-Driven Multi-Person Conversational Video Generation Audio-driven dubbing for user generated contents via style-aware semi-parametric synthesis.IEEE TCSVT, 33(3):1247– 1261, 2022
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation b146ccef-7509-4d14-9277-f7205658eaec · outbound
Let Them Talk: Audio-Driven Multi-Person Conversational Video Generation AniPortrait: Audio-Driven Synthesis of Photorealistic Portrait Animation
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ed285332-15ad-49eb-8270-340bcb57a288 · outbound
Let Them Talk: Audio-Driven Multi-Person Conversational Video Generation Sonic: Shifting Focus to Global Audio Perception in Portrait Animation
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c09bd768-67e9-454f-ad5b-96dc55584f40 · outbound
Let Them Talk: Audio-Driven Multi-Person Conversational Video Generation High- resolution image synthesis with latent diffusion models
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bc0ffce2-f473-4b86-8b1a-a197b4d7db29 · outbound
Let Them Talk: Audio-Driven Multi-Person Conversational Video Generation SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0ecc0f12-bfbd-4fed-9640-38d0985b3b6c · outbound
Let Them Talk: Audio-Driven Multi-Person Conversational Video Generation Dreambooth: Fine tuning text-to-image diffusion models for subject-driven generation
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 371b3d49-83ea-48ec-92bf-3dd33e1180fd · outbound
Let Them Talk: Audio-Driven Multi-Person Conversational Video Generation Omg: Occlusion-friendly personalized multi-concept generation in diffusion models
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 9670f83a-ce27-43d2-a9b0-b70fee369f18 · outbound
Let Them Talk: Audio-Driven Multi-Person Conversational Video Generation VideoCrafter1: Open Diffusion Models for High-Quality Video Generation
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 45142d58-ff24-4dff-b3ce-ff0fafedebe3 · outbound
Let Them Talk: Audio-Driven Multi-Person Conversational Video Generation Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 99fcc958-4322-4fa7-b2df-68e4f9319c6d · outbound
Let Them Talk: Audio-Driven Multi-Person Conversational Video Generation CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f6fc5fa5-7aa1-4b42-b00d-4e50b51b8941 · outbound
Let Them Talk: Audio-Driven Multi-Person Conversational Video Generation HunyuanVideo: A Systematic Framework For Large Video Generative Models
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d2cfe75d-30fa-4cbe-aa61-1bcdc41eaec0 · outbound
Let Them Talk: Audio-Driven Multi-Person Conversational Video Generation Scalable diffusion models with transformers
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9768261e-b0d3-480e-a0a8-7c803596f44d · outbound
Let Them Talk: Audio-Driven Multi-Person Conversational Video Generation StereoCrafter: Diffusion-based Generation of Long and High-fidelity Stereoscopic 3D from Monocular Videos
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1d053117-d560-4002-9f3d-e633d01bf2db · outbound
Let Them Talk: Audio-Driven Multi-Person Conversational Video Generation StyleMaster: Stylize Your Video with Artistic Generation and Translation
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c0459e79-81c5-434a-9579-81238439c0fa · outbound
Let Them Talk: Audio-Driven Multi-Person Conversational Video Generation Towards multiple character image animation through enhancing implicit decoupling
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 7d1f711e-748c-4da5-af35-5bfa901d2e78 · outbound
Let Them Talk: Audio-Driven Multi-Person Conversational Video Generation Generative Inbetweening: Adapting Image-to-Video Models for Keyframe Interpolation
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d5eba661-413c-4173-a00e-c5026e2133ff · outbound
Let Them Talk: Audio-Driven Multi-Person Conversational Video Generation Learning transferable visual models from natural language supervision
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f41f52a3-af98-4264-9419-6d1c277d33ad · outbound
Let Them Talk: Audio-Driven Multi-Person Conversational Video Generation wav2vec 2.0: A framework for self-supervised learning of speech representations
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation bd1241d6-cccc-42ee-87d2-93771b447322 · outbound
Let Them Talk: Audio-Driven Multi-Person Conversational Video Generation Roformer: Enhanced transformer with rotary position embedding.Neurocomputing, 568:127063, 2024
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d6462992-add8-41c3-810f-1ac97f91f31b · outbound
Let Them Talk: Audio-Driven Multi-Person Conversational Video Generation Flow-guided one-shot talking face generation with a high-resolution audio-visual dataset
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 0d4da94a-4e14-4506-8e7f-7ba252bfc1f2 · outbound
Let Them Talk: Audio-Driven Multi-Person Conversational Video Generation CelebV-HQ: A large-scale video facial attributes dataset
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 322e3111-5c0e-4bfc-b587-f592a7b76e39 · outbound
Let Them Talk: Audio-Driven Multi-Person Conversational Video Generation Gans trained by a two time-scale update rule converge to a local nash equilibrium
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation a7c1b579-1744-4a02-83cc-50eef747966b · outbound
Let Them Talk: Audio-Driven Multi-Person Conversational Video Generation Fvd: A new metric for video generation
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3312cc87-6f2e-4766-a6cb-0a5fca454df8 · outbound
Let Them Talk: Audio-Driven Multi-Person Conversational Video Generation Out of time: automated lip sync in the wild
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation e616d77c-07ab-44ad-94e3-9d28ef19142c · inbound
FantasyPortrait: Enhancing Multi-Character Portrait Animation with Expression-Augmented Diffusion Transformers Let Them Talk: Audio-Driven Multi-Person Conversational Video Generation
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8b35010b-3fba-47ee-be13-7448fe47632f · inbound
FantasyTalking2: Timestep-Layer Adaptive Preference Optimization for Audio-Driven Portrait Animation Let Them Talk: Audio-Driven Multi-Person Conversational Video Generation
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dc40bea0-8f73-4da4-a842-e54029d6a5f6 · inbound
InfiniteTalk: Audio-driven Video Generation for Sparse-Frame Video Dubbing Let Them Talk: Audio-Driven Multi-Person Conversational Video Generation
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b7dc1f40-57f5-44e4-a77f-f664e284d62b · inbound
MIDAS: Multimodal Interactive Digital-humAn Synthesis via Real-time Autoregressive Video Generation Let Them Talk: Audio-Driven Multi-Person Conversational Video Generation
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2b649be8-018a-4de3-b28b-05f9b18454b2 · inbound
InfinityHuman: Towards Long-Term Audio-Driven Human Let Them Talk: Audio-Driven Multi-Person Conversational Video Generation
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8a4458db-a56d-4c30-b58b-9acbacd29f39 · inbound
iMontage: Unified, Versatile, Highly Dynamic Many-to-many Image Generation Let Them Talk: Audio-Driven Multi-Person Conversational Video Generation
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e141d963-46eb-4e0b-9dd4-a4da10d6e3bc · inbound
AUHead: Realistic Emotional Talking Head Generation via Action Units Control Let Them Talk: Audio-Driven Multi-Person Conversational Video Generation
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 8a522e79-6f7f-40c3-9027-d8ac000648e5 · inbound
OmniShow: Unifying Multimodal Conditions for Human-Object Interaction Video Generation Let Them Talk: Audio-Driven Multi-Person Conversational Video Generation
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation f203a03a-a1d4-41f3-8e1a-4be122b0e434 · inbound
TurboTalk: Progressive Distillation for One-Step Audio-Driven Talking Avatar Generation Let Them Talk: Audio-Driven Multi-Person Conversational Video Generation
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation e91fb8d3-1b09-4292-9d6b-8aa21c96918c · inbound
PresentAgent-2: Towards Generalist Multimodal Presentation Agents Let Them Talk: Audio-Driven Multi-Person Conversational Video Generation
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 4465b051-f896-42db-a793-18df0d46d9ca · inbound
Test-Time Self-Adaptive Conditioning for Stable Audio-Driven Talking-Head Generation Let Them Talk: Audio-Driven Multi-Person Conversational Video Generation
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 4373d6ea-fa3d-45f8-8968-bf36cdb2981b · inbound
InteractiveAvatar: Real-Time Streaming Video Generation for Consistent and Intent-Aware Avatars Let Them Talk: Audio-Driven Multi-Person Conversational Video Generation
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 5b418ee3-d410-4682-877f-e5aac23dcbe5 · inbound
InteractiveAvatar: Real-Time Streaming Video Generation for Consistent and Intent-Aware Avatars Let Them Talk: Audio-Driven Multi-Person Conversational Video Generation
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 46353343-5273-4f76-8a0e-0cb7f3a567fe · inbound
Towards Flexible, Natural, Efficient Interaction for Conversational Talking Face Generation Let Them Talk: Audio-Driven Multi-Person Conversational Video Generation
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation fb4228e5-dd58-4905-9692-3d1b0e345dce · inbound
Towards Flexible, Natural, Efficient Interaction for Conversational Talking Face Generation Let Them Talk: Audio-Driven Multi-Person Conversational Video Generation
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e825ffa0-523b-44a7-ac06-98abd718733f · inbound
OmniMate: Open-Ended Real-Time Streaming Audio-Visual Generation for Interactive Avatars Let Them Talk: Audio-Driven Multi-Person Conversational Video Generation
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.