Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-10T19:07:45.221095Z
Paper Citation Record · LEDGER
As of 19 August 2026, this Paper Citation Record lists 52 of 52 outbound references and 18 inbound Pith citation observations for arXiv:2501.10687.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-10T19:07:45.221095Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-15T18:47:48.993817Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-20T15:08:24.974505Z
52 of 52 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation d46a8603-6c07-49b5-a5bc-106dfed95916 · outbound
EMO2: End-Effector Guided Audio-Driven Avatar Video Generation write newline
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8386d5a9-e677-48c2-9ec2-6f4df61936ea · outbound
EMO2: End-Effector Guided Audio-Driven Avatar Video Generation HiT-DVAE: Human Motion Generation via Hierarchical Transformer Dynamical VAE
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 67bedf89-1649-43e9-a78a-80484fddd179 · outbound
EMO2: End-Effector Guided Audio-Driven Avatar Video Generation Animated conversation: rule-based generation of facial expression, gesture & spoken intonation for multiple conversational agents
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8d0d59f4-3c71-449e-a2f9-c54b379dff2b · outbound
EMO2: End-Effector Guided Audio-Driven Avatar Video Generation PixArt-$\alpha$: Fast Training of Diffusion Transformer for Photorealistic Text-to-Image Synthesis
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e9f926d6-563b-44c8-bef6-38164e480f43 · outbound
EMO2: End-Effector Guided Audio-Driven Avatar Video Generation DiffSHEG: A Diffusion-Based Approach for Real-Time Speech-driven Holistic 3D Expression and Gesture Generation
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 228d129c-39ad-4c3b-8d2f-d83ba539746d · outbound
EMO2: End-Effector Guided Audio-Driven Avatar Video Generation Unresolved cited work
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation fbe43582-34c0-4a03-ad4e-35f01f63e6ae · outbound
EMO2: End-Effector Guided Audio-Driven Avatar Video Generation G., Kolotouros, N., Alldieck, T., and Sminchisescu, C
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation e84914e9-4063-40e7-b73d-38cc5cf1f9d1 · outbound
EMO2: End-Effector Guided Audio-Driven Avatar Video Generation Hallo2: Long-duration and high-resolution audio-driven portrait image animation, 2024 a
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation f1d74624-a967-4c73-96e0-e68fc9589b12 · outbound
EMO2: End-Effector Guided Audio-Driven Avatar Video Generation Hallo3: Highly dynamic and realistic portrait image animation with diffusion transformer networks, 2024 b
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 463275a5-827c-46b9-8f82-60d495c0c97e · outbound
EMO2: End-Effector Guided Audio-Driven Avatar Video Generation Looking to Listen at the Cocktail Party: A Speaker-Independent Audio-Visual Model for Speech Separation
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a1c1f1e4-48d8-4edf-962a-488e34fbd9fe · outbound
EMO2: End-Effector Guided Audio-Driven Avatar Video Generation Faceformer: Speech-driven 3d facial animation with transformers
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 9199f9a0-291c-4c1d-acde-7716533e5321 · outbound
EMO2: End-Effector Guided Audio-Driven Avatar Video Generation AnimateDiff: Animate Your Personalized Text-to-Image Diffusion Models without Specific Tuning
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f11fa156-1471-4c4e-ac04-000e416ee358 · outbound
EMO2: End-Effector Guided Audio-Driven Avatar Video Generation Learning Speech-driven 3D Conversational Gestures from Video
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0074fee3-c4c0-46d7-aba1-2efbe560883b · outbound
EMO2: End-Effector Guided Audio-Driven Avatar Video Generation LTX-Video: Realtime Video Latent Diffusion
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4e1c25aa-af23-454c-ad45-253e11896186 · outbound
EMO2: End-Effector Guided Audio-Driven Avatar Video Generation Deep Residual Learning for Image Recognition
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 18ebd840-2307-4262-a2dc-302873d4a2f3 · outbound
EMO2: End-Effector Guided Audio-Driven Avatar Video Generation Co-Speech Gesture Video Generation via Motion-Decoupled Diffusion Model
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 88223648-b9dd-4629-8a27-f99dc26dd8d2 · outbound
EMO2: End-Effector Guided Audio-Driven Avatar Video Generation Gans trained by a two time-scale update rule converge to a local nash equilibrium
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4ec958fc-4bda-44d6-870c-1fc7f41911da · outbound
EMO2: End-Effector Guided Audio-Driven Avatar Video Generation Denoising diffusion probabilistic models
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b951ccd4-c850-4f44-a14e-91003d0ca9bc · outbound
EMO2: End-Effector Guided Audio-Driven Avatar Video Generation Diffted: One-shot audio-driven ted talk video generation with diffusion-based co-speech gestures
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 5937b6f3-11fc-417b-ba9f-bf98de64a44b · outbound
EMO2: End-Effector Guided Audio-Driven Avatar Video Generation and Ziou, D
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 77191f27-01d1-46e8-b4b5-aec3957d9b58 · outbound
EMO2: End-Effector Guided Audio-Driven Avatar Video Generation Animate Anyone: Consistent and Controllable Image-to-Video Synthesis for Character Animation
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 44580a3f-96d6-49fd-a056-12ac5ca1ad4d · outbound
EMO2: End-Effector Guided Audio-Driven Avatar Video Generation Nmpc-mp: Real-time nonlinear model predictive control for safe motion planning in manipulator teleoperation
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f4e2a91a-174d-4779-b697-3027103a9d9e · outbound
EMO2: End-Effector Guided Audio-Driven Avatar Video Generation AdvMT: Adversarial Motion Transformer for Long-term Human Motion Prediction
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6506d9d8-ce50-460e-85ec-ad7e171e272d · outbound
EMO2: End-Effector Guided Audio-Driven Avatar Video Generation Loopy: Taming Audio-Driven Portrait Avatar with Long-Term Motion Dependency
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1849de9f-8928-4482-91a6-580475186483 · outbound
EMO2: End-Effector Guided Audio-Driven Avatar Video Generation Unresolved cited work
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fb2e0638-a41c-4f23-8e1d-64b24ffc7402 · outbound
EMO2: End-Effector Guided Audio-Driven Avatar Video Generation Gesture generation by imitation: from human behavior to computer character animation
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation c03c436e-cf20-42e4-8987-2c72d569920b · outbound
EMO2: End-Effector Guided Audio-Driven Avatar Video Generation HunyuanVideo: A Systematic Framework For Large Video Generative Models
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a1c41a43-2780-4aa2-8759-0477c9e9f580 · outbound
EMO2: End-Effector Guided Audio-Driven Avatar Video Generation CyberHost: Taming Audio-driven Avatar Diffusion Model with Region Codebook Attention
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c82a9412-fc89-4c4b-a45d-0a9c67a0331c · outbound
EMO2: End-Effector Guided Audio-Driven Avatar Video Generation Echomimicv2: Towards striking, simplified, and semi-body human animation, 2024
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2f71b423-7134-46b6-b857-4a00e3d21275 · outbound
EMO2: End-Effector Guided Audio-Driven Avatar Video Generation Unresolved cited work
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 5c10ff8d-e5f5-4378-bc79-8ba0ce10f97c · outbound
EMO2: End-Effector Guided Audio-Driven Avatar Video Generation Scalable Diffusion Models with Transformers
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ebf9c7e3-b4dd-443e-adaf-850fa77ef16c · outbound
EMO2: End-Effector Guided Audio-Driven Avatar Video Generation CoCoGesture: Toward Coherent Co-speech 3D Gesture Generation in the Wild
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 53768143-aadb-48e6-b2d2-f74faa1f65b5 · outbound
EMO2: End-Effector Guided Audio-Driven Avatar Video Generation W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J., et al
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e4439a15-853c-4678-a293-ed74f93ac61f · outbound
EMO2: End-Effector Guided Audio-Driven Avatar Video Generation Relaxedik: Real-time synthesis of accurate and feasible robot arm motion
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 8e331457-34f7-4c0c-9a66-4cdf6328bc5b · outbound
EMO2: End-Effector Guided Audio-Driven Avatar Video Generation High-resolution image synthesis with latent diffusion models
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 505057dc-8778-48c5-bf85-6b9a6c0ed7f7 · outbound
EMO2: End-Effector Guided Audio-Driven Avatar Video Generation Unresolved cited work
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 74578a9c-e32c-4bd7-bd41-94f6772050f6 · outbound
EMO2: End-Effector Guided Audio-Driven Avatar Video Generation wav2vec: Unsupervised pre-training for speech recognition
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 96157ebc-aeee-4c4e-b14a-d6431e19f78b · outbound
EMO2: End-Effector Guided Audio-Driven Avatar Video Generation Vote for grasp poses from noisy point sets by learning from human
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 2f562a17-fefc-400b-a4fc-f4ba3fe11830 · outbound
EMO2: End-Effector Guided Audio-Driven Avatar Video Generation Emo: Emote portrait alive generating expressive portrait videos with audio2video diffusion model under weak conditions
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation f361dbad-56d2-4954-af85-754193358a02 · outbound
EMO2: End-Effector Guided Audio-Driven Avatar Video Generation Fvd: A new metric for video generation
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 7acbd6ea-72f6-4077-b533-af518c20911d · outbound
EMO2: End-Effector Guided Audio-Driven Avatar Video Generation N., Kaiser, L
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3d7b7b86-38c4-4bfa-90be-dca40b5adbe2 · outbound
EMO2: End-Effector Guided Audio-Driven Avatar Video Generation Gesture and speech in interaction: An overview
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1fb83bd0-0a4d-4a93-bfa4-51b0f1839c9e · outbound
EMO2: End-Effector Guided Audio-Driven Avatar Video Generation Image quality assessment: from error visibility to structural similarity
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 27509708-6de1-4049-9fe8-9b3ed70fc39a · outbound
EMO2: End-Effector Guided Audio-Driven Avatar Video Generation Unresolved cited work
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation ec6b8765-5dda-42fc-af05-9e26dc6cfdc6 · outbound
EMO2: End-Effector Guided Audio-Driven Avatar Video Generation CodeTalker: Speech-Driven 3D Facial Animation with Discrete Motion Prior
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9f381e3f-ee27-49df-8c49-31809076feb7 · outbound
EMO2: End-Effector Guided Audio-Driven Avatar Video Generation Hallo: Hierarchical Audio-Driven Visual Synthesis for Portrait Image Animation
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ab07bcb2-4c5d-4596-b37d-d283235d03a1 · outbound
EMO2: End-Effector Guided Audio-Driven Avatar Video Generation CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7d5fb159-6158-441c-be39-308b42c5aa0a · outbound
EMO2: End-Effector Guided Audio-Driven Avatar Video Generation Unresolved cited work
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation a1057e2a-9171-468a-98ff-830d0206817b · outbound
EMO2: End-Effector Guided Audio-Driven Avatar Video Generation B., Liang, P
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6c86fe9d-7707-456b-a7fc-a79696b8ca8b · outbound
EMO2: End-Effector Guided Audio-Driven Avatar Video Generation MimicMotion: High-Quality Human Motion Video Generation with Confidence-aware Pose Guidance
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8df9ff49-93dd-43e5-a63e-7face30e9237 · outbound
EMO2: End-Effector Guided Audio-Driven Avatar Video Generation Taming diffusion models for audio-driven co-speech gesture generation
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation b47c117a-3a55-44a0-a5c1-8043a8a15a17 · outbound
EMO2: End-Effector Guided Audio-Driven Avatar Video Generation Tryondiffusion: A tale of two unets
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 5afc09ef-6566-4fdd-991a-13e24d1e0f52 · inbound
OmniHuman-1: Rethinking the Scaling-Up of One-Stage Conditioned Human Animation Models EMO2: End-Effector Guided Audio-Driven Avatar Video Generation
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3db3c4ac-a43b-4926-b51d-162747b2aa91 · inbound
DanceTogether! Identity-Preserving Multi-Person Interactive Video Generation EMO2: End-Effector Guided Audio-Driven Avatar Video Generation
Reference 78
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f2aa37e3-d4ee-4e5b-8ad2-d4f0164a6dd9 · inbound
Let Them Talk: Audio-Driven Multi-Person Conversational Video Generation EMO2: End-Effector Guided Audio-Driven Avatar Video Generation
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 71ee290a-41c2-40a6-a35c-52be9b948044 · inbound
AlignHuman: Improving Motion and Fidelity via Timestep-Segment Preference Optimization for Audio-Driven Human Animation EMO2: End-Effector Guided Audio-Driven Avatar Video Generation
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 577e9e6f-0a67-4ca5-8604-873d2f065486 · inbound
OmniAvatar: Efficient Audio-Driven Avatar Video Generation with Adaptive Body Animation EMO2: End-Effector Guided Audio-Driven Avatar Video Generation
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation da3f1d5b-9ad4-4a4b-a287-15db42aa1353 · inbound
Bind-Your-Avatar: Multi-Talking-Character Video Generation with Dynamic 3D-mask-based Embedding Router EMO2: End-Effector Guided Audio-Driven Avatar Video Generation
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6edac27d-81d2-4a35-8790-f58beb70f529 · inbound
MoDA: Multi-modal Diffusion Architecture for Talking Head Generation EMO2: End-Effector Guided Audio-Driven Avatar Video Generation
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9bcad11c-f830-4166-bf18-e135b4ea1500 · inbound
StableAvatar: Infinite-Length Audio-Driven Avatar Video Generation EMO2: End-Effector Guided Audio-Driven Avatar Video Generation
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 66570f13-5e7a-4442-9bca-2b84f46e50b1 · inbound
InfiniteTalk: Audio-driven Video Generation for Sparse-Frame Video Dubbing EMO2: End-Effector Guided Audio-Driven Avatar Video Generation
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6897fdfb-9ade-43ff-a776-5b1078b0f010 · inbound
Wan-S2V: Audio-Driven Cinematic Video Generation EMO2: End-Effector Guided Audio-Driven Avatar Video Generation
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b09d4e2d-a50d-4672-8d2f-6686a90fcc50 · inbound
OmniHuman-1.5: Instilling an Active Mind in Avatars via Cognitive Simulation EMO2: End-Effector Guided Audio-Driven Avatar Video Generation
Reference 66
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c3902634-2cbb-4768-91eb-efa0626f97cf · inbound
MIDAS: Multimodal Interactive Digital-humAn Synthesis via Real-time Autoregressive Video Generation EMO2: End-Effector Guided Audio-Driven Avatar Video Generation
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5d297850-8bc0-4a21-9835-3968dbf9b933 · inbound
ControlHair: Synergizing Physics Simulator and Video Diffusion for Controllable Dynamic Hair Rendering EMO2: End-Effector Guided Audio-Driven Avatar Video Generation
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 87f6a9f6-5adb-4993-9e9b-a852732d3937 · inbound
AsymTalker: Identity-Consistent Long-Term Talking Head Generation via Asymmetric Distillation EMO2: End-Effector Guided Audio-Driven Avatar Video Generation
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 05e49444-5004-4e83-a3f8-573441607464 · inbound
AsymTalker: Identity-Consistent Long-Term Talking Head Generation via Asymmetric Distillation EMO2: End-Effector Guided Audio-Driven Avatar Video Generation
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 2935ab38-1ffd-44af-a3e3-0a45cc479da8 · inbound
Image-to-Video Diffusion: From Foundations to Open Frontiers EMO2: End-Effector Guided Audio-Driven Avatar Video Generation
Reference 81
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 10763ea5-837b-46f1-9cae-d67df8a3e3ae · inbound
InstructAV2AV: Instruction-Guided Audio-Video Joint Editing EMO2: End-Effector Guided Audio-Driven Avatar Video Generation
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 74502ed6-8987-428d-8011-e947ddedcbc0 · inbound
EchoCache: Energy-Guided Cross-Modal Caching for Efficient Audio-Driven Video Generation EMO2: End-Effector Guided Audio-Driven Avatar Video Generation
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.