Pith. sign in

Paper Citation Record · LEDGER

Let Them Talk: Audio-Driven Multi-Person Conversational Video Generation

As of 15 August 2026, this Paper Citation Record lists 46 of 46 outbound references and 20 inbound Pith citation observations for arXiv:2505.22647.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.22647 v1

Coverage vector

measured 46 of 46 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:07:35.362688Z

measured 66 of 66 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 20 of 20 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T18:47:48.965076Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T10:09:44.572116Z

Reference resolution

46 of 46 outbound references displayed

  • verified exact0
  • verified fuzzy17
  • unresolved29
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e4bf8d19-192d-4813-9711-d75132ca8e69 · outbound

This paper cites Emo: Emote portrait alive generating expressive portrait videos with audio2video diffusion model under weak conditions.

Let Them Talk: Audio-Driven Multi-Person Conversational Video Generation Emo: Emote portrait alive generating expressive portrait videos with audio2video diffusion model under weak conditions

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T13:07:31.369913Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:07:31.369913Z digest=sha256:57007617187112e4e92ea26c35e2375936259be84be1c133ca6d87ad7652b7a6

Observation 7ab23656-362f-4d99-bfd5-0ca9b01b2dd1 · outbound

This paper cites Hallo: Hierarchical Audio-Driven Visual Synthesis for Portrait Image Animation.

Let Them Talk: Audio-Driven Multi-Person Conversational Video Generation Hallo: Hierarchical Audio-Driven Visual Synthesis for Portrait Image Animation

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T13:07:31.440529Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:07:31.440529Z digest=sha256:7e3fd50e2a1e288b6e33b82819105438eefe2ad86c2a34a450a6363bd1679377

Observation 126822ad-cdb2-4732-b1c0-b087d9131350 · outbound

This paper cites Hallo3: Highly Dynamic and Realistic Portrait Image Animation with Video Diffusion Transformer.

Let Them Talk: Audio-Driven Multi-Person Conversational Video Generation Hallo3: Highly Dynamic and Realistic Portrait Image Animation with Video Diffusion Transformer

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T13:07:31.596042Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:07:31.596042Z digest=sha256:367a0b8bc1f8fb243c12baa71dd6d5c7bddafeea2bf27c2bb40d612f95079788

Observation 83457919-cecf-463f-8ee5-730e3cc5059c · outbound

This paper cites Echomimic: Lifelike audio-driven portrait animations through editable landmark conditions.

Let Them Talk: Audio-Driven Multi-Person Conversational Video Generation Echomimic: Lifelike audio-driven portrait animations through editable landmark conditions

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:07:38.450148Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T13:07:31.759871Z digest=sha256:e323d9a3c5e685a06ba1d907844b59427b9c5e339b594a4f513f03204a5964cd

Observation 8f6b2128-93cc-4554-990d-7c199ea931a3 · outbound

This paper cites LatentSync: Taming Audio-Conditioned Latent Diffusion Models for Lip Sync with SyncNet Supervision.

Let Them Talk: Audio-Driven Multi-Person Conversational Video Generation LatentSync: Taming Audio-Conditioned Latent Diffusion Models for Lip Sync with SyncNet Supervision

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T13:07:31.880403Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:07:31.880403Z digest=sha256:b8f846f326e1ca6321a9f00bd8c9ccb0cd2f62ec1eba70f87e78d1b767f4837d

Observation 54b61ee0-15ce-4696-8d2b-805a45979bb4 · outbound

This paper cites Loopy: Taming Audio-Driven Portrait Avatar with Long-Term Motion Dependency.

Let Them Talk: Audio-Driven Multi-Person Conversational Video Generation Loopy: Taming Audio-Driven Portrait Avatar with Long-Term Motion Dependency

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T13:07:31.988771Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:07:31.988771Z digest=sha256:5a33d18dd5924960bf7f2e0290f99f95b9e4f934d57990e36e7a979b75ac59e6

Observation 603278bc-f543-4d50-9d0a-4cb48ffa9b45 · outbound

This paper cites Cyberhost: A one-stage diffusion framework for audio-driven talking body generation.

Let Them Talk: Audio-Driven Multi-Person Conversational Video Generation Cyberhost: A one-stage diffusion framework for audio-driven talking body generation

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:07:38.322693Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T13:07:32.073177Z digest=sha256:c549fc29a2e378eb7d3d821eef22e9fc322c8677d84088c3d04dc91b809dd581

Observation 146516e4-b405-484e-8be2-5eff66dcbb08 · outbound

This paper cites OmniHuman-1: Rethinking the Scaling-Up of One-Stage Conditioned Human Animation Models.

Let Them Talk: Audio-Driven Multi-Person Conversational Video Generation OmniHuman-1: Rethinking the Scaling-Up of One-Stage Conditioned Human Animation Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T13:07:32.220659Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:07:32.220659Z digest=sha256:d4e5811d76d3f04b0e9606ffb79138effcf98dc0d03d07cea754f2d8ffcd0f32

Observation f2aa37e3-d4ee-4e5b-8ad2-d4f0164a6dd9 · outbound

This paper cites EMO2: End-Effector Guided Audio-Driven Avatar Video Generation.

Let Them Talk: Audio-Driven Multi-Person Conversational Video Generation EMO2: End-Effector Guided Audio-Driven Avatar Video Generation

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T13:07:32.372279Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:07:32.372279Z digest=sha256:e863d90bc0e763d87a1f3fe53b4badc43bc0ff1416f917b69e5c60573463adb5

Observation 7ed5f4a0-ce47-4f08-8461-6d8df225a658 · outbound

This paper cites Echomimicv2: Towards striking, simplified, and semi-body human animation.arXiv preprint arXiv:2411.10061, 2024.

Let Them Talk: Audio-Driven Multi-Person Conversational Video Generation Echomimicv2: Towards striking, simplified, and semi-body human animation.arXiv preprint arXiv:2411.10061, 2024

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T13:07:32.506861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:07:32.506861Z digest=sha256:f2ef622b0c7b0d435431b59dd403e544c05c52a463b306edc435144604758e68

Observation 259b03b1-ff70-43b9-b91d-4b75e3601cab · outbound

This paper cites FantasyTalking: Realistic Talking Portrait Generation via Coherent Motion Synthesis.

Let Them Talk: Audio-Driven Multi-Person Conversational Video Generation FantasyTalking: Realistic Talking Portrait Generation via Coherent Motion Synthesis

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T13:07:32.607328Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:07:32.607328Z digest=sha256:9e1fe5965cfb7c5d605235e85966ac1ca38852b49f0a30d70b302ebf1cffa46f

Observation 1261ea0a-3beb-4122-8082-c182da4cbd1f · outbound

This paper cites AnimateDiff: Animate Your Personalized Text-to-Image Diffusion Models without Specific Tuning.

Let Them Talk: Audio-Driven Multi-Person Conversational Video Generation AnimateDiff: Animate Your Personalized Text-to-Image Diffusion Models without Specific Tuning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T13:07:32.720783Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:07:32.720783Z digest=sha256:e3027cd4d6cf03a7e1a9a1a36a0a6b9533379f099c51d152593f08100ac0b06f

Observation e308acb5-2e92-4ad9-938c-4c6047b69b48 · outbound

This paper cites Diffusion adversarial post-training for one-step video generation.arXiv preprint arXiv:2501.08316, 2025.

Let Them Talk: Audio-Driven Multi-Person Conversational Video Generation Diffusion adversarial post-training for one-step video generation.arXiv preprint arXiv:2501.08316, 2025

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T13:07:32.854272Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:07:32.854272Z digest=sha256:9c96ff89aa5524eacdd6187d45222ec9b8fb9b55a1ca6815bd2d44c9c47341a3

Observation 7f95a48d-b2d2-458a-a096-8f3ffc45e4d6 · outbound

This paper cites Wan: Open and Advanced Large-Scale Video Generative Models.

Let Them Talk: Audio-Driven Multi-Person Conversational Video Generation Wan: Open and Advanced Large-Scale Video Generative Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T13:07:33.012708Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:07:33.012708Z digest=sha256:f4be2853ac471ed18e9d34d63ae16d00717fd25dba07f503a1978f90be17f32b

Observation be6adcbb-8649-49f2-9bb9-d6df34a74a93 · outbound

This paper cites Stylesync: High-fidelity generalized and personalized lip sync in style-based generator.

Let Them Talk: Audio-Driven Multi-Person Conversational Video Generation Stylesync: High-fidelity generalized and personalized lip sync in style-based generator

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:07:38.151957Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T13:07:33.101465Z digest=sha256:a8f0bba83a54a705a061796c9b06d0a223923257c8f774a2c74aad196b864596

Observation b3e515dd-9aa9-4a8b-af68-e07dbfa7e31a · outbound

This paper cites Sadtalker: Learning realistic 3d motion coefficients for stylized audio-driven single image talking face animation.

Let Them Talk: Audio-Driven Multi-Person Conversational Video Generation Sadtalker: Learning realistic 3d motion coefficients for stylized audio-driven single image talking face animation

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:07:37.978666Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T13:07:33.182183Z digest=sha256:29c73ea9bb5214723a341c9b372c2fbaf6506acf471666e4198f98bf9dccbc9f

Observation 6276d75d-38de-4d24-9024-a13fab61bf94 · outbound

This paper cites Videoretalking: Audio-based lip synchronization for talking head video editing in the wild.

Let Them Talk: Audio-Driven Multi-Person Conversational Video Generation Videoretalking: Audio-based lip synchronization for talking head video editing in the wild

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:07:37.812816Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T13:07:33.306501Z digest=sha256:50576040f799f0e9e43f9779817e679418d4ea49f63c4e59f8723b6c21806a56

Observation 207c8e16-14de-4c8b-a087-5f80423b576c · outbound

This paper cites Dpe: Disentanglement of pose and expression for general video portrait editing.

Let Them Talk: Audio-Driven Multi-Person Conversational Video Generation Dpe: Disentanglement of pose and expression for general video portrait editing

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:07:37.623316Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T13:07:33.386221Z digest=sha256:6a14de5f3c137e939235ab420363f60b79bd962e5d6b63ee60b1060b0ca492ce

Observation c5908bc9-7adf-42bb-8e7e-a33af17923d8 · outbound

This paper cites Styleheat: One-shot high-resolution editable talking face generation via pre-trained stylegan.

Let Them Talk: Audio-Driven Multi-Person Conversational Video Generation Styleheat: One-shot high-resolution editable talking face generation via pre-trained stylegan

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:07:37.496222Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T13:07:33.450158Z digest=sha256:c15e4db38ada4bf02836d552e4b783eb93efcfac8f6bfa800f4389f3684a9816

Observation f61d0710-f06b-4136-8457-72c43342f32a · outbound

This paper cites Toontalker: Cross-domain face reenactment.

Let Them Talk: Audio-Driven Multi-Person Conversational Video Generation Toontalker: Cross-domain face reenactment

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:07:37.265272Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T13:07:33.519967Z digest=sha256:609951b797b6fe0fafcd90b1e7172486ad06e8b71f6ee8c0fe2205bdcc42ca51

Observation 74043573-0dc9-4a75-ae0e-87a21f8a5153 · outbound

This paper cites V-Express: Conditional Dropout for Progressive Training of Portrait Video Generation.

Let Them Talk: Audio-Driven Multi-Person Conversational Video Generation V-Express: Conditional Dropout for Progressive Training of Portrait Video Generation

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T13:07:33.665330Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:07:33.665330Z digest=sha256:32481cf2d637730d0c1e1c754be972f0e045e85f99f2a6fcab6d774a7f482976

Observation 495fb50a-5248-408a-a27b-c6bf9729b9ae · outbound

This paper cites Nonlinear 3d face morphable model.

Let Them Talk: Audio-Driven Multi-Person Conversational Video Generation Nonlinear 3d face morphable model

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:07:37.138422Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T13:07:33.768430Z digest=sha256:3bf7bddbf31245b447fc489908973e3d496e5059d37bd161513ca725e26ecbcc

Observation d9bf26c2-7b55-4e30-b37a-59afa8d46960 · outbound

This paper cites Audio-driven dubbing for user generated contents via style-aware semi-parametric synthesis.IEEE TCSVT, 33(3):1247– 1261, 2022.

Let Them Talk: Audio-Driven Multi-Person Conversational Video Generation Audio-driven dubbing for user generated contents via style-aware semi-parametric synthesis.IEEE TCSVT, 33(3):1247– 1261, 2022

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:07:36.972021Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T13:07:33.896039Z digest=sha256:be22a1f6913751211b554edb6a66d73aa4a17e7a6dce309ef1ce38d3cd6ed69f

Observation b146ccef-7509-4d14-9277-f7205658eaec · outbound

This paper cites AniPortrait: Audio-Driven Synthesis of Photorealistic Portrait Animation.

Let Them Talk: Audio-Driven Multi-Person Conversational Video Generation AniPortrait: Audio-Driven Synthesis of Photorealistic Portrait Animation

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T13:07:34.059612Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:07:34.059612Z digest=sha256:000057957b30525901092a2c9b46f4ef54c850ee9d0757f7011d34eb80eac8e6

Observation ed285332-15ad-49eb-8270-340bcb57a288 · outbound

This paper cites Sonic: Shifting Focus to Global Audio Perception in Portrait Animation.

Let Them Talk: Audio-Driven Multi-Person Conversational Video Generation Sonic: Shifting Focus to Global Audio Perception in Portrait Animation

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T13:07:34.169979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:07:34.169979Z digest=sha256:0c0cdcfd6835b54a3ef4884f8dcb9617c7a782b2a4f2e4841ac85fe335cf0926

Observation c09bd768-67e9-454f-ad5b-96dc55584f40 · outbound

This paper cites High- resolution image synthesis with latent diffusion models.

Let Them Talk: Audio-Driven Multi-Person Conversational Video Generation High- resolution image synthesis with latent diffusion models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T13:07:34.216777Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:07:34.216777Z digest=sha256:9a0f1a5fec287761a653d0638afce81fc9783a34d1024f6ba8636c3376bc5333

Observation bc0ffce2-f473-4b86-8b1a-a197b4d7db29 · outbound

This paper cites SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis.

Let Them Talk: Audio-Driven Multi-Person Conversational Video Generation SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T13:07:34.276289Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:07:34.276289Z digest=sha256:6b771872da7e3b3817f6d6c1e821fc0bb36ebf9e463352656cc2a35148d55330

Observation 0ecc0f12-bfbd-4fed-9640-38d0985b3b6c · outbound

This paper cites Dreambooth: Fine tuning text-to-image diffusion models for subject-driven generation.

Let Them Talk: Audio-Driven Multi-Person Conversational Video Generation Dreambooth: Fine tuning text-to-image diffusion models for subject-driven generation

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:07:36.846962Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T13:07:34.329887Z digest=sha256:750422b7850d7a812b3fba2d8dbb69799f830d340490f24bd9d030755fea7564

Observation 371b3d49-83ea-48ec-92bf-3dd33e1180fd · outbound

This paper cites Omg: Occlusion-friendly personalized multi-concept generation in diffusion models.

Let Them Talk: Audio-Driven Multi-Person Conversational Video Generation Omg: Occlusion-friendly personalized multi-concept generation in diffusion models

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:07:36.657064Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T13:07:34.358893Z digest=sha256:941d72908151ed0ad7e42fbf1820aebb01ff088248aec34ffd8d28cfcaff54a3

Observation 9670f83a-ce27-43d2-a9b0-b70fee369f18 · outbound

This paper cites VideoCrafter1: Open Diffusion Models for High-Quality Video Generation.

Let Them Talk: Audio-Driven Multi-Person Conversational Video Generation VideoCrafter1: Open Diffusion Models for High-Quality Video Generation

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T13:07:34.412479Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:07:34.412479Z digest=sha256:4cd650bc82261841fc75fded6abf029e1f188bd9f397e5fc7cf2f4c678c8363d

Observation 45142d58-ff24-4dff-b3ce-ff0fafedebe3 · outbound

This paper cites Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets.

Let Them Talk: Audio-Driven Multi-Person Conversational Video Generation Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T13:07:34.493506Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:07:34.493506Z digest=sha256:e258a46b02a97a3bcfd0a5df74be57fa5d806ae64842afcca7e04f97e45bc528

Observation 99fcc958-4322-4fa7-b2df-68e4f9319c6d · outbound

This paper cites CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer.

Let Them Talk: Audio-Driven Multi-Person Conversational Video Generation CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T13:07:34.552987Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:07:34.552987Z digest=sha256:6abe050fe80381dbbdd5e479f00e9c8b4fbc20e080c38056aa8ffe002e8d9037

Observation f6fc5fa5-7aa1-4b42-b00d-4e50b51b8941 · outbound

This paper cites HunyuanVideo: A Systematic Framework For Large Video Generative Models.

Let Them Talk: Audio-Driven Multi-Person Conversational Video Generation HunyuanVideo: A Systematic Framework For Large Video Generative Models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T13:07:34.579279Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:07:34.579279Z digest=sha256:a9a9882dbb928431fe5026bb3251346f0da59cb6614b523b2eae2bf61f9f7a1b

Observation d2cfe75d-30fa-4cbe-aa61-1bcdc41eaec0 · outbound

This paper cites Scalable diffusion models with transformers.

Let Them Talk: Audio-Driven Multi-Person Conversational Video Generation Scalable diffusion models with transformers

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T13:07:34.614272Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:07:34.614272Z digest=sha256:9757ccec7f576c0fb40cd0563a65f65fef2db035ccce8d0e32cf2128716790fd

Observation 9768261e-b0d3-480e-a0a8-7c803596f44d · outbound

This paper cites StereoCrafter: Diffusion-based Generation of Long and High-fidelity Stereoscopic 3D from Monocular Videos.

Let Them Talk: Audio-Driven Multi-Person Conversational Video Generation StereoCrafter: Diffusion-based Generation of Long and High-fidelity Stereoscopic 3D from Monocular Videos

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T13:07:34.672108Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:07:34.672108Z digest=sha256:3c3112196129eae50d50236aeb47e0e78634ba5cfdf2ba416a2ab4500143b135

Observation 1d053117-d560-4002-9f3d-e633d01bf2db · outbound

This paper cites StyleMaster: Stylize Your Video with Artistic Generation and Translation.

Let Them Talk: Audio-Driven Multi-Person Conversational Video Generation StyleMaster: Stylize Your Video with Artistic Generation and Translation

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T13:07:34.731259Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:07:34.731259Z digest=sha256:954ec57f8bec0d7ff8fd01dff5fdf457759a5503e78051e1fe5a672ef1034676

Observation c0459e79-81c5-434a-9579-81238439c0fa · outbound

This paper cites Towards multiple character image animation through enhancing implicit decoupling.

Let Them Talk: Audio-Driven Multi-Person Conversational Video Generation Towards multiple character image animation through enhancing implicit decoupling

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:07:36.497400Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T13:07:34.777121Z digest=sha256:6bef78b9ab0e6ef1981a0b168cb02fa8b958ad7c424012ef03e1a51d2bb94528

Observation 7d1f711e-748c-4da5-af35-5bfa901d2e78 · outbound

This paper cites Generative Inbetweening: Adapting Image-to-Video Models for Keyframe Interpolation.

Let Them Talk: Audio-Driven Multi-Person Conversational Video Generation Generative Inbetweening: Adapting Image-to-Video Models for Keyframe Interpolation

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T13:07:34.859398Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:07:34.859398Z digest=sha256:77635257f8871597ff0b20fabb63f167a8036d5f71349ddee7ade6a9b6ea191a

Observation d5eba661-413c-4173-a00e-c5026e2133ff · outbound

This paper cites Learning transferable visual models from natural language supervision.

Let Them Talk: Audio-Driven Multi-Person Conversational Video Generation Learning transferable visual models from natural language supervision

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T13:07:34.923083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:07:34.923083Z digest=sha256:2312ee388de5b43b25817d2b4c0ba12dfc1c8bc3f7c8a11c3b990770170eb9fc

Observation f41f52a3-af98-4264-9419-6d1c277d33ad · outbound

This paper cites wav2vec 2.0: A framework for self-supervised learning of speech representations.

Let Them Talk: Audio-Driven Multi-Person Conversational Video Generation wav2vec 2.0: A framework for self-supervised learning of speech representations

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:07:36.305862Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T13:07:34.967027Z digest=sha256:42cfa9d930882f72da98085bb16dd7b17729f25ed39c14b819b6a332a44fc240

Observation bd1241d6-cccc-42ee-87d2-93771b447322 · outbound

This paper cites Roformer: Enhanced transformer with rotary position embedding.Neurocomputing, 568:127063, 2024.

Let Them Talk: Audio-Driven Multi-Person Conversational Video Generation Roformer: Enhanced transformer with rotary position embedding.Neurocomputing, 568:127063, 2024

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T13:07:35.020211Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:07:35.020211Z digest=sha256:4fc06caa09e7a3353e6a11ae1b492ee41c8e0cc2a3db311d1b8dbd8dcefb1d96

Observation d6462992-add8-41c3-810f-1ac97f91f31b · outbound

This paper cites Flow-guided one-shot talking face generation with a high-resolution audio-visual dataset.

Let Them Talk: Audio-Driven Multi-Person Conversational Video Generation Flow-guided one-shot talking face generation with a high-resolution audio-visual dataset

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:07:36.142026Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T13:07:35.062209Z digest=sha256:52b60d6c90b455c42c005190201cd00de8756e146ce3ef983e1b17b29dffaa7a

Observation 0d4da94a-4e14-4506-8e7f-7ba252bfc1f2 · outbound

This paper cites CelebV-HQ: A large-scale video facial attributes dataset.

Let Them Talk: Audio-Driven Multi-Person Conversational Video Generation CelebV-HQ: A large-scale video facial attributes dataset

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T13:07:35.138198Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:07:35.138198Z digest=sha256:75c129cec02e1265dce30f584ff007acb5c8da94577089e37f7ab91b1e5d2da0

Observation 322e3111-5c0e-4bfc-b587-f592a7b76e39 · outbound

This paper cites Gans trained by a two time-scale update rule converge to a local nash equilibrium.

Let Them Talk: Audio-Driven Multi-Person Conversational Video Generation Gans trained by a two time-scale update rule converge to a local nash equilibrium

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:07:36.007740Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T13:07:35.210087Z digest=sha256:2f74a68d280c31ca0ed0a0427b3d4780aaa2ee6ddabc1d138bf1a906e122b385

Observation a7c1b579-1744-4a02-83cc-50eef747966b · outbound

This paper cites Fvd: A new metric for video generation.

Let Them Talk: Audio-Driven Multi-Person Conversational Video Generation Fvd: A new metric for video generation

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T13:07:35.267621Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:07:35.267621Z digest=sha256:c92617c980a2b236ddebe55d8136a55bde138d1f95c044e68d7f58f2249c6df5

Observation 3312cc87-6f2e-4766-a6cb-0a5fca454df8 · outbound

This paper cites Out of time: automated lip sync in the wild.

Let Them Talk: Audio-Driven Multi-Person Conversational Video Generation Out of time: automated lip sync in the wild

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:07:35.864051Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T13:07:35.362688Z digest=sha256:fc1db41ad4607d9df5df912bbcb5e5a11cd776da1256d1cd2909713bc45b578a

Pith citing papers

Observation 35bc076f-164b-4f9d-9aee-9b426dc0248a · inbound

OmniAvatar: Efficient Audio-Driven Avatar Video Generation with Adaptive Body Animation cites this paper.

OmniAvatar: Efficient Audio-Driven Avatar Video Generation with Adaptive Body Animation Let Them Talk: Audio-Driven Multi-Person Conversational Video Generation

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-15T18:47:48.965076Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:47:48.965076Z digest=sha256:12436bb86b49b6e14d01468fcf2780bb433d781e600803acb41886eadb137c78

Observation aa4d0478-a48b-47e1-9c22-bf1c1b355111 · inbound

Bind-Your-Avatar: Multi-Talking-Character Video Generation with Dynamic 3D-mask-based Embedding Router cites this paper.

Bind-Your-Avatar: Multi-Talking-Character Video Generation with Dynamic 3D-mask-based Embedding Router Let Them Talk: Audio-Driven Multi-Person Conversational Video Generation

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-15T18:29:41.011307Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:29:41.011307Z digest=sha256:4a2fe43f03367a204fd53a011b55a6ab9016ff9b5405b57a7fd9fbfb80911a1d

Observation e616d77c-07ab-44ad-94e3-9d28ef19142c · inbound

FantasyPortrait: Enhancing Multi-Character Portrait Animation with Expression-Augmented Diffusion Transformers cites this paper.

FantasyPortrait: Enhancing Multi-Character Portrait Animation with Expression-Augmented Diffusion Transformers Let Them Talk: Audio-Driven Multi-Person Conversational Video Generation

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T16:39:35.477467Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:39:35.477467Z digest=sha256:efd4c70f4fefa0a9e5d2b31e02fdc2a19871d60342ee61d743f0afa266ef3484

Observation 9138e043-e8ec-4bc5-b3e6-77ca87adc231 · inbound

StableAvatar: Infinite-Length Audio-Driven Avatar Video Generation cites this paper.

StableAvatar: Infinite-Length Audio-Driven Avatar Video Generation Let Them Talk: Audio-Driven Multi-Person Conversational Video Generation

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-15T17:42:41.185614Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:42:41.185614Z digest=sha256:1d5eaac0a2fe5af819940cf5f73dc2651bbc2b2ba685242fbd64318b2aef901e

Observation 8b35010b-3fba-47ee-be13-7448fe47632f · inbound

FantasyTalking2: Timestep-Layer Adaptive Preference Optimization for Audio-Driven Portrait Animation cites this paper.

FantasyTalking2: Timestep-Layer Adaptive Preference Optimization for Audio-Driven Portrait Animation Let Them Talk: Audio-Driven Multi-Person Conversational Video Generation

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-05T20:06:47.005370Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:06:47.005370Z digest=sha256:20cc2d3cf80092673e85d3649f636f5b125f1c2f356cdec9a68f35a35b1463b7

Observation dc40bea0-8f73-4da4-a842-e54029d6a5f6 · inbound

InfiniteTalk: Audio-driven Video Generation for Sparse-Frame Video Dubbing cites this paper.

InfiniteTalk: Audio-driven Video Generation for Sparse-Frame Video Dubbing Let Them Talk: Audio-Driven Multi-Person Conversational Video Generation

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-05T18:50:16.246901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T18:50:16.246901Z digest=sha256:5bc0cc9c073c8c7555dee09294197fa3ada25c3c8c6c1d67972e3a1cf045c233

Observation 314e670b-57e9-4958-980c-3f1c7d83ef65 · inbound

OmniHuman-1.5: Instilling an Active Mind in Avatars via Cognitive Simulation cites this paper.

OmniHuman-1.5: Instilling an Active Mind in Avatars via Cognitive Simulation Let Them Talk: Audio-Driven Multi-Person Conversational Video Generation

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-15T16:57:29.599781Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:57:29.599781Z digest=sha256:353a24f98235909cb2c8bac436f0b14a613e72d23bafceac9256a1e7b3503bf4

Observation b7dc1f40-57f5-44e4-a77f-f664e284d62b · inbound

MIDAS: Multimodal Interactive Digital-humAn Synthesis via Real-time Autoregressive Video Generation cites this paper.

MIDAS: Multimodal Interactive Digital-humAn Synthesis via Real-time Autoregressive Video Generation Let Them Talk: Audio-Driven Multi-Person Conversational Video Generation

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-05T16:01:30.762208Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:01:30.762208Z digest=sha256:66f3b8413169a02194dbebb15574d3490081d9c5eaf09e18e21723b493c18fbf

Observation 2b649be8-018a-4de3-b28b-05f9b18454b2 · inbound

InfinityHuman: Towards Long-Term Audio-Driven Human cites this paper.

InfinityHuman: Towards Long-Term Audio-Driven Human Let Them Talk: Audio-Driven Multi-Person Conversational Video Generation

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-05T15:19:37.168237Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:19:37.168237Z digest=sha256:53fae6decc4c5aa64b1576da13ca87098b15b848eebfa7ae21518236975844c0

Observation 8a4458db-a56d-4c30-b58b-9acbacd29f39 · inbound

iMontage: Unified, Versatile, Highly Dynamic Many-to-many Image Generation cites this paper.

iMontage: Unified, Versatile, Highly Dynamic Many-to-many Image Generation Let Them Talk: Audio-Driven Multi-Person Conversational Video Generation

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-04T06:47:05.973725Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:47:05.973725Z digest=sha256:2cd61004dba0a8f7da7d3402383acd9ee28e649e7451d4615fe2005184b3b0e9

Observation e141d963-46eb-4e0b-9dd4-a4da10d6e3bc · inbound

AUHead: Realistic Emotional Talking Head Generation via Action Units Control cites this paper.

AUHead: Realistic Emotional Talking Head Generation via Action Units Control Let Them Talk: Audio-Driven Multi-Person Conversational Video Generation

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-16T05:50:40.255880Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-16T05:49:15.734418Z digest=sha256:3fb2934de6683c1561f60cfb9b60a92cb26a48bb6b9e03c0392fe7cd320f1ac0

Observation 8a522e79-6f7f-40c3-9027-d8ac000648e5 · inbound

OmniShow: Unifying Multimodal Conditions for Human-Object Interaction Video Generation cites this paper.

OmniShow: Unifying Multimodal Conditions for Human-Object Interaction Video Generation Let Them Talk: Audio-Driven Multi-Person Conversational Video Generation

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-11T11:06:05.818834Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-10T15:09:02.727887Z digest=sha256:2f19c715407aa8e1b584fde2878cbe205c787c31952a88f28adcf55c8c5a3c55

Observation f203a03a-a1d4-41f3-8e1a-4be122b0e434 · inbound

TurboTalk: Progressive Distillation for One-Step Audio-Driven Talking Avatar Generation cites this paper.

TurboTalk: Progressive Distillation for One-Step Audio-Driven Talking Avatar Generation Let Them Talk: Audio-Driven Multi-Person Conversational Video Generation

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-10T11:15:10.341482Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-10T11:13:27.689539Z digest=sha256:d941ef042987749e749c35a25035068dea21a8e2eb24a3e88b74188647509f09

Observation e91fb8d3-1b09-4292-9d6b-8aa21c96918c · inbound

PresentAgent-2: Towards Generalist Multimodal Presentation Agents cites this paper.

PresentAgent-2: Towards Generalist Multimodal Presentation Agents Let Them Talk: Audio-Driven Multi-Person Conversational Video Generation

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-13T02:32:06.441607Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-13T02:29:42.157339Z digest=sha256:5989b018e025e9f860b5d976f30118396f8aa70d8066c37ef4805f8bd8e5bbe3

Observation 4465b051-f896-42db-a793-18df0d46d9ca · inbound

Test-Time Self-Adaptive Conditioning for Stable Audio-Driven Talking-Head Generation cites this paper.

Test-Time Self-Adaptive Conditioning for Stable Audio-Driven Talking-Head Generation Let Them Talk: Audio-Driven Multi-Person Conversational Video Generation

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-06-29T22:44:01.724087Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-29T22:36:53.138354Z digest=sha256:1686e950ac2284c04a3c1d22e0fcf72d3883a686c99e9c9fc0bb6720c81e886c

Observation 4373d6ea-fa3d-45f8-8968-bf36cdb2981b · inbound

InteractiveAvatar: Real-Time Streaming Video Generation for Consistent and Intent-Aware Avatars cites this paper.

InteractiveAvatar: Real-Time Streaming Video Generation for Consistent and Intent-Aware Avatars Let Them Talk: Audio-Driven Multi-Person Conversational Video Generation

Reference 15

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T10:09:44.573639Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-26T09:09:06.925645Z digest=sha256:866101ac871c4e5a863086a55fc50ebe8fad94d6467320bbcd3a7d2953c8c753

Observation 5b418ee3-d410-4682-877f-e5aac23dcbe5 · inbound

InteractiveAvatar: Real-Time Streaming Video Generation for Consistent and Intent-Aware Avatars cites this paper.

InteractiveAvatar: Real-Time Streaming Video Generation for Consistent and Intent-Aware Avatars Let Them Talk: Audio-Driven Multi-Person Conversational Video Generation

Reference 16

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T07:05:29.129385Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-07-01T07:00:53.496569Z digest=sha256:0edd5e9f69f69e1c3e5cd7ead5ed4c38fffd319bd518b3c2eca041d2c4c4518c

Observation 46353343-5273-4f76-8a0e-0cb7f3a567fe · inbound

Towards Flexible, Natural, Efficient Interaction for Conversational Talking Face Generation cites this paper.

Towards Flexible, Natural, Efficient Interaction for Conversational Talking Face Generation Let Them Talk: Audio-Driven Multi-Person Conversational Video Generation

Reference 14

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T09:35:40.311712Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-07-01T06:26:20.283349Z digest=sha256:9e1fe0f6caa807e59ab6125771ef667fd63e2142e9019775b81b22f04f71dac3

Observation fb4228e5-dd58-4905-9692-3d1b0e345dce · inbound

Towards Flexible, Natural, Efficient Interaction for Conversational Talking Face Generation cites this paper.

Towards Flexible, Natural, Efficient Interaction for Conversational Talking Face Generation Let Them Talk: Audio-Driven Multi-Person Conversational Video Generation

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-02T09:27:55.316754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T09:27:55.316754Z digest=sha256:c88c609974423c91e65e776c81536d0ca553ac21b4d7417cc26a2e33159206b4

Observation e825ffa0-523b-44a7-ac06-98abd718733f · inbound

OmniMate: Open-Ended Real-Time Streaming Audio-Visual Generation for Interactive Avatars cites this paper.

OmniMate: Open-Ended Real-Time Streaming Audio-Visual Generation for Interactive Avatars Let Them Talk: Audio-Driven Multi-Person Conversational Video Generation

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-01T03:52:55.358485Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T03:52:55.358485Z digest=sha256:57b69077cac781d57ddb7bd53b4947e9679b91144a73653d533a72ec093955e4