Pith. sign in

Paper Citation Record · LEDGER

Let Them Talk: Audio-Driven Multi-Person Conversational Video Generation

As of 9 August 2026, this Paper Citation Record lists 46 of 46 outbound references and 16 inbound Pith citation observations for arXiv:2505.22647.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.22647 v1

Coverage vector

measured 46 of 46 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:07:35.362688Z

measured 62 of 62 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 16 of 16 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T16:39:35.477467Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T10:09:44.572116Z

Reference resolution

46 of 46 outbound references displayed

  • verified exact0
  • verified fuzzy17
  • unresolved29
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e4bf8d19-192d-4813-9711-d75132ca8e69 · outbound

This paper cites Emo: Emote portrait alive generating expressive portrait videos with audio2video diffusion model under weak conditions.

Let Them Talk: Audio-Driven Multi-Person Conversational Video Generation Emo: Emote portrait alive generating expressive portrait videos with audio2video diffusion model under weak conditions

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T13:07:31.369913Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:07:31.369913Z digest=sha256:57007617187112e4e92ea26c35e2375936259be84be1c133ca6d87ad7652b7a6

Observation 7ab23656-362f-4d99-bfd5-0ca9b01b2dd1 · outbound

This paper cites Hallo: Hierarchical Audio-Driven Visual Synthesis for Portrait Image Animation.

Let Them Talk: Audio-Driven Multi-Person Conversational Video Generation Hallo: Hierarchical Audio-Driven Visual Synthesis for Portrait Image Animation

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T13:07:31.440529Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:07:31.440529Z digest=sha256:da5bfd5017bc0833f59b5dde70ec54fdb323394966112ca81d705e571f5f21f0

Observation 126822ad-cdb2-4732-b1c0-b087d9131350 · outbound

This paper cites Hallo3: Highly Dynamic and Realistic Portrait Image Animation with Video Diffusion Transformer.

Let Them Talk: Audio-Driven Multi-Person Conversational Video Generation Hallo3: Highly Dynamic and Realistic Portrait Image Animation with Video Diffusion Transformer

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T13:07:31.596042Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:07:31.596042Z digest=sha256:116c9f1fd6ecd8980ca63e8d9b22b229d2731b2af2685099bae4900ef62d6e8d

Observation 83457919-cecf-463f-8ee5-730e3cc5059c · outbound

This paper cites Echomimic: Lifelike audio-driven portrait animations through editable landmark conditions.

Let Them Talk: Audio-Driven Multi-Person Conversational Video Generation Echomimic: Lifelike audio-driven portrait animations through editable landmark conditions

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:07:38.450148Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:07:31.759871Z digest=sha256:03c407276755830b02e172306fd3edecef7564b71b5dc30d2f4712ff6c0e15a3

Observation 8f6b2128-93cc-4554-990d-7c199ea931a3 · outbound

This paper cites LatentSync: Taming Audio-Conditioned Latent Diffusion Models for Lip Sync with SyncNet Supervision.

Let Them Talk: Audio-Driven Multi-Person Conversational Video Generation LatentSync: Taming Audio-Conditioned Latent Diffusion Models for Lip Sync with SyncNet Supervision

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T13:07:31.880403Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:07:31.880403Z digest=sha256:8a704a06f9619e359890ad8935a80b841fbfd7497bb40cd5a3f6239dcf903b0d

Observation 54b61ee0-15ce-4696-8d2b-805a45979bb4 · outbound

This paper cites Loopy: Taming Audio-Driven Portrait Avatar with Long-Term Motion Dependency.

Let Them Talk: Audio-Driven Multi-Person Conversational Video Generation Loopy: Taming Audio-Driven Portrait Avatar with Long-Term Motion Dependency

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T13:07:31.988771Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:07:31.988771Z digest=sha256:23a21e6df94274459405922717262e37cb6904af95972d7e9ec23449d3a39600

Observation 603278bc-f543-4d50-9d0a-4cb48ffa9b45 · outbound

This paper cites Cyberhost: A one-stage diffusion framework for audio-driven talking body generation.

Let Them Talk: Audio-Driven Multi-Person Conversational Video Generation Cyberhost: A one-stage diffusion framework for audio-driven talking body generation

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:07:38.322693Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:07:32.073177Z digest=sha256:7754658e46d60c313668ffe2a12900a308ac43744289b33c9da6057b6af6f330

Observation 146516e4-b405-484e-8be2-5eff66dcbb08 · outbound

This paper cites OmniHuman-1: Rethinking the Scaling-Up of One-Stage Conditioned Human Animation Models.

Let Them Talk: Audio-Driven Multi-Person Conversational Video Generation OmniHuman-1: Rethinking the Scaling-Up of One-Stage Conditioned Human Animation Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T13:07:32.220659Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:07:32.220659Z digest=sha256:1313948a0980ff88eacab72a713822ea01b464e293da5ff1fb66ade29dd293bf

Observation f2aa37e3-d4ee-4e5b-8ad2-d4f0164a6dd9 · outbound

This paper cites EMO2: End-Effector Guided Audio-Driven Avatar Video Generation.

Let Them Talk: Audio-Driven Multi-Person Conversational Video Generation EMO2: End-Effector Guided Audio-Driven Avatar Video Generation

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T13:07:32.372279Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:07:32.372279Z digest=sha256:557465261c00d7bb351d85db1f614e758fbd0fd79b6bff4fa25b31d65e699a48

Observation 7ed5f4a0-ce47-4f08-8461-6d8df225a658 · outbound

This paper cites Echomimicv2: Towards striking, simplified, and semi-body human animation.arXiv preprint arXiv:2411.10061, 2024.

Let Them Talk: Audio-Driven Multi-Person Conversational Video Generation Echomimicv2: Towards striking, simplified, and semi-body human animation.arXiv preprint arXiv:2411.10061, 2024

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T13:07:32.506861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:07:32.506861Z digest=sha256:f2ef622b0c7b0d435431b59dd403e544c05c52a463b306edc435144604758e68

Observation 259b03b1-ff70-43b9-b91d-4b75e3601cab · outbound

This paper cites FantasyTalking: Realistic Talking Portrait Generation via Coherent Motion Synthesis.

Let Them Talk: Audio-Driven Multi-Person Conversational Video Generation FantasyTalking: Realistic Talking Portrait Generation via Coherent Motion Synthesis

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T13:07:32.607328Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:07:32.607328Z digest=sha256:9e1fe5965cfb7c5d605235e85966ac1ca38852b49f0a30d70b302ebf1cffa46f

Observation 1261ea0a-3beb-4122-8082-c182da4cbd1f · outbound

This paper cites AnimateDiff: Animate Your Personalized Text-to-Image Diffusion Models without Specific Tuning.

Let Them Talk: Audio-Driven Multi-Person Conversational Video Generation AnimateDiff: Animate Your Personalized Text-to-Image Diffusion Models without Specific Tuning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T13:07:32.720783Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:07:32.720783Z digest=sha256:6e55c6fecb9a644be5b035d6bf471f23fb086db33001bdef26c8f79e3d72fb72

Observation e308acb5-2e92-4ad9-938c-4c6047b69b48 · outbound

This paper cites Diffusion adversarial post-training for one-step video generation.arXiv preprint arXiv:2501.08316, 2025.

Let Them Talk: Audio-Driven Multi-Person Conversational Video Generation Diffusion adversarial post-training for one-step video generation.arXiv preprint arXiv:2501.08316, 2025

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T13:07:32.854272Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:07:32.854272Z digest=sha256:9c96ff89aa5524eacdd6187d45222ec9b8fb9b55a1ca6815bd2d44c9c47341a3

Observation 7f95a48d-b2d2-458a-a096-8f3ffc45e4d6 · outbound

This paper cites Wan: Open and Advanced Large-Scale Video Generative Models.

Let Them Talk: Audio-Driven Multi-Person Conversational Video Generation Wan: Open and Advanced Large-Scale Video Generative Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T13:07:33.012708Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:07:33.012708Z digest=sha256:269c7194a721e2307aef7dda4d9d5e48e366d431581a88e14c71787f13b2cdee

Observation be6adcbb-8649-49f2-9bb9-d6df34a74a93 · outbound

This paper cites Stylesync: High-fidelity generalized and personalized lip sync in style-based generator.

Let Them Talk: Audio-Driven Multi-Person Conversational Video Generation Stylesync: High-fidelity generalized and personalized lip sync in style-based generator

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:07:38.151957Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:07:33.101465Z digest=sha256:180d50a93b81fb338d34a9cea77f597462ebda6552aeb0219192dd5a904921af

Observation b3e515dd-9aa9-4a8b-af68-e07dbfa7e31a · outbound

This paper cites Sadtalker: Learning realistic 3d motion coefficients for stylized audio-driven single image talking face animation.

Let Them Talk: Audio-Driven Multi-Person Conversational Video Generation Sadtalker: Learning realistic 3d motion coefficients for stylized audio-driven single image talking face animation

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:07:37.978666Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:07:33.182183Z digest=sha256:a20489dcf9f955650d6790a044abc60ed62cd8e6692fd5380254ef8bc326322c

Observation 6276d75d-38de-4d24-9024-a13fab61bf94 · outbound

This paper cites Videoretalking: Audio-based lip synchronization for talking head video editing in the wild.

Let Them Talk: Audio-Driven Multi-Person Conversational Video Generation Videoretalking: Audio-based lip synchronization for talking head video editing in the wild

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:07:37.812816Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:07:33.306501Z digest=sha256:aa3ec899a92cc7874adbb30b462ac3ab941e4ed676ccfb761acfeab708d22f47

Observation 207c8e16-14de-4c8b-a087-5f80423b576c · outbound

This paper cites Dpe: Disentanglement of pose and expression for general video portrait editing.

Let Them Talk: Audio-Driven Multi-Person Conversational Video Generation Dpe: Disentanglement of pose and expression for general video portrait editing

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:07:37.623316Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:07:33.386221Z digest=sha256:b406956a0b78e82f27c873197443d501ea3e67fb9ec67c28eabb368cbc6ff073

Observation c5908bc9-7adf-42bb-8e7e-a33af17923d8 · outbound

This paper cites Styleheat: One-shot high-resolution editable talking face generation via pre-trained stylegan.

Let Them Talk: Audio-Driven Multi-Person Conversational Video Generation Styleheat: One-shot high-resolution editable talking face generation via pre-trained stylegan

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:07:37.496222Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:07:33.450158Z digest=sha256:f40b97272d1bdea2b206678a4199e597d0e473cdfd0ecc1aef6c02d44bc3e522

Observation f61d0710-f06b-4136-8457-72c43342f32a · outbound

This paper cites Toontalker: Cross-domain face reenactment.

Let Them Talk: Audio-Driven Multi-Person Conversational Video Generation Toontalker: Cross-domain face reenactment

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:07:37.265272Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:07:33.519967Z digest=sha256:98afcaee6bb80c818cc36e7fe144aefb35e4ff8aaa05506efd706dfbce3bdada

Observation 74043573-0dc9-4a75-ae0e-87a21f8a5153 · outbound

This paper cites V-Express: Conditional Dropout for Progressive Training of Portrait Video Generation.

Let Them Talk: Audio-Driven Multi-Person Conversational Video Generation V-Express: Conditional Dropout for Progressive Training of Portrait Video Generation

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T13:07:33.665330Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:07:33.665330Z digest=sha256:ef14193f95a095a8acf5417651a132589f6d5b4e89dbf068fc7d2097b496bd9f

Observation 495fb50a-5248-408a-a27b-c6bf9729b9ae · outbound

This paper cites Nonlinear 3d face morphable model.

Let Them Talk: Audio-Driven Multi-Person Conversational Video Generation Nonlinear 3d face morphable model

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:07:37.138422Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:07:33.768430Z digest=sha256:0b606c6df4ae2e96d571e6809da613f4898e0645ffaccaa6993a6cd44966532a

Observation d9bf26c2-7b55-4e30-b37a-59afa8d46960 · outbound

This paper cites Audio-driven dubbing for user generated contents via style-aware semi-parametric synthesis.IEEE TCSVT, 33(3):1247– 1261, 2022.

Let Them Talk: Audio-Driven Multi-Person Conversational Video Generation Audio-driven dubbing for user generated contents via style-aware semi-parametric synthesis.IEEE TCSVT, 33(3):1247– 1261, 2022

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:07:36.972021Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:07:33.896039Z digest=sha256:9ea69f3308fad41b0b173741a44172c0797634d5bfe0d596a397f9dfb8287226

Observation b146ccef-7509-4d14-9277-f7205658eaec · outbound

This paper cites AniPortrait: Audio-Driven Synthesis of Photorealistic Portrait Animation.

Let Them Talk: Audio-Driven Multi-Person Conversational Video Generation AniPortrait: Audio-Driven Synthesis of Photorealistic Portrait Animation

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T13:07:34.059612Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:07:34.059612Z digest=sha256:38656b6271c532a66005e5e058e1b473a34a0a268733fd2b5a8a12e4f42bdc05

Observation ed285332-15ad-49eb-8270-340bcb57a288 · outbound

This paper cites Sonic: Shifting Focus to Global Audio Perception in Portrait Animation.

Let Them Talk: Audio-Driven Multi-Person Conversational Video Generation Sonic: Shifting Focus to Global Audio Perception in Portrait Animation

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T13:07:34.169979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:07:34.169979Z digest=sha256:2dec7ae8e3e6e390774162d1f0c7e67ab7f44d6a56558c75f07a95eb3695f4b8

Observation c09bd768-67e9-454f-ad5b-96dc55584f40 · outbound

This paper cites High- resolution image synthesis with latent diffusion models.

Let Them Talk: Audio-Driven Multi-Person Conversational Video Generation High- resolution image synthesis with latent diffusion models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T13:07:34.216777Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:07:34.216777Z digest=sha256:9a0f1a5fec287761a653d0638afce81fc9783a34d1024f6ba8636c3376bc5333

Observation bc0ffce2-f473-4b86-8b1a-a197b4d7db29 · outbound

This paper cites SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis.

Let Them Talk: Audio-Driven Multi-Person Conversational Video Generation SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T13:07:34.276289Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:07:34.276289Z digest=sha256:623a0064b8a144e1badcb031c0f3b259fd2941b59ca1c357ffc98bccca1df6a0

Observation 0ecc0f12-bfbd-4fed-9640-38d0985b3b6c · outbound

This paper cites Dreambooth: Fine tuning text-to-image diffusion models for subject-driven generation.

Let Them Talk: Audio-Driven Multi-Person Conversational Video Generation Dreambooth: Fine tuning text-to-image diffusion models for subject-driven generation

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:07:36.846962Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:07:34.329887Z digest=sha256:0e3e66cfc7975d9b6a50fde48cbd9b434d203bfdc884abfa8cf305e1a2d3da7b

Observation 371b3d49-83ea-48ec-92bf-3dd33e1180fd · outbound

This paper cites Omg: Occlusion-friendly personalized multi-concept generation in diffusion models.

Let Them Talk: Audio-Driven Multi-Person Conversational Video Generation Omg: Occlusion-friendly personalized multi-concept generation in diffusion models

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:07:36.657064Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:07:34.358893Z digest=sha256:bab952b874e895c9e37bc9cc9d140a92d74855bff606ee80691248beaf9cf11a

Observation 9670f83a-ce27-43d2-a9b0-b70fee369f18 · outbound

This paper cites VideoCrafter1: Open Diffusion Models for High-Quality Video Generation.

Let Them Talk: Audio-Driven Multi-Person Conversational Video Generation VideoCrafter1: Open Diffusion Models for High-Quality Video Generation

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T13:07:34.412479Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:07:34.412479Z digest=sha256:6cfaf503b2ccdd4839b2690ee0d031490e97825530f179ca66dfb468c7902634

Observation 45142d58-ff24-4dff-b3ce-ff0fafedebe3 · outbound

This paper cites Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets.

Let Them Talk: Audio-Driven Multi-Person Conversational Video Generation Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T13:07:34.493506Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:07:34.493506Z digest=sha256:e258a46b02a97a3bcfd0a5df74be57fa5d806ae64842afcca7e04f97e45bc528

Observation 99fcc958-4322-4fa7-b2df-68e4f9319c6d · outbound

This paper cites CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer.

Let Them Talk: Audio-Driven Multi-Person Conversational Video Generation CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T13:07:34.552987Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:07:34.552987Z digest=sha256:6abe050fe80381dbbdd5e479f00e9c8b4fbc20e080c38056aa8ffe002e8d9037

Observation f6fc5fa5-7aa1-4b42-b00d-4e50b51b8941 · outbound

This paper cites HunyuanVideo: A Systematic Framework For Large Video Generative Models.

Let Them Talk: Audio-Driven Multi-Person Conversational Video Generation HunyuanVideo: A Systematic Framework For Large Video Generative Models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T13:07:34.579279Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:07:34.579279Z digest=sha256:877e86684503068a4d7acc4272dcfea1a760558495f9c1d15a01f23f8c935e4f

Observation d2cfe75d-30fa-4cbe-aa61-1bcdc41eaec0 · outbound

This paper cites Scalable diffusion models with transformers.

Let Them Talk: Audio-Driven Multi-Person Conversational Video Generation Scalable diffusion models with transformers

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T13:07:34.614272Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:07:34.614272Z digest=sha256:9757ccec7f576c0fb40cd0563a65f65fef2db035ccce8d0e32cf2128716790fd

Observation 9768261e-b0d3-480e-a0a8-7c803596f44d · outbound

This paper cites StereoCrafter: Diffusion-based Generation of Long and High-fidelity Stereoscopic 3D from Monocular Videos.

Let Them Talk: Audio-Driven Multi-Person Conversational Video Generation StereoCrafter: Diffusion-based Generation of Long and High-fidelity Stereoscopic 3D from Monocular Videos

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T13:07:34.672108Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:07:34.672108Z digest=sha256:33abb4afe363e34359d97b6d72b14b154ec455d71d534469e09cb2be813f8db4

Observation 1d053117-d560-4002-9f3d-e633d01bf2db · outbound

This paper cites StyleMaster: Stylize Your Video with Artistic Generation and Translation.

Let Them Talk: Audio-Driven Multi-Person Conversational Video Generation StyleMaster: Stylize Your Video with Artistic Generation and Translation

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T13:07:34.731259Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:07:34.731259Z digest=sha256:bbcd31b7d877af8c42bef515ad35b7311e9864d4ef02ebf87904a771e6402233

Observation c0459e79-81c5-434a-9579-81238439c0fa · outbound

This paper cites Towards multiple character image animation through enhancing implicit decoupling.

Let Them Talk: Audio-Driven Multi-Person Conversational Video Generation Towards multiple character image animation through enhancing implicit decoupling

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:07:36.497400Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:07:34.777121Z digest=sha256:f97e7e6e388794604cce9e65a58ef759f5e8bf4a35e215a8390ac284a91bb971

Observation 7d1f711e-748c-4da5-af35-5bfa901d2e78 · outbound

This paper cites Generative Inbetweening: Adapting Image-to-Video Models for Keyframe Interpolation.

Let Them Talk: Audio-Driven Multi-Person Conversational Video Generation Generative Inbetweening: Adapting Image-to-Video Models for Keyframe Interpolation

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T13:07:34.859398Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:07:34.859398Z digest=sha256:4a2c6dbf3f65c9414be98ae5510e68b6e3719a1413986868f2642a82ebcb97a1

Observation d5eba661-413c-4173-a00e-c5026e2133ff · outbound

This paper cites Learning transferable visual models from natural language supervision.

Let Them Talk: Audio-Driven Multi-Person Conversational Video Generation Learning transferable visual models from natural language supervision

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T13:07:34.923083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:07:34.923083Z digest=sha256:2312ee388de5b43b25817d2b4c0ba12dfc1c8bc3f7c8a11c3b990770170eb9fc

Observation f41f52a3-af98-4264-9419-6d1c277d33ad · outbound

This paper cites wav2vec 2.0: A framework for self-supervised learning of speech representations.

Let Them Talk: Audio-Driven Multi-Person Conversational Video Generation wav2vec 2.0: A framework for self-supervised learning of speech representations

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:07:36.305862Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:07:34.967027Z digest=sha256:64fcf23d1dbe216a47eb03f8445b8a8cabecced7f3c2fb9d1c5b668c5e6084fd

Observation bd1241d6-cccc-42ee-87d2-93771b447322 · outbound

This paper cites Roformer: Enhanced transformer with rotary position embedding.Neurocomputing, 568:127063, 2024.

Let Them Talk: Audio-Driven Multi-Person Conversational Video Generation Roformer: Enhanced transformer with rotary position embedding.Neurocomputing, 568:127063, 2024

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T13:07:35.020211Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:07:35.020211Z digest=sha256:4fc06caa09e7a3353e6a11ae1b492ee41c8e0cc2a3db311d1b8dbd8dcefb1d96

Observation d6462992-add8-41c3-810f-1ac97f91f31b · outbound

This paper cites Flow-guided one-shot talking face generation with a high-resolution audio-visual dataset.

Let Them Talk: Audio-Driven Multi-Person Conversational Video Generation Flow-guided one-shot talking face generation with a high-resolution audio-visual dataset

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:07:36.142026Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:07:35.062209Z digest=sha256:cc301078a67bb9c5f4eb85f23a8db43036cd590e40c7ef96b9bb2523c11e7f7e

Observation 0d4da94a-4e14-4506-8e7f-7ba252bfc1f2 · outbound

This paper cites CelebV-HQ: A large-scale video facial attributes dataset.

Let Them Talk: Audio-Driven Multi-Person Conversational Video Generation CelebV-HQ: A large-scale video facial attributes dataset

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T13:07:35.138198Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:07:35.138198Z digest=sha256:75c129cec02e1265dce30f584ff007acb5c8da94577089e37f7ab91b1e5d2da0

Observation 322e3111-5c0e-4bfc-b587-f592a7b76e39 · outbound

This paper cites Gans trained by a two time-scale update rule converge to a local nash equilibrium.

Let Them Talk: Audio-Driven Multi-Person Conversational Video Generation Gans trained by a two time-scale update rule converge to a local nash equilibrium

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:07:36.007740Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:07:35.210087Z digest=sha256:c3cd6d7daa501c5adc128d2f1fdc1b9045fc5947dfccaa842505bb89496fa807

Observation a7c1b579-1744-4a02-83cc-50eef747966b · outbound

This paper cites Fvd: A new metric for video generation.

Let Them Talk: Audio-Driven Multi-Person Conversational Video Generation Fvd: A new metric for video generation

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T13:07:35.267621Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:07:35.267621Z digest=sha256:c92617c980a2b236ddebe55d8136a55bde138d1f95c044e68d7f58f2249c6df5

Observation 3312cc87-6f2e-4766-a6cb-0a5fca454df8 · outbound

This paper cites Out of time: automated lip sync in the wild.

Let Them Talk: Audio-Driven Multi-Person Conversational Video Generation Out of time: automated lip sync in the wild

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:07:35.864051Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:07:35.362688Z digest=sha256:c2ef958ed149257647380c95f3796ffa304e0850b04a39d094ae2754514ea238

Pith citing papers

Observation e616d77c-07ab-44ad-94e3-9d28ef19142c · inbound

FantasyPortrait: Enhancing Multi-Character Portrait Animation with Expression-Augmented Diffusion Transformers cites this paper.

FantasyPortrait: Enhancing Multi-Character Portrait Animation with Expression-Augmented Diffusion Transformers Let Them Talk: Audio-Driven Multi-Person Conversational Video Generation

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T16:39:35.477467Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:39:35.477467Z digest=sha256:ab73c14a89f8a893e069fb30751de96785f801483625ed1cac1f46d8f0486e04

Observation 8b35010b-3fba-47ee-be13-7448fe47632f · inbound

FantasyTalking2: Timestep-Layer Adaptive Preference Optimization for Audio-Driven Portrait Animation cites this paper.

FantasyTalking2: Timestep-Layer Adaptive Preference Optimization for Audio-Driven Portrait Animation Let Them Talk: Audio-Driven Multi-Person Conversational Video Generation

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-05T20:06:47.005370Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:06:47.005370Z digest=sha256:67ab02da6fb69d77e73ee6579e612281e0b692eb73d78c98ad31eb1f3348d848

Observation dc40bea0-8f73-4da4-a842-e54029d6a5f6 · inbound

InfiniteTalk: Audio-driven Video Generation for Sparse-Frame Video Dubbing cites this paper.

InfiniteTalk: Audio-driven Video Generation for Sparse-Frame Video Dubbing Let Them Talk: Audio-Driven Multi-Person Conversational Video Generation

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-05T18:50:16.246901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T18:50:16.246901Z digest=sha256:e6c4d2720de20687ef74bdd2e613b9917bd0727537d04788d94e48b92b90bf31

Observation b7dc1f40-57f5-44e4-a77f-f664e284d62b · inbound

MIDAS: Multimodal Interactive Digital-humAn Synthesis via Real-time Autoregressive Video Generation cites this paper.

MIDAS: Multimodal Interactive Digital-humAn Synthesis via Real-time Autoregressive Video Generation Let Them Talk: Audio-Driven Multi-Person Conversational Video Generation

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-05T16:01:30.762208Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:01:30.762208Z digest=sha256:811ade0cc86457b45797207b8a4cb6ad184f16b8e9189d86b818750f89976234

Observation 2b649be8-018a-4de3-b28b-05f9b18454b2 · inbound

InfinityHuman: Towards Long-Term Audio-Driven Human cites this paper.

InfinityHuman: Towards Long-Term Audio-Driven Human Let Them Talk: Audio-Driven Multi-Person Conversational Video Generation

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-05T15:19:37.168237Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:19:37.168237Z digest=sha256:204b16dc681b264a38dfeb2b85104e0aa35b8b78ffbf6432ab8124c64dc496dc

Observation 8a4458db-a56d-4c30-b58b-9acbacd29f39 · inbound

iMontage: Unified, Versatile, Highly Dynamic Many-to-many Image Generation cites this paper.

iMontage: Unified, Versatile, Highly Dynamic Many-to-many Image Generation Let Them Talk: Audio-Driven Multi-Person Conversational Video Generation

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-04T06:47:05.973725Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:47:05.973725Z digest=sha256:e29a0c8aaae9e03fac3241d60d40fb97046de4ef7a65c9f29b6fa7f34777d110

Observation e141d963-46eb-4e0b-9dd4-a4da10d6e3bc · inbound

AUHead: Realistic Emotional Talking Head Generation via Action Units Control cites this paper.

AUHead: Realistic Emotional Talking Head Generation via Action Units Control Let Them Talk: Audio-Driven Multi-Person Conversational Video Generation

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-16T05:50:40.255880Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-16T05:49:15.734418Z digest=sha256:ccd53d7ca6a99e09acc51d4a8335d1876f38a53c29d1b387fa8b0e3972d104ab

Observation 8a522e79-6f7f-40c3-9027-d8ac000648e5 · inbound

OmniShow: Unifying Multimodal Conditions for Human-Object Interaction Video Generation cites this paper.

OmniShow: Unifying Multimodal Conditions for Human-Object Interaction Video Generation Let Them Talk: Audio-Driven Multi-Person Conversational Video Generation

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-11T11:06:05.818834Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T15:09:02.727887Z digest=sha256:ad61847092172425b41c7e0f91dede4c980d03eebd38d9d4018507ab9bf0a36b

Observation f203a03a-a1d4-41f3-8e1a-4be122b0e434 · inbound

TurboTalk: Progressive Distillation for One-Step Audio-Driven Talking Avatar Generation cites this paper.

TurboTalk: Progressive Distillation for One-Step Audio-Driven Talking Avatar Generation Let Them Talk: Audio-Driven Multi-Person Conversational Video Generation

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-10T11:15:10.341482Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T11:13:27.689539Z digest=sha256:a0aeebe954bb63b917e6ddf0676b2ab3a142069b960c18fa14e8b84f308c3908

Observation e91fb8d3-1b09-4292-9d6b-8aa21c96918c · inbound

PresentAgent-2: Towards Generalist Multimodal Presentation Agents cites this paper.

PresentAgent-2: Towards Generalist Multimodal Presentation Agents Let Them Talk: Audio-Driven Multi-Person Conversational Video Generation

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-13T02:32:06.441607Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-13T02:29:42.157339Z digest=sha256:737ecb1a81a58167995dd987746e641ae212de4bf0f6befb8eccb5c116dba9ec

Observation 4465b051-f896-42db-a793-18df0d46d9ca · inbound

Test-Time Self-Adaptive Conditioning for Stable Audio-Driven Talking-Head Generation cites this paper.

Test-Time Self-Adaptive Conditioning for Stable Audio-Driven Talking-Head Generation Let Them Talk: Audio-Driven Multi-Person Conversational Video Generation

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-06-29T22:44:01.724087Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-29T22:36:53.138354Z digest=sha256:c138e922c40ea915bd56acea1f31cae2583a4d887989f4a89d123688fd1c70bd

Observation 4373d6ea-fa3d-45f8-8968-bf36cdb2981b · inbound

InteractiveAvatar: Real-Time Streaming Video Generation for Consistent and Intent-Aware Avatars cites this paper.

InteractiveAvatar: Real-Time Streaming Video Generation for Consistent and Intent-Aware Avatars Let Them Talk: Audio-Driven Multi-Person Conversational Video Generation

Reference 15

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T10:09:44.573639Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-26T09:09:06.925645Z digest=sha256:603f0cde0f1fc48ab0d7185eead01c8c4ceaf0924ef6fb6de3e99db591b74da1

Observation 5b418ee3-d410-4682-877f-e5aac23dcbe5 · inbound

InteractiveAvatar: Real-Time Streaming Video Generation for Consistent and Intent-Aware Avatars cites this paper.

InteractiveAvatar: Real-Time Streaming Video Generation for Consistent and Intent-Aware Avatars Let Them Talk: Audio-Driven Multi-Person Conversational Video Generation

Reference 16

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T07:05:29.129385Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-01T07:00:53.496569Z digest=sha256:be3e2f292937869f990839d495471cdecac56cd098c89eb1fea0363863c1fef9

Observation 46353343-5273-4f76-8a0e-0cb7f3a567fe · inbound

Towards Flexible, Natural, Efficient Interaction for Conversational Talking Face Generation cites this paper.

Towards Flexible, Natural, Efficient Interaction for Conversational Talking Face Generation Let Them Talk: Audio-Driven Multi-Person Conversational Video Generation

Reference 14

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T09:35:40.311712Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-01T06:26:20.283349Z digest=sha256:a5e3a458f5e06bf1a3bb543f37fde715eec5804de701cb3866621c9c83d5f345

Observation fb4228e5-dd58-4905-9692-3d1b0e345dce · inbound

Towards Flexible, Natural, Efficient Interaction for Conversational Talking Face Generation cites this paper.

Towards Flexible, Natural, Efficient Interaction for Conversational Talking Face Generation Let Them Talk: Audio-Driven Multi-Person Conversational Video Generation

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-02T09:27:55.316754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T09:27:55.316754Z digest=sha256:a435303599e36b48a6476bd0e849162442430add47f13a342238406139836e0e

Observation e825ffa0-523b-44a7-ac06-98abd718733f · inbound

OmniMate: Open-Ended Real-Time Streaming Audio-Visual Generation for Interactive Avatars cites this paper.

OmniMate: Open-Ended Real-Time Streaming Audio-Visual Generation for Interactive Avatars Let Them Talk: Audio-Driven Multi-Person Conversational Video Generation

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-01T03:52:55.358485Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T03:52:55.358485Z digest=sha256:57b69077cac781d57ddb7bd53b4947e9679b91144a73653d533a72ec093955e4