Pith. sign in

Paper Citation Record · LEDGER

LLIA -- Enabling Low-Latency Interactive Avatars: Real-Time Audio-Driven Portrait Video Generation with Diffusion Models

As of 8 August 2026, this Paper Citation Record lists 40 of 40 outbound references and 3 inbound Pith citation observations for arXiv:2506.05806.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.05806 v1

Coverage vector

measured 40 of 40 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T10:18:27.810835Z

measured 43 of 43 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T06:45:49.611770Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-17T01:58:51.404766Z

Reference resolution

40 of 40 outbound references displayed

  • verified exact0
  • verified fuzzy7
  • unresolved33
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation ca17bc80-f60d-4a04-98ef-1412b9481454 · outbound

This paper cites A lip sync expert is all you need for speech to lip generation in the wild.

LLIA -- Enabling Low-Latency Interactive Avatars: Real-Time Audio-Driven Portrait Video Generation with Diffusion Models A lip sync expert is all you need for speech to lip generation in the wild

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T10:18:23.196349Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:18:23.196349Z digest=sha256:83afde7b27b89826930c2003d1c12626cc64cf4428db1638923e565d1dc718b0

Observation 414366dc-3bdd-410c-bb28-36aa5f59796e · outbound

This paper cites Emo: Emote portrait alive generating expressive portrait videos with audio2video diffusion model under weak conditions.

LLIA -- Enabling Low-Latency Interactive Avatars: Real-Time Audio-Driven Portrait Video Generation with Diffusion Models Emo: Emote portrait alive generating expressive portrait videos with audio2video diffusion model under weak conditions

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T10:18:23.303790Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:18:23.303790Z digest=sha256:752c25bb02d3c03a8050f1ca7a76a4a22f5f24a91871e50e6e850ec0ca7998ab

Observation f43dcb10-69d6-49e7-bdeb-7d8b0775adde · outbound

This paper cites Vasa-1: Lifelike audio-driven talking faces generated in real time.Advances in Neural Information Processing Systems, 37:660–684, 2024.

LLIA -- Enabling Low-Latency Interactive Avatars: Real-Time Audio-Driven Portrait Video Generation with Diffusion Models Vasa-1: Lifelike audio-driven talking faces generated in real time.Advances in Neural Information Processing Systems, 37:660–684, 2024

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T10:18:23.353953Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:18:23.353953Z digest=sha256:a67973c846b949a7811da66a4e403b9d9bcf8f13aee4134cbc8e08797901aa21

Observation bff6e7e8-659e-4164-a4ff-5516565eb2fe · outbound

This paper cites Sadtalker: Learning realistic 3d motion coefficients for stylized audio-driven single image talking face animation.

LLIA -- Enabling Low-Latency Interactive Avatars: Real-Time Audio-Driven Portrait Video Generation with Diffusion Models Sadtalker: Learning realistic 3d motion coefficients for stylized audio-driven single image talking face animation

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:18:29.735534Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:18:23.468957Z digest=sha256:ba9f5bf2a2a60633d12c3a133751665f7ab29e9163bb9c1339fa73e369ce8977

Observation 67a55ddd-daf9-4f89-b61c-93d999784cdc · outbound

This paper cites Denoising Diffusion Implicit Models.

LLIA -- Enabling Low-Latency Interactive Avatars: Real-Time Audio-Driven Portrait Video Generation with Diffusion Models Denoising Diffusion Implicit Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T10:18:23.581989Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:18:23.581989Z digest=sha256:01baacbb2f2176956a22c5b1b7105aa7c21f1fba3c406ef75c29d456a0176810

Observation c55c8c15-75b7-4eff-bc7f-ef6a04219bd3 · outbound

This paper cites Denoising diffusion probabilistic models.Advances in neural information processing systems, 33:6840–6851, 2020.

LLIA -- Enabling Low-Latency Interactive Avatars: Real-Time Audio-Driven Portrait Video Generation with Diffusion Models Denoising diffusion probabilistic models.Advances in neural information processing systems, 33:6840–6851, 2020

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T10:18:23.702045Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:18:23.702045Z digest=sha256:4070e21424d3d5d11025350ae46728d9e29cad9927153fa9d0a74dba7b82f138

Observation 3dd28e69-e6fd-4311-9756-8f0bc0ac0d65 · outbound

This paper cites Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets.

LLIA -- Enabling Low-Latency Interactive Avatars: Real-Time Audio-Driven Portrait Video Generation with Diffusion Models Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T10:18:23.849703Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:18:23.849703Z digest=sha256:4aec4304bcdbac0874aaddeb2052d53ac04ea63857a59152d17e01596ca3615d

Observation 7d26a6e7-3c58-4ad6-a627-9edd312330b6 · outbound

This paper cites Megaportraits: One-shot megapixel neural head avatars.

LLIA -- Enabling Low-Latency Interactive Avatars: Real-Time Audio-Driven Portrait Video Generation with Diffusion Models Megaportraits: One-shot megapixel neural head avatars

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:18:29.479575Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:18:23.901840Z digest=sha256:44b86810d4d4247dd25442353fc0e909f4d4698ef36e79f5bdc9311887e8f2b7

Observation bc068e31-25cc-4c12-bd41-c1c26721d8de · outbound

This paper cites INFP: Audio-Driven Interactive Head Generation in Dyadic Conversations.

LLIA -- Enabling Low-Latency Interactive Avatars: Real-Time Audio-Driven Portrait Video Generation with Diffusion Models INFP: Audio-Driven Interactive Head Generation in Dyadic Conversations

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T10:18:23.943246Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:18:23.943246Z digest=sha256:fe13fe39c06f2a39b46cdd595bf8021c6dbe79af25d4b245f7f32f2df1d6ba10

Observation 81f47936-4d82-4270-8276-134aa87531d0 · outbound

This paper cites Echomimic: Lifelike audio-driven portrait animations through editable landmark conditions.

LLIA -- Enabling Low-Latency Interactive Avatars: Real-Time Audio-Driven Portrait Video Generation with Diffusion Models Echomimic: Lifelike audio-driven portrait animations through editable landmark conditions

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T10:18:24.000147Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:18:24.000147Z digest=sha256:d0982bad7a7610ee4892dc47f453ba784893095fcbb610ed918fd3e4a5b9fd58

Observation ddb92de9-fd10-43c3-beb7-8a690711cdf2 · outbound

This paper cites Hallo: Hierarchical Audio-Driven Visual Synthesis for Portrait Image Animation.

LLIA -- Enabling Low-Latency Interactive Avatars: Real-Time Audio-Driven Portrait Video Generation with Diffusion Models Hallo: Hierarchical Audio-Driven Visual Synthesis for Portrait Image Animation

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T10:18:24.119456Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:18:24.119456Z digest=sha256:38d5ad395b16d43e8f69333f37c583703c609513316137c56d6f4295f027f9f2

Observation be5e88b4-c072-4605-8b1e-ab752b2d2f3e · outbound

This paper cites Hallo2: Long-Duration and High-Resolution Audio-Driven Portrait Image Animation.

LLIA -- Enabling Low-Latency Interactive Avatars: Real-Time Audio-Driven Portrait Video Generation with Diffusion Models Hallo2: Long-Duration and High-Resolution Audio-Driven Portrait Image Animation

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T10:18:24.230525Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:18:24.230525Z digest=sha256:bfc254704b915a39416dfd18531aae388e9486fff28846f0fcff948f2b6f863d

Observation e92e6c34-72c1-457f-94e0-37c71a4b7bc5 · outbound

This paper cites Hallo3: Highly Dynamic and Realistic Portrait Image Animation with Video Diffusion Transformer.

LLIA -- Enabling Low-Latency Interactive Avatars: Real-Time Audio-Driven Portrait Video Generation with Diffusion Models Hallo3: Highly Dynamic and Realistic Portrait Image Animation with Video Diffusion Transformer

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T10:18:24.346992Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:18:24.346992Z digest=sha256:6e0c9da3e4d33e9d1320055cd1727ed669d6e7e6dc90dd1b91ee193b54012d63

Observation 8dce6ce4-191d-456a-94f1-149be2ffcbab · outbound

This paper cites AniPortrait: Audio-Driven Synthesis of Photorealistic Portrait Animation.

LLIA -- Enabling Low-Latency Interactive Avatars: Real-Time Audio-Driven Portrait Video Generation with Diffusion Models AniPortrait: Audio-Driven Synthesis of Photorealistic Portrait Animation

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T10:18:24.450551Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:18:24.450551Z digest=sha256:963ce44915fdf9a4223b5908c4909f8443c06a1634003bb48ab0bf002a1c3d10

Observation 4a76ebc3-cbf2-4246-8763-529b4b596b24 · outbound

This paper cites CyberHost: Taming Audio-driven Avatar Diffusion Model with Region Codebook Attention.

LLIA -- Enabling Low-Latency Interactive Avatars: Real-Time Audio-Driven Portrait Video Generation with Diffusion Models CyberHost: Taming Audio-driven Avatar Diffusion Model with Region Codebook Attention

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T10:18:24.579865Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:18:24.579865Z digest=sha256:98aa8dfa4cd25c42b121013d9872fcafd8db5b78d81d3acb54c21366e07c213f

Observation 33133a8a-1ce5-4dcc-8d69-a991d324d416 · outbound

This paper cites Loopy: Taming Audio-Driven Portrait Avatar with Long-Term Motion Dependency.

LLIA -- Enabling Low-Latency Interactive Avatars: Real-Time Audio-Driven Portrait Video Generation with Diffusion Models Loopy: Taming Audio-Driven Portrait Avatar with Long-Term Motion Dependency

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T10:18:24.712720Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:18:24.712720Z digest=sha256:253b2e9b46faed1330f05b5a6a0b5d963ceca22ab9e9450d3622258d153565ea

Observation 697f31a6-831e-4965-8661-c50adda35c8e · outbound

This paper cites Consistency models.

LLIA -- Enabling Low-Latency Interactive Avatars: Real-Time Audio-Driven Portrait Video Generation with Diffusion Models Consistency models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T10:18:24.846533Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:18:24.846533Z digest=sha256:7bad8358d78a90af551436b47abdf3dd7cbd56529c4baf06649df15905f13620

Observation bc55c5a8-c0ba-4577-ae14-afadc0db8298 · outbound

This paper cites Animate anyone: Consistent and controllable image-to-video synthesis for character animation.

LLIA -- Enabling Low-Latency Interactive Avatars: Real-Time Audio-Driven Portrait Video Generation with Diffusion Models Animate anyone: Consistent and controllable image-to-video synthesis for character animation

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T10:18:24.944161Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:18:24.944161Z digest=sha256:b0ba10dbed6e6fe927093f62664a0efffc2f6acedb90a32a282456052b323d6f

Observation e5eaeb69-6529-43ee-972c-34fcf1bf1952 · outbound

This paper cites LivePortrait: Efficient Portrait Animation with Stitching and Retargeting Control.

LLIA -- Enabling Low-Latency Interactive Avatars: Real-Time Audio-Driven Portrait Video Generation with Diffusion Models LivePortrait: Efficient Portrait Animation with Stitching and Retargeting Control

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T10:18:25.039278Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:18:25.039278Z digest=sha256:8adea129bb20d6d40b24009d1d6c3fd152ee73996571e4a5be36cd00fee75be4

Observation 39e9e265-0112-47c7-9195-558aaf3c522d · outbound

This paper cites X-portrait: Expressive portrait animation with hierarchical motion attention.

LLIA -- Enabling Low-Latency Interactive Avatars: Real-Time Audio-Driven Portrait Video Generation with Diffusion Models X-portrait: Expressive portrait animation with hierarchical motion attention

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:18:29.186798Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:18:25.147513Z digest=sha256:110dfc4c6072e3d61a56e9face1596d578e2266fe49aac6fa43f36c566e7a8e2

Observation f43c7580-7b67-4ca7-ace9-e864facf826f · outbound

This paper cites AnimateDiff: Animate Your Personalized Text-to-Image Diffusion Models without Specific Tuning.

LLIA -- Enabling Low-Latency Interactive Avatars: Real-Time Audio-Driven Portrait Video Generation with Diffusion Models AnimateDiff: Animate Your Personalized Text-to-Image Diffusion Models without Specific Tuning

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T10:18:25.256345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:18:25.256345Z digest=sha256:0130a08349ff1a32c7337c81372a0f4904eefbaa18ea19aab6f037099de1d370

Observation 1dc0c154-2846-459b-86c5-c2fdb5b91be2 · outbound

This paper cites First order motion model for image animation.Advances in neural information processing systems, 32, 2019.

LLIA -- Enabling Low-Latency Interactive Avatars: Real-Time Audio-Driven Portrait Video Generation with Diffusion Models First order motion model for image animation.Advances in neural information processing systems, 32, 2019

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T10:18:25.377340Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:18:25.377340Z digest=sha256:34fffae486cd21e90047248f3c350ceabbd7f063cd51357a9bf27b8153579c98

Observation 8c23bccc-b356-4b6b-bf42-54674df3b85a · outbound

This paper cites OmniTalker: One-shot Real-time Text-Driven Talking Audio-Video Generation With Multimodal Style Mimicking.

LLIA -- Enabling Low-Latency Interactive Avatars: Real-Time Audio-Driven Portrait Video Generation with Diffusion Models OmniTalker: One-shot Real-time Text-Driven Talking Audio-Video Generation With Multimodal Style Mimicking

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T10:18:25.487102Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:18:25.487102Z digest=sha256:32b04101b257dcfda276186197dbbb5d0097fc75f57c280b7fb32f546c0a987c

Observation adf63a13-6dec-4a02-bbb7-b95d647a2d93 · outbound

This paper cites ChatAnyone: Stylized Real-time Portrait Video Generation with Hierarchical Motion Diffusion Model.

LLIA -- Enabling Low-Latency Interactive Avatars: Real-Time Audio-Driven Portrait Video Generation with Diffusion Models ChatAnyone: Stylized Real-time Portrait Video Generation with Hierarchical Motion Diffusion Model

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T10:18:25.637600Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:18:25.637600Z digest=sha256:ea903b7b017bfd0f92032777b37f89b87a58201158406fa9a21d8b23c3a3301a

Observation 717ece2f-dd5b-439a-aad0-e5076817d8f4 · outbound

This paper cites Learning transferable visual models from natural language supervision.

LLIA -- Enabling Low-Latency Interactive Avatars: Real-Time Audio-Driven Portrait Video Generation with Diffusion Models Learning transferable visual models from natural language supervision

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T10:18:25.735445Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:18:25.735445Z digest=sha256:be0d52052f98523839c098c1e6d62708b51b600847be97382c5bacd0b323f19c

Observation 0717dae3-1cfa-4ab1-995d-90e61e9b060a · outbound

This paper cites GPT-4o System Card.

LLIA -- Enabling Low-Latency Interactive Avatars: Real-Time Audio-Driven Portrait Video Generation with Diffusion Models GPT-4o System Card

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T10:18:25.836802Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:18:25.836802Z digest=sha256:9a9bb436956a43af03ca8994f3b9d9783e889140ce3d9177d3fd59ba763ab12b

Observation 968c5f37-a31d-4231-9f4e-a29f2ae3147a · outbound

This paper cites wav2vec 2.0: A framework for self-supervised learning of speech representations.Advances in neural information processing systems, 33:12449– 12460, 2020.

LLIA -- Enabling Low-Latency Interactive Avatars: Real-Time Audio-Driven Portrait Video Generation with Diffusion Models wav2vec 2.0: A framework for self-supervised learning of speech representations.Advances in neural information processing systems, 33:12449– 12460, 2020

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T10:18:25.969650Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:18:25.969650Z digest=sha256:3be78a760ea400496f3c0273c65de850b6b0a77cc7bf11dda7b1136c9b8b3ba4

Observation fbf8f96b-d270-49c2-93b7-92681762c5ff · outbound

This paper cites wav2vec: Unsupervised Pre-training for Speech Recognition.

LLIA -- Enabling Low-Latency Interactive Avatars: Real-Time Audio-Driven Portrait Video Generation with Diffusion Models wav2vec: Unsupervised Pre-training for Speech Recognition

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T10:18:26.115489Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:18:26.115489Z digest=sha256:7284968013bf77bd7525261758b6f5a9721e0423f4052a3d7cb93f700ff2cf24

Observation 5943c728-876a-4f5d-874f-dffb4ad6e053 · outbound

This paper cites chinese speech pretrain, 2022.

LLIA -- Enabling Low-Latency Interactive Avatars: Real-Time Audio-Driven Portrait Video Generation with Diffusion Models chinese speech pretrain, 2022

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:18:29.032142Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:18:26.276961Z digest=sha256:367d22016c9cbd820e6ebdfa96b4a6675f10d7e7aac5350b25b390eb56bac8d1

Observation 2b9f9eef-7a4e-4e63-a5ac-b45a97da4f5b · outbound

This paper cites Hsemotion: High-speed emotion recognition library.Software Impacts, 14:100433, 2022.

LLIA -- Enabling Low-Latency Interactive Avatars: Real-Time Audio-Driven Portrait Video Generation with Diffusion Models Hsemotion: High-speed emotion recognition library.Software Impacts, 14:100433, 2022

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:18:28.831253Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:18:26.409346Z digest=sha256:97e5affdcad451ba330547d147287783acd6ce8ecd5b043b50a0ba1ab38d05cb

Observation a47c9991-4a64-4092-987f-714aa37a07b6 · outbound

This paper cites AnimateDiff-Lightning: Cross-Model Diffusion Distillation.

LLIA -- Enabling Low-Latency Interactive Avatars: Real-Time Audio-Driven Portrait Video Generation with Diffusion Models AnimateDiff-Lightning: Cross-Model Diffusion Distillation

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T10:18:26.580353Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:18:26.580353Z digest=sha256:19c862dbdbc911e9179aaf501b8efd317c462b9aedbd44cb4e516617a246d300

Observation 50b578e2-7349-4c67-9925-0e2ecbd5ec06 · outbound

This paper cites SDXL-Lightning: Progressive Adversarial Diffusion Distillation.

LLIA -- Enabling Low-Latency Interactive Avatars: Real-Time Audio-Driven Portrait Video Generation with Diffusion Models SDXL-Lightning: Progressive Adversarial Diffusion Distillation

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T10:18:26.743525Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:18:26.743525Z digest=sha256:2a4fb699a16cdd0b566ebde505cf748cadc15d3034fe8bc4fc7640e702e3a2f4

Observation b2e01aa1-ed59-4949-a6f3-8843792ae7a4 · outbound

This paper cites Generative adversarial networks.Communications of the ACM, 63(11):139–144, 2020.

LLIA -- Enabling Low-Latency Interactive Avatars: Real-Time Audio-Driven Portrait Video Generation with Diffusion Models Generative adversarial networks.Communications of the ACM, 63(11):139–144, 2020

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T10:18:26.967027Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:18:26.967027Z digest=sha256:686c805eb254ed57ac3c1465ba81096cd352e1721959df04195710ab19f23ab8

Observation 7dacbc6b-1c76-4e54-a9b1-191da6f89350 · outbound

This paper cites Animatelcm: Computation-efficient personalized style video generation without personalized video data.

LLIA -- Enabling Low-Latency Interactive Avatars: Real-Time Audio-Driven Portrait Video Generation with Diffusion Models Animatelcm: Computation-efficient personalized style video generation without personalized video data

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:18:28.603473Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:18:27.146314Z digest=sha256:af5db7cde7135f47228439cd57096585cc19e9d32224f5cdd118bbde1032c381

Observation e31a5c23-a23b-4ce4-9c00-b9aefb053ddf · outbound

This paper cites Latent Consistency Models: Synthesizing High-Resolution Images with Few-Step Inference.

LLIA -- Enabling Low-Latency Interactive Avatars: Real-Time Audio-Driven Portrait Video Generation with Diffusion Models Latent Consistency Models: Synthesizing High-Resolution Images with Few-Step Inference

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T10:18:27.295219Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:18:27.295219Z digest=sha256:c0819d6d858bd89c7bffef07c41605614ed1307d125e9e93215055948d65f42a

Observation dac69101-cd27-46dd-8147-7781b68b1b9a · outbound

This paper cites Mead: A large-scale audio-visual dataset for emotional talking-face generation.

LLIA -- Enabling Low-Latency Interactive Avatars: Real-Time Audio-Driven Portrait Video Generation with Diffusion Models Mead: A large-scale audio-visual dataset for emotional talking-face generation

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:18:28.369641Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T10:18:27.412988Z digest=sha256:1f8e6c44271dc51010782061cf40767f98c9778960eda66ac1029468fea7648d

Observation 6fab25fc-0325-4226-9fef-b30f63d9c48c · outbound

This paper cites VoxCeleb2: Deep Speaker Recognition.

LLIA -- Enabling Low-Latency Interactive Avatars: Real-Time Audio-Driven Portrait Video Generation with Diffusion Models VoxCeleb2: Deep Speaker Recognition

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T10:18:27.509209Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:18:27.509209Z digest=sha256:714f98ee4e6b9914e2971ce704ce13e96527e53fd27cd513bdd4b055745514be

Observation e97bcf03-adde-49d0-9043-b325031729a4 · outbound

This paper cites Celebv-hq: A large-scale video facial attributes dataset.

LLIA -- Enabling Low-Latency Interactive Avatars: Real-Time Audio-Driven Portrait Video Generation with Diffusion Models Celebv-hq: A large-scale video facial attributes dataset

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T10:18:27.601101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:18:27.601101Z digest=sha256:3bbbf2136b8733d1402770a0ff99452d8b52f25c34ce145ed668adcf82e8ce7d

Observation 548058e6-ad0b-41b6-8251-ab0ed6dd69a3 · outbound

This paper cites YOLOv6 v3.0: A Full-Scale Reloading.

LLIA -- Enabling Low-Latency Interactive Avatars: Real-Time Audio-Driven Portrait Video Generation with Diffusion Models YOLOv6 v3.0: A Full-Scale Reloading

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T10:18:27.700080Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:18:27.700080Z digest=sha256:2067884344f5f50ff4fda8ff77ee432bff84e87acb587f0e48e288f5e6e9d970

Observation b2da14e8-311a-41ca-a8be-501cc43112c5 · outbound

This paper cites MediaPipe: A Framework for Building Perception Pipelines.

LLIA -- Enabling Low-Latency Interactive Avatars: Real-Time Audio-Driven Portrait Video Generation with Diffusion Models MediaPipe: A Framework for Building Perception Pipelines

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T10:18:27.810835Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:18:27.810835Z digest=sha256:c03c3a70abf4f29fa1fe49f433fea5484f4a282f970c018bb9c527f8354c1388

Pith citing papers

Observation d188d4fa-de9f-4e3e-9ba6-660af4482247 · inbound

Live Avatar: Streaming Real-time Audio-Driven Avatar Generation with Infinite Length cites this paper.

Live Avatar: Streaming Real-time Audio-Driven Avatar Generation with Infinite Length LLIA -- Enabling Low-Latency Interactive Avatars: Real-Time Audio-Driven Portrait Video Generation with Diffusion Models

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-05-17T01:58:51.407987Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-17T01:56:44.123092Z digest=sha256:aadb4bc1d8cf03b1cd55fb5aab9e9bc0bfb82cc7ab89cff151b910e00a09c475

Observation fb09a11e-2136-4d79-978a-df66026497ea · inbound

Live Avatar: Streaming Real-time Audio-Driven Avatar Generation with Infinite Length cites this paper.

Live Avatar: Streaming Real-time Audio-Driven Avatar Generation with Infinite Length LLIA -- Enabling Low-Latency Interactive Avatars: Real-Time Audio-Driven Portrait Video Generation with Diffusion Models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-03T18:39:43.646679Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:39:43.646679Z digest=sha256:be9b9a980a892765b10dc9427d8636ef2b674fcbe1513ac2b10239b4540cf9e3

Observation 72ceac6e-82b4-470d-b963-349472f310d9 · inbound

EchoCache: Energy-Guided Cross-Modal Caching for Efficient Audio-Driven Video Generation cites this paper.

EchoCache: Energy-Guided Cross-Modal Caching for Efficient Audio-Driven Video Generation LLIA -- Enabling Low-Latency Interactive Avatars: Real-Time Audio-Driven Portrait Video Generation with Diffusion Models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-04T06:45:49.611770Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:45:49.611770Z digest=sha256:6bec0bda2e1d71b2903fe08323c9fee938adc89a3a6eb3a39f86d27fc9d41a2a