Pith. sign in

Paper Citation Record · LEDGER

Hallo: Hierarchical Audio-Driven Visual Synthesis for Portrait Image Animation

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 61 inbound Pith citation observations for arXiv:2406.08801.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2406.08801 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 61 of 61 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 61 of 61 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T21:11:25.346813Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-01T22:06:16.996788Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 454d8982-09dc-40b5-b107-adef64a65c38 · inbound

Multimodal Diffusion Transformer with Memory Bank for Scalable Long-Duration Talking Video Generation cites this paper.

Multimodal Diffusion Transformer with Memory Bank for Scalable Long-Duration Talking Video Generation Hallo: Hierarchical Audio-Driven Visual Synthesis for Portrait Image Animation

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-23T17:23:15.339474Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-23T17:19:51.411937Z digest=sha256:895c3899f185aa9cd24ce00af2a102816b7ad8fc90d0f70b29d9a4aaa76fcc84

Observation 782611a5-f612-4cd7-9bab-0569d37f67e8 · inbound

Long-Term TalkingFace Generation via Motion-Prior Conditional Diffusion Model cites this paper.

Long-Term TalkingFace Generation via Motion-Prior Conditional Diffusion Model Hallo: Hierarchical Audio-Driven Visual Synthesis for Portrait Image Animation

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T21:11:25.346813Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T21:11:25.346813Z digest=sha256:c117251b8417faddfc020bf71a94aa4e205349ae4b4da060d8bbdc0618b965f3

Observation 0bc3d107-8172-44b4-a7e9-6efaf9a46f08 · inbound

HunyuanVideo-Avatar: High-Fidelity Audio-Driven Human Animation for Multiple Characters cites this paper.

HunyuanVideo-Avatar: High-Fidelity Audio-Driven Human Animation for Multiple Characters Hallo: Hierarchical Audio-Driven Visual Synthesis for Portrait Image Animation

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T14:02:30.045227Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:02:30.045227Z digest=sha256:87e164e2a878fddb34e930d9a83d17beddfc862900c1dfb785839c89cf5b2a22

Observation 2bc0af3b-65de-4f02-bee0-74fa0848f548 · inbound

Exploring Timeline Control for Facial Motion Generation cites this paper.

Exploring Timeline Control for Facial Motion Generation Hallo: Hierarchical Audio-Driven Visual Synthesis for Portrait Image Animation

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T13:49:53.135171Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:49:53.135171Z digest=sha256:630dfc38e00a4204f0ce10a64d4ab27ade7cd26ebee6e9d2753b96b3824396b0

Observation bda4b030-9459-4d3e-a797-136c19c67721 · inbound

FaceEditTalker: Controllable Talking Head Generation with Facial Attribute Editing cites this paper.

FaceEditTalker: Controllable Talking Head Generation with Facial Attribute Editing Hallo: Hierarchical Audio-Driven Visual Synthesis for Portrait Image Animation

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T13:21:47.698228Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:21:47.698228Z digest=sha256:40a781689bea2c239dceaf97d37a9b3bfe10d31ffe34b03db4315f34874a8451

Observation 7ab23656-362f-4d99-bfd5-0ca9b01b2dd1 · inbound

Let Them Talk: Audio-Driven Multi-Person Conversational Video Generation cites this paper.

Let Them Talk: Audio-Driven Multi-Person Conversational Video Generation Hallo: Hierarchical Audio-Driven Visual Synthesis for Portrait Image Animation

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T13:07:31.440529Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:07:31.440529Z digest=sha256:da5bfd5017bc0833f59b5dde70ec54fdb323394966112ca81d705e571f5f21f0

Observation 68bc8d76-28d7-45f1-b939-899835f93638 · inbound

Silence is Golden: Leveraging Adversarial Examples to Nullify Audio Control in LDM-based Talking-Head Generation cites this paper.

Silence is Golden: Leveraging Adversarial Examples to Nullify Audio Control in LDM-based Talking-Head Generation Hallo: Hierarchical Audio-Driven Visual Synthesis for Portrait Image Animation

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T11:44:16.996589Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:44:16.996589Z digest=sha256:af05c3dc424ad2be4bca952c7f2ab29a149272586434797f5c66517c8a34e490

Observation aabd3e5e-2151-4911-ae38-133b9270bda3 · inbound

Speaking images. A novel framework for the automated self-description of artworks cites this paper.

Speaking images. A novel framework for the automated self-description of artworks Hallo: Hierarchical Audio-Driven Visual Synthesis for Portrait Image Animation

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T13:17:51.804052Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:17:51.804052Z digest=sha256:73bc574200df803282a69d4d0bf0641ed4287f4c806529411a1dc76a281338ee

Observation ddb92de9-fd10-43c3-beb7-8a690711cdf2 · inbound

LLIA -- Enabling Low-Latency Interactive Avatars: Real-Time Audio-Driven Portrait Video Generation with Diffusion Models cites this paper.

LLIA -- Enabling Low-Latency Interactive Avatars: Real-Time Audio-Driven Portrait Video Generation with Diffusion Models Hallo: Hierarchical Audio-Driven Visual Synthesis for Portrait Image Animation

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T10:18:24.119456Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:18:24.119456Z digest=sha256:38d5ad395b16d43e8f69333f37c583703c609513316137c56d6f4295f027f9f2

Observation 3bb33f79-b67e-4935-a2f6-dd17e8501b7c · inbound

SyncTalk++: High-Fidelity and Efficient Synchronized Talking Heads Synthesis Using Gaussian Splatting cites this paper.

SyncTalk++: High-Fidelity and Efficient Synchronized Talking Heads Synthesis Using Gaussian Splatting Hallo: Hierarchical Audio-Driven Visual Synthesis for Portrait Image Animation

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T00:15:44.531892Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:15:44.531892Z digest=sha256:9059f7c5718b3904011fc8b8d3addd447e6b3d6bc64db949aedf86081c490d07

Observation c3770234-4008-42b3-8d16-364b7c7e8b9f · inbound

GGTalker: Talking Head Systhesis with Generalizable Gaussian Priors and Identity-Specific Adaptation cites this paper.

GGTalker: Talking Head Systhesis with Generalizable Gaussian Priors and Identity-Specific Adaptation Hallo: Hierarchical Audio-Driven Visual Synthesis for Portrait Image Animation

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T22:28:22.248495Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:28:22.248495Z digest=sha256:1f06e367e8953833ca4859907ac3e3bc210972b8bbcb9530586b679a7a3d0f54

Observation e33b25f5-35a6-477b-9fe0-d98e240dc3eb · inbound

MirrorMe: Towards Realtime and High Fidelity Audio-Driven Halfbody Animation cites this paper.

MirrorMe: Towards Realtime and High Fidelity Audio-Driven Halfbody Animation Hallo: Hierarchical Audio-Driven Visual Synthesis for Portrait Image Animation

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T22:15:10.341196Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:15:10.341196Z digest=sha256:69ee7c203285dabe09db523946e8a614db088fe6bd3f43b986421ef36a7b8296

Observation c1f186e3-bbdf-4141-bd1b-e49110f87c7b · inbound

JAM-Flow: Joint Audio-Motion Synthesis with Flow Matching cites this paper.

JAM-Flow: Joint Audio-Motion Synthesis with Flow Matching Hallo: Hierarchical Audio-Driven Visual Synthesis for Portrait Image Animation

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-22T00:50:51.078255Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-22T00:46:39.196042Z digest=sha256:30ea7943c75555b49bf5d45a0eaed2e6acf69ae4df78f8122c51268b388c1d40

Observation 2352e022-fa20-4464-95aa-1a22924a12d1 · inbound

ARIG: Autoregressive Interactive Head Generation for Real-time Conversations cites this paper.

ARIG: Autoregressive Interactive Head Generation for Real-time Conversations Hallo: Hierarchical Audio-Driven Visual Synthesis for Portrait Image Animation

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T21:19:15.990935Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:19:15.990935Z digest=sha256:b053b25c41359ada3a58aaedd5ee712c2dc709a094494df4c35ac6cedb437f8b

Observation 7ee5b291-450d-4725-a774-594001b100eb · inbound

FixTalk: Taming Identity Leakage for High-Quality Talking Head Generation in Extreme Cases cites this paper.

FixTalk: Taming Identity Leakage for High-Quality Talking Head Generation in Extreme Cases Hallo: Hierarchical Audio-Driven Visual Synthesis for Portrait Image Animation

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-06T20:58:30.681138Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:58:30.681138Z digest=sha256:91b6da7cac18b891bf06f1f7f37fa26b5e65464184d0b3276f5dfe3e1c2b242c

Observation 8748f2a9-ef36-44b3-aa33-9157174c1930 · inbound

MoDA: Multi-modal Diffusion Architecture for Talking Head Generation cites this paper.

MoDA: Multi-modal Diffusion Architecture for Talking Head Generation Hallo: Hierarchical Audio-Driven Visual Synthesis for Portrait Image Animation

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T20:23:39.562824Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:23:39.562824Z digest=sha256:70bdc7a411252680400ffe832e1549b105b84ade22f6f8a0bb451ce30bd9be09

Observation 4f8d4c40-5d73-4182-855e-64329f5e6bcd · inbound

MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation cites this paper.

MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation Hallo: Hierarchical Audio-Driven Visual Synthesis for Portrait Image Animation

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T19:38:20.565284Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:38:20.565284Z digest=sha256:a0cd0a75cc4a7daa78939e7fc0e2fdcfedb4af02c2d91a8b753b2cd14e906815

Observation 9d2589d2-9f06-4d19-9e32-6192e00b73a8 · inbound

Think-Before-Draw: Decomposing Emotion Semantics & Fine-Grained Controllable Expressive Talking Head Generation cites this paper.

Think-Before-Draw: Decomposing Emotion Semantics & Fine-Grained Controllable Expressive Talking Head Generation Hallo: Hierarchical Audio-Driven Visual Synthesis for Portrait Image Animation

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T16:42:42.932426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:42:42.932426Z digest=sha256:7d80717a5c44099be67871f8bb57a32a141b8c477c6c66471600a6ab94340b30

Observation bd06b022-ed52-47f7-8dd7-2b69bb19bac8 · inbound

MagicAnime: A Hierarchically Annotated, Multimodal and Multitasking Dataset with Benchmarks for Cartoon Animation Generation cites this paper.

MagicAnime: A Hierarchically Annotated, Multimodal and Multitasking Dataset with Benchmarks for Cartoon Animation Generation Hallo: Hierarchical Audio-Driven Visual Synthesis for Portrait Image Animation

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T13:37:09.440363Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:37:09.440363Z digest=sha256:2f3e6db8dc060ded29e864762c647007a6d46db10a7e41c10970f4c81b54d392

Observation 776c542c-0ebf-4ce1-8283-319b259dc0f0 · inbound

JOLT3D: Joint Learning of Talking Heads and 3DMM Parameters with Application to Lip-Sync cites this paper.

JOLT3D: Joint Learning of Talking Heads and 3DMM Parameters with Application to Lip-Sync Hallo: Hierarchical Audio-Driven Visual Synthesis for Portrait Image Animation

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-06T13:39:34.642522Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:39:34.642522Z digest=sha256:92f18aae3413b6364f390ef4c1f0087661665f828c9f38ce8a3d9850be893bc1

Observation f42f9ad9-2c8b-4f31-b1c6-2c9b39b12e6c · inbound

X-NeMo: Expressive Neural Motion Reenactment via Disentangled Latent Attention cites this paper.

X-NeMo: Expressive Neural Motion Reenactment via Disentangled Latent Attention Hallo: Hierarchical Audio-Driven Visual Synthesis for Portrait Image Animation

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T11:03:44.888102Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:03:44.888102Z digest=sha256:2c7297b2d6dbb7fe3b38c88c01eec5cdac20310dedb92b5ae2157540348eee49

Observation 0c392671-72de-46ef-b7e9-d6338604c9e0 · inbound

Text2Lip: Progressive Lip-Synced Talking Face Generation from Text via Viseme-Guided Rendering cites this paper.

Text2Lip: Progressive Lip-Synced Talking Face Generation from Text via Viseme-Guided Rendering Hallo: Hierarchical Audio-Driven Visual Synthesis for Portrait Image Animation

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T05:05:49.218801Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:05:49.218801Z digest=sha256:c5ea4e244af9ce00b07801843cb165621e1adead2f718de37141b400de5d03a0

Observation 9da71d78-ef20-4e73-bb0d-8d84202ec70f · inbound

DiTalker: A Unified DiT-based Framework for High-Quality and Speaking Styles Controllable Portrait Animation cites this paper.

DiTalker: A Unified DiT-based Framework for High-Quality and Speaking Styles Controllable Portrait Animation Hallo: Hierarchical Audio-Driven Visual Synthesis for Portrait Image Animation

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T12:43:23.782760Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:43:23.782760Z digest=sha256:dbfa4ff351dc9bb5b23ab55be89e977277b06be9d57ea024db14f64e5ac745cc

Observation d1361a53-2ccb-4a5e-b35b-efc01a35f5ba · inbound

EDTalk++: Full Disentanglement for Controllable Talking Head Synthesis cites this paper.

EDTalk++: Full Disentanglement for Controllable Talking Head Synthesis Hallo: Hierarchical Audio-Driven Visual Synthesis for Portrait Image Animation

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-05T19:07:42.180912Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:07:42.180912Z digest=sha256:d6cf241780b0966536838f61a1bc31cd18f6a1318cf1a6be673cc35451bedbda

Observation 19ef1e5d-8a17-44dc-b467-54d2949b301b · inbound

InfiniteTalk: Audio-driven Video Generation for Sparse-Frame Video Dubbing cites this paper.

InfiniteTalk: Audio-driven Video Generation for Sparse-Frame Video Dubbing Hallo: Hierarchical Audio-Driven Visual Synthesis for Portrait Image Animation

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-05T18:50:16.297962Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T18:50:16.297962Z digest=sha256:e6a4b09c65c6785b4f24d0b4f12a9e7a12daced718846cefd2b8d2356335fc05

Observation 62b552ec-c597-4ea0-8810-5b3f829d2845 · inbound

InfinityHuman: Towards Long-Term Audio-Driven Human cites this paper.

InfinityHuman: Towards Long-Term Audio-Driven Human Hallo: Hierarchical Audio-Driven Visual Synthesis for Portrait Image Animation

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-05T15:19:37.259854Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:19:37.259854Z digest=sha256:c24db42d58d73e5060b97bb060aae7a5aaef8760f2e35267928c8f13180458d3

Observation e0f06b7c-dffa-4f88-a39b-534522cb11ef · inbound

Human Motion Video Generation: A Survey cites this paper.

Human Motion Video Generation: A Survey Hallo: Hierarchical Audio-Driven Visual Synthesis for Portrait Image Animation

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-05T10:36:52.656632Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:36:52.656632Z digest=sha256:c0ad7a7a72ecef8a41c0763b85dc2a99932e4d6c765836d4459ea5a05c3c2652

Observation d0d0c22e-34d9-4ad0-9c96-9e2eb28e041c · inbound

FluentAvatar: Flicker-Free Talking-Head Animation via Phoneme-Guided Autoregressive Modeling cites this paper.

FluentAvatar: Flicker-Free Talking-Head Animation via Phoneme-Guided Autoregressive Modeling Hallo: Hierarchical Audio-Driven Visual Synthesis for Portrait Image Animation

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-18T16:42:43.789808Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-18T16:42:25.803856Z digest=sha256:2af46deb03b17acbc0223d06a223b970f001cf9f55de0798fadc2f7b52ad1ce0

Observation 697de81c-4365-4064-9772-1108bc316543 · inbound

STARCaster: Spatio-Temporal AutoRegressive Video Diffusion for Identity- and View-Aware Talking Portraits cites this paper.

STARCaster: Spatio-Temporal AutoRegressive Video Diffusion for Identity- and View-Aware Talking Portraits Hallo: Hierarchical Audio-Driven Visual Synthesis for Portrait Image Animation

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-03T16:30:39.500409Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:30:39.500409Z digest=sha256:d4ea62a8423cb957b0b9c1818ace0d4573d04540760197f302d375f1ec14d1c2

Observation 17b76d59-bf52-4010-a9e8-b08a79c75706 · inbound

AvatarPointillist: AutoRegressive 4D Gaussian Avatarization cites this paper.

AvatarPointillist: AutoRegressive 4D Gaussian Avatarization Hallo: Hierarchical Audio-Driven Visual Synthesis for Portrait Image Animation

Reference 71

Resolution
verified exact
arxiv_id, observed 2026-05-10T23:55:51.152270Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T18:47:37.605996Z digest=sha256:27d1bf44fd7fd844c6cc03c1670dd57b654c80daaed4a72ffc7614ec55f4b714

Observation 0e727eef-15e8-4aac-b21d-edae96a3c738 · inbound

SyncBreaker:Stage-Aware Multimodal Adversarial Attacks on Audio-Driven Talking Head Generation cites this paper.

SyncBreaker:Stage-Aware Multimodal Adversarial Attacks on Audio-Driven Talking Head Generation Hallo: Hierarchical Audio-Driven Visual Synthesis for Portrait Image Animation

Reference 57

Resolution
verified exact
arxiv_id, observed 2026-05-11T06:41:25.776975Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T17:30:36.462567Z digest=sha256:7830c1fadcec1229a26e3a166a2b34ddc6867c177250ffb26f78b8a60bc633a3

Observation 205693a5-d40b-47e9-9cea-eadd417d5a37 · inbound

SyncBreaker:Stage-Aware Multimodal Adversarial Attacks on Audio-Driven Talking Head Generation cites this paper.

SyncBreaker:Stage-Aware Multimodal Adversarial Attacks on Audio-Driven Talking Head Generation Hallo: Hierarchical Audio-Driven Visual Synthesis for Portrait Image Animation

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-02T16:38:16.303378Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T16:38:16.303378Z digest=sha256:35acf8395fc5f94023d229b2650df8e6c761971138cca3ec2dd2192072423900

Observation b7ead486-41cd-413e-8ce4-48e880f4463d · inbound

Beyond Monologue: Interactive Talking-Listening Avatar Generation with Conversational Audio Context-Aware Kernels cites this paper.

Beyond Monologue: Interactive Talking-Listening Avatar Generation with Conversational Audio Context-Aware Kernels Hallo: Hierarchical Audio-Driven Visual Synthesis for Portrait Image Animation

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-05-11T10:56:01.763121Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T15:16:53.603971Z digest=sha256:0840b00345e1d9c7c87d20f562d49281dbab8e27655539784d6a6956a4274157

Observation 898331d4-11ad-4304-b41f-de82d2d0d867 · inbound

Direct Discrepancy Replay: Distribution-Discrepancy Condensation and Manifold-Consistent Replay for Continual Face Forgery Detection cites this paper.

Direct Discrepancy Replay: Distribution-Discrepancy Condensation and Manifold-Consistent Replay for Continual Face Forgery Detection Hallo: Hierarchical Audio-Driven Visual Synthesis for Portrait Image Animation

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-11T11:31:01.252071Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T14:50:39.077649Z digest=sha256:c04515a5e67884da79e43f3a3022aaa6a21260591375607ef30d1e1fff20693a

Observation 1b4cb1a1-22f5-4bab-8424-996a9907d853 · inbound

Giving Faces Their Feelings Back: Explicit Emotion Control for Feedforward Single-Image 3D Head Avatars cites this paper.

Giving Faces Their Feelings Back: Explicit Emotion Control for Feedforward Single-Image 3D Head Avatars Hallo: Hierarchical Audio-Driven Visual Synthesis for Portrait Image Animation

Reference 84

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T11:50:20.275679Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-10T11:47:42.039240Z digest=sha256:04f7ee70e625bc0a9e338fe187ce0aba1f7896bc22df5cb19d621999d5fa588f

Observation ed2442b5-6753-4c39-b7af-53b558c18d0d · inbound

TurboTalk: Progressive Distillation for One-Step Audio-Driven Talking Avatar Generation cites this paper.

TurboTalk: Progressive Distillation for One-Step Audio-Driven Talking Avatar Generation Hallo: Hierarchical Audio-Driven Visual Synthesis for Portrait Image Animation

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-10T11:15:10.394608Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T11:13:27.689539Z digest=sha256:d22db69ebea0d21187b0eba00b5ca93c367eca10beceeaff86ff1e6d7886f0c9

Observation c6fe6c3d-71db-4fad-9345-e425bcfd5ca5 · inbound

MMControl: Unified Multi-Modal Control for Joint Audio-Video Generation cites this paper.

MMControl: Unified Multi-Modal Control for Joint Audio-Video Generation Hallo: Hierarchical Audio-Driven Visual Synthesis for Portrait Image Animation

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-11T12:46:28.097078Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T02:55:09.008954Z digest=sha256:7aa5d3456d68af98d1d9a251b5a5c00ab7fb937675a5939c070d4eac2ca10f5f

Observation c2658e7b-c65f-431d-be8b-547f74962c57 · inbound

MeshLAM: Feed-Forward One-Shot Animatable Textured Mesh Avatar Reconstruction cites this paper.

MeshLAM: Feed-Forward One-Shot Animatable Textured Mesh Avatar Reconstruction Hallo: Hierarchical Audio-Driven Visual Synthesis for Portrait Image Animation

Reference 68

Resolution
verified exact
arxiv_id, observed 2026-05-11T14:16:27.786083Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-09T22:06:58.480052Z digest=sha256:c86bb5f9aea6fd802bf72cd6c2dcc3125ff86147079d82fa7bdaa7c0998427f7

Observation 6178aacc-2956-4397-afa6-f704ee068401 · inbound

Talker-T2AV: Joint Talking Audio-Video Generation with Autoregressive Diffusion Modeling cites this paper.

Talker-T2AV: Joint Talking Audio-Video Generation with Autoregressive Diffusion Modeling Hallo: Hierarchical Audio-Driven Visual Synthesis for Portrait Image Animation

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-11T21:06:14.299885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-08T06:44:36.000353Z digest=sha256:70ad6cc4ebb1e1ac72bcae56627a1081b4c14ed3707cbd130b4d39815533171e

Observation f0d3deec-29e2-4350-b3ab-4b1dd035a100 · inbound

Hallo-Live: Real-Time Streaming Joint Audio-Video Avatar Generation with Asynchronous Dual-Stream and Human-Centric Preference Distillation cites this paper.

Hallo-Live: Real-Time Streaming Joint Audio-Video Avatar Generation with Asynchronous Dual-Stream and Human-Centric Preference Distillation Hallo: Hierarchical Audio-Driven Visual Synthesis for Portrait Image Animation

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-05-11T21:06:12.720703Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-08T06:56:19.795651Z digest=sha256:6933414d3c2877206ac2a1e9ae67ab168aeecd05dfece2da6bd37de805044dbb

Observation a13e0a97-4797-4d0e-95fa-2842b1557e0b · inbound

Do Protective Perturbations Really Protect Portrait Privacy under Real-world Image Transformations? cites this paper.

Do Protective Perturbations Really Protect Portrait Privacy under Real-world Image Transformations? Hallo: Hierarchical Audio-Driven Visual Synthesis for Portrait Image Animation

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-11T21:06:14.711337Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-08T06:42:31.456648Z digest=sha256:d6e034440ca2212fb3914907c4b407d55af609517adb87bc2cd49a39e8b85c85

Observation 91275adb-8fef-4224-9e85-f1acd0d5c7ad · inbound

Generate Your Talking Avatar from Video Reference cites this paper.

Generate Your Talking Avatar from Video Reference Hallo: Hierarchical Audio-Driven Visual Synthesis for Portrait Image Animation

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-05-12T10:36:29.833106Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-07T05:32:04.519820Z digest=sha256:f0c560b5a57ea5a3af0f156731ba1d5d672a80d5f9be35ae275288fe8377ceaa

Observation bd95ef94-b0d9-4823-b202-c18502e0671e · inbound

AsymTalker: Identity-Consistent Long-Term Talking Head Generation via Asymmetric Distillation cites this paper.

AsymTalker: Identity-Consistent Long-Term Talking Head Generation via Asymmetric Distillation Hallo: Hierarchical Audio-Driven Visual Synthesis for Portrait Image Animation

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-11T15:16:11.240985Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-09T20:19:39.565157Z digest=sha256:aef25c3b6271a1556cb73da4a7ce3c39f267202ef46af52621950f7e6c800fe8

Observation 93601e6a-008c-4ed1-b8af-6296ce99a58b · inbound

AsymTalker: Identity-Consistent Long-Term Talking Head Generation via Asymmetric Distillation cites this paper.

AsymTalker: Identity-Consistent Long-Term Talking Head Generation via Asymmetric Distillation Hallo: Hierarchical Audio-Driven Visual Synthesis for Portrait Image Animation

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-11T03:45:57.523880Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-11T02:18:58.996355Z digest=sha256:95c532eb828d0717cb265e540bb3d2d86551de4dacc95d4959fbc14ca63acc20

Observation 09b818d7-6573-4f55-936e-f1ba2c36e39e · inbound

Eulerian Motion Guidance: Robust Image Animation via Bidirectional Geometric Consistency cites this paper.

Eulerian Motion Guidance: Robust Image Animation via Bidirectional Geometric Consistency Hallo: Hierarchical Audio-Driven Visual Synthesis for Portrait Image Animation

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-11T18:51:08.915687Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-08T13:30:47.547011Z digest=sha256:6ec4d472236eca6316368557b61c03bdeb2325d02a996bf8d3fd8075e8fdb3a4

Observation 86e8d2fc-022b-4df2-8423-3a3f04da3434 · inbound

Eulerian Motion Guidance: Robust Image Animation via Bidirectional Geometric Consistency cites this paper.

Eulerian Motion Guidance: Robust Image Animation via Bidirectional Geometric Consistency Hallo: Hierarchical Audio-Driven Visual Synthesis for Portrait Image Animation

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-11T03:55:54.099448Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-11T02:10:23.616363Z digest=sha256:787f2fc5697001768725f97f55f2a6054b5d1e968d43780645c750a38cba6845

Observation cae1bfc1-313a-4c5a-abd6-e0c6a72ec99c · inbound

Eulerian Motion Guidance: Robust Image Animation via Bidirectional Geometric Consistency cites this paper.

Eulerian Motion Guidance: Robust Image Animation via Bidirectional Geometric Consistency Hallo: Hierarchical Audio-Driven Visual Synthesis for Portrait Image Animation

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-13T07:02:27.640716Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-13T06:58:57.520289Z digest=sha256:cae8cabb940adb08c47619d432fa18476eef65d2f267d1860b5cf99628958e9d

Observation 92272046-ae8d-4b36-a705-39063cb981dd · inbound

Eulerian Motion Guidance: Robust Image Animation via Bidirectional Geometric Consistency cites this paper.

Eulerian Motion Guidance: Robust Image Animation via Bidirectional Geometric Consistency Hallo: Hierarchical Audio-Driven Visual Synthesis for Portrait Image Animation

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-07-01T13:25:44.775363Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-30T23:25:13.965611Z digest=sha256:576b43aa86e32222dd511a38e26937af33732ee359e71270560897f701c7dfbf

Observation 10a4157b-6876-4985-9995-0de49e6314b6 · inbound

HighSync: High-Quality Lip Synchronization via Latent Diffusion Models cites this paper.

HighSync: High-Quality Lip Synchronization via Latent Diffusion Models Hallo: Hierarchical Audio-Driven Visual Synthesis for Portrait Image Animation

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-19T21:17:48.650577Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-19T21:14:25.606097Z digest=sha256:46419a3afa02c31c6aec030cea85473a12e11f861cd42e1799bb55f3052ea665

Observation 08535c7e-7e6c-4f05-835b-ab2fd1ed2a52 · inbound

Image-to-Video Diffusion: From Foundations to Open Frontiers cites this paper.

Image-to-Video Diffusion: From Foundations to Open Frontiers Hallo: Hierarchical Audio-Driven Visual Synthesis for Portrait Image Animation

Reference 73

Resolution
verified exact
arxiv_id, observed 2026-05-20T15:08:24.871201Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-20T15:06:02.084336Z digest=sha256:31ec770b6c2f37c62f68d80b2e177b3c0719eb146d6e1c0b27d5afdef90fe464

Observation f34ed0c2-6d49-439c-9a29-b38f52266d0d · inbound

Loki: Representation over Architecture for Diffusion-Based Portrait Animation cites this paper.

Loki: Representation over Architecture for Diffusion-Based Portrait Animation Hallo: Hierarchical Audio-Driven Visual Synthesis for Portrait Image Animation

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-06-30T15:24:49.848910Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-30T15:23:48.899214Z digest=sha256:124c9ad2351ed2416c79bdaf8d04413e63cce401aa5272f176256579664dde93

Observation 4d58aa8a-0465-4083-80dd-4555587343bd · inbound

Test-Time Self-Adaptive Conditioning for Stable Audio-Driven Talking-Head Generation cites this paper.

Test-Time Self-Adaptive Conditioning for Stable Audio-Driven Talking-Head Generation Hallo: Hierarchical Audio-Driven Visual Synthesis for Portrait Image Animation

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-06-29T22:44:01.687595Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-29T22:36:53.138354Z digest=sha256:a496efeeba86187d0f408238bb24cc3f9bfd0699785727c84c59b08b00aedb34

Observation a89fcbd1-5187-4cb5-abf3-3dc878572ccf · inbound

IP-Adapter Is All You Need: Towards Fine-Tuning-Free Diffusion-Based Talking Face Generation cites this paper.

IP-Adapter Is All You Need: Towards Fine-Tuning-Free Diffusion-Based Talking Face Generation Hallo: Hierarchical Audio-Driven Visual Synthesis for Portrait Image Animation

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-06-29T07:53:13.366896Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-29T07:50:56.947671Z digest=sha256:011acea27b2b702eca0fc84eb86dee094e7045ce79bb696147cacc065b1d3a9f

Observation 210e1107-9f6b-4bd8-b811-d59d4911bc8e · inbound

Temporally-Aligned Evaluation for Audio-Driven Talking Head Generation cites this paper.

Temporally-Aligned Evaluation for Audio-Driven Talking Head Generation Hallo: Hierarchical Audio-Driven Visual Synthesis for Portrait Image Animation

Reference 43

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T21:46:15.588139Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-28T16:19:55.048831Z digest=sha256:11a6845e72c1efd428089e4f65aa2296a2a2bbe8d9be8988bec976e2535d7d7c

Observation 1a0bf6b9-050b-4aae-9500-3c2f8ad4609f · inbound

Real-Time Generation of Streamable Talking Portrait Video with Reference-Guided Deep Compression VAEs cites this paper.

Real-Time Generation of Streamable Talking Portrait Video with Reference-Guided Deep Compression VAEs Hallo: Hierarchical Audio-Driven Visual Synthesis for Portrait Image Animation

Reference 58

Resolution
verified exact
arxiv_id, observed 2026-07-01T22:06:16.998665Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-28T15:41:59.683979Z digest=sha256:8f1cd4532832cad8b51221fd780b5d983359fb6c40f9ff58e58343e8f0a3711b

Observation 9344f1e9-3dae-430c-a33d-06ccf449486f · inbound

SyncCache: Exploiting Asymmetric Dynamics for Fast Audio-Driven Portrait Animation cites this paper.

SyncCache: Exploiting Asymmetric Dynamics for Fast Audio-Driven Portrait Animation Hallo: Hierarchical Audio-Driven Visual Synthesis for Portrait Image Animation

Reference 40

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T09:45:40.302530Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-01T06:13:11.504526Z digest=sha256:a3ff28a8c3d88cc1ed84808be35c0c353a7fef376e7e0b508c5b90b3cdb81841

Observation b493bb94-9218-4ea3-8a09-b0dfb81ca557 · inbound

FA-LAM: Focus-Aware Large Avatar Model for One-Shot 4D Animatable Gaussian Head cites this paper.

FA-LAM: Focus-Aware Large Avatar Model for One-Shot 4D Animatable Gaussian Head Hallo: Hierarchical Audio-Driven Visual Synthesis for Portrait Image Animation

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-01T09:02:19.386267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T09:02:19.386267Z digest=sha256:a13a76bc198629128fc20b349fbb732ef0f83348015eed28ed492bc1a2edfcf8

Observation 6038c66d-f163-47bc-bfd3-a30d05155090 · inbound

ViDS: Video Diffusion Shader using 3D Face Tracking cites this paper.

ViDS: Video Diffusion Shader using 3D Face Tracking Hallo: Hierarchical Audio-Driven Visual Synthesis for Portrait Image Animation

Reference 75

Resolution
unresolved
no resolver link, observed 2026-07-31T23:01:32.039915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T23:01:32.039915Z digest=sha256:635cedbb728df2925918ee6d7d9b166afab22347abac5884b238bca04dcff0a6

Observation 44928e9b-9bca-4652-8a70-7e1e460e7b88 · inbound

TaoMate: Anchor-Guided Memory Bridging Evolving and Reference States for Real-Time Audio-Video Digital Human Generation cites this paper.

TaoMate: Anchor-Guided Memory Bridging Evolving and Reference States for Real-Time Audio-Video Digital Human Generation Hallo: Hierarchical Audio-Driven Visual Synthesis for Portrait Image Animation

Reference 48

Resolution
unresolved
no resolver link, observed 2026-07-31T17:08:14.198579Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T17:08:14.198579Z digest=sha256:daac3580bcd065d97cc80317bc93038085dbe432d16ccaaf2195b41f4ad34445

Observation 8192666c-149d-491b-9cfc-40000194f655 · inbound

Geometry-guided Emotion Modulation for Controllable and Photorealistic Emotional Talking Face Generation cites this paper.

Geometry-guided Emotion Modulation for Controllable and Photorealistic Emotional Talking Face Generation Hallo: Hierarchical Audio-Driven Visual Synthesis for Portrait Image Animation

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-04T01:03:13.201630Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:03:13.201630Z digest=sha256:c081b8ed6c1b6f8b0fa6b1e919cc0e5d7043b0fc6469637bf114fd546d3f1e20

Observation b22aefd0-67be-4417-bbb4-4566ef4c7540 · inbound

EmpaAva: An Open-source Agentic 3D-Avatar Empathetic Live Chatbot cites this paper.

EmpaAva: An Open-source Agentic 3D-Avatar Empathetic Live Chatbot Hallo: Hierarchical Audio-Driven Visual Synthesis for Portrait Image Animation

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T18:28:59.349570Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:28:59.349570Z digest=sha256:22675842ff68a46a88b0da78fc481a82cd6484896b6cedc74afad1c155706326