Pith. sign in

Paper Citation Record · LEDGER

AniPortrait: Audio-Driven Synthesis of Photorealistic Portrait Animation

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 58 inbound Pith citation observations for arXiv:2403.17694.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2403.17694 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 58 of 58 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 58 of 58 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T21:11:25.341871Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T02:16:26.575398Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 01855afa-757d-419b-9cb5-40e207cc27d3 · inbound

JoyVASA: Portrait and Animal Image Animation with Diffusion-Based Audio-Driven Facial Dynamics and Head Motion Generation cites this paper.

JoyVASA: Portrait and Animal Image Animation with Diffusion-Based Audio-Driven Facial Dynamics and Head Motion Generation AniPortrait: Audio-Driven Synthesis of Photorealistic Portrait Animation

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-23T17:03:12.558326Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-23T16:59:57.727035Z digest=sha256:844482ee526540d5c3f6fa6520e4a0a0282b1df4cea58b48591af3d0bdecb90a

Observation a0475f1b-34d4-4481-a77b-500538431ca7 · inbound

Multimodal Diffusion Transformer with Memory Bank for Scalable Long-Duration Talking Video Generation cites this paper.

Multimodal Diffusion Transformer with Memory Bank for Scalable Long-Duration Talking Video Generation AniPortrait: Audio-Driven Synthesis of Photorealistic Portrait Animation

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-23T17:23:15.361627Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-23T17:19:51.411937Z digest=sha256:86f6f7b6dd89e7f7e3fa110fe3f07b489806d0a075ce5a8cc0a195f82544d21f

Observation 88ec1c50-ff40-442b-8334-1dada8c4bcc6 · inbound

Long-Term TalkingFace Generation via Motion-Prior Conditional Diffusion Model cites this paper.

Long-Term TalkingFace Generation via Motion-Prior Conditional Diffusion Model AniPortrait: Audio-Driven Synthesis of Photorealistic Portrait Animation

Reference 2004

Resolution
unresolved
no resolver link, observed 2026-08-07T21:11:25.341871Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T21:11:25.341871Z digest=sha256:88afa82781ea52944915c7897145abc6396993da535dc774a0246a090b03067c

Observation 1f0bcfd1-d675-4f87-ae55-6c336198b589 · inbound

Exploring Timeline Control for Facial Motion Generation cites this paper.

Exploring Timeline Control for Facial Motion Generation AniPortrait: Audio-Driven Synthesis of Photorealistic Portrait Animation

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T13:49:52.852981Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:49:52.852981Z digest=sha256:ab0ef451219d63a556d51dc605e91493224d457fe41b440ebeb531c20033004b

Observation 0d178024-8051-4ec1-a070-bf1c516af756 · inbound

FaceEditTalker: Controllable Talking Head Generation with Facial Attribute Editing cites this paper.

FaceEditTalker: Controllable Talking Head Generation with Facial Attribute Editing AniPortrait: Audio-Driven Synthesis of Photorealistic Portrait Animation

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T13:21:46.497816Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:21:46.497816Z digest=sha256:ec11e9223def77b56140121cc226fb77c1b6bdb8bed4e7c643cc4f8286c4a6f0

Observation b146ccef-7509-4d14-9277-f7205658eaec · inbound

Let Them Talk: Audio-Driven Multi-Person Conversational Video Generation cites this paper.

Let Them Talk: Audio-Driven Multi-Person Conversational Video Generation AniPortrait: Audio-Driven Synthesis of Photorealistic Portrait Animation

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T13:07:34.059612Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:07:34.059612Z digest=sha256:38656b6271c532a66005e5e058e1b473a34a0a268733fd2b5a8a12e4f42bdc05

Observation 67cedca9-14dc-44f2-b242-aff70eb235b5 · inbound

SkyReels-Audio: Omni Audio-Conditioned Talking Portraits in Video Diffusion Transformers cites this paper.

SkyReels-Audio: Omni Audio-Conditioned Talking Portraits in Video Diffusion Transformers AniPortrait: Audio-Driven Synthesis of Photorealistic Portrait Animation

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-07T12:01:08.767656Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:01:08.767656Z digest=sha256:4eb33d89c050df1f232e4fcb614bf19478e2e54c000ff0d7288ceadbe6f28716

Observation 9ac92218-4b33-48cd-bef9-b7b07723083f · inbound

Speaking images. A novel framework for the automated self-description of artworks cites this paper.

Speaking images. A novel framework for the automated self-description of artworks AniPortrait: Audio-Driven Synthesis of Photorealistic Portrait Animation

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-07T13:17:56.065494Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:17:56.065494Z digest=sha256:8e197e96907e3b3075313ad1d4a2bd0a53db15e0183a998e9fb9298120b3d5aa

Observation 8dce6ce4-191d-456a-94f1-149be2ffcbab · inbound

LLIA -- Enabling Low-Latency Interactive Avatars: Real-Time Audio-Driven Portrait Video Generation with Diffusion Models cites this paper.

LLIA -- Enabling Low-Latency Interactive Avatars: Real-Time Audio-Driven Portrait Video Generation with Diffusion Models AniPortrait: Audio-Driven Synthesis of Photorealistic Portrait Animation

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T10:18:24.450551Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:18:24.450551Z digest=sha256:963ce44915fdf9a4223b5908c4909f8443c06a1634003bb48ab0bf002a1c3d10

Observation c3a1929b-d64c-4a7e-b2a7-33342cbf61a2 · inbound

Audio-Sync Video Generation with Multi-Stream Temporal Control cites this paper.

Audio-Sync Video Generation with Multi-Stream Temporal Control AniPortrait: Audio-Driven Synthesis of Photorealistic Portrait Animation

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T05:24:30.408110Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:24:30.408110Z digest=sha256:c1ea5bee8aaa75ce63493d53e7ef4b32392931bfd0d2c7435436d5dff731f091

Observation e7768a50-d62b-4394-bd18-c1521d85f809 · inbound

Controllable and Expressive One-Shot Video Head Swapping cites this paper.

Controllable and Expressive One-Shot Video Head Swapping AniPortrait: Audio-Driven Synthesis of Photorealistic Portrait Animation

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-06T23:42:37.820190Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:42:37.820190Z digest=sha256:18a1627ebbe151c38732652b92ae813950a0e1f47e1f8e4e6a31ddbce1126d09

Observation 042642b1-40a0-4eb1-a81b-f1ab250bac94 · inbound

JAM-Flow: Joint Audio-Motion Synthesis with Flow Matching cites this paper.

JAM-Flow: Joint Audio-Motion Synthesis with Flow Matching AniPortrait: Audio-Driven Synthesis of Photorealistic Portrait Animation

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-22T00:50:51.067493Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-22T00:46:39.196042Z digest=sha256:aa58f62607f332453b1ac629523686899e57a0fabf34ef889249eebac4d64bfa

Observation b3eff734-9988-4b39-90e7-632f17a4d897 · inbound

FixTalk: Taming Identity Leakage for High-Quality Talking Head Generation in Extreme Cases cites this paper.

FixTalk: Taming Identity Leakage for High-Quality Talking Head Generation in Extreme Cases AniPortrait: Audio-Driven Synthesis of Photorealistic Portrait Animation

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-06T20:58:30.336835Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:58:30.336835Z digest=sha256:383f8f54e4a00ce13a723f98eae30d9a087c7e597b70e0e1ecd511937990dcce

Observation 4654cc85-ca5c-4d2b-8181-5c1c68606216 · inbound

MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation cites this paper.

MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation AniPortrait: Audio-Driven Synthesis of Photorealistic Portrait Animation

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T19:38:20.556209Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:38:20.556209Z digest=sha256:e3d314ad30236bab0aea8b1f8e3ec9865cde9039d69abd45ec2e9f5546b9c0d4

Observation 2b7a5166-c530-4d93-928f-bb09d9a70d22 · inbound

Democratizing High-Fidelity Co-Speech Gesture Video Generation cites this paper.

Democratizing High-Fidelity Co-Speech Gesture Video Generation AniPortrait: Audio-Driven Synthesis of Photorealistic Portrait Animation

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T19:00:03.247582Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:00:03.247582Z digest=sha256:3362825de82767b10122e246ae206cb0ca66be2de8d87cc147a3e420b3ff1e1d

Observation 8e5fc1b8-01d5-4dad-9028-01b91495ccec · inbound

HairShifter: Consistent and High-Fidelity Video Hair Transfer via Anchor-Guided Animation cites this paper.

HairShifter: Consistent and High-Fidelity Video Hair Transfer via Anchor-Guided Animation AniPortrait: Audio-Driven Synthesis of Photorealistic Portrait Animation

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T16:44:38.865416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:44:38.865416Z digest=sha256:6362a7b6d1beb9b363c5a7902816a051cac7bb766a2f8dab6d17a760f00e7e14

Observation 451ecb1f-5204-42bb-bf13-a7b8433af352 · inbound

MagicAnime: A Hierarchically Annotated, Multimodal and Multitasking Dataset with Benchmarks for Cartoon Animation Generation cites this paper.

MagicAnime: A Hierarchically Annotated, Multimodal and Multitasking Dataset with Benchmarks for Cartoon Animation Generation AniPortrait: Audio-Driven Synthesis of Photorealistic Portrait Animation

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T13:37:09.367614Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:37:09.367614Z digest=sha256:25fa361301fdb608bbd1280efbf1d5d9e3feaa20659c62720bf4232a07163451

Observation 57b2f6d9-df91-483c-b7ec-79d5c4258cc8 · inbound

X-NeMo: Expressive Neural Motion Reenactment via Disentangled Latent Attention cites this paper.

X-NeMo: Expressive Neural Motion Reenactment via Disentangled Latent Attention AniPortrait: Audio-Driven Synthesis of Photorealistic Portrait Animation

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T11:03:44.884845Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:03:44.884845Z digest=sha256:6d859fb1e59d3056d4e56131c1a072e0537b8961dcfc895e9baf243b768ffd36

Observation 247b5a6b-0495-44f0-bebb-21bf2c18ce6e · inbound

Text2Lip: Progressive Lip-Synced Talking Face Generation from Text via Viseme-Guided Rendering cites this paper.

Text2Lip: Progressive Lip-Synced Talking Face Generation from Text via Viseme-Guided Rendering AniPortrait: Audio-Driven Synthesis of Photorealistic Portrait Animation

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T05:05:49.133791Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:05:49.133791Z digest=sha256:f50dedcbdfe2b7f4a8ee10bdda4b8dbb995eab0a16685ec6fe1d7c90bc6c316a

Observation f0271651-de71-42ed-baa9-1453bd3231fa · inbound

PoseGuard: Pose-Guided Generation with Safety Guardrails cites this paper.

PoseGuard: Pose-Guided Generation with Safety Guardrails AniPortrait: Audio-Driven Synthesis of Photorealistic Portrait Animation

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T05:02:04.490234Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T05:02:04.490234Z digest=sha256:ca86ca7b89139daa8e6882a32faf2cb51a2fd3f4c89a540d89b4d5c5bbe89fd2

Observation 499289a0-6f66-421d-af3e-61575593d220 · inbound

DiTalker: A Unified DiT-based Framework for High-Quality and Speaking Styles Controllable Portrait Animation cites this paper.

DiTalker: A Unified DiT-based Framework for High-Quality and Speaking Styles Controllable Portrait Animation AniPortrait: Audio-Driven Synthesis of Photorealistic Portrait Animation

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T12:43:23.684599Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:43:23.684599Z digest=sha256:9beda31f46becc79c01c37a0f4ccd7df53b552b6fe5e809166ee7dbd1b5191fd

Observation b0f0a0fb-6853-4d42-82b3-bdd80fcd9c44 · inbound

LaVieID: Local Autoregressive Diffusion Transformers for Identity-Preserving Video Creation cites this paper.

LaVieID: Local Autoregressive Diffusion Transformers for Identity-Preserving Video Creation AniPortrait: Audio-Driven Synthesis of Photorealistic Portrait Animation

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-05T22:04:35.670348Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T22:04:35.670348Z digest=sha256:51e1d4641327513ad5be5a38a7316eb5f7827182d0c6ed45f58a58d7db273e7c

Observation e80c8543-d238-47bf-8d0b-1784c00321a0 · inbound

FantasyTalking2: Timestep-Layer Adaptive Preference Optimization for Audio-Driven Portrait Animation cites this paper.

FantasyTalking2: Timestep-Layer Adaptive Preference Optimization for Audio-Driven Portrait Animation AniPortrait: Audio-Driven Synthesis of Photorealistic Portrait Animation

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-05T20:06:49.448579Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:06:49.448579Z digest=sha256:9417ddad1c361c741e0b83b44f215196a7e4ea0c0641768a784257fc946fcbaa

Observation 4023f325-6310-4ae0-81b6-c88c7ac8add9 · inbound

EDTalk++: Full Disentanglement for Controllable Talking Head Synthesis cites this paper.

EDTalk++: Full Disentanglement for Controllable Talking Head Synthesis AniPortrait: Audio-Driven Synthesis of Photorealistic Portrait Animation

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-05T19:07:36.310450Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:07:36.310450Z digest=sha256:5d50dc8f199ebb9bf039613a36318dece0e05ff400d4d0c9aa51c3628ea8e59a

Observation 579ea692-c4df-4f52-9229-a7b404e6c94e · inbound

InfiniteTalk: Audio-driven Video Generation for Sparse-Frame Video Dubbing cites this paper.

InfiniteTalk: Audio-driven Video Generation for Sparse-Frame Video Dubbing AniPortrait: Audio-Driven Synthesis of Photorealistic Portrait Animation

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-05T18:50:16.290012Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T18:50:16.290012Z digest=sha256:b5af2a2a9b41bab9512399fac42c7c70a5b027c29e740a77c3d32956d35d6a1a

Observation 2728a245-1ee2-4b8b-9c40-ab9ca16ade0a · inbound

InfinityHuman: Towards Long-Term Audio-Driven Human cites this paper.

InfinityHuman: Towards Long-Term Audio-Driven Human AniPortrait: Audio-Driven Synthesis of Photorealistic Portrait Animation

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-05T15:19:37.244894Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T15:19:37.244894Z digest=sha256:2b3f8a639ab5dcb8b4d3e53d592435dc1bbb606d2dc76d8741f259cb93cedf91

Observation d021f543-f425-431b-8fc4-9e227d662046 · inbound

Human Motion Video Generation: A Survey cites this paper.

Human Motion Video Generation: A Survey AniPortrait: Audio-Driven Synthesis of Photorealistic Portrait Animation

Reference 137

Resolution
unresolved
no resolver link, observed 2026-08-05T10:36:57.303214Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T10:36:57.303214Z digest=sha256:b42272b5997e87cf9f9847e7c4719d26a93eb797199b324d53047fd2d43e2d84

Observation e89d7844-db38-42da-8e1b-1700a9abf86a · inbound

FluentAvatar: Flicker-Free Talking-Head Animation via Phoneme-Guided Autoregressive Modeling cites this paper.

FluentAvatar: Flicker-Free Talking-Head Animation via Phoneme-Guided Autoregressive Modeling AniPortrait: Audio-Driven Synthesis of Photorealistic Portrait Animation

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-18T16:42:43.813294Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-18T16:42:25.803856Z digest=sha256:67a6972f2c39c76eca4bd8da1dd7d43b1bcfffd4c31b3d3b1143f213ba215584

Observation ff743ba2-f1ba-4f2e-a92c-58fd02cb41be · inbound

STARCaster: Spatio-Temporal AutoRegressive Video Diffusion for Identity- and View-Aware Talking Portraits cites this paper.

STARCaster: Spatio-Temporal AutoRegressive Video Diffusion for Identity- and View-Aware Talking Portraits AniPortrait: Audio-Driven Synthesis of Photorealistic Portrait Animation

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-03T16:30:39.321612Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:30:39.321612Z digest=sha256:3b9ccd3e1cfa11b440497b57ccaf09ac6690cd31c9cf78cc7ac2059b361c0964

Observation be3101ce-5b9a-497a-a0f2-5d72983df5a3 · inbound

ViBES: A Conversational Agent with Behaviorally-Intelligent 3D Virtual Body cites this paper.

ViBES: A Conversational Agent with Behaviorally-Intelligent 3D Virtual Body AniPortrait: Audio-Driven Synthesis of Photorealistic Portrait Animation

Reference 113

Resolution
verified exact
arxiv_id, observed 2026-05-16T22:08:36.327070Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T22:04:07.403410Z digest=sha256:a48a7da59be4a287068c5525ff89563d35a94839659581c9bdca692cfd8b3b6a

Observation 01552f3e-7d80-4027-934f-60e289d09be6 · inbound

Instant Expressive Gaussian Head Avatars at Over 100 FPS cites this paper.

Instant Expressive Gaussian Head Avatars at Over 100 FPS AniPortrait: Audio-Driven Synthesis of Photorealistic Portrait Animation

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-03T15:28:59.211811Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T15:28:59.211811Z digest=sha256:4b1959457aa20ae80f94d3fa2eb21cf3ee81434f80a41e36d62a2d935843bfb9

Observation 75c95618-e92d-4942-898b-04cdcaf3db5c · inbound

UIKA: Fast Universal Head Avatar from Pose-Free Images cites this paper.

UIKA: Fast Universal Head Avatar from Pose-Free Images AniPortrait: Audio-Driven Synthesis of Photorealistic Portrait Animation

Reference 79

Resolution
verified exact
arxiv_id, observed 2026-05-22T12:04:51.573998Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-22T12:02:27.828968Z digest=sha256:836b74947cde6756daafff4779def216e5c1e11e09b864616fe888dba9498a7d

Observation ef2d74b1-b261-492d-934c-d8f159487fab · inbound

EchoTorrent: Towards Swift, Sustained, and Streaming Multi-Modal Video Generation cites this paper.

EchoTorrent: Towards Swift, Sustained, and Streaming Multi-Modal Video Generation AniPortrait: Audio-Driven Synthesis of Photorealistic Portrait Animation

Reference 86

Resolution
verified exact
arxiv_id, observed 2026-05-15T22:20:22.464159Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-15T22:20:16.320171Z digest=sha256:02547b99f307c59545225157aa6b3ca128fe188acff6e40ee448f0dbae3bcc15

Observation 1b11bd3d-be73-40cd-916d-4bf807868b56 · inbound

AHOY! Animatable Humans under Occlusion from YouTube Videos with Gaussian Splatting and Video Diffusion Priors cites this paper.

AHOY! Animatable Humans under Occlusion from YouTube Videos with Gaussian Splatting and Video Diffusion Priors AniPortrait: Audio-Driven Synthesis of Photorealistic Portrait Animation

Reference 129

Resolution
unresolved
no resolver link, observed 2026-07-13T22:49:03.259461Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T22:49:03.259461Z digest=sha256:e1e16e1fe17d6dbdc0b8bac6d720e42fcdc9d29c8938e6fd1b1b374dbf9b6dbe

Observation 6b77c536-6fe4-4752-b714-7494b8aa2b33 · inbound

AvatarPointillist: AutoRegressive 4D Gaussian Avatarization cites this paper.

AvatarPointillist: AutoRegressive 4D Gaussian Avatarization AniPortrait: Audio-Driven Synthesis of Photorealistic Portrait Animation

Reference 67

Resolution
verified exact
arxiv_id, observed 2026-05-10T23:55:51.059816Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T18:47:37.605996Z digest=sha256:3c81b031a220a585221a3ba408d2d2401ef234c29e857fe231807f32b1e15b14

Observation 869f6de3-8dd9-4338-8b44-166d6aa071da · inbound

SyncBreaker:Stage-Aware Multimodal Adversarial Attacks on Audio-Driven Talking Head Generation cites this paper.

SyncBreaker:Stage-Aware Multimodal Adversarial Attacks on Audio-Driven Talking Head Generation AniPortrait: Audio-Driven Synthesis of Photorealistic Portrait Animation

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-05-11T06:41:25.219400Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T17:30:36.462567Z digest=sha256:147e28598fad154c40c383eb579741eb3bd61062803e87f8924c22436fdac998

Observation 84eddbc5-c482-4812-a9d6-7c416a6dbf5b · inbound

SyncBreaker:Stage-Aware Multimodal Adversarial Attacks on Audio-Driven Talking Head Generation cites this paper.

SyncBreaker:Stage-Aware Multimodal Adversarial Attacks on Audio-Driven Talking Head Generation AniPortrait: Audio-Driven Synthesis of Photorealistic Portrait Animation

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-02T16:38:16.290404Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T16:38:16.290404Z digest=sha256:5628ec4f840d63fc5492f17b72cc548d1b2281a78c6a6d7b0b2885ed864f61de

Observation c194592f-4807-4c2f-931d-091a5e279976 · inbound

PianoFlow: Music-Aware Streaming Piano Motion Generation with Bimanual Coordination cites this paper.

PianoFlow: Music-Aware Streaming Piano Motion Generation with Bimanual Coordination AniPortrait: Audio-Driven Synthesis of Photorealistic Portrait Animation

Reference 64

Resolution
verified exact
arxiv_id, observed 2026-05-11T09:41:00.065751Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T15:54:30.725820Z digest=sha256:27c68f0beab83bc91e390d21de406fa8801de14c104c245af3b96a0e4629b47e

Observation e4cc6ec4-2ecb-41f5-a7de-4b26bff0e8cf · inbound

Giving Faces Their Feelings Back: Explicit Emotion Control for Feedforward Single-Image 3D Head Avatars cites this paper.

Giving Faces Their Feelings Back: Explicit Emotion Control for Feedforward Single-Image 3D Head Avatars AniPortrait: Audio-Driven Synthesis of Photorealistic Portrait Animation

Reference 78

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T11:50:20.281786Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-10T11:47:42.039240Z digest=sha256:ad0fc10af08efca585f00cb16d008658b3095b4778b8ddcde5a53a456643b3c5

Observation edf47450-dea3-49f0-ac3b-54d3b7f004f0 · inbound

TurboTalk: Progressive Distillation for One-Step Audio-Driven Talking Avatar Generation cites this paper.

TurboTalk: Progressive Distillation for One-Step Audio-Driven Talking Avatar Generation AniPortrait: Audio-Driven Synthesis of Photorealistic Portrait Animation

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-10T11:15:10.370873Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T11:13:27.689539Z digest=sha256:7b35455375d1ed1b091b3ae7bc25507fd1546ad5059f877f48e61de78f9eb919

Observation 00aa68bb-6a1b-4202-a546-18bad7f72798 · inbound

PortraitDirector: A Hierarchical Disentanglement Framework for Controllable and Real-time Facial Reenactment cites this paper.

PortraitDirector: A Hierarchical Disentanglement Framework for Controllable and Real-time Facial Reenactment AniPortrait: Audio-Driven Synthesis of Photorealistic Portrait Animation

Reference 59

Resolution
verified exact
arxiv_id, observed 2026-05-11T12:31:04.133105Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T03:30:18.645845Z digest=sha256:3c02b16feccd5f3ad9155a91cef0919c0692a7965b6436ec654c50e0240ac922

Observation 9d3dab7b-01be-4261-b5db-b5f24d0e5bc0 · inbound

MMControl: Unified Multi-Modal Control for Joint Audio-Video Generation cites this paper.

MMControl: Unified Multi-Modal Control for Joint Audio-Video Generation AniPortrait: Audio-Driven Synthesis of Photorealistic Portrait Animation

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-11T12:46:28.066848Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T02:55:09.008954Z digest=sha256:2ce0e74c81b78a4545bd8d8cdaf2539b4d208583794aa5c6b91a7ca8c8d16d61

Observation 9f3a0d5e-ad8e-414e-bf73-ac6bf6b8856e · inbound

Omni-Fake: Benchmarking Unified Multimodal Social Media Deepfake Detection cites this paper.

Omni-Fake: Benchmarking Unified Multimodal Social Media Deepfake Detection AniPortrait: Audio-Driven Synthesis of Photorealistic Portrait Animation

Reference 83

Resolution
verified exact
arxiv_id, observed 2026-05-11T17:06:04.288003Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-09T14:04:52.065878Z digest=sha256:f8d9f585c3c512b80913e45f36fc7cb03d3a696f5ae570d1cdd30593e1671126

Observation b429ade0-fbb9-496e-8c84-a4209b84db4a · inbound

AsymTalker: Identity-Consistent Long-Term Talking Head Generation via Asymmetric Distillation cites this paper.

AsymTalker: Identity-Consistent Long-Term Talking Head Generation via Asymmetric Distillation AniPortrait: Audio-Driven Synthesis of Photorealistic Portrait Animation

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-11T15:16:11.234958Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-09T20:19:39.565157Z digest=sha256:f380cdab70e5da9a698ae880914384b6cd50b623c6a3f57a5dc7f531c26eea32

Observation 8671e9e2-b04c-4b2b-89e0-90fa9bde142e · inbound

AsymTalker: Identity-Consistent Long-Term Talking Head Generation via Asymmetric Distillation cites this paper.

AsymTalker: Identity-Consistent Long-Term Talking Head Generation via Asymmetric Distillation AniPortrait: Audio-Driven Synthesis of Photorealistic Portrait Animation

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-11T03:45:57.546592Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-11T02:18:58.996355Z digest=sha256:f3478b69f7f2539c2e1cc45245754d410f68ce9981d1eaa19f332608e62366be

Observation 3c1edea4-37c8-4276-bb89-6a154abf2069 · inbound

Image-to-Video Diffusion: From Foundations to Open Frontiers cites this paper.

Image-to-Video Diffusion: From Foundations to Open Frontiers AniPortrait: Audio-Driven Synthesis of Photorealistic Portrait Animation

Reference 77

Resolution
verified exact
arxiv_id, observed 2026-05-20T15:08:24.949010Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-20T15:06:02.084336Z digest=sha256:417155b056139b664b695a13ed133a6f4dd88cfbc2b8ce8a6f7884f8f569f0d5

Observation ebfd67df-04cf-4a8c-b3bc-87646700981d · inbound

Loki: Representation over Architecture for Diffusion-Based Portrait Animation cites this paper.

Loki: Representation over Architecture for Diffusion-Based Portrait Animation AniPortrait: Audio-Driven Synthesis of Photorealistic Portrait Animation

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-06-30T15:24:49.856815Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-30T15:23:48.899214Z digest=sha256:a80bbf728b64d10ca7d3422f12a0a0f3d1269ab37609eedd2cd7c2960582e56d

Observation 4d174dec-3c35-407f-95c6-8f7806df1f53 · inbound

Test-Time Self-Adaptive Conditioning for Stable Audio-Driven Talking-Head Generation cites this paper.

Test-Time Self-Adaptive Conditioning for Stable Audio-Driven Talking-Head Generation AniPortrait: Audio-Driven Synthesis of Photorealistic Portrait Animation

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-06-29T22:44:01.748177Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-29T22:36:53.138354Z digest=sha256:cb66b2d83eae423eba730da1d20ee8661685597b776348a3df7eede816149ef8

Observation 190bd956-2db7-4ee8-a375-ab9eeb6c1bfa · inbound

CogPortrait: Fine-Grained Eye-Region Control in Portrait Animation via Hierarchical Agent Planning cites this paper.

CogPortrait: Fine-Grained Eye-Region Control in Portrait Animation via Hierarchical Agent Planning AniPortrait: Audio-Driven Synthesis of Photorealistic Portrait Animation

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-06-29T13:13:27.648129Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-29T13:03:27.312548Z digest=sha256:bf95e2d1be2cb572cf0e6151c4a909e0644a6af63fd12f752a5cd53d215a8177

Observation 35c1c419-e08c-47c2-a075-f516d9e9a118 · inbound

IP-Adapter Is All You Need: Towards Fine-Tuning-Free Diffusion-Based Talking Face Generation cites this paper.

IP-Adapter Is All You Need: Towards Fine-Tuning-Free Diffusion-Based Talking Face Generation AniPortrait: Audio-Driven Synthesis of Photorealistic Portrait Animation

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-06-29T07:53:13.369443Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-29T07:50:56.947671Z digest=sha256:906d4c509d942c3194dd12db595d3d7f2fc30718276c0bb77ede299c83383b0d

Observation 8294845f-7655-4acb-a80a-8ef13ad32a3b · inbound

Archon: A Unified Multimodal Model for Holistic Digital Human Generation cites this paper.

Archon: A Unified Multimodal Model for Holistic Digital Human Generation AniPortrait: Audio-Driven Synthesis of Photorealistic Portrait Animation

Reference 52

Resolution
verified exact
arxiv_id, observed 2026-06-29T08:03:13.783826Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-29T08:03:04.294439Z digest=sha256:b08abf91081aefa89e132340c787fc57f28731e91e06d581ae9e3c9002137894

Observation 33df6ba3-c176-445b-8002-068b531b6d8c · inbound

Temporally-Aligned Evaluation for Audio-Driven Talking Head Generation cites this paper.

Temporally-Aligned Evaluation for Audio-Driven Talking Head Generation AniPortrait: Audio-Driven Synthesis of Photorealistic Portrait Animation

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T21:46:15.580499Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-28T16:19:55.048831Z digest=sha256:fd0878c46b3192f71c5b9cb9cfb0d17a0ac0731a382ab24cb555e0f490e5bf38

Observation 679e79a2-4c15-4b52-8c6a-2414065c442f · inbound

Real-Time Generation of Streamable Talking Portrait Video with Reference-Guided Deep Compression VAEs cites this paper.

Real-Time Generation of Streamable Talking Portrait Video with Reference-Guided Deep Compression VAEs AniPortrait: Audio-Driven Synthesis of Photorealistic Portrait Animation

Reference 56

Resolution
verified exact
arxiv_id, observed 2026-07-01T22:06:17.010720Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-28T15:41:59.683979Z digest=sha256:5818385cecdc51796fc6e5b2c47c42d460c6f0cf46416f586f0254a666c3fd9d

Observation eba3d6fb-668c-474d-a03d-982536240543 · inbound

Mamba-Enhanced Implicit Motion Learning for Audio-Driven Portrait Animation cites this paper.

Mamba-Enhanced Implicit Motion Learning for Audio-Driven Portrait Animation AniPortrait: Audio-Driven Synthesis of Photorealistic Portrait Animation

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-07-02T02:16:26.577129Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-28T11:07:20.577007Z digest=sha256:0b7426b388fa17980bcb0fdb95fec11911bb31958a2a16b5bf2c67d8dd5a7385

Observation f3d16bdd-0726-4ba9-91db-abde66eb85a1 · inbound

SyncCache: Exploiting Asymmetric Dynamics for Fast Audio-Driven Portrait Animation cites this paper.

SyncCache: Exploiting Asymmetric Dynamics for Fast Audio-Driven Portrait Animation AniPortrait: Audio-Driven Synthesis of Photorealistic Portrait Animation

Reference 37

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T09:45:40.320788Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-01T06:13:11.504526Z digest=sha256:cccfc81776eadcafa6718b2d7571348b473dff5bb5fca12f55db11c90ad210ac

Observation 2656af68-61a0-4795-8024-6653e9887fee · inbound

ViDS: Video Diffusion Shader using 3D Face Tracking cites this paper.

ViDS: Video Diffusion Shader using 3D Face Tracking AniPortrait: Audio-Driven Synthesis of Photorealistic Portrait Animation

Reference 70

Resolution
unresolved
no resolver link, observed 2026-07-31T23:01:31.435147Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T23:01:31.435147Z digest=sha256:27d1d2f320b086b9a9e98112d5c0bc421fc45f0304e7a8c1b855ef5b4d986720

Observation 4f7baa9b-77d8-4a3e-8cf2-10a4f0f8b209 · inbound

LeapTalk: Breaking the Latency-Quality Trade-off in Talking Head Generation cites this paper.

LeapTalk: Breaking the Latency-Quality Trade-off in Talking Head Generation AniPortrait: Audio-Driven Synthesis of Photorealistic Portrait Animation

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-04T01:32:12.829607Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T01:32:12.829607Z digest=sha256:9281e6d7413d01618332f32cead87965f69b59ff52ffe2349f2bd4c2e458f2bd

Observation 5056342d-4155-41e5-8f60-43ca33456263 · inbound

Geometry-guided Emotion Modulation for Controllable and Photorealistic Emotional Talking Face Generation cites this paper.

Geometry-guided Emotion Modulation for Controllable and Photorealistic Emotional Talking Face Generation AniPortrait: Audio-Driven Synthesis of Photorealistic Portrait Animation

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-04T01:03:12.618881Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:03:12.618881Z digest=sha256:9774811cb76a7fd161d8c1520ed4a7bfa3a84617d5860d151e7d086682944054