Pith. sign in

Paper Citation Record · LEDGER

Audio2Head: Audio-driven One-shot Talking-head Generation with Natural Head Motion

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 10 inbound Pith citation observations for arXiv:2107.09293.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2107.09293 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 10 of 10 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 10 of 10 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-09T18:23:42.444845Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-01T21:46:15.581615Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation cb04df4c-a84f-4ec1-93a3-b2fd16aa261e · inbound

EmoTalkingGaussian: Continuous Emotion-conditioned Talking Head Synthesis cites this paper.

EmoTalkingGaussian: Continuous Emotion-conditioned Talking Head Synthesis Audio2Head: Audio-driven One-shot Talking-head Generation with Natural Head Motion

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-09T18:16:40.933559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T18:16:40.933559Z digest=sha256:deb992e56770127e07cca4caab428e98520b47c93c9de8521ba3f45da4b3647e

Observation c6111573-4bb5-4e1a-a5fb-e502009c9c16 · inbound

Emotion Recognition and Generation: A Comprehensive Review of Face, Speech, and Text Modalities cites this paper.

Emotion Recognition and Generation: A Comprehensive Review of Face, Speech, and Text Modalities Audio2Head: Audio-driven One-shot Talking-head Generation with Natural Head Motion

Reference 153

Resolution
unresolved
no resolver link, observed 2026-08-09T18:23:42.444845Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T18:23:42.444845Z digest=sha256:1fada8d9586d4444e8713d8c4be5cb54aef41c2df8e538f3471d4d33c2e826b2

Observation 4990ae01-00a6-440d-94d2-4aee3b75ffcb · inbound

Long-Term TalkingFace Generation via Motion-Prior Conditional Diffusion Model cites this paper.

Long-Term TalkingFace Generation via Motion-Prior Conditional Diffusion Model Audio2Head: Audio-driven One-shot Talking-head Generation with Natural Head Motion

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T21:11:25.336787Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T21:11:25.336787Z digest=sha256:75f2f1eccb9812561de0e7554015cbcfd24163346caa3990d8c95ccd37da745e

Observation 49565245-a006-4168-b27b-68b18c85b1c1 · inbound

NTIRE 2025 XGC Quality Assessment Challenge: Methods and Results cites this paper.

NTIRE 2025 XGC Quality Assessment Challenge: Methods and Results Audio2Head: Audio-driven One-shot Talking-head Generation with Natural Head Motion

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-07T11:20:51.176183Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:20:51.176183Z digest=sha256:cc5343e7f216091d0a054454805a1d8742e808639b3846baab5e3247a72e7cdb

Observation 5f3739da-5e95-40cc-ad27-7280efcb2de9 · inbound

SyncTalk++: High-Fidelity and Efficient Synchronized Talking Heads Synthesis Using Gaussian Splatting cites this paper.

SyncTalk++: High-Fidelity and Efficient Synchronized Talking Heads Synthesis Using Gaussian Splatting Audio2Head: Audio-driven One-shot Talking-head Generation with Natural Head Motion

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T00:15:44.518622Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:15:44.518622Z digest=sha256:4e45867506b494952092f6ad57b5bd5cfd4ddff8ffe746b36b98695510f7302e

Observation b549b02b-96c3-43dc-8314-61caf0adf92d · inbound

Who is a Better Talker: Subjective and Objective Quality Assessment for AI-Generated Talking Heads cites this paper.

Who is a Better Talker: Subjective and Objective Quality Assessment for AI-Generated Talking Heads Audio2Head: Audio-driven One-shot Talking-head Generation with Natural Head Motion

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T10:53:19.120381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T10:53:19.120381Z digest=sha256:1abbda4fffc3db1fbcdc5fad780c50129d3bea5c5babe5f4e00bff026222c8b1

Observation 1b9799a9-e6e2-4e32-af47-421865ccbb43 · inbound

AUHead: Realistic Emotional Talking Head Generation via Action Units Control cites this paper.

AUHead: Realistic Emotional Talking Head Generation via Action Units Control Audio2Head: Audio-driven One-shot Talking-head Generation with Natural Head Motion

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-16T05:50:40.246667Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-16T05:49:15.734418Z digest=sha256:036115328619ee62bdb116b4a6e2dd0e408a2aa463e10cd27ebf752be3bafaf7

Observation 707c910a-ce2d-4375-ba78-2e5ec50a3b58 · inbound

CoInteract: Physically-Consistent Human-Object Interaction Video Synthesis via Spatially-Structured Co-Generation cites this paper.

CoInteract: Physically-Consistent Human-Object Interaction Video Synthesis via Spatially-Structured Co-Generation Audio2Head: Audio-driven One-shot Talking-head Generation with Natural Head Motion

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-05-11T12:41:04.425900Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T03:14:45.834520Z digest=sha256:d484414629a840db135b02c08ff035e44f2f475f027b4999be73967ad2d8e240

Observation 98ed7b23-2b3a-4cd5-9d47-00028f0b7d70 · inbound

Temporally-Aligned Evaluation for Audio-Driven Talking Head Generation cites this paper.

Temporally-Aligned Evaluation for Audio-Driven Talking Head Generation Audio2Head: Audio-driven One-shot Talking-head Generation with Natural Head Motion

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-07-01T21:46:15.583006Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-28T16:19:55.048831Z digest=sha256:efab6495d4f6d3b2c88a55128f1892827581d621d177ca08a9b7770013d3a773

Observation dbf346d2-dc65-4f91-9949-97b7be03f722 · inbound

Proxy Avatar Meets Low-Rank Caching: Real-Time One-Shot Emotion-Controllable Portrait Animation cites this paper.

Proxy Avatar Meets Low-Rank Caching: Real-Time One-Shot Emotion-Controllable Portrait Animation Audio2Head: Audio-driven One-shot Talking-head Generation with Natural Head Motion

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-04T17:26:50.150881Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T17:26:50.150881Z digest=sha256:694587da307223438c16dd476682c75af00066c0aeaa83b37f3f3b6f6acd2757