Pith. sign in

Paper Citation Record · LEDGER

INFP: Audio-Driven Interactive Head Generation in Dyadic Conversations

As of 21 August 2026, this Paper Citation Record lists 61 of 61 outbound references and 5 inbound Pith citation observations for arXiv:2412.04037.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.04037 v1

Coverage vector

measured 61 of 61 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T21:53:11.966188Z

measured 66 of 66 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 5 of 5 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T18:29:41.247052Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-06T04:46:05.518160Z

Reference resolution

61 of 61 outbound references displayed

  • verified exact1
  • verified fuzzy35
  • unresolved25
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 9dfbd778-91a0-4e1a-b756-f41b373cabdb · outbound

This paper cites Talknet: Fully-convolutional non-autoregressive speech syn- thesis model, 2020.

INFP: Audio-Driven Interactive Head Generation in Dyadic Conversations Talknet: Fully-convolutional non-autoregressive speech syn- thesis model, 2020

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:53:12.662393Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T21:53:11.736835Z digest=sha256:c893ab4499d3688fe695883bf0850e46b7c8150f3ccc719a4ba0b0f358b7a137

Observation e01a7de3-3a67-43a5-80b4-4f2f332e41cb · outbound

This paper cites EchoMimic: Lifelike Audio-Driven Portrait Animations through Editable Landmark Conditions.

INFP: Audio-Driven Interactive Head Generation in Dyadic Conversations EchoMimic: Lifelike Audio-Driven Portrait Animations through Editable Landmark Conditions

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T21:53:11.741339Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:53:11.741339Z digest=sha256:21c3bc38188e842aac194855dac929d205bec232b1fd2ee934bebf87432eecac

Observation 8c939fb3-5021-474a-a984-8d0821648288 · outbound

This paper cites an unresolved cited work.

INFP: Audio-Driven Interactive Head Generation in Dyadic Conversations Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-11T21:53:12.651792Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T21:53:11.745431Z digest=sha256:2f5a6e90c722a12c2a2ecbdb34efec4d3ee66dd91cdd95af671dd956b44975ac

Observation 1d51c5e5-81bf-4137-86c4-16bb6e96f1b8 · outbound

This paper cites Arcface: Additive angular margin loss for deep face recognition.

INFP: Audio-Driven Interactive Head Generation in Dyadic Conversations Arcface: Additive angular margin loss for deep face recognition

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:53:12.641213Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T21:53:11.749257Z digest=sha256:e5dee11974d71b4dbc0cbb67055189f13eca80187151129e941c16b46368cefe

Observation 08c07cb8-55ad-46f0-b586-bd5b8030d8f4 · outbound

This paper cites Accurate 3d face reconstruction with weakly-supervised learning: From single image to image set.

INFP: Audio-Driven Interactive Head Generation in Dyadic Conversations Accurate 3d face reconstruction with weakly-supervised learning: From single image to image set

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:53:12.629328Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T21:53:11.753360Z digest=sha256:7d68c497a92086bd9d2f8250c23aae301f5b2a44dfedd8a35191e43a8d1b0fd7

Observation af9b6e5d-18ec-444a-84a9-a2c8769f56a4 · outbound

This paper cites Megaportraits: One-shot megapixel neural head avatars.

INFP: Audio-Driven Interactive Head Generation in Dyadic Conversations Megaportraits: One-shot megapixel neural head avatars

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:53:12.617980Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T21:53:11.757387Z digest=sha256:243bf38536c05e5fcf3615b6f62bae34eb0a4cab296b43fd6e15b6afe8149eaf

Observation e4c6fd79-5ff4-4c1e-80f9-c40506c89119 · outbound

This paper cites Visual Speech-Aware Perceptual 3D Facial Expression Reconstruction from Videos.

INFP: Audio-Driven Interactive Head Generation in Dyadic Conversations Visual Speech-Aware Perceptual 3D Facial Expression Reconstruction from Videos

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T21:53:11.761550Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:53:11.761550Z digest=sha256:0384810741f31711616251fc57dedcb1dfed1cc91bec4bd4f19bc68ff8b191c0

Observation 3f5d3a74-cdb1-4e34-8cbe-fce70d047100 · outbound

This paper cites Affective faces for goal-driven dyadic communication.

INFP: Audio-Driven Interactive Head Generation in Dyadic Conversations Affective faces for goal-driven dyadic communication

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:53:12.605474Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T21:53:11.765905Z digest=sha256:72c65877b1f66243c691d65399cae2ff54ae4f8968349bbc9a62d8bf91f1d524

Observation d0a58d21-4ac4-4bf8-a592-1fdf781060f1 · outbound

This paper cites Stylesync: High-fidelity generalized and personalized lip sync in style-based genera- tor.

INFP: Audio-Driven Interactive Head Generation in Dyadic Conversations Stylesync: High-fidelity generalized and personalized lip sync in style-based genera- tor

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T21:53:11.769746Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:53:11.769746Z digest=sha256:9f9ae74b2de4f03472871407ca4147da3d6abd5802168b5e37f4428a2defde35

Observation 64457f55-f057-4b5d-9b20-366deb89aea7 · outbound

This paper cites Animatediff: Animate your personalized text-to- image diffusion models without specific tuning.

INFP: Audio-Driven Interactive Head Generation in Dyadic Conversations Animatediff: Animate your personalized text-to- image diffusion models without specific tuning

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:53:12.587350Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T21:53:11.773705Z digest=sha256:187015c2e67479400e3610fcb50d5d0fbfa8be46c9353e69c33ccabec3cef4b3

Observation d46729b8-d32c-4717-9bde-47efb366501e · outbound

This paper cites Denoising dif- fusion probabilistic models.

INFP: Audio-Driven Interactive Head Generation in Dyadic Conversations Denoising dif- fusion probabilistic models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T21:53:11.777648Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:53:11.777648Z digest=sha256:5d6b13492ac6a9b692bc3dbf9a1ac9ad816f1a5e160df5d88884becaebaaf654

Observation 3ea19e3d-1b9a-444c-84f8-3b943d9c3b02 · outbound

This paper cites Hubert: Self-supervised speech representation learning by masked prediction of hidden units.

INFP: Audio-Driven Interactive Head Generation in Dyadic Conversations Hubert: Self-supervised speech representation learning by masked prediction of hidden units

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:53:12.570776Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T21:53:11.781659Z digest=sha256:7f65c03a852b2ed5c43af4d7967a6558a554ad1cb125523b0f8966524df65928

Observation 4cd2147a-360d-47ac-a1ff-c919f54cc49c · outbound

This paper cites Percep- tual conversational head generation with regularized driver and enhanced renderer.

INFP: Audio-Driven Interactive Head Generation in Dyadic Conversations Percep- tual conversational head generation with regularized driver and enhanced renderer

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:53:12.560307Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T21:53:11.786102Z digest=sha256:5be00206c389d3ee6618249cd1bc24b2122a6e518f6a4f9672fe9dc104c83823

Observation f876d396-533d-4fa7-a2af-2d1b9e52ea9b · outbound

This paper cites Interact: Capture and modelling of realistic, ex- pressive and interactive activities between two persons in daily scenarios.

INFP: Audio-Driven Interactive Head Generation in Dyadic Conversations Interact: Capture and modelling of realistic, ex- pressive and interactive activities between two persons in daily scenarios

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:53:12.549397Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T21:53:11.790062Z digest=sha256:9ffd7f3d59e81e598257c045ff0ca5af2665d4c9ab7ff48c437af05720d11189

Observation 4c5b979d-c718-4557-a079-32305fd51330 · outbound

This paper cites Analyzing and improving the image quality of StyleGAN.

INFP: Audio-Driven Interactive Head Generation in Dyadic Conversations Analyzing and improving the image quality of StyleGAN

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:53:12.537303Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T21:53:11.793901Z digest=sha256:d6a8c14a801c97230ca725eaad0d23d5d570b4e6f6f462568c3ac41cd75f2c89

Observation 80ee6fb2-eae4-4aa6-93ef-5b5a6110388a · outbound

This paper cites Mfr-net: Multi-faceted responsive listening head generation via denoising diffusion model.

INFP: Audio-Driven Interactive Head Generation in Dyadic Conversations Mfr-net: Multi-faceted responsive listening head generation via denoising diffusion model

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:53:12.524818Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T21:53:11.798431Z digest=sha256:6d40b0162ab6eab7656540095768934b56eb5a1492f879c6d6020420f30ea076

Observation 0a198528-c71e-4090-b58a-2a5be7fb51e0 · outbound

This paper cites Anitalker: Animate vivid and di- verse talking faces through identity-decoupled facial motion encoding, 2024.

INFP: Audio-Driven Interactive Head Generation in Dyadic Conversations Anitalker: Animate vivid and di- verse talking faces through identity-decoupled facial motion encoding, 2024

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:53:12.512443Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T21:53:11.802718Z digest=sha256:866b153e3e0bd92073552f3464dd7e3c6912523d4eb87a48fed127a1aa925c6e

Observation 1bcd68d9-ed2a-4e36-b3a7-d0d0cf83888f · outbound

This paper cites Customlistener: Text-guided responsive inter- action for user-friendly listening head generation.

INFP: Audio-Driven Interactive Head Generation in Dyadic Conversations Customlistener: Text-guided responsive inter- action for user-friendly listening head generation

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:53:12.500099Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T21:53:11.806502Z digest=sha256:627aa75f105c55ae1d39e36fc93144ea96c5f1dfd56feba823db0c493eb69643

Observation 92457ba9-b2f2-426b-b2e2-a57c29c27bdf · outbound

This paper cites Decoupled weight decay regularization, 2019.

INFP: Audio-Driven Interactive Head Generation in Dyadic Conversations Decoupled weight decay regularization, 2019

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T21:53:11.810244Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:53:11.810244Z digest=sha256:255bb720b0eb07aac40c3feaac1e57b30ac5f111c996a340fc3fa723aa8bdc8f

Observation 4d2cc443-95cc-4ba6-b0e8-8f657a0ff0fa · outbound

This paper cites Mediapipe: A framework for perceiving and processing reality.

INFP: Audio-Driven Interactive Head Generation in Dyadic Conversations Mediapipe: A framework for perceiving and processing reality

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:53:12.478745Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T21:53:11.814002Z digest=sha256:8a0aca395e16cd449d14630b65e70cd5f89b2faa9d1c979f968d4006a0784178

Observation 17a75fd3-f8c2-4b18-9369-64d289ad0d98 · outbound

This paper cites Styletalk: one-shot talking head generation with controllable speaking styles.

INFP: Audio-Driven Interactive Head Generation in Dyadic Conversations Styletalk: one-shot talking head generation with controllable speaking styles

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:53:12.466465Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T21:53:11.817967Z digest=sha256:81a8aca95805d09283f17ac3954081e29dc31f667e82462a75ab1ac6db41129d

Observation 90b8a417-58e2-4082-a8b7-fb75f5d16a88 · outbound

This paper cites Follow-Your-Emoji: Fine-Controllable and Expressive Freestyle Portrait Animation.

INFP: Audio-Driven Interactive Head Generation in Dyadic Conversations Follow-Your-Emoji: Fine-Controllable and Expressive Freestyle Portrait Animation

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T21:53:11.821517Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:53:11.821517Z digest=sha256:ac608c42221134f63f19f855e8f30acbed036dd0cfdcedeecc9cbd846a177dd4

Observation c3b53312-fc9d-4762-b338-396a61ab3e84 · outbound

This paper cites Learning to listen: Modeling non-deterministic dyadic facial motion.

INFP: Audio-Driven Interactive Head Generation in Dyadic Conversations Learning to listen: Modeling non-deterministic dyadic facial motion

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:53:12.455133Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T21:53:11.826076Z digest=sha256:3c6154b35fc05b99f1fb62b71d4115abc2e2cfa9ab188ec894d74d0d3801e15c

Observation 34ff912f-bca7-4d07-8d0f-a8975091555f · outbound

This paper cites Can language 9 models learn to listen? In Proceedings of the IEEE/CVF In- ternational Conference on Computer Vision , pages 10083– 10093, 2023.

INFP: Audio-Driven Interactive Head Generation in Dyadic Conversations Can language 9 models learn to listen? In Proceedings of the IEEE/CVF In- ternational Conference on Computer Vision , pages 10083– 10093, 2023

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:53:12.445047Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T21:53:11.829510Z digest=sha256:2695cc2adac9834741cf6f04ea6213c692eb8a010c4305a22f0beb5909f5e14b

Observation 1bb1ca10-fd45-4e29-b21e-510fe3a8ceb7 · outbound

This paper cites From audio to photoreal embodiment: Synthesizing humans in conversations.

INFP: Audio-Driven Interactive Head Generation in Dyadic Conversations From audio to photoreal embodiment: Synthesizing humans in conversations

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T21:53:11.832950Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:53:11.832950Z digest=sha256:c260a82653bb1daf298f93d6b8f2173c711d7bb67c9ff6a9fae72287b1c33c4e

Observation 828926df-fee3-4fa5-aa54-cc8362e75f70 · outbound

This paper cites Let's Go Real Talk: Spoken Dialogue Model for Face-to-Face Conversation.

INFP: Audio-Driven Interactive Head Generation in Dyadic Conversations Let's Go Real Talk: Spoken Dialogue Model for Face-to-Face Conversation

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T21:53:11.836010Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:53:11.836010Z digest=sha256:fa5c7b7160bbb557625bbbdc4a490956e6e8c3076e87a71a38530d474877f12f

Observation 26962c05-95d5-4fc1-ab7c-e7f69bf8a61a · outbound

This paper cites Nambood- iri, and C.V.

INFP: Audio-Driven Interactive Head Generation in Dyadic Conversations Nambood- iri, and C.V

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:53:12.429424Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T21:53:11.839810Z digest=sha256:e984826a8f5b01472bb3744af7b005c0e90612326baff327ab49807f4491144a

Observation 4167e2e0-7650-4ac9-ae74-1d60ebefa28e · outbound

This paper cites Li, and Shan Liu.

INFP: Audio-Driven Interactive Head Generation in Dyadic Conversations Li, and Shan Liu

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:53:12.419924Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T21:53:11.843204Z digest=sha256:81a8e295bdb0fd7864316c0a814f8baa299f948fe8a4a0d41a6e8f62e273bea8

Observation 00fd7ca6-e607-405b-8a26-32954fd78f03 · outbound

This paper cites High-resolution image syn- thesis with latent diffusion models, 2021.

INFP: Audio-Driven Interactive Head Generation in Dyadic Conversations High-resolution image syn- thesis with latent diffusion models, 2021

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:53:12.407735Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T21:53:11.846509Z digest=sha256:c1be557cdc0d3e4bfc55515888a24abc11592ef25c2a313478ffb276e1a3d979

Observation c1e83505-e230-48f0-9ace-d61588619bd3 · outbound

This paper cites Denoising Diffusion Implicit Models.

INFP: Audio-Driven Interactive Head Generation in Dyadic Conversations Denoising Diffusion Implicit Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-11T21:53:11.849442Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:53:11.849442Z digest=sha256:d9164af32c669288c38750448c179329d0b6c7495350252d47381be88665b347

Observation 5ba9a542-de87-4e9b-8dc8-2bb4b7183f42 · outbound

This paper cites Emotional listener portrait: Realistic listener motion simulation in conversation.

INFP: Audio-Driven Interactive Head Generation in Dyadic Conversations Emotional listener portrait: Realistic listener motion simulation in conversation

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:53:12.394845Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T21:53:11.853496Z digest=sha256:fcda3480c4200799bd15c086246a32dab2f61d6f0278dfd3f8a7d0a1423043b7

Observation f49d2435-db0a-451f-8d7c-0425748ea3b4 · outbound

This paper cites React 2024: the second multiple appropriate facial reaction generation challenge, 2024.

INFP: Audio-Driven Interactive Head Generation in Dyadic Conversations React 2024: the second multiple appropriate facial reaction generation challenge, 2024

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:53:12.383278Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T21:53:11.857028Z digest=sha256:83f20cf2caf7bbbc268613b2585a3509052a2353df386a4eec975f7a83ded5e4

Observation 44fce604-1b28-4b16-a51f-e5b43513cf80 · outbound

This paper cites Diffused heads: Diffusion models beat gans on talking-face genera- tion.

INFP: Audio-Driven Interactive Head Generation in Dyadic Conversations Diffused heads: Diffusion models beat gans on talking-face genera- tion

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:53:12.370436Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T21:53:11.860695Z digest=sha256:f40946bc8032742c472ded87eefd57920a80bf8840a3506284f67ae13e358928

Observation bc4d08c7-6a15-489d-845c-6610fe9d8394 · outbound

This paper cites Beyond Talking -- Generating Holistic 3D Human Dyadic Motion for Communication.

INFP: Audio-Driven Interactive Head Generation in Dyadic Conversations Beyond Talking -- Generating Holistic 3D Human Dyadic Motion for Communication

Reference 34

Resolution
verified exact
local_arxiv, observed 2026-08-11T21:53:12.138969Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T21:53:11.864164Z digest=sha256:6d3169ce7b2882b9ef7caf69ce17fd643e99e2885c84877e9cbbfc92edae8aad

Observation c36eb33e-eb22-470d-aa2f-1244e748e648 · outbound

This paper cites Diffposetalk: Speech-driven stylistic 3d facial animation and head pose generation via diffusion models.

INFP: Audio-Driven Interactive Head Generation in Dyadic Conversations Diffposetalk: Speech-driven stylistic 3d facial animation and head pose generation via diffusion models

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:53:12.358339Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T21:53:11.868282Z digest=sha256:c552e6405378231b92785b6d1e73cf70cf8ed3c57918078bfdddbde740d35899

Observation b7ae3069-786c-4293-940e-d229b14366d4 · outbound

This paper cites Real-time Neural Radiance Talking Portrait Synthesis via Audio-spatial Decomposition.

INFP: Audio-Driven Interactive Head Generation in Dyadic Conversations Real-time Neural Radiance Talking Portrait Synthesis via Audio-spatial Decomposition

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-11T21:53:11.871802Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:53:11.871802Z digest=sha256:5cadc96fe023b818ac9664847c505f7aed897dce5e3705dba07ab44977b3b915

Observation 5422135d-daf7-4ae8-9c3f-f71349a2abcb · outbound

This paper cites EMO: Emote Portrait Alive -- Generating Expressive Portrait Videos with Audio2Video Diffusion Model under Weak Conditions.

INFP: Audio-Driven Interactive Head Generation in Dyadic Conversations EMO: Emote Portrait Alive -- Generating Expressive Portrait Videos with Audio2Video Diffusion Model under Weak Conditions

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-11T21:53:11.875876Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:53:11.875876Z digest=sha256:8921d4ab5ad55f8deb9cd5016736ab9b40a8d34991fbaf0bb431ebd88626c0b0

Observation 44299887-66e1-4064-9b30-a94d1897ee86 · outbound

This paper cites The information bottleneck method.

INFP: Audio-Driven Interactive Head Generation in Dyadic Conversations The information bottleneck method

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-11T21:53:11.879569Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:53:11.879569Z digest=sha256:74ab4691e13738ca5c5f9511caa5be94cd10ba3f87734444d21164dcfaae3d13

Observation 15856dd0-0eb4-4200-8c0a-0fe06384bf69 · outbound

This paper cites Dyadic Interaction Modeling for Social Behavior Generation.

INFP: Audio-Driven Interactive Head Generation in Dyadic Conversations Dyadic Interaction Modeling for Social Behavior Generation

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-11T21:53:11.883377Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:53:11.883377Z digest=sha256:ed0e654300647835266e4838132a53a082b995c4aba4cc49beb8e34538f4bb3a

Observation f55b42de-9ac8-47a2-a1f2-d080df0dae48 · outbound

This paper cites Neural discrete representation learning,.

INFP: Audio-Driven Interactive Head Generation in Dyadic Conversations Neural discrete representation learning,

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-11T21:53:11.887127Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:53:11.887127Z digest=sha256:33e6db01d4f9803ecaa0dd1f93f663c200ec8c3352b5f4a657ced617eed67d9e

Observation 38bd2266-8725-4b43-b102-90c96203977f · outbound

This paper cites V-Express: Conditional Dropout for Progressive Training of Portrait Video Generation.

INFP: Audio-Driven Interactive Head Generation in Dyadic Conversations V-Express: Conditional Dropout for Progressive Training of Portrait Video Generation

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-11T21:53:11.890901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:53:11.890901Z digest=sha256:408fdcaacaab39d290cf94a42d2081ec6aeddd416eb1a344327f4f14d9339d0e

Observation 5b61c8f8-b7cb-4adb-beca-738db3cbe68f · outbound

This paper cites Dis- entangling planning, driving and rendering for photorealistic avatar agents.

INFP: Audio-Driven Interactive Head Generation in Dyadic Conversations Dis- entangling planning, driving and rendering for photorealistic avatar agents

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:53:12.341725Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T21:53:11.895026Z digest=sha256:021f7cff61209ea80481a773359284e4e8626c3ae97bb19efb4af7bc462cae4a

Observation fb60fb04-14c2-42f7-a499-1b56bb6f36da · outbound

This paper cites Progressive disentangled representation learning for fine-grained controllable talking head synthesis.

INFP: Audio-Driven Interactive Head Generation in Dyadic Conversations Progressive disentangled representation learning for fine-grained controllable talking head synthesis

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:53:12.331269Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T21:53:11.898857Z digest=sha256:7a69759821dde32ef00622384f8f12da7376472c6985091ce70b3c31adfc46f1

Observation c4bf814d-f4fb-4ae2-a766-52e77c0130c6 · outbound

This paper cites Latent Image Animator: Learning to Animate Images via Latent Space Navigation.

INFP: Audio-Driven Interactive Head Generation in Dyadic Conversations Latent Image Animator: Learning to Animate Images via Latent Space Navigation

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-11T21:53:11.902478Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:53:11.902478Z digest=sha256:ad6d236ff02b3f80880c850446094e8868d1172b6759b0f0769ea43c3a7ebbdf

Observation 3dc37f7c-6fb2-4cbf-b667-81018ea390d6 · outbound

This paper cites Hallo: Hierarchical Audio-Driven Visual Synthesis for Portrait Image Animation.

INFP: Audio-Driven Interactive Head Generation in Dyadic Conversations Hallo: Hierarchical Audio-Driven Visual Synthesis for Portrait Image Animation

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-11T21:53:11.906895Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:53:11.906895Z digest=sha256:c40d6ab488e53cc4a2eaa17eee71256d3346ec8d3ef54aae611f242f172e96dc

Observation 190e1e3a-50df-4b69-a36a-75126bd6c944 · outbound

This paper cites VASA-1: Lifelike Audio-Driven Talking Faces Generated in Real Time.

INFP: Audio-Driven Interactive Head Generation in Dyadic Conversations VASA-1: Lifelike Audio-Driven Talking Faces Generated in Real Time

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-11T21:53:11.911780Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:53:11.911780Z digest=sha256:cbe3890f54a1b781cdc608a5af5706a75228460579d59fc9ef3680caed9861fc

Observation 2f065508-d3ba-46c9-9ade-ff4086ac8fe5 · outbound

This paper cites Dialoguenerf: Towards realistic avatar face- to-face conversation video generation, 2023.

INFP: Audio-Driven Interactive Head Generation in Dyadic Conversations Dialoguenerf: Towards realistic avatar face- to-face conversation video generation, 2023

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:53:12.320642Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T21:53:11.916197Z digest=sha256:d3d1c71197df197a4faca84d3e1ad92e2f761b6d671b93403fbc10c26dcebe5c

Observation d274d88d-da5d-4bc9-a55b-d7a30f776d41 · outbound

This paper cites DREAM-Talk: Diffusion-based Realistic Emotional Audio-driven Method for Single Image Talking Face Generation.

INFP: Audio-Driven Interactive Head Generation in Dyadic Conversations DREAM-Talk: Diffusion-based Realistic Emotional Audio-driven Method for Single Image Talking Face Generation

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-11T21:53:11.920401Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:53:11.920401Z digest=sha256:073ab974779336f5b2f8018638c4130392e9fbefed13ad0e6b52159462981094

Observation d59b9a30-b0f5-4193-9e94-cdbed22f3d6d · outbound

This paper cites PersonaTalk: Bring Attention to Your Persona in Visual Dubbing.

INFP: Audio-Driven Interactive Head Generation in Dyadic Conversations PersonaTalk: Bring Attention to Your Persona in Visual Dubbing

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-11T21:53:11.924761Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:53:11.924761Z digest=sha256:29afd710b90fee3fb2729005f04300bd3a2ec67d96fbb58d4b9ed4a825e7913e

Observation 337fa33e-41e9-4670-bb57-838f5fff805f · outbound

This paper cites The unreasonable effectiveness of deep features as a perceptual metric.

INFP: Audio-Driven Interactive Head Generation in Dyadic Conversations The unreasonable effectiveness of deep features as a perceptual metric

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-11T21:53:11.928742Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:53:11.928742Z digest=sha256:2a802859d915b0c4106ac94b9e0a7b61bb7a013d2ca602b4be0a6232a41b2122

Observation d69ba8ac-5856-4051-abbf-0a4f49d87ea5 · outbound

This paper cites an unresolved cited work.

INFP: Audio-Driven Interactive Head Generation in Dyadic Conversations Unresolved cited work

Reference 51

Resolution
unresolved
raw_fallback, observed 2026-08-11T21:53:12.302455Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T21:53:11.932444Z digest=sha256:220dbf833d06aa2b0d4d5dc2e5af60a520617ae253a0fe5ec9c276cbf53fefd3

Observation ec525900-31ea-4a90-a3fc-7c55d91ef699 · outbound

This paper cites Sadtalker: Learning realistic 3d motion coefficients for stylized audio- driven single image talking face animation.

INFP: Audio-Driven Interactive Head Generation in Dyadic Conversations Sadtalker: Learning realistic 3d motion coefficients for stylized audio- driven single image talking face animation

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:53:12.291723Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T21:53:11.935710Z digest=sha256:1c269d9355de1ed2d884d810908c381f8a5e23f2d7cd4fef4c10f1db3d410e05

Observation 9d72ee49-1f4e-44a9-88b4-b06b855aac9a · outbound

This paper cites Flow-guided one-shot talking face generation with a high-resolution audio-visual dataset.

INFP: Audio-Driven Interactive Head Generation in Dyadic Conversations Flow-guided one-shot talking face generation with a high-resolution audio-visual dataset

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:53:12.280781Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T21:53:11.939319Z digest=sha256:1e2bbfc3f5f52956b2133662d3c78f488e9b9ba58b5b6512ec421314ee16b6d5

Observation cba50b78-9cf3-4440-8865-b82f94135e34 · outbound

This paper cites Mossformer2: Combining transformer and rnn-free recurrent network for enhanced time-domain monaural speech separation, 2023.

INFP: Audio-Driven Interactive Head Generation in Dyadic Conversations Mossformer2: Combining transformer and rnn-free recurrent network for enhanced time-domain monaural speech separation, 2023

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:53:12.270963Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T21:53:11.942072Z digest=sha256:e4326e1fb31fdc2e1b2f71b1f2f5173612becbaf572377f479dd776dffb6de9d

Observation e01301c8-85e0-43fd-95db-9deb7568ab80 · outbound

This paper cites Semantic-aware responsive listener head syn- thesis.

INFP: Audio-Driven Interactive Head Generation in Dyadic Conversations Semantic-aware responsive listener head syn- thesis

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:53:12.258543Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T21:53:11.944923Z digest=sha256:ff421178d759bed56484a96848621dc1c9c73b839408b1d6c3939eddeac014ff

Observation ea97939c-5eeb-4b4e-96d6-562d765352ff · outbound

This paper cites Pose-controllable talking face generation by implicitly modularized audio-visual rep- resentation.

INFP: Audio-Driven Interactive Head Generation in Dyadic Conversations Pose-controllable talking face generation by implicitly modularized audio-visual rep- resentation

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:53:12.246872Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T21:53:11.948116Z digest=sha256:3ed8e38f75bab4380a4cf83072fb85de781c7aaafebe0bbe3e732c0a8f4282d0

Observation 7a6df66d-d046-4ea9-97dd-b41d5db23498 · outbound

This paper cites Vico-x: Multimodal conversation dataset.

INFP: Audio-Driven Interactive Head Generation in Dyadic Conversations Vico-x: Multimodal conversation dataset

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:53:12.235650Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T21:53:11.951252Z digest=sha256:3d5a16ade9d9ecec941bb41e1704e33c26fa37f66acb3e4c969127e35b56bf26

Observation 38ddad94-5703-4d7c-a4fe-5dbafca72a3e · outbound

This paper cites Responsive listening head generation: A benchmark dataset and baseline.

INFP: Audio-Driven Interactive Head Generation in Dyadic Conversations Responsive listening head generation: A benchmark dataset and baseline

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:53:12.213548Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T21:53:11.958153Z digest=sha256:edc913c5e934f85b1e3c2b855dcccd27c21bb9b8af664e01637107dac835a9c8

Observation aea57f22-9287-4d71-8a75-dc7fa33b229c · outbound

This paper cites Interactive Conversational Head Generation.

INFP: Audio-Driven Interactive Head Generation in Dyadic Conversations Interactive Conversational Head Generation

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-11T21:53:11.961583Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:53:11.961583Z digest=sha256:705533cc6773be9565e2ad354332aaefc567a1145e1b7f3bdf9afc43cc786c2d

Observation 21ee5012-1755-4aec-a448-072998ad928d · outbound

This paper cites Makelttalk: speaker-aware talking-head animation.

INFP: Audio-Driven Interactive Head Generation in Dyadic Conversations Makelttalk: speaker-aware talking-head animation

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:53:12.200475Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T21:53:11.966188Z digest=sha256:7e3a5056656dd62c4edaaa58fc7129f6b1290b0a5afe17f3fa18e1b8f694d94b

Observation b748deca-3023-4fc7-a8a8-84f33a166cc1 · outbound

This paper cites an unresolved cited work.

INFP: Audio-Driven Interactive Head Generation in Dyadic Conversations Unresolved cited work

Reference 2022

Resolution
unresolved
raw_fallback, observed 2026-08-11T21:53:12.224929Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-11T21:53:11.954498Z digest=sha256:053309a6d73fce021e3b90844d94b18e7e027b09cf77e7a19935de867c368918

Pith citing papers

Observation bc068e31-25cc-4c12-bd41-c1c26721d8de · inbound

LLIA -- Enabling Low-Latency Interactive Avatars: Real-Time Audio-Driven Portrait Video Generation with Diffusion Models cites this paper.

LLIA -- Enabling Low-Latency Interactive Avatars: Real-Time Audio-Driven Portrait Video Generation with Diffusion Models INFP: Audio-Driven Interactive Head Generation in Dyadic Conversations

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T10:18:23.943246Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:18:23.943246Z digest=sha256:90fb7af0fb33869eb270a1129fa98728a8ae00b30a5e498d95213b09054f7207

Observation 65f06d46-f32a-4f29-bb42-726d23356af8 · inbound

Bind-Your-Avatar: Multi-Talking-Character Video Generation with Dynamic 3D-mask-based Embedding Router cites this paper.

Bind-Your-Avatar: Multi-Talking-Character Video Generation with Dynamic 3D-mask-based Embedding Router INFP: Audio-Driven Interactive Head Generation in Dyadic Conversations

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-15T18:29:41.247052Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:29:41.247052Z digest=sha256:824f1b7c8576d876f764973fbcd82c5e591e16c434a8b89068fbe6eb02603dc5

Observation 3239a036-a110-470e-84cc-4cd7c2e1fd95 · inbound

ARIG: Autoregressive Interactive Head Generation for Real-time Conversations cites this paper.

ARIG: Autoregressive Interactive Head Generation for Real-time Conversations INFP: Audio-Driven Interactive Head Generation in Dyadic Conversations

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T21:19:17.092084Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:19:17.092084Z digest=sha256:9e99ca8c3346cbbc6df1873745450ef4388e34ff25458474eee09eff099b39bd

Observation 466d4ee9-2b49-4c9d-871e-bd7dd98597c9 · inbound

Real-time Generation of Various Types of Nodding for Avatar Attentive Listening System cites this paper.

Real-time Generation of Various Types of Nodding for Avatar Attentive Listening System INFP: Audio-Driven Interactive Head Generation in Dyadic Conversations

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T10:58:01.712367Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T10:58:01.712367Z digest=sha256:2f1ad8b97852076e0c8faf91fc3b0e1df22b0aaba97219d2c891bc9b7fca92cd

Observation 8569be9f-6540-496b-9667-ed88b5a6cfde · inbound

Multi-human Interactive Talking Dataset cites this paper.

Multi-human Interactive Talking Dataset INFP: Audio-Driven Interactive Head Generation in Dyadic Conversations

Reference 60

Resolution
verified exact
local_arxiv, observed 2026-08-06T04:46:05.601149Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T04:46:05.382405Z digest=sha256:45ad0a77121bd8fe445e73e4f850addb583f35649424682abb7ee834855c3dd4