Pith. sign in

Paper Citation Record · LEDGER

MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation

As of 8 August 2026, this Paper Citation Record lists 54 of 54 outbound references and 0 inbound Pith citation observations for arXiv:2507.05092.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.05092 v1

Coverage vector

measured 54 of 54 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T19:38:20.590974Z

measured 54 of 54 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

54 of 54 outbound references displayed

  • verified exact3
  • verified fuzzy37
  • unresolved14
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 8cee1c3d-4d3b-4c04-bcf2-587a78f80c85 · outbound

This paper cites A morphable model for the synthesis of 3d faces.

MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation A morphable model for the synthesis of 3d faces

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:38:30.736538Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T19:38:20.430133Z digest=sha256:5891c97e3eefac9b215369e25e7f97166ff1544c26836fa54930e05dc972aef6

Observation b4f8edd8-0cc6-47df-afb1-b065500818bb · outbound

This paper cites Hierarchical cross-modal talking face generation with dynamic pixel-wise loss.

MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation Hierarchical cross-modal talking face generation with dynamic pixel-wise loss

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:38:30.513004Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T19:38:20.433575Z digest=sha256:165230b566cb777c46cc2ab7c00a313772342393c9cd4cafc1355c3591ab4c4e

Observation fc95a1b9-b64f-41a4-bf41-efd907fe8f5c · outbound

This paper cites You said that?.

MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation You said that?

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-08-06T19:38:21.116132Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T19:38:20.437079Z digest=sha256:ba17ebb96534b4dac9888addfea4d029c5abf9d041fe849b8c346b6f48d8ac4e

Observation b4d553b6-7fa0-4490-9a37-b2ebf5c3621f · outbound

This paper cites Lip reading sentences in the wild.

MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation Lip reading sentences in the wild

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:38:30.189584Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T19:38:20.440540Z digest=sha256:12bb8469c05873f19bc1f755e835c5717b66824d28848fca27ed7d57f40e149f

Observation 23f4c44d-469a-4e9a-879e-8adeea596953 · outbound

This paper cites Accurate 3d face reconstruction with weakly-supervised learning: From single image to image set.

MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation Accurate 3d face reconstruction with weakly-supervised learning: From single image to image set

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:38:29.858375Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T19:38:20.443919Z digest=sha256:69652bb71cd80eeb5d1f2a2b75cd3c8150785172d2aa8a9d2f84ec87ffd3d9f3

Observation 1d9fa577-2ed1-4238-a56d-6653311e1b01 · outbound

This paper cites Dae-talker: High fidelity speech-driven talking face generation with diffusion autoen- coder.

MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation Dae-talker: High fidelity speech-driven talking face generation with diffusion autoen- coder

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:38:29.574230Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T19:38:20.446762Z digest=sha256:f7b53893245832d4f9776bf31329ab2f7a62623750408b89162d2f13142cca86

Observation d3848b8e-6786-49dc-92a1-bafbcdb3aedf · outbound

This paper cites Generative adversarial networks.

MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation Generative adversarial networks

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T19:38:20.449788Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:38:20.449788Z digest=sha256:c352fad97cbf1004dc8e8e78e6048030ef660885ba8cbe4149ede9c59959d90c

Observation 5fe11f5c-a2a4-4837-a253-a6c786e08c48 · outbound

This paper cites Ad-nerf: Audio driven neural radi- ance fields for talking head synthesis.

MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation Ad-nerf: Audio driven neural radi- ance fields for talking head synthesis

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:38:29.242496Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T19:38:20.452970Z digest=sha256:a4aacc1a099b884c318479453f66dce8b045c493ec12cb1400748b628ca6f65b

Observation 03a3c0e8-84a5-4ecf-810d-a754cbdf8bd8 · outbound

This paper cites Facexhubert: Text-less speech-driven e (x) pressive 3d facial animation synthesis using self-supervised speech representation learn- ing.

MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation Facexhubert: Text-less speech-driven e (x) pressive 3d facial animation synthesis using self-supervised speech representation learn- ing

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:38:28.887210Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T19:38:20.455770Z digest=sha256:05d7484c90e9e1b9b425dc47f66c67d18bb828fb6cd683505ba21c919c749b30

Observation 7c1bd97c-5501-4a64-aa92-1ca5b3f01f88 · outbound

This paper cites GAIA: Zero-shot Talking Avatar Generation.

MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation GAIA: Zero-shot Talking Avatar Generation

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-08-06T19:38:20.866922Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T19:38:20.458871Z digest=sha256:c008e821c195fc2e39f4b6bc1d330a958ae3d6cc18c26788592100f8db664ad9

Observation f1a09b4b-2371-44f8-b2bc-3d9669cd06b6 · outbound

This paper cites Gans trained by a two time-scale update rule converge to a local nash equilib- rium.

MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation Gans trained by a two time-scale update rule converge to a local nash equilib- rium

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:38:28.593769Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T19:38:20.462422Z digest=sha256:737d7bd9c1133bbcdb08326ec8a19761edae10c239bb6b5bb4a31245e85ed4aa

Observation 0328e1ad-3b23-491a-b9fe-563c3d8b4f84 · outbound

This paper cites Implicit identity representation conditioned memory compensation network for talking head video generation.

MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation Implicit identity representation conditioned memory compensation network for talking head video generation

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:38:28.281921Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T19:38:20.465191Z digest=sha256:41fd58bbf06271197a5776d9ba09e642e1d7419b4be0c5782e665fac3692f064

Observation fe1ef8fa-8c66-4cda-a6bb-dd4350f8f3f8 · outbound

This paper cites DaGAN++: Depth-Aware Generative Adversarial Network for Talking Head Video Generation.

MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation DaGAN++: Depth-Aware Generative Adversarial Network for Talking Head Video Generation

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-08-06T19:38:20.756341Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T19:38:20.467907Z digest=sha256:6617641c9da515e48ac1367250461147f529e1fefc5b61c5428051119ec834dd

Observation 167f1b13-1462-49b2-9c92-1bc5f0dac928 · outbound

This paper cites Depth-aware generative adversarial network for talking head video generation.

MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation Depth-aware generative adversarial network for talking head video generation

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:38:28.009063Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T19:38:20.471010Z digest=sha256:ecb025fde2ecff60de3dc56c562dd2bf09a9daebdf88be39956287223ca15914

Observation aa640803-062e-4a2a-9704-f5eec816209f · outbound

This paper cites Audio-visual Controlled Video Diffusion with Masked Selective State Spaces Modeling for Natural Talking Head Generation.

MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation Audio-visual Controlled Video Diffusion with Masked Selective State Spaces Modeling for Natural Talking Head Generation

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T19:38:20.474478Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:38:20.474478Z digest=sha256:8f7ebbe3ef3ed1a47179d687a124382fd526115f27676e6fda62c440ffbbe3db

Observation 84f58e79-35e0-4779-a77f-f2f8913b2b09 · outbound

This paper cites Audio-driven emotional video portraits.

MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation Audio-driven emotional video portraits

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:38:27.781678Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T19:38:20.481151Z digest=sha256:5ea4b17570bcdc5e3713b6211d681d1030334d0cad52e4a9d756bf27c7d5161e

Observation 8243d0ef-8a90-4b24-a7cf-e347d3ca4fa7 · outbound

This paper cites Transformers in vision: A survey.

MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation Transformers in vision: A survey

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:38:27.506549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T19:38:20.484001Z digest=sha256:6247c15779785081abf391462677b105b92fbf61dee43615d258f0ed15cad37f

Observation d4d0c738-bfa5-4595-a262-3ad63505262d · outbound

This paper cites Deep video portraits.

MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation Deep video portraits

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:38:27.275743Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T19:38:20.486816Z digest=sha256:5103e685e41dc9b615780f14e0bec21ea13851f1f397bc160b44021ad5e61e53

Observation b6204282-38eb-45d7-8ddf-72ca548fb991 · outbound

This paper cites Auto-Encoding Variational Bayes.

MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation Auto-Encoding Variational Bayes

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T19:38:20.489513Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:38:20.489513Z digest=sha256:d44cf6d203adac5e96cc7efb6ed579cbee704b451d9d003a388edba96b492319

Observation e7716d64-440d-4cf7-8da5-33d4e7156fef · outbound

This paper cites AniTalker: Animate Vivid and Diverse Talking Faces through Identity-Decoupled Facial Motion Encoding.

MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation AniTalker: Animate Vivid and Diverse Talking Faces through Identity-Decoupled Facial Motion Encoding

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T19:38:20.492669Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:38:20.492669Z digest=sha256:4455cd22414985238ac0e9b0e95826e497bcc38e8c9c2614168349c0e9f5af72

Observation 6230e1cb-f0f8-44d4-910c-4e30cc395eec · outbound

This paper cites Moda: Mapping-once audio-driven portrait animation with dual attentions.

MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation Moda: Mapping-once audio-driven portrait animation with dual attentions

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:38:26.994177Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T19:38:20.496048Z digest=sha256:3292c718baea50e111e9af7b525eb98751df5bea5c54c933ebda4e525229adcd

Observation cc2d224c-570b-4ae8-9600-2000bb2cc771 · outbound

This paper cites Live speech por- traits: real-time photorealistic talking-head animation.

MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation Live speech por- traits: real-time photorealistic talking-head animation

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:38:26.776730Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T19:38:20.498870Z digest=sha256:03a21a72e1da75dd7b73c8eb7b8ba8ad3b03530b100f40daa0b5c3ad3095a8eb

Observation 602d3749-d771-43e3-9f6e-6b43b772e56d · outbound

This paper cites Training strategies for improved lip- reading.

MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation Training strategies for improved lip- reading

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:38:26.437494Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T19:38:20.501957Z digest=sha256:c6ce1c72b7b4081ca799d467b66e4d5f951e246530fa9c4b4dede8a58be42609

Observation 3d446572-f400-4b64-92f0-d6db6a9cd57a · outbound

This paper cites DreamTalk: When Emotional Talking Head Generation Meets Diffusion Probabilistic Models.

MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation DreamTalk: When Emotional Talking Head Generation Meets Diffusion Probabilistic Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T19:38:20.504641Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:38:20.504641Z digest=sha256:9cdbc93214aab33e8ca2dc6d2bdbb188dbb2f9d15a910327e5cc9a8e05318750

Observation 9c2e43db-8b8a-431b-b31b-810aef76b4fc · outbound

This paper cites DiffSpeaker: Speech-Driven 3D Facial Animation with Diffusion Transformer.

MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation DiffSpeaker: Speech-Driven 3D Facial Animation with Diffusion Transformer

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T19:38:20.507706Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:38:20.507706Z digest=sha256:41280ba533283f873cea52e208367eef03406edf7f91ac7489b44ec17dc9580f

Observation 0049d14a-a49a-4224-91ae-ef785403e837 · outbound

This paper cites Librispeech: An asr corpus based on public do- main audio books.

MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation Librispeech: An asr corpus based on public do- main audio books

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:38:26.182893Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T19:38:20.510785Z digest=sha256:d3f4864c6ab3242119481c28f523ee60c30b5379216586a674ebf8623a84a6b5

Observation 66454920-a701-4407-b670-4873c58a61a8 · outbound

This paper cites A lip sync expert is all you need for speech to lip generation in the wild.

MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation A lip sync expert is all you need for speech to lip generation in the wild

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:38:25.916830Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T19:38:20.513555Z digest=sha256:be23ac6ce727a44e303541e1249c7b9f3b04ea39cdb943245de3a23bf986a06d

Observation 4d84c517-e4cb-4281-bbe7-e67d57a618de · outbound

This paper cites High-resolution image syn- thesis with latent diffusion models.

MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation High-resolution image syn- thesis with latent diffusion models

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:38:25.657706Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T19:38:20.516395Z digest=sha256:cd235cd753b7103f86192b1a4b662bd94e922db36e7fb331a6234cd3cbbfc1a5

Observation 90859342-df51-45a5-affb-1005a4294afc · outbound

This paper cites pytorch-fid: FID Score for PyTorch.

MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation pytorch-fid: FID Score for PyTorch

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T19:38:20.519303Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:38:20.519303Z digest=sha256:77e840f33eadd02f736a32358fac3c07970972ca27927d5054e54912e0d4a265

Observation 182a4583-fc6d-40a7-98af-a6d4a49ca6cb · outbound

This paper cites Difftalk: Crafting diffusion models for generalized audio-driven portraits animation.

MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation Difftalk: Crafting diffusion models for generalized audio-driven portraits animation

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:38:25.321835Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T19:38:20.522395Z digest=sha256:85cfa2a8efb85fb75f7ed834487aaf0ee050e17b1807ff2b498063f1fcca25da

Observation eab7c841-fa7b-4c58-96b0-e37a38906e01 · outbound

This paper cites First order motion model for image animation.

MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation First order motion model for image animation

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:38:25.005868Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T19:38:20.525008Z digest=sha256:a1a0a8708051dc977c707c500c51b9296ab3994d43f7c6539985db1009a5f7c9

Observation 6b328f7e-e45c-473a-b45c-e6a49c2fb524 · outbound

This paper cites Denoising Diffusion Implicit Models.

MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation Denoising Diffusion Implicit Models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T19:38:20.527573Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:38:20.527573Z digest=sha256:e5231328c7c7cd524994e820ad12239d1a4da23b35eac3d9d4ddfc7d4ca5cd0e

Observation db8b93ee-57c1-445e-8a74-f7e50d88e5a7 · outbound

This paper cites Talking Face Generation by Conditional Recurrent Adversarial Network.

MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation Talking Face Generation by Conditional Recurrent Adversarial Network

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T19:38:20.530316Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:38:20.530316Z digest=sha256:3a5c12010dc535104ae47a0deb42ef9d392dadf2272cb45f76266650973dffa5

Observation 5e5556e3-48c2-4283-9fd3-23c6d9d927bd · outbound

This paper cites Diffused heads: Diffusion models beat gans on talking-face genera- tion.

MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation Diffused heads: Diffusion models beat gans on talking-face genera- tion

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:38:24.766718Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T19:38:20.533486Z digest=sha256:8b0eec07ffebc92953927af8be8e5116bb3d5a104323b4e796841a4cafa79872

Observation d31bcbb6-b8f8-42cb-b8a4-c4d34335d420 · outbound

This paper cites VividTalk: One-Shot Audio-Driven Talking Head Generation Based on 3D Hybrid Prior.

MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation VividTalk: One-Shot Audio-Driven Talking Head Generation Based on 3D Hybrid Prior

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T19:38:20.536182Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:38:20.536182Z digest=sha256:275b119128784bea4401582303a150874849a150a902a17dbf86133e1b71737b

Observation 3c3df29d-292f-4f0d-a801-1f09a062b89d · outbound

This paper cites Masked lip-sync prediction by audio-visual contextual exploitation in transformers.

MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation Masked lip-sync prediction by audio-visual contextual exploitation in transformers

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:38:24.596939Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T19:38:20.539264Z digest=sha256:dd81bfeebb9b59852557691efbfe802e313cad58b5a590e598ecbb53cbed29e3

Observation 3d2b0f69-728b-419e-8371-57b07ec3b469 · outbound

This paper cites Synthesizing obama: learn- ing lip sync from audio.

MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation Synthesizing obama: learn- ing lip sync from audio

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:38:24.416179Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T19:38:20.542272Z digest=sha256:e9a1da53857a91c304fb919048358e77edebe6d58ce4afe6182941d934293d3d

Observation e12a29a5-f076-4969-97b5-4afc1ff8e784 · outbound

This paper cites Human-centric founda- tion models: Perception, generation and agentic modeling.

MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation Human-centric founda- tion models: Perception, generation and agentic modeling

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:38:24.155157Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T19:38:20.545160Z digest=sha256:6ea7b70035fce99fd61c67da51caaf7d98ab0d3bf7e3af0a12eea9e0fb981f92

Observation 595754d1-34d3-431e-bb73-9d6e071cbc06 · outbound

This paper cites Neural voice puppetry: Audio-driven facial reenactment.

MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation Neural voice puppetry: Audio-driven facial reenactment

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:38:23.933754Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T19:38:20.547678Z digest=sha256:a2c6f3d4183c28a17990b053f5c6113c509b52c75cac34ab3b83706567774d9c

Observation 4a77f82c-db0b-40b5-b1ce-92998ccadec3 · outbound

This paper cites EMO: Emote Portrait Alive -- Generating Expressive Portrait Videos with Audio2Video Diffusion Model under Weak Conditions.

MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation EMO: Emote Portrait Alive -- Generating Expressive Portrait Videos with Audio2Video Diffusion Model under Weak Conditions

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T19:38:20.550356Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:38:20.550356Z digest=sha256:9e8967ed762c277d55572e41a7866dd6514a96ea9e57b6eb16562dd1be1683be

Observation 19dcfa08-89cb-4ef2-88dc-3f3cc904297e · outbound

This paper cites Xintao Wang, Honglun Zhang, Chao Dong, and Ying Shan.

MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation Xintao Wang, Honglun Zhang, Chao Dong, and Ying Shan

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:38:23.698180Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T19:38:20.553474Z digest=sha256:fd04091603a08ae220136f9ead63cc4aa28f48b06711e05970b8cd0dffd819f5

Observation 4654cc85-ca5c-4d2b-8181-5c1c68606216 · outbound

This paper cites AniPortrait: Audio-Driven Synthesis of Photorealistic Portrait Animation.

MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation AniPortrait: Audio-Driven Synthesis of Photorealistic Portrait Animation

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T19:38:20.556209Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:38:20.556209Z digest=sha256:f040334b61e62cd784c14bde2a78c328df7df9d37cc9a379cc5df2053be8cb38

Observation 507c5c7c-e994-4ad0-91d6-a0c91649c2de · outbound

This paper cites Photorealistic audio-driven video portraits.

MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation Photorealistic audio-driven video portraits

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:38:23.520340Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T19:38:20.559258Z digest=sha256:5547d2aa6b0862928d8c325a425de8fd7759984ecfc2d62b1169cf827bff0680

Observation bc2f2c89-b3ef-4d21-8fc4-bdd1eb374986 · outbound

This paper cites Monocular depth estimation using multi-scale continuous crfs as sequential deep networks.

MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation Monocular depth estimation using multi-scale continuous crfs as sequential deep networks

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:38:23.248574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T19:38:20.562097Z digest=sha256:4f676fab44c29341a74f0d62ff2c0bbcef69bfd2080fb902ef45f542faa709c1

Observation 4f8d4c40-5d73-4182-855e-64329f5e6bcd · outbound

This paper cites Hallo: Hierarchical Audio-Driven Visual Synthesis for Portrait Image Animation.

MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation Hallo: Hierarchical Audio-Driven Visual Synthesis for Portrait Image Animation

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T19:38:20.565284Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:38:20.565284Z digest=sha256:be44e9403e5221fb0e1ef9413fdd528222a6d820c1a851d6e38bc696d578219b

Observation d12e33c3-f04c-447c-9d90-d17c40df500b · outbound

This paper cites Jointly attentive spatial-temporal pooling networks for video-based person re-identification.

MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation Jointly attentive spatial-temporal pooling networks for video-based person re-identification

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:38:23.054479Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T19:38:20.568145Z digest=sha256:1456c79ae68ff175f0ca539f38c53d5815ba53ea4b2725ec98db161d68ad8b30

Observation dca78684-f3f4-463c-bd87-e6cd7a31d531 · outbound

This paper cites VASA-1: Lifelike Audio-Driven Talking Faces Generated in Real Time.

MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation VASA-1: Lifelike Audio-Driven Talking Faces Generated in Real Time

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-06T19:38:20.570734Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:38:20.570734Z digest=sha256:97915f38de18b145785082c0b72d9b49fc1b8abf790925ba08e8f919200c5a5b

Observation dff6ae6e-ce79-4f35-a1a7-bbd36534cfdc · outbound

This paper cites Sadtalker: Learning realistic 3d motion coefficients for stylized audio- driven single image talking face animation.

MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation Sadtalker: Learning realistic 3d motion coefficients for stylized audio- driven single image talking face animation

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:38:22.788464Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T19:38:20.573841Z digest=sha256:b17abaa04a8c8b1371b25954f7b575d52ca132b53d89cb88c84bc16b717bab55

Observation 75e3770d-6dcf-4c39-a418-8e21e67f5dbd · outbound

This paper cites Flow-guided one-shot talking face generation with a high- resolution audio-visual dataset.

MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation Flow-guided one-shot talking face generation with a high- resolution audio-visual dataset

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:38:22.536446Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T19:38:20.576418Z digest=sha256:b7f428cd73d5417eeb375ae8dbc4d5b6a23515f8cfa4050996440f47bb36a31b

Observation 5b552dea-6b52-4284-a90d-f2bb8f8158d8 · outbound

This paper cites Thin-plate spline motion model for image animation.

MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation Thin-plate spline motion model for image animation

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:38:22.300827Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T19:38:20.579388Z digest=sha256:f74b8a0e7b2f739d0a29da2176f04c204569dc594cd55799e8f3d8bc38909dbf

Observation be9364e6-806f-4294-96bd-d1152f147fb2 · outbound

This paper cites Synergizing motion and appearance: Multi-scale com- pensatory codebooks for talking head video generation.

MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation Synergizing motion and appearance: Multi-scale com- pensatory codebooks for talking head video generation

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:38:22.071391Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T19:38:20.582069Z digest=sha256:f6fe80bcd8e9d8b22d286cc71df5ddb7d1d218cda2889936b134fa33d4c0b87a

Observation 8869ee41-c585-4503-9ad9-5b28eb0f11cc · outbound

This paper cites Talking face generation by adversarially disentangled audio-visual representation.

MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation Talking face generation by adversarially disentangled audio-visual representation

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:38:21.866644Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T19:38:20.585041Z digest=sha256:20b0f9c46829a7d3b11d81ed47f222251fd01d64825553e9efd1e56c1a03ca3c

Observation 689b50cd-62c7-4faf-8a0e-6f3f39699a20 · outbound

This paper cites Pose-controllable talking face generation by implicitly modularized audio-visual rep- resentation.

MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation Pose-controllable talking face generation by implicitly modularized audio-visual rep- resentation

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:38:21.588387Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T19:38:20.588203Z digest=sha256:d7d93e792983a1d05e198e3b9223513a235fcd913ebad6c01d36ff23f4b0913c

Observation 16ce3c0d-c8d6-4417-98a8-ce28ee3543ca · outbound

This paper cites Makelttalk: speaker-aware talking-head animation.

MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation Makelttalk: speaker-aware talking-head animation

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:38:21.405568Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T19:38:20.590974Z digest=sha256:5f02c76c8e0ecb0f497d9b12f2ce4c32036af727e7e6ce02e2ee6f340bddd061

Pith citing papers

No inbound Pith citation observations are available.