Pith. sign in

Paper Citation Record · LEDGER

MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation

As of 10 August 2026, this Paper Citation Record lists 54 of 54 outbound references and 0 inbound Pith citation observations for arXiv:2507.05092.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.05092 v1

Coverage vector

measured 54 of 54 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T19:38:20.590974Z

measured 54 of 54 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

54 of 54 outbound references displayed

  • verified exact3
  • verified fuzzy37
  • unresolved14
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 8cee1c3d-4d3b-4c04-bcf2-587a78f80c85 · outbound

This paper cites A morphable model for the synthesis of 3d faces.

MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation A morphable model for the synthesis of 3d faces

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:38:30.736538Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T19:38:20.430133Z digest=sha256:b1d37dc0bd55c15d4d2b27038354c76bb48ec25c1c07e37ab0edf1a246c925f1

Observation b4f8edd8-0cc6-47df-afb1-b065500818bb · outbound

This paper cites Hierarchical cross-modal talking face generation with dynamic pixel-wise loss.

MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation Hierarchical cross-modal talking face generation with dynamic pixel-wise loss

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:38:30.513004Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T19:38:20.433575Z digest=sha256:964a80bb713a59479c05eab8143791f3fc5149cc09886b489bb377cb917050d3

Observation fc95a1b9-b64f-41a4-bf41-efd907fe8f5c · outbound

This paper cites You said that?.

MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation You said that?

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-08-06T19:38:21.116132Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T19:38:20.437079Z digest=sha256:91b8a8a15e67128f79c2517ebd7410c0edb92d2bd93f2c0e605d6051aec96897

Observation b4d553b6-7fa0-4490-9a37-b2ebf5c3621f · outbound

This paper cites Lip reading sentences in the wild.

MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation Lip reading sentences in the wild

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:38:30.189584Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T19:38:20.440540Z digest=sha256:7c6e6a664cfbcc41d41a61806a268f3cdecc23255f439812586052b958f692ce

Observation 23f4c44d-469a-4e9a-879e-8adeea596953 · outbound

This paper cites Accurate 3d face reconstruction with weakly-supervised learning: From single image to image set.

MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation Accurate 3d face reconstruction with weakly-supervised learning: From single image to image set

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:38:29.858375Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T19:38:20.443919Z digest=sha256:664123679f937e363af9630a5a93335f94fe0b6547664b8353c575ace5b7f85e

Observation 1d9fa577-2ed1-4238-a56d-6653311e1b01 · outbound

This paper cites Dae-talker: High fidelity speech-driven talking face generation with diffusion autoen- coder.

MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation Dae-talker: High fidelity speech-driven talking face generation with diffusion autoen- coder

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:38:29.574230Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T19:38:20.446762Z digest=sha256:085716485966f9ec7517a949a208688e5ac124422f66f09890c37e2a78f79a35

Observation d3848b8e-6786-49dc-92a1-bafbcdb3aedf · outbound

This paper cites Generative adversarial networks.

MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation Generative adversarial networks

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T19:38:20.449788Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:38:20.449788Z digest=sha256:761563d0302ab0d664da857597f5c2cc052b71489800b88af6f8532288fdef88

Observation 5fe11f5c-a2a4-4837-a253-a6c786e08c48 · outbound

This paper cites Ad-nerf: Audio driven neural radi- ance fields for talking head synthesis.

MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation Ad-nerf: Audio driven neural radi- ance fields for talking head synthesis

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:38:29.242496Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T19:38:20.452970Z digest=sha256:4540894ebe04868e8d2bda0e275b514f1cdab270678065ffc2629d21a81b0728

Observation 03a3c0e8-84a5-4ecf-810d-a754cbdf8bd8 · outbound

This paper cites Facexhubert: Text-less speech-driven e (x) pressive 3d facial animation synthesis using self-supervised speech representation learn- ing.

MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation Facexhubert: Text-less speech-driven e (x) pressive 3d facial animation synthesis using self-supervised speech representation learn- ing

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:38:28.887210Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T19:38:20.455770Z digest=sha256:9d3e4129341af9bd479ab780fd39ef1607d07502701c7e5076c7a4f560a2cf15

Observation 7c1bd97c-5501-4a64-aa92-1ca5b3f01f88 · outbound

This paper cites GAIA: Zero-shot Talking Avatar Generation.

MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation GAIA: Zero-shot Talking Avatar Generation

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-08-06T19:38:20.866922Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T19:38:20.458871Z digest=sha256:dbac38070650436becda31c25157988036e486c8c42416f718a93abc4602cbcc

Observation f1a09b4b-2371-44f8-b2bc-3d9669cd06b6 · outbound

This paper cites Gans trained by a two time-scale update rule converge to a local nash equilib- rium.

MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation Gans trained by a two time-scale update rule converge to a local nash equilib- rium

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:38:28.593769Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T19:38:20.462422Z digest=sha256:0d7a148a06b187a80e79ab3bc5859a088202a1d833e58d56517687464277e489

Observation 0328e1ad-3b23-491a-b9fe-563c3d8b4f84 · outbound

This paper cites Implicit identity representation conditioned memory compensation network for talking head video generation.

MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation Implicit identity representation conditioned memory compensation network for talking head video generation

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:38:28.281921Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T19:38:20.465191Z digest=sha256:a9b6601e809de8f7dd9588edec17613632a7a052ae31e22327660f15124295f6

Observation fe1ef8fa-8c66-4cda-a6bb-dd4350f8f3f8 · outbound

This paper cites DaGAN++: Depth-Aware Generative Adversarial Network for Talking Head Video Generation.

MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation DaGAN++: Depth-Aware Generative Adversarial Network for Talking Head Video Generation

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-08-06T19:38:20.756341Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T19:38:20.467907Z digest=sha256:ad677cbd939e362dfa474598ae6bbda403385bb20bb7a8bc345e753d943c095d

Observation 167f1b13-1462-49b2-9c92-1bc5f0dac928 · outbound

This paper cites Depth-aware generative adversarial network for talking head video generation.

MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation Depth-aware generative adversarial network for talking head video generation

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:38:28.009063Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T19:38:20.471010Z digest=sha256:1ac50b0cdb5b657c3cbd047e189f3f71d81b3a63b20168526284abb647347b89

Observation aa640803-062e-4a2a-9704-f5eec816209f · outbound

This paper cites Audio-visual Controlled Video Diffusion with Masked Selective State Spaces Modeling for Natural Talking Head Generation.

MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation Audio-visual Controlled Video Diffusion with Masked Selective State Spaces Modeling for Natural Talking Head Generation

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T19:38:20.474478Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:38:20.474478Z digest=sha256:669a0754aebc72bd99ac49a7d768761ea1282407becfc305277d5a1687124068

Observation 84f58e79-35e0-4779-a77f-f2f8913b2b09 · outbound

This paper cites Audio-driven emotional video portraits.

MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation Audio-driven emotional video portraits

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:38:27.781678Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T19:38:20.481151Z digest=sha256:ba4721130c7e14c278e32ded22e6b5cb936304e08e4c2b2307d1ca538e39d9f7

Observation 8243d0ef-8a90-4b24-a7cf-e347d3ca4fa7 · outbound

This paper cites Transformers in vision: A survey.

MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation Transformers in vision: A survey

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:38:27.506549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T19:38:20.484001Z digest=sha256:25012e2501921ad44d6178a7a8024a59cba116049454f00f060aa8eeb2c6823c

Observation d4d0c738-bfa5-4595-a262-3ad63505262d · outbound

This paper cites Deep video portraits.

MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation Deep video portraits

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:38:27.275743Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T19:38:20.486816Z digest=sha256:0aa5dfaa6c4cbc124bb9215d4069a407541e3ffcdcfe467e84adea06a8ba699e

Observation b6204282-38eb-45d7-8ddf-72ca548fb991 · outbound

This paper cites Auto-Encoding Variational Bayes.

MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation Auto-Encoding Variational Bayes

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T19:38:20.489513Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:38:20.489513Z digest=sha256:b130c477cc6be75e6f649512f1ffd8d57f5511bf8a6b34cd9835d98f912c95c9

Observation e7716d64-440d-4cf7-8da5-33d4e7156fef · outbound

This paper cites AniTalker: Animate Vivid and Diverse Talking Faces through Identity-Decoupled Facial Motion Encoding.

MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation AniTalker: Animate Vivid and Diverse Talking Faces through Identity-Decoupled Facial Motion Encoding

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T19:38:20.492669Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:38:20.492669Z digest=sha256:e5c366020d00eeb1a2ab04a5e379c18c576ca7da75938ee2c1bf06a9c478e243

Observation 6230e1cb-f0f8-44d4-910c-4e30cc395eec · outbound

This paper cites Moda: Mapping-once audio-driven portrait animation with dual attentions.

MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation Moda: Mapping-once audio-driven portrait animation with dual attentions

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:38:26.994177Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T19:38:20.496048Z digest=sha256:32a7e6220b0031b0cdff2ccb1c7e0c4d5cd2d6877a130c16c839635b0b1fdff4

Observation cc2d224c-570b-4ae8-9600-2000bb2cc771 · outbound

This paper cites Live speech por- traits: real-time photorealistic talking-head animation.

MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation Live speech por- traits: real-time photorealistic talking-head animation

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:38:26.776730Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T19:38:20.498870Z digest=sha256:40da9a2b6f80170cbf462916559b5c5f3d727c907baf3e2b84abe8647efe4d68

Observation 602d3749-d771-43e3-9f6e-6b43b772e56d · outbound

This paper cites Training strategies for improved lip- reading.

MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation Training strategies for improved lip- reading

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:38:26.437494Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T19:38:20.501957Z digest=sha256:2b8fa219871519b9a646747868632c0782d0b716a7e90d3c8118f84f198ae14d

Observation 3d446572-f400-4b64-92f0-d6db6a9cd57a · outbound

This paper cites DreamTalk: When Emotional Talking Head Generation Meets Diffusion Probabilistic Models.

MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation DreamTalk: When Emotional Talking Head Generation Meets Diffusion Probabilistic Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T19:38:20.504641Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:38:20.504641Z digest=sha256:9e275e737f157bd91f15d703c7872210660995afb133cfffb2e7a4918d10062d

Observation 9c2e43db-8b8a-431b-b31b-810aef76b4fc · outbound

This paper cites DiffSpeaker: Speech-Driven 3D Facial Animation with Diffusion Transformer.

MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation DiffSpeaker: Speech-Driven 3D Facial Animation with Diffusion Transformer

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T19:38:20.507706Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:38:20.507706Z digest=sha256:e28d3543cc4aa19536f2cd409f2e3683a9934332ccc25f1ef420ce3cb7b7d9a0

Observation 0049d14a-a49a-4224-91ae-ef785403e837 · outbound

This paper cites Librispeech: An asr corpus based on public do- main audio books.

MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation Librispeech: An asr corpus based on public do- main audio books

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:38:26.182893Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T19:38:20.510785Z digest=sha256:5ba351b5a5963e48fae5fdc2c60cfe0e4a67af1aed443d0f0ab742088d71f00f

Observation 66454920-a701-4407-b670-4873c58a61a8 · outbound

This paper cites A lip sync expert is all you need for speech to lip generation in the wild.

MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation A lip sync expert is all you need for speech to lip generation in the wild

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:38:25.916830Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T19:38:20.513555Z digest=sha256:97084b33ff912bde537152b342943d3d27d0fc3e13f06c5fdddfc88105a64540

Observation 4d84c517-e4cb-4281-bbe7-e67d57a618de · outbound

This paper cites High-resolution image syn- thesis with latent diffusion models.

MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation High-resolution image syn- thesis with latent diffusion models

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:38:25.657706Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T19:38:20.516395Z digest=sha256:a3debb588baf87938c4bfff09e276e40d2c38d2ff84a242648a12f236d872da4

Observation 90859342-df51-45a5-affb-1005a4294afc · outbound

This paper cites pytorch-fid: FID Score for PyTorch.

MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation pytorch-fid: FID Score for PyTorch

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T19:38:20.519303Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:38:20.519303Z digest=sha256:29a8f82acd77d014788c8e45669a86e5e2c7b41de25ecf5a7e83d955daeeca66

Observation 182a4583-fc6d-40a7-98af-a6d4a49ca6cb · outbound

This paper cites Difftalk: Crafting diffusion models for generalized audio-driven portraits animation.

MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation Difftalk: Crafting diffusion models for generalized audio-driven portraits animation

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:38:25.321835Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T19:38:20.522395Z digest=sha256:3645dee8bc414c07d829e99715a8939cd9d7e2871720e1b77186e9920215e627

Observation eab7c841-fa7b-4c58-96b0-e37a38906e01 · outbound

This paper cites First order motion model for image animation.

MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation First order motion model for image animation

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:38:25.005868Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T19:38:20.525008Z digest=sha256:3b9c7d8663319365512de431f20eaf9f804eadbdbd2fcdabf6e8a3f99b28c093

Observation 6b328f7e-e45c-473a-b45c-e6a49c2fb524 · outbound

This paper cites Denoising Diffusion Implicit Models.

MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation Denoising Diffusion Implicit Models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T19:38:20.527573Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:38:20.527573Z digest=sha256:7383cb134d070878636b063b74d38e537acf4efe80e1064edf92d62c13c07019

Observation db8b93ee-57c1-445e-8a74-f7e50d88e5a7 · outbound

This paper cites Talking Face Generation by Conditional Recurrent Adversarial Network.

MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation Talking Face Generation by Conditional Recurrent Adversarial Network

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T19:38:20.530316Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:38:20.530316Z digest=sha256:c39af15f2154b7bc1be0225af8cc1ec43f4f8f1c27bc2fc83095cf1dd7ca2a28

Observation 5e5556e3-48c2-4283-9fd3-23c6d9d927bd · outbound

This paper cites Diffused heads: Diffusion models beat gans on talking-face genera- tion.

MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation Diffused heads: Diffusion models beat gans on talking-face genera- tion

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:38:24.766718Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T19:38:20.533486Z digest=sha256:cfc8680932fbf169e482931f526d495e1ecdc2031b6e38dbeac15a53bd91bc84

Observation d31bcbb6-b8f8-42cb-b8a4-c4d34335d420 · outbound

This paper cites VividTalk: One-Shot Audio-Driven Talking Head Generation Based on 3D Hybrid Prior.

MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation VividTalk: One-Shot Audio-Driven Talking Head Generation Based on 3D Hybrid Prior

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T19:38:20.536182Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:38:20.536182Z digest=sha256:12cc918c2ed4f1234b91dda4d251ed5889ac29368f2a1337b8ea0ceef8f45376

Observation 3c3df29d-292f-4f0d-a801-1f09a062b89d · outbound

This paper cites Masked lip-sync prediction by audio-visual contextual exploitation in transformers.

MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation Masked lip-sync prediction by audio-visual contextual exploitation in transformers

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:38:24.596939Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T19:38:20.539264Z digest=sha256:d0f554aef87793cce3f9ff6595c6e10dd7800c3caad75b700bf09c519cbc26c9

Observation 3d2b0f69-728b-419e-8371-57b07ec3b469 · outbound

This paper cites Synthesizing obama: learn- ing lip sync from audio.

MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation Synthesizing obama: learn- ing lip sync from audio

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:38:24.416179Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T19:38:20.542272Z digest=sha256:084f5b008b0784664041e46ac2927794c046eb06a3d5ef391d8ed3936933f6d5

Observation e12a29a5-f076-4969-97b5-4afc1ff8e784 · outbound

This paper cites Human-centric founda- tion models: Perception, generation and agentic modeling.

MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation Human-centric founda- tion models: Perception, generation and agentic modeling

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:38:24.155157Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T19:38:20.545160Z digest=sha256:2e36621e23dc9acac808c5f5966fdf398a16131da556b7fc282e571445f66a4f

Observation 595754d1-34d3-431e-bb73-9d6e071cbc06 · outbound

This paper cites Neural voice puppetry: Audio-driven facial reenactment.

MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation Neural voice puppetry: Audio-driven facial reenactment

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:38:23.933754Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T19:38:20.547678Z digest=sha256:773fed88dd6ec630b751ba38144f3c43f7f75748fdf935a0fd753977ac583d2a

Observation 4a77f82c-db0b-40b5-b1ce-92998ccadec3 · outbound

This paper cites EMO: Emote Portrait Alive -- Generating Expressive Portrait Videos with Audio2Video Diffusion Model under Weak Conditions.

MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation EMO: Emote Portrait Alive -- Generating Expressive Portrait Videos with Audio2Video Diffusion Model under Weak Conditions

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T19:38:20.550356Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:38:20.550356Z digest=sha256:e5542d9d34efdd4371d9e5e41151828faeb1bc2cd0c3cbb36cff943897d0453a

Observation 19dcfa08-89cb-4ef2-88dc-3f3cc904297e · outbound

This paper cites Xintao Wang, Honglun Zhang, Chao Dong, and Ying Shan.

MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation Xintao Wang, Honglun Zhang, Chao Dong, and Ying Shan

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:38:23.698180Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T19:38:20.553474Z digest=sha256:a7db56ae38db7099e51ad55e4857a772ea2024c1dcd4230ac978b53a5d616c6a

Observation 4654cc85-ca5c-4d2b-8181-5c1c68606216 · outbound

This paper cites AniPortrait: Audio-Driven Synthesis of Photorealistic Portrait Animation.

MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation AniPortrait: Audio-Driven Synthesis of Photorealistic Portrait Animation

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T19:38:20.556209Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:38:20.556209Z digest=sha256:e3d314ad30236bab0aea8b1f8e3ec9865cde9039d69abd45ec2e9f5546b9c0d4

Observation 507c5c7c-e994-4ad0-91d6-a0c91649c2de · outbound

This paper cites Photorealistic audio-driven video portraits.

MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation Photorealistic audio-driven video portraits

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:38:23.520340Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T19:38:20.559258Z digest=sha256:e3372dcf1cb4c71108fd62693fea6458af1ead9a5ad5b2132f5d37cf3cc1e161

Observation bc2f2c89-b3ef-4d21-8fc4-bdd1eb374986 · outbound

This paper cites Monocular depth estimation using multi-scale continuous crfs as sequential deep networks.

MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation Monocular depth estimation using multi-scale continuous crfs as sequential deep networks

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:38:23.248574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T19:38:20.562097Z digest=sha256:e3bb595e5544dd9de0a1838868d412499a282bfe06518d9caa96c33e63e603f1

Observation 4f8d4c40-5d73-4182-855e-64329f5e6bcd · outbound

This paper cites Hallo: Hierarchical Audio-Driven Visual Synthesis for Portrait Image Animation.

MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation Hallo: Hierarchical Audio-Driven Visual Synthesis for Portrait Image Animation

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T19:38:20.565284Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:38:20.565284Z digest=sha256:bdda89520597cb54ddfdab568f1da031b60e8e524722dd33327d1b7b98a6a175

Observation d12e33c3-f04c-447c-9d90-d17c40df500b · outbound

This paper cites Jointly attentive spatial-temporal pooling networks for video-based person re-identification.

MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation Jointly attentive spatial-temporal pooling networks for video-based person re-identification

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:38:23.054479Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T19:38:20.568145Z digest=sha256:e0393e0fb4de20289ca906f79454032b98227eb6b1e2d61686b55e1bbfa57797

Observation dca78684-f3f4-463c-bd87-e6cd7a31d531 · outbound

This paper cites VASA-1: Lifelike Audio-Driven Talking Faces Generated in Real Time.

MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation VASA-1: Lifelike Audio-Driven Talking Faces Generated in Real Time

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-06T19:38:20.570734Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:38:20.570734Z digest=sha256:6ef92ef0223a6c5f360877ef41e6876bde82ab1f4e654da3d5fc0e28b364ecf3

Observation dff6ae6e-ce79-4f35-a1a7-bbd36534cfdc · outbound

This paper cites Sadtalker: Learning realistic 3d motion coefficients for stylized audio- driven single image talking face animation.

MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation Sadtalker: Learning realistic 3d motion coefficients for stylized audio- driven single image talking face animation

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:38:22.788464Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T19:38:20.573841Z digest=sha256:58f2d6da0cf759a4aef1b2b5bb9c6920b9b86d906067dc20db827484f6e21271

Observation 75e3770d-6dcf-4c39-a418-8e21e67f5dbd · outbound

This paper cites Flow-guided one-shot talking face generation with a high- resolution audio-visual dataset.

MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation Flow-guided one-shot talking face generation with a high- resolution audio-visual dataset

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:38:22.536446Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T19:38:20.576418Z digest=sha256:0efe9ef7128a614eecaaf7331849a1794d3c749bc9744d30074044a9d1e250b7

Observation 5b552dea-6b52-4284-a90d-f2bb8f8158d8 · outbound

This paper cites Thin-plate spline motion model for image animation.

MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation Thin-plate spline motion model for image animation

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:38:22.300827Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T19:38:20.579388Z digest=sha256:eccac65c62d49928e7fdb1e9bccba8bff88c717dbe4fdb4dff5654f056c1fe68

Observation be9364e6-806f-4294-96bd-d1152f147fb2 · outbound

This paper cites Synergizing motion and appearance: Multi-scale com- pensatory codebooks for talking head video generation.

MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation Synergizing motion and appearance: Multi-scale com- pensatory codebooks for talking head video generation

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:38:22.071391Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T19:38:20.582069Z digest=sha256:ca3551afd7c2eb8c6d9f3adc11a4e8aa02e89c12b764ad42b8b79e0b86eb6134

Observation 8869ee41-c585-4503-9ad9-5b28eb0f11cc · outbound

This paper cites Talking face generation by adversarially disentangled audio-visual representation.

MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation Talking face generation by adversarially disentangled audio-visual representation

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:38:21.866644Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T19:38:20.585041Z digest=sha256:32ccb4e4b69760bfaa9b364fc6c095ae81259e0edae07e4de1db50fbccea0dd3

Observation 689b50cd-62c7-4faf-8a0e-6f3f39699a20 · outbound

This paper cites Pose-controllable talking face generation by implicitly modularized audio-visual rep- resentation.

MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation Pose-controllable talking face generation by implicitly modularized audio-visual rep- resentation

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:38:21.588387Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T19:38:20.588203Z digest=sha256:53682b0362fcb46d31381c4ca048a9bc93aacd87ebc11eb5f9fce2284feb7be1

Observation 16ce3c0d-c8d6-4417-98a8-ce28ee3543ca · outbound

This paper cites Makelttalk: speaker-aware talking-head animation.

MoDiT: Learning Highly Consistent 3D Motion Coefficients with Diffusion Transformer for Talking Head Generation Makelttalk: speaker-aware talking-head animation

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:38:21.405568Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T19:38:20.590974Z digest=sha256:28c41a03537886d501dad72742f1777d30db8da5b61b7ab4953a65562fd30565

Pith citing papers

No inbound Pith citation observations are available.