Pith. sign in

Paper Citation Record · LEDGER

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation

As of 15 August 2026, this Paper Citation Record lists 100 of 112 outbound references and 0 inbound Pith citation observations for arXiv:2507.20953.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.20953 v1

Coverage vector

measured 100 of 112 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T13:10:19.735725Z

measured 100 of 100 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

100 of 112 outbound references displayed

  • verified exact7
  • verified fuzzy49
  • unresolved44
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c9836d1c-40de-4247-93fa-1901ad46d875 · outbound

This paper cites Deep audio-visual speech recognition.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Deep audio-visual speech recognition

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T13:10:18.321025Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:10:18.321025Z digest=sha256:ef9a2ceea0e81ad88102e74a1694d22d7ecbb4dce70383a531c2b0a87b1314f1

Observation e44250eb-0e52-4ba4-9828-3f35e7705e3a · outbound

This paper cites Self-supervised learning of audio- visual objects from video.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Self-supervised learning of audio- visual objects from video

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T13:10:18.394341Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:10:18.394341Z digest=sha256:19a2e7f3974c3593c04d0fce722787a57ca158bfdaba163f9b7abd8a5bdc3327

Observation a26a3b85-3e13-4604-9d4e-300a40fe3a18 · outbound

This paper cites A morphable model for the synthesis of 3d faces.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation A morphable model for the synthesis of 3d faces

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T13:10:18.435036Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:10:18.435036Z digest=sha256:87b2c7b6bf3bf261760247e93f30e1e613b13a263aa240c9b256e458ce504d74

Observation 47cec664-b9c1-47b1-8ea0-2516f105dff5 · outbound

This paper cites Large scale 3d mor- phable models.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Large scale 3d mor- phable models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T13:10:18.457608Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:10:18.457608Z digest=sha256:582fbc5c5090e9a08a7b7368c373869809a01beed0330da835d1a338034e638f

Observation aa7ec033-a8cd-4eb7-9a73-a9b8dc04ff0c · outbound

This paper cites V oice puppetry.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation V oice puppetry

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T13:10:18.673658Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:10:18.673658Z digest=sha256:a22898811b0a0ae33659337952ca879f521382760b6310cf16a504f3de9d5547

Observation 1247365f-f787-4d8f-ad42-c20fb3f75153 · outbound

This paper cites How far are we from solving the 2d & 3d face alignment problem? (and a dataset of 230,000 3d facial landmarks).

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation How far are we from solving the 2d & 3d face alignment problem? (and a dataset of 230,000 3d facial landmarks)

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T13:10:18.825656Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:10:18.825656Z digest=sha256:50aa8a0c4d3bae77219d30bf7ebaba04f5db159f3a044a202035ac8bf5be5f94

Observation 48638b90-8a57-4852-ad8b-26736454c962 · outbound

This paper cites JEAN: Joint Expression and Audio-guided NeRF-based Talking Face Generation.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation JEAN: Joint Expression and Audio-guided NeRF-based Talking Face Generation

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-08-06T13:10:19.966755Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T13:10:18.912385Z digest=sha256:0809e2c7cf5268a07dd9d8ccdeadde1590a4835726fe443d4d524d010d79a81f

Observation a2f7b28a-1a6f-4e2e-a351-f81ad5ff73ed · outbound

This paper cites TalkinNeRF: Animatable Neural Fields for Full-Body Talking Humans.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation TalkinNeRF: Animatable Neural Fields for Full-Body Talking Humans

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-08-06T13:10:19.954691Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T13:10:18.981498Z digest=sha256:a2e250a884cef53755d74d8bf8961b1d5f615c88180b5abce5c119d9728a16ed

Observation 2295c8f3-91b3-4243-955b-983ef24f91a1 · outbound

This paper cites Implicit neural head synthesis via controllable lo- cal deformation fields.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Implicit neural head synthesis via controllable lo- cal deformation fields

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T13:10:19.010196Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:10:19.010196Z digest=sha256:f740ba4b8de34126422f279704125e8f657cc22a4a66c86b094bf8a88ff9a285

Observation 73213bc5-db62-49a3-96f6-7311a9ae1884 · outbound

This paper cites Audio-Visual Synchronisation in the wild.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Audio-Visual Synchronisation in the wild

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T13:10:19.013378Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:10:19.013378Z digest=sha256:12e03e391d75f5a7104f3eb9478032035374ba4cc02d63c13e56daf3f534999a

Observation 3a41c1df-a95c-4af2-b709-2002971afbc2 · outbound

This paper cites Hierarchical cross-modal talking face generation with dynamic pixel-wise loss.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Hierarchical cross-modal talking face generation with dynamic pixel-wise loss

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T13:10:19.016638Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:10:19.016638Z digest=sha256:7f502b4e52b14ed30964929020fde0b56b71c13bd0fd179e515119483e2844a4

Observation 64367a9b-4d71-4470-8663-8f071082d597 · outbound

This paper cites Videoretalking: Audio-based lip synchronization for talking head video editing in the wild.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Videoretalking: Audio-based lip synchronization for talking head video editing in the wild

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T13:10:19.039371Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:10:19.039371Z digest=sha256:2cbf9b73cd33cfc1fbf2790f6ca00a8a7fe50b038691380bc1b5a79b1b40968b

Observation bff3e6aa-a911-44ba-a68b-1396f64f1d27 · outbound

This paper cites GPAvatar: Generalizable and Precise Head Avatar from Image(s).

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation GPAvatar: Generalizable and Precise Head Avatar from Image(s)

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T13:10:19.117914Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:10:19.117914Z digest=sha256:89422b2113e97ead60cdd66a48b71580832040d44ade9e4b8a7a1ff009a325ef

Observation ee007793-6166-4ab5-9bfc-17b23c410ea5 · outbound

This paper cites Lip reading in the wild.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Lip reading in the wild

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T13:10:19.197057Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:10:19.197057Z digest=sha256:d3071c8378f71b79eabfef5ad19cea63ee555b93ca10f9f5eb25be47535d176d

Observation de5759f6-daf1-4d47-ba93-af9702d9fe3b · outbound

This paper cites Out of time: auto- mated lip sync in the wild.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Out of time: auto- mated lip sync in the wild

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T13:10:19.251086Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:10:19.251086Z digest=sha256:efe63bc37b8850c41287f016d73b3458a87dbdfb3f6d1b8b37068dbe4b1ae7cf

Observation d7077656-b694-4b0e-a0b6-485bdf88205b · outbound

This paper cites Perfect match: Improved cross-modal embeddings for audio-visual synchronisation.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Perfect match: Improved cross-modal embeddings for audio-visual synchronisation

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T13:10:19.326692Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:10:19.326692Z digest=sha256:d978b835e3d1e1aba1d780401e29fc4ddf6b34ebcc0303d30477c04efbb8a6a0

Observation 7e629fbd-0271-4e8c-b7ea-2228696613d8 · outbound

This paper cites Speech-driven facial animation us- ing cascaded gans for learning of motion and texture.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Speech-driven facial animation us- ing cascaded gans for learning of motion and texture

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T13:10:19.385049Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:10:19.385049Z digest=sha256:8da205152bcbe8abdc3a557c95880cd7cb4bfcdfe38d6886939b7c249c6c666d

Observation c63e7167-c726-4d70-b92c-2039a471566e · outbound

This paper cites Imagenet: A large-scale hierarchical im- age database.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Imagenet: A large-scale hierarchical im- age database

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T13:10:19.459990Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:10:19.459990Z digest=sha256:ec5911fb19af07f6ab6d2ff1d6426a8b29be01198c493981d5e3e6b317b0dc51

Observation 947f87c2-5421-482a-8e89-119d03e4e5b5 · outbound

This paper cites Arcface: Additive angular margin loss for deep face recognition.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Arcface: Additive angular margin loss for deep face recognition

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T13:10:19.502108Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:10:19.502108Z digest=sha256:7cf8e7ca599e7b7e49360ff83907983cb4e34bf623ba5bc19e73ad7e0a54e744

Observation 8c0b5163-0e1f-4848-9af1-704a4430b9f4 · outbound

This paper cites End-to-end generation of talking faces from noisy speech.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation End-to-end generation of talking faces from noisy speech

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T13:10:19.508637Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:10:19.508637Z digest=sha256:c20277d9fad939f10ccd95ac7c863849dde583fd7a241744adebb8cd71a07c31

Observation b4dce184-be8c-4a22-89d1-ff2bb1490e0d · outbound

This paper cites Efficient emotional adaptation for audio- driven talking-head generation.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Efficient emotional adaptation for audio- driven talking-head generation

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T13:10:19.511638Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:10:19.511638Z digest=sha256:27db0e71995943eaa6135e08c0e4d59021a1875aa1faafb1626614bd2e801558

Observation 1b9b07f6-4d3c-4ce0-859b-12423c898c07 · outbound

This paper cites Generative adversarial nets.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Generative adversarial nets

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T13:10:19.514379Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:10:19.514379Z digest=sha256:bb35f9a564687092aa796a488b861b415d23622dc6b3e91e0d6e492c5af89309

Observation 9595d764-85bd-42be-9a1e-da5ec0d9a827 · outbound

This paper cites Stylesync: High-fidelity generalized and personalized lip sync in style-based gener- ator.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Stylesync: High-fidelity generalized and personalized lip sync in style-based gener- ator

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T13:10:19.517398Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:10:19.517398Z digest=sha256:fe35b29bc2c6702fce08d728bab3229ab1fe795e4567ee4ddef24fe8cd9415c5

Observation 7908e6e8-4050-4c77-9977-7b8a4c4f7175 · outbound

This paper cites Ad-nerf: Audio driven neural radiance fields for talking head synthesis.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Ad-nerf: Audio driven neural radiance fields for talking head synthesis

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T13:10:19.520614Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:10:19.520614Z digest=sha256:0400dc414aedc627a075889496e8b4fbaf2c72e93e447f7e5768c9a287ee9b20

Observation 93ac78b8-73f8-4c80-b831-ffbf666981d6 · outbound

This paper cites Audio vision: Using 9 audio-visual synchrony to locate sounds.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Audio vision: Using 9 audio-visual synchrony to locate sounds

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T13:10:19.523372Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:10:19.523372Z digest=sha256:98f57d87ac37baa640f3f4a2f6038b99386b22b4ef730440a2c06122cfe03eac

Observation 7b289745-c745-4903-9d2d-1fbc51f4e216 · outbound

This paper cites Gans trained by a two time-scale update rule converge to a local nash equilib- rium.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Gans trained by a two time-scale update rule converge to a local nash equilib- rium

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T13:10:19.525963Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:10:19.525963Z digest=sha256:ba614ddfb08e1039e179a0ba320c01c2437bb3b072d284b0f040cff70bdf2835

Observation d8848ea0-459e-4dc8-b06b-cf5cf66b28ee · outbound

This paper cites Diffted: One-shot audio-driven ted talk video generation with diffusion-based co-speech gestures.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Diffted: One-shot audio-driven ted talk video generation with diffusion-based co-speech gestures

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T13:10:19.528609Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:10:19.528609Z digest=sha256:edc6680d9ecafdba39ae08d15dc98129bc18fe4279022dbdb072746b5109d33d

Observation c59b91eb-fd0a-40c5-abaf-ef8fbfac12ae · outbound

This paper cites Arbitrary style transfer in real-time with adaptive instance normalization.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Arbitrary style transfer in real-time with adaptive instance normalization

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T13:10:19.531664Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:10:19.531664Z digest=sha256:d8d91d14ffe961b26c2d4d10d13cf8a1b5269cd559be8f3f47231adb4c51b1ec

Observation 04ca43ee-d2a3-4adf-b2d8-13e0320596ca · outbound

This paper cites Discohead: audio-and- video-driven talking head generation by disentangled con- trol of head pose and facial expressions.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Discohead: audio-and- video-driven talking head generation by disentangled con- trol of head pose and facial expressions

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T13:10:19.534433Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:10:19.534433Z digest=sha256:3b4e5687f0a222ac24f0f2d28a6fc5d1aa040bec982531cf656da68c3adb2fb4

Observation 96076c35-38b5-4dcd-b088-ec50f6f83c1c · outbound

This paper cites Sparse in Space and Time: Audio-visual Synchronisation with Trainable Selectors.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Sparse in Space and Time: Audio-visual Synchronisation with Trainable Selectors

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-08-06T13:10:19.927064Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T13:10:19.536817Z digest=sha256:d8890eda69e54836a0b0a54d423d51f25bf534d65ab5ac842b0d36f2089cfedc

Observation 2ac0bac6-07cc-42c9-b2b6-78c4779ac027 · outbound

This paper cites Batch normalization: Accelerating deep network training by reducing internal co- variate shift.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Batch normalization: Accelerating deep network training by reducing internal co- variate shift

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T13:10:19.539510Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:10:19.539510Z digest=sha256:9e423fae932bf44d01606b4abee28aafa6541c09621dcedeea77697036525354

Observation bf1738ef-787b-4d2b-9561-0d3ae75af01b · outbound

This paper cites You said that?: Synthesising talking faces from audio.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation You said that?: Synthesising talking faces from audio

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T13:10:19.542094Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:10:19.542094Z digest=sha256:b45c1c058224800c2ff91fdffd5b0c20b321a0c397660a1a217b4a414aee0d44

Observation f0bb7da1-c2ff-4dd5-afa3-d1d15ebcb4ea · outbound

This paper cites Audio-driven emotional video portraits.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Audio-driven emotional video portraits

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T13:10:19.544758Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:10:19.544758Z digest=sha256:651ef8801eb76ba53c69a0b930a89d4ab20c5d17206bfd7e0691023be5eb3c67

Observation 54b7fb74-5b2d-4f1e-bdc7-1389719fc14a · outbound

This paper cites Eamm: One-shot emotional talking face via audio-based emotion-aware motion model.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Eamm: One-shot emotional talking face via audio-based emotion-aware motion model

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T13:10:19.547228Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:10:19.547228Z digest=sha256:692f1f73caaaa45490481601e020620fc65c61a573f14811eb8752be6db9845e

Observation 50611682-95c7-4edc-a162-fca7cf7d985a · outbound

This paper cites Audio-driven facial an- imation with deep learning: A survey.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Audio-driven facial an- imation with deep learning: A survey

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T13:10:19.549789Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:10:19.549789Z digest=sha256:0df1963c3166025ebe7ec35e20875c7445ffa33dc5d8440201f0e49bd3c463df

Observation 74a4bae3-77ca-42db-85e4-5f3bc5a3a944 · outbound

This paper cites Percep- tual losses for real-time style transfer and super-resolution.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Percep- tual losses for real-time style transfer and super-resolution

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T13:10:19.552282Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:10:19.552282Z digest=sha256:6f85a5e1e3cad9e935558e332382848ab3b6c53da2c262559560a3e8678193d6

Observation 603d623b-998b-49fd-91cd-1f36294bb2f4 · outbound

This paper cites VocaLiST: An Audio-Visual Synchronisation Model for Lips and Voices.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation VocaLiST: An Audio-Visual Synchronisation Model for Lips and Voices

Reference 37

Resolution
verified exact
local_arxiv, observed 2026-08-06T13:10:19.916020Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T13:10:19.554988Z digest=sha256:17f5f65397bfe988cb3138e3b15d2bac43366a78cf58fc411fddc55d2ef439d2

Observation 63d3afe6-400e-415d-9453-343bf0cd2f9c · outbound

This paper cites NeRFFaceSpeech: One-shot Audio-driven 3D Talking Head Synthesis via Generative Prior.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation NeRFFaceSpeech: One-shot Audio-driven 3D Talking Head Synthesis via Generative Prior

Reference 38

Resolution
verified exact
local_arxiv, observed 2026-08-06T13:10:19.904549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T13:10:19.557924Z digest=sha256:2fa0c517ba16a30cfddd84c3ad674650fc2d97bdd3ad6874d6b90dc5f3461781

Observation 7d5ff702-a208-41f0-982e-a75fda6385cf · outbound

This paper cites End-to-end lip synchronisation based on pattern classification.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation End-to-end lip synchronisation based on pattern classification

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T13:10:19.561409Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:10:19.561409Z digest=sha256:32a19c41510c240180316d8f20ece27f1cdb1b70dfcd64f11cbb13d41522a64b

Observation 4fda48fa-ff56-435a-8872-805a678b6d6b · outbound

This paper cites Towards automatic face-to-face translation.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Towards automatic face-to-face translation

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:10:20.461263Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T13:10:19.564369Z digest=sha256:10879fa420df3074271267a7375ab9753fdcb566573e0fe1b5268f3bca0bfcaf

Observation 13883951-8b0d-44a1-9327-244bf9e0391e · outbound

This paper cites Imagenet classification with deep convolutional neural net- works.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Imagenet classification with deep convolutional neural net- works

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:10:20.453293Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T13:10:19.566904Z digest=sha256:a9a1f16b6cbafabca540442c1acd1c5dd0fa1a9e7590191d8dad700456f0c249

Observation 1d741386-e3a3-445c-960f-b667e71f8da9 · outbound

This paper cites Layer normalization.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Layer normalization

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:10:20.445443Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T13:10:19.569647Z digest=sha256:c8d588ffe40047eca4a5e4525996b294cd43b9e9aa38e39c8afc52d84df807c5

Observation 707b5b6f-89ec-46c4-8d69-dea8ab395aed · outbound

This paper cites Efficient region-aware neural radiance fields for high- fidelity talking portrait synthesis.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Efficient region-aware neural radiance fields for high- fidelity talking portrait synthesis

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:10:20.437092Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T13:10:19.572738Z digest=sha256:7085bd7f23326b7e484eddf64cde15dd6aa87900dcabcab1b6035893a08f1542

Observation cadf50fa-9c70-4cc5-a319-565e6e90dc60 · outbound

This paper cites One-shot high- fidelity talking-head synthesis with deformable neural ra- diance field.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation One-shot high- fidelity talking-head synthesis with deformable neural ra- diance field

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:10:20.429530Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T13:10:19.575782Z digest=sha256:c199d8e3756823cc4c02bf84688e4d86985444331dba9317ab3011eeb612b834

Observation b0690c38-67e4-4175-99a8-2eb20b746046 · outbound

This paper cites Expressive talking head generation with granular audio-visual control.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Expressive talking head generation with granular audio-visual control

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:10:20.420942Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T13:10:19.578661Z digest=sha256:aa53f7aa1ad88c287e48540355d22f1059887e5f830bfaff699465df5f7e23e7

Observation 54f85be3-39e3-498a-930b-e85474827d81 · outbound

This paper cites Font: Flow-guided one-shot talking head generation with natural head motions.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Font: Flow-guided one-shot talking head generation with natural head motions

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:10:20.412736Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T13:10:19.581524Z digest=sha256:0535c6b7b6422c3f30ef67b66f0165fcf8148e0be7bbd7c84e63bbdfe449b901

Observation 9320ef49-2fa4-4e36-a7b6-dfa1cd2b1997 · outbound

This paper cites Opt: One-shot pose- controllable talking head generation.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Opt: One-shot pose- controllable talking head generation

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:10:20.404441Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T13:10:19.585066Z digest=sha256:7ee3ee35cb399cfb52e6d792ceb5d42e0514a7250d3f262d5554582c2cd4a90c

Observation 34290a6a-be79-4b19-ac6f-e795ca15481a · outbound

This paper cites Semantic-aware implicit neural audio- driven video portrait generation.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Semantic-aware implicit neural audio- driven video portrait generation

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:10:20.396341Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T13:10:19.587601Z digest=sha256:94b6ba54f7c7a5539df083b2d42589aa5521ac6f7815f62f6eaa1e8a2c1f6903

Observation 926e99f2-d6c7-4f41-81ca-136480043d3d · outbound

This paper cites Moda: Mapping-once audio-driven portrait animation with dual attentions.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Moda: Mapping-once audio-driven portrait animation with dual attentions

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:10:20.389007Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T13:10:19.590321Z digest=sha256:60059a83b324c071e46be36992ac0d7a715f47f8a963ee1a78269631aba28b7b

Observation 94b78c77-e869-430a-a66f-388a4e2f3f1d · outbound

This paper cites MediaPipe: A Framework for Building Perception Pipelines.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation MediaPipe: A Framework for Building Perception Pipelines

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T13:10:19.592928Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:10:19.592928Z digest=sha256:aedd914ed110e96f5cf0c4f07dc4c9ed9b183e345fba3dac0b8e0d4c0cf862f3

Observation 884beea5-1ed5-4ba5-8f21-4e6dd446f871 · outbound

This paper cites Cvthead: One-shot controllable head avatar with vertex-feature transformer.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Cvthead: One-shot controllable head avatar with vertex-feature transformer

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:10:20.380403Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T13:10:19.596067Z digest=sha256:c70f72986cd7fbd08cb67294da6d1cd504325b021f82054e7ce0895226c4e964

Observation ba0a46d6-2ca8-4a44-85cf-0847073d3a51 · outbound

This paper cites Styletalk: One-shot talking head generation with controllable speak- ing styles.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Styletalk: One-shot talking head generation with controllable speak- ing styles

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:10:20.372221Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T13:10:19.598971Z digest=sha256:fe7d4d2a5f0407c1c621e53a1e3753acc38f3cb840f45af32a8ef0a9b4dfe534

Observation 288b8791-b9ae-48dd-b171-d59d1f9503f7 · outbound

This paper cites DreamTalk: When Emotional Talking Head Generation Meets Diffusion Probabilistic Models.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation DreamTalk: When Emotional Talking Head Generation Meets Diffusion Probabilistic Models

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-06T13:10:19.601406Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:10:19.601406Z digest=sha256:26ed946963b376b506ab5a30a4ff0627907fcb453e137c1f26e9dce8987dafd7

Observation 523113b3-b116-4e6e-9da3-443bf7f20502 · outbound

This paper cites Otavatar: One-shot talking face avatar with control- lable tri-plane rendering.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Otavatar: One-shot talking face avatar with control- lable tri-plane rendering

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:10:20.363238Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T13:10:19.604026Z digest=sha256:1033d66d11e1eaf7fcc0a3b9b612e0ab4f7324b0007ba5f423e13b94297273db

Observation f4e159a3-aadf-4dd3-8212-a8334a1fd1c0 · outbound

This paper cites Sidgan: High-resolution dubbed video generation via shift-invariant learning.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Sidgan: High-resolution dubbed video generation via shift-invariant learning

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:10:20.355365Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T13:10:19.607458Z digest=sha256:1ad6e9b2080ac108e9afaf7630bbc815387189cfba97e1f614c7874679275470

Observation 97f21d94-2da1-4794-be23-23bb058d756f · outbound

This paper cites Diff2lip: Audio conditioned dif- fusion models for lip-synchronization.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Diff2lip: Audio conditioned dif- fusion models for lip-synchronization

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:10:20.347539Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T13:10:19.609925Z digest=sha256:820942ffd38ad312e01bc3f83d7cfead752a4de837e9a0304d583028e3ecbafe

Observation e2764bf5-eb83-40be-88ca-e845cc200340 · outbound

This paper cites Rectified linear units improve restricted boltzmann machines.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Rectified linear units improve restricted boltzmann machines

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:10:20.339828Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T13:10:19.612889Z digest=sha256:ae8fc9893925f930eb8d47a1de7ad952d0b98dbbf8aafa675ed4685f26dea310

Observation dd662cb7-fed5-4e7c-a4bf-d28150096b28 · outbound

This paper cites Audio-visual scene analysis with self-supervised multisensory features.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Audio-visual scene analysis with self-supervised multisensory features

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:10:20.332060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T13:10:19.615652Z digest=sha256:49ec0bbac77b9dc91cff0e8f57fe8c15bfe13edda13f27499f292e26c529fe43

Observation ef00998f-538b-4937-add0-0ec3cf2e446e · outbound

This paper cites in-the-wild.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation in-the-wild

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:10:20.324159Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T13:10:19.619193Z digest=sha256:c80232e0285ea5698d606e77df6125723bbd08533fa7958ce9cf0ddfe65368a2

Observation 2f85de8b-b5c0-4c33-9255-2a73ed77ca1d · outbound

This paper cites Synctalkface: Talking face generation with precise lip-syncing via audio-lip memory.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Synctalkface: Talking face generation with precise lip-syncing via audio-lip memory

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:10:20.315937Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T13:10:19.621947Z digest=sha256:0453bb097c6375fa0cfa7cc778809f1d7e6e3cad738268008f94b772bf4d204c

Observation fd35efd7-792e-44ca-9eba-57fd936d6b3b · outbound

This paper cites Semantic image synthesis with spatially-adaptive normalization.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Semantic image synthesis with spatially-adaptive normalization

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:10:20.308577Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T13:10:19.624656Z digest=sha256:24a77539e71f385575e9aae35554b5fac6b1f11f15d629e849ab553cd2f46824

Observation ade94ba0-b672-461c-b057-199276a2257f · outbound

This paper cites Synctalk: The devil is in the synchronization for talking head synthesis.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Synctalk: The devil is in the synchronization for talking head synthesis

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:10:20.300514Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T13:10:19.627651Z digest=sha256:987dcd6fd2d12247a94bd4a1ca5fe14432c3076319fc014ce50e6f921fe0c14c

Observation 9cadb96b-feec-4e4b-b72f-224f52d8bfae · outbound

This paper cites A lip sync expert is all you need for speech to lip generation in the wild.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation A lip sync expert is all you need for speech to lip generation in the wild

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:10:20.292496Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T13:10:19.630189Z digest=sha256:acfc40edc05080309d252728763f3c5998a0622fed27406838e837bb782ad477

Observation 2da4e927-122b-47b4-b155-b6c22f891c9b · outbound

This paper cites U- net: Convolutional networks for biomedical image seg- mentation.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation U- net: Convolutional networks for biomedical image seg- mentation

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:10:20.284326Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T13:10:19.632941Z digest=sha256:6001c03df32a690fb38ca30b911117be75bb6f35289542b77b2da2b0054c8c08

Observation ad693093-af0d-40a4-bfb5-f85717dbeadc · outbound

This paper cites Learning dynamic facial radiance fields for few-shot talking head synthesis.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Learning dynamic facial radiance fields for few-shot talking head synthesis

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:10:20.275948Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T13:10:19.635850Z digest=sha256:6e30cc9a5b8807bf88cd4ffb20dd3734db3c06c54db45781c360b8f77418c4c3

Observation a28a1ad9-22f4-44ef-8aae-b48360df52ce · outbound

This paper cites Difftalk: Crafting dif- fusion models for generalized audio-driven portraits anima- tion.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Difftalk: Crafting dif- fusion models for generalized audio-driven portraits anima- tion

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:10:20.268356Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T13:10:19.639035Z digest=sha256:b489d5f873c5766ed60f04e06cf02156c0cc2f410efcce64b38426da7f0ab8e1

Observation 31a4aa87-2b76-4f62-ae70-d3c3653e449b · outbound

This paper cites Real-time single image and video super-resolution using an efficient sub-pixel convolutional neural network.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Real-time single image and video super-resolution using an efficient sub-pixel convolutional neural network

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:10:20.260481Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T13:10:19.642120Z digest=sha256:afd79c2fb6303c8184d6999b424985d53bec6648723c49a94c918ed75dc18c77

Observation fcc24d0e-7f31-4adc-8fb3-e90e3d0b7940 · outbound

This paper cites Very Deep Convolutional Networks for Large-Scale Image Recognition.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Very Deep Convolutional Networks for Large-Scale Image Recognition

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-06T13:10:19.645004Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:10:19.645004Z digest=sha256:52ecbf058fe7fa44632e1ea410dbe178a4836f17a40f091c79518645c3733298

Observation 4f7efac3-601e-40f6-be5a-1d29a0daefba · outbound

This paper cites Facesync: A linear operator for measuring synchronization of video facial im- ages and audio tracks.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Facesync: A linear operator for measuring synchronization of video facial im- ages and audio tracks

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:10:20.252056Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T13:10:19.647521Z digest=sha256:7e4405d72b596e408630bfa2606d9fda36f7d1f7a4b8a8f96d5796c013c5f082

Observation 56d1ce50-6a9a-4143-9da0-33072d0e0c4e · outbound

This paper cites Audio-driven dubbing for user generated contents via style-aware semi-parametric synthesis.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Audio-driven dubbing for user generated contents via style-aware semi-parametric synthesis

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:10:20.244244Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T13:10:19.650253Z digest=sha256:2e1aadb49e60c907649ef20f48e7d7aaaeeca2d273969525166e6ebf7528237a

Observation 13db41f6-5ae3-4c5c-986f-d5b61442860f · outbound

This paper cites Everybody’s talkin’: Let me talk as you want.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Everybody’s talkin’: Let me talk as you want

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:10:20.236266Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T13:10:19.653102Z digest=sha256:1670c7ebcc502e92a927228670d703f7b9f0d91d758f6bb8948a6fd681fcea55

Observation 3a95139c-8a27-4893-82b8-0682e13e607a · outbound

This paper cites Talking Face Generation by Conditional Recurrent Adversarial Network.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Talking Face Generation by Conditional Recurrent Adversarial Network

Reference 72

Resolution
verified exact
local_arxiv, observed 2026-08-06T13:10:19.869488Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T13:10:19.656203Z digest=sha256:2bbeffd532a0a8d29237b36afc586d2c2f40814ce8d526fe38e6ed6e8c4091ae

Observation 9c298006-4cc7-430b-9dd0-4d3cf4ff1767 · outbound

This paper cites Diffused 11 heads: Diffusion models beat gans on talking-face gen- eration.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Diffused 11 heads: Diffusion models beat gans on talking-face gen- eration

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:10:20.227469Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T13:10:19.659807Z digest=sha256:b8f7337afbce1d616550c093570b6ef2a94d9b8ebbfb907299c04b21b572dabf

Observation fa57b284-95aa-4720-8122-48f2245aff9c · outbound

This paper cites VividTalk: One-Shot Audio-Driven Talking Head Generation Based on 3D Hybrid Prior.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation VividTalk: One-Shot Audio-Driven Talking Head Generation Based on 3D Hybrid Prior

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-06T13:10:19.662786Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:10:19.662786Z digest=sha256:fbf7961a1e20f8950ebb774fceddb962b40955f6148f2d42402d10e1a8f26d2e

Observation 0af2d784-07f0-4183-afb5-af0fa7269a19 · outbound

This paper cites Masked lip-sync predic- tion by audio-visual contextual exploitation in transform- ers.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Masked lip-sync predic- tion by audio-visual contextual exploitation in transform- ers

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:10:20.218728Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T13:10:19.666211Z digest=sha256:82397ba9821ac52763587e7cd4dbd748b68e136649ff9c66d1aacf52baed622f

Observation 71220db6-5298-4d11-b709-face2ee95da0 · outbound

This paper cites Synthesizing obama: learning lip sync from audio.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Synthesizing obama: learning lip sync from audio

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:10:20.210090Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T13:10:19.668698Z digest=sha256:c2c6318a9f9108ce9b6bb658988fb018dc99402fc792b6cbc504c471a80f56c4

Observation 89baec36-0a0d-465a-9ae3-b5a4fbc037db · outbound

This paper cites Rethinking the inception architecture for computer vision.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Rethinking the inception architecture for computer vision

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:10:20.200831Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T13:10:19.671548Z digest=sha256:474538ac8930652b0674a3a9e9dcb7e98876c88ea7f58ae973789c2bfa248378

Observation 43a03a3b-b351-4e10-a744-7273bc44ee51 · outbound

This paper cites Emmn: Emotional mo- tion memory network for audio-driven emotional talking face generation.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Emmn: Emotional mo- tion memory network for audio-driven emotional talking face generation

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:10:20.192518Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T13:10:19.674028Z digest=sha256:7f66ae5a61954718b59725b914a0235898ab3c3aa4a4ddc5d7a832defceb574e

Observation 77cbb8cc-050a-41ed-9803-552063712921 · outbound

This paper cites Real-time Neural Radiance Talking Portrait Synthesis via Audio-spatial Decomposition.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Real-time Neural Radiance Talking Portrait Synthesis via Audio-spatial Decomposition

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-06T13:10:19.676863Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:10:19.676863Z digest=sha256:6b959221a4e72f9a5df7aa3c457425b14ce133c1e4666898b95758bb5f79e0ad

Observation ce345719-f4e4-4368-ab1e-eaafe44a5c34 · outbound

This paper cites Neural voice puppetry: Audio-driven facial reenactment.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Neural voice puppetry: Audio-driven facial reenactment

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:10:20.184100Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T13:10:19.679847Z digest=sha256:c22fddf06d2eef02ea252fc14e95f680b6cff0f6f29318a953c0249d2d4b7a9e

Observation ffec64f7-ca81-4640-bf4b-13c69139ce6f · outbound

This paper cites EMO: Emote Portrait Alive -- Generating Expressive Portrait Videos with Audio2Video Diffusion Model under Weak Conditions.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation EMO: Emote Portrait Alive -- Generating Expressive Portrait Videos with Audio2Video Diffusion Model under Weak Conditions

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-06T13:10:19.683024Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:10:19.683024Z digest=sha256:a4cef8ef7a2404e4b85f70a9ab33020bd3a8b5f93be9e97412b441091b775b13

Observation 5390878b-824a-4c0a-9588-953d1edbec11 · outbound

This paper cites Attention is all you need.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Attention is all you need

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-06T13:10:19.686259Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:10:19.686259Z digest=sha256:cd8b861257089a8206d0dd79988fd55f7ae669232e4fe287d2103948662de998

Observation 2c1cdc65-6158-44dc-8663-cc88fc4c5d81 · outbound

This paper cites Realistic speech-driven facial animation with gans.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Realistic speech-driven facial animation with gans

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:10:20.170908Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T13:10:19.688989Z digest=sha256:14b4ee3e78b29153243bf7e80f1df2361ed5899777f54dc0fb1f9bc9927806c5

Observation 67bbe2e9-9fd3-4856-b015-21f45598c9ab · outbound

This paper cites Progressive disentangled representa- tion learning for fine-grained controllable talking head syn- thesis.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Progressive disentangled representa- tion learning for fine-grained controllable talking head syn- thesis

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:10:20.162335Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T13:10:19.691728Z digest=sha256:f1f4ffc248f290a358ad8e849eaf613ef4b8120755eea81143332977fc516bbf

Observation e85864e7-bb3a-437a-8389-fef662dabc56 · outbound

This paper cites Seeing what you said: Talking face gen- eration guided by a lip reading expert.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Seeing what you said: Talking face gen- eration guided by a lip reading expert

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:10:20.154256Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T13:10:19.694660Z digest=sha256:ddce7d556fa7968696f1273c341cd0478b1ef71d4eecdedf4a67276d784ab35b

Observation 65ea565c-68cf-404b-a3ef-15deae8292cb · outbound

This paper cites Lipformer: High- fidelity and generalizable talking face generation with a pre- learned facial codebook.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Lipformer: High- fidelity and generalizable talking face generation with a pre- learned facial codebook

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:10:20.146253Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T13:10:19.697320Z digest=sha256:ae2a2d417ccc6b50ca290a4c5a29c29ea2c0ec0cd439abc0de421694cac3a221

Observation 8bdd2035-f607-4f0a-a3be-173ac2636cab · outbound

This paper cites Styletalk++: A unified framework for controlling the speaking styles of talking heads.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Styletalk++: A unified framework for controlling the speaking styles of talking heads

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:10:20.138768Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T13:10:19.699987Z digest=sha256:6d4ae0877d82e2cef18ce341fb48074edf6937ad8a9f72515e971af01d2d0340

Observation 0a3835a7-da50-4a24-8ade-78355e5da511 · outbound

This paper cites High-resolution image synthesis and semantic manipulation with conditional gans.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation High-resolution image synthesis and semantic manipulation with conditional gans

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:10:20.131528Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T13:10:19.702818Z digest=sha256:4407938d2f7ab3ebc1de7857d5ef5be58f8da5454b8641547843a212b03d2c91

Observation 4aba5721-85e4-468a-b946-2292b91a0e6b · outbound

This paper cites Image quality assessment: from error visibility to structural similarity.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Image quality assessment: from error visibility to structural similarity

Reference 89

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:10:20.123405Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T13:10:19.706016Z digest=sha256:4230978654dfd591238080c254275b2da9ef78e9a6a1cc647a13ce03e9d4b212

Observation 5c688292-ed0f-45b3-8365-a39e22ecab81 · outbound

This paper cites Imitating arbitrary talking style for realistic audio-driven talking face synthesis.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Imitating arbitrary talking style for realistic audio-driven talking face synthesis

Reference 90

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:10:20.115487Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T13:10:19.708385Z digest=sha256:07c3e46f71dde5a606f03001097e9316e7d7f8bedfe22ea59aaeebc7dcd21e39

Observation e2f6bcc4-c996-4de0-894e-64a8dfd382c9 · outbound

This paper cites Ganhead: Towards generative animatable neural head avatars.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Ganhead: Towards generative animatable neural head avatars

Reference 91

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:10:20.107368Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T13:10:19.711283Z digest=sha256:01b13967c7fe46c7821b8e001a8bf508a603e889d97abcb084132a26820440d5

Observation 29b07864-bf89-421b-b1bb-15f1d78bcb0d · outbound

This paper cites High-fidelity generalized emotional talking face generation with multi-modal emotion space learning.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation High-fidelity generalized emotional talking face generation with multi-modal emotion space learning

Reference 92

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:10:20.099524Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T13:10:19.713845Z digest=sha256:7d3e8fbb8e9fa3ffdcdb49b22299b0e765985d040a097be5612502dca0115c02

Observation ad5563d8-07ce-40b8-a753-308cf54b5653 · outbound

This paper cites Audio-visual speech representation expert for enhanced talking face video generation and evaluation.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Audio-visual speech representation expert for enhanced talking face video generation and evaluation

Reference 93

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:10:20.092115Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T13:10:19.716184Z digest=sha256:895d9efce517efe3ca264006d040fe34544ede69a6db0e97a319ebb3efcd0589

Observation 2342329e-7e22-4ced-8090-c566a2e7bfde · outbound

This paper cites Audio-driven Talking Face Generation with Stabilized Synchronization Loss.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Audio-driven Talking Face Generation with Stabilized Synchronization Loss

Reference 94

Resolution
verified exact
local_arxiv, observed 2026-08-06T13:10:19.832891Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T13:10:19.718719Z digest=sha256:9cdce03806755e5610e18d4a3582ba2ac14242ecfae5a22b692b4d3a96f0552b

Observation 354d3f85-f09d-4b74-8eef-cf5b8850e70e · outbound

This paper cites DFA-NeRF: Personalized Talking Head Generation via Disentangled Face Attributes Neural Rendering.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation DFA-NeRF: Personalized Talking Head Generation via Disentangled Face Attributes Neural Rendering

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-06T13:10:19.721450Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:10:19.721450Z digest=sha256:42507132c8aeac667a449abe9ef48875bd4cd882eecf8e057fae6dc29411f23f

Observation 47656887-e25c-4bb2-be3e-a6f50dc87109 · outbound

This paper cites GeneFace: Generalized and High-Fidelity Audio-Driven 3D Talking Face Synthesis.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation GeneFace: Generalized and High-Fidelity Audio-Driven 3D Talking Face Synthesis

Reference 96

Resolution
unresolved
no resolver link, observed 2026-08-06T13:10:19.724040Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:10:19.724040Z digest=sha256:a215401bd3132bf067dffcfa229f4b6c7859c2eddb92eca4dd4ccf8723a4983c

Observation 06b9f2ff-c777-4e0c-b011-a9a0bec6255e · outbound

This paper cites Real3D-Portrait: One-shot Realistic 3D Talking Portrait Synthesis.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Real3D-Portrait: One-shot Realistic 3D Talking Portrait Synthesis

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-06T13:10:19.726995Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:10:19.726995Z digest=sha256:2c73dc5d94bdfa22031b28005f81343f4c594fa5735c6c5f091f752896ba932f

Observation d27549de-85a5-4933-be4e-b151ba5d8dcd · outbound

This paper cites Quantitative association of vocal-tract and facial behavior.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Quantitative association of vocal-tract and facial behavior

Reference 98

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:10:20.083847Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T13:10:19.729858Z digest=sha256:0da2ed48bee476724a6125e883974ccf85503fa893289ea5e94c25bde07f9968

Observation 16907925-7056-447c-990a-e7fc25da1dee · outbound

This paper cites Styleheat: One-shot high-resolution editable talking face generation via pre-trained stylegan.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Styleheat: One-shot high-resolution editable talking face generation via pre-trained stylegan

Reference 99

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:10:20.076042Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T13:10:19.732934Z digest=sha256:e5862c2406ea44486a8b9404311c4633eaf5d41576749463e88d8b4289f86fb3

Observation cdd47e1d-af3e-44e5-bc7b-6df6cbaf5e97 · outbound

This paper cites Multimodal image synthesis and editing: A survey and taxonomy.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Multimodal image synthesis and editing: A survey and taxonomy

Reference 100

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:10:20.068461Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T13:10:19.735725Z digest=sha256:81a9e30bdf137a4a9dd7adb4e20b1ed5213e8325ece6e457348ffed77a871d72

Pith citing papers

No inbound Pith citation observations are available.