Pith. sign in

Paper Citation Record · LEDGER

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation

As of 18 August 2026, this Paper Citation Record lists 100 of 112 outbound references and 0 inbound Pith citation observations for arXiv:2507.20953.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.20953 v1

Coverage vector

measured 100 of 112 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T13:10:19.735725Z

measured 100 of 100 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

100 of 112 outbound references displayed

  • verified exact7
  • verified fuzzy49
  • unresolved44
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c9836d1c-40de-4247-93fa-1901ad46d875 · outbound

This paper cites Deep audio-visual speech recognition.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Deep audio-visual speech recognition

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T13:10:18.321025Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:10:18.321025Z digest=sha256:ef9a2ceea0e81ad88102e74a1694d22d7ecbb4dce70383a531c2b0a87b1314f1

Observation e44250eb-0e52-4ba4-9828-3f35e7705e3a · outbound

This paper cites Self-supervised learning of audio- visual objects from video.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Self-supervised learning of audio- visual objects from video

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T13:10:18.394341Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:10:18.394341Z digest=sha256:19a2e7f3974c3593c04d0fce722787a57ca158bfdaba163f9b7abd8a5bdc3327

Observation a26a3b85-3e13-4604-9d4e-300a40fe3a18 · outbound

This paper cites A morphable model for the synthesis of 3d faces.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation A morphable model for the synthesis of 3d faces

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T13:10:18.435036Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:10:18.435036Z digest=sha256:87b2c7b6bf3bf261760247e93f30e1e613b13a263aa240c9b256e458ce504d74

Observation 47cec664-b9c1-47b1-8ea0-2516f105dff5 · outbound

This paper cites Large scale 3d mor- phable models.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Large scale 3d mor- phable models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T13:10:18.457608Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:10:18.457608Z digest=sha256:582fbc5c5090e9a08a7b7368c373869809a01beed0330da835d1a338034e638f

Observation aa7ec033-a8cd-4eb7-9a73-a9b8dc04ff0c · outbound

This paper cites V oice puppetry.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation V oice puppetry

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T13:10:18.673658Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:10:18.673658Z digest=sha256:a22898811b0a0ae33659337952ca879f521382760b6310cf16a504f3de9d5547

Observation 1247365f-f787-4d8f-ad42-c20fb3f75153 · outbound

This paper cites How far are we from solving the 2d & 3d face alignment problem? (and a dataset of 230,000 3d facial landmarks).

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation How far are we from solving the 2d & 3d face alignment problem? (and a dataset of 230,000 3d facial landmarks)

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T13:10:18.825656Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:10:18.825656Z digest=sha256:50aa8a0c4d3bae77219d30bf7ebaba04f5db159f3a044a202035ac8bf5be5f94

Observation 48638b90-8a57-4852-ad8b-26736454c962 · outbound

This paper cites JEAN: Joint Expression and Audio-guided NeRF-based Talking Face Generation.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation JEAN: Joint Expression and Audio-guided NeRF-based Talking Face Generation

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-08-06T13:10:19.966755Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T13:10:18.912385Z digest=sha256:041acb1d121f0fc394097d779a2814783ac0fe8828d74d9c44cb088de57980f1

Observation a2f7b28a-1a6f-4e2e-a351-f81ad5ff73ed · outbound

This paper cites TalkinNeRF: Animatable Neural Fields for Full-Body Talking Humans.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation TalkinNeRF: Animatable Neural Fields for Full-Body Talking Humans

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-08-06T13:10:19.954691Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T13:10:18.981498Z digest=sha256:09a77d7f34f6cfc7edbc32cae926bb99958b65f86183ac9ea66ead11f29e61a0

Observation 2295c8f3-91b3-4243-955b-983ef24f91a1 · outbound

This paper cites Implicit neural head synthesis via controllable lo- cal deformation fields.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Implicit neural head synthesis via controllable lo- cal deformation fields

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T13:10:19.010196Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:10:19.010196Z digest=sha256:f740ba4b8de34126422f279704125e8f657cc22a4a66c86b094bf8a88ff9a285

Observation 73213bc5-db62-49a3-96f6-7311a9ae1884 · outbound

This paper cites Audio-Visual Synchronisation in the wild.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Audio-Visual Synchronisation in the wild

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T13:10:19.013378Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:10:19.013378Z digest=sha256:cc7da65c0cc8ead1f30daed4379db8b39847ad5ff8e88d75ed4ed821e8a978b8

Observation 3a41c1df-a95c-4af2-b709-2002971afbc2 · outbound

This paper cites Hierarchical cross-modal talking face generation with dynamic pixel-wise loss.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Hierarchical cross-modal talking face generation with dynamic pixel-wise loss

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T13:10:19.016638Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:10:19.016638Z digest=sha256:7f502b4e52b14ed30964929020fde0b56b71c13bd0fd179e515119483e2844a4

Observation 64367a9b-4d71-4470-8663-8f071082d597 · outbound

This paper cites Videoretalking: Audio-based lip synchronization for talking head video editing in the wild.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Videoretalking: Audio-based lip synchronization for talking head video editing in the wild

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T13:10:19.039371Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:10:19.039371Z digest=sha256:2cbf9b73cd33cfc1fbf2790f6ca00a8a7fe50b038691380bc1b5a79b1b40968b

Observation bff3e6aa-a911-44ba-a68b-1396f64f1d27 · outbound

This paper cites GPAvatar: Generalizable and Precise Head Avatar from Image(s).

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation GPAvatar: Generalizable and Precise Head Avatar from Image(s)

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T13:10:19.117914Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:10:19.117914Z digest=sha256:ab8c029a44dc8a6b4c060050e6c178d208a5997867687c327ef38363eb636256

Observation ee007793-6166-4ab5-9bfc-17b23c410ea5 · outbound

This paper cites Lip reading in the wild.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Lip reading in the wild

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T13:10:19.197057Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:10:19.197057Z digest=sha256:d3071c8378f71b79eabfef5ad19cea63ee555b93ca10f9f5eb25be47535d176d

Observation de5759f6-daf1-4d47-ba93-af9702d9fe3b · outbound

This paper cites Out of time: auto- mated lip sync in the wild.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Out of time: auto- mated lip sync in the wild

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T13:10:19.251086Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:10:19.251086Z digest=sha256:efe63bc37b8850c41287f016d73b3458a87dbdfb3f6d1b8b37068dbe4b1ae7cf

Observation d7077656-b694-4b0e-a0b6-485bdf88205b · outbound

This paper cites Perfect match: Improved cross-modal embeddings for audio-visual synchronisation.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Perfect match: Improved cross-modal embeddings for audio-visual synchronisation

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T13:10:19.326692Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:10:19.326692Z digest=sha256:d978b835e3d1e1aba1d780401e29fc4ddf6b34ebcc0303d30477c04efbb8a6a0

Observation 7e629fbd-0271-4e8c-b7ea-2228696613d8 · outbound

This paper cites Speech-driven facial animation us- ing cascaded gans for learning of motion and texture.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Speech-driven facial animation us- ing cascaded gans for learning of motion and texture

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T13:10:19.385049Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:10:19.385049Z digest=sha256:8da205152bcbe8abdc3a557c95880cd7cb4bfcdfe38d6886939b7c249c6c666d

Observation c63e7167-c726-4d70-b92c-2039a471566e · outbound

This paper cites Imagenet: A large-scale hierarchical im- age database.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Imagenet: A large-scale hierarchical im- age database

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T13:10:19.459990Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:10:19.459990Z digest=sha256:ec5911fb19af07f6ab6d2ff1d6426a8b29be01198c493981d5e3e6b317b0dc51

Observation 947f87c2-5421-482a-8e89-119d03e4e5b5 · outbound

This paper cites Arcface: Additive angular margin loss for deep face recognition.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Arcface: Additive angular margin loss for deep face recognition

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T13:10:19.502108Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:10:19.502108Z digest=sha256:7cf8e7ca599e7b7e49360ff83907983cb4e34bf623ba5bc19e73ad7e0a54e744

Observation 8c0b5163-0e1f-4848-9af1-704a4430b9f4 · outbound

This paper cites End-to-end generation of talking faces from noisy speech.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation End-to-end generation of talking faces from noisy speech

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T13:10:19.508637Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:10:19.508637Z digest=sha256:c20277d9fad939f10ccd95ac7c863849dde583fd7a241744adebb8cd71a07c31

Observation b4dce184-be8c-4a22-89d1-ff2bb1490e0d · outbound

This paper cites Efficient emotional adaptation for audio- driven talking-head generation.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Efficient emotional adaptation for audio- driven talking-head generation

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T13:10:19.511638Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:10:19.511638Z digest=sha256:27db0e71995943eaa6135e08c0e4d59021a1875aa1faafb1626614bd2e801558

Observation 1b9b07f6-4d3c-4ce0-859b-12423c898c07 · outbound

This paper cites Generative adversarial nets.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Generative adversarial nets

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T13:10:19.514379Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:10:19.514379Z digest=sha256:bb35f9a564687092aa796a488b861b415d23622dc6b3e91e0d6e492c5af89309

Observation 9595d764-85bd-42be-9a1e-da5ec0d9a827 · outbound

This paper cites Stylesync: High-fidelity generalized and personalized lip sync in style-based gener- ator.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Stylesync: High-fidelity generalized and personalized lip sync in style-based gener- ator

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T13:10:19.517398Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:10:19.517398Z digest=sha256:fe35b29bc2c6702fce08d728bab3229ab1fe795e4567ee4ddef24fe8cd9415c5

Observation 7908e6e8-4050-4c77-9977-7b8a4c4f7175 · outbound

This paper cites Ad-nerf: Audio driven neural radiance fields for talking head synthesis.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Ad-nerf: Audio driven neural radiance fields for talking head synthesis

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T13:10:19.520614Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:10:19.520614Z digest=sha256:0400dc414aedc627a075889496e8b4fbaf2c72e93e447f7e5768c9a287ee9b20

Observation 93ac78b8-73f8-4c80-b831-ffbf666981d6 · outbound

This paper cites Audio vision: Using 9 audio-visual synchrony to locate sounds.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Audio vision: Using 9 audio-visual synchrony to locate sounds

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T13:10:19.523372Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:10:19.523372Z digest=sha256:98f57d87ac37baa640f3f4a2f6038b99386b22b4ef730440a2c06122cfe03eac

Observation 7b289745-c745-4903-9d2d-1fbc51f4e216 · outbound

This paper cites Gans trained by a two time-scale update rule converge to a local nash equilib- rium.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Gans trained by a two time-scale update rule converge to a local nash equilib- rium

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T13:10:19.525963Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:10:19.525963Z digest=sha256:ba614ddfb08e1039e179a0ba320c01c2437bb3b072d284b0f040cff70bdf2835

Observation d8848ea0-459e-4dc8-b06b-cf5cf66b28ee · outbound

This paper cites Diffted: One-shot audio-driven ted talk video generation with diffusion-based co-speech gestures.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Diffted: One-shot audio-driven ted talk video generation with diffusion-based co-speech gestures

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T13:10:19.528609Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:10:19.528609Z digest=sha256:edc6680d9ecafdba39ae08d15dc98129bc18fe4279022dbdb072746b5109d33d

Observation c59b91eb-fd0a-40c5-abaf-ef8fbfac12ae · outbound

This paper cites Arbitrary style transfer in real-time with adaptive instance normalization.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Arbitrary style transfer in real-time with adaptive instance normalization

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T13:10:19.531664Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:10:19.531664Z digest=sha256:d8d91d14ffe961b26c2d4d10d13cf8a1b5269cd559be8f3f47231adb4c51b1ec

Observation 04ca43ee-d2a3-4adf-b2d8-13e0320596ca · outbound

This paper cites Discohead: audio-and- video-driven talking head generation by disentangled con- trol of head pose and facial expressions.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Discohead: audio-and- video-driven talking head generation by disentangled con- trol of head pose and facial expressions

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T13:10:19.534433Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:10:19.534433Z digest=sha256:3b4e5687f0a222ac24f0f2d28a6fc5d1aa040bec982531cf656da68c3adb2fb4

Observation 96076c35-38b5-4dcd-b088-ec50f6f83c1c · outbound

This paper cites Sparse in Space and Time: Audio-visual Synchronisation with Trainable Selectors.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Sparse in Space and Time: Audio-visual Synchronisation with Trainable Selectors

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-08-06T13:10:19.927064Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T13:10:19.536817Z digest=sha256:63e589318bd37ece5ac5590134eea3d01c0c7ff367ac7b82235e099b32309e6b

Observation 2ac0bac6-07cc-42c9-b2b6-78c4779ac027 · outbound

This paper cites Batch normalization: Accelerating deep network training by reducing internal co- variate shift.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Batch normalization: Accelerating deep network training by reducing internal co- variate shift

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T13:10:19.539510Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:10:19.539510Z digest=sha256:9e423fae932bf44d01606b4abee28aafa6541c09621dcedeea77697036525354

Observation bf1738ef-787b-4d2b-9561-0d3ae75af01b · outbound

This paper cites You said that?: Synthesising talking faces from audio.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation You said that?: Synthesising talking faces from audio

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T13:10:19.542094Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:10:19.542094Z digest=sha256:b45c1c058224800c2ff91fdffd5b0c20b321a0c397660a1a217b4a414aee0d44

Observation f0bb7da1-c2ff-4dd5-afa3-d1d15ebcb4ea · outbound

This paper cites Audio-driven emotional video portraits.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Audio-driven emotional video portraits

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T13:10:19.544758Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:10:19.544758Z digest=sha256:651ef8801eb76ba53c69a0b930a89d4ab20c5d17206bfd7e0691023be5eb3c67

Observation 54b7fb74-5b2d-4f1e-bdc7-1389719fc14a · outbound

This paper cites Eamm: One-shot emotional talking face via audio-based emotion-aware motion model.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Eamm: One-shot emotional talking face via audio-based emotion-aware motion model

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T13:10:19.547228Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:10:19.547228Z digest=sha256:692f1f73caaaa45490481601e020620fc65c61a573f14811eb8752be6db9845e

Observation 50611682-95c7-4edc-a162-fca7cf7d985a · outbound

This paper cites Audio-driven facial an- imation with deep learning: A survey.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Audio-driven facial an- imation with deep learning: A survey

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T13:10:19.549789Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:10:19.549789Z digest=sha256:0df1963c3166025ebe7ec35e20875c7445ffa33dc5d8440201f0e49bd3c463df

Observation 74a4bae3-77ca-42db-85e4-5f3bc5a3a944 · outbound

This paper cites Percep- tual losses for real-time style transfer and super-resolution.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Percep- tual losses for real-time style transfer and super-resolution

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T13:10:19.552282Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:10:19.552282Z digest=sha256:6f85a5e1e3cad9e935558e332382848ab3b6c53da2c262559560a3e8678193d6

Observation 603d623b-998b-49fd-91cd-1f36294bb2f4 · outbound

This paper cites VocaLiST: An Audio-Visual Synchronisation Model for Lips and Voices.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation VocaLiST: An Audio-Visual Synchronisation Model for Lips and Voices

Reference 37

Resolution
verified exact
local_arxiv, observed 2026-08-06T13:10:19.916020Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T13:10:19.554988Z digest=sha256:8e7f8bbf02993743f04226fb15daf2a35f73ccd5dcd29d8cf2082da9e81aa674

Observation 63d3afe6-400e-415d-9453-343bf0cd2f9c · outbound

This paper cites NeRFFaceSpeech: One-shot Audio-driven 3D Talking Head Synthesis via Generative Prior.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation NeRFFaceSpeech: One-shot Audio-driven 3D Talking Head Synthesis via Generative Prior

Reference 38

Resolution
verified exact
local_arxiv, observed 2026-08-06T13:10:19.904549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T13:10:19.557924Z digest=sha256:364ac782c3e62e106a486562fd1f49c051380bdf0d9f979120a71f020ca5409b

Observation 7d5ff702-a208-41f0-982e-a75fda6385cf · outbound

This paper cites End-to-end lip synchronisation based on pattern classification.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation End-to-end lip synchronisation based on pattern classification

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T13:10:19.561409Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:10:19.561409Z digest=sha256:32a19c41510c240180316d8f20ece27f1cdb1b70dfcd64f11cbb13d41522a64b

Observation 4fda48fa-ff56-435a-8872-805a678b6d6b · outbound

This paper cites Towards automatic face-to-face translation.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Towards automatic face-to-face translation

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:10:20.461263Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T13:10:19.564369Z digest=sha256:f6961c492608ee2d970a7b111776d35a6772934b7a3d91352ca8a861b0bda551

Observation 13883951-8b0d-44a1-9327-244bf9e0391e · outbound

This paper cites Imagenet classification with deep convolutional neural net- works.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Imagenet classification with deep convolutional neural net- works

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:10:20.453293Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T13:10:19.566904Z digest=sha256:d1406561241f5bafc7555f08fd8f2842373b5cb6f04ad5bb8d62fbbe6dc69cec

Observation 1d741386-e3a3-445c-960f-b667e71f8da9 · outbound

This paper cites Layer normalization.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Layer normalization

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:10:20.445443Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T13:10:19.569647Z digest=sha256:24b6a2c72f99ac73025cac95b5db3da3b953644c18b65857e9e9839d50230ba9

Observation 707b5b6f-89ec-46c4-8d69-dea8ab395aed · outbound

This paper cites Efficient region-aware neural radiance fields for high- fidelity talking portrait synthesis.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Efficient region-aware neural radiance fields for high- fidelity talking portrait synthesis

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:10:20.437092Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T13:10:19.572738Z digest=sha256:68ed9f66550a4cc860434f1674f757a6f0e83f2455a89e24abc331b44e2ca9db

Observation cadf50fa-9c70-4cc5-a319-565e6e90dc60 · outbound

This paper cites One-shot high- fidelity talking-head synthesis with deformable neural ra- diance field.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation One-shot high- fidelity talking-head synthesis with deformable neural ra- diance field

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:10:20.429530Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T13:10:19.575782Z digest=sha256:12a53c74af9666d95eb6b7f9506acf098eae9a50887569fa67a223f1f11cdeae

Observation b0690c38-67e4-4175-99a8-2eb20b746046 · outbound

This paper cites Expressive talking head generation with granular audio-visual control.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Expressive talking head generation with granular audio-visual control

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:10:20.420942Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T13:10:19.578661Z digest=sha256:f04b79a96f9203b50f7fef7ba16411f664bbd7e3e651e342b55d7ed67aae8bd8

Observation 54f85be3-39e3-498a-930b-e85474827d81 · outbound

This paper cites Font: Flow-guided one-shot talking head generation with natural head motions.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Font: Flow-guided one-shot talking head generation with natural head motions

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:10:20.412736Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T13:10:19.581524Z digest=sha256:b93e98956777852065300115f12d5b9190bd2cfe9c21e563ed1845f72e43eeb2

Observation 9320ef49-2fa4-4e36-a7b6-dfa1cd2b1997 · outbound

This paper cites Opt: One-shot pose- controllable talking head generation.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Opt: One-shot pose- controllable talking head generation

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:10:20.404441Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T13:10:19.585066Z digest=sha256:b5e4a802ebe015eafedb51331a49bdaba4fdf5d25489f3d6f14fb574ac3689ed

Observation 34290a6a-be79-4b19-ac6f-e795ca15481a · outbound

This paper cites Semantic-aware implicit neural audio- driven video portrait generation.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Semantic-aware implicit neural audio- driven video portrait generation

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:10:20.396341Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T13:10:19.587601Z digest=sha256:9cf9805f7f7c12ebd789b97ffc411ed2cb37ec840441428a46fd1ea5d89f9f27

Observation 926e99f2-d6c7-4f41-81ca-136480043d3d · outbound

This paper cites Moda: Mapping-once audio-driven portrait animation with dual attentions.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Moda: Mapping-once audio-driven portrait animation with dual attentions

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:10:20.389007Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T13:10:19.590321Z digest=sha256:8b72dd9a1aa26d20ed6e1ec0905ad6134435ec9221d5618da9f4217a63b8dea2

Observation 94b78c77-e869-430a-a66f-388a4e2f3f1d · outbound

This paper cites MediaPipe: A Framework for Building Perception Pipelines.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation MediaPipe: A Framework for Building Perception Pipelines

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T13:10:19.592928Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:10:19.592928Z digest=sha256:adc029d969b6492384dff7713e196f4578661a84be6df02b1363ae28f8b22877

Observation 884beea5-1ed5-4ba5-8f21-4e6dd446f871 · outbound

This paper cites Cvthead: One-shot controllable head avatar with vertex-feature transformer.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Cvthead: One-shot controllable head avatar with vertex-feature transformer

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:10:20.380403Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T13:10:19.596067Z digest=sha256:cfc1e3b51fbf87e3bcd9d7c29abd03701a5fcbc0ea4f3a3b612dfacda0bd39bf

Observation ba0a46d6-2ca8-4a44-85cf-0847073d3a51 · outbound

This paper cites Styletalk: One-shot talking head generation with controllable speak- ing styles.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Styletalk: One-shot talking head generation with controllable speak- ing styles

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:10:20.372221Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T13:10:19.598971Z digest=sha256:59db75492c8b13aae1b5d1976bd3c3957d80ab60355327ce57cfe848eea995e1

Observation 288b8791-b9ae-48dd-b171-d59d1f9503f7 · outbound

This paper cites DreamTalk: When Emotional Talking Head Generation Meets Diffusion Probabilistic Models.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation DreamTalk: When Emotional Talking Head Generation Meets Diffusion Probabilistic Models

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-06T13:10:19.601406Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:10:19.601406Z digest=sha256:82ae1da6617ecc1b92e8fdce77203a36e246949ea716c63b1ad0abf59b21bcf4

Observation 523113b3-b116-4e6e-9da3-443bf7f20502 · outbound

This paper cites Otavatar: One-shot talking face avatar with control- lable tri-plane rendering.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Otavatar: One-shot talking face avatar with control- lable tri-plane rendering

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:10:20.363238Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T13:10:19.604026Z digest=sha256:e1cc26c313e47cb73d5a1af730f49c2e2c5dc97e68d659b83a779461c8a138ce

Observation f4e159a3-aadf-4dd3-8212-a8334a1fd1c0 · outbound

This paper cites Sidgan: High-resolution dubbed video generation via shift-invariant learning.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Sidgan: High-resolution dubbed video generation via shift-invariant learning

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:10:20.355365Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T13:10:19.607458Z digest=sha256:eec9206364efc2481951e3ca5e47f4b645fbd44a7970061e96b7ee4a65ad19e7

Observation 97f21d94-2da1-4794-be23-23bb058d756f · outbound

This paper cites Diff2lip: Audio conditioned dif- fusion models for lip-synchronization.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Diff2lip: Audio conditioned dif- fusion models for lip-synchronization

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:10:20.347539Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T13:10:19.609925Z digest=sha256:1f0e562e06a7138f6707eda76ff2d46b18264e3f3581a49b689c250aac06ee26

Observation e2764bf5-eb83-40be-88ca-e845cc200340 · outbound

This paper cites Rectified linear units improve restricted boltzmann machines.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Rectified linear units improve restricted boltzmann machines

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:10:20.339828Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T13:10:19.612889Z digest=sha256:f566d4bedac5f7e8010de7d44b47ee19a2b4f7358f18f1a252a5a59b5a232a41

Observation dd662cb7-fed5-4e7c-a4bf-d28150096b28 · outbound

This paper cites Audio-visual scene analysis with self-supervised multisensory features.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Audio-visual scene analysis with self-supervised multisensory features

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:10:20.332060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T13:10:19.615652Z digest=sha256:e268132b79c9fcbea727bd19b2374535f38010446809fc847f54ed6301a2548c

Observation ef00998f-538b-4937-add0-0ec3cf2e446e · outbound

This paper cites in-the-wild.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation in-the-wild

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:10:20.324159Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T13:10:19.619193Z digest=sha256:e97d82a5b1b5770ef2369dc63f88338848d070ba3c337158e737366995f78a93

Observation 2f85de8b-b5c0-4c33-9255-2a73ed77ca1d · outbound

This paper cites Synctalkface: Talking face generation with precise lip-syncing via audio-lip memory.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Synctalkface: Talking face generation with precise lip-syncing via audio-lip memory

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:10:20.315937Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T13:10:19.621947Z digest=sha256:25b47296bddf4347c24d9c80418614d30d105837947b9054507b5c273785fc06

Observation fd35efd7-792e-44ca-9eba-57fd936d6b3b · outbound

This paper cites Semantic image synthesis with spatially-adaptive normalization.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Semantic image synthesis with spatially-adaptive normalization

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:10:20.308577Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T13:10:19.624656Z digest=sha256:97de3a94b4747f3a09c4671f0885bf6cdf05bf3cd97e0cccd0c88c5b40fd5edd

Observation ade94ba0-b672-461c-b057-199276a2257f · outbound

This paper cites Synctalk: The devil is in the synchronization for talking head synthesis.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Synctalk: The devil is in the synchronization for talking head synthesis

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:10:20.300514Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T13:10:19.627651Z digest=sha256:e17b7d0362fbc53485be91610672ada09b58482873e492c47b626a493faa49f3

Observation 9cadb96b-feec-4e4b-b72f-224f52d8bfae · outbound

This paper cites A lip sync expert is all you need for speech to lip generation in the wild.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation A lip sync expert is all you need for speech to lip generation in the wild

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:10:20.292496Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T13:10:19.630189Z digest=sha256:87fc5dd22578f8f8d4ffb7a92bf0edfbdcb050f7bcdc3d46c214b9caf95f579e

Observation 2da4e927-122b-47b4-b155-b6c22f891c9b · outbound

This paper cites U- net: Convolutional networks for biomedical image seg- mentation.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation U- net: Convolutional networks for biomedical image seg- mentation

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:10:20.284326Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T13:10:19.632941Z digest=sha256:0b8b56e5ed0b4043184363a1cf3cdb67c728a2c93660675b5ebaa0f9134dd738

Observation ad693093-af0d-40a4-bfb5-f85717dbeadc · outbound

This paper cites Learning dynamic facial radiance fields for few-shot talking head synthesis.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Learning dynamic facial radiance fields for few-shot talking head synthesis

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:10:20.275948Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T13:10:19.635850Z digest=sha256:b8818f92f6686f83753381f722da71eade63299522cb79c47e8533f8f42233d9

Observation a28a1ad9-22f4-44ef-8aae-b48360df52ce · outbound

This paper cites Difftalk: Crafting dif- fusion models for generalized audio-driven portraits anima- tion.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Difftalk: Crafting dif- fusion models for generalized audio-driven portraits anima- tion

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:10:20.268356Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T13:10:19.639035Z digest=sha256:1775aa5516932b41def2f8b7a1fba0b01cff4a17730227724c858464eccfea51

Observation 31a4aa87-2b76-4f62-ae70-d3c3653e449b · outbound

This paper cites Real-time single image and video super-resolution using an efficient sub-pixel convolutional neural network.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Real-time single image and video super-resolution using an efficient sub-pixel convolutional neural network

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:10:20.260481Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T13:10:19.642120Z digest=sha256:a07be9e9ef5df518777ebf902a2f581336cbba5c6db0310fbae9baa6f0de4393

Observation fcc24d0e-7f31-4adc-8fb3-e90e3d0b7940 · outbound

This paper cites Very Deep Convolutional Networks for Large-Scale Image Recognition.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Very Deep Convolutional Networks for Large-Scale Image Recognition

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-06T13:10:19.645004Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:10:19.645004Z digest=sha256:d04521c5578b4512c46fb1c9a7970f6e6b2303380cdac9d6e2d0a54ccec45d47

Observation 4f7efac3-601e-40f6-be5a-1d29a0daefba · outbound

This paper cites Facesync: A linear operator for measuring synchronization of video facial im- ages and audio tracks.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Facesync: A linear operator for measuring synchronization of video facial im- ages and audio tracks

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:10:20.252056Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T13:10:19.647521Z digest=sha256:9ab5d0d93727aff6319f099b6dc2aa9324c5ae6a5b99a19a3cfb463990077e1a

Observation 56d1ce50-6a9a-4143-9da0-33072d0e0c4e · outbound

This paper cites Audio-driven dubbing for user generated contents via style-aware semi-parametric synthesis.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Audio-driven dubbing for user generated contents via style-aware semi-parametric synthesis

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:10:20.244244Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T13:10:19.650253Z digest=sha256:59ea1513efd0ee6413be573b80670540fc9911f83a33e3b0d20420b5eed4250d

Observation 13db41f6-5ae3-4c5c-986f-d5b61442860f · outbound

This paper cites Everybody’s talkin’: Let me talk as you want.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Everybody’s talkin’: Let me talk as you want

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:10:20.236266Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T13:10:19.653102Z digest=sha256:7b5b5430758e8869820294cb88888f5d736c3c662a1ad633ead389b9f1956b3d

Observation 3a95139c-8a27-4893-82b8-0682e13e607a · outbound

This paper cites Talking Face Generation by Conditional Recurrent Adversarial Network.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Talking Face Generation by Conditional Recurrent Adversarial Network

Reference 72

Resolution
verified exact
local_arxiv, observed 2026-08-06T13:10:19.869488Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T13:10:19.656203Z digest=sha256:8727dda02fd0f5bf5e49a9567119a6161394ff04d708458b51522b0d6d32a100

Observation 9c298006-4cc7-430b-9dd0-4d3cf4ff1767 · outbound

This paper cites Diffused 11 heads: Diffusion models beat gans on talking-face gen- eration.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Diffused 11 heads: Diffusion models beat gans on talking-face gen- eration

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:10:20.227469Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T13:10:19.659807Z digest=sha256:b51354d56a716ea0e244611678d96fd58bc5b13dd8ce67fd3ba36bd0618a9208

Observation fa57b284-95aa-4720-8122-48f2245aff9c · outbound

This paper cites VividTalk: One-Shot Audio-Driven Talking Head Generation Based on 3D Hybrid Prior.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation VividTalk: One-Shot Audio-Driven Talking Head Generation Based on 3D Hybrid Prior

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-06T13:10:19.662786Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:10:19.662786Z digest=sha256:44c9f88ae6bb19d7d92b4bca73eced874a6cb455a6dcd0dec4202b59b10f6683

Observation 0af2d784-07f0-4183-afb5-af0fa7269a19 · outbound

This paper cites Masked lip-sync predic- tion by audio-visual contextual exploitation in transform- ers.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Masked lip-sync predic- tion by audio-visual contextual exploitation in transform- ers

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:10:20.218728Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T13:10:19.666211Z digest=sha256:8ea48468104f5cc93209d2f0e0f3279671e505cdb5c7c258b6e562b1baa425af

Observation 71220db6-5298-4d11-b709-face2ee95da0 · outbound

This paper cites Synthesizing obama: learning lip sync from audio.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Synthesizing obama: learning lip sync from audio

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:10:20.210090Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T13:10:19.668698Z digest=sha256:6626a62bd49f38257d5570a4ddb6ec054de2d338aa9982150ad45e7f97c96462

Observation 89baec36-0a0d-465a-9ae3-b5a4fbc037db · outbound

This paper cites Rethinking the inception architecture for computer vision.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Rethinking the inception architecture for computer vision

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:10:20.200831Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T13:10:19.671548Z digest=sha256:37387d2044f0b72c32279f52567f207752364f52f7354ac525c6ad0821b5601d

Observation 43a03a3b-b351-4e10-a744-7273bc44ee51 · outbound

This paper cites Emmn: Emotional mo- tion memory network for audio-driven emotional talking face generation.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Emmn: Emotional mo- tion memory network for audio-driven emotional talking face generation

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:10:20.192518Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T13:10:19.674028Z digest=sha256:2fcb706eface3e9527ea72bad9d5cc8fdc60830fa2ec0baa9412b1ea2a0a8d24

Observation 77cbb8cc-050a-41ed-9803-552063712921 · outbound

This paper cites Real-time Neural Radiance Talking Portrait Synthesis via Audio-spatial Decomposition.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Real-time Neural Radiance Talking Portrait Synthesis via Audio-spatial Decomposition

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-06T13:10:19.676863Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:10:19.676863Z digest=sha256:02b23ad09428fb2e51ee00d266f71b03d7de0e862f81108f45dfdba1fd1c24aa

Observation ce345719-f4e4-4368-ab1e-eaafe44a5c34 · outbound

This paper cites Neural voice puppetry: Audio-driven facial reenactment.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Neural voice puppetry: Audio-driven facial reenactment

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:10:20.184100Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T13:10:19.679847Z digest=sha256:bfe50c2058e8fac68e49c668202d40551a408b9764199f58337536fba095434e

Observation ffec64f7-ca81-4640-bf4b-13c69139ce6f · outbound

This paper cites EMO: Emote Portrait Alive -- Generating Expressive Portrait Videos with Audio2Video Diffusion Model under Weak Conditions.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation EMO: Emote Portrait Alive -- Generating Expressive Portrait Videos with Audio2Video Diffusion Model under Weak Conditions

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-06T13:10:19.683024Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:10:19.683024Z digest=sha256:2fd7f303707081cb45ec3366eb6475bc2813ce7248281af921efd8532bb19621

Observation 5390878b-824a-4c0a-9588-953d1edbec11 · outbound

This paper cites Attention is all you need.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Attention is all you need

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-06T13:10:19.686259Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:10:19.686259Z digest=sha256:cd8b861257089a8206d0dd79988fd55f7ae669232e4fe287d2103948662de998

Observation 2c1cdc65-6158-44dc-8663-cc88fc4c5d81 · outbound

This paper cites Realistic speech-driven facial animation with gans.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Realistic speech-driven facial animation with gans

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:10:20.170908Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T13:10:19.688989Z digest=sha256:8dc11340fb959fe7af160e5b3cf914a82117007679784a4c53f98d635eb13b65

Observation 67bbe2e9-9fd3-4856-b015-21f45598c9ab · outbound

This paper cites Progressive disentangled representa- tion learning for fine-grained controllable talking head syn- thesis.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Progressive disentangled representa- tion learning for fine-grained controllable talking head syn- thesis

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:10:20.162335Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T13:10:19.691728Z digest=sha256:37e2a87397b49a12287400925c4ae8cefd6b0ded796b06688026fea4a4896d32

Observation e85864e7-bb3a-437a-8389-fef662dabc56 · outbound

This paper cites Seeing what you said: Talking face gen- eration guided by a lip reading expert.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Seeing what you said: Talking face gen- eration guided by a lip reading expert

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:10:20.154256Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T13:10:19.694660Z digest=sha256:8f5ee2d83e6311df088072035ec5456754e577fb0890bdf9a961540a86c77e31

Observation 65ea565c-68cf-404b-a3ef-15deae8292cb · outbound

This paper cites Lipformer: High- fidelity and generalizable talking face generation with a pre- learned facial codebook.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Lipformer: High- fidelity and generalizable talking face generation with a pre- learned facial codebook

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:10:20.146253Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T13:10:19.697320Z digest=sha256:1c32d978aeae625e12661ed62e77e4b095f0b3ea1f521b918b1c297f131546fb

Observation 8bdd2035-f607-4f0a-a3be-173ac2636cab · outbound

This paper cites Styletalk++: A unified framework for controlling the speaking styles of talking heads.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Styletalk++: A unified framework for controlling the speaking styles of talking heads

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:10:20.138768Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T13:10:19.699987Z digest=sha256:09988572faa8df77acb67c849151670c80a0ca5506e62706cdf4d15b2022bd84

Observation 0a3835a7-da50-4a24-8ade-78355e5da511 · outbound

This paper cites High-resolution image synthesis and semantic manipulation with conditional gans.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation High-resolution image synthesis and semantic manipulation with conditional gans

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:10:20.131528Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T13:10:19.702818Z digest=sha256:c681f949e707452c476243449e997d2aa247c7ec74907cff82cc1100f2ab3094

Observation 4aba5721-85e4-468a-b946-2292b91a0e6b · outbound

This paper cites Image quality assessment: from error visibility to structural similarity.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Image quality assessment: from error visibility to structural similarity

Reference 89

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:10:20.123405Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T13:10:19.706016Z digest=sha256:3da996744caf5b85c26722e899db9ce1dea75e551d0614bd02e178cd8dbad5fe

Observation 5c688292-ed0f-45b3-8365-a39e22ecab81 · outbound

This paper cites Imitating arbitrary talking style for realistic audio-driven talking face synthesis.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Imitating arbitrary talking style for realistic audio-driven talking face synthesis

Reference 90

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:10:20.115487Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T13:10:19.708385Z digest=sha256:b58930cc8fa5fc639faad798c9390891ab24909d40f2ba2e79c3f18ca1076b69

Observation e2f6bcc4-c996-4de0-894e-64a8dfd382c9 · outbound

This paper cites Ganhead: Towards generative animatable neural head avatars.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Ganhead: Towards generative animatable neural head avatars

Reference 91

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:10:20.107368Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T13:10:19.711283Z digest=sha256:ed7bfa31ba99453014f870f19d3bf163ec46b2901fee00312d9a43911c56537a

Observation 29b07864-bf89-421b-b1bb-15f1d78bcb0d · outbound

This paper cites High-fidelity generalized emotional talking face generation with multi-modal emotion space learning.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation High-fidelity generalized emotional talking face generation with multi-modal emotion space learning

Reference 92

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:10:20.099524Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T13:10:19.713845Z digest=sha256:af52c83ff8d4dccbc862edfea2412361396296fb815834e58edd98e133091fa3

Observation ad5563d8-07ce-40b8-a753-308cf54b5653 · outbound

This paper cites Audio-visual speech representation expert for enhanced talking face video generation and evaluation.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Audio-visual speech representation expert for enhanced talking face video generation and evaluation

Reference 93

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:10:20.092115Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T13:10:19.716184Z digest=sha256:ea87b78df0436e49ec5d9d35f7b71d44965b71daea044eac47c212d1525b96c0

Observation 2342329e-7e22-4ced-8090-c566a2e7bfde · outbound

This paper cites Audio-driven Talking Face Generation with Stabilized Synchronization Loss.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Audio-driven Talking Face Generation with Stabilized Synchronization Loss

Reference 94

Resolution
verified exact
local_arxiv, observed 2026-08-06T13:10:19.832891Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T13:10:19.718719Z digest=sha256:063d91c299a8f22e6c368609bff7579a1febb0e9cadf5baa9e474982156fe2a0

Observation 354d3f85-f09d-4b74-8eef-cf5b8850e70e · outbound

This paper cites DFA-NeRF: Personalized Talking Head Generation via Disentangled Face Attributes Neural Rendering.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation DFA-NeRF: Personalized Talking Head Generation via Disentangled Face Attributes Neural Rendering

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-06T13:10:19.721450Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:10:19.721450Z digest=sha256:55957690ee40b6237e1f9c44d7c04b211b31f6a792f07b4cf07cc6e2afa9c22a

Observation 47656887-e25c-4bb2-be3e-a6f50dc87109 · outbound

This paper cites GeneFace: Generalized and High-Fidelity Audio-Driven 3D Talking Face Synthesis.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation GeneFace: Generalized and High-Fidelity Audio-Driven 3D Talking Face Synthesis

Reference 96

Resolution
unresolved
no resolver link, observed 2026-08-06T13:10:19.724040Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:10:19.724040Z digest=sha256:788c88a03666ceb1771a807ab2f1ce61f2d404a70830171139def41d11cace86

Observation 06b9f2ff-c777-4e0c-b011-a9a0bec6255e · outbound

This paper cites Real3D-Portrait: One-shot Realistic 3D Talking Portrait Synthesis.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Real3D-Portrait: One-shot Realistic 3D Talking Portrait Synthesis

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-06T13:10:19.726995Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:10:19.726995Z digest=sha256:0d992a8fe3357e602f959eb1f009453346d616f61117e7359e52f65c497b2c38

Observation d27549de-85a5-4933-be4e-b151ba5d8dcd · outbound

This paper cites Quantitative association of vocal-tract and facial behavior.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Quantitative association of vocal-tract and facial behavior

Reference 98

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:10:20.083847Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T13:10:19.729858Z digest=sha256:df3a23d3b44e343fd5d16df3027f3e454deb7a280cfe764108092b5317c085ff

Observation 16907925-7056-447c-990a-e7fc25da1dee · outbound

This paper cites Styleheat: One-shot high-resolution editable talking face generation via pre-trained stylegan.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Styleheat: One-shot high-resolution editable talking face generation via pre-trained stylegan

Reference 99

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:10:20.076042Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T13:10:19.732934Z digest=sha256:56622725116720cd918ea81ee56da569cfc839247081e3bda238793b170017e9

Observation cdd47e1d-af3e-44e5-bc7b-6df6cbaf5e97 · outbound

This paper cites Multimodal image synthesis and editing: A survey and taxonomy.

Mask-Free Audio-driven Talking Face Generation for Enhanced Visual Quality and Identity Preservation Multimodal image synthesis and editing: A survey and taxonomy

Reference 100

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T13:10:20.068461Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T13:10:19.735725Z digest=sha256:49d87115c6aabfa5cc1f5daf1b1688df63cbc6743941a3f2715e2957ef638565

Pith citing papers

No inbound Pith citation observations are available.