Pith. sign in

Paper Citation Record · LEDGER

Joint Co-Speech Gesture and Expressive Talking Face Generation using Diffusion with Adapters

As of 20 August 2026, this Paper Citation Record lists 46 of 46 outbound references and 0 inbound Pith citation observations for arXiv:2412.14333.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.14333 v1

Coverage vector

measured 46 of 46 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T12:23:32.145017Z

measured 46 of 46 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

46 of 46 outbound references displayed

  • verified exact0
  • verified fuzzy36
  • unresolved10
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation de17da21-11c8-4624-9345-2081b1d77823 · outbound

This paper cites Style-controllable speech-driven gesture synthesis using normalising flows.

Joint Co-Speech Gesture and Expressive Talking Face Generation using Diffusion with Adapters Style-controllable speech-driven gesture synthesis using normalising flows

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:23:34.067094Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T12:23:31.067700Z digest=sha256:6dddc09ddc8bd6d2d3bf97a1fde2caa2e4b8464f2342601c9adc0a41db205b17

Observation a9721ad2-9eb5-4a52-aa1f-eb44464dc39c · outbound

This paper cites Rhythmic gesticulator: Rhythm-aware co-speech gesture synthesis with hierarchical neural embeddings.TOG,.

Joint Co-Speech Gesture and Expressive Talking Face Generation using Diffusion with Adapters Rhythmic gesticulator: Rhythm-aware co-speech gesture synthesis with hierarchical neural embeddings.TOG,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:23:34.055410Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T12:23:31.071913Z digest=sha256:a29991723fb53310a3691f238fc6f0ce968672f88b0c4622e42182a33f96e636

Observation ec3ddc86-7edc-4399-9931-f85138335ec9 · outbound

This paper cites Gesturediffuclip: Gesture diffusion model with clip latents.

Joint Co-Speech Gesture and Expressive Talking Face Generation using Diffusion with Adapters Gesturediffuclip: Gesture diffusion model with clip latents

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:23:34.042875Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T12:23:31.146258Z digest=sha256:9fc92e5f64c5ccfde89a6b866ae4d6cce2952327c202f4184387714692fa9b5c

Observation 2bac6dbe-5ece-4b5c-98fa-d48cb7f3b6e6 · outbound

This paper cites an unresolved cited work.

Joint Co-Speech Gesture and Expressive Talking Face Generation using Diffusion with Adapters Unresolved cited work

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T12:23:31.244880Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:23:31.244880Z digest=sha256:54cfdac102f4f8183700749ea32709e02e22bd69e9f255305a907ab39660e95f

Observation 557dc6e7-913f-4e05-ae09-43013793edcb · outbound

This paper cites Diffsheg: A diffusion-based approach for real-time speech-driven holistic 3d expression and ges- ture generation.

Joint Co-Speech Gesture and Expressive Talking Face Generation using Diffusion with Adapters Diffsheg: A diffusion-based approach for real-time speech-driven holistic 3d expression and ges- ture generation

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:23:34.024075Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T12:23:31.351151Z digest=sha256:f279ab0e2b2d9a89a55e955e3856c2807ca88349f81cafdae20b6b905939ae16

Observation c6580a61-4b30-4065-9d75-372e39b19e30 · outbound

This paper cites Hierarchical cross-modal talking face generation with dynamic pixel-wise loss.

Joint Co-Speech Gesture and Expressive Talking Face Generation using Diffusion with Adapters Hierarchical cross-modal talking face generation with dynamic pixel-wise loss

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:23:34.011852Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T12:23:31.404157Z digest=sha256:05391cc5997669078f6f4986e28fecbfb6a5775b650c6457a1d6416396d47839

Observation cdf013cc-808f-4df3-a5d0-c63da60afecd · outbound

This paper cites Diffusion-based co-speech gesture genera- tion using joint text and audio representation.

Joint Co-Speech Gesture and Expressive Talking Face Generation using Diffusion with Adapters Diffusion-based co-speech gesture genera- tion using joint text and audio representation

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:23:33.962297Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T12:23:31.438146Z digest=sha256:983d60f3731184d524c0bf7f09e317f502f397cd6999d4af0057fa0538340244

Observation b8210488-e2d2-44bc-80a3-217d6e708d0f · outbound

This paper cites Faceformer: Speech-driven 3d facial anima- tion with transformers.

Joint Co-Speech Gesture and Expressive Talking Face Generation using Diffusion with Adapters Faceformer: Speech-driven 3d facial anima- tion with transformers

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:23:33.786933Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T12:23:31.443053Z digest=sha256:d9e4061362236ac1ef846cc8f7da1ba15e9a58c9d1cfe33afef3d1938f36fa39

Observation 615e3994-babb-4e20-8c07-67c13b8abc7f · outbound

This paper cites Learning individual styles of conversational gesture.

Joint Co-Speech Gesture and Expressive Talking Face Generation using Diffusion with Adapters Learning individual styles of conversational gesture

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:23:33.747078Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T12:23:31.446861Z digest=sha256:fee417a09dde437537043f683dfde140265a83014736e512d19d058dc1f25acc

Observation 48695b40-abba-43ad-9950-ea13fdcc28b5 · outbound

This paper cites Learning speech-driven 3d conversational gestures from video.

Joint Co-Speech Gesture and Expressive Talking Face Generation using Diffusion with Adapters Learning speech-driven 3d conversational gestures from video

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:23:33.735805Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T12:23:31.451665Z digest=sha256:01ac55ebc69cbad72a0f6a21dd407990939dee61cebb820d32a64ab2d9aaaf45

Observation 11677b1a-bee1-497e-8e91-c0593f927221 · outbound

This paper cites Deber- tav3: Improving deberta using electra-style pre-training with gradient-disentangled embedding sharing, 2021.

Joint Co-Speech Gesture and Expressive Talking Face Generation using Diffusion with Adapters Deber- tav3: Improving deberta using electra-style pre-training with gradient-disentangled embedding sharing, 2021

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:23:33.723157Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T12:23:31.455595Z digest=sha256:f1f7d236f41066b8200492ca613dc48539af4c0647ff0d46c6c1a59fcedb1c01

Observation 7c9035d2-24d3-4caa-8233-764d5fed8cf5 · outbound

This paper cites Gans trained by a two time-scale update rule converge to a local nash equilib- rium.

Joint Co-Speech Gesture and Expressive Talking Face Generation using Diffusion with Adapters Gans trained by a two time-scale update rule converge to a local nash equilib- rium

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:23:33.708910Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T12:23:31.458972Z digest=sha256:10fb09290eafe28f9015a0f9704455e880fa186914cac4382e4cb803384727b2

Observation de3028ce-9680-4d18-a986-8b3fd4d38f68 · outbound

This paper cites Denoising diffu- sion probabilistic models.

Joint Co-Speech Gesture and Expressive Talking Face Generation using Diffusion with Adapters Denoising diffu- sion probabilistic models

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:23:33.558058Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T12:23:31.463531Z digest=sha256:e72bbd6ffc33997d3e18903afc6cde790c78cb8a2b2a33764781a5044f6d22d0

Observation 20c1385d-4ad5-44bf-9093-074fb337faf9 · outbound

This paper cites Parameter-efficient transfer learning for nlp, 2019.

Joint Co-Speech Gesture and Expressive Talking Face Generation using Diffusion with Adapters Parameter-efficient transfer learning for nlp, 2019

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:23:33.493892Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T12:23:31.467792Z digest=sha256:8176358d60b8a03b818e7f206f320d88a759b1de2af05a2db0d141dda4f74c3f

Observation 21fb5a73-da6c-4413-82df-55c24e98143c · outbound

This paper cites Hubert: Self-supervised speech representation learning by masked prediction of hidden units, 2021.

Joint Co-Speech Gesture and Expressive Talking Face Generation using Diffusion with Adapters Hubert: Self-supervised speech representation learning by masked prediction of hidden units, 2021

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:23:33.481714Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T12:23:31.472014Z digest=sha256:9d4c8710c4b98324d5887cdbaa0d8fa04d688292c470d2ba1b2263b4ead20073

Observation 21b5c66a-9c29-4f13-ab8d-5ca186e16be5 · outbound

This paper cites LoRA: Low-rank adaptation of large language models.

Joint Co-Speech Gesture and Expressive Talking Face Generation using Diffusion with Adapters LoRA: Low-rank adaptation of large language models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T12:23:31.475850Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:23:31.475850Z digest=sha256:11afd51cfab925eb42a27ba166574545b2c6c3f6b88b3ded0b35c756b5128005

Observation d2bf0af0-ec1b-479e-9319-3e959240a69c · outbound

This paper cites Audio2gestures: Generating diverse gestures from speech audio with conditional varia- tional autoencoders.

Joint Co-Speech Gesture and Expressive Talking Face Generation using Diffusion with Adapters Audio2gestures: Generating diverse gestures from speech audio with conditional varia- tional autoencoders

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:23:33.461578Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T12:23:31.480398Z digest=sha256:3bca4e22dc06169b3ab3a4286b2aac4ca4cface8b694b8a280fd307e396d245b

Observation 3beac673-783f-4fb4-b1fa-07eaa142fbad · outbound

This paper cites Speech2video synthesis with 3d skeleton regularization and expressive body poses.

Joint Co-Speech Gesture and Expressive Talking Face Generation using Diffusion with Adapters Speech2video synthesis with 3d skeleton regularization and expressive body poses

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:23:33.344564Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T12:23:31.485230Z digest=sha256:0802b9c9b40f742eeffa7f010eafa71e4112852634e1a3f74d0d34f831483104

Observation f9b12401-f0dd-4737-8d41-8203a1085381 · outbound

This paper cites Vision transformers are parameter-efficient audio- visual learners.

Joint Co-Speech Gesture and Expressive Talking Face Generation using Diffusion with Adapters Vision transformers are parameter-efficient audio- visual learners

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:23:33.229564Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T12:23:31.509484Z digest=sha256:9358fa2c8cd3a2b4f333d938f8d9ce4190e017f51453cb2884b0dc982ff179db

Observation 713894c4-4ce2-447a-a95f-c7c77a8c0f5c · outbound

This paper cites Exploring Versatile Generative Language Model Via Parameter-Efficient Transfer Learning.

Joint Co-Speech Gesture and Expressive Talking Face Generation using Diffusion with Adapters Exploring Versatile Generative Language Model Via Parameter-Efficient Transfer Learning

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T12:23:31.559189Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:23:31.559189Z digest=sha256:4aff910a251d029d2333c1f446490a03048d5836a7dca8a5735747aa7ee679fd

Observation c6c3e714-6832-4b67-b065-f4c78800753d · outbound

This paper cites an unresolved cited work.

Joint Co-Speech Gesture and Expressive Talking Face Generation using Diffusion with Adapters Unresolved cited work

Reference 21

Resolution
unresolved
raw_fallback, observed 2026-08-11T12:23:33.217489Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T12:23:31.625723Z digest=sha256:bf97fc68553e39d1970fe8d3e303e425b5e92fecb123bdd6fcb75597b9cbf385

Observation b4d58b59-2201-4373-a1b0-cfdc17ff6b80 · outbound

This paper cites Beat: A large-scale semantic and emotional multi-modal dataset for conversational gestures synthesis, 2022.

Joint Co-Speech Gesture and Expressive Talking Face Generation using Diffusion with Adapters Beat: A large-scale semantic and emotional multi-modal dataset for conversational gestures synthesis, 2022

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:23:33.201901Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T12:23:31.630788Z digest=sha256:5e9edd20f58fc637045cb4f842192a161253b56dfd79f7cdf308d40c2e1a04be

Observation dc00ff0c-139b-4f82-a4e2-d8ff44e7e33a · outbound

This paper cites Audio-driven co-speech gesture video generation.

Joint Co-Speech Gesture and Expressive Talking Face Generation using Diffusion with Adapters Audio-driven co-speech gesture video generation

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:23:33.187432Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T12:23:31.635416Z digest=sha256:79929133e231ff9ff3a5e3726f25e9888a96114a8767e501c7350e39c54543d4

Observation 7b77b202-17d8-499d-8376-c9bdbddc2b0a · outbound

This paper cites Learning hierarchical cross-modal association for co- speech gesture generation.

Joint Co-Speech Gesture and Expressive Talking Face Generation using Diffusion with Adapters Learning hierarchical cross-modal association for co- speech gesture generation

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:23:33.172649Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T12:23:31.639709Z digest=sha256:805b45b0a3008f8295cb6b3a31f6aae6e6fd45616da2115b11b5698b39833a1e

Observation f53deabb-5f77-46f0-ad37-9ff8c9a2eaa2 · outbound

This paper cites Roberta: A robustly optimized bert pretraining approach, 2019.

Joint Co-Speech Gesture and Expressive Talking Face Generation using Diffusion with Adapters Roberta: A robustly optimized bert pretraining approach, 2019

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:23:33.160042Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T12:23:31.644071Z digest=sha256:fa29652da1e71bc8ed48cc6cfa4e864bd945af8030bbbf523bc6e28e91342a89

Observation 9a85b79f-4e6d-40e0-a575-d987a1b2d6d5 · outbound

This paper cites RePaint: Inpainting using Denoising Diffusion Probabilistic Models.

Joint Co-Speech Gesture and Expressive Talking Face Generation using Diffusion with Adapters RePaint: Inpainting using Denoising Diffusion Probabilistic Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T12:23:31.648637Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:23:31.648637Z digest=sha256:4b360b84a5997625447d489cac26022726d12b1831ff71c01c8f77b60c7833b6

Observation 07f3fe2b-a012-478f-af69-f311ecb3beda · outbound

This paper cites Lcm-lora: A universal stable-diffusion acceleration module, 2023.

Joint Co-Speech Gesture and Expressive Talking Face Generation using Diffusion with Adapters Lcm-lora: A universal stable-diffusion acceleration module, 2023

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:23:33.146076Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T12:23:31.652690Z digest=sha256:788f394df4bbd0e75a13b21889d654f994572ebdf74d45cbbec0524cee09632e

Observation 552df057-bb53-4636-ada1-7b0a805806bf · outbound

This paper cites Bodyformer: Semantics-guided 3d body gesture synthesis with transformer.

Joint Co-Speech Gesture and Expressive Talking Face Generation using Diffusion with Adapters Bodyformer: Semantics-guided 3d body gesture synthesis with transformer

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:23:33.132275Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T12:23:31.656916Z digest=sha256:985892ca80b4157d93dbccbadf122f8a32a6ddcb3412e279e1c8c8a5bc8dd5bb

Observation a70b11a5-f6d2-4532-9a90-6999b548749f · outbound

This paper cites Sdxl: Improving latent diffusion models for high-resolution image synthesis, 2023.

Joint Co-Speech Gesture and Expressive Talking Face Generation using Diffusion with Adapters Sdxl: Improving latent diffusion models for high-resolution image synthesis, 2023

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-11T12:23:31.661292Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:23:31.661292Z digest=sha256:05584960f47ba6f5fff58ec58e2a9816ede2db16f2e173ec9c996353dc918dae

Observation 80da1581-1b76-42e4-acd6-8771970f700e · outbound

This paper cites Speech drives templates: Co-speech gesture synthesis with learned templates.

Joint Co-Speech Gesture and Expressive Talking Face Generation using Diffusion with Adapters Speech drives templates: Co-speech gesture synthesis with learned templates

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:23:32.945360Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T12:23:31.665586Z digest=sha256:57746b7da5f5ad998f90727c312dd4d8787e7853f58cf3bf763dbd759948d9bd

Observation 03897cd4-7687-4f4b-b9a7-92e269978dba · outbound

This paper cites Hierarchical text-conditional image gener- ation with clip latents, 2022.

Joint Co-Speech Gesture and Expressive Talking Face Generation using Diffusion with Adapters Hierarchical text-conditional image gener- ation with clip latents, 2022

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:23:32.908236Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T12:23:31.669295Z digest=sha256:2a47b4225c7dcd2ecd3da9d6c811e6706ce396d9a0686725428bb0e674562a16

Observation 8d306477-907f-4374-9411-b591ef803086 · outbound

This paper cites Bilen, and A.

Joint Co-Speech Gesture and Expressive Talking Face Generation using Diffusion with Adapters Bilen, and A

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:23:32.896346Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T12:23:31.674290Z digest=sha256:14370cf0c170ac86a0cc95222646e34117fae68a350df5dfc378bda5be6d2c5b

Observation c89492ae-6577-45de-baf7-dc664f340397 · outbound

This paper cites Difftalk: Crafting diffusion models for generalized audio-driven portraits animation.

Joint Co-Speech Gesture and Expressive Talking Face Generation using Diffusion with Adapters Difftalk: Crafting diffusion models for generalized audio-driven portraits animation

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-11T12:23:31.735738Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:23:31.735738Z digest=sha256:4457bd9e015d46c55ce17fed44f0216519f12a66bd57951833e7cf99209dcfdf

Observation 5f0b0f00-ad23-4374-a57f-fe707b79a10c · outbound

This paper cites Co-speech gesture synthesis by reinforcement learning with contrastive pre- trained rewards.

Joint Co-Speech Gesture and Expressive Talking Face Generation using Diffusion with Adapters Co-speech gesture synthesis by reinforcement learning with contrastive pre- trained rewards

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:23:32.876344Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T12:23:31.860240Z digest=sha256:6c16220355e8b42aad287da28cc1b99af2d75f0a42753cd3c1c6aeb6d404d3f9

Observation 87c1cf54-64d6-473f-bf6c-16a79e484b47 · outbound

This paper cites Human motion diffu- sion model.

Joint Co-Speech Gesture and Expressive Talking Face Generation using Diffusion with Adapters Human motion diffu- sion model

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:23:32.865114Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T12:23:31.997537Z digest=sha256:feaa01ed7308fc1304c057b15e2ecfee1ad2a7ec2b279d17601b0f4b349f4aa7

Observation 71eb2280-5116-44be-a9c9-f36253d81bf9 · outbound

This paper cites Imitator: Personalized speech-driven 3d facial animation.

Joint Co-Speech Gesture and Expressive Talking Face Generation using Diffusion with Adapters Imitator: Personalized speech-driven 3d facial animation

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:23:32.853046Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T12:23:32.055724Z digest=sha256:349ca5a0a30143df994546e3ad59c688b4c375642ecd92aafb3c30b4d60271db

Observation 863fe772-c987-4f60-94c2-2a0dd2a8a784 · outbound

This paper cites Edge: Editable dance generation from music.

Joint Co-Speech Gesture and Expressive Talking Face Generation using Diffusion with Adapters Edge: Editable dance generation from music

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:23:32.631868Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T12:23:32.107837Z digest=sha256:8f81314cf77ebc33566552d98905a10f0795e2993d60ae81f226e7008656fcdb

Observation 60783e95-6633-4e84-90b1-dd7ec2560c3f · outbound

This paper cites FVD: A new metric for video generation, 2019.

Joint Co-Speech Gesture and Expressive Talking Face Generation using Diffusion with Adapters FVD: A new metric for video generation, 2019

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:23:32.568816Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T12:23:32.112009Z digest=sha256:e8ed1ba0e4439f36475112e8cce849c5866cc879386b3d7b88ece184776eb030

Observation a3dd28c5-08eb-40f2-8194-3abf1315680c · outbound

This paper cites Codetalker: Speech-driven 3d facial animation with discrete motion prior.

Joint Co-Speech Gesture and Expressive Talking Face Generation using Diffusion with Adapters Codetalker: Speech-driven 3d facial animation with discrete motion prior

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-11T12:23:32.116235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:23:32.116235Z digest=sha256:656b6de4a0424c1b90ebe83334258c325666836b643744b3dd312032490b8cbc

Observation b9b8f6bb-cecd-4398-a466-ecbe5407879e · outbound

This paper cites Diffus- estylegesture: Stylized audio-driven co-speech gesture gen- eration with diffusion models.

Joint Co-Speech Gesture and Expressive Talking Face Generation using Diffusion with Adapters Diffus- estylegesture: Stylized audio-driven co-speech gesture gen- eration with diffusion models

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:23:32.549848Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T12:23:32.120331Z digest=sha256:e817ddd5d4fcd5d66a696ef0fd8feca6ab0486c42d414e8e30d183c48580fc2f

Observation cedc24dd-d7d9-43d2-8cb5-8f859c98dbd8 · outbound

This paper cites Audio-driven stylized gesture generation with flow-based model.

Joint Co-Speech Gesture and Expressive Talking Face Generation using Diffusion with Adapters Audio-driven stylized gesture generation with flow-based model

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:23:32.537909Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T12:23:32.123995Z digest=sha256:717884be7dfbe9f81b47a16a6e3e82acdbc02a5c559b3ac0e21317ea94a1fa60

Observation 4a377408-9a49-42e6-87b7-3fee53773318 · outbound

This paper cites Generating holistic 3d human motion from speech.

Joint Co-Speech Gesture and Expressive Talking Face Generation using Diffusion with Adapters Generating holistic 3d human motion from speech

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-11T12:23:32.128384Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:23:32.128384Z digest=sha256:ab06870d75284b7a21178796f0ed153552b9a745c40b2c52ed26d9b87fb8fc6a

Observation d4ce4ed1-0bf4-4622-87c7-c2da91e4a42d · outbound

This paper cites Speech ges- ture generation from the trimodal context of text, audio, and speaker identity.

Joint Co-Speech Gesture and Expressive Talking Face Generation using Diffusion with Adapters Speech ges- ture generation from the trimodal context of text, audio, and speaker identity

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:23:32.518223Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T12:23:32.132159Z digest=sha256:59bc80a072f7ede91e740c90fb508a184679d76ae3a1e12b6d9c4495bcf74939

Observation 5fef76ce-c3ee-49b9-8c4c-49c02e99bac4 · outbound

This paper cites Robots learn social skills: End-to-end learning of co-speech gesture generation for hu- manoid robots.

Joint Co-Speech Gesture and Expressive Talking Face Generation using Diffusion with Adapters Robots learn social skills: End-to-end learning of co-speech gesture generation for hu- manoid robots

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:23:32.505102Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T12:23:32.135661Z digest=sha256:5f6850c5cc6953631bcf7f94ddd6be51a7346392a56cecc6259a4048a7d0a12c

Observation d7a1bf6b-3431-4b71-996a-e67f7e06493d · outbound

This paper cites SadTalker: Learning Realistic 3D Motion Coefficients for Stylized Audio-Driven Single Image Talking Face Animation.

Joint Co-Speech Gesture and Expressive Talking Face Generation using Diffusion with Adapters SadTalker: Learning Realistic 3D Motion Coefficients for Stylized Audio-Driven Single Image Talking Face Animation

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-11T12:23:32.140530Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:23:32.140530Z digest=sha256:b0c18efe0105e6da9a5c87409c9c4f2022a73f0d14e98cf5a367d8d7c190e6ac

Observation d7c2bcfc-c6d6-4d40-bbaf-aee0cc5155d1 · outbound

This paper cites Taming diffusion models for audio-driven co-speech gesture generation.

Joint Co-Speech Gesture and Expressive Talking Face Generation using Diffusion with Adapters Taming diffusion models for audio-driven co-speech gesture generation

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:23:32.283367Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-11T12:23:32.145017Z digest=sha256:774d1ec85335378ea3b5a6484e108a78613042b142c34b2214c276311357b808

Pith citing papers

No inbound Pith citation observations are available.