Pith. sign in

Paper Citation Record · LEDGER

Joint Co-Speech Gesture and Expressive Talking Face Generation using Diffusion with Adapters

As of 20 August 2026, this Paper Citation Record lists 46 of 46 outbound references and 0 inbound Pith citation observations for arXiv:2412.14333.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.14333 v1

Coverage vector

measured 46 of 46 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T12:23:32.145017Z

measured 46 of 46 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

46 of 46 outbound references displayed

  • verified exact0
  • verified fuzzy36
  • unresolved10
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation de17da21-11c8-4624-9345-2081b1d77823 · outbound

This paper cites Style-controllable speech-driven gesture synthesis using normalising flows.

Joint Co-Speech Gesture and Expressive Talking Face Generation using Diffusion with Adapters Style-controllable speech-driven gesture synthesis using normalising flows

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:23:34.067094Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T12:23:31.067700Z digest=sha256:1f50b918ed05e608d9a2fa3eb7c1eedb7bc69e9952595bd1fb239ca397c6c76b

Observation a9721ad2-9eb5-4a52-aa1f-eb44464dc39c · outbound

This paper cites Rhythmic gesticulator: Rhythm-aware co-speech gesture synthesis with hierarchical neural embeddings.TOG,.

Joint Co-Speech Gesture and Expressive Talking Face Generation using Diffusion with Adapters Rhythmic gesticulator: Rhythm-aware co-speech gesture synthesis with hierarchical neural embeddings.TOG,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:23:34.055410Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T12:23:31.071913Z digest=sha256:536e0be280f6589c773b3aac1c74e1d4c331837edb3f334c9324c4af40a3c9a3

Observation ec3ddc86-7edc-4399-9931-f85138335ec9 · outbound

This paper cites Gesturediffuclip: Gesture diffusion model with clip latents.

Joint Co-Speech Gesture and Expressive Talking Face Generation using Diffusion with Adapters Gesturediffuclip: Gesture diffusion model with clip latents

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:23:34.042875Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T12:23:31.146258Z digest=sha256:40deaf24c1d7cd0bdf0eafbe89d92d2e88d992e13490d79b7b600a191b4dec34

Observation 2bac6dbe-5ece-4b5c-98fa-d48cb7f3b6e6 · outbound

This paper cites an unresolved cited work.

Joint Co-Speech Gesture and Expressive Talking Face Generation using Diffusion with Adapters Unresolved cited work

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T12:23:31.244880Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:23:31.244880Z digest=sha256:54cfdac102f4f8183700749ea32709e02e22bd69e9f255305a907ab39660e95f

Observation 557dc6e7-913f-4e05-ae09-43013793edcb · outbound

This paper cites Diffsheg: A diffusion-based approach for real-time speech-driven holistic 3d expression and ges- ture generation.

Joint Co-Speech Gesture and Expressive Talking Face Generation using Diffusion with Adapters Diffsheg: A diffusion-based approach for real-time speech-driven holistic 3d expression and ges- ture generation

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:23:34.024075Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T12:23:31.351151Z digest=sha256:9a1ffea1a982cb838a0b056bc96b8440ad03fe99403ade5c93be58321e2d5fa2

Observation c6580a61-4b30-4065-9d75-372e39b19e30 · outbound

This paper cites Hierarchical cross-modal talking face generation with dynamic pixel-wise loss.

Joint Co-Speech Gesture and Expressive Talking Face Generation using Diffusion with Adapters Hierarchical cross-modal talking face generation with dynamic pixel-wise loss

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:23:34.011852Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T12:23:31.404157Z digest=sha256:9a7e1a2984d8c7d9d696f0b318a93a066fea04c792942a7f7bd94a7b51a5560b

Observation cdf013cc-808f-4df3-a5d0-c63da60afecd · outbound

This paper cites Diffusion-based co-speech gesture genera- tion using joint text and audio representation.

Joint Co-Speech Gesture and Expressive Talking Face Generation using Diffusion with Adapters Diffusion-based co-speech gesture genera- tion using joint text and audio representation

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:23:33.962297Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T12:23:31.438146Z digest=sha256:7623bb3948cf61a9c1b212595a191369d53bc0478a7300b2d6fc089cb874c989

Observation b8210488-e2d2-44bc-80a3-217d6e708d0f · outbound

This paper cites Faceformer: Speech-driven 3d facial anima- tion with transformers.

Joint Co-Speech Gesture and Expressive Talking Face Generation using Diffusion with Adapters Faceformer: Speech-driven 3d facial anima- tion with transformers

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:23:33.786933Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T12:23:31.443053Z digest=sha256:36553896062c62381cc088306b8dc50b74563ef2f24170b888946e9dd3254829

Observation 615e3994-babb-4e20-8c07-67c13b8abc7f · outbound

This paper cites Learning individual styles of conversational gesture.

Joint Co-Speech Gesture and Expressive Talking Face Generation using Diffusion with Adapters Learning individual styles of conversational gesture

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:23:33.747078Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T12:23:31.446861Z digest=sha256:b786e9aad875c43c70ae5162eabb9c3210af2d52e6ea642dac52ce14bf946e2b

Observation 48695b40-abba-43ad-9950-ea13fdcc28b5 · outbound

This paper cites Learning speech-driven 3d conversational gestures from video.

Joint Co-Speech Gesture and Expressive Talking Face Generation using Diffusion with Adapters Learning speech-driven 3d conversational gestures from video

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:23:33.735805Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T12:23:31.451665Z digest=sha256:64c0cacf599dcb1a97ff496add0787873f8aea13453d246ea4c825e6e82f4acd

Observation 11677b1a-bee1-497e-8e91-c0593f927221 · outbound

This paper cites Deber- tav3: Improving deberta using electra-style pre-training with gradient-disentangled embedding sharing, 2021.

Joint Co-Speech Gesture and Expressive Talking Face Generation using Diffusion with Adapters Deber- tav3: Improving deberta using electra-style pre-training with gradient-disentangled embedding sharing, 2021

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:23:33.723157Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T12:23:31.455595Z digest=sha256:0aab397d23093710817066fe3b9dc5aa90746f9a8dd0287af3295c6975ce4063

Observation 7c9035d2-24d3-4caa-8233-764d5fed8cf5 · outbound

This paper cites Gans trained by a two time-scale update rule converge to a local nash equilib- rium.

Joint Co-Speech Gesture and Expressive Talking Face Generation using Diffusion with Adapters Gans trained by a two time-scale update rule converge to a local nash equilib- rium

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:23:33.708910Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T12:23:31.458972Z digest=sha256:25ac9ee42d906c2f62a1451629befffa84659fd16a22d0d69db27707dd643ddb

Observation de3028ce-9680-4d18-a986-8b3fd4d38f68 · outbound

This paper cites Denoising diffu- sion probabilistic models.

Joint Co-Speech Gesture and Expressive Talking Face Generation using Diffusion with Adapters Denoising diffu- sion probabilistic models

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:23:33.558058Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T12:23:31.463531Z digest=sha256:8c81e6fef6d93e5f9f1f11d43b3f54b00ecf24fcfe43dfa44b7766c43db0ef02

Observation 20c1385d-4ad5-44bf-9093-074fb337faf9 · outbound

This paper cites Parameter-efficient transfer learning for nlp, 2019.

Joint Co-Speech Gesture and Expressive Talking Face Generation using Diffusion with Adapters Parameter-efficient transfer learning for nlp, 2019

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:23:33.493892Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T12:23:31.467792Z digest=sha256:ea00ac19200c2ca2191c1ec77093e077c4b16b2f1867b9bf0f818ddc50373ee4

Observation 21fb5a73-da6c-4413-82df-55c24e98143c · outbound

This paper cites Hubert: Self-supervised speech representation learning by masked prediction of hidden units, 2021.

Joint Co-Speech Gesture and Expressive Talking Face Generation using Diffusion with Adapters Hubert: Self-supervised speech representation learning by masked prediction of hidden units, 2021

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:23:33.481714Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T12:23:31.472014Z digest=sha256:babe75c839f2ce3ecdd2530c2295a5fafcac83e8ae0eb8e3926e83f68fe61766

Observation 21b5c66a-9c29-4f13-ab8d-5ca186e16be5 · outbound

This paper cites LoRA: Low-rank adaptation of large language models.

Joint Co-Speech Gesture and Expressive Talking Face Generation using Diffusion with Adapters LoRA: Low-rank adaptation of large language models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T12:23:31.475850Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:23:31.475850Z digest=sha256:11afd51cfab925eb42a27ba166574545b2c6c3f6b88b3ded0b35c756b5128005

Observation d2bf0af0-ec1b-479e-9319-3e959240a69c · outbound

This paper cites Audio2gestures: Generating diverse gestures from speech audio with conditional varia- tional autoencoders.

Joint Co-Speech Gesture and Expressive Talking Face Generation using Diffusion with Adapters Audio2gestures: Generating diverse gestures from speech audio with conditional varia- tional autoencoders

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:23:33.461578Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T12:23:31.480398Z digest=sha256:65faedc0789a1fe9d7caa54e32952b795354f1f659b29faa7756390a2c15fb88

Observation 3beac673-783f-4fb4-b1fa-07eaa142fbad · outbound

This paper cites Speech2video synthesis with 3d skeleton regularization and expressive body poses.

Joint Co-Speech Gesture and Expressive Talking Face Generation using Diffusion with Adapters Speech2video synthesis with 3d skeleton regularization and expressive body poses

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:23:33.344564Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T12:23:31.485230Z digest=sha256:ed14355c7b08d8ded8b8849e3dc7491a32b41e20dc1cc662e8166540eb606092

Observation f9b12401-f0dd-4737-8d41-8203a1085381 · outbound

This paper cites Vision transformers are parameter-efficient audio- visual learners.

Joint Co-Speech Gesture and Expressive Talking Face Generation using Diffusion with Adapters Vision transformers are parameter-efficient audio- visual learners

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:23:33.229564Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T12:23:31.509484Z digest=sha256:69dd93ae6b4d0dd6a82035745840fb2722d1e4db228f176fc65ca6b8c43d1382

Observation 713894c4-4ce2-447a-a95f-c7c77a8c0f5c · outbound

This paper cites Exploring Versatile Generative Language Model Via Parameter-Efficient Transfer Learning.

Joint Co-Speech Gesture and Expressive Talking Face Generation using Diffusion with Adapters Exploring Versatile Generative Language Model Via Parameter-Efficient Transfer Learning

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T12:23:31.559189Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:23:31.559189Z digest=sha256:4aff910a251d029d2333c1f446490a03048d5836a7dca8a5735747aa7ee679fd

Observation c6c3e714-6832-4b67-b065-f4c78800753d · outbound

This paper cites an unresolved cited work.

Joint Co-Speech Gesture and Expressive Talking Face Generation using Diffusion with Adapters Unresolved cited work

Reference 21

Resolution
unresolved
raw_fallback, observed 2026-08-11T12:23:33.217489Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T12:23:31.625723Z digest=sha256:3e68b9e114399dbdbc48061a94c5b9691c3eed685336675f2f836494a8c816ef

Observation b4d58b59-2201-4373-a1b0-cfdc17ff6b80 · outbound

This paper cites Beat: A large-scale semantic and emotional multi-modal dataset for conversational gestures synthesis, 2022.

Joint Co-Speech Gesture and Expressive Talking Face Generation using Diffusion with Adapters Beat: A large-scale semantic and emotional multi-modal dataset for conversational gestures synthesis, 2022

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:23:33.201901Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T12:23:31.630788Z digest=sha256:a4e4530c110d599e884c09ea75974809b99f773cc3393b2ee8ca6bb8b9613466

Observation dc00ff0c-139b-4f82-a4e2-d8ff44e7e33a · outbound

This paper cites Audio-driven co-speech gesture video generation.

Joint Co-Speech Gesture and Expressive Talking Face Generation using Diffusion with Adapters Audio-driven co-speech gesture video generation

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:23:33.187432Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T12:23:31.635416Z digest=sha256:624c41928b5c1621187883851e6e0a58399b21723fd984953eb2d50ecd27438a

Observation 7b77b202-17d8-499d-8376-c9bdbddc2b0a · outbound

This paper cites Learning hierarchical cross-modal association for co- speech gesture generation.

Joint Co-Speech Gesture and Expressive Talking Face Generation using Diffusion with Adapters Learning hierarchical cross-modal association for co- speech gesture generation

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:23:33.172649Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T12:23:31.639709Z digest=sha256:f5c4e97a81b1db83dd12522497c72a374f14a3ff0447410bda8628df0a3ef643

Observation f53deabb-5f77-46f0-ad37-9ff8c9a2eaa2 · outbound

This paper cites Roberta: A robustly optimized bert pretraining approach, 2019.

Joint Co-Speech Gesture and Expressive Talking Face Generation using Diffusion with Adapters Roberta: A robustly optimized bert pretraining approach, 2019

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:23:33.160042Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T12:23:31.644071Z digest=sha256:bdc19f50133bac884c8292eef4fb78bacc959ae9213261e58e8fb72c990b466b

Observation 9a85b79f-4e6d-40e0-a575-d987a1b2d6d5 · outbound

This paper cites RePaint: Inpainting using Denoising Diffusion Probabilistic Models.

Joint Co-Speech Gesture and Expressive Talking Face Generation using Diffusion with Adapters RePaint: Inpainting using Denoising Diffusion Probabilistic Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T12:23:31.648637Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:23:31.648637Z digest=sha256:4b360b84a5997625447d489cac26022726d12b1831ff71c01c8f77b60c7833b6

Observation 07f3fe2b-a012-478f-af69-f311ecb3beda · outbound

This paper cites Lcm-lora: A universal stable-diffusion acceleration module, 2023.

Joint Co-Speech Gesture and Expressive Talking Face Generation using Diffusion with Adapters Lcm-lora: A universal stable-diffusion acceleration module, 2023

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:23:33.146076Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T12:23:31.652690Z digest=sha256:063b2e241b2bf49a50df25a9601487b120a59678b2f96cb64bc746d6db09ef3a

Observation 552df057-bb53-4636-ada1-7b0a805806bf · outbound

This paper cites Bodyformer: Semantics-guided 3d body gesture synthesis with transformer.

Joint Co-Speech Gesture and Expressive Talking Face Generation using Diffusion with Adapters Bodyformer: Semantics-guided 3d body gesture synthesis with transformer

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:23:33.132275Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T12:23:31.656916Z digest=sha256:d4a032a639ba956434edee2a0ac2da69687d0248fd9cdf1cea05db98df78dec7

Observation a70b11a5-f6d2-4532-9a90-6999b548749f · outbound

This paper cites Sdxl: Improving latent diffusion models for high-resolution image synthesis, 2023.

Joint Co-Speech Gesture and Expressive Talking Face Generation using Diffusion with Adapters Sdxl: Improving latent diffusion models for high-resolution image synthesis, 2023

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-11T12:23:31.661292Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:23:31.661292Z digest=sha256:05584960f47ba6f5fff58ec58e2a9816ede2db16f2e173ec9c996353dc918dae

Observation 80da1581-1b76-42e4-acd6-8771970f700e · outbound

This paper cites Speech drives templates: Co-speech gesture synthesis with learned templates.

Joint Co-Speech Gesture and Expressive Talking Face Generation using Diffusion with Adapters Speech drives templates: Co-speech gesture synthesis with learned templates

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:23:32.945360Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T12:23:31.665586Z digest=sha256:ec15cc2812dc8aa106297a35b678ef9440e831537bc1efd0d7502b3a487862ab

Observation 03897cd4-7687-4f4b-b9a7-92e269978dba · outbound

This paper cites Hierarchical text-conditional image gener- ation with clip latents, 2022.

Joint Co-Speech Gesture and Expressive Talking Face Generation using Diffusion with Adapters Hierarchical text-conditional image gener- ation with clip latents, 2022

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:23:32.908236Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T12:23:31.669295Z digest=sha256:9b51124ab867e164070b8db002edfedc18f97adc33290c2176a30b40d4902c86

Observation 8d306477-907f-4374-9411-b591ef803086 · outbound

This paper cites Bilen, and A.

Joint Co-Speech Gesture and Expressive Talking Face Generation using Diffusion with Adapters Bilen, and A

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:23:32.896346Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T12:23:31.674290Z digest=sha256:6daf7f148cad39f963c18633b8ece8440c262b93ea34c316784727ffd259a4a6

Observation c89492ae-6577-45de-baf7-dc664f340397 · outbound

This paper cites Difftalk: Crafting diffusion models for generalized audio-driven portraits animation.

Joint Co-Speech Gesture and Expressive Talking Face Generation using Diffusion with Adapters Difftalk: Crafting diffusion models for generalized audio-driven portraits animation

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-11T12:23:31.735738Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:23:31.735738Z digest=sha256:4457bd9e015d46c55ce17fed44f0216519f12a66bd57951833e7cf99209dcfdf

Observation 5f0b0f00-ad23-4374-a57f-fe707b79a10c · outbound

This paper cites Co-speech gesture synthesis by reinforcement learning with contrastive pre- trained rewards.

Joint Co-Speech Gesture and Expressive Talking Face Generation using Diffusion with Adapters Co-speech gesture synthesis by reinforcement learning with contrastive pre- trained rewards

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:23:32.876344Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T12:23:31.860240Z digest=sha256:e6e8b166791efcee95c912ce86ceee78abf2c02a1708ce67c7bdb4416608ca3d

Observation 87c1cf54-64d6-473f-bf6c-16a79e484b47 · outbound

This paper cites Human motion diffu- sion model.

Joint Co-Speech Gesture and Expressive Talking Face Generation using Diffusion with Adapters Human motion diffu- sion model

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:23:32.865114Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T12:23:31.997537Z digest=sha256:d3b3b3c246e5fb051ded11c59b7b2d1169c8beb9c6ff8827e48c63ae87d0bca4

Observation 71eb2280-5116-44be-a9c9-f36253d81bf9 · outbound

This paper cites Imitator: Personalized speech-driven 3d facial animation.

Joint Co-Speech Gesture and Expressive Talking Face Generation using Diffusion with Adapters Imitator: Personalized speech-driven 3d facial animation

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:23:32.853046Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T12:23:32.055724Z digest=sha256:846fe4c929911a2f05280f4d401f30f8dd7f8a696195560abba8d49bf0757abb

Observation 863fe772-c987-4f60-94c2-2a0dd2a8a784 · outbound

This paper cites Edge: Editable dance generation from music.

Joint Co-Speech Gesture and Expressive Talking Face Generation using Diffusion with Adapters Edge: Editable dance generation from music

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:23:32.631868Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T12:23:32.107837Z digest=sha256:9ecb0f8db7cb0cbd0884e8fb6a6a4854c37d970330100f21ab258ed76101b438

Observation 60783e95-6633-4e84-90b1-dd7ec2560c3f · outbound

This paper cites FVD: A new metric for video generation, 2019.

Joint Co-Speech Gesture and Expressive Talking Face Generation using Diffusion with Adapters FVD: A new metric for video generation, 2019

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:23:32.568816Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T12:23:32.112009Z digest=sha256:c1f3aa6fe4714fac5b332d35259449d8d0f31b2e11d67ecb57a28eb3ae06fd7e

Observation a3dd28c5-08eb-40f2-8194-3abf1315680c · outbound

This paper cites Codetalker: Speech-driven 3d facial animation with discrete motion prior.

Joint Co-Speech Gesture and Expressive Talking Face Generation using Diffusion with Adapters Codetalker: Speech-driven 3d facial animation with discrete motion prior

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-11T12:23:32.116235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:23:32.116235Z digest=sha256:656b6de4a0424c1b90ebe83334258c325666836b643744b3dd312032490b8cbc

Observation b9b8f6bb-cecd-4398-a466-ecbe5407879e · outbound

This paper cites Diffus- estylegesture: Stylized audio-driven co-speech gesture gen- eration with diffusion models.

Joint Co-Speech Gesture and Expressive Talking Face Generation using Diffusion with Adapters Diffus- estylegesture: Stylized audio-driven co-speech gesture gen- eration with diffusion models

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:23:32.549848Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T12:23:32.120331Z digest=sha256:b8d5b7c3ad653024e51cd16d987450035e1a5eff85a1ae3edff73c191e4211d1

Observation cedc24dd-d7d9-43d2-8cb5-8f859c98dbd8 · outbound

This paper cites Audio-driven stylized gesture generation with flow-based model.

Joint Co-Speech Gesture and Expressive Talking Face Generation using Diffusion with Adapters Audio-driven stylized gesture generation with flow-based model

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:23:32.537909Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T12:23:32.123995Z digest=sha256:afe40a6b8963e7ece74ec518e46b96080111ebe2a5607cae54970cca2b028c15

Observation 4a377408-9a49-42e6-87b7-3fee53773318 · outbound

This paper cites Generating holistic 3d human motion from speech.

Joint Co-Speech Gesture and Expressive Talking Face Generation using Diffusion with Adapters Generating holistic 3d human motion from speech

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-11T12:23:32.128384Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:23:32.128384Z digest=sha256:ab06870d75284b7a21178796f0ed153552b9a745c40b2c52ed26d9b87fb8fc6a

Observation d4ce4ed1-0bf4-4622-87c7-c2da91e4a42d · outbound

This paper cites Speech ges- ture generation from the trimodal context of text, audio, and speaker identity.

Joint Co-Speech Gesture and Expressive Talking Face Generation using Diffusion with Adapters Speech ges- ture generation from the trimodal context of text, audio, and speaker identity

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:23:32.518223Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T12:23:32.132159Z digest=sha256:d644549ea600213c1033c789005dac0d889a07d4149bdf509c944eb28dc405c2

Observation 5fef76ce-c3ee-49b9-8c4c-49c02e99bac4 · outbound

This paper cites Robots learn social skills: End-to-end learning of co-speech gesture generation for hu- manoid robots.

Joint Co-Speech Gesture and Expressive Talking Face Generation using Diffusion with Adapters Robots learn social skills: End-to-end learning of co-speech gesture generation for hu- manoid robots

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:23:32.505102Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T12:23:32.135661Z digest=sha256:1582d81d0b21aba68305a7b0f2ce67a4247df261068937132477d06aa1ad12c6

Observation d7a1bf6b-3431-4b71-996a-e67f7e06493d · outbound

This paper cites SadTalker: Learning Realistic 3D Motion Coefficients for Stylized Audio-Driven Single Image Talking Face Animation.

Joint Co-Speech Gesture and Expressive Talking Face Generation using Diffusion with Adapters SadTalker: Learning Realistic 3D Motion Coefficients for Stylized Audio-Driven Single Image Talking Face Animation

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-11T12:23:32.140530Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:23:32.140530Z digest=sha256:b0c18efe0105e6da9a5c87409c9c4f2022a73f0d14e98cf5a367d8d7c190e6ac

Observation d7c2bcfc-c6d6-4d40-bbaf-aee0cc5155d1 · outbound

This paper cites Taming diffusion models for audio-driven co-speech gesture generation.

Joint Co-Speech Gesture and Expressive Talking Face Generation using Diffusion with Adapters Taming diffusion models for audio-driven co-speech gesture generation

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:23:32.283367Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T12:23:32.145017Z digest=sha256:1fd81dedfae9bce45d83577b88b10de03c3f9021341a57751fb598e1efffcf20

Pith citing papers

No inbound Pith citation observations are available.