Pith. sign in

Paper Citation Record · LEDGER

Multi-human Interactive Talking Dataset

As of 9 August 2026, this Paper Citation Record lists 60 of 60 outbound references and 0 inbound Pith citation observations for arXiv:2508.03050.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.03050 v1

Coverage vector

measured 60 of 60 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T04:46:05.382405Z

measured 60 of 60 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

60 of 60 outbound references displayed

  • verified exact3
  • verified fuzzy28
  • unresolved29
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 0cdfd198-ace5-49ca-938f-ec0ff35a64bd · outbound

This paper cites GitHub repository.

Multi-human Interactive Talking Dataset GitHub repository

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T04:46:10.599728Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T04:46:04.006639Z digest=sha256:06cfe6acc48988c189f925c38bb045f6607498d9e387960b4042d47d7f06e47f

Observation 3acb0df4-a24a-403b-9b3e-0e9e8fe8397f · outbound

This paper cites wav2vec 2.0: A framework for self-supervised learning of speech representations.

Multi-human Interactive Talking Dataset wav2vec 2.0: A framework for self-supervised learning of speech representations

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T04:46:10.437319Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T04:46:04.012092Z digest=sha256:0ab00e6d5c20522fca9c4fc5f851f54bbad0cd3a81b37893a13d7f3cac605c44

Observation 2c37f9f6-713c-4fea-96d7-9182bcb6f0a7 · outbound

This paper cites TalkNet: Fully-Convolutional Non-Autoregressive Speech Synthesis Model.

Multi-human Interactive Talking Dataset TalkNet: Fully-Convolutional Non-Autoregressive Speech Synthesis Model

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-08-06T04:46:06.309021Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T04:46:04.017758Z digest=sha256:fa5d3dbb76623848633970e74f8da6150b405333c032cdaad66fdb0deb27678a

Observation 4bd0cc6d-2ddf-452a-8569-95a465a7cd43 · outbound

This paper cites Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets.

Multi-human Interactive Talking Dataset Stable Video Diffusion: Scaling Latent Video Diffusion Models to Large Datasets

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T04:46:04.022782Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:46:04.022782Z digest=sha256:97521f4cbafcc318a6a1f7bd9dae7a4ba6b90e269bc0d6e69fa2300729c81e5d

Observation 7d291381-c3c7-4d22-a4da-7a1cf7ef60c9 · outbound

This paper cites Magicdance: Realistic human dance video generation with motions & facial expressions transfer.

Multi-human Interactive Talking Dataset Magicdance: Realistic human dance video generation with motions & facial expressions transfer

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T04:46:10.271706Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T04:46:04.027925Z digest=sha256:4b2bb019423700873d86de0882f3a5f01a693572dedcd0d22a43b65f2ccbe72a

Observation 0a5c21c2-a261-4a67-b674-442421b2a455 · outbound

This paper cites VideoCrafter1: Open Diffusion Models for High-Quality Video Generation.

Multi-human Interactive Talking Dataset VideoCrafter1: Open Diffusion Models for High-Quality Video Generation

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T04:46:04.032620Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:46:04.032620Z digest=sha256:0d0f705f49529d5c1b5a0d442b7aed68fa7e057f08b240b78f2c3a3b2d978c6e

Observation 04945ce8-166c-4c18-b9a1-af83f84ca9a4 · outbound

This paper cites EchoMimic: Lifelike Audio-Driven Portrait Animations through Editable Landmark Conditions.

Multi-human Interactive Talking Dataset EchoMimic: Lifelike Audio-Driven Portrait Animations through Editable Landmark Conditions

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T04:46:04.037670Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:46:04.037670Z digest=sha256:49fbcf5b73f7afb45b66564c0e758d14a0f1a6fd13f8eba95aca20174369dab6

Observation 4105ac0e-03fb-4bd3-b2b7-cd02a2c2c66a · outbound

This paper cites Out of time: automated lip sync in the wild.

Multi-human Interactive Talking Dataset Out of time: automated lip sync in the wild

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T04:46:04.042957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:46:04.042957Z digest=sha256:a775967f6199c77579252f795ff65119044c6d13cbbaaf18981d542edf4ea1fb

Observation 0f8c61c2-97e5-4f21-9eab-7ef66fa10556 · outbound

This paper cites VoxCeleb2: Deep Speaker Recognition.

Multi-human Interactive Talking Dataset VoxCeleb2: Deep Speaker Recognition

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T04:46:04.049199Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:46:04.049199Z digest=sha256:9fab80fd9fad8382b59cb1f892cb6ed90f398273508dcf3bf36b7ad88f45152a

Observation 6e201cba-1307-4505-960c-56ce975e0a70 · outbound

This paper cites Hallo2: Long-Duration and High-Resolution Audio-Driven Portrait Image Animation.

Multi-human Interactive Talking Dataset Hallo2: Long-Duration and High-Resolution Audio-Driven Portrait Image Animation

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T04:46:04.054306Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:46:04.054306Z digest=sha256:83d5d6d925f4d51f6d0e64feed182ae2ed34eb87ef0965013c66e5ff392f3fb0

Observation dafcda37-a4f0-4a18-8f21-1a03e6c906f9 · outbound

This paper cites DreaMoving: A Human Video Generation Framework based on Diffusion Models.

Multi-human Interactive Talking Dataset DreaMoving: A Human Video Generation Framework based on Diffusion Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T04:46:04.059505Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:46:04.059505Z digest=sha256:9b9d96643c1dcbd85b67856684bf24f5ab17da68eb1b8d9a3943fdc008fdd06f

Observation 5b10e9c2-488c-4edc-934e-7963998b7bd4 · outbound

This paper cites Affective Faces for Goal-Driven Dyadic Communication.

Multi-human Interactive Talking Dataset Affective Faces for Goal-Driven Dyadic Communication

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T04:46:04.064242Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:46:04.064242Z digest=sha256:8538af00245cc6a17f289722a12511363c6ce5eb73d2863df84580a6fdf56a1d

Observation 4a6ebae4-3608-4ccb-8724-bdf0b704c5ae · outbound

This paper cites Learning individual styles of conversational gesture.

Multi-human Interactive Talking Dataset Learning individual styles of conversational gesture

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T04:46:10.130062Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T04:46:04.069026Z digest=sha256:d4abad1ad3b0713c44bba720a70bd5d448ed73f25c89733cfe83599ce0dcfea8

Observation 525143e0-4471-4645-9138-0b81f1c55acd · outbound

This paper cites AnimateDiff: Animate Your Personalized Text-to-Image Diffusion Models without Specific Tuning.

Multi-human Interactive Talking Dataset AnimateDiff: Animate Your Personalized Text-to-Image Diffusion Models without Specific Tuning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T04:46:04.073413Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:46:04.073413Z digest=sha256:169b33965f7e75b27f11969141800d84942ee130adf0c0b3919fa9b2a71d098e

Observation 30321dbd-b666-4811-9777-d4e665d03ccd · outbound

This paper cites Co-speech gesture video generation via motion- decoupled diffusion model.

Multi-human Interactive Talking Dataset Co-speech gesture video generation via motion- decoupled diffusion model

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T04:46:10.044867Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T04:46:04.078056Z digest=sha256:98e770e146ab74e720407991e558f8a72617765c2f1a9ce668b207e0f3a89e0e

Observation dc9fa1c7-c6c9-4b7e-b603-3e1e9081c766 · outbound

This paper cites Animate anyone: Consistent and controllable image-to-video synthesis for character animation.

Multi-human Interactive Talking Dataset Animate anyone: Consistent and controllable image-to-video synthesis for character animation

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T04:46:09.900096Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T04:46:04.082151Z digest=sha256:19a2bc2b92bcb3d731b39543e155060f343616bc12aae5bf172575b1619ba499

Observation c3cbb69b-bda5-4073-830d-9de0291b2d3c · outbound

This paper cites Whisperv.

Multi-human Interactive Talking Dataset Whisperv

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T04:46:09.772447Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T04:46:04.086464Z digest=sha256:6f0e2a920be9ed97a7fe859fc45c12ea8797d12397a093f8fbb6fa76d03b2b11

Observation dd891624-2dfa-48d8-8afb-3f224a0dd716 · outbound

This paper cites Perceptual conversational head generation with regularized driver and enhanced renderer.

Multi-human Interactive Talking Dataset Perceptual conversational head generation with regularized driver and enhanced renderer

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T04:46:09.659633Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T04:46:04.090821Z digest=sha256:e9332f40a928df715c7e74ed009c3e23ed02ab248dba9113c41e079e1cf156c6

Observation 8ac39e5e-8a09-4631-bb69-4d28ba9a7697 · outbound

This paper cites Vbench: Comprehensive benchmark suite for video generative models.

Multi-human Interactive Talking Dataset Vbench: Comprehensive benchmark suite for video generative models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T04:46:04.095209Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:46:04.095209Z digest=sha256:2c447b5851eb9be50683cef628cda61a6e7bf8376cb857bbf450c6de72ac862c

Observation 12f37f79-c744-4589-aa63-959fc0534037 · outbound

This paper cites Loopy: Taming Audio-Driven Portrait Avatar with Long-Term Motion Dependency.

Multi-human Interactive Talking Dataset Loopy: Taming Audio-Driven Portrait Avatar with Long-Term Motion Dependency

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T04:46:04.100602Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:46:04.100602Z digest=sha256:1703b17e341954de6e1817dcd286a29880a93a3d14792c8dae659193e589a145

Observation 4b8685a9-ffe7-4a01-9409-9ecc6342e2cf · outbound

This paper cites Text2performer: Text-driven human video generation.

Multi-human Interactive Talking Dataset Text2performer: Text-driven human video generation

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T04:46:09.498775Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T04:46:04.106115Z digest=sha256:83ae37ed4c6b8f7132ffd0dd68e686c6c8616f2d459061bb224779b7657c5d37

Observation 6b6eb7d4-1ed5-4b79-a431-773593f37b62 · outbound

This paper cites Whole-body human pose estimation in the wild.

Multi-human Interactive Talking Dataset Whole-body human pose estimation in the wild

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T04:46:09.341158Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T04:46:04.110498Z digest=sha256:db66e56be7ba0552b318fb33d89dfa34b9b0e5e3fd0c696b32c985eec0150748

Observation 34967ca4-0aae-41e9-8366-dd2a3c3760c0 · outbound

This paper cites Sapiens: Foundation for human vision models.

Multi-human Interactive Talking Dataset Sapiens: Foundation for human vision models

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T04:46:09.208983Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T04:46:04.114657Z digest=sha256:a6352cb61779bec8f9172631c6356830c392543027e53b9a502b2e44d559b9a1

Observation 8873e075-bb5b-4ce9-8ca0-52ff8e6e0a6b · outbound

This paper cites A Comprehensive Survey on Human Video Generation: Challenges, Methods, and Insights.

Multi-human Interactive Talking Dataset A Comprehensive Survey on Human Video Generation: Challenges, Methods, and Insights

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T04:46:04.118923Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:46:04.118923Z digest=sha256:1eb47aad6112abbbf2ef43b7232e530c93c1ab3b605013ce463edb39fedeecf8

Observation 782103c6-d6b4-4152-bf72-c0d101291091 · outbound

This paper cites OpenHumanVid: A Large-Scale High-Quality Dataset for Enhancing Human-Centric Video Generation.

Multi-human Interactive Talking Dataset OpenHumanVid: A Large-Scale High-Quality Dataset for Enhancing Human-Centric Video Generation

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T04:46:04.123309Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:46:04.123309Z digest=sha256:a4f1f0cd04f6f65bec1ea508e04cf04e2b320ff5938703128295d5a7f8d30c20

Observation 9370856a-bb87-47d2-87dd-38fe730e3878 · outbound

This paper cites TANGO: Co-Speech Gesture Video Reenactment with Hierarchical Audio Motion Embedding and Diffusion Interpolation.

Multi-human Interactive Talking Dataset TANGO: Co-Speech Gesture Video Reenactment with Hierarchical Audio Motion Embedding and Diffusion Interpolation

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T04:46:04.127815Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:46:04.127815Z digest=sha256:dd58e3f509f19d2a6b317a795dfad5359836fa16493ee823ce8929e82d56532e

Observation 5fb68e29-6299-4801-804f-d9269e44cf0e · outbound

This paper cites Customlistener: Text-guided responsive interaction for user-friendly listening head generation.

Multi-human Interactive Talking Dataset Customlistener: Text-guided responsive interaction for user-friendly listening head generation

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T04:46:09.041918Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T04:46:04.135037Z digest=sha256:346102b933c0b44efa1fad4509c4e9f7be500e08b17d46ebf849ec71d8bdb4ab

Observation 89b6971c-5938-4a0a-9e98-bee54b5bd831 · outbound

This paper cites Learning hierarchical cross-modal association for co-speech gesture generation.

Multi-human Interactive Talking Dataset Learning hierarchical cross-modal association for co-speech gesture generation

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T04:46:08.912736Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T04:46:04.148338Z digest=sha256:03f7f0f5d76098e128e17d5c8a212b0a9a79866ee689e47f3d107992f33d2f5e

Observation 4848fdaa-7499-4bce-8010-c945cdee0496 · outbound

This paper cites Follow your pose: Pose-guided text-to-video generation using pose-free videos.

Multi-human Interactive Talking Dataset Follow your pose: Pose-guided text-to-video generation using pose-free videos

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T04:46:08.773653Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T04:46:04.157908Z digest=sha256:9c92f1f5b6eb3524c9681baf6789da1e30ce671d553299cf8c67f6d0ee278fdb

Observation 5610ee8e-98c2-4014-9d1b-1a3db1726d28 · outbound

This paper cites Learning to listen: Modeling non-deterministic dyadic facial motion.

Multi-human Interactive Talking Dataset Learning to listen: Modeling non-deterministic dyadic facial motion

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T04:46:04.167430Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:46:04.167430Z digest=sha256:287ea218e24e00b28508c796847b6ed646deeb2b1b015c7b3d71c8ac244a992f

Observation 67282c5a-0dc5-4cc9-8554-8913de99d050 · outbound

This paper cites Scalable diffusion models with transformers.

Multi-human Interactive Talking Dataset Scalable diffusion models with transformers

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T04:46:04.179142Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:46:04.179142Z digest=sha256:d3c1c93cc9c991d7507bd31f3f0b3d21fe8c7d0bce3bb13721c8464a96114583

Observation 882adab5-7d3a-4c8a-9520-a8f3cedb2778 · outbound

This paper cites ControlNeXt: Powerful and Efficient Control for Image and Video Generation.

Multi-human Interactive Talking Dataset ControlNeXt: Powerful and Efficient Control for Image and Video Generation

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T04:46:04.188293Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:46:04.188293Z digest=sha256:48bb57fe1b40fec350a095c52927adca0fff466fe84790515955700f6563f12a

Observation 44850ca4-9bf1-4f9c-9384-6555bf76b903 · outbound

This paper cites A lip sync expert is all you need for speech to lip generation in the wild.

Multi-human Interactive Talking Dataset A lip sync expert is all you need for speech to lip generation in the wild

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T04:46:04.197991Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:46:04.197991Z digest=sha256:8d45cce97c476bc5e55d5cc8869230a49bddaf801690055f36ee58fb986cc4f9

Observation ed8189e5-f776-49dd-993b-93e31b7833c8 · outbound

This paper cites Speech drives templates: Co- speech gesture synthesis with learned templates.

Multi-human Interactive Talking Dataset Speech drives templates: Co- speech gesture synthesis with learned templates

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T04:46:08.607230Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T04:46:04.208672Z digest=sha256:80f0456736e8c7f3be44947450e531b1cbede4354dfa96a0e2d0908a7616a3d7

Observation 80a096b3-0296-45be-a6e4-6eb1d2eae47a · outbound

This paper cites High- resolution image synthesis with latent diffusion models.

Multi-human Interactive Talking Dataset High- resolution image synthesis with latent diffusion models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T04:46:04.220246Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:46:04.220246Z digest=sha256:bb3465a79e988aa7b71eda98c0577027c5bd8409d56f33cd46c8cc45982e1d6b

Observation b0b63739-a436-4cef-9232-56b0f0fa4212 · outbound

This paper cites MediConfusion: Can you trust your AI radiologist? Probing the reliability of multimodal medical foundation models.

Multi-human Interactive Talking Dataset MediConfusion: Can you trust your AI radiologist? Probing the reliability of multimodal medical foundation models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T04:46:04.233222Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:46:04.233222Z digest=sha256:c1a67bc5d9c900d302634fdf128e050e1de231c57d33a7a750da8573f3b2a0a0

Observation c89c97e4-3d0e-475b-add1-2a0b11bc546d · outbound

This paper cites Make-A-Video: Text-to-Video Generation without Text-Video Data.

Multi-human Interactive Talking Dataset Make-A-Video: Text-to-Video Generation without Text-Video Data

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T04:46:04.239375Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:46:04.239375Z digest=sha256:806970ca5d3c8a147284eb6253ffa22c2adc70b2504362d861584691b1f14de6

Observation 52ef254d-7061-4244-ac43-24e4e974b57d · outbound

This paper cites Lip reading sentences in the wild.

Multi-human Interactive Talking Dataset Lip reading sentences in the wild

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T04:46:04.248960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:46:04.248960Z digest=sha256:0e247cbda7bb98f65cd404fb82b1de4ab57fc55ff67c64c04a68ee590d5422b7

Observation f110161b-2b44-47a6-848e-45469ae2262d · outbound

This paper cites Llama Learns to Direct: DirectorLLM for Human-Centric Video Generation.

Multi-human Interactive Talking Dataset Llama Learns to Direct: DirectorLLM for Human-Centric Video Generation

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T04:46:04.259024Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:46:04.259024Z digest=sha256:9ed7bac967abc63f5f46f003fef2f21d7356eb8f6572162ff34d98b24381767d

Observation 49d27cb2-3f02-41a5-a0a7-1e78d1932a71 · outbound

This paper cites Diffused heads: Diffusion models beat gans on talking-face generation.

Multi-human Interactive Talking Dataset Diffused heads: Diffusion models beat gans on talking-face generation

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T04:46:08.416341Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T04:46:04.267724Z digest=sha256:2872eb87210deef9ed48318d27a7dce00a58bc919dbebc4cfe2a93d21fc50098

Observation 1c5507d2-af67-4c7a-9fe8-8d898218fdb2 · outbound

This paper cites MultiTalk: Enhancing 3D Talking Head Generation Across Languages with Multilingual Video Dataset.

Multi-human Interactive Talking Dataset MultiTalk: Enhancing 3D Talking Head Generation Across Languages with Multilingual Video Dataset

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T04:46:04.277714Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:46:04.277714Z digest=sha256:c373ac1c871f9f32b703bf8b5d63058e32d7f6231fb469dccd6241532edb9d25

Observation d0c1f5e6-d51c-4daf-b9f5-2d5f6367ca36 · outbound

This paper cites Edtalk: Efficient disentanglement for emotional talking head synthesis.

Multi-human Interactive Talking Dataset Edtalk: Efficient disentanglement for emotional talking head synthesis

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T04:46:08.289330Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T04:46:04.288655Z digest=sha256:0b34c7dcf923d71f58a33c776d9bdbd03cc5bc80e314c73eefc5bfc8d514a48f

Observation a3925a25-8fac-4919-8aa6-840e946351e4 · outbound

This paper cites Dyadic Interaction Modeling for Social Behavior Generation.

Multi-human Interactive Talking Dataset Dyadic Interaction Modeling for Social Behavior Generation

Reference 43

Resolution
verified exact
local_arxiv, observed 2026-08-06T04:46:05.855275Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T04:46:04.305034Z digest=sha256:494558ae264138c6200fe4b1248b31ec19a6678699a89ed4088d26e95567da0b

Observation ec413d04-8900-4808-9b44-97b1fa0e29ae · outbound

This paper cites Realistic speech-driven facial animation with gans.

Multi-human Interactive Talking Dataset Realistic speech-driven facial animation with gans

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T04:46:08.112266Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T04:46:04.317437Z digest=sha256:c879e5a50ff681abd2cb4202fda288acd7dcb912f3195a93bf678db30f099cb7

Observation accf5f45-1e4d-4eda-b2a4-6609cff35088 · outbound

This paper cites AgentAvatar: Disentangling Planning, Driving and Rendering for Photorealistic Avatar Agents.

Multi-human Interactive Talking Dataset AgentAvatar: Disentangling Planning, Driving and Rendering for Photorealistic Avatar Agents

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T04:46:04.346797Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:46:04.346797Z digest=sha256:fce224073df92b5d67b7f9a58de25f9e82b091fdfba8f01873159a160eec2ed4

Observation 3894480e-4fe8-4be3-8cb4-4d99737df128 · outbound

This paper cites EmotiveTalk: Expressive Talking Head Generation through Audio Information Decoupling and Emotional Video Diffusion.

Multi-human Interactive Talking Dataset EmotiveTalk: Expressive Talking Head Generation through Audio Information Decoupling and Emotional Video Diffusion

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T04:46:04.405476Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:46:04.405476Z digest=sha256:9799c248daf6e663b345199eac6b12dee88307e632042c710c73f54ac4dce92f

Observation c01e8b70-a59c-4002-baac-3ec27f6e87c4 · outbound

This paper cites Mead: A large-scale audio-visual dataset for emotional talking-face generation.

Multi-human Interactive Talking Dataset Mead: A large-scale audio-visual dataset for emotional talking-face generation

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T04:46:07.975733Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T04:46:04.507674Z digest=sha256:b391012e44b5fb9003a4637c457258b9d674f4753de9dc55bee4b081dd721948

Observation c791ccfc-cd29-40f0-84ed-f9e2f1b706e4 · outbound

This paper cites Draganything: Motion control for anything using entity representation.

Multi-human Interactive Talking Dataset Draganything: Motion control for anything using entity representation

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T04:46:07.805118Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T04:46:04.599303Z digest=sha256:ce9965c2341ae47360b9c8e95fa91c184133e9b1b143665f70c1bf6e0a07cb15

Observation cde52439-12ea-4f8b-9e43-9842e2f3b7b2 · outbound

This paper cites Magicanimate: Temporally consistent human image animation using diffusion model.

Multi-human Interactive Talking Dataset Magicanimate: Temporally consistent human image animation using diffusion model

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-06T04:46:04.669820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:46:04.669820Z digest=sha256:c53a32f03b4540353df03391300fc4c4adb06dd0c4259d2adbc495adbbc818de

Observation 194abaf8-9235-4387-ae97-1d68579224ba · outbound

This paper cites Cogvideox: Text-to-video diffusion models with an expert transformer.

Multi-human Interactive Talking Dataset Cogvideox: Text-to-video diffusion models with an expert transformer

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T04:46:07.641043Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T04:46:04.734189Z digest=sha256:28cfecfa59c697b0e91857a807869b85eeffce91242c925d1871c10310c1ef69

Observation 93cc9205-cab3-4e42-acf8-19831313e13f · outbound

This paper cites Show-1: Marrying pixel and latent diffusion models for text-to-video generation.

Multi-human Interactive Talking Dataset Show-1: Marrying pixel and latent diffusion models for text-to-video generation

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T04:46:07.488950Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T04:46:04.832246Z digest=sha256:68d955647941e3956ea7da5bc24c5f3fe721587c8ee53878f520e6ed6d845ca6

Observation 8b1f1e3a-64e0-491e-8084-8de15dbb6c95 · outbound

This paper cites Adding conditional control to text-to-image diffusion models.

Multi-human Interactive Talking Dataset Adding conditional control to text-to-image diffusion models

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-06T04:46:04.910982Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:46:04.910982Z digest=sha256:314bb6e6a34b74afc785af241a07e02dc18ef1282a3b21a1197bcfb13ad91814

Observation 08bde3ec-3b0c-4783-a73b-0dd0982db4bb · outbound

This paper cites Sadtalker: Learning realistic 3d motion coefficients for stylized audio-driven single image talking face animation.

Multi-human Interactive Talking Dataset Sadtalker: Learning realistic 3d motion coefficients for stylized audio-driven single image talking face animation

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T04:46:07.337650Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T04:46:04.954087Z digest=sha256:12a2f6b17897319ee184d299620691a843f9e80c99d3acefaae3b83a58567ee2

Observation 4fb099b9-1173-4346-84d0-3e624918e8a1 · outbound

This paper cites Flow-guided one-shot talking face generation with a high-resolution audio-visual dataset.

Multi-human Interactive Talking Dataset Flow-guided one-shot talking face generation with a high-resolution audio-visual dataset

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T04:46:07.173922Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T04:46:05.010072Z digest=sha256:e85b8a4069d03cfcae85873467de9de91211480cecf038eab776d6e54e16f11f

Observation a8d6a4bf-0880-40a8-be54-6ae719648052 · outbound

This paper cites Responsive listening head generation: a benchmark dataset and baseline.

Multi-human Interactive Talking Dataset Responsive listening head generation: a benchmark dataset and baseline

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T04:46:07.012112Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T04:46:05.079666Z digest=sha256:4e7f1aa16aae3b4637e14624718d86613e179e3eaf635f9afb2bf4e35c67997e

Observation a2503984-6b32-4526-a0ac-39f27d8b765f · outbound

This paper cites Interactive Conversational Head Generation.

Multi-human Interactive Talking Dataset Interactive Conversational Head Generation

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-06T04:46:05.142153Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:46:05.142153Z digest=sha256:b8012025b031f003db1a435576fb1659874c66db432dfd4e3e34848b277c7ce1

Observation 714cd6f0-3ada-4664-ae7a-1e3308d5ebf0 · outbound

This paper cites Audio-driven neural gesture reenactment with video motion graphs.

Multi-human Interactive Talking Dataset Audio-driven neural gesture reenactment with video motion graphs

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T04:46:06.843301Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T04:46:05.203642Z digest=sha256:048181da8c664c63cec1c1f8a5e17c819394449f6a3653d3f3745ec95ba11241

Observation a22c984e-8a12-4baa-bdc1-e6c6364a5ffb · outbound

This paper cites Celebv-hq: A large-scale video facial attributes dataset.

Multi-human Interactive Talking Dataset Celebv-hq: A large-scale video facial attributes dataset

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T04:46:06.678256Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T04:46:05.271574Z digest=sha256:bed8180cb68bef1486f8640374b1618bbb1a9062e0bf2a03e4a7ea3fc4cd250c

Observation 94f3564d-f034-4b36-8adb-d63cb3e3a24d · outbound

This paper cites Taming diffusion models for audio-driven co-speech gesture generation.

Multi-human Interactive Talking Dataset Taming diffusion models for audio-driven co-speech gesture generation

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T04:46:06.486569Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T04:46:05.333389Z digest=sha256:dce20f750323acb8d79555ec2a93bc3c51f17d2dc5d915b94958752c62a009f7

Observation 8569be9f-6540-496b-9667-ed88b5a6cfde · outbound

This paper cites INFP: Audio-Driven Interactive Head Generation in Dyadic Conversations.

Multi-human Interactive Talking Dataset INFP: Audio-Driven Interactive Head Generation in Dyadic Conversations

Reference 60

Resolution
verified exact
local_arxiv, observed 2026-08-06T04:46:05.601149Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T04:46:05.382405Z digest=sha256:064190b96ef80b5472eb63cdd81ce5baeb4fd8922c3330243c9f33bc7ead00e5

Pith citing papers

No inbound Pith citation observations are available.