Pith. sign in

Paper Citation Record · LEDGER

Bind-Your-Avatar: Multi-Talking-Character Video Generation with Dynamic 3D-mask-based Embedding Router

As of 18 August 2026, this Paper Citation Record lists 56 of 56 outbound references and 4 inbound Pith citation observations for arXiv:2506.19833.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.19833 v1

Coverage vector

measured 56 of 56 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T18:29:41.247052Z

measured 60 of 60 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T20:06:46.703429Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-12T07:21:23.863541Z

Reference resolution

56 of 56 outbound references displayed

  • verified exact2
  • verified fuzzy25
  • unresolved29
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 9ae6cfe5-ec19-4624-9fb5-bce06d374d66 · outbound

This paper cites wav2vec 2.0: A framework for self- supervised learning of speech representations.Advances in neural information processing systems, 2020.

Bind-Your-Avatar: Multi-Talking-Character Video Generation with Dynamic 3D-mask-based Embedding Router wav2vec 2.0: A framework for self- supervised learning of speech representations.Advances in neural information processing systems, 2020

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:29:42.536440Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T18:29:40.874718Z digest=sha256:ee307dabb793f1d202dc8a99ef4013216b67ce15f703a9b39b1d14e6e1798293

Observation 67cbf082-6afe-47f7-92ae-964a83bc28a0 · outbound

This paper cites Qwen Technical Report.

Bind-Your-Avatar: Multi-Talking-Character Video Generation with Dynamic 3D-mask-based Embedding Router Qwen Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-15T18:29:40.882526Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:29:40.882526Z digest=sha256:89093ef4397c537260ccc4deb9b0b1f662e5ddf271b79cf4d674a163a1e58258

Observation fc12339c-6837-4ef5-b3dc-a834be500af1 · outbound

This paper cites DanceTogether! Identity-Preserving Multi-Person Interactive Video Generation.

Bind-Your-Avatar: Multi-Talking-Character Video Generation with Dynamic 3D-mask-based Embedding Router DanceTogether! Identity-Preserving Multi-Person Interactive Video Generation

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T18:29:40.891406Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:29:40.891406Z digest=sha256:200fe85db33e5f7022c8c139b8444acbf606b948ce1ac06f930e2fdaf624fa78

Observation bbce61ed-b268-474d-b389-49af8bed024c · outbound

This paper cites HunyuanVideo-Avatar: High-Fidelity Audio-Driven Human Animation for Multiple Characters.

Bind-Your-Avatar: Multi-Talking-Character Video Generation with Dynamic 3D-mask-based Embedding Router HunyuanVideo-Avatar: High-Fidelity Audio-Driven Human Animation for Multiple Characters

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-15T18:29:40.897824Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:29:40.897824Z digest=sha256:086385f0d0a8e6d8e63012c65964c0eab78153f5eafed4c7ecb4747c6fbdcb59

Observation 071dee7c-97a3-44ba-a21b-8e192002fdca · outbound

This paper cites EchoMimic: Lifelike Audio-Driven Portrait Animations through Editable Landmark Conditions.

Bind-Your-Avatar: Multi-Talking-Character Video Generation with Dynamic 3D-mask-based Embedding Router EchoMimic: Lifelike Audio-Driven Portrait Animations through Editable Landmark Conditions

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-15T18:29:40.904311Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:29:40.904311Z digest=sha256:697d66f2d8f84a4e930301350041abf1a61637df6ee106eee568e5da2db6d110

Observation 4303f6c1-5ee0-41ad-8b16-bdf48627e9e0 · outbound

This paper cites Out of time: Automated lip sync in the wild.

Bind-Your-Avatar: Multi-Talking-Character Video Generation with Dynamic 3D-mask-based Embedding Router Out of time: Automated lip sync in the wild

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:29:42.517595Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T18:29:40.911198Z digest=sha256:167f1fa8cfc26f59ac56d91ce196bcc268d171833066b322c572083e86d11e17

Observation 413fcc20-f2e7-4be2-896b-3263e2d8569d · outbound

This paper cites Hallo2: Long-Duration and High-Resolution Audio-Driven Portrait Image Animation.

Bind-Your-Avatar: Multi-Talking-Character Video Generation with Dynamic 3D-mask-based Embedding Router Hallo2: Long-Duration and High-Resolution Audio-Driven Portrait Image Animation

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T18:29:40.918924Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:29:40.918924Z digest=sha256:71e53aee73b8e009d77c28af749eb8e10ec04c45db9a40f22c2913415812de01

Observation 894cea0e-1e9c-4818-b037-35899f37e355 · outbound

This paper cites Hallo3: Highly Dynamic and Realistic Portrait Image Animation with Video Diffusion Transformer.

Bind-Your-Avatar: Multi-Talking-Character Video Generation with Dynamic 3D-mask-based Embedding Router Hallo3: Highly Dynamic and Realistic Portrait Image Animation with Video Diffusion Transformer

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T18:29:40.930102Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:29:40.930102Z digest=sha256:152653bc521c6c0d8da18f26d6e216962b7fe62236e6daf618169efc99ac89c8

Observation a3cddfa5-64db-4e63-b678-8e547a6ddc66 · outbound

This paper cites Insightface: An open-source 2d and 3d deep face analysis toolkit.https://github.com/ deepinsight / insightface, 2024.

Bind-Your-Avatar: Multi-Talking-Character Video Generation with Dynamic 3D-mask-based Embedding Router Insightface: An open-source 2d and 3d deep face analysis toolkit.https://github.com/ deepinsight / insightface, 2024

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:29:42.499171Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T18:29:40.937172Z digest=sha256:feb4c0dffe273e9dbddcaa3c1438f86f2d251e8e156fb1cc38b0977c17290ef3

Observation e11e0319-0115-4f3c-872c-b6a6bf87de3a · outbound

This paper cites MAGREF: Masked Guidance for Any-Reference Video Generation.arXiv preprint arXiv:2505.23742,.

Bind-Your-Avatar: Multi-Talking-Character Video Generation with Dynamic 3D-mask-based Embedding Router MAGREF: Masked Guidance for Any-Reference Video Generation.arXiv preprint arXiv:2505.23742,

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T18:29:40.943044Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:29:40.943044Z digest=sha256:8cff1f3cda739fa74a122b63bdb604ddf7d909b9ef4098b074516a36dc0db641

Observation 51c4967b-8fd6-4bf9-89cc-0a5730644613 · outbound

This paper cites CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models.

Bind-Your-Avatar: Multi-Talking-Character Video Generation with Dynamic 3D-mask-based Embedding Router CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T18:29:40.948861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:29:40.948861Z digest=sha256:34b12042af512c41d0d3bc35e826cb95bcbe26a166ecedc6e0ba21ede426fd2d

Observation 1d2e6acc-5680-44a5-9f9d-1102e2ec2c34 · outbound

This paper cites Ingredients: Blending custom photos with video diffusion transformers, 2025.

Bind-Your-Avatar: Multi-Talking-Character Video Generation with Dynamic 3D-mask-based Embedding Router Ingredients: Blending custom photos with video diffusion transformers, 2025

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:29:42.475568Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T18:29:40.955567Z digest=sha256:512bf02da52e8037ea602fb2729e5832f39b3de6290e4d3005f561926c6e7421

Observation c8588426-c846-4c61-b75a-53637396cf4f · outbound

This paper cites AD-NeRF: Audio Driven Neural Radiance Fields for Talking Head Synthesis.

Bind-Your-Avatar: Multi-Talking-Character Video Generation with Dynamic 3D-mask-based Embedding Router AD-NeRF: Audio Driven Neural Radiance Fields for Talking Head Synthesis

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:29:42.457046Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T18:29:40.960991Z digest=sha256:66e0577b5b692bfab2926b6b915824e0fdcfd5f6a3e0d31c0e7060b27734a4be

Observation 123a781b-c7b8-438f-a9fb-3a403dc6500a · outbound

This paper cites Gans trained by a two time-scale update rule converge to a local nash equilibrium.Advances in neural information processing systems, 2017.

Bind-Your-Avatar: Multi-Talking-Character Video Generation with Dynamic 3D-mask-based Embedding Router Gans trained by a two time-scale update rule converge to a local nash equilibrium.Advances in neural information processing systems, 2017

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:29:42.437516Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T18:29:40.967072Z digest=sha256:0d7634e8c2d37288be81ff93cc41f6b86006e6b9bd9f673f9bbf205bba36ee5e

Observation 0f2a3271-9c90-48cc-9389-d450073965bd · outbound

This paper cites Classifier-Free Diffusion Guidance.

Bind-Your-Avatar: Multi-Talking-Character Video Generation with Dynamic 3D-mask-based Embedding Router Classifier-Free Diffusion Guidance

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-15T18:29:40.972801Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:29:40.972801Z digest=sha256:56c750e8cc3c8b366566823b2c5f9841d9616b81390addc74cc063b35964afcc

Observation 9455f4c6-90f8-491f-8e1c-509e2c678394 · outbound

This paper cites Depth-Aware Generative Adversarial Network for Talking Head Video Generation.

Bind-Your-Avatar: Multi-Talking-Character Video Generation with Dynamic 3D-mask-based Embedding Router Depth-Aware Generative Adversarial Network for Talking Head Video Generation

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-08-15T18:29:41.818978Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T18:29:40.979832Z digest=sha256:be014b68f158267e1e2a38b3b3830bcb0797e12f47d99eb4d7e438d07f1c7449

Observation 0ac8b7b0-0015-45a8-816a-4c24221f424b · outbound

This paper cites Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen.

Bind-Your-Avatar: Multi-Talking-Character Video Generation with Dynamic 3D-mask-based Embedding Router Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:29:42.419051Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T18:29:40.985848Z digest=sha256:6a4deb14980509733cf775d7c7ccb2b67f92750e5d0b93378cbfda1a5967bb77

Observation e1cc1add-0c90-4bf9-8013-b43a171df0f9 · outbound

This paper cites ConceptMaster: Multi-Concept Video Customization on Diffusion Transformer Models Without Test-Time Tuning.

Bind-Your-Avatar: Multi-Talking-Character Video Generation with Dynamic 3D-mask-based Embedding Router ConceptMaster: Multi-Concept Video Customization on Diffusion Transformer Models Without Test-Time Tuning

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-15T18:29:40.991681Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:29:40.991681Z digest=sha256:f60b29bd1fa68be0f8b71527e0b8f32372c73500ab7a95d5e0235c25e6b11b53

Observation 3185e00d-a467-465d-8335-82e1606546b4 · outbound

This paper cites Sonic: Shifting Focus to Global Audio Perception in Portrait Animation.

Bind-Your-Avatar: Multi-Talking-Character Video Generation with Dynamic 3D-mask-based Embedding Router Sonic: Shifting Focus to Global Audio Perception in Portrait Animation

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-15T18:29:40.998391Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:29:40.998391Z digest=sha256:ee408f119649e93e8efb3f83d1b288f72079a94c2cea71540d40afe10cbc6720

Observation 15a9ad3f-e860-469a-8b65-14ddb0865d37 · outbound

This paper cites Ultralytics yolov8, 2023.

Bind-Your-Avatar: Multi-Talking-Character Video Generation with Dynamic 3D-mask-based Embedding Router Ultralytics yolov8, 2023

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:29:42.402576Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T18:29:41.005080Z digest=sha256:ae6a95a5f1358e4f53f67ee4a42be5cd42468bf04b51872b0d5f67402337fabe

Observation aa4d0478-a48b-47e1-9c22-bf1c1b355111 · outbound

This paper cites Let Them Talk: Audio-Driven Multi-Person Conversational Video Generation.

Bind-Your-Avatar: Multi-Talking-Character Video Generation with Dynamic 3D-mask-based Embedding Router Let Them Talk: Audio-Driven Multi-Person Conversational Video Generation

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-15T18:29:41.011307Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:29:41.011307Z digest=sha256:5f8682044b9b38f7ea101d433375d445925b1cbf0eea69c5a96e0a4a7a1f3bca

Observation 0ae4bcb1-cfd6-436e-84f5-b1099d0c7f87 · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

Bind-Your-Avatar: Multi-Talking-Character Video Generation with Dynamic 3D-mask-based Embedding Router Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:29:42.385262Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T18:29:41.017052Z digest=sha256:e0b73019f0d6407b065e1169db5cae800bc2414173645b2beb840f13d3826c3b

Observation 61c42d9a-1c42-4ce1-8918-8f2b5726e6f3 · outbound

This paper cites OmniHuman-1: Rethinking the Scaling-Up of One-Stage Conditioned Human Animation Models.

Bind-Your-Avatar: Multi-Talking-Character Video Generation with Dynamic 3D-mask-based Embedding Router OmniHuman-1: Rethinking the Scaling-Up of One-Stage Conditioned Human Animation Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-15T18:29:41.022920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:29:41.022920Z digest=sha256:fb4cc516000d4a43973961d25476d810341a4b5106480e74f2a628bbc7a470f9

Observation 5a88c2ad-a372-417a-9364-9ad441156dad · outbound

This paper cites EchoMimicV2: Towards Striking, Simpli- fied, and Semi-Body Human Animation.arXiv preprint arXiv:2411.10061, 2024.

Bind-Your-Avatar: Multi-Talking-Character Video Generation with Dynamic 3D-mask-based Embedding Router EchoMimicV2: Towards Striking, Simpli- fied, and Semi-Body Human Animation.arXiv preprint arXiv:2411.10061, 2024

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-15T18:29:41.029029Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:29:41.029029Z digest=sha256:6f8a20d75f646e968ff98a77bea773a6bb20d295968ae3b3e2b065be82a685ae

Observation 439bd3d7-e367-4777-9bec-566685bc2968 · outbound

This paper cites Let's Go Real Talk: Spoken Dialogue Model for Face-to-Face Conversation.

Bind-Your-Avatar: Multi-Talking-Character Video Generation with Dynamic 3D-mask-based Embedding Router Let's Go Real Talk: Spoken Dialogue Model for Face-to-Face Conversation

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-15T18:29:41.035031Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:29:41.035031Z digest=sha256:e9fdff4d1703f28e7660ec862bb4360227849b95310a808ceb9c4c8f950a4316

Observation 37fc7fac-fece-4695-a863-abbfb37dc2ae · outbound

This paper cites an unresolved cited work.

Bind-Your-Avatar: Multi-Talking-Character Video Generation with Dynamic 3D-mask-based Embedding Router Unresolved cited work

Reference 26

Resolution
unresolved
raw_fallback, observed 2026-08-15T18:29:42.368470Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T18:29:41.040389Z digest=sha256:543fbd35245b0b17db967d80148b9df5a32a666b5a576d6398918d448ab1480d

Observation b0eb454a-880c-4762-8bb2-5e3c32d7d3e7 · outbound

This paper cites ChatAnyone: Stylized Real-time Portrait Video Generation with Hierarchical Motion Diffusion Model.

Bind-Your-Avatar: Multi-Talking-Character Video Generation with Dynamic 3D-mask-based Embedding Router ChatAnyone: Stylized Real-time Portrait Video Generation with Hierarchical Motion Diffusion Model

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-15T18:29:41.046871Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:29:41.046871Z digest=sha256:7c6d2ef8eef4eccaed1029fa2165b301b03e8ef553d3e88ef47d3a545be185de

Observation 1089d791-55eb-4a4b-8a70-e20c46eda12b · outbound

This paper cites Learning transferable visual models from natural lan- guage supervision.

Bind-Your-Avatar: Multi-Talking-Character Video Generation with Dynamic 3D-mask-based Embedding Router Learning transferable visual models from natural lan- guage supervision

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:29:42.351972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T18:29:41.053142Z digest=sha256:fa186e645dd9977e0a8b050fa7d0f05ef1ba1fd0027b77c428afa8c00fd50962

Observation ecd26ade-a319-47b2-a691-85ced96058b3 · outbound

This paper cites Exploring the limits of transfer learn- ing with a unified text-to-text transformer.Journal of machine learning research, 2020.

Bind-Your-Avatar: Multi-Talking-Character Video Generation with Dynamic 3D-mask-based Embedding Router Exploring the limits of transfer learn- ing with a unified text-to-text transformer.Journal of machine learning research, 2020

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:29:42.333141Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T18:29:41.060900Z digest=sha256:340e67257361b1dda3a6be1037bc54883f201838fe039685868a16e4803c2969

Observation 82b4e6de-7bc0-4979-bdc7-19f390b510a2 · outbound

This paper cites SAM 2: Segment Anything in Images and Videos.

Bind-Your-Avatar: Multi-Talking-Character Video Generation with Dynamic 3D-mask-based Embedding Router SAM 2: Segment Anything in Images and Videos

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-15T18:29:41.066092Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:29:41.066092Z digest=sha256:864e0878d47e1af10c814d38443a2303e13c6510ca43fe4b641825afeb882942

Observation 185b3d47-9e4b-4acd-948a-ea5257ad4149 · outbound

This paper cites Facenet: A unified embedding for face recog- nition and clustering.

Bind-Your-Avatar: Multi-Talking-Character Video Generation with Dynamic 3D-mask-based Embedding Router Facenet: A unified embedding for face recog- nition and clustering

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:29:42.316829Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T18:29:41.071353Z digest=sha256:926b687ff15c0e84c4d79138da691a97f741ff6ca821ab59f3613213887b2f12

Observation 3bcd6596-ef3a-4c17-8633-ccbc8491f0b8 · outbound

This paper cites Clearvoice: An open-source audio-visual speech processing toolkit.

Bind-Your-Avatar: Multi-Talking-Character Video Generation with Dynamic 3D-mask-based Embedding Router Clearvoice: An open-source audio-visual speech processing toolkit

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:29:42.300001Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T18:29:41.076852Z digest=sha256:5c242113c4ea7cb1b80dc04bbeddcaaffc7c8dcad7e223dae1f004a96e83d427

Observation 3209e392-4c96-4471-94d8-462e36ccfac2 · outbound

This paper cites Roformer: Enhanced trans- former with rotary position embedding.Neurocomput- ing, 2024.

Bind-Your-Avatar: Multi-Talking-Character Video Generation with Dynamic 3D-mask-based Embedding Router Roformer: Enhanced trans- former with rotary position embedding.Neurocomput- ing, 2024

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:29:42.279039Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T18:29:41.081947Z digest=sha256:0af12d4292a760e96ba1474f6ec91f18a208cb8aed14b64ee3c141294bcae16b

Observation ea381be4-5efc-4aaa-8261-4c3570f8d594 · outbound

This paper cites Rethinking the inception architecture for computer vision.

Bind-Your-Avatar: Multi-Talking-Character Video Generation with Dynamic 3D-mask-based Embedding Router Rethinking the inception architecture for computer vision

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:29:42.261882Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T18:29:41.087505Z digest=sha256:e900278b9e789ec0f8fbbea8aa73c11a10549234c46c6550a09f1c81c41f1b3d

Observation 1465142b-8e95-4a94-9831-939cb069a48c · outbound

This paper cites EMO: Emote Portrait Alive Generating Expressive Por- trait Videos with Audio2Video Diffusion Model Under Weak Conditions.European Conference on Computer Vision, 2024.

Bind-Your-Avatar: Multi-Talking-Character Video Generation with Dynamic 3D-mask-based Embedding Router EMO: Emote Portrait Alive Generating Expressive Por- trait Videos with Audio2Video Diffusion Model Under Weak Conditions.European Conference on Computer Vision, 2024

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:29:42.244304Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T18:29:41.092984Z digest=sha256:e0423f761ef4141bb066236a90d3b381eb244c3a4864d06b423ea3d3e217c126

Observation da3f1d5b-9ad4-4a4b-a287-15db42aa1353 · outbound

This paper cites EMO2: End-Effector Guided Audio-Driven Avatar Video Generation.

Bind-Your-Avatar: Multi-Talking-Character Video Generation with Dynamic 3D-mask-based Embedding Router EMO2: End-Effector Guided Audio-Driven Avatar Video Generation

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-15T18:29:41.099508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:29:41.099508Z digest=sha256:669801c917bf27e456a6218c1119d659bbbdb2585b0816b8a0e0d8b17e1f0caa

Observation f81d63b5-77a2-43fb-8adf-0d56fb2f5862 · outbound

This paper cites Fvd: A new metric for video generation.Inter- national Conference on Learning Representations, 2019.

Bind-Your-Avatar: Multi-Talking-Character Video Generation with Dynamic 3D-mask-based Embedding Router Fvd: A new metric for video generation.Inter- national Conference on Learning Representations, 2019

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:29:42.224889Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T18:29:41.106135Z digest=sha256:127201c4c5873b5fabb3f7af9ee4b979350e9c304faf31d3ae3206c40dea0162

Observation 6635004c-1458-4206-9d79-89e730e8ec3a · outbound

This paper cites FantasyTalking: Realistic Talking Portrait Generation via Coherent Motion Synthesis.

Bind-Your-Avatar: Multi-Talking-Character Video Generation with Dynamic 3D-mask-based Embedding Router FantasyTalking: Realistic Talking Portrait Generation via Coherent Motion Synthesis

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-15T18:29:41.112453Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:29:41.112453Z digest=sha256:742afaf4c74cae468c3f20112afb2f44aa9dcd9c3717bb6bcab280dcfae66058

Observation f4977eb1-9df5-4fb5-afbb-6613645e2b27 · outbound

This paper cites JoyGen: Audio-Driven 3D Depth-Aware Talking-Face Video Editing.

Bind-Your-Avatar: Multi-Talking-Character Video Generation with Dynamic 3D-mask-based Embedding Router JoyGen: Audio-Driven 3D Depth-Aware Talking-Face Video Editing

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-15T18:29:41.120881Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:29:41.120881Z digest=sha256:ad436eda76d896a33f917815ac02b55d22c726b52d2212923cc378a96ab676f4

Observation 94a023f8-43c8-4607-8568-0501ea41f3a1 · outbound

This paper cites Friends-mmc: A dataset for multi-modal multi-party conversation under- standing.

Bind-Your-Avatar: Multi-Talking-Character Video Generation with Dynamic 3D-mask-based Embedding Router Friends-mmc: A dataset for multi-modal multi-party conversation under- standing

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:29:42.205860Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T18:29:41.128270Z digest=sha256:a6811afb2568f75189fc20ad0941eb8ddaf13c621aa6d0f79dd1bf861eeea243

Observation eb926631-9f96-4f14-965e-53cf9066cb53 · outbound

This paper cites MoCha: Towards Movie-Grade Talking Character Synthesis.

Bind-Your-Avatar: Multi-Talking-Character Video Generation with Dynamic 3D-mask-based Embedding Router MoCha: Towards Movie-Grade Talking Character Synthesis

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-15T18:29:41.134588Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:29:41.134588Z digest=sha256:697385e1afa1833a175b5b7536ee09873c04df18bee421a979f8b43695bab4c0

Observation 526a7123-c880-447b-b8fc-f246009736d0 · outbound

This paper cites Seeking the Shape of Sound: An Adaptive Framework for Learning V oice-Face Association.Computer Vision and Pattern Recognition, 2021.

Bind-Your-Avatar: Multi-Talking-Character Video Generation with Dynamic 3D-mask-based Embedding Router Seeking the Shape of Sound: An Adaptive Framework for Learning V oice-Face Association.Computer Vision and Pattern Recognition, 2021

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:29:42.188175Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T18:29:41.141552Z digest=sha256:830a38938ffd912ea91b950632d92bac00b837365338696d4fea0f7121e9779b

Observation c6eef73f-22e4-4800-92df-6f11d9dc3317 · outbound

This paper cites Sophia Koepke, and Andrew Zisserman.

Bind-Your-Avatar: Multi-Talking-Character Video Generation with Dynamic 3D-mask-based Embedding Router Sophia Koepke, and Andrew Zisserman

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:29:42.166403Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T18:29:41.147803Z digest=sha256:222f6fe1e5082c3bce235ca0a556ddb4596e408a44ce24a4622d6e3ef1f1b973

Observation 8d64fd63-8ca5-424f-bfd0-f54732f4fe4c · outbound

This paper cites Vfhq: A high-quality dataset and benchmark for video face super-resolution.

Bind-Your-Avatar: Multi-Talking-Character Video Generation with Dynamic 3D-mask-based Embedding Router Vfhq: A high-quality dataset and benchmark for video face super-resolution

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:29:42.142862Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T18:29:41.152970Z digest=sha256:9eec319e7aa71e4bb139dc51fb9ba75f9ddcb54a347cd6e3a20c2f3e2f371f7b

Observation 46fe26e0-e612-4faf-914e-64ce8570f192 · outbound

This paper cites MegActor-$\Sigma$: Unlocking Flexible Mixed-Modal Control in Portrait Animation with Diffusion Transformer.

Bind-Your-Avatar: Multi-Talking-Character Video Generation with Dynamic 3D-mask-based Embedding Router MegActor-$\Sigma$: Unlocking Flexible Mixed-Modal Control in Portrait Animation with Diffusion Transformer

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-15T18:29:41.160355Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:29:41.160355Z digest=sha256:11ea08f7a1a5f1cceeca023eb1b46023df03cfba24484001c0881f5bdb4cd414

Observation 5490a590-c3e7-428c-96b6-2d729b5d2735 · outbound

This paper cites CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer.

Bind-Your-Avatar: Multi-Talking-Character Video Generation with Dynamic 3D-mask-based Embedding Router CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-15T18:29:41.166356Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:29:41.166356Z digest=sha256:12d54ab21f7d2efa2e4852b0fa652f38d1577df2f2405a3b0eb2d350e11ad8c2

Observation f6fefac0-7751-4e12-b0d0-766902a442ef · outbound

This paper cites MagicInfinite: Generating Infinite Talking Videos with Your Words and Voice.

Bind-Your-Avatar: Multi-Talking-Character Video Generation with Dynamic 3D-mask-based Embedding Router MagicInfinite: Generating Infinite Talking Videos with Your Words and Voice

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-15T18:29:41.174301Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:29:41.174301Z digest=sha256:c7703ef37eca33e80c8e30e31d3f06254c2dbd25d232ba2455b5f79f43be7f21

Observation a60fa5e2-e5a0-4658-bb2f-0dffa95a4423 · outbound

This paper cites Identity-preserving text-to-video generation by frequency decomposition.

Bind-Your-Avatar: Multi-Talking-Character Video Generation with Dynamic 3D-mask-based Embedding Router Identity-preserving text-to-video generation by frequency decomposition

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:29:42.122815Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T18:29:41.180742Z digest=sha256:585201d6b89483fb5372fa998748cbbdeffa0ec96361962eca9be937034a66bd

Observation 603caf57-e8fd-4926-947d-21f2bc562819 · outbound

This paper cites Multimodal Diffusion Transformer with Memory Bank for Scalable Long-Duration Talking Video Generation.

Bind-Your-Avatar: Multi-Talking-Character Video Generation with Dynamic 3D-mask-based Embedding Router Multimodal Diffusion Transformer with Memory Bank for Scalable Long-Duration Talking Video Generation

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-15T18:29:41.186964Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:29:41.186964Z digest=sha256:47c8e11c85c481aa75ecf1cf0c39f9d6da7cf4e1fb47646a7e71f958fb3d8bc2

Observation 9693cf19-839d-469d-9428-0eb9033de889 · outbound

This paper cites Flow-guided one-shot talking face generation with a high-resolution audio-visual dataset.

Bind-Your-Avatar: Multi-Talking-Character Video Generation with Dynamic 3D-mask-based Embedding Router Flow-guided one-shot talking face generation with a high-resolution audio-visual dataset

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:29:42.102635Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T18:29:41.194524Z digest=sha256:2f43ab37e9eddf8b2e59c123aa3f03d8586cb57d8ca9a3c2ec2cdf337599ef73

Observation ff7a4695-16b7-4eb7-a8de-d792b9ad4f75 · outbound

This paper cites Teller: Real-Time Streaming Audio-Driven Portrait Animation with Autoregressive Motion Generation.

Bind-Your-Avatar: Multi-Talking-Character Video Generation with Dynamic 3D-mask-based Embedding Router Teller: Real-Time Streaming Audio-Driven Portrait Animation with Autoregressive Motion Generation

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-15T18:29:41.203323Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:29:41.203323Z digest=sha256:50d070bfd4b55fcae70303fcb1888aa71a6e30351cf8a9cbdc37818eafbe112e

Observation f35c5b6b-082a-4bdf-9cd7-1ee8743b90de · outbound

This paper cites Identity-Preserving Talking Face Generation with Landmark and Appearance Priors.

Bind-Your-Avatar: Multi-Talking-Character Video Generation with Dynamic 3D-mask-based Embedding Router Identity-Preserving Talking Face Generation with Landmark and Appearance Priors

Reference 52

Resolution
verified exact
local_arxiv, observed 2026-08-15T18:29:41.349435Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T18:29:41.211423Z digest=sha256:f09a727f4a2df1288d98c0b9e667d2d42f21db1824e183c417f3edc1e554b370

Observation a55f6395-40da-4d8c-b826-c1f05efd0f3b · outbound

This paper cites Concat-ID: Towards Universal Identity-Preserving Video Synthesis.

Bind-Your-Avatar: Multi-Talking-Character Video Generation with Dynamic 3D-mask-based Embedding Router Concat-ID: Towards Universal Identity-Preserving Video Synthesis

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-15T18:29:41.221434Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:29:41.221434Z digest=sha256:e72aac6cdbb7c3c31b5f3d97273826fd37df26238204f579e5f226f62ac87d17

Observation ff41877d-6a52-4214-82ca-ba08d7e8b4c0 · outbound

This paper cites Responsive listening head gener- ation: a benchmark dataset and baseline.

Bind-Your-Avatar: Multi-Talking-Character Video Generation with Dynamic 3D-mask-based Embedding Router Responsive listening head gener- ation: a benchmark dataset and baseline

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:29:42.084227Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T18:29:41.232003Z digest=sha256:a09afa01f4e7f995fe7b6d1d78d748015a6968b00501967a5726b47c5ad2046d

Observation 206ac9ac-1855-4936-931f-e1632cf09446 · outbound

This paper cites Makeittalk: Speaker-aware talking-head animation.

Bind-Your-Avatar: Multi-Talking-Character Video Generation with Dynamic 3D-mask-based Embedding Router Makeittalk: Speaker-aware talking-head animation

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:29:42.065317Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T18:29:41.240795Z digest=sha256:057e28bbdb27b0e84bfc0d348ee652e25cf54d436b522fbb9794068ad0ff73b2

Observation 65f06d46-f32a-4f29-bb42-726d23356af8 · outbound

This paper cites INFP: Audio-Driven Interactive Head Generation in Dyadic Conversations.

Bind-Your-Avatar: Multi-Talking-Character Video Generation with Dynamic 3D-mask-based Embedding Router INFP: Audio-Driven Interactive Head Generation in Dyadic Conversations

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-15T18:29:41.247052Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:29:41.247052Z digest=sha256:824f1b7c8576d876f764973fbcd82c5e591e16c434a8b89068fbe6eb02603dc5

Pith citing papers

Observation 6972d79b-1c29-4e4c-962c-3ad9928f8c59 · inbound

FantasyTalking2: Timestep-Layer Adaptive Preference Optimization for Audio-Driven Portrait Animation cites this paper.

FantasyTalking2: Timestep-Layer Adaptive Preference Optimization for Audio-Driven Portrait Animation Bind-Your-Avatar: Multi-Talking-Character Video Generation with Dynamic 3D-mask-based Embedding Router

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-05T20:06:46.703429Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:06:46.703429Z digest=sha256:272ab16698d0beab453a5fa494498dbc605bb4126737319233a160e1a5aa53ee

Observation 637da00f-17f4-4838-b577-1697b3dc2884 · inbound

MIDAS: Multimodal Interactive Digital-humAn Synthesis via Real-time Autoregressive Video Generation cites this paper.

MIDAS: Multimodal Interactive Digital-humAn Synthesis via Real-time Autoregressive Video Generation Bind-Your-Avatar: Multi-Talking-Character Video Generation with Dynamic 3D-mask-based Embedding Router

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-05T16:01:30.754418Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:01:30.754418Z digest=sha256:cc1298bafceb63c74b62f1d6648169061921f9ee08d80bd1cfd32e486c8f7c2e

Observation 1000af5a-71de-40fa-a0fb-7c935a561a92 · inbound

SocialDirector: Training-Free Social Interaction Control for Multi-Person Video Generation cites this paper.

SocialDirector: Training-Free Social Interaction Control for Multi-Person Video Generation Bind-Your-Avatar: Multi-Talking-Character Video Generation with Dynamic 3D-mask-based Embedding Router

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-12T07:21:23.866335Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-12T03:28:30.471533Z digest=sha256:a7c61337fa66d03f4b28ae71f8a4672b25c6580bcf1f790d63d0db0a7aa1670a

Observation 6b4962d2-9e61-441b-ad8b-528320b0f77e · inbound

EchoCache: Energy-Guided Cross-Modal Caching for Efficient Audio-Driven Video Generation cites this paper.

EchoCache: Energy-Guided Cross-Modal Caching for Efficient Audio-Driven Video Generation Bind-Your-Avatar: Multi-Talking-Character Video Generation with Dynamic 3D-mask-based Embedding Router

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-04T06:45:47.112983Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:45:47.112983Z digest=sha256:2df9151108d62748403e1bf0af5ce52d0a29a48571ebc7c984950f775a82e10f