Pith. sign in

Paper Citation Record · LEDGER

MuseTalk: Real-Time High-Fidelity Video Dubbing via Spatio-Temporal Sampling

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 19 inbound Pith citation observations for arXiv:2410.10122.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2410.10122 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 19 of 19 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 19 of 19 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T22:20:13.390763Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T04:47:38.011253Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 3bf6e830-fa96-46aa-8213-d2887cacab16 · inbound

RiverEcho: Real-Time Interactive Digital System for Ancient Yellow River Culture cites this paper.

RiverEcho: Real-Time Interactive Digital System for Ancient Yellow River Culture MuseTalk: Real-Time High-Fidelity Video Dubbing via Spatio-Temporal Sampling

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T22:20:13.390763Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:20:13.390763Z digest=sha256:ff89a7b40d04d8197ffb75354287311ddc65c5a70b94434dd46c495580f3a3df

Observation c045272e-bab1-4bf1-a1f3-b3c58fd6f7c3 · inbound

Fine-Grained Zero-Shot Object Detection cites this paper.

Fine-Grained Zero-Shot Object Detection MuseTalk: Real-Time High-Fidelity Video Dubbing via Spatio-Temporal Sampling

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-06T17:41:02.903909Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:41:02.903909Z digest=sha256:b4bba05e0ac93fc417524e90492a28e09b1cfabeb96930a6f80039377d29601d

Observation 74d21d0f-12d3-4872-b90f-f88063bff5ab · inbound

MS-DETR: Towards Effective Video Moment Retrieval and Highlight Detection by Joint Motion-Semantic Learning cites this paper.

MS-DETR: Towards Effective Video Moment Retrieval and Highlight Detection by Joint Motion-Semantic Learning MuseTalk: Real-Time High-Fidelity Video Dubbing via Spatio-Temporal Sampling

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-06T17:01:57.033054Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:01:57.033054Z digest=sha256:bcac0b6aa2027ecdcbb5eaee0762e39867a87b544fdc9c4d92f7aec7eb09eeb7

Observation 9811c709-e290-485d-8a75-120a73b6c8ff · inbound

JOLT3D: Joint Learning of Talking Heads and 3DMM Parameters with Application to Lip-Sync cites this paper.

JOLT3D: Joint Learning of Talking Heads and 3DMM Parameters with Application to Lip-Sync MuseTalk: Real-Time High-Fidelity Video Dubbing via Spatio-Temporal Sampling

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-06T13:39:35.346930Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:39:35.346930Z digest=sha256:133ed1b7faa5b3f654c5442f02e488c65fdd5998ac8b77bd2cce566738cba521

Observation 554011ca-c2db-4c6d-bc9a-557ade5b7d1b · inbound

Who is a Better Talker: Subjective and Objective Quality Assessment for AI-Generated Talking Heads cites this paper.

Who is a Better Talker: Subjective and Objective Quality Assessment for AI-Generated Talking Heads MuseTalk: Real-Time High-Fidelity Video Dubbing via Spatio-Temporal Sampling

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T10:53:19.154197Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T10:53:19.154197Z digest=sha256:a664d3499b113356761bf8ecc9eb6e71c9039a9d1eaae949cce70e630a1337b5

Observation 4a19a2d5-4538-4199-8f15-632ebf821315 · inbound

InfiniteTalk: Audio-driven Video Generation for Sparse-Frame Video Dubbing cites this paper.

InfiniteTalk: Audio-driven Video Generation for Sparse-Frame Video Dubbing MuseTalk: Real-Time High-Fidelity Video Dubbing via Spatio-Temporal Sampling

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-05T18:50:16.311819Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T18:50:16.311819Z digest=sha256:b4553253409b8f46283ca0f58a2faa565fe38cfe3e72f662c6fec1dd0c6dac58

Observation 90ff5ada-57db-41fe-bf4f-0ecb1f85de0d · inbound

FluentAvatar: Flicker-Free Talking-Head Animation via Phoneme-Guided Autoregressive Modeling cites this paper.

FluentAvatar: Flicker-Free Talking-Head Animation via Phoneme-Guided Autoregressive Modeling MuseTalk: Real-Time High-Fidelity Video Dubbing via Spatio-Temporal Sampling

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-18T16:42:43.818772Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-18T16:42:25.803856Z digest=sha256:6685f4041e7a6d1ed858307fc2f4740cd4e6a2ce565cf630b0d5d2cb7e00e819

Observation 88af652b-816c-426b-9966-cf5ed06e921f · inbound

MIND-V: Hierarchical World Model for Long-Horizon Robotic Manipulation with RL-based Physical Alignment cites this paper.

MIND-V: Hierarchical World Model for Long-Horizon Robotic Manipulation with RL-based Physical Alignment MuseTalk: Real-Time High-Fidelity Video Dubbing via Spatio-Temporal Sampling

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-03T18:09:52.881855Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:09:52.881855Z digest=sha256:490e28cd8dc1b951e703ed25dd73875fca8afc474899a72372d29b90935593dd

Observation a37a3846-f6ce-4a5e-98c0-87cdc70ff99f · inbound

EAD-Net: Emotion-Aware Talking Head Generation with Spatial Refinement and Temporal Coherence cites this paper.

EAD-Net: Emotion-Aware Talking Head Generation with Spatial Refinement and Temporal Coherence MuseTalk: Real-Time High-Fidelity Video Dubbing via Spatio-Temporal Sampling

Reference 55

Resolution
verified exact
arxiv_id, observed 2026-05-11T20:36:11.231717Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-08T08:27:41.839123Z digest=sha256:9e982842546b33ac2e25b744887a6d9d011632459f40fef02632004685d21674

Observation 2ac94527-4d4f-4665-8b36-b66dede390be · inbound

Hallo-Live: Real-Time Streaming Joint Audio-Video Avatar Generation with Asynchronous Dual-Stream and Human-Centric Preference Distillation cites this paper.

Hallo-Live: Real-Time Streaming Joint Audio-Video Avatar Generation with Asynchronous Dual-Stream and Human-Centric Preference Distillation MuseTalk: Real-Time High-Fidelity Video Dubbing via Spatio-Temporal Sampling

Reference 52

Resolution
verified exact
arxiv_id, observed 2026-05-11T21:06:12.895871Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-08T06:56:19.795651Z digest=sha256:1143637a1114f38ba71d662de123ddeeedb66002d01feef0a2f2167d3466696c

Observation 8a48b7fb-336e-4f81-b0eb-43feeb2e5363 · inbound

Fre-Res: Frequency-Residual Video Token Compression for Efficient Video MLLMs cites this paper.

Fre-Res: Frequency-Residual Video Token Compression for Efficient Video MLLMs MuseTalk: Real-Time High-Fidelity Video Dubbing via Spatio-Temporal Sampling

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-20T23:03:50.705507Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-20T23:01:15.886434Z digest=sha256:05ea80bcfcf09c8923ed9d80f0fdf763d09e551d6a0e1f0c6be64352db35c0ee

Observation f0bcc245-6bbf-47ea-985c-c15129de1490 · inbound

Fre-Res: Frequency-Residual Video Token Compression for Efficient Video MLLMs cites this paper.

Fre-Res: Frequency-Residual Video Token Compression for Efficient Video MLLMs MuseTalk: Real-Time High-Fidelity Video Dubbing via Spatio-Temporal Sampling

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-02T14:32:54.590138Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T14:32:54.590138Z digest=sha256:3ddac6d65f8aeff33f3f8766f2371e9913a89a5547cdf0670a09c58a5c5353b6

Observation 5914fffe-a6e0-4506-8cd8-579769c5f438 · inbound

HighSync: High-Quality Lip Synchronization via Latent Diffusion Models cites this paper.

HighSync: High-Quality Lip Synchronization via Latent Diffusion Models MuseTalk: Real-Time High-Fidelity Video Dubbing via Spatio-Temporal Sampling

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-19T21:17:48.632095Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-19T21:14:25.606097Z digest=sha256:b908fc1cd0a8c4dee43a280cc4f6be20255468e7c45d74c3aa1c5640065f58e8

Observation 5d5100b1-824e-4d2e-83d1-65e4672194d3 · inbound

Lip Forcing: Few-Step Autoregressive Diffusion for Real-time Lip Synchronization cites this paper.

Lip Forcing: Few-Step Autoregressive Diffusion for Real-time Lip Synchronization MuseTalk: Real-Time High-Fidelity Video Dubbing via Spatio-Temporal Sampling

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-07-03T04:47:38.012668Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-27T13:40:14.506288Z digest=sha256:cbbce905ef323e6a89f331d0361f3bfcb976b32095e98ce82d595fe04eae2f2d

Observation 661c553b-0133-47e6-8d93-e98fdc407a11 · inbound

MindFlow: Harmonizing Cognitive Semantics and Acoustic Dynamics for Facial Animation Generation in Dyadic Conversations cites this paper.

MindFlow: Harmonizing Cognitive Semantics and Acoustic Dynamics for Facial Animation Generation in Dyadic Conversations MuseTalk: Real-Time High-Fidelity Video Dubbing via Spatio-Temporal Sampling

Reference 47

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T18:33:51.213900Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-29T05:04:31.366845Z digest=sha256:f6d7c742c9b610ecf341dfca9ae2d4f5dcd651bf0a04fd8e8c5165f7d040870d

Observation 511abe14-d690-4b64-8f87-673f607e6ae4 · inbound

KM-Speaker: Keypoint-Based Style Control for High-Quality Speech-Driven 3D Facial Animation and Dialogue Localization cites this paper.

KM-Speaker: Keypoint-Based Style Control for High-Quality Speech-Driven 3D Facial Animation and Dialogue Localization MuseTalk: Real-Time High-Fidelity Video Dubbing via Spatio-Temporal Sampling

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-07-01T15:45:48.569495Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-30T01:07:49.064263Z digest=sha256:708f66160dc92900449f954d3a68b93151102e8c4e53dba5f4c262321a52219b

Observation 820bd89c-a5ba-4416-a4d3-cb2c25c59470 · inbound

FlashRT: Agent Harness for Guiding Agents to Deploy Real-Time Multimodal Applications cites this paper.

FlashRT: Agent Harness for Guiding Agents to Deploy Real-Time Multimodal Applications MuseTalk: Real-Time High-Fidelity Video Dubbing via Spatio-Temporal Sampling

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-01T15:52:51.792816Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:52:51.792816Z digest=sha256:79e6adb3f1845253612bbdd4eedc412a8d9500ad8f156ed6e95bb1c9fdaff7b1

Observation 21d2069e-e574-49a4-8cdf-c505334d2d27 · inbound

LeapTalk: Breaking the Latency-Quality Trade-off in Talking Head Generation cites this paper.

LeapTalk: Breaking the Latency-Quality Trade-off in Talking Head Generation MuseTalk: Real-Time High-Fidelity Video Dubbing via Spatio-Temporal Sampling

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-04T01:32:12.897470Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T01:32:12.897470Z digest=sha256:79af09988bb905755306fc06c89168df0c6034632b5f8276c17dcb65e36c0230

Observation 2a6fab00-b4cf-4027-81d2-52a40deb78b8 · inbound

Geometry-guided Emotion Modulation for Controllable and Photorealistic Emotional Talking Face Generation cites this paper.

Geometry-guided Emotion Modulation for Controllable and Photorealistic Emotional Talking Face Generation MuseTalk: Real-Time High-Fidelity Video Dubbing via Spatio-Temporal Sampling

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-04T01:03:13.376570Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:03:13.376570Z digest=sha256:a82e0e9bad57bb249ab24a8c0330e50f87d7da3027eb7aff677c7b008130308e