Pith. sign in

Paper Citation Record · LEDGER

MuseTalk: Real-Time High-Fidelity Video Dubbing via Spatio-Temporal Sampling

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 19 inbound Pith citation observations for arXiv:2410.10122.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2410.10122 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 19 of 19 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 19 of 19 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T22:20:13.390763Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T04:47:38.011253Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 3bf6e830-fa96-46aa-8213-d2887cacab16 · inbound

RiverEcho: Real-Time Interactive Digital System for Ancient Yellow River Culture cites this paper.

RiverEcho: Real-Time Interactive Digital System for Ancient Yellow River Culture MuseTalk: Real-Time High-Fidelity Video Dubbing via Spatio-Temporal Sampling

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T22:20:13.390763Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:20:13.390763Z digest=sha256:3f82913b871b5ee821581900792d4e0575cb99b0316085f20faeb611df8562db

Observation c045272e-bab1-4bf1-a1f3-b3c58fd6f7c3 · inbound

Fine-Grained Zero-Shot Object Detection cites this paper.

Fine-Grained Zero-Shot Object Detection MuseTalk: Real-Time High-Fidelity Video Dubbing via Spatio-Temporal Sampling

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-06T17:41:02.903909Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:41:02.903909Z digest=sha256:a20215d670f4e9c36ddbf29af3555e84b7aa1fdf37f5f12affc7c7ce1e28cdbc

Observation 74d21d0f-12d3-4872-b90f-f88063bff5ab · inbound

MS-DETR: Towards Effective Video Moment Retrieval and Highlight Detection by Joint Motion-Semantic Learning cites this paper.

MS-DETR: Towards Effective Video Moment Retrieval and Highlight Detection by Joint Motion-Semantic Learning MuseTalk: Real-Time High-Fidelity Video Dubbing via Spatio-Temporal Sampling

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-06T17:01:57.033054Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:01:57.033054Z digest=sha256:bcac0b6aa2027ecdcbb5eaee0762e39867a87b544fdc9c4d92f7aec7eb09eeb7

Observation 9811c709-e290-485d-8a75-120a73b6c8ff · inbound

JOLT3D: Joint Learning of Talking Heads and 3DMM Parameters with Application to Lip-Sync cites this paper.

JOLT3D: Joint Learning of Talking Heads and 3DMM Parameters with Application to Lip-Sync MuseTalk: Real-Time High-Fidelity Video Dubbing via Spatio-Temporal Sampling

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-06T13:39:35.346930Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T13:39:35.346930Z digest=sha256:b3f3730041fa368b4d711ecb8f11ca1eca887b0a5cd9abaf4c360fed8788d48a

Observation 554011ca-c2db-4c6d-bc9a-557ade5b7d1b · inbound

Who is a Better Talker: Subjective and Objective Quality Assessment for AI-Generated Talking Heads cites this paper.

Who is a Better Talker: Subjective and Objective Quality Assessment for AI-Generated Talking Heads MuseTalk: Real-Time High-Fidelity Video Dubbing via Spatio-Temporal Sampling

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T10:53:19.154197Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T10:53:19.154197Z digest=sha256:a664d3499b113356761bf8ecc9eb6e71c9039a9d1eaae949cce70e630a1337b5

Observation 4a19a2d5-4538-4199-8f15-632ebf821315 · inbound

InfiniteTalk: Audio-driven Video Generation for Sparse-Frame Video Dubbing cites this paper.

InfiniteTalk: Audio-driven Video Generation for Sparse-Frame Video Dubbing MuseTalk: Real-Time High-Fidelity Video Dubbing via Spatio-Temporal Sampling

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-05T18:50:16.311819Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T18:50:16.311819Z digest=sha256:b4553253409b8f46283ca0f58a2faa565fe38cfe3e72f662c6fec1dd0c6dac58

Observation 90ff5ada-57db-41fe-bf4f-0ecb1f85de0d · inbound

FluentAvatar: Flicker-Free Talking-Head Animation via Phoneme-Guided Autoregressive Modeling cites this paper.

FluentAvatar: Flicker-Free Talking-Head Animation via Phoneme-Guided Autoregressive Modeling MuseTalk: Real-Time High-Fidelity Video Dubbing via Spatio-Temporal Sampling

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-18T16:42:43.818772Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-18T16:42:25.803856Z digest=sha256:6669a98d7c396e2cc7c891efa596a09e836b85d8847cc2f93d6150fd1a60aa7b

Observation 88af652b-816c-426b-9966-cf5ed06e921f · inbound

MIND-V: Hierarchical World Model for Long-Horizon Robotic Manipulation with RL-based Physical Alignment cites this paper.

MIND-V: Hierarchical World Model for Long-Horizon Robotic Manipulation with RL-based Physical Alignment MuseTalk: Real-Time High-Fidelity Video Dubbing via Spatio-Temporal Sampling

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-03T18:09:52.881855Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:09:52.881855Z digest=sha256:d10231cc0c93d4f5fb29573dbe69dcc644cf5a2c6c5d3f1e1d81f59a9d6e7e82

Observation a37a3846-f6ce-4a5e-98c0-87cdc70ff99f · inbound

EAD-Net: Emotion-Aware Talking Head Generation with Spatial Refinement and Temporal Coherence cites this paper.

EAD-Net: Emotion-Aware Talking Head Generation with Spatial Refinement and Temporal Coherence MuseTalk: Real-Time High-Fidelity Video Dubbing via Spatio-Temporal Sampling

Reference 55

Resolution
verified exact
arxiv_id, observed 2026-05-11T20:36:11.231717Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-08T08:27:41.839123Z digest=sha256:4c85fd4920b9b2c4027a7111b835c484030ce293a8ed1c41ae0cb1c3cb243da2

Observation 2ac94527-4d4f-4665-8b36-b66dede390be · inbound

Hallo-Live: Real-Time Streaming Joint Audio-Video Avatar Generation with Asynchronous Dual-Stream and Human-Centric Preference Distillation cites this paper.

Hallo-Live: Real-Time Streaming Joint Audio-Video Avatar Generation with Asynchronous Dual-Stream and Human-Centric Preference Distillation MuseTalk: Real-Time High-Fidelity Video Dubbing via Spatio-Temporal Sampling

Reference 52

Resolution
verified exact
arxiv_id, observed 2026-05-11T21:06:12.895871Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-08T06:56:19.795651Z digest=sha256:a0e20c0fd25fd6307387fd194eb2b9c60bd9cd384aa8197848f1bcab28a7ff8e

Observation 8a48b7fb-336e-4f81-b0eb-43feeb2e5363 · inbound

Fre-Res: Frequency-Residual Video Token Compression for Efficient Video MLLMs cites this paper.

Fre-Res: Frequency-Residual Video Token Compression for Efficient Video MLLMs MuseTalk: Real-Time High-Fidelity Video Dubbing via Spatio-Temporal Sampling

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-20T23:03:50.705507Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-20T23:01:15.886434Z digest=sha256:46ed7811eee2a4111a306dfa048654fb2d778f67d231633214d02ad5f04b9a25

Observation f0bcc245-6bbf-47ea-985c-c15129de1490 · inbound

Fre-Res: Frequency-Residual Video Token Compression for Efficient Video MLLMs cites this paper.

Fre-Res: Frequency-Residual Video Token Compression for Efficient Video MLLMs MuseTalk: Real-Time High-Fidelity Video Dubbing via Spatio-Temporal Sampling

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-02T14:32:54.590138Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T14:32:54.590138Z digest=sha256:3ddac6d65f8aeff33f3f8766f2371e9913a89a5547cdf0670a09c58a5c5353b6

Observation 5914fffe-a6e0-4506-8cd8-579769c5f438 · inbound

HighSync: High-Quality Lip Synchronization via Latent Diffusion Models cites this paper.

HighSync: High-Quality Lip Synchronization via Latent Diffusion Models MuseTalk: Real-Time High-Fidelity Video Dubbing via Spatio-Temporal Sampling

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-19T21:17:48.632095Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-19T21:14:25.606097Z digest=sha256:0dd841a5950fb9f7928cf909f679de12fb68c5138a5fe4b36e696e4b23b50d97

Observation 5d5100b1-824e-4d2e-83d1-65e4672194d3 · inbound

Lip Forcing: Few-Step Autoregressive Diffusion for Real-time Lip Synchronization cites this paper.

Lip Forcing: Few-Step Autoregressive Diffusion for Real-time Lip Synchronization MuseTalk: Real-Time High-Fidelity Video Dubbing via Spatio-Temporal Sampling

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-07-03T04:47:38.012668Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-27T13:40:14.506288Z digest=sha256:ce2e918a7796258a022da97b69c4a5e0e71f8a149c7da2212ead7428912d75da

Observation 661c553b-0133-47e6-8d93-e98fdc407a11 · inbound

MindFlow: Harmonizing Cognitive Semantics and Acoustic Dynamics for Facial Animation Generation in Dyadic Conversations cites this paper.

MindFlow: Harmonizing Cognitive Semantics and Acoustic Dynamics for Facial Animation Generation in Dyadic Conversations MuseTalk: Real-Time High-Fidelity Video Dubbing via Spatio-Temporal Sampling

Reference 47

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T18:33:51.213900Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-29T05:04:31.366845Z digest=sha256:fe8c78f298b73f756050e02a7564203a76eade0d667eed9ba57f8d15ab2effcb

Observation 511abe14-d690-4b64-8f87-673f607e6ae4 · inbound

KM-Speaker: Keypoint-Based Style Control for High-Quality Speech-Driven 3D Facial Animation and Dialogue Localization cites this paper.

KM-Speaker: Keypoint-Based Style Control for High-Quality Speech-Driven 3D Facial Animation and Dialogue Localization MuseTalk: Real-Time High-Fidelity Video Dubbing via Spatio-Temporal Sampling

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-07-01T15:45:48.569495Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-30T01:07:49.064263Z digest=sha256:25b9c336be7389c2bda04087c7ace8125a3167bd33854a26da14391e3f10c7b9

Observation 820bd89c-a5ba-4416-a4d3-cb2c25c59470 · inbound

FlashRT: Agent Harness for Guiding Agents to Deploy Real-Time Multimodal Applications cites this paper.

FlashRT: Agent Harness for Guiding Agents to Deploy Real-Time Multimodal Applications MuseTalk: Real-Time High-Fidelity Video Dubbing via Spatio-Temporal Sampling

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-01T15:52:51.792816Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:52:51.792816Z digest=sha256:79e6adb3f1845253612bbdd4eedc412a8d9500ad8f156ed6e95bb1c9fdaff7b1

Observation 21d2069e-e574-49a4-8cdf-c505334d2d27 · inbound

LeapTalk: Breaking the Latency-Quality Trade-off in Talking Head Generation cites this paper.

LeapTalk: Breaking the Latency-Quality Trade-off in Talking Head Generation MuseTalk: Real-Time High-Fidelity Video Dubbing via Spatio-Temporal Sampling

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-04T01:32:12.897470Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T01:32:12.897470Z digest=sha256:79af09988bb905755306fc06c89168df0c6034632b5f8276c17dcb65e36c0230

Observation 2a6fab00-b4cf-4027-81d2-52a40deb78b8 · inbound

Geometry-guided Emotion Modulation for Controllable and Photorealistic Emotional Talking Face Generation cites this paper.

Geometry-guided Emotion Modulation for Controllable and Photorealistic Emotional Talking Face Generation MuseTalk: Real-Time High-Fidelity Video Dubbing via Spatio-Temporal Sampling

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-04T01:03:13.376570Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T01:03:13.376570Z digest=sha256:82e70f83543c28de3ac229bdf8d9c56a19c331f3ea1fa155957056c4a807da09