Pith. sign in

Paper Citation Record · LEDGER

MM-MovieDubber: Towards Multi-Modal Learning for Multi-Modal Movie Dubbing

As of 19 August 2026, this Paper Citation Record lists 33 of 33 outbound references and 1 inbound Pith citation observation for arXiv:2505.16279.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.16279 v1

Coverage vector

measured 33 of 33 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:07:10.491851Z

measured 34 of 34 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:07:07.961535Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-07T15:07:10.886803Z

Reference resolution

33 of 33 outbound references displayed

  • verified exact2
  • verified fuzzy20
  • unresolved10
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c74c4cec-69f9-42aa-861a-f36edf271657 · outbound

This paper cites Exist- ing dubbing methods can be categorized into two groups, each focusing on learning different styles of key prior information to generate high-quality voices.

MM-MovieDubber: Towards Multi-Modal Learning for Multi-Modal Movie Dubbing Exist- ing dubbing methods can be categorized into two groups, each focusing on learning different styles of key prior information to generate high-quality voices

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:07:13.014975Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T15:07:07.840682Z digest=sha256:56b23da8ffbc118542450eff504f90f0f9ab5e9231514d90c168dc2c1ddfdd30

Observation 74e62b26-d73e-4307-998b-7010f449b1a2 · outbound

This paper cites MM-MovieDubber: Towards Multi-Modal Learning for Multi-Modal Movie Dubbing.

MM-MovieDubber: Towards Multi-Modal Learning for Multi-Modal Movie Dubbing MM-MovieDubber: Towards Multi-Modal Learning for Multi-Modal Movie Dubbing

Reference 2

Resolution
malformed identifier
local_arxiv, observed 2026-08-07T15:07:10.927698Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T15:07:07.961535Z digest=sha256:fcfe0bd6285fc7f7458d42baa1a182e58c64bcb090f3e5c514c9cd81aa6da7c7

Observation 907a7eee-cea2-48df-a5a4-d728b32c66ec · outbound

This paper cites Datasets Emilia is a comprehensive multilingual speech generation dataset containing a total of 101,654 hours of speech data across six languages [21].

MM-MovieDubber: Towards Multi-Modal Learning for Multi-Modal Movie Dubbing Datasets Emilia is a comprehensive multilingual speech generation dataset containing a total of 101,654 hours of speech data across six languages [21]

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:07:12.775612Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T15:07:08.031502Z digest=sha256:19930d3252eb118bc48aab9a11092e7a6d86a576d4656dc7be9633d799cc38a5

Observation 375117c2-b77b-41b1-86d7-a80476fd686f · outbound

This paper cites To as- sess pronunciation accuracy, we use Word Error Rate (WER) with Whisper-V3[24] as the ASR model.

MM-MovieDubber: Towards Multi-Modal Learning for Multi-Modal Movie Dubbing To as- sess pronunciation accuracy, we use Word Error Rate (WER) with Whisper-V3[24] as the ASR model

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:07:12.637611Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T15:07:08.114058Z digest=sha256:adadbfe794aeeeb3ccf56e84369162b9007df6a2d75c53b404bf4d40039f4059

Observation 498d685b-d814-4216-89ac-e165952d658e · outbound

This paper cites Additionally, we have de- veloped a movie dubbing dataset with multi-type annotations to enhance movie understanding and improve dubbing quality.

MM-MovieDubber: Towards Multi-Modal Learning for Multi-Modal Movie Dubbing Additionally, we have de- veloped a movie dubbing dataset with multi-type annotations to enhance movie understanding and improve dubbing quality

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:07:12.543048Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T15:07:08.198474Z digest=sha256:e15d7d420284841c9e69d765dd8f0ba697ac0be1cbd3515d44a45e7a810e2c79

Observation 0aea5d8d-5aee-43db-9ce9-3407e0333e0f · outbound

This paper cites V2c: Vi- sual voice cloning,.

MM-MovieDubber: Towards Multi-Modal Learning for Multi-Modal Movie Dubbing V2c: Vi- sual voice cloning,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:07:12.439554Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T15:07:08.292579Z digest=sha256:68148cd70b87ae3ace727dfcf98996b0a0914ca427385b6965bcc5c074393c57

Observation fc319574-6ad6-4352-ad12-923d5cb6d196 · outbound

This paper cites More than words: In-the-wild visually-driven prosody for text-to-speech,.

MM-MovieDubber: Towards Multi-Modal Learning for Multi-Modal Movie Dubbing More than words: In-the-wild visually-driven prosody for text-to-speech,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:07:12.336180Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T15:07:08.360687Z digest=sha256:641ca21fb2ee1b9c44ca688a96133fb39327d5dd510e6be5b3d1c88cead4d1de

Observation 575950f6-6b37-4f81-b5a9-6886c6669eb4 · outbound

This paper cites Generalized end-to-end loss for speaker verification,.

MM-MovieDubber: Towards Multi-Modal Learning for Multi-Modal Movie Dubbing Generalized end-to-end loss for speaker verification,

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T15:07:08.454604Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:07:08.454604Z digest=sha256:8730a57d3fc1984a39fa71f6beb57eb45f5b5fa501cf03da359ac4065b1cacab

Observation d6475ee3-f567-450d-b6ad-8bbabb0b15de · outbound

This paper cites Learning to dub movies via hierarchical prosody models,.

MM-MovieDubber: Towards Multi-Modal Learning for Multi-Modal Movie Dubbing Learning to dub movies via hierarchical prosody models,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:07:12.220221Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T15:07:08.557957Z digest=sha256:07a418f16c27a1d6dffaef11be9efba32090c07f0c3f87703212516e7980419c

Observation 52d242fa-daf6-4206-9957-d93433bf20b2 · outbound

This paper cites Neu- ral dubber: Dubbing for videos according to scripts,.

MM-MovieDubber: Towards Multi-Modal Learning for Multi-Modal Movie Dubbing Neu- ral dubber: Dubbing for videos according to scripts,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:07:12.046317Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T15:07:08.676630Z digest=sha256:78fc4d38fb3333de2e354439967ac03f57477cc7cfffac9598d91a72abecc4d7

Observation bc5635ba-34aa-4c19-942f-ce1a908c9175 · outbound

This paper cites Imaginary voice: Face- styled diffusion model for text-to-speech,.

MM-MovieDubber: Towards Multi-Modal Learning for Multi-Modal Movie Dubbing Imaginary voice: Face- styled diffusion model for text-to-speech,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:07:11.944032Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T15:07:08.818516Z digest=sha256:90265ba8ca69e8e6b7839b18c71332366087b584fcb328bfe2751cf727196e2c

Observation 0912692c-d95d-4188-8855-7853c0c477ef · outbound

This paper cites Mcdubber: Multimodal context-aware expressive video dubbing,.

MM-MovieDubber: Towards Multi-Modal Learning for Multi-Modal Movie Dubbing Mcdubber: Multimodal context-aware expressive video dubbing,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:07:11.867601Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T15:07:08.943533Z digest=sha256:009c405150cdab7495a509115c3d5dd8f31d0ae952ace2d41feb547b285df25f

Observation 4c9175e7-1513-408d-8ac9-b6bb9cec84cd · outbound

This paper cites Audiopedia: Audio qa with knowledge,.

MM-MovieDubber: Towards Multi-Modal Learning for Multi-Modal Movie Dubbing Audiopedia: Audio qa with knowledge,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:07:11.766439Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T15:07:09.054916Z digest=sha256:4cd0cb5a05342df0c96d33c2d34a9359e2b1f29b9f7abd2143a40c7b5121b1bd

Observation e501b4aa-ab4b-48e9-bb12-9dc7b57355a7 · outbound

This paper cites Insight-V: Exploring Long-Chain Visual Reasoning with Multimodal Large Language Models.

MM-MovieDubber: Towards Multi-Modal Learning for Multi-Modal Movie Dubbing Insight-V: Exploring Long-Chain Visual Reasoning with Multimodal Large Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T15:07:09.148574Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:07:09.148574Z digest=sha256:adacc841102c46d4f14d39ecc429096a8711280f9ed34db72d5205fefb938726

Observation f8be3954-69be-438e-9ad5-4432531ef75d · outbound

This paper cites VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs.

MM-MovieDubber: Towards Multi-Modal Learning for Multi-Modal Movie Dubbing VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T15:07:09.258423Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:07:09.258423Z digest=sha256:646df8192221a8088bf78f571a656df1a113f859ae363f0b44f44bdb645f3748

Observation 6f78254c-f256-495b-aee1-8f76d4b0dadb · outbound

This paper cites Learning to dub movies via hierarchical prosody models,.

MM-MovieDubber: Towards Multi-Modal Learning for Multi-Modal Movie Dubbing Learning to dub movies via hierarchical prosody models,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:07:11.669537Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T15:07:09.359041Z digest=sha256:2efe954fbd2becfe02b0f44942aab797e8f6d85eb0ceb0a1b8394ac50ced9ef7

Observation 37ce9657-f384-4ed4-a705-5791653bab47 · outbound

This paper cites StyleDubber: Towards Multi-Scale Style Learning for Movie Dubbing.

MM-MovieDubber: Towards Multi-Modal Learning for Multi-Modal Movie Dubbing StyleDubber: Towards Multi-Scale Style Learning for Movie Dubbing

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T15:07:09.453123Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:07:09.453123Z digest=sha256:b44564fefee80e42db2a945bd8712fc4ca0333e04909114aa9d5e4851540a677

Observation 3a0d21ad-481e-4ab7-aa36-f256148f54d8 · outbound

This paper cites From speaker to dubber: Movie dubbing with prosody and duration consistency learning,.

MM-MovieDubber: Towards Multi-Modal Learning for Multi-Modal Movie Dubbing From speaker to dubber: Movie dubbing with prosody and duration consistency learning,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:07:11.589518Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T15:07:09.527245Z digest=sha256:8d4a5b453aa734327b28dc601cba3ee420fe74ca27a06e954a04d9f0c750ffe8

Observation a2ca5ff3-e452-482c-a4cc-9c0d122a62cd · outbound

This paper cites Visual instruction tuning,.

MM-MovieDubber: Towards Multi-Modal Learning for Multi-Modal Movie Dubbing Visual instruction tuning,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:07:11.508937Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T15:07:09.645485Z digest=sha256:f3529374c182cc3e14014f106737a3940794da00bb51d729b6c056ac0201b8ee

Observation fb820020-97f2-4de1-9232-86c7858f292d · outbound

This paper cites LLaVA-CoT: Let Vision Language Models Reason Step-by-Step.

MM-MovieDubber: Towards Multi-Modal Learning for Multi-Modal Movie Dubbing LLaVA-CoT: Let Vision Language Models Reason Step-by-Step

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T15:07:09.737133Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:07:09.737133Z digest=sha256:0943199ff19786b81870a4512b45307d850ff6455a75b38abaf60df1316364c0

Observation b22e3ad7-cc45-48fe-b3c1-c8e53935a9a1 · outbound

This paper cites F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching.

MM-MovieDubber: Towards Multi-Modal Learning for Multi-Modal Movie Dubbing F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T15:07:09.853775Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:07:09.853775Z digest=sha256:8127f45c54758ecc6671837f21a6c29356cbd23a52cee8663e7662e8e205c6ca

Observation 297f42d7-bd54-44b4-aed5-644aafba4cb8 · outbound

This paper cites Flow matching for generative modeling,.

MM-MovieDubber: Towards Multi-Modal Learning for Multi-Modal Movie Dubbing Flow matching for generative modeling,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:07:11.422727Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T15:07:09.981897Z digest=sha256:882c695944ece9a2fbd885ef53517cf2a9079c83c64ee32b149f905bc975d6ed

Observation e275e094-90a4-4712-9e20-fe0c7af3ff27 · outbound

This paper cites An audio-visual corpus for speech perception and automatic speech recognition,.

MM-MovieDubber: Towards Multi-Modal Learning for Multi-Modal Movie Dubbing An audio-visual corpus for speech perception and automatic speech recognition,

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T15:07:10.333160Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:07:10.333160Z digest=sha256:56f87306b3d9320056bd7fab9268c78d9989d914a38b840d78289567ff21a844

Observation 3d6cc985-0e82-459d-850e-bb737c7a2e05 · outbound

This paper cites Exploring the limits of transfer learning with a unified text-to-text transformer,.

MM-MovieDubber: Towards Multi-Modal Learning for Multi-Modal Movie Dubbing Exploring the limits of transfer learning with a unified text-to-text transformer,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:07:11.331277Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T15:07:10.067509Z digest=sha256:37d637eeaf7c0afd668da74d387bb34e3c359af794da880b441aab136b4bc46f

Observation 0cdbc5a6-b3fb-4b09-aebe-f4d579bae2a7 · outbound

This paper cites Learning transferable visual models from natural language supervision,.

MM-MovieDubber: Towards Multi-Modal Learning for Multi-Modal Movie Dubbing Learning transferable visual models from natural language supervision,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:07:11.240478Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T15:07:10.120717Z digest=sha256:03454d2a4efc97490222cc91ecbfea0289036859364392f5e0dc08b109fb546e

Observation 710d353d-232f-42a6-875b-bfde1a713ea6 · outbound

This paper cites DiVE: Dit-based video generation with enhanced control,.

MM-MovieDubber: Towards Multi-Modal Learning for Multi-Modal Movie Dubbing DiVE: Dit-based video generation with enhanced control,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:07:11.170141Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T15:07:10.174456Z digest=sha256:d7db64f98b0e810bc7b96d20e883c1de5bfb4b505c13aba74c2dc63ecca4c8c1

Observation f3fe19ef-6c80-4bbd-9835-209e6f266201 · outbound

This paper cites Emilia: An Extensive, Multilingual, and Diverse Speech Dataset for Large-Scale Speech Generation.

MM-MovieDubber: Towards Multi-Modal Learning for Multi-Modal Movie Dubbing Emilia: An Extensive, Multilingual, and Diverse Speech Dataset for Large-Scale Speech Generation

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T15:07:10.212017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:07:10.212017Z digest=sha256:245b00da0afdacc7dcedc9c79f63c352b716aa9318a542f771b5b485da9a4359

Observation 5e30b8f5-a9a6-4481-8b01-4d2924fa4b96 · outbound

This paper cites V2C: Visual Voice Cloning.

MM-MovieDubber: Towards Multi-Modal Learning for Multi-Modal Movie Dubbing V2C: Visual Voice Cloning

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-08-07T15:07:10.775143Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T15:07:10.279629Z digest=sha256:97f991fcac516d6c368ed34a65d45c65c4627c4ef9ac0367d78dda03ba4c0c58

Observation 092bb8b5-818c-403a-9762-39b6ad937b9a · outbound

This paper cites Robust Speech Recognition via Large-Scale Weak Supervision.

MM-MovieDubber: Towards Multi-Modal Learning for Multi-Modal Movie Dubbing Robust Speech Recognition via Large-Scale Weak Supervision

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T15:07:10.363576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:07:10.363576Z digest=sha256:45f42074365c007216f7ed9deab818393bc94e1c2dbb2322af4743e127e431fe

Observation fd3fd5fe-5aa3-4ab2-b700-eb822cd43a8a · outbound

This paper cites Location-Relative Attention Mechanisms For Robust Long-Form Speech Synthesis.

MM-MovieDubber: Towards Multi-Modal Learning for Multi-Modal Movie Dubbing Location-Relative Attention Mechanisms For Robust Long-Form Speech Synthesis

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-08-07T15:07:10.612708Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T15:07:10.404068Z digest=sha256:b930d56e614003e6ca4565503d379106bf6224e88f7a4b1a2dd17f4769022cc3

Observation 9edd3850-dac4-4e3a-a2ff-8b7a03a45d5b · outbound

This paper cites Tem- poral modeling matters: A novel temporal emotional modeling approach for speech emotion recognition,.

MM-MovieDubber: Towards Multi-Modal Learning for Multi-Modal Movie Dubbing Tem- poral modeling matters: A novel temporal emotional modeling approach for speech emotion recognition,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:07:11.100841Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T15:07:10.445186Z digest=sha256:2b996f2df80f8fb8f8938ba3da59cda80332228a77abe9b167a8b1288e7cf840

Observation 0c2c5a5c-a7cc-448a-aa58-3a770be52bba · outbound

This paper cites Out of time: automated lip sync in the wild,.

MM-MovieDubber: Towards Multi-Modal Learning for Multi-Modal Movie Dubbing Out of time: automated lip sync in the wild,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:07:11.031819Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T15:07:10.491851Z digest=sha256:f05ef6d18037d28e79e670662ae22e306e062412348e86da0f790030e7446118

Observation afeebc99-ff03-4cc8-8e88-c14dd6001225 · outbound

This paper cites Available: https://openreview.net/forum?id= PqvMRDCJT9t.

MM-MovieDubber: Towards Multi-Modal Learning for Multi-Modal Movie Dubbing Available: https://openreview.net/forum?id= PqvMRDCJT9t

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-07T15:07:10.028412Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:07:10.028412Z digest=sha256:7d4ca1507a44c94f8f3545d8cdaf3d107878165d136982fd8fdf1ac6b7e4d83e

Pith citing papers

Observation 74e62b26-d73e-4307-998b-7010f449b1a2 · inbound

MM-MovieDubber: Towards Multi-Modal Learning for Multi-Modal Movie Dubbing cites this paper.

MM-MovieDubber: Towards Multi-Modal Learning for Multi-Modal Movie Dubbing MM-MovieDubber: Towards Multi-Modal Learning for Multi-Modal Movie Dubbing

Reference 2

Resolution
malformed identifier
local_arxiv, observed 2026-08-07T15:07:10.927698Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T15:07:07.961535Z digest=sha256:fcfe0bd6285fc7f7458d42baa1a182e58c64bcb090f3e5c514c9cd81aa6da7c7