Pith. sign in

Paper Citation Record · LEDGER

Towards Expressive Video Dubbing with Multiscale Multimodal Context Interaction

As of 13 August 2026, this Paper Citation Record lists 37 of 37 outbound references and 0 inbound Pith citation observations for arXiv:2412.18748.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.18748 v2

Coverage vector

measured 37 of 37 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T04:34:07.920960Z

measured 37 of 37 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

37 of 37 outbound references displayed

  • verified exact5
  • verified fuzzy14
  • unresolved18
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 1764f373-12b8-4dc3-9752-73911e284520 · outbound

This paper cites Neural dubber: Dubbing for videos according to scripts,.

Towards Expressive Video Dubbing with Multiscale Multimodal Context Interaction Neural dubber: Dubbing for videos according to scripts,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:34:08.436892Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T04:34:07.758539Z digest=sha256:1661825f8bf6077bb16c5a43bb2567836f183427589318fc9d35cf18e7e30005

Observation ec54db63-1103-45d6-a4df-f48b18939241 · outbound

This paper cites Prosody Modeling with 3D Visual Information for Expressive Video Dubbing,.

Towards Expressive Video Dubbing with Multiscale Multimodal Context Interaction Prosody Modeling with 3D Visual Information for Expressive Video Dubbing,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:34:08.423769Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T04:34:07.763464Z digest=sha256:f834b5cf713e5ea231de5b9a6ad86b9ea8beacf53e79d6c14e15f7f5aa50d07c

Observation 532688cd-8e25-4821-9be8-8dccc87e6896 · outbound

This paper cites StyleDubber: Towards Multi-Scale Style Learning for Movie Dubbing.

Towards Expressive Video Dubbing with Multiscale Multimodal Context Interaction StyleDubber: Towards Multi-Scale Style Learning for Movie Dubbing

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T04:34:07.768994Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:34:07.768994Z digest=sha256:ca9e0c55946f11a510b608bacd77e3b54e77356c558cebc5386d3a072fde974f

Observation d8e07d3b-dd93-47d1-a79c-27271ebc6bc6 · outbound

This paper cites V2c: Visual voice cloning,.

Towards Expressive Video Dubbing with Multiscale Multimodal Context Interaction V2c: Visual voice cloning,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:34:08.410226Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T04:34:07.774060Z digest=sha256:f00fcb0e42300e426f3f61aec16e403b330eecc0a568ab5b8bf3f2d76b65f227

Observation 6d8cef30-e21c-4608-acba-195e24dab3aa · outbound

This paper cites From speaker to dubber: Movie dubbing with prosody and duration consistency learning,.

Towards Expressive Video Dubbing with Multiscale Multimodal Context Interaction From speaker to dubber: Movie dubbing with prosody and duration consistency learning,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:34:08.397469Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T04:34:07.778636Z digest=sha256:e40684c26c3246cb22e9e074f13915b2ff9736b18e15eee422d06f8285a44cab

Observation 3d42e29e-c8bc-407d-8f4b-c075f902d825 · outbound

This paper cites EmoDubber: Towards High Quality and Emotion Controllable Movie Dubbing.

Towards Expressive Video Dubbing with Multiscale Multimodal Context Interaction EmoDubber: Towards High Quality and Emotion Controllable Movie Dubbing

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T04:34:07.783191Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:34:07.783191Z digest=sha256:9d2d6284eeceaf60d5354568fa6775f4c2c25345cedaa79042d9d22a2f99711a

Observation 076a3700-5358-47d4-b854-4be31dabfaa0 · outbound

This paper cites High-Quality Automatic Voice Over with Accurate Alignment: Supervision through Self-Supervised Discrete Speech Units.

Towards Expressive Video Dubbing with Multiscale Multimodal Context Interaction High-Quality Automatic Voice Over with Accurate Alignment: Supervision through Self-Supervised Discrete Speech Units

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-08-11T04:34:08.115409Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T04:34:07.788553Z digest=sha256:9c6160ae424a2a82dca3043c1189c11a8973a259336d76f716394e70372d03b7

Observation a9d75b0b-ad46-4407-adea-1b9b0a18b657 · outbound

This paper cites DubWise: Video-Guided Speech Duration Control in Multimodal LLM-based Text-to-Speech for Dubbing.

Towards Expressive Video Dubbing with Multiscale Multimodal Context Interaction DubWise: Video-Guided Speech Duration Control in Multimodal LLM-based Text-to-Speech for Dubbing

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T04:34:07.793247Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:34:07.793247Z digest=sha256:a76fa2d784e76bbfc51e498202648f55f3168906d80cd47c67ff19616c66ba41

Observation 10c9ed64-3d2a-46ff-baba-9e02046012bd · outbound

This paper cites More than words: In-the-wild visually-driven prosody for text-to-speech,.

Towards Expressive Video Dubbing with Multiscale Multimodal Context Interaction More than words: In-the-wild visually-driven prosody for text-to-speech,

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T04:34:07.797858Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:34:07.797858Z digest=sha256:089c6397dcbfd6b539a339a7c573fceed28b6cdd5aac5d9783ea292ce2c7c0a7

Observation ec185281-1a87-4f95-95d9-0d25343fd108 · outbound

This paper cites Learning to dub movies via hierarchical prosody models,.

Towards Expressive Video Dubbing with Multiscale Multimodal Context Interaction Learning to dub movies via hierarchical prosody models,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:34:08.374498Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T04:34:07.803591Z digest=sha256:59290aa2d25bdf9733f687f231f26262ced2f61f2d6b45ac02bc2537020fd666

Observation eb396945-041d-4a16-9cc0-c33c93495eeb · outbound

This paper cites MCDubber: Multimodal Context-Aware Expressive Video Dubbing.

Towards Expressive Video Dubbing with Multiscale Multimodal Context Interaction MCDubber: Multimodal Context-Aware Expressive Video Dubbing

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T04:34:07.808156Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:34:07.808156Z digest=sha256:3dbcc0c07aa530d5b0447202b53e3b91b4a991f10720606637234e8abd084454

Observation 5b6c150b-4206-40f6-9818-d1ee99cd7d13 · outbound

This paper cites To- wards Multi-Scale Speaking Style Modelling with Hierarchical Context Information for Mandarin Speech Synthesis,.

Towards Expressive Video Dubbing with Multiscale Multimodal Context Interaction To- wards Multi-Scale Speaking Style Modelling with Hierarchical Context Information for Mandarin Speech Synthesis,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:34:08.359596Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T04:34:07.813123Z digest=sha256:25eb66a18caf8fc6d0d907b27e86d2b071bd627a10481a3e3702954c9a314b24

Observation 0d71bcdc-de81-40a7-9b32-2f5b9f125c49 · outbound

This paper cites Msstyletts: Multi-scale style modeling with hierarchical context infor- mation for expressive speech synthesis,.

Towards Expressive Video Dubbing with Multiscale Multimodal Context Interaction Msstyletts: Multi-scale style modeling with hierarchical context infor- mation for expressive speech synthesis,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:34:08.345484Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T04:34:07.817279Z digest=sha256:5dd49892a358e986f5276f35fe0e4454cc9113b721b1a187bc0657541f22d011

Observation 688ab624-d08f-4aa7-9a80-136595686d56 · outbound

This paper cites Unsupervised multi-scale expressive speaking style modeling with hierarchical context information for audiobook speech synthesis,.

Towards Expressive Video Dubbing with Multiscale Multimodal Context Interaction Unsupervised multi-scale expressive speaking style modeling with hierarchical context information for audiobook speech synthesis,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:34:08.331259Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T04:34:07.821403Z digest=sha256:3a9f342f7508a143504d313989e358f1d9fba15af9f49afa787daff9c005cb20

Observation edfbc176-425b-45b1-8274-b302725aed01 · outbound

This paper cites Mae-dfer: Efficient masked au- toencoder for self-supervised dynamic facial expression recognition,.

Towards Expressive Video Dubbing with Multiscale Multimodal Context Interaction Mae-dfer: Efficient masked au- toencoder for self-supervised dynamic facial expression recognition,

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T04:34:07.825600Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:34:07.825600Z digest=sha256:e7b47afb46d5d1f993dd177da846957c2fc38e9527798c0d1586699422666fe7

Observation 37cda586-b2a8-43b7-a36e-3a46b127c0eb · outbound

This paper cites Dfew: A large-scale database for recognizing dynamic facial expressions in the wild,.

Towards Expressive Video Dubbing with Multiscale Multimodal Context Interaction Dfew: A large-scale database for recognizing dynamic facial expressions in the wild,

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T04:34:07.830150Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:34:07.830150Z digest=sha256:de3259af4c3c4c8ec2254b86941a2b09eea9a2e71f3aa697b3e6b117c8ef4542

Observation 0a5f69df-dd96-4e6b-99c6-b53bed2baedc · outbound

This paper cites Estimation of continuous valence and arousal levels from faces in naturalistic conditions,.

Towards Expressive Video Dubbing with Multiscale Multimodal Context Interaction Estimation of continuous valence and arousal levels from faces in naturalistic conditions,

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T04:34:07.834265Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:34:07.834265Z digest=sha256:2a8c29b14a458bb7c5e2b75caea3ff124359c670f2a5c1dad6f29082d93a29fc

Observation ce456576-3b5c-4c7f-83c9-3b088757a057 · outbound

This paper cites Towards Multi-Scale Style Control for Expressive Speech Synthesis.

Towards Expressive Video Dubbing with Multiscale Multimodal Context Interaction Towards Multi-Scale Style Control for Expressive Speech Synthesis

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-08-11T04:34:08.066385Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T04:34:07.838476Z digest=sha256:6f865f4df86106608a27a813fe46cfc375ce2c70f47e0d64714b1de5446717ae

Observation f6da0022-bbc2-459a-b689-5e853cf2a636 · outbound

This paper cites Fastspeech: Fast, robust and controllable text to speech,.

Towards Expressive Video Dubbing with Multiscale Multimodal Context Interaction Fastspeech: Fast, robust and controllable text to speech,

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T04:34:07.843110Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:34:07.843110Z digest=sha256:84d0a0651d2ac5aa9ff09a71eba23c1cfa21ff28d582623675d7c92a12a189b9

Observation 60da494e-3e91-4d06-9a2b-37f6263c0626 · outbound

This paper cites Emotion-english-roberta-large,.

Towards Expressive Video Dubbing with Multiscale Multimodal Context Interaction Emotion-english-roberta-large,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:34:08.281979Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T04:34:07.847568Z digest=sha256:8f38e857af75f83b974f3712a9688a3b3cad350dce6ef082000641c336901ea2

Observation 8bd7cb49-60bb-4b7c-af54-203db6ed751b · outbound

This paper cites FastSpeech 2: Fast and High-Quality End-to-End Text to Speech.

Towards Expressive Video Dubbing with Multiscale Multimodal Context Interaction FastSpeech 2: Fast and High-Quality End-to-End Text to Speech

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T04:34:07.851676Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:34:07.851676Z digest=sha256:84dff0642b239d2db7efc7f6db586e44ad2f680a9087ea149200b50bf27d1e15

Observation bf5a25f9-586a-43d5-b5cb-77c6bbd199c2 · outbound

This paper cites wav2vec 2.0: A framework for self-supervised learning of speech representations,.

Towards Expressive Video Dubbing with Multiscale Multimodal Context Interaction wav2vec 2.0: A framework for self-supervised learning of speech representations,

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T04:34:07.856108Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:34:07.856108Z digest=sha256:00425ebfa6f974c1ca5880636c2b846d8da9df98f77407880eb52e8de6887750

Observation 0f6185fe-dba4-47dd-b786-bceae8be4482 · outbound

This paper cites Iemocap: Interactive emotional dyadic motion capture database,.

Towards Expressive Video Dubbing with Multiscale Multimodal Context Interaction Iemocap: Interactive emotional dyadic motion capture database,

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T04:34:07.860218Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:34:07.860218Z digest=sha256:503e45096783f2c949836e6d0cea2413219df393b2e5f97bf2d66918d16c2211

Observation 8f05115e-8c98-41a0-8900-5ef51a094696 · outbound

This paper cites Emotion-Aware Speech Self-Supervised Representation Learning with Intensity Knowledge.

Towards Expressive Video Dubbing with Multiscale Multimodal Context Interaction Emotion-Aware Speech Self-Supervised Representation Learning with Intensity Knowledge

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-08-11T04:34:08.031588Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T04:34:07.863878Z digest=sha256:051f6660474ee0ca11b33125378e9c8755b0f4c4fa25e5853274838f48ffa726

Observation 98ea45bd-36ee-4d5e-83fa-c11bcb31e692 · outbound

This paper cites Graph Attention Networks.

Towards Expressive Video Dubbing with Multiscale Multimodal Context Interaction Graph Attention Networks

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T04:34:07.868501Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:34:07.868501Z digest=sha256:276bdae8588e12838aef12cd5b3da2c7096ab2a0266f0ecedcc87f36d0dfc913

Observation e8d6a142-b366-41a0-9501-e5f13c3e9403 · outbound

This paper cites Hifi-gan: Generative adversarial networks for efficient and high fidelity speech synthesis,.

Towards Expressive Video Dubbing with Multiscale Multimodal Context Interaction Hifi-gan: Generative adversarial networks for efficient and high fidelity speech synthesis,

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T04:34:07.872926Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:34:07.872926Z digest=sha256:0946adf287ff75444a1b6e1d60cd9c4dddffb0648cab7a1d96c260bb08e8e219

Observation 2f32eebf-0417-427d-917a-9b4282893bbb · outbound

This paper cites A method for fundamental frequency estimation and voicing decision: Application to infant utterances recorded in real acoustical environments,.

Towards Expressive Video Dubbing with Multiscale Multimodal Context Interaction A method for fundamental frequency estimation and voicing decision: Application to infant utterances recorded in real acoustical environments,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:34:08.241427Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T04:34:07.877073Z digest=sha256:a27ac1ac230634c56e7b561adf69331430f32905475d59571bfc3c187478724c

Observation bd2a6eaf-eeac-401a-9bc7-d319b635a0c3 · outbound

This paper cites Reducing f0 frame error of f0 tracking algorithms under noisy conditions with an unvoiced/voiced classification frontend,.

Towards Expressive Video Dubbing with Multiscale Multimodal Context Interaction Reducing f0 frame error of f0 tracking algorithms under noisy conditions with an unvoiced/voiced classification frontend,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:34:08.226734Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T04:34:07.881464Z digest=sha256:d9212cbe28e20ac5e5a5bc932709fc9103d51b86a0c42f93bedf6924a52e6b89

Observation 6543205f-33e0-49c2-a923-6013d7cf04da · outbound

This paper cites Out of time: automated lip sync in the wild,.

Towards Expressive Video Dubbing with Multiscale Multimodal Context Interaction Out of time: automated lip sync in the wild,

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-11T04:34:07.885819Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:34:07.885819Z digest=sha256:df94f0e5d093d09c1e86e53e7049647671665d5aa60efe1a0517bfb840b1d81d

Observation 05c1dd58-9811-46ec-964f-baf9744ac89a · outbound

This paper cites A lip sync expert is all you need for speech to lip generation in the wild,.

Towards Expressive Video Dubbing with Multiscale Multimodal Context Interaction A lip sync expert is all you need for speech to lip generation in the wild,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:34:08.203889Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T04:34:07.890030Z digest=sha256:32a2c823b7189c0f4a366d7cf0be931d1410f8109f19e4f7260d8a99d992ade6

Observation 94fe868d-8b6d-4530-a256-6b0d5e4926d7 · outbound

This paper cites Robust speech recognition via large-scale weak supervi- sion,.

Towards Expressive Video Dubbing with Multiscale Multimodal Context Interaction Robust speech recognition via large-scale weak supervi- sion,

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-11T04:34:07.894310Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:34:07.894310Z digest=sha256:3fbe0a32fca96dbc113a0f57ed94033ff0b9c637d11f40bb5690ab55d6185c26

Observation d49774c0-f3e4-4156-ad62-43b9ea9708e5 · outbound

This paper cites Text-to-speech for low-resource agglutinative language with morphology-aware language model pre-training,.

Towards Expressive Video Dubbing with Multiscale Multimodal Context Interaction Text-to-speech for low-resource agglutinative language with morphology-aware language model pre-training,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:34:08.181646Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T04:34:07.898590Z digest=sha256:f2ce798c1fdda1d7d514d7ff4b92ce631160dfcf65eebe36b341785a0b5bfaea

Observation 0deab94c-ff93-4883-8c05-0eb527827fc1 · outbound

This paper cites Multi-Source Spatial Knowledge Understanding for Immersive Visual Text-to-Speech.

Towards Expressive Video Dubbing with Multiscale Multimodal Context Interaction Multi-Source Spatial Knowledge Understanding for Immersive Visual Text-to-Speech

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-08-11T04:34:07.997986Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T04:34:07.902789Z digest=sha256:717399aeb33be7e983c0ba7a38b9ca7d6774db03896d36886e06b7894d8d8364

Observation 6ece83f0-cdae-48d8-b707-a9faaa591b34 · outbound

This paper cites Multi-modal and Multi-scale Spatial Environment Understanding for Immersive Visual Text-to-Speech.

Towards Expressive Video Dubbing with Multiscale Multimodal Context Interaction Multi-modal and Multi-scale Spatial Environment Understanding for Immersive Visual Text-to-Speech

Reference 34

Resolution
verified exact
local_arxiv, observed 2026-08-11T04:34:07.977433Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T04:34:07.907336Z digest=sha256:b6b685c04b8209f128c8cff9e590d88fb37cdc3068d8b9917505674829dd2a34

Observation 6be81104-14f3-459f-ab46-f3adbaabcb2e · outbound

This paper cites Emphasis Rendering for Conversational Text-to-Speech with Multi-modal Multi-scale Context Modeling.

Towards Expressive Video Dubbing with Multiscale Multimodal Context Interaction Emphasis Rendering for Conversational Text-to-Speech with Multi-modal Multi-scale Context Modeling

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-11T04:34:07.911899Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:34:07.911899Z digest=sha256:6ba48641c7300f6fbb49383da578372d1e2c55b4171cc41640b675148ce1fc50

Observation e5ba2496-6766-4eea-8122-51f259e4ce91 · outbound

This paper cites Generative expressive conversational speech synthesis,.

Towards Expressive Video Dubbing with Multiscale Multimodal Context Interaction Generative expressive conversational speech synthesis,

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-11T04:34:07.916601Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:34:07.916601Z digest=sha256:3527b0c24f1f5829be435c0b84c01afed0ca4d5a6f4f9287296a9d121cedbb0a

Observation ee44814d-3815-4c23-8edb-3d0aa133750c · outbound

This paper cites Fctalker: Fine and coarse grained context modeling for expressive conversational speech synthesis,.

Towards Expressive Video Dubbing with Multiscale Multimodal Context Interaction Fctalker: Fine and coarse grained context modeling for expressive conversational speech synthesis,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:34:08.159255Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T04:34:07.920960Z digest=sha256:87803eea2043bfe36c828ea03819887c0064f0fc5d0fe39f81d651cad2f22592

Pith citing papers

No inbound Pith citation observations are available.