Pith. sign in

Paper Citation Record · LEDGER

Towards Expressive Video Dubbing with Multiscale Multimodal Context Interaction

As of 13 August 2026, this Paper Citation Record lists 37 of 37 outbound references and 0 inbound Pith citation observations for arXiv:2412.18748.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.18748 v2

Coverage vector

measured 37 of 37 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T04:34:07.920960Z

measured 37 of 37 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

37 of 37 outbound references displayed

  • verified exact5
  • verified fuzzy14
  • unresolved18
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 1764f373-12b8-4dc3-9752-73911e284520 · outbound

This paper cites Neural dubber: Dubbing for videos according to scripts,.

Towards Expressive Video Dubbing with Multiscale Multimodal Context Interaction Neural dubber: Dubbing for videos according to scripts,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:34:08.436892Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T04:34:07.758539Z digest=sha256:9b96b66d0134d483b13942d406d565bf7a31d7b6783e0a04570ad4d9daec3b97

Observation ec54db63-1103-45d6-a4df-f48b18939241 · outbound

This paper cites Prosody Modeling with 3D Visual Information for Expressive Video Dubbing,.

Towards Expressive Video Dubbing with Multiscale Multimodal Context Interaction Prosody Modeling with 3D Visual Information for Expressive Video Dubbing,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:34:08.423769Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T04:34:07.763464Z digest=sha256:efc302a324cc93034d30fa6d6382b50e07a9440d21e088456d41c06ad23c0c95

Observation 532688cd-8e25-4821-9be8-8dccc87e6896 · outbound

This paper cites StyleDubber: Towards Multi-Scale Style Learning for Movie Dubbing.

Towards Expressive Video Dubbing with Multiscale Multimodal Context Interaction StyleDubber: Towards Multi-Scale Style Learning for Movie Dubbing

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T04:34:07.768994Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:34:07.768994Z digest=sha256:857f038e9b94f6ecdfaca20ae37874c61d7ea2020e4ec0f9a7cb6a139d729763

Observation d8e07d3b-dd93-47d1-a79c-27271ebc6bc6 · outbound

This paper cites V2c: Visual voice cloning,.

Towards Expressive Video Dubbing with Multiscale Multimodal Context Interaction V2c: Visual voice cloning,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:34:08.410226Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T04:34:07.774060Z digest=sha256:838292676240fdb0fe91ade5c961e1cffd3e0c73463343447a253512247341b8

Observation 6d8cef30-e21c-4608-acba-195e24dab3aa · outbound

This paper cites From speaker to dubber: Movie dubbing with prosody and duration consistency learning,.

Towards Expressive Video Dubbing with Multiscale Multimodal Context Interaction From speaker to dubber: Movie dubbing with prosody and duration consistency learning,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:34:08.397469Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T04:34:07.778636Z digest=sha256:5a2b3e6c3b40716d8f79c6582ecde9945f14668e0ae0f3748fb0ed2db4521a06

Observation 3d42e29e-c8bc-407d-8f4b-c075f902d825 · outbound

This paper cites EmoDubber: Towards High Quality and Emotion Controllable Movie Dubbing.

Towards Expressive Video Dubbing with Multiscale Multimodal Context Interaction EmoDubber: Towards High Quality and Emotion Controllable Movie Dubbing

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T04:34:07.783191Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:34:07.783191Z digest=sha256:a481f2d1450a23dd5129aa8858ec2c19d687daa87c9f2412e43eccfe3f0b780c

Observation 076a3700-5358-47d4-b854-4be31dabfaa0 · outbound

This paper cites High-Quality Automatic Voice Over with Accurate Alignment: Supervision through Self-Supervised Discrete Speech Units.

Towards Expressive Video Dubbing with Multiscale Multimodal Context Interaction High-Quality Automatic Voice Over with Accurate Alignment: Supervision through Self-Supervised Discrete Speech Units

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-08-11T04:34:08.115409Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T04:34:07.788553Z digest=sha256:a0cf1038c3f1e525cf229f9a08524bf66c0f8af163cf516c8acce6e8e6c760ee

Observation a9d75b0b-ad46-4407-adea-1b9b0a18b657 · outbound

This paper cites DubWise: Video-Guided Speech Duration Control in Multimodal LLM-based Text-to-Speech for Dubbing.

Towards Expressive Video Dubbing with Multiscale Multimodal Context Interaction DubWise: Video-Guided Speech Duration Control in Multimodal LLM-based Text-to-Speech for Dubbing

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T04:34:07.793247Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:34:07.793247Z digest=sha256:d671c051b1c9a9beca94ee8b9a7f4bb277419a6ea6148376b89abd429850a864

Observation 10c9ed64-3d2a-46ff-baba-9e02046012bd · outbound

This paper cites More than words: In-the-wild visually-driven prosody for text-to-speech,.

Towards Expressive Video Dubbing with Multiscale Multimodal Context Interaction More than words: In-the-wild visually-driven prosody for text-to-speech,

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T04:34:07.797858Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:34:07.797858Z digest=sha256:c8d7d189df1b143e60f72bcba73bcdc88872371464f04301ba8d8c4ba1a6dbca

Observation ec185281-1a87-4f95-95d9-0d25343fd108 · outbound

This paper cites Learning to dub movies via hierarchical prosody models,.

Towards Expressive Video Dubbing with Multiscale Multimodal Context Interaction Learning to dub movies via hierarchical prosody models,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:34:08.374498Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T04:34:07.803591Z digest=sha256:f492560443ad9c3987c09e7d270ebf88d1ec3f7310fe26f133cfa9b40c365bab

Observation eb396945-041d-4a16-9cc0-c33c93495eeb · outbound

This paper cites MCDubber: Multimodal Context-Aware Expressive Video Dubbing.

Towards Expressive Video Dubbing with Multiscale Multimodal Context Interaction MCDubber: Multimodal Context-Aware Expressive Video Dubbing

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T04:34:07.808156Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:34:07.808156Z digest=sha256:4827ca0df053ad3df1fc797dee25c90c8d6b937092e50a0fcd618514da2beea4

Observation 5b6c150b-4206-40f6-9818-d1ee99cd7d13 · outbound

This paper cites To- wards Multi-Scale Speaking Style Modelling with Hierarchical Context Information for Mandarin Speech Synthesis,.

Towards Expressive Video Dubbing with Multiscale Multimodal Context Interaction To- wards Multi-Scale Speaking Style Modelling with Hierarchical Context Information for Mandarin Speech Synthesis,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:34:08.359596Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T04:34:07.813123Z digest=sha256:d73517f86e5bf425b2968e3e6d7164c58d4426145026089cb1fb7baad3f38b33

Observation 0d71bcdc-de81-40a7-9b32-2f5b9f125c49 · outbound

This paper cites Msstyletts: Multi-scale style modeling with hierarchical context infor- mation for expressive speech synthesis,.

Towards Expressive Video Dubbing with Multiscale Multimodal Context Interaction Msstyletts: Multi-scale style modeling with hierarchical context infor- mation for expressive speech synthesis,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:34:08.345484Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T04:34:07.817279Z digest=sha256:563be89f0d1f3e8591ef1b69b3efe2fe2c28bafba9445cffbc3d8ffba90b1a4f

Observation 688ab624-d08f-4aa7-9a80-136595686d56 · outbound

This paper cites Unsupervised multi-scale expressive speaking style modeling with hierarchical context information for audiobook speech synthesis,.

Towards Expressive Video Dubbing with Multiscale Multimodal Context Interaction Unsupervised multi-scale expressive speaking style modeling with hierarchical context information for audiobook speech synthesis,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:34:08.331259Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T04:34:07.821403Z digest=sha256:50e78852e33c1eeae2d8c3d481dd435699a6824a92abadad75a1d807ff828f67

Observation edfbc176-425b-45b1-8274-b302725aed01 · outbound

This paper cites Mae-dfer: Efficient masked au- toencoder for self-supervised dynamic facial expression recognition,.

Towards Expressive Video Dubbing with Multiscale Multimodal Context Interaction Mae-dfer: Efficient masked au- toencoder for self-supervised dynamic facial expression recognition,

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T04:34:07.825600Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:34:07.825600Z digest=sha256:e9fb3be4b6badd0c9e53fce727ac570369eb04300b93984325f1ee38ea8f8e6d

Observation 37cda586-b2a8-43b7-a36e-3a46b127c0eb · outbound

This paper cites Dfew: A large-scale database for recognizing dynamic facial expressions in the wild,.

Towards Expressive Video Dubbing with Multiscale Multimodal Context Interaction Dfew: A large-scale database for recognizing dynamic facial expressions in the wild,

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T04:34:07.830150Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:34:07.830150Z digest=sha256:663e246c0cd79c1c99e019eef49add1450057da2317f0ac409e22244cabf425e

Observation 0a5f69df-dd96-4e6b-99c6-b53bed2baedc · outbound

This paper cites Estimation of continuous valence and arousal levels from faces in naturalistic conditions,.

Towards Expressive Video Dubbing with Multiscale Multimodal Context Interaction Estimation of continuous valence and arousal levels from faces in naturalistic conditions,

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T04:34:07.834265Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:34:07.834265Z digest=sha256:20598b93970869da76412a910d835320279b98c80fc023728f833331b8278f77

Observation ce456576-3b5c-4c7f-83c9-3b088757a057 · outbound

This paper cites Towards Multi-Scale Style Control for Expressive Speech Synthesis.

Towards Expressive Video Dubbing with Multiscale Multimodal Context Interaction Towards Multi-Scale Style Control for Expressive Speech Synthesis

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-08-11T04:34:08.066385Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T04:34:07.838476Z digest=sha256:b2ac9941eec9712f14bb7675726d7212ab242b1bf036596a6cc1c386cb42e4aa

Observation f6da0022-bbc2-459a-b689-5e853cf2a636 · outbound

This paper cites Fastspeech: Fast, robust and controllable text to speech,.

Towards Expressive Video Dubbing with Multiscale Multimodal Context Interaction Fastspeech: Fast, robust and controllable text to speech,

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T04:34:07.843110Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:34:07.843110Z digest=sha256:65af9c46e5a759f016f8c2960bb3b24f496f3f2a153b7dd476ba884f42efff4e

Observation 60da494e-3e91-4d06-9a2b-37f6263c0626 · outbound

This paper cites Emotion-english-roberta-large,.

Towards Expressive Video Dubbing with Multiscale Multimodal Context Interaction Emotion-english-roberta-large,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:34:08.281979Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T04:34:07.847568Z digest=sha256:481ddc9d7aceb80d8fcf2180cd301a67fa5972a3ed9349d0569c23520668263e

Observation 8bd7cb49-60bb-4b7c-af54-203db6ed751b · outbound

This paper cites FastSpeech 2: Fast and High-Quality End-to-End Text to Speech.

Towards Expressive Video Dubbing with Multiscale Multimodal Context Interaction FastSpeech 2: Fast and High-Quality End-to-End Text to Speech

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T04:34:07.851676Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:34:07.851676Z digest=sha256:95402ee84640866c4636b4207c16dc27e6c7389021f17a2f0e4287829a1b1203

Observation bf5a25f9-586a-43d5-b5cb-77c6bbd199c2 · outbound

This paper cites wav2vec 2.0: A framework for self-supervised learning of speech representations,.

Towards Expressive Video Dubbing with Multiscale Multimodal Context Interaction wav2vec 2.0: A framework for self-supervised learning of speech representations,

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T04:34:07.856108Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:34:07.856108Z digest=sha256:bbaba712e3c017c773cc6457f5cd16a8d72699a73c6ecc54e7701f41b45f0375

Observation 0f6185fe-dba4-47dd-b786-bceae8be4482 · outbound

This paper cites Iemocap: Interactive emotional dyadic motion capture database,.

Towards Expressive Video Dubbing with Multiscale Multimodal Context Interaction Iemocap: Interactive emotional dyadic motion capture database,

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T04:34:07.860218Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:34:07.860218Z digest=sha256:076fd9ae8f0056a4a85b02c68887b0e6550ae2d9dbe843db38097a01387dadc1

Observation 8f05115e-8c98-41a0-8900-5ef51a094696 · outbound

This paper cites Emotion-Aware Speech Self-Supervised Representation Learning with Intensity Knowledge.

Towards Expressive Video Dubbing with Multiscale Multimodal Context Interaction Emotion-Aware Speech Self-Supervised Representation Learning with Intensity Knowledge

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-08-11T04:34:08.031588Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T04:34:07.863878Z digest=sha256:0c25f9e0e782c2e34c1823a10731a53960630075f6ec525f9564d3a8febfb9ae

Observation 98ea45bd-36ee-4d5e-83fa-c11bcb31e692 · outbound

This paper cites Graph Attention Networks.

Towards Expressive Video Dubbing with Multiscale Multimodal Context Interaction Graph Attention Networks

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T04:34:07.868501Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:34:07.868501Z digest=sha256:7be409ef3f1708df1ed5b9467ddaf997dc10be471e03bcb34c591d4f226d8f59

Observation e8d6a142-b366-41a0-9501-e5f13c3e9403 · outbound

This paper cites Hifi-gan: Generative adversarial networks for efficient and high fidelity speech synthesis,.

Towards Expressive Video Dubbing with Multiscale Multimodal Context Interaction Hifi-gan: Generative adversarial networks for efficient and high fidelity speech synthesis,

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T04:34:07.872926Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:34:07.872926Z digest=sha256:ea710b5c454887da00c39f11c71ce260abf9577bd060eee45222fffefe9a1339

Observation 2f32eebf-0417-427d-917a-9b4282893bbb · outbound

This paper cites A method for fundamental frequency estimation and voicing decision: Application to infant utterances recorded in real acoustical environments,.

Towards Expressive Video Dubbing with Multiscale Multimodal Context Interaction A method for fundamental frequency estimation and voicing decision: Application to infant utterances recorded in real acoustical environments,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:34:08.241427Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T04:34:07.877073Z digest=sha256:67a8ca29d4cde56049dc120e34e45229466d890cc1d121f89a71bf0f41370e37

Observation bd2a6eaf-eeac-401a-9bc7-d319b635a0c3 · outbound

This paper cites Reducing f0 frame error of f0 tracking algorithms under noisy conditions with an unvoiced/voiced classification frontend,.

Towards Expressive Video Dubbing with Multiscale Multimodal Context Interaction Reducing f0 frame error of f0 tracking algorithms under noisy conditions with an unvoiced/voiced classification frontend,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:34:08.226734Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T04:34:07.881464Z digest=sha256:aa203ea9cb53cb373bc29fe9d68d054bb4feff2ebf2713d7f1f5badaeebe6d25

Observation 6543205f-33e0-49c2-a923-6013d7cf04da · outbound

This paper cites Out of time: automated lip sync in the wild,.

Towards Expressive Video Dubbing with Multiscale Multimodal Context Interaction Out of time: automated lip sync in the wild,

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-11T04:34:07.885819Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:34:07.885819Z digest=sha256:09e0e05bafcce5d5e1182a499ba1be03c2696ebd4cf5c5ffee569e0ab0c26cb3

Observation 05c1dd58-9811-46ec-964f-baf9744ac89a · outbound

This paper cites A lip sync expert is all you need for speech to lip generation in the wild,.

Towards Expressive Video Dubbing with Multiscale Multimodal Context Interaction A lip sync expert is all you need for speech to lip generation in the wild,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:34:08.203889Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T04:34:07.890030Z digest=sha256:a1d276975152819f21864e5c1b136a2c6f1a84a3ac3c4f6c4c5c3af3dd73d6be

Observation 94fe868d-8b6d-4530-a256-6b0d5e4926d7 · outbound

This paper cites Robust speech recognition via large-scale weak supervi- sion,.

Towards Expressive Video Dubbing with Multiscale Multimodal Context Interaction Robust speech recognition via large-scale weak supervi- sion,

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-11T04:34:07.894310Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:34:07.894310Z digest=sha256:b95247e583290b49fcff4d9f411e16d7a80075ef386705fc2e7c759c86bbf9e8

Observation d49774c0-f3e4-4156-ad62-43b9ea9708e5 · outbound

This paper cites Text-to-speech for low-resource agglutinative language with morphology-aware language model pre-training,.

Towards Expressive Video Dubbing with Multiscale Multimodal Context Interaction Text-to-speech for low-resource agglutinative language with morphology-aware language model pre-training,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:34:08.181646Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T04:34:07.898590Z digest=sha256:380cf86ac47ac4674dc9c4dae3cab2d5ad3a832ebad7b6354fa386c7d5a1ef1e

Observation 0deab94c-ff93-4883-8c05-0eb527827fc1 · outbound

This paper cites Multi-Source Spatial Knowledge Understanding for Immersive Visual Text-to-Speech.

Towards Expressive Video Dubbing with Multiscale Multimodal Context Interaction Multi-Source Spatial Knowledge Understanding for Immersive Visual Text-to-Speech

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-08-11T04:34:07.997986Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T04:34:07.902789Z digest=sha256:3a89cb56a3cfcb750cfa36ecd537034fb2b3e670b790f8d7a3b7da1157cc4e94

Observation 6ece83f0-cdae-48d8-b707-a9faaa591b34 · outbound

This paper cites Multi-modal and Multi-scale Spatial Environment Understanding for Immersive Visual Text-to-Speech.

Towards Expressive Video Dubbing with Multiscale Multimodal Context Interaction Multi-modal and Multi-scale Spatial Environment Understanding for Immersive Visual Text-to-Speech

Reference 34

Resolution
verified exact
local_arxiv, observed 2026-08-11T04:34:07.977433Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T04:34:07.907336Z digest=sha256:3ab9cbd75667f3967898f9b644f16a77a4fd22db83b761707dcf0f2662f169a2

Observation 6be81104-14f3-459f-ab46-f3adbaabcb2e · outbound

This paper cites Emphasis Rendering for Conversational Text-to-Speech with Multi-modal Multi-scale Context Modeling.

Towards Expressive Video Dubbing with Multiscale Multimodal Context Interaction Emphasis Rendering for Conversational Text-to-Speech with Multi-modal Multi-scale Context Modeling

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-11T04:34:07.911899Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:34:07.911899Z digest=sha256:8a5aa039bba06eee8130bd6eaed49ddd59a767c8a71392ae954f2b2c5c74898c

Observation e5ba2496-6766-4eea-8122-51f259e4ce91 · outbound

This paper cites Generative expressive conversational speech synthesis,.

Towards Expressive Video Dubbing with Multiscale Multimodal Context Interaction Generative expressive conversational speech synthesis,

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-11T04:34:07.916601Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:34:07.916601Z digest=sha256:f2c1f95b3d330bce911b2c31550e043d6ade0ea0e78160d1ff2cb6a10cb19b09

Observation ee44814d-3815-4c23-8edb-3d0aa133750c · outbound

This paper cites Fctalker: Fine and coarse grained context modeling for expressive conversational speech synthesis,.

Towards Expressive Video Dubbing with Multiscale Multimodal Context Interaction Fctalker: Fine and coarse grained context modeling for expressive conversational speech synthesis,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:34:08.159255Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T04:34:07.920960Z digest=sha256:0e74f88f1d13c1b1d0154a3a844581c9600ca151e0d345beb457153ad28d6861

Pith citing papers

No inbound Pith citation observations are available.