Pith. sign in

Paper Citation Record · LEDGER

DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis

As of 9 August 2026, this Paper Citation Record lists 62 of 62 outbound references and 1 inbound Pith citation observation for arXiv:2507.10109.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.10109 v1

Coverage vector

measured 62 of 62 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T17:46:19.243232Z

measured 63 of 63 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T06:04:29.936928Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-06T06:04:30.077501Z

Reference resolution

62 of 62 outbound references displayed

  • verified exact2
  • verified fuzzy7
  • unresolved53
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 37ec9f19-0267-400f-8dc4-186cffcba3fa · outbound

This paper cites an unresolved cited work.

DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis Unresolved cited work

Reference 1

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:46:26.620782Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T17:46:12.972404Z digest=sha256:f4b451e78f184e77139016082abccf432e16b19215c75de6e1fcf427f3f90892

Observation 33e3d267-b0bd-4db1-9e4a-d1cc222ff02c · outbound

This paper cites an unresolved cited work.

DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:46:26.409135Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T17:46:13.063080Z digest=sha256:b77cddaa71ed5f03d3739d5f89db454ba7eb390a9e59e28a6a4ed1c30cfdc66c

Observation 279008d1-a795-4170-97b3-afdac7bdc411 · outbound

This paper cites an unresolved cited work.

DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:46:26.181674Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T17:46:13.181371Z digest=sha256:0c54e4fd5bce7ec113fa0fe2205abd9a192c4f61ceafe7ce2075c72ac969fedb

Observation e6801440-a98a-4f92-acfd-75a1da97f3bc · outbound

This paper cites BEATs: Audio Pre-Training with Acoustic Tokenizers.

DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis BEATs: Audio Pre-Training with Acoustic Tokenizers

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T17:46:13.348025Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:46:13.348025Z digest=sha256:a21bbdbee37c457c352cabc3d044036c326d67e9043fbe22ab493c54a98de960

Observation 20b2da09-6e34-4360-a3a7-76032610e109 · outbound

This paper cites F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching.

DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T17:46:13.509047Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:46:13.509047Z digest=sha256:781bf583f18c0aabead8cbdaa4b9dea426d28dcd84813e0b9136c48ac398205e

Observation 5fb1a958-27a4-49d4-a92e-1e2e09f2bb59 · outbound

This paper cites an unresolved cited work.

DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:46:26.008575Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T17:46:13.584990Z digest=sha256:7cb87e6955cd7d63a2d049c7aae388aaaf569469606c0b640925aafd20efade2

Observation 713fd661-c934-40de-b955-c4023f195efc · outbound

This paper cites MMAudio: Taming Multimodal Joint Training for High-Quality Video-to-Audio Synthesis.

DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis MMAudio: Taming Multimodal Joint Training for High-Quality Video-to-Audio Synthesis

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T17:46:13.686519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:46:13.686519Z digest=sha256:2410e86029db851a5dcdae6aa6b44a55f1dc4f50ac579dcb74f280663355d548

Observation 96d79e06-cee1-4956-9569-cf243ea2cd07 · outbound

This paper cites an unresolved cited work.

DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:46:25.846602Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T17:46:13.792235Z digest=sha256:12b16312f64fc0f1bf2ca3a517e6a07d5f84a5e78c4262ef4f7014ea3ff2608e

Observation af46798e-68a2-4bfc-a683-538e14f5b39d · outbound

This paper cites V2SFlow: Video-to-Speech Generation with Speech Decomposition and Rectified Flow.

DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis V2SFlow: Video-to-Speech Generation with Speech Decomposition and Rectified Flow

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-08-06T17:46:19.957973Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T17:46:13.885497Z digest=sha256:f66a4a36baef89f430db12718c220e78382ebc67264913ad973e7b20dcd39739

Observation 3539a8e8-b6b9-422e-b7e8-535a6d171926 · outbound

This paper cites an unresolved cited work.

DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:46:25.693597Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T17:46:13.957998Z digest=sha256:5da93ac2139efa914db1a1e3441060203fb57773d90391b70531bf34793e1ecf

Observation bd1e16fc-e7a3-459a-8aa2-8b710ccffd68 · outbound

This paper cites an unresolved cited work.

DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:46:25.478203Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T17:46:14.028538Z digest=sha256:943124b50d8ca6235127cf912d6273741ff1ca1ee96e703a71e593666c1d9105

Observation 05516ffb-ffcb-411c-b448-4b524f94e2d8 · outbound

This paper cites Russell, and Andrew Owens.

DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis Russell, and Andrew Owens

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:46:25.270345Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T17:46:14.097517Z digest=sha256:4430ef9b448e0e01cebabac3db544ffce5f34ffa014e2a9e1f13022208dbc205

Observation 7500dc1f-d9f1-4e22-816f-8eff8d1f2d58 · outbound

This paper cites CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models.

DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T17:46:14.237637Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:46:14.237637Z digest=sha256:c38da99dacc16f2adda7197e07944ff1bf641434e6dc7973bfc0bfcdf98d117e

Observation f60c6bcb-e0e9-4e14-a6af-32311c28ccdc · outbound

This paper cites an unresolved cited work.

DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis Unresolved cited work

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T17:46:14.325703Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:46:14.325703Z digest=sha256:5d82435564b5f02ff7670810db75883382fe4ee4cf31f61e546469ff22251077

Observation da33c1e8-ba39-4f81-b34e-d7a3e86c388c · outbound

This paper cites an unresolved cited work.

DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis Unresolved cited work

Reference 15

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:46:24.620552Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T17:46:14.470174Z digest=sha256:a460b870636a2e87fe0c38825641801ed4ad890354a2f6d94dea9648f21f5c71

Observation 93b9eb47-acfb-442e-98db-12b9101fda90 · outbound

This paper cites In IEEE International Conference on Acoustics, Speech and Signal Processing ICASSP 2023, Rhodes Island, Greece, June 4-10, 2023.

DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis In IEEE International Conference on Acoustics, Speech and Signal Processing ICASSP 2023, Rhodes Island, Greece, June 4-10, 2023

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:46:24.828783Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T17:46:14.399881Z digest=sha256:93e5c0bd6ff992c5e7a2cbce2e0258391839218e014ea1944d2100ee72238b34

Observation 3e0bb405-fcd7-4792-a562-d92b225e889c · outbound

This paper cites MINT: a Multi-modal Image and Narrative Text Dubbing Dataset for Foley Audio Content Planning and Generation.

DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis MINT: a Multi-modal Image and Narrative Text Dubbing Dataset for Foley Audio Content Planning and Generation

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T17:46:14.616924Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:46:14.616924Z digest=sha256:cb38c5d440a0db2679177078b4c13d81039b7729353b72492207ee524b48c04a

Observation c03ef0c9-698a-47ef-b11b-dc059e450910 · outbound

This paper cites Hawley, and Jordi Pons.

DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis Hawley, and Jordi Pons

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:46:24.392742Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T17:46:14.542109Z digest=sha256:e8d3c150fa32e43b0be2674ef35bdfd3a509fa72e117eac52a58603167175056

Observation abba1354-4e34-4433-88c0-93456f3e50af · outbound

This paper cites Contrastive Audio-Visual Masked Autoencoder.

DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis Contrastive Audio-Visual Masked Autoencoder

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T17:46:14.781335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:46:14.781335Z digest=sha256:947d9d11a89fd3ac54eb2c6fa8682529f18bc967c16c59d256e827f025f67cc4

Observation b9b38237-e27d-4dac-af82-ec1e839fb9f0 · outbound

This paper cites an unresolved cited work.

DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:46:24.175716Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T17:46:14.711057Z digest=sha256:a072a799b902a0d2d1870242fc8e30234469aa6d287b9cb3c09d5d3740625b7d

Observation ff0220a8-2b66-47d3-abd6-7c0f316c06f6 · outbound

This paper cites an unresolved cited work.

DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis Unresolved cited work

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T17:46:14.936094Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:46:14.936094Z digest=sha256:de99fa1700bcaa1d1c96308cc3dfcf335a0a48466b57f80f408e4d3d9589aa4c

Observation 1eda6f2e-526f-4acb-b809-f56b05765310 · outbound

This paper cites an unresolved cited work.

DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis Unresolved cited work

Reference 22

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:46:23.960830Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T17:46:14.853327Z digest=sha256:d01352d641e96d597bab49d6b98b67b93b7ac90c33071c0dedc3edc188879272

Observation 1ee185fd-e353-4680-b406-11bc1b508eea · outbound

This paper cites an unresolved cited work.

DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis Unresolved cited work

Reference 23

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:46:23.412143Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T17:46:15.132879Z digest=sha256:eb32d2102117ba1fde5fb25ed86974c27f2a2028dccee982a1f3e7a3f7415939

Observation ce3f8763-5543-4aa8-a484-95c08a948cd2 · outbound

This paper cites an unresolved cited work.

DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:46:23.702173Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T17:46:15.031709Z digest=sha256:d2f37e94a1eedf2b5ec6700db92e5b0d3452a8c69270140d8acebefcd0aece9c

Observation a7780ee1-ce41-48bc-aa46-3f2993a269f1 · outbound

This paper cites Mega-TTS 2: Boosting Prompting Mechanisms for Zero-Shot Speech Synthesis.

DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis Mega-TTS 2: Boosting Prompting Mechanisms for Zero-Shot Speech Synthesis

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T17:46:15.334498Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:46:15.334498Z digest=sha256:901e4d4929d656177d2baab4b84c50d1597a3acf23335986e4064e4bd5ee2aa1

Observation 2292ea90-dbba-4505-9b79-ede17c98baab · outbound

This paper cites WavTokenizer: an Efficient Acoustic Discrete Codec Tokenizer for Audio Language Modeling.

DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis WavTokenizer: an Efficient Acoustic Discrete Codec Tokenizer for Audio Language Modeling

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T17:46:15.230611Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:46:15.230611Z digest=sha256:f383230a52cad3da03019f018956b3567124e1a7d6f5cdee49f14239ecce4c7d

Observation 03647407-1f64-4c07-840f-723e7258cd1f · outbound

This paper cites Fr\'echet Audio Distance: A Metric for Evaluating Music Enhancement Algorithms.

DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis Fr\'echet Audio Distance: A Metric for Evaluating Music Enhancement Algorithms

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T17:46:15.533022Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:46:15.533022Z digest=sha256:2db2b4269fa9eb3069dc469bff18813a3245dd4e8489a389e0ce1852f38bcb43

Observation 1183abe4-53c7-4251-8fd3-81917221eb3f · outbound

This paper cites an unresolved cited work.

DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis Unresolved cited work

Reference 28

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:46:23.177628Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T17:46:15.431630Z digest=sha256:3dd6a8fbfee396e63ff158fd3614d95027b94747d436417e592d901df994f40d

Observation cb3d7ec8-f2d5-4e92-acaf-e0b4150495b1 · outbound

This paper cites Salik, Rajiv Ratn Shah, Yifang Yin, and Roger Zimmermann.

DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis Salik, Rajiv Ratn Shah, Yifang Yin, and Roger Zimmermann

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:46:22.955376Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T17:46:15.732760Z digest=sha256:5483ce4b5ebe6f8d2706be7501ce8f369cd28f163d0c4a93d92adcf004cbbf5e

Observation cfe93e57-9358-4e6b-ae3c-1fd69c32cc39 · outbound

This paper cites an unresolved cited work.

DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis Unresolved cited work

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T17:46:15.631928Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:46:15.631928Z digest=sha256:a976b9d99b03ba6441f221d110a9620019875d781f6fcd1ad5af14677c36569e

Observation c1b4caf5-dafc-4d0a-bb99-b0f39eb64b0c · outbound

This paper cites an unresolved cited work.

DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis Unresolved cited work

Reference 31

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:46:22.591730Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T17:46:15.965671Z digest=sha256:fe28bb9935bbbe512290a891cb38569090077e8131ec869af79ebe460919e614

Observation 6f9c274d-f627-426c-829c-1842c25a0dfc · outbound

This paper cites an unresolved cited work.

DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis Unresolved cited work

Reference 32

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:46:22.788610Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T17:46:15.862042Z digest=sha256:9113ff291455d401093c947b1fe634c711aeed62738064f87caeeda1195368ab

Observation 3f69a0db-57dc-429a-8de0-b4d50dccc736 · outbound

This paper cites Mandic, Wenwu Wang, and Mark D.

DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis Mandic, Wenwu Wang, and Mark D

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:46:22.193356Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T17:46:16.165396Z digest=sha256:76d7de0e13f9f1e989a3b21fb74b74d4571f5dd41d6a9cc6efcd2ce18750d7bc

Observation 9cd59548-046e-4908-977a-3d14562cecff · outbound

This paper cites Raghavan, Gavin Mischler, and Nima Mesgarani.

DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis Raghavan, Gavin Mischler, and Nima Mesgarani

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:46:22.365402Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T17:46:16.059867Z digest=sha256:4bff47c4e20d5f37f3484e3d08b9b5c9b5ce000463f23e20522a12293b0ce66d

Observation e3194c8a-04cd-421e-adef-e2557f37f3b6 · outbound

This paper cites Autoregressive Speech Synthesis without Vector Quantization.

DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis Autoregressive Speech Synthesis without Vector Quantization

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T17:46:16.401709Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:46:16.401709Z digest=sha256:ec24b6d4b85caa62a5b3f37eb7347f6249fe3f6b2119fd7497fea9788f43a2f1

Observation 99e73ac7-28ec-4a05-8240-968939682a8b · outbound

This paper cites an unresolved cited work.

DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis Unresolved cited work

Reference 36

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:46:21.977075Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T17:46:16.281334Z digest=sha256:4b3dd8e8af2f15e19bd1584e1b1fa03b5f9ec95d0c2d879521ec077d38942c5b

Observation 3af9a247-6149-4362-b3ec-3b5024433dbf · outbound

This paper cites an unresolved cited work.

DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis Unresolved cited work

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T17:46:16.612710Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:46:16.612710Z digest=sha256:0599ebc28758ddc6214b2b743ef380d055e406e34b797b929dda91af7225fcf9

Observation 42533c02-709b-444c-ba82-f108d0b1d43e · outbound

This paper cites an unresolved cited work.

DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis Unresolved cited work

Reference 38

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:46:21.844207Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T17:46:16.513575Z digest=sha256:90950a6ef9dc65e6dc26f1e620a8ec0e47ea8ea475c5cf3056dc589a0fefe77d

Observation 0c3fb29b-10e6-4788-9da2-c122d56c57c6 · outbound

This paper cites an unresolved cited work.

DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis Unresolved cited work

Reference 39

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:46:21.699975Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T17:46:16.856077Z digest=sha256:abf352538161ef235fffb35250b5166c56dc61b3e78bd781cc04dc51ba3850d3

Observation bf519b0d-1df2-45f7-a72c-fb11017f577f · outbound

This paper cites UTMOS: UTokyo-SaruLab System for VoiceMOS Challenge 2022.

DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis UTMOS: UTokyo-SaruLab System for VoiceMOS Challenge 2022

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T17:46:16.729243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:46:16.729243Z digest=sha256:0336b518263644b7bfc0addcea4f806347332e2cf7cc458d41a81ffd72b59a5b

Observation 6d151fd1-b8cb-4da5-b6dd-dafea2042c69 · outbound

This paper cites an unresolved cited work.

DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis Unresolved cited work

Reference 41

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:46:21.376640Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T17:46:17.074451Z digest=sha256:563868a8f8ce3a200ba40f386be1b78b36b7428ff29fdf45e69c1e0cd00c0024

Observation d5aff840-67a9-4dbe-861c-bc0ed3ecf6f6 · outbound

This paper cites an unresolved cited work.

DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis Unresolved cited work

Reference 42

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:46:21.517553Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T17:46:16.972078Z digest=sha256:90578683521b89727fa7cde00795348ace56ddea44cb4e4d5701b228677bb8f9

Observation c778f8fa-b4e8-4720-a550-e1c19f05d7d0 · outbound

This paper cites FELLE: Autoregressive Speech Synthesis with Token-Wise Coarse-to-Fine Flow Matching.

DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis FELLE: Autoregressive Speech Synthesis with Token-Wise Coarse-to-Fine Flow Matching

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T17:46:17.295613Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:46:17.295613Z digest=sha256:3a2c82db318dd98803f560504d63a871751c0336416335a59455d133409fbe82

Observation 13942b3b-21ae-497b-8103-3869f8f17fe3 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis LLaMA: Open and Efficient Foundation Language Models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T17:46:17.194390Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:46:17.194390Z digest=sha256:02a591a357158f897611256917a7f74630fbf5e12d532d168dc8543e4ecc590f

Observation 5aadd354-21aa-434d-b3e4-e651d848f1c5 · outbound

This paper cites Mel-Band RoFormer for Music Source Separation.

DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis Mel-Band RoFormer for Music Source Separation

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T17:46:17.731914Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:46:17.731914Z digest=sha256:f27f3c1491179734c969a0d7197230f5bad2314d65bafd94b016a00f7a9b70de

Observation 80d1a0b5-080a-4f04-8e92-8ef224a181b4 · outbound

This paper cites an unresolved cited work.

DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis Unresolved cited work

Reference 46

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:46:21.242636Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T17:46:17.454006Z digest=sha256:2d71a24a03709f05645c9d0f8b6a16f7bb267ac6aa7285585e344900451acfff

Observation 25de76f6-9bd3-4faa-b791-113c7b55626b · outbound

This paper cites an unresolved cited work.

DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis Unresolved cited work

Reference 47

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:46:20.768252Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T17:46:18.070098Z digest=sha256:3ab3299caa566bb1ef2588ab64a46474c0debf70db22764ab0e5fd2f75020e68

Observation 3869cc5b-f98e-484f-8fbd-91cb5e62436a · outbound

This paper cites an unresolved cited work.

DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis Unresolved cited work

Reference 48

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:46:20.618868Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T17:46:18.195413Z digest=sha256:9c9b4be7d49834a6118b95a2bc97990c3c50180715fed97b9030572a1622415e

Observation 91a24ba9-2e29-43ef-827b-95867ffd5fa3 · outbound

This paper cites an unresolved cited work.

DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis Unresolved cited work

Reference 49

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:46:20.915025Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T17:46:17.855498Z digest=sha256:7a928dcdf94d35d5bc22758ec4e6e503f1b3aff224ea6d619ef4de663dddd7c8

Observation c6e79082-fc53-44c9-8f1e-6ec750618d11 · outbound

This paper cites UniAudio: An Audio Foundation Model Toward Universal Audio Generation.

DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis UniAudio: An Audio Foundation Model Toward Universal Audio Generation

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T17:46:18.500277Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:46:18.500277Z digest=sha256:48207959a5c486383219d6a1b8d609e6f971cd6f93f69fd35d3d490e2aa9f528

Observation 5232a114-2d96-40a1-b022-b42e42fff0c2 · outbound

This paper cites an unresolved cited work.

DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis Unresolved cited work

Reference 51

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:46:20.489871Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T17:46:18.587362Z digest=sha256:98dd98aa01322274e3363f1dbdb0a23dedcc710cb99ddbe6370a342bccd4be65

Observation b3fbe57a-f389-445e-a3c7-8911c34f56f5 · outbound

This paper cites Qwen2 Technical Report.

DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis Qwen2 Technical Report

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-06T17:46:18.341711Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:46:18.341711Z digest=sha256:e2bef48da1ad86bf065dd910f0f7f512d691cdeb5991a80541ef67477492b215

Observation 7eda949d-c7f0-4718-9603-63ce8d097b60 · outbound

This paper cites LibriTTS: A Corpus Derived from LibriSpeech for Text-to-Speech.

DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis LibriTTS: A Corpus Derived from LibriSpeech for Text-to-Speech

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-06T17:46:18.798620Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:46:18.798620Z digest=sha256:f5e17793a4d8e841cbec5d2921131e757040bbf984f51d3d6227d024d4aedfa9

Observation 66285bcf-7fed-49f0-8900-4f4d5a0d744f · outbound

This paper cites FoleyCrafter: Bring Silent Videos to Life with Lifelike and Synchronized Sounds.

DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis FoleyCrafter: Bring Silent Videos to Life with Lifelike and Synchronized Sounds

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-06T17:46:18.900460Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:46:18.900460Z digest=sha256:ec5dbec9beb052ffd3c09a83ea0192b2785e2802a3e5d3371ad177a2015ef9f2

Observation 7417fdab-75af-4ed8-918e-d769dee767a7 · outbound

This paper cites Diverse and Aligned Audio-to-Video Generation via Text-to-Video Model Adaptation.

DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis Diverse and Aligned Audio-to-Video Generation via Text-to-Video Model Adaptation

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-06T17:46:18.671928Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:46:18.671928Z digest=sha256:36fcf61c06ba769868cc631a45db9ec5426f21e660363e01006dad961a142f25

Observation f08d860c-4432-4423-a6b6-64609e10eb18 · outbound

This paper cites Llasa: Scaling Train-Time and Inference-Time Compute for Llama-based Speech Synthesis.

DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis Llasa: Scaling Train-Time and Inference-Time Compute for Llama-based Speech Synthesis

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-06T17:46:18.709302Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:46:18.709302Z digest=sha256:9a6ee6c9d32de14314a9d78adbac3d256c24a0309fe4cc468ff3e5933c5645ae

Observation e39c3107-8a78-4cd1-b18b-1e2b4f16793d · outbound

This paper cites CosyAudio: Improving Audio Generation with Confidence Scores and Synthetic Captions.

DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis CosyAudio: Improving Audio Generation with Confidence Scores and Synthetic Captions

Reference 57

Resolution
verified exact
local_arxiv, observed 2026-08-06T17:46:19.565364Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T17:46:19.155991Z digest=sha256:2d6bd474a68eed05d9223612c9d41209f69a70bea235cf7c148d93a35c4131a9

Observation 5b4e0a7c-e4c5-4cd3-bead-1dee93540884 · outbound

This paper cites an unresolved cited work.

DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis Unresolved cited work

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-06T17:46:19.243232Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:46:19.243232Z digest=sha256:6bd058ff8c8aed59a079542a96618cfd99a5fa344c5ce6ad3ba41e6e5b289985

Observation 952abe25-feca-40f1-a9b9-985be6455714 · outbound

This paper cites an unresolved cited work.

DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis Unresolved cited work

Reference 59

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:46:20.344490Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T17:46:18.985234Z digest=sha256:6f615cdcac9fed74895437a4ca49f2d24c17325b0d279add3e084f9a94930d9f

Observation 6bed8f7c-9191-4d2a-b1a7-59f1dd70c6e5 · outbound

This paper cites an unresolved cited work.

DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis Unresolved cited work

Reference 60

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:46:20.150751Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T17:46:19.066327Z digest=sha256:54dc3810ab3db7f1009c4eb308a35faee2bc0ed0664144db4d3e013fbacc76f9

Observation 946540d6-4f50-4a0a-bf43-337744e816ff · outbound

This paper cites In IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2023, Vancouver, BC, Canada, June 17-24, 2023.

DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis In IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2023, Vancouver, BC, Canada, June 17-24, 2023

Reference 2023

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:46:25.054650Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T17:46:14.166030Z digest=sha256:86e0f0d4708c0e3f6e58d856a44f46c49d9bfea2ddf85db96de4682d9a819626

Observation 6ecb8ab1-8022-41c7-bd52-9c10e0998921 · outbound

This paper cites an unresolved cited work.

DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis Unresolved cited work

Reference 2024

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:46:21.053627Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T17:46:17.612626Z digest=sha256:8c774cbeff58930885b976b592f114480b4a695828cd04204f8ba324509dbfa9

Pith citing papers

Observation 9f73405d-b2d7-4205-bea6-b06ea22e0311 · inbound

AudioGen-Omni: A Unified Multimodal Diffusion Transformer for Video-Synchronized Audio, Speech, and Song Generation cites this paper.

AudioGen-Omni: A Unified Multimodal Diffusion Transformer for Video-Synchronized Audio, Speech, and Song Generation DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-08-06T06:04:30.083418Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T06:04:29.936928Z digest=sha256:30e4e88b97d2c911b8a45bca9f6fed6ac697e23255f070529706ae6531f78355