Pith. sign in

Paper Citation Record · LEDGER

DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis

As of 15 August 2026, this Paper Citation Record lists 62 of 62 outbound references and 1 inbound Pith citation observation for arXiv:2507.10109.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.10109 v1

Coverage vector

measured 62 of 62 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T17:46:19.243232Z

measured 63 of 63 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T06:04:29.936928Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-06T06:04:30.077501Z

Reference resolution

62 of 62 outbound references displayed

  • verified exact2
  • verified fuzzy7
  • unresolved53
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 37ec9f19-0267-400f-8dc4-186cffcba3fa · outbound

This paper cites an unresolved cited work.

DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis Unresolved cited work

Reference 1

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:46:26.620782Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T17:46:12.972404Z digest=sha256:b82d3771a21e759680342e6dbf8686ad68e05fc84b72c4f32cedff4603b28825

Observation 33e3d267-b0bd-4db1-9e4a-d1cc222ff02c · outbound

This paper cites an unresolved cited work.

DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:46:26.409135Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T17:46:13.063080Z digest=sha256:4dac07c94e595b94b6742f7d36b839c4d9ce7d362497589306745dbca230cd6c

Observation 279008d1-a795-4170-97b3-afdac7bdc411 · outbound

This paper cites an unresolved cited work.

DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:46:26.181674Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T17:46:13.181371Z digest=sha256:b20c0feb511b4efcc3c757bce11e4e5cdfe1f2d72f68507e14de208d77c84c01

Observation e6801440-a98a-4f92-acfd-75a1da97f3bc · outbound

This paper cites BEATs: Audio Pre-Training with Acoustic Tokenizers.

DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis BEATs: Audio Pre-Training with Acoustic Tokenizers

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T17:46:13.348025Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:46:13.348025Z digest=sha256:9bee5f7df7488550303a66bde955c63e797d048088905d1123cb9bd080b01527

Observation 20b2da09-6e34-4360-a3a7-76032610e109 · outbound

This paper cites F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching.

DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T17:46:13.509047Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:46:13.509047Z digest=sha256:0e2263752196db99b38e868413565bbc3be8cc1bbc6119b0bd6771e3425a3c5d

Observation 5fb1a958-27a4-49d4-a92e-1e2e09f2bb59 · outbound

This paper cites an unresolved cited work.

DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:46:26.008575Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T17:46:13.584990Z digest=sha256:77ffb43807a335330eaa91f03d51dc19f367c16361417c11340600fbcdd04a16

Observation 713fd661-c934-40de-b955-c4023f195efc · outbound

This paper cites MMAudio: Taming Multimodal Joint Training for High-Quality Video-to-Audio Synthesis.

DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis MMAudio: Taming Multimodal Joint Training for High-Quality Video-to-Audio Synthesis

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T17:46:13.686519Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:46:13.686519Z digest=sha256:3a40bcf9bf257964dfb63838f6bdbe1e84df3a9590c6755c952e68c54e515357

Observation 96d79e06-cee1-4956-9569-cf243ea2cd07 · outbound

This paper cites an unresolved cited work.

DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:46:25.846602Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T17:46:13.792235Z digest=sha256:225120e394f4142c8ba6a756327597d1f6b475574624aaeee9daf17dc0771e2f

Observation af46798e-68a2-4bfc-a683-538e14f5b39d · outbound

This paper cites V2SFlow: Video-to-Speech Generation with Speech Decomposition and Rectified Flow.

DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis V2SFlow: Video-to-Speech Generation with Speech Decomposition and Rectified Flow

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-08-06T17:46:19.957973Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T17:46:13.885497Z digest=sha256:4369f255930b870aaa2942d5c8faec5e6571d523708c208f1afa3ad50dcfb356

Observation 3539a8e8-b6b9-422e-b7e8-535a6d171926 · outbound

This paper cites an unresolved cited work.

DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:46:25.693597Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T17:46:13.957998Z digest=sha256:97cee7e8c2c22662102c2a98576a35a11373f1874c29f6aa74bb4d17cd5fd083

Observation bd1e16fc-e7a3-459a-8aa2-8b710ccffd68 · outbound

This paper cites an unresolved cited work.

DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:46:25.478203Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T17:46:14.028538Z digest=sha256:41612af8ed23bcbddfebffbfd991d2dd5fcdcd5f073d846f839658e06bb3c81d

Observation 05516ffb-ffcb-411c-b448-4b524f94e2d8 · outbound

This paper cites Russell, and Andrew Owens.

DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis Russell, and Andrew Owens

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:46:25.270345Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T17:46:14.097517Z digest=sha256:98945495b46b3515d6feef58d80eadcf76c785f1b18881af384f6eef918190b1

Observation 7500dc1f-d9f1-4e22-816f-8eff8d1f2d58 · outbound

This paper cites CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models.

DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T17:46:14.237637Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:46:14.237637Z digest=sha256:fc21014b8a1009f73eb6483d8edc59f6e58ffda16a6cd21c320fa771d5dd9c51

Observation f60c6bcb-e0e9-4e14-a6af-32311c28ccdc · outbound

This paper cites an unresolved cited work.

DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis Unresolved cited work

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T17:46:14.325703Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:46:14.325703Z digest=sha256:b038f1bdf4ff2e2ef0f6883e39bdcce705dea8a5c6a40a028c33e694a9b179aa

Observation da33c1e8-ba39-4f81-b34e-d7a3e86c388c · outbound

This paper cites an unresolved cited work.

DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis Unresolved cited work

Reference 15

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:46:24.620552Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T17:46:14.470174Z digest=sha256:7ed73abc687800b9cf565195163b5649e68651fda7a2d803d04c439bf1ecdb26

Observation 93b9eb47-acfb-442e-98db-12b9101fda90 · outbound

This paper cites In IEEE International Conference on Acoustics, Speech and Signal Processing ICASSP 2023, Rhodes Island, Greece, June 4-10, 2023.

DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis In IEEE International Conference on Acoustics, Speech and Signal Processing ICASSP 2023, Rhodes Island, Greece, June 4-10, 2023

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:46:24.828783Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T17:46:14.399881Z digest=sha256:4df6430a443ace060b27133f81241eecf9d97aefe9d21379cc6238760606e8d4

Observation 3e0bb405-fcd7-4792-a562-d92b225e889c · outbound

This paper cites MINT: a Multi-modal Image and Narrative Text Dubbing Dataset for Foley Audio Content Planning and Generation.

DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis MINT: a Multi-modal Image and Narrative Text Dubbing Dataset for Foley Audio Content Planning and Generation

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T17:46:14.616924Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:46:14.616924Z digest=sha256:d346876612aa6f964855e6b6ff4ef7597b73e1858c712f180cfd01e450af8cdb

Observation c03ef0c9-698a-47ef-b11b-dc059e450910 · outbound

This paper cites Hawley, and Jordi Pons.

DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis Hawley, and Jordi Pons

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:46:24.392742Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T17:46:14.542109Z digest=sha256:bede0ec5668fd7e1b5b9c41a6e0e40c8561744cf7b94ccb6291ae7c342537dd9

Observation abba1354-4e34-4433-88c0-93456f3e50af · outbound

This paper cites Contrastive Audio-Visual Masked Autoencoder.

DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis Contrastive Audio-Visual Masked Autoencoder

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T17:46:14.781335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:46:14.781335Z digest=sha256:4399932107843827a084f0590269df5092f41b7ef019856c74d85711eeea7e0c

Observation b9b38237-e27d-4dac-af82-ec1e839fb9f0 · outbound

This paper cites an unresolved cited work.

DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:46:24.175716Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T17:46:14.711057Z digest=sha256:28f5567c1138adfcba8def19384f69851cbc1b53e38f727280814095978a16ae

Observation ff0220a8-2b66-47d3-abd6-7c0f316c06f6 · outbound

This paper cites an unresolved cited work.

DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis Unresolved cited work

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T17:46:14.936094Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:46:14.936094Z digest=sha256:09c03694284fbde821adc49a90b3050bb14a711441f0dae6be92597f055803f7

Observation 1eda6f2e-526f-4acb-b809-f56b05765310 · outbound

This paper cites an unresolved cited work.

DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis Unresolved cited work

Reference 22

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:46:23.960830Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T17:46:14.853327Z digest=sha256:a922ffc7a4f97056caf4aa15459648d2355774726bfb2eebddc223ebe8d90b03

Observation 1ee185fd-e353-4680-b406-11bc1b508eea · outbound

This paper cites an unresolved cited work.

DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis Unresolved cited work

Reference 23

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:46:23.412143Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T17:46:15.132879Z digest=sha256:19a13be5bca8cbc0782fcc254102f15a85073890caa81429c3cabababfb7adf7

Observation ce3f8763-5543-4aa8-a484-95c08a948cd2 · outbound

This paper cites an unresolved cited work.

DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:46:23.702173Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T17:46:15.031709Z digest=sha256:f88779363366fbf639e76a22728e4f833a662f71219d9fa539d476c5b7075bc5

Observation a7780ee1-ce41-48bc-aa46-3f2993a269f1 · outbound

This paper cites Mega-TTS 2: Boosting Prompting Mechanisms for Zero-Shot Speech Synthesis.

DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis Mega-TTS 2: Boosting Prompting Mechanisms for Zero-Shot Speech Synthesis

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T17:46:15.334498Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:46:15.334498Z digest=sha256:406ef15686d040b507ee17b17dcd3b243e45ebad1fe0bc3961b47b8af05517a7

Observation 2292ea90-dbba-4505-9b79-ede17c98baab · outbound

This paper cites WavTokenizer: an Efficient Acoustic Discrete Codec Tokenizer for Audio Language Modeling.

DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis WavTokenizer: an Efficient Acoustic Discrete Codec Tokenizer for Audio Language Modeling

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T17:46:15.230611Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:46:15.230611Z digest=sha256:3f856db8048588aa5a3a14ba3508fcffdf730ec37264b2eb2c1e76d909b135b4

Observation 03647407-1f64-4c07-840f-723e7258cd1f · outbound

This paper cites Fr\'echet Audio Distance: A Metric for Evaluating Music Enhancement Algorithms.

DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis Fr\'echet Audio Distance: A Metric for Evaluating Music Enhancement Algorithms

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T17:46:15.533022Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:46:15.533022Z digest=sha256:be5472177838632d861200316f375d2a6a72042090aee91e8765a33dc6aa38f9

Observation 1183abe4-53c7-4251-8fd3-81917221eb3f · outbound

This paper cites an unresolved cited work.

DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis Unresolved cited work

Reference 28

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:46:23.177628Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T17:46:15.431630Z digest=sha256:8340b1d155109ec3b4bda9e4e97fef48872aa01747aa9705a2f98b835af2813e

Observation cb3d7ec8-f2d5-4e92-acaf-e0b4150495b1 · outbound

This paper cites Salik, Rajiv Ratn Shah, Yifang Yin, and Roger Zimmermann.

DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis Salik, Rajiv Ratn Shah, Yifang Yin, and Roger Zimmermann

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:46:22.955376Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T17:46:15.732760Z digest=sha256:b4f06926955dbd000a18200889dec91105092accbb28f9907370426a1d9c4a8a

Observation cfe93e57-9358-4e6b-ae3c-1fd69c32cc39 · outbound

This paper cites an unresolved cited work.

DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis Unresolved cited work

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T17:46:15.631928Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:46:15.631928Z digest=sha256:f8075c5b3edb4c616c564db49a6777fda54a329dfae4cfd1ab9e5e48bb073939

Observation c1b4caf5-dafc-4d0a-bb99-b0f39eb64b0c · outbound

This paper cites an unresolved cited work.

DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis Unresolved cited work

Reference 31

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:46:22.591730Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T17:46:15.965671Z digest=sha256:7528ff8b2c34fad12376a62408f74321b49ec02fa12703cb0e53c8aca9171e1d

Observation 6f9c274d-f627-426c-829c-1842c25a0dfc · outbound

This paper cites an unresolved cited work.

DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis Unresolved cited work

Reference 32

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:46:22.788610Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T17:46:15.862042Z digest=sha256:ba9035412fc62c6398ea5d49cfa5c20c3f138ae18c6a0b6fbd323ecae2685723

Observation 3f69a0db-57dc-429a-8de0-b4d50dccc736 · outbound

This paper cites Mandic, Wenwu Wang, and Mark D.

DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis Mandic, Wenwu Wang, and Mark D

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:46:22.193356Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T17:46:16.165396Z digest=sha256:818bddff11e6492581c712115eec9b13b8800084f66ec8b136d9d432d2e3dca8

Observation 9cd59548-046e-4908-977a-3d14562cecff · outbound

This paper cites Raghavan, Gavin Mischler, and Nima Mesgarani.

DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis Raghavan, Gavin Mischler, and Nima Mesgarani

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:46:22.365402Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T17:46:16.059867Z digest=sha256:1886291dcd072d9429534faf578bee2fd4f2a78feaea5672a790f77ff6958f4c

Observation e3194c8a-04cd-421e-adef-e2557f37f3b6 · outbound

This paper cites Autoregressive Speech Synthesis without Vector Quantization.

DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis Autoregressive Speech Synthesis without Vector Quantization

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T17:46:16.401709Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:46:16.401709Z digest=sha256:c5ab9adae699a32fc003b6bffe24b0c16f3c2f970e65cf3c9d222069916cd939

Observation 99e73ac7-28ec-4a05-8240-968939682a8b · outbound

This paper cites an unresolved cited work.

DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis Unresolved cited work

Reference 36

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:46:21.977075Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T17:46:16.281334Z digest=sha256:748ecb0a2d9dc8497ed2a8cc9256046951c2afd7231bfe26faf3fb1ce1348a06

Observation 3af9a247-6149-4362-b3ec-3b5024433dbf · outbound

This paper cites an unresolved cited work.

DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis Unresolved cited work

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T17:46:16.612710Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:46:16.612710Z digest=sha256:85f68c015e3223c94b6b27cc0d9399dc6c11467fd17ed04014e503e3558158b0

Observation 42533c02-709b-444c-ba82-f108d0b1d43e · outbound

This paper cites an unresolved cited work.

DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis Unresolved cited work

Reference 38

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:46:21.844207Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T17:46:16.513575Z digest=sha256:de59f2d8770cf0803b33ba56cfe97c11af30780c8c12dbb9db2ca7f8b4e0ed3d

Observation 0c3fb29b-10e6-4788-9da2-c122d56c57c6 · outbound

This paper cites an unresolved cited work.

DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis Unresolved cited work

Reference 39

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:46:21.699975Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T17:46:16.856077Z digest=sha256:fbecdb42e5818646657bf2991dc22b9db6e354592d96cd0184fefb784698d778

Observation bf519b0d-1df2-45f7-a72c-fb11017f577f · outbound

This paper cites UTMOS: UTokyo-SaruLab System for VoiceMOS Challenge 2022.

DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis UTMOS: UTokyo-SaruLab System for VoiceMOS Challenge 2022

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T17:46:16.729243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:46:16.729243Z digest=sha256:25195be1a069a67b5bb0907432ac5fb657a20bfafc49514995c888035b7fef39

Observation 6d151fd1-b8cb-4da5-b6dd-dafea2042c69 · outbound

This paper cites an unresolved cited work.

DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis Unresolved cited work

Reference 41

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:46:21.376640Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T17:46:17.074451Z digest=sha256:e24cbe56099d40d10989c35d67fc4e6619784bed579944b00a5524dbf2a257ea

Observation d5aff840-67a9-4dbe-861c-bc0ed3ecf6f6 · outbound

This paper cites an unresolved cited work.

DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis Unresolved cited work

Reference 42

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:46:21.517553Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T17:46:16.972078Z digest=sha256:f418e216094f3fb5279b504ed35cb65749c996e5d81598a0a43915f32826b5be

Observation c778f8fa-b4e8-4720-a550-e1c19f05d7d0 · outbound

This paper cites FELLE: Autoregressive Speech Synthesis with Token-Wise Coarse-to-Fine Flow Matching.

DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis FELLE: Autoregressive Speech Synthesis with Token-Wise Coarse-to-Fine Flow Matching

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T17:46:17.295613Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:46:17.295613Z digest=sha256:2e78eaad60656b27c02470b5179ff0bd9ec9b94da5b90ef43f4cc217a78f3b9e

Observation 13942b3b-21ae-497b-8103-3869f8f17fe3 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis LLaMA: Open and Efficient Foundation Language Models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T17:46:17.194390Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:46:17.194390Z digest=sha256:f8f526158f6bba3580f7f0f6ebbcc48cdbf7c20f5a837e4a5df959e8974db5a1

Observation 5aadd354-21aa-434d-b3e4-e651d848f1c5 · outbound

This paper cites Mel-Band RoFormer for Music Source Separation.

DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis Mel-Band RoFormer for Music Source Separation

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T17:46:17.731914Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:46:17.731914Z digest=sha256:7d67524db7e41b72b1df8984f67d355e773b9f6debfdb1b76a97fe687cb036b7

Observation 80d1a0b5-080a-4f04-8e92-8ef224a181b4 · outbound

This paper cites an unresolved cited work.

DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis Unresolved cited work

Reference 46

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:46:21.242636Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T17:46:17.454006Z digest=sha256:0408f774f3b2913873d584878295c7374851abbcad1e56d4ccbd5db74cffeb25

Observation 25de76f6-9bd3-4faa-b791-113c7b55626b · outbound

This paper cites an unresolved cited work.

DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis Unresolved cited work

Reference 47

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:46:20.768252Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T17:46:18.070098Z digest=sha256:58dd7365430365df91271aeaeb4c42f6abc83397bbf6c7cf98e7751f6e69efe9

Observation 3869cc5b-f98e-484f-8fbd-91cb5e62436a · outbound

This paper cites an unresolved cited work.

DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis Unresolved cited work

Reference 48

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:46:20.618868Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T17:46:18.195413Z digest=sha256:04d8af014be8b069c6c21824952b14b7d4cefb59fcfb30f82929bb56f994ef50

Observation 91a24ba9-2e29-43ef-827b-95867ffd5fa3 · outbound

This paper cites an unresolved cited work.

DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis Unresolved cited work

Reference 49

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:46:20.915025Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T17:46:17.855498Z digest=sha256:21dd0282ebae04fa58035821721f8411212c6108f68fceb7c8aaaf4d14e63657

Observation c6e79082-fc53-44c9-8f1e-6ec750618d11 · outbound

This paper cites UniAudio: An Audio Foundation Model Toward Universal Audio Generation.

DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis UniAudio: An Audio Foundation Model Toward Universal Audio Generation

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T17:46:18.500277Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:46:18.500277Z digest=sha256:d80917f86553f4d743f35cc5543d0aa9f1a5cf08e9627fe18003e00695813ef9

Observation 5232a114-2d96-40a1-b022-b42e42fff0c2 · outbound

This paper cites an unresolved cited work.

DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis Unresolved cited work

Reference 51

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:46:20.489871Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T17:46:18.587362Z digest=sha256:06e59b16d9bcd2dda0d440e21071f1a294b95498f2ff5faf47a318900953fea4

Observation b3fbe57a-f389-445e-a3c7-8911c34f56f5 · outbound

This paper cites Qwen2 Technical Report.

DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis Qwen2 Technical Report

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-06T17:46:18.341711Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:46:18.341711Z digest=sha256:f4ca4adf5049d706d837919a7a22911ffa9d05bcde5961fc9c9974dcd7b2386d

Observation 7eda949d-c7f0-4718-9603-63ce8d097b60 · outbound

This paper cites LibriTTS: A Corpus Derived from LibriSpeech for Text-to-Speech.

DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis LibriTTS: A Corpus Derived from LibriSpeech for Text-to-Speech

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-06T17:46:18.798620Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:46:18.798620Z digest=sha256:787a3efec339c51b2d5cce8c045cd2ae81c01b9366301badc73713655969eecb

Observation 66285bcf-7fed-49f0-8900-4f4d5a0d744f · outbound

This paper cites FoleyCrafter: Bring Silent Videos to Life with Lifelike and Synchronized Sounds.

DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis FoleyCrafter: Bring Silent Videos to Life with Lifelike and Synchronized Sounds

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-06T17:46:18.900460Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:46:18.900460Z digest=sha256:999de092e4031a35cc140e83b406f18a603681760bb8def136c54abd54efd75a

Observation 7417fdab-75af-4ed8-918e-d769dee767a7 · outbound

This paper cites Diverse and Aligned Audio-to-Video Generation via Text-to-Video Model Adaptation.

DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis Diverse and Aligned Audio-to-Video Generation via Text-to-Video Model Adaptation

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-06T17:46:18.671928Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:46:18.671928Z digest=sha256:5bf37c35c0c7d5a00c757d0afdee5cbb769a0f96702f65c34b4b17888c7bbe2c

Observation f08d860c-4432-4423-a6b6-64609e10eb18 · outbound

This paper cites Llasa: Scaling Train-Time and Inference-Time Compute for Llama-based Speech Synthesis.

DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis Llasa: Scaling Train-Time and Inference-Time Compute for Llama-based Speech Synthesis

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-06T17:46:18.709302Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:46:18.709302Z digest=sha256:da634a0d037d84a1c387b303dec98f02afb4f5e368bcaf158f7df1b17f8ac3b5

Observation e39c3107-8a78-4cd1-b18b-1e2b4f16793d · outbound

This paper cites CosyAudio: Improving Audio Generation with Confidence Scores and Synthetic Captions.

DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis CosyAudio: Improving Audio Generation with Confidence Scores and Synthetic Captions

Reference 57

Resolution
verified exact
local_arxiv, observed 2026-08-06T17:46:19.565364Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T17:46:19.155991Z digest=sha256:31a9e1bc574550c3dc48b4b9d56aabb3153fc34663c900b2ed9c5afaa1cd2bff

Observation 5b4e0a7c-e4c5-4cd3-bead-1dee93540884 · outbound

This paper cites an unresolved cited work.

DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis Unresolved cited work

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-06T17:46:19.243232Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:46:19.243232Z digest=sha256:f8ce30bef62d23f66627fc9091098e17e6139737df9ff3a2e536cd62aa694e47

Observation 952abe25-feca-40f1-a9b9-985be6455714 · outbound

This paper cites an unresolved cited work.

DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis Unresolved cited work

Reference 59

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:46:20.344490Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T17:46:18.985234Z digest=sha256:9a7070361f6c872714ded2dac0af823f66ee991da08e58c00e12e07b509e8ab4

Observation 6bed8f7c-9191-4d2a-b1a7-59f1dd70c6e5 · outbound

This paper cites an unresolved cited work.

DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis Unresolved cited work

Reference 60

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:46:20.150751Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T17:46:19.066327Z digest=sha256:e409cd266f84bf11c3b5820e47f9aa6123e2ee0515be6025987cc8ad65c858b5

Observation 946540d6-4f50-4a0a-bf43-337744e816ff · outbound

This paper cites In IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2023, Vancouver, BC, Canada, June 17-24, 2023.

DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis In IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2023, Vancouver, BC, Canada, June 17-24, 2023

Reference 2023

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T17:46:25.054650Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T17:46:14.166030Z digest=sha256:66222619ad82abed2b311ae10bf56fa4170aeb7623293b99a9b191cb2256971f

Observation 6ecb8ab1-8022-41c7-bd52-9c10e0998921 · outbound

This paper cites an unresolved cited work.

DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis Unresolved cited work

Reference 2024

Resolution
unresolved
raw_fallback, observed 2026-08-06T17:46:21.053627Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T17:46:17.612626Z digest=sha256:49c299fcbf078bde562f040e7b033961c2fc7d07d7a1572c2ade66b8e1db9e6c

Pith citing papers

Observation 9f73405d-b2d7-4205-bea6-b06ea22e0311 · inbound

AudioGen-Omni: A Unified Multimodal Diffusion Transformer for Video-Synchronized Audio, Speech, and Song Generation cites this paper.

AudioGen-Omni: A Unified Multimodal Diffusion Transformer for Video-Synchronized Audio, Speech, and Song Generation DualDub: Video-to-Soundtrack Generation via Joint Speech and Background Audio Synthesis

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-08-06T06:04:30.083418Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T06:04:29.936928Z digest=sha256:da5ea92584b793dcdf658d806f2bf92ce0843e8d2232e80b34974ac7fa3a59fc