Pith. sign in

Paper Citation Record · LEDGER

UmbraTTS: Adapting Text-to-Speech to Environmental Contexts with Flow Matching

As of 14 August 2026, this Paper Citation Record lists 16 of 16 outbound references and 1 inbound Pith citation observation for arXiv:2506.09874.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.09874 v2

Coverage vector

measured 16 of 16 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T04:45:19.307090Z

measured 17 of 17 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-28T21:12:19.893944Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-01T20:26:12.658368Z

Reference resolution

16 of 16 outbound references displayed

  • verified exact1
  • verified fuzzy3
  • unresolved11
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 78962cc4-95b9-4726-bf81-1735024ca5ba · outbound

This paper cites Audiobox: Unified Audio Generation with Natural Language Prompts.

UmbraTTS: Adapting Text-to-Speech to Environmental Contexts with Flow Matching Audiobox: Unified Audio Generation with Natural Language Prompts

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T04:45:19.234441Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:45:19.234441Z digest=sha256:318e6d5300e48730cad166de36d6254f40a50de544da6593c3bf6b8ac13eaebe

Observation 41accd41-1833-4715-b628-8871c45e6645 · outbound

This paper cites Fr\'echet Audio Distance: A Metric for Evaluating Music Enhancement Algorithms.

UmbraTTS: Adapting Text-to-Speech to Environmental Contexts with Flow Matching Fr\'echet Audio Distance: A Metric for Evaluating Music Enhancement Algorithms

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T04:45:18.687085Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:45:18.687085Z digest=sha256:340662aea47cbc8d39851082809b178a446ab598a1ff2a88f61a0f9554c0291a

Observation dd1ab871-4070-40d1-bc54-9ce02b7be156 · outbound

This paper cites AudioGen: Textually Guided Audio Generation.

UmbraTTS: Adapting Text-to-Speech to Environmental Contexts with Flow Matching AudioGen: Textually Guided Audio Generation

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T04:45:18.835009Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:45:18.835009Z digest=sha256:807afa202dbba8cc763e0718c215e1aed215d3a233ac4be851ca83a2cd920eed

Observation 61ac979c-25d9-4ad5-aa99-8b313ae3ec8d · outbound

This paper cites Flow Matching for Generative Modeling.

UmbraTTS: Adapting Text-to-Speech to Environmental Contexts with Flow Matching Flow Matching for Generative Modeling

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T04:45:18.891032Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:45:18.891032Z digest=sha256:d0727817c093d15f9f996819213c07ccec2a8a7cf8dd850a2341c868b5eeeec9

Observation 87fd9334-847c-4c91-abcb-77ad8bf938a9 · outbound

This paper cites AudioLDM: Text-to-Audio Generation with Latent Diffusion Models.

UmbraTTS: Adapting Text-to-Speech to Environmental Contexts with Flow Matching AudioLDM: Text-to-Audio Generation with Latent Diffusion Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T04:45:18.973851Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:45:18.973851Z digest=sha256:6d7142b40bd8a8f9f5dbdd53c9b81c9eb4f68a508287f56d0f6b837342039e90

Observation d4c66cb4-ace9-4a88-8fcf-9256a36ca013 · outbound

This paper cites FlowTSE: Target Speaker Extraction with Flow Matching.

UmbraTTS: Adapting Text-to-Speech to Environmental Contexts with Flow Matching FlowTSE: Target Speaker Extraction with Flow Matching

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-08-07T04:45:19.443377Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T04:45:19.031576Z digest=sha256:def9cfe11e81c4f8dd53a8e76ddbad7ded17d76afd47d05f174eaa4d50c89649

Observation e0cd1bd0-62d1-4a9c-8e8c-c581427ab987 · outbound

This paper cites FastSpeech 2: Fast and High-Quality End-to-End Text to Speech.

UmbraTTS: Adapting Text-to-Speech to Environmental Contexts with Flow Matching FastSpeech 2: Fast and High-Quality End-to-End Text to Speech

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T04:45:19.164512Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:45:19.164512Z digest=sha256:b33c512869f31cc40dcdb9586d5e759846b02488e53d7d830e6ffc697a8f8289

Observation 397edf0a-3a63-4222-8951-32e8d2ce87a7 · outbound

This paper cites A Survey on Neural Speech Synthesis.

UmbraTTS: Adapting Text-to-Speech to Environmental Contexts with Flow Matching A Survey on Neural Speech Synthesis

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T04:45:19.182281Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:45:19.182281Z digest=sha256:27ff07403bc19ae448d583b0615e6a1ca46d7b0033d205e02f24d1c18ea85433

Observation 32340c6e-8d6a-47fb-92eb-cd7e6ca085cd · outbound

This paper cites Large-scale contrastive language-audio pretraining with feature fusion and keyword-to-caption augmentation.

UmbraTTS: Adapting Text-to-Speech to Environmental Contexts with Flow Matching Large-scale contrastive language-audio pretraining with feature fusion and keyword-to-caption augmentation

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:45:19.578299Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T04:45:19.283562Z digest=sha256:45cf8a453acd3938c11638ad5e8893082885f32c13ba8dae5cf666cba7495a1d

Observation ba40f52d-e387-48a7-9d1b-dcec9d791bdd · outbound

This paper cites A Survey on Audio Diffusion Models: Text To Speech Synthesis and Enhancement in Generative AI.

UmbraTTS: Adapting Text-to-Speech to Environmental Contexts with Flow Matching A Survey on Audio Diffusion Models: Text To Speech Synthesis and Enhancement in Generative AI

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T04:45:19.307090Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:45:19.307090Z digest=sha256:c9be57db6d76bd28fb14884d85e8e576b49a264c101f163700acf6ec2cfe198b

Observation 453f421e-7868-45c0-af7c-3bc0739412a9 · outbound

This paper cites Enhancing Speech Intelligibility in Text-To-Speech Synthesis using Speaking Style Conversion.

UmbraTTS: Adapting Text-to-Speech to Environmental Contexts with Flow Matching Enhancing Speech Intelligibility in Text-To-Speech Synthesis using Speaking Style Conversion

Reference 2015

Resolution
unresolved
no resolver link, observed 2026-08-07T04:45:19.104573Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:45:19.104573Z digest=sha256:a3c1690efc65319b9c21ca4343a47383d0b5a64a18a946c5ba2bfec7be2eac4d

Observation 80e90982-39fd-4958-8789-b936820b4f61 · outbound

This paper cites Emilia: An extensive, multi- lingual, and diverse speech dataset for large-scale speech generation.

UmbraTTS: Adapting Text-to-Speech to Environmental Contexts with Flow Matching Emilia: An extensive, multi- lingual, and diverse speech dataset for large-scale speech generation

Reference 2017

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:45:19.593624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T04:45:18.667628Z digest=sha256:b269b67289c9862fffdf547d7a087543223850dbb8c3af4dbc764bfe9d18f354

Observation 850f0b23-de79-4904-b426-71fde83fc81d · outbound

This paper cites Speak in the Scene: Diffusion-based Acoustic Scene Transfer toward Immersive Speech Generation.

UmbraTTS: Adapting Text-to-Speech to Environmental Contexts with Flow Matching Speak in the Scene: Diffusion-based Acoustic Scene Transfer toward Immersive Speech Generation

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-07T04:45:18.755238Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:45:18.755238Z digest=sha256:63d7cf54a0d502495bfdac753de2e858fbc6d2b9d2b1f204a5b4f1e3d90b56ec

Observation f960d625-ac9f-4934-ac86-60a071372484 · outbound

This paper cites F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching.

UmbraTTS: Adapting Text-to-Speech to Environmental Contexts with Flow Matching F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-07T04:45:18.608842Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:45:18.608842Z digest=sha256:ed0c2a24fc4432b932df5c097b8dfb061c566b53d6df1e06abd4d4370b08e248

Observation 0176a76c-24a2-4bec-b123-ff19bb18f4bf · outbound

This paper cites VoiceDiT: Dual-Condition Diffusion Transformer for Environment-Aware Speech Synthesis.

UmbraTTS: Adapting Text-to-Speech to Environmental Contexts with Flow Matching VoiceDiT: Dual-Condition Diffusion Transformer for Environment-Aware Speech Synthesis

Reference 2023

Resolution
metadata mismatch
local_arxiv, observed 2026-08-07T04:45:19.546105Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T04:45:18.672137Z digest=sha256:c696a5f3677fec90428f048a63ec6d9a3a8d3360542d003c10b8551d7315e318

Observation 7b16013c-daf3-4120-bcbe-ffb91856b964 · outbound

This paper cites E., Wang, X., Thakker, M., Li, C., Tsai, C.-H., Xiao, Z., Yang, H., Zhu, Z., Tang, M., Tan, X., et al.

UmbraTTS: Adapting Text-to-Speech to Environmental Contexts with Flow Matching E., Wang, X., Thakker, M., Li, C., Tsai, C.-H., Xiao, Z., Yang, H., Zhu, Z., Tang, M., Tan, X., et al

Reference 2024

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:45:19.608252Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T04:45:18.647201Z digest=sha256:fc69b5f3935cb67d70eabeed84877e35138db591c055e94a90ffe06fa6825e68

Pith citing papers

Observation 48b72c44-7294-484b-9ac6-93694d3d2945 · inbound

ImmersiveTTS: Environment-Aware Text-to-Speech with Multimodal Diffusion Transformer and Domain-Specific Representation Alignment cites this paper.

ImmersiveTTS: Environment-Aware Text-to-Speech with Multimodal Diffusion Transformer and Domain-Specific Representation Alignment UmbraTTS: Adapting Text-to-Speech to Environmental Contexts with Flow Matching

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-07-01T20:26:12.661007Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-06-28T21:12:19.893944Z digest=sha256:861997ae16a0ff21b65475b54c6a3a006a3ff6703f7447a062ed8cd7ca55c782