Pith. sign in

Paper Citation Record · LEDGER

UmbraTTS: Adapting Text-to-Speech to Environmental Contexts with Flow Matching

As of 9 August 2026, this Paper Citation Record lists 16 of 16 outbound references and 1 inbound Pith citation observation for arXiv:2506.09874.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.09874 v2

Coverage vector

measured 16 of 16 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T04:45:19.307090Z

measured 17 of 17 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-28T21:12:19.893944Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-01T20:26:12.658368Z

Reference resolution

16 of 16 outbound references displayed

  • verified exact1
  • verified fuzzy3
  • unresolved11
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 78962cc4-95b9-4726-bf81-1735024ca5ba · outbound

This paper cites Audiobox: Unified Audio Generation with Natural Language Prompts.

UmbraTTS: Adapting Text-to-Speech to Environmental Contexts with Flow Matching Audiobox: Unified Audio Generation with Natural Language Prompts

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T04:45:19.234441Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:45:19.234441Z digest=sha256:7adcbdcf0a7e004abb97c238f8679d8682811e2af2d3d25278fef3e83040a488

Observation 41accd41-1833-4715-b628-8871c45e6645 · outbound

This paper cites Fr\'echet Audio Distance: A Metric for Evaluating Music Enhancement Algorithms.

UmbraTTS: Adapting Text-to-Speech to Environmental Contexts with Flow Matching Fr\'echet Audio Distance: A Metric for Evaluating Music Enhancement Algorithms

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T04:45:18.687085Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:45:18.687085Z digest=sha256:3297676ff8da7689bc4c80fa818d36005533a9156b44b7a7af0b75307f692077

Observation dd1ab871-4070-40d1-bc54-9ce02b7be156 · outbound

This paper cites AudioGen: Textually Guided Audio Generation.

UmbraTTS: Adapting Text-to-Speech to Environmental Contexts with Flow Matching AudioGen: Textually Guided Audio Generation

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T04:45:18.835009Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:45:18.835009Z digest=sha256:19bc987ba9a60190a4326c516dde860dc4193c25bbf44ffe489b09bfb505e432

Observation 61ac979c-25d9-4ad5-aa99-8b313ae3ec8d · outbound

This paper cites Flow Matching for Generative Modeling.

UmbraTTS: Adapting Text-to-Speech to Environmental Contexts with Flow Matching Flow Matching for Generative Modeling

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T04:45:18.891032Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:45:18.891032Z digest=sha256:d0727817c093d15f9f996819213c07ccec2a8a7cf8dd850a2341c868b5eeeec9

Observation 87fd9334-847c-4c91-abcb-77ad8bf938a9 · outbound

This paper cites AudioLDM: Text-to-Audio Generation with Latent Diffusion Models.

UmbraTTS: Adapting Text-to-Speech to Environmental Contexts with Flow Matching AudioLDM: Text-to-Audio Generation with Latent Diffusion Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T04:45:18.973851Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:45:18.973851Z digest=sha256:9a06e81066cc0f562929a49a768b3fa0932e204bb3db028fd7ed142430bd2921

Observation d4c66cb4-ace9-4a88-8fcf-9256a36ca013 · outbound

This paper cites FlowTSE: Target Speaker Extraction with Flow Matching.

UmbraTTS: Adapting Text-to-Speech to Environmental Contexts with Flow Matching FlowTSE: Target Speaker Extraction with Flow Matching

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-08-07T04:45:19.443377Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T04:45:19.031576Z digest=sha256:b9610b888cbba0e8fc7890948e55e8a57c9a567fc460ebcb8e67531e3dacba80

Observation e0cd1bd0-62d1-4a9c-8e8c-c581427ab987 · outbound

This paper cites FastSpeech 2: Fast and High-Quality End-to-End Text to Speech.

UmbraTTS: Adapting Text-to-Speech to Environmental Contexts with Flow Matching FastSpeech 2: Fast and High-Quality End-to-End Text to Speech

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T04:45:19.164512Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:45:19.164512Z digest=sha256:c200cb4cb1e2ce61a71fab4fabd3f11dada32b3838e67507420427b5fa4880c0

Observation 397edf0a-3a63-4222-8951-32e8d2ce87a7 · outbound

This paper cites A Survey on Neural Speech Synthesis.

UmbraTTS: Adapting Text-to-Speech to Environmental Contexts with Flow Matching A Survey on Neural Speech Synthesis

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T04:45:19.182281Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:45:19.182281Z digest=sha256:ffd4d26174f54946bf3a0663a40c4832424beda7d92ad174483ee05f68f6449c

Observation 32340c6e-8d6a-47fb-92eb-cd7e6ca085cd · outbound

This paper cites Large-scale contrastive language-audio pretraining with feature fusion and keyword-to-caption augmentation.

UmbraTTS: Adapting Text-to-Speech to Environmental Contexts with Flow Matching Large-scale contrastive language-audio pretraining with feature fusion and keyword-to-caption augmentation

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:45:19.578299Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T04:45:19.283562Z digest=sha256:af3459ffe0abdb4f0871474338f9d59fbf2096833df688398d2d29a676031ac5

Observation ba40f52d-e387-48a7-9d1b-dcec9d791bdd · outbound

This paper cites A Survey on Audio Diffusion Models: Text To Speech Synthesis and Enhancement in Generative AI.

UmbraTTS: Adapting Text-to-Speech to Environmental Contexts with Flow Matching A Survey on Audio Diffusion Models: Text To Speech Synthesis and Enhancement in Generative AI

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T04:45:19.307090Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:45:19.307090Z digest=sha256:39ed86cdb57479cba38c5f6ddcc969fad8e35cfd7a3e2ab95da5495f942da446

Observation 453f421e-7868-45c0-af7c-3bc0739412a9 · outbound

This paper cites Enhancing Speech Intelligibility in Text-To-Speech Synthesis using Speaking Style Conversion.

UmbraTTS: Adapting Text-to-Speech to Environmental Contexts with Flow Matching Enhancing Speech Intelligibility in Text-To-Speech Synthesis using Speaking Style Conversion

Reference 2015

Resolution
unresolved
no resolver link, observed 2026-08-07T04:45:19.104573Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:45:19.104573Z digest=sha256:a09e34d9e505984e68c57acf843d4a3f5ddb60ccb5d7b253f631053df77d02cb

Observation 80e90982-39fd-4958-8789-b936820b4f61 · outbound

This paper cites Emilia: An extensive, multi- lingual, and diverse speech dataset for large-scale speech generation.

UmbraTTS: Adapting Text-to-Speech to Environmental Contexts with Flow Matching Emilia: An extensive, multi- lingual, and diverse speech dataset for large-scale speech generation

Reference 2017

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:45:19.593624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T04:45:18.667628Z digest=sha256:f00383a5587d97b1afa4b03096b2b052dc230dcb55fcbd7f06e7477214916990

Observation 850f0b23-de79-4904-b426-71fde83fc81d · outbound

This paper cites Speak in the Scene: Diffusion-based Acoustic Scene Transfer toward Immersive Speech Generation.

UmbraTTS: Adapting Text-to-Speech to Environmental Contexts with Flow Matching Speak in the Scene: Diffusion-based Acoustic Scene Transfer toward Immersive Speech Generation

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-07T04:45:18.755238Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:45:18.755238Z digest=sha256:0e9a8de10ae803a0f7077f74870b41c24acf4c2aa14cdc406ff6745d0fa8dfa6

Observation f960d625-ac9f-4934-ac86-60a071372484 · outbound

This paper cites F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching.

UmbraTTS: Adapting Text-to-Speech to Environmental Contexts with Flow Matching F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-07T04:45:18.608842Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:45:18.608842Z digest=sha256:ed0c2a24fc4432b932df5c097b8dfb061c566b53d6df1e06abd4d4370b08e248

Observation 0176a76c-24a2-4bec-b123-ff19bb18f4bf · outbound

This paper cites VoiceDiT: Dual-Condition Diffusion Transformer for Environment-Aware Speech Synthesis.

UmbraTTS: Adapting Text-to-Speech to Environmental Contexts with Flow Matching VoiceDiT: Dual-Condition Diffusion Transformer for Environment-Aware Speech Synthesis

Reference 2023

Resolution
metadata mismatch
local_arxiv, observed 2026-08-07T04:45:19.546105Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T04:45:18.672137Z digest=sha256:015227cfa2d50ed089dce717eac44f40e12a708e07079f2cc6e7ca2d1ab8ed85

Observation 7b16013c-daf3-4120-bcbe-ffb91856b964 · outbound

This paper cites E., Wang, X., Thakker, M., Li, C., Tsai, C.-H., Xiao, Z., Yang, H., Zhu, Z., Tang, M., Tan, X., et al.

UmbraTTS: Adapting Text-to-Speech to Environmental Contexts with Flow Matching E., Wang, X., Thakker, M., Li, C., Tsai, C.-H., Xiao, Z., Yang, H., Zhu, Z., Tang, M., Tan, X., et al

Reference 2024

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:45:19.608252Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T04:45:18.647201Z digest=sha256:1cfe3d8fc70828455b2f9d15b661e98f35dc0579341f0e617810c4fa68d9c788

Pith citing papers

Observation 48b72c44-7294-484b-9ac6-93694d3d2945 · inbound

ImmersiveTTS: Environment-Aware Text-to-Speech with Multimodal Diffusion Transformer and Domain-Specific Representation Alignment cites this paper.

ImmersiveTTS: Environment-Aware Text-to-Speech with Multimodal Diffusion Transformer and Domain-Specific Representation Alignment UmbraTTS: Adapting Text-to-Speech to Environmental Contexts with Flow Matching

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-07-01T20:26:12.661007Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-28T21:12:19.893944Z digest=sha256:7663d3944d6f4f0dac228eef5c04c805b2bc35f0e7f850e94090b317546f3374