Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T00:50:14.477869Z
Paper Citation Record · LEDGER
As of 18 August 2026, this Paper Citation Record lists 47 of 47 outbound references and 1 inbound Pith citation observation for arXiv:2412.19259.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T00:50:14.477869Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T04:45:18.672137Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-07T04:45:19.541066Z
47 of 47 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 6b32beb8-aab1-4eac-bc9f-74fb71444762 · outbound
VoiceDiT: Dual-Condition Diffusion Transformer for Environment-Aware Speech Synthesis Diffsound: Discrete diffusion model for text-to-sound generation,
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 094f55ab-107d-41e0-9c48-8717f5112feb · outbound
VoiceDiT: Dual-Condition Diffusion Transformer for Environment-Aware Speech Synthesis Audiogen: T extually guided audio generation,
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 3ad2f980-d9d0-4127-a929-4660690f676a · outbound
VoiceDiT: Dual-Condition Diffusion Transformer for Environment-Aware Speech Synthesis Make-an-audio: Text-to-audio generation with prompt-enhanced diffusion models,
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation a5f91e7a-310c-4ed2-8aee-62127335f342 · outbound
VoiceDiT: Dual-Condition Diffusion Transformer for Environment-Aware Speech Synthesis Text-to-audio generation using instruction-tuned llm and latent diffusion model,
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation f7fb3555-bf53-43f4-914c-512933627bd9 · outbound
VoiceDiT: Dual-Condition Diffusion Transformer for Environment-Aware Speech Synthesis Fast timing-conditioned latent audio diffusion,
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 5826aceb-20be-4475-9594-38094d696869 · outbound
VoiceDiT: Dual-Condition Diffusion Transformer for Environment-Aware Speech Synthesis Audiolcm: T ext-to-audio generation with latent consistency models,
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 96161442-fdb3-4d2c-a669-0a7587eedd84 · outbound
VoiceDiT: Dual-Condition Diffusion Transformer for Environment-Aware Speech Synthesis Auffusion: Leveraging the power of diffusion and large language models for text-to-audio generation,
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 866bf51f-7348-4d09-b6a3-f98108cbacce · outbound
VoiceDiT: Dual-Condition Diffusion Transformer for Environment-Aware Speech Synthesis Grad-tts: A diffusion probabilistic model for text-to-speech,
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation b83ad586-66bd-442f-aa2f-57da8d966bdf · outbound
VoiceDiT: Dual-Condition Diffusion Transformer for Environment-Aware Speech Synthesis Glow-tts: A generative flow for text-to-speech via monotonic alignment search,
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 3570f116-6a1e-4efe-91f6-b6190acf802a · outbound
VoiceDiT: Dual-Condition Diffusion Transformer for Environment-Aware Speech Synthesis Guided-tts: A diffusion model for text-to-speech via classifier guidance,
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation a07c7764-289b-4eb4-8d11-7ed93878560e · outbound
VoiceDiT: Dual-Condition Diffusion Transformer for Environment-Aware Speech Synthesis Prosody- tts: Improving prosody with masked autoencoder and conditional diffusion model for expressive text-to-speech,
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 05595412-d585-489b-a585-a3e52d1ffac4 · outbound
VoiceDiT: Dual-Condition Diffusion Transformer for Environment-Aware Speech Synthesis Crossspeech: Speaker-independent acoustic representation for cross-lingual speech synthesis,
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 8d2cf736-2145-48f0-8635-ceb07a5fbd8a · outbound
VoiceDiT: Dual-Condition Diffusion Transformer for Environment-Aware Speech Synthesis Fregrad: Lightweight and fast frequency-aware diffusion vocoder,
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 5a9ab71d-c656-4bb3-a954-3049b1c08c85 · outbound
VoiceDiT: Dual-Condition Diffusion Transformer for Environment-Aware Speech Synthesis Naturalspeech 3: Zero-shot speech synthesis with factorized codec and diffusion models,
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 8b8859f6-5518-4147-b656-d74c0e71beab · outbound
VoiceDiT: Dual-Condition Diffusion Transformer for Environment-Aware Speech Synthesis DiTTo-TTS: Diffusion Transformers for Scalable Text-to-Speech without Domain-Specific Factors
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 054bde7a-7c6c-4ef6-be1c-69dc370506ee · outbound
VoiceDiT: Dual-Condition Diffusion Transformer for Environment-Aware Speech Synthesis High-resolution image synthesis with latent diffusion models,
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 5b16b0d9-ce4a-441c-aec6-5333b0a9a7fd · outbound
VoiceDiT: Dual-Condition Diffusion Transformer for Environment-Aware Speech Synthesis Score-based generative modeling through stochastic differential equations,
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 2f27dd73-d8ea-47e4-888a-189716c5e863 · outbound
VoiceDiT: Dual-Condition Diffusion Transformer for Environment-Aware Speech Synthesis Denoising diffusion probabilistic models,
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 0dde5b27-b2bd-42bf-af60-0cb16a1b94d0 · outbound
VoiceDiT: Dual-Condition Diffusion Transformer for Environment-Aware Speech Synthesis Audioldm: Text-to-audio generation with latent diffusion models,
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 97ac5869-fa74-40a9-90c8-df348e62accd · outbound
VoiceDiT: Dual-Condition Diffusion Transformer for Environment-Aware Speech Synthesis V oiceldm: T ext-to-speech with environmental context,
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 61f37537-7038-4ae3-bf5a-78615cee647e · outbound
VoiceDiT: Dual-Condition Diffusion Transformer for Environment-Aware Speech Synthesis Audiobox: Unified Audio Generation with Natural Language Prompts
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation df5d3f5d-cf88-4c99-ba04-0127df0485c3 · outbound
VoiceDiT: Dual-Condition Diffusion Transformer for Environment-Aware Speech Synthesis Audi- oldm 2: Learning holistic audio generation with self-supervised pretraining,
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation a18ce097-5470-44c5-ae28-1137c2850153 · outbound
VoiceDiT: Dual-Condition Diffusion Transformer for Environment-Aware Speech Synthesis UniAudio: An Audio Foundation Model Toward Universal Audio Generation
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 59acc64c-a9fd-4fcb-8d5f-e372abf275de · outbound
VoiceDiT: Dual-Condition Diffusion Transformer for Environment-Aware Speech Synthesis Large-scale contrastive language-audio pretraining with feature fusion and keyword-to-caption augmentation,
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 417bc99a-a84f-4c07-9d38-b7b5b5545f87 · outbound
VoiceDiT: Dual-Condition Diffusion Transformer for Environment-Aware Speech Synthesis Libritts: A corpus derived from librispeech for text-to-speech,
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation eb003fc7-c6f4-4ad0-8706-355b6141c441 · outbound
VoiceDiT: Dual-Condition Diffusion Transformer for Environment-Aware Speech Synthesis Common Voice: A Massively-Multilingual Speech Corpus
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f48a203c-1e89-48c3-af62-8f8be0893b2e · outbound
VoiceDiT: Dual-Condition Diffusion Transformer for Environment-Aware Speech Synthesis VoxCeleb: a large-scale speaker identification dataset
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1291cacf-1c3f-403c-b3b0-680ff8d89c55 · outbound
VoiceDiT: Dual-Condition Diffusion Transformer for Environment-Aware Speech Synthesis V oxceleb2: Deep speaker recognition,
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 073bab1a-186b-4873-9d8f-be69211a50e8 · outbound
VoiceDiT: Dual-Condition Diffusion Transformer for Environment-Aware Speech Synthesis V oxmm: Rich transcription of conversations in the wild,
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 13e62add-ba5f-4567-b03b-edbb80afbbda · outbound
VoiceDiT: Dual-Condition Diffusion Transformer for Environment-Aware Speech Synthesis W avcaps: A chatgpt-assisted weakly-labelled audio captioning dataset for audio-language multimodal research,
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation e317840c-6275-4d3c-a7b4-45af724f4a20 · outbound
VoiceDiT: Dual-Condition Diffusion Transformer for Environment-Aware Speech Synthesis Libritts-r: A restored multi-speaker text-to-speech corpus,
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 5a5c5051-1778-4204-baf2-c1a8100db60b · outbound
VoiceDiT: Dual-Condition Diffusion Transformer for Environment-Aware Speech Synthesis Metric learning for user-defined keyword spotting,
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation ea8cbf64-ed8d-4d1b-99b0-a6c1311caa85 · outbound
VoiceDiT: Dual-Condition Diffusion Transformer for Environment-Aware Speech Synthesis Robust speech recognition via large-scale weak supervision,
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation da3a4373-5002-4315-b025-6d3554c11bd4 · outbound
VoiceDiT: Dual-Condition Diffusion Transformer for Environment-Aware Speech Synthesis Scalable diffusion models with transformers,
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation b67da38a-ce2c-4d72-b657-6310103773a3 · outbound
VoiceDiT: Dual-Condition Diffusion Transformer for Environment-Aware Speech Synthesis PixArt-$\alpha$: Fast Training of Diffusion Transformer for Photorealistic Text-to-Image Synthesis
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 01d1a473-d10b-49d5-829f-0ee88de9051c · outbound
VoiceDiT: Dual-Condition Diffusion Transformer for Environment-Aware Speech Synthesis Longformer: The Long-Document Transformer
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 27796288-61a8-4e45-9c4b-e705bfd00286 · outbound
VoiceDiT: Dual-Condition Diffusion Transformer for Environment-Aware Speech Synthesis Attention is all you need,
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 7f9b4b55-19c7-44e7-bb70-2f55a367775e · outbound
VoiceDiT: Dual-Condition Diffusion Transformer for Environment-Aware Speech Synthesis Film: Visual reasoning with a general conditioning layer,
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 88288ab1-4d2f-4530-916b-5ed8639edebf · outbound
VoiceDiT: Dual-Condition Diffusion Transformer for Environment-Aware Speech Synthesis V2a-mapper: A lightweight solution for vision-to-audio generation by connecting foundation models,
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 1c70662f-3fda-493c-9a53-3376fb623a9c · outbound
VoiceDiT: Dual-Condition Diffusion Transformer for Environment-Aware Speech Synthesis Learning transferable visual models from natural language supervision,
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 8e850ab3-abff-44bd-b0cb-c6e4e8a0759c · outbound
VoiceDiT: Dual-Condition Diffusion Transformer for Environment-Aware Speech Synthesis Hifi-gan: Generative adversarial networks for efficient and high fidelity speech synthesis,
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 325ff93c-b15e-402e-ae75-a33f32770a4e · outbound
VoiceDiT: Dual-Condition Diffusion Transformer for Environment-Aware Speech Synthesis W avlm: Large-scale self-supervised pre-training for full stack speech processing,
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 92de90e0-cda9-4b98-97db-384300238ca1 · outbound
VoiceDiT: Dual-Condition Diffusion Transformer for Environment-Aware Speech Synthesis Espnet-spk: full pipeline speaker embedding toolkit with reproducible recipes, self-supervised front-ends, and off-the-shelf models,
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation d0e2697d-9dc6-444d-be79-6328668b6409 · outbound
VoiceDiT: Dual-Condition Diffusion Transformer for Environment-Aware Speech Synthesis Decoupled weight decay regularization,
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation c405fef8-c725-4275-a956-6b3578d0e8b1 · outbound
VoiceDiT: Dual-Condition Diffusion Transformer for Environment-Aware Speech Synthesis Fr\'echet Audio Distance: A Metric for Evaluating Music Enhancement Algorithms
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8aff6a2c-f666-456c-bdb0-52c603beb41d · outbound
VoiceDiT: Dual-Condition Diffusion Transformer for Environment-Aware Speech Synthesis I hear your true colors: Image guided audio generation,
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation a1446cfc-b7be-493a-b313-a8e7e3f96f16 · outbound
VoiceDiT: Dual-Condition Diffusion Transformer for Environment-Aware Speech Synthesis T aming visually guided sound generation,
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 0176a76c-24a2-4bec-b123-ff19bb18f4bf · inbound
UmbraTTS: Adapting Text-to-Speech to Environmental Contexts with Flow Matching VoiceDiT: Dual-Condition Diffusion Transformer for Environment-Aware Speech Synthesis
Reference 2023
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.