Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T14:15:55.714499Z
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 45 of 45 outbound references and 1 inbound Pith citation observation for arXiv:2505.19595.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T14:15:55.714499Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T14:15:52.829136Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-07T14:15:55.827429Z
45 of 45 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation bb7fc26a-ef85-4b17-9cd1-4f8198f10e5c · outbound
Accelerating Diffusion-based Text-to-Speech Model Training with Dual Modality Alignment Unresolved cited work
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 7613c17a-498e-42a3-8d17-dd44120f1d7c · outbound
Accelerating Diffusion-based Text-to-Speech Model Training with Dual Modality Alignment Accelerating Diffusion-based Text-to-Speech Model Training with Dual Modality Alignment
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 01dc9593-8e57-4e1e-a048-8da73aba0b7c · outbound
Accelerating Diffusion-based Text-to-Speech Model Training with Dual Modality Alignment Dataset We train our model on the LibriTTS [31] dataset, a multi- speaker English corpus containing approximately 585 hours of read speech sampled at 24 kHz
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 702f54ac-37a2-4adc-a8b5-3a746a42d4ac · outbound
Accelerating Diffusion-based Text-to-Speech Model Training with Dual Modality Alignment speech align
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 0faadb96-b8ce-493a-b09d-7274d9b0a12a · outbound
Accelerating Diffusion-based Text-to-Speech Model Training with Dual Modality Alignment This ap- proach significantly accelerates model convergence while en- hancing the quality of generated speech
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 98f04b34-3288-47ce-a147-dd04790540d3 · outbound
Accelerating Diffusion-based Text-to-Speech Model Training with Dual Modality Alignment In future work, we plan to further op- timize the framework by reducing model complexity and accel- erating inference
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 258d8797-2481-4966-a231-dbe012515b96 · outbound
Accelerating Diffusion-based Text-to-Speech Model Training with Dual Modality Alignment Choi, J.-H
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 6cc1c89c-7940-405c-8b17-9a854be9677f · outbound
Accelerating Diffusion-based Text-to-Speech Model Training with Dual Modality Alignment Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1d4d6018-6c07-457a-8bcd-920d8e4b9f3c · outbound
Accelerating Diffusion-based Text-to-Speech Model Training with Dual Modality Alignment Seed-TTS: A Family of High-Quality Versatile Speech Generation Models
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d2e23561-cb4c-4429-b577-f92f37f4fae7 · outbound
Accelerating Diffusion-based Text-to-Speech Model Training with Dual Modality Alignment V oicecraft: Zero-shot speech editing and text-to-speech in the wild,
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation e6714b00-7472-44e4-8dd7-2f684454f6c5 · outbound
Accelerating Diffusion-based Text-to-Speech Model Training with Dual Modality Alignment BASE TTS: Lessons from building a billion-parameter Text-to-Speech model on 100K hours of data
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 94639417-d75f-4193-b59b-dd326c03cc02 · outbound
Accelerating Diffusion-based Text-to-Speech Model Training with Dual Modality Alignment Autoregressive Speech Synthesis without Vector Quantization
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e41f997a-0485-4e23-98ed-f51d6bd54884 · outbound
Accelerating Diffusion-based Text-to-Speech Model Training with Dual Modality Alignment V oice- box: Text-guided multilingual universal speech generation at scale,
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 27bab5ff-4fe7-4713-8a79-6587a2a54bbc · outbound
Accelerating Diffusion-based Text-to-Speech Model Training with Dual Modality Alignment Naturalspeech 2: Latent diffusion models are natural and zero-shot speech and singing synthesizers,
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 5c1f7394-f16b-432a-8e9c-fe282429a3cf · outbound
Accelerating Diffusion-based Text-to-Speech Model Training with Dual Modality Alignment E2 tts: Embarrassingly easy fully non-autoregressive zero-shot tts,
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 9596f8ff-8a8d-4e9a-8684-6b308bf5e031 · outbound
Accelerating Diffusion-based Text-to-Speech Model Training with Dual Modality Alignment F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c507b2b2-efe6-46d7-8b42-232f6c1a8f0c · outbound
Accelerating Diffusion-based Text-to-Speech Model Training with Dual Modality Alignment Ditto-tts: Efficient and scalable zero-shot text-to-speech with diffusion transformer,
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 5f3b1eba-920e-48f8-b4c0-9c14da64bfbe · outbound
Accelerating Diffusion-based Text-to-Speech Model Training with Dual Modality Alignment Accelerating codec-based speech synthesis with multi- token prediction and speculative decoding,
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 578d911b-13af-4ec7-899d-516c8e495232 · outbound
Accelerating Diffusion-based Text-to-Speech Model Training with Dual Modality Alignment VALL-E R: Robust and Efficient Zero-Shot Text-to-Speech Synthesis via Monotonic Alignment
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 37e878e4-81ed-4917-b258-be620ad63813 · outbound
Accelerating Diffusion-based Text-to-Speech Model Training with Dual Modality Alignment Vall-t: Decoder-only generative transducer for robust and decoding-controllable text-to-speech,
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation e4889d69-d0b6-4e18-94b5-6db5a6357eba · outbound
Accelerating Diffusion-based Text-to-Speech Model Training with Dual Modality Alignment Flow matching for generative modeling,
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8e022587-4e6e-4444-8ab2-63d0c33e41af · outbound
Accelerating Diffusion-based Text-to-Speech Model Training with Dual Modality Alignment Scalable diffusion models with transform- ers,
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bf7e6c63-0eae-4340-ae2d-04e1ecd0ec97 · outbound
Accelerating Diffusion-based Text-to-Speech Model Training with Dual Modality Alignment Efficient diffusion training via min-snr weighting strat- egy,
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 7e709a39-c473-4410-b4e4-20bf79475049 · outbound
Accelerating Diffusion-based Text-to-Speech Model Training with Dual Modality Alignment Fasterdit: Towards faster diffusion transformers training without architecture modification,
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation e9e50506-c6bb-495b-98d0-e26ff9e6bf4d · outbound
Accelerating Diffusion-based Text-to-Speech Model Training with Dual Modality Alignment Representation alignment for generation: Training diffu- sion transformers is easier than you think,
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation c2174071-aecc-4622-8714-9a44830d98e4 · outbound
Accelerating Diffusion-based Text-to-Speech Model Training with Dual Modality Alignment U-net: Convolutional networks for biomedical image segmentation,
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 8d6618dd-846d-4943-9b46-f4f7c3b948a5 · outbound
Accelerating Diffusion-based Text-to-Speech Model Training with Dual Modality Alignment Attention is all you need,
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a3b9ff3e-2fcf-4bf6-b73e-6b2a7743199f · outbound
Accelerating Diffusion-based Text-to-Speech Model Training with Dual Modality Alignment Con- nectionist temporal classification: Labelling unsegmented se- quence data with recurrent neural networks,
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 9f3a332f-2942-4462-a4da-9d66d61fe7de · outbound
Accelerating Diffusion-based Text-to-Speech Model Training with Dual Modality Alignment Intermediate loss regularization for ctc- based speech recognition,
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 64a576d3-3564-423d-a154-bf4847c85770 · outbound
Accelerating Diffusion-based Text-to-Speech Model Training with Dual Modality Alignment Audio-visual efficient conformer for robust speech recognition,
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 12af8411-e6e6-4a18-afc5-e109460d64a5 · outbound
Accelerating Diffusion-based Text-to-Speech Model Training with Dual Modality Alignment Speechtok- enizer: Unified speech tokenizer for speech large language mod- els,
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation dfce062a-4b05-4457-ad8e-f42ebb911f27 · outbound
Accelerating Diffusion-based Text-to-Speech Model Training with Dual Modality Alignment Naturalspeech 3: Zero-shot speech synthesis with factorized codec and diffusion models,
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation afad2a94-9925-4ab6-bb5f-c3a1503d52ce · outbound
Accelerating Diffusion-based Text-to-Speech Model Training with Dual Modality Alignment Hubert: Self-supervised speech represen- tation learning by masked prediction of hidden units,
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 803692c3-015b-466b-8cbe-2b7465842a7f · outbound
Accelerating Diffusion-based Text-to-Speech Model Training with Dual Modality Alignment Wavlm: Large-scale self- supervised pre-training for full stack speech processing,
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 9d8afcab-636c-4abe-8b1d-ef87767ad15e · outbound
Accelerating Diffusion-based Text-to-Speech Model Training with Dual Modality Alignment W2v-bert: Combining contrastive learning and masked language modeling for self-supervised speech pre-training,
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation e3641322-96fe-4ba8-8d41-73fac6cef310 · outbound
Accelerating Diffusion-based Text-to-Speech Model Training with Dual Modality Alignment CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 594d5d61-61ff-4221-becd-05815f436794 · outbound
Accelerating Diffusion-based Text-to-Speech Model Training with Dual Modality Alignment FireRedTTS: A Foundation Text-To-Speech Framework for Industry-Level Generative Speech Applications
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 35f5bfe1-7bfc-40ea-98c9-3231adacf766 · outbound
Accelerating Diffusion-based Text-to-Speech Model Training with Dual Modality Alignment Libritts: A corpus derived from librispeech for text- to-speech,
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 80e28793-afc4-4762-8796-e84d592efcdd · outbound
Accelerating Diffusion-based Text-to-Speech Model Training with Dual Modality Alignment Lib- rispeech: An asr corpus based on public domain audio books,
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 697a294c-cd83-46f6-bcda-2ebe01d793fb · outbound
Accelerating Diffusion-based Text-to-Speech Model Training with Dual Modality Alignment Convnext v2: Co-designing and scaling convnets with masked autoencoders,
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 53c028ee-ac48-4ef9-be5f-9f625d05ba04 · outbound
Accelerating Diffusion-based Text-to-Speech Model Training with Dual Modality Alignment Decoupled weight decay regulariza- tion,
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f1a62001-924d-4777-95bb-7a89c40b6917 · outbound
Accelerating Diffusion-based Text-to-Speech Model Training with Dual Modality Alignment V ocos: Closing the gap between time-domain and fourier-based neural vocoders for high-quality audio synthesis,
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 03fb12a6-8584-435b-a51b-19c107223dbd · outbound
Accelerating Diffusion-based Text-to-Speech Model Training with Dual Modality Alignment Librispeech-pc: Benchmark for evaluation of punctuation and capitalization capabilities of end-to-end asr mod- els,
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation f1d8e692-c73e-4f5f-b7e5-6a6bb1df7fa6 · outbound
Accelerating Diffusion-based Text-to-Speech Model Training with Dual Modality Alignment Robust speech recognition via large-scale weak su- pervision,
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 9292437f-a354-4340-a186-9c6b2b4d6cf5 · outbound
Accelerating Diffusion-based Text-to-Speech Model Training with Dual Modality Alignment Utmos: Utokyo-sarulab system for voicemos challenge 2022,
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 7613c17a-498e-42a3-8d17-dd44120f1d7c · inbound
Accelerating Diffusion-based Text-to-Speech Model Training with Dual Modality Alignment Accelerating Diffusion-based Text-to-Speech Model Training with Dual Modality Alignment
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.