Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-16T05:37:50.085103Z
Paper Citation Record · LEDGER
As of 23 August 2026, this Paper Citation Record lists 36 of 36 outbound references and 2 inbound Pith citation observations for arXiv:2504.20334.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-16T05:37:50.085103Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-05-12T01:43:48.555523Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-12T01:46:13.927285Z
36 of 36 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation b5a93af7-4a77-4f06-acc9-21a5cbde97d0 · outbound
Towards Flow-Matching-based TTS without Classifier-Free Guidance FunAudioLLM: Voice Understanding and Generation Foundation Models for Natural Interaction Between Humans and LLMs
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7e79d1bf-d1da-42a3-8634-00854d067bfa · outbound
Towards Flow-Matching-based TTS without Classifier-Free Guidance CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ab803573-9161-44e9-a68d-ec4f9f1e898f · outbound
Towards Flow-Matching-based TTS without Classifier-Free Guidance VioLA: Unified Codec Language Models for Speech Recognition, Synthesis, and Translation
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 432bd714-4042-4d4c-8fbb-1fde08e92eb6 · outbound
Towards Flow-Matching-based TTS without Classifier-Free Guidance NaturalSpeech 3: Zero-Shot Speech Synthesis with Factorized Codec and Diffusion Models
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation da9754b8-7a3f-49e2-ae31-f88bc3a5ac2e · outbound
Towards Flow-Matching-based TTS without Classifier-Free Guidance E2 TTS: Embarrassingly Easy Fully Non-Autoregressive Zero-Shot TTS
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 725c402b-4271-4d0a-9668-e5e7b11984ec · outbound
Towards Flow-Matching-based TTS without Classifier-Free Guidance F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eb2b5291-506a-4704-bab0-a95331474ca2 · outbound
Towards Flow-Matching-based TTS without Classifier-Free Guidance Matcha- TTS: A fast TTS architecture with conditional flow matching,
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation c3d6e5ad-4bf4-436b-94c9-88e7c7cf771f · outbound
Towards Flow-Matching-based TTS without Classifier-Free Guidance Language mod- els are few-shot learners,
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cf13ed11-3bdb-4ef9-850d-11b37b8169d7 · outbound
Towards Flow-Matching-based TTS without Classifier-Free Guidance Diffusion models beat gans on image synthesis,
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 1faa44f5-9b63-44b2-bcb2-3f798b55384e · outbound
Towards Flow-Matching-based TTS without Classifier-Free Guidance Denoising diffusion probabilistic models,
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6d879e7f-d793-4fdd-813f-f08d12e1e117 · outbound
Towards Flow-Matching-based TTS without Classifier-Free Guidance Denoising Diffusion Implicit Models
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7821f77c-b65a-4c72-a510-a34e4c2d208c · outbound
Towards Flow-Matching-based TTS without Classifier-Free Guidance Improved denoising diffusion probabilis- tic models,
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 36e995a2-b52f-488e-a9aa-289eb79580f9 · outbound
Towards Flow-Matching-based TTS without Classifier-Free Guidance Flow matching for generative modeling,
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 88b36217-93b5-4688-8090-15f435e51147 · outbound
Towards Flow-Matching-based TTS without Classifier-Free Guidance FastDiff: A Fast Conditional Diffusion Model for High-Quality Speech Synthesis
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0dfddaa3-b92d-4651-8ccf-11b176cf83b5 · outbound
Towards Flow-Matching-based TTS without Classifier-Free Guidance ProDiff: Progressive fast diffusion model for high-quality text-to-speech,
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 28eeff79-995a-4718-8680-a967211824a5 · outbound
Towards Flow-Matching-based TTS without Classifier-Free Guidance Diff-TTS: A Denoising Diffusion Model for Text-to-Speech
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8629ada8-baf1-4dc8-be8e-088778faba65 · outbound
Towards Flow-Matching-based TTS without Classifier-Free Guidance CoMoSpeech: One-step speech and singing voice synthesis via consistency model,
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 69059aa1-d9d0-43df-b600-84a239fe3092 · outbound
Towards Flow-Matching-based TTS without Classifier-Free Guidance Reflow- TTS: A rectified flow model for high-fidelity text-to-speech,
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 210e8011-bc56-4a6e-a271-563bb7fb207b · outbound
Towards Flow-Matching-based TTS without Classifier-Free Guidance FlashSpeech: Efficient zero-shot speech synthesis,
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation c5b798eb-04c4-41e4-b6d2-c138d4416705 · outbound
Towards Flow-Matching-based TTS without Classifier-Free Guidance Diffusion Models without Classifier-free Guidance
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 00164ab1-6cdd-41d5-b386-2ba88884d514 · outbound
Towards Flow-Matching-based TTS without Classifier-Free Guidance Classifier-Free Diffusion Guidance
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a09220d1-a841-42d4-9a19-30e7d1bf5cec · outbound
Towards Flow-Matching-based TTS without Classifier-Free Guidance Score-based generative modeling through stochastic differ- ential equations,
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 423d892d-2067-4160-ba2d-65fa3f11d895 · outbound
Towards Flow-Matching-based TTS without Classifier-Free Guidance Improving and generalizing flow-based generative models with minibatch optimal transport
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 784585eb-55bd-4336-9dec-dc1b837f796c · outbound
Towards Flow-Matching-based TTS without Classifier-Free Guidance V oicebox: Text-guided mul- tilingual universal speech generation at scale,
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 98ce7513-ac39-40de-80c6-0dd9c84df21b · outbound
Towards Flow-Matching-based TTS without Classifier-Free Guidance Seed-TTS: A Family of High-Quality Versatile Speech Generation Models
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bdad7f93-a227-4a8c-b6d5-1943da6d15a4 · outbound
Towards Flow-Matching-based TTS without Classifier-Free Guidance FireRedTTS: A Foundation Text-To-Speech Framework for Industry-Level Generative Speech Applications
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0cbb6ed1-5194-4700-92dc-7d086b5965a2 · outbound
Towards Flow-Matching-based TTS without Classifier-Free Guidance Autoregressive Diffusion Transformer for Text-to-Speech Synthesis
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 852afc43-d050-4183-ae6c-628669d09cf7 · outbound
Towards Flow-Matching-based TTS without Classifier-Free Guidance DiTAR: Diffusion Transformer Autoregressive Modeling for Speech Generation,
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a7b95ccc-1d80-48bc-addc-2b338ffdbc6a · outbound
Towards Flow-Matching-based TTS without Classifier-Free Guidance Frieren: Efficient video-to-audio generation network with rectified flow matching,
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 1c8ae97d-9119-4354-a485-507f1924ccbd · outbound
Towards Flow-Matching-based TTS without Classifier-Free Guidance Null- text inversion for editing real images using guided diffusion models,
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 2299b1f6-a57e-4d82-b9fe-7a14181d294f · outbound
Towards Flow-Matching-based TTS without Classifier-Free Guidance CFG++: Manifold-constrained Classifier Free Guidance for Diffusion Models
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f979415a-bb07-4956-93d4-b4d7156e149c · outbound
Towards Flow-Matching-based TTS without Classifier-Free Guidance CFG-Zero*: Improved Classifier-Free Guidance for Flow Matching Models
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 67e21c6f-1cfe-4778-91dc-685514139090 · outbound
Towards Flow-Matching-based TTS without Classifier-Free Guidance Classifier-Free Guidance is a Predictor-Corrector
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 796dea0a-6d1f-4805-9a0c-05c3ccfacecd · outbound
Towards Flow-Matching-based TTS without Classifier-Free Guidance LibriTTS: A Corpus Derived from LibriSpeech for Text-to-Speech
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 73812b68-fb1a-4a83-948a-59b4d20d498b · outbound
Towards Flow-Matching-based TTS without Classifier-Free Guidance Large-scale self-supervised speech representation learning for automatic speaker verification,
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 1926ba2a-b443-45ad-9567-ba7ca98d6dc9 · outbound
Towards Flow-Matching-based TTS without Classifier-Free Guidance NISQA: A Deep CNN-Self-Attention Model for Multidimensional Speech Quality Prediction with Crowdsourced Datasets
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b7dc7364-8203-41b1-95d8-21e77211b331 · inbound
X-Voice: Enabling Everyone to Speak 30 Languages via Zero-Shot Cross-Lingual Voice Cloning Towards Flow-Matching-based TTS without Classifier-Free Guidance
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.
Observation 7963b79a-0732-4403-8abf-362e45720cc5 · inbound
X-Voice: Enabling Everyone to Speak 30 Languages via Zero-Shot Cross-Lingual Voice Cloning Towards Flow-Matching-based TTS without Classifier-Free Guidance
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.