Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 22 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 68 inbound Pith citation observations for arXiv:2006.04558.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-16T11:24:26.891680Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-07-08T18:15:21.399873Z
0 of 0 outbound references displayed
External citation measurements
No source-named external measurement is stored.
No outbound reference observations are available for this paper version.
Observation e3ced301-f751-4fa9-a4b5-f08d66b409ff · inbound
Seed-TTS: A Family of High-Quality Versatile Speech Generation Models FastSpeech 2: Fast and High-Quality End-to-End Text to Speech
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation f9c429f0-e244-418f-b4c4-40b9260a9347 · inbound
F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching FastSpeech 2: Fast and High-Quality End-to-End Text to Speech
Reference 134
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 7d176ddb-9b26-4ff6-935c-4a9d5e2baec5 · inbound
I2TTS: Image-indicated Immersive Text-to-speech Synthesis with Spatial Perception FastSpeech 2: Fast and High-Quality End-to-End Text to Speech
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aab0afda-e84e-46d6-8d38-4f2dba8dd974 · inbound
WavChat: A Survey of Spoken Dialogue Models FastSpeech 2: Fast and High-Quality End-to-End Text to Speech
Reference 177
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ff9c97c0-aaa2-497e-b7b9-dd602b6bb6bc · inbound
Analytic Study of Text-Free Speech Synthesis for Raw Audio using a Self-Supervised Learning Model FastSpeech 2: Fast and High-Quality End-to-End Text to Speech
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2ec1ae0d-b63b-44fb-a6c7-f2372c92b755 · inbound
Aligner-Guided Training Paradigm: Advancing Text-to-Speech Models with Aligner Guided Duration FastSpeech 2: Fast and High-Quality End-to-End Text to Speech
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 46ba7a1e-1aaf-4d62-9f4f-a52768b8babd · inbound
LatentSpeech: Latent Diffusion for Text-To-Speech Generation FastSpeech 2: Fast and High-Quality End-to-End Text to Speech
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4c17a08f-eb8d-4e28-a643-276f24aa57c7 · inbound
YingSound: Video-Guided Sound Effects Generation with Multi-modal Chain-of-Thought Controls FastSpeech 2: Fast and High-Quality End-to-End Text to Speech
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2ebe90da-d97e-4d12-a128-be2ba4309274 · inbound
Libri2Vox Dataset: Target Speaker Extraction with Diverse Speaker Conditions and Synthetic Data FastSpeech 2: Fast and High-Quality End-to-End Text to Speech
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8bd7cb49-60bb-4b7c-af54-203db6ed751b · inbound
Towards Expressive Video Dubbing with Multiscale Multimodal Context Interaction FastSpeech 2: Fast and High-Quality End-to-End Text to Speech
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 68c217f8-b477-4d00-a6f6-1b230cea27e9 · inbound
FaceSpeak: Expressive and High-Quality Speech Synthesis from Human Portraits of Different Styles FastSpeech 2: Fast and High-Quality End-to-End Text to Speech
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 304b3469-c214-49b8-b21e-bf8811bbf1c3 · inbound
PROEMO: Prompt-Driven Text-to-Speech Synthesis Based on Emotion and Intensity Control FastSpeech 2: Fast and High-Quality End-to-End Text to Speech
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 25f1d255-6d72-4087-87ff-f4a4ff22e5fa · inbound
A Non-autoregressive Model for Joint STT and TTS FastSpeech 2: Fast and High-Quality End-to-End Text to Speech
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b3401b66-54ee-4800-a4b7-f3fed282a958 · inbound
VisualSpeech: Enhancing Prosody Modeling in TTS Using Video FastSpeech 2: Fast and High-Quality End-to-End Text to Speech
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 30595c40-9843-41ae-99a8-e0c97676db0f · inbound
Continuous Autoregressive Modeling with Stochastic Monotonic Alignment for Speech Synthesis FastSpeech 2: Fast and High-Quality End-to-End Text to Speech
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cfe4fe30-e4b7-4e03-be0a-e17f861970d9 · inbound
Metis: A Foundation Speech Generation Model with Masked Generative Pre-training FastSpeech 2: Fast and High-Quality End-to-End Text to Speech
Reference 78
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c19b05be-bb84-4943-a068-019d9501fb15 · inbound
VocalCrypt: Novel Active Defense Against Deepfake Voice Based on Masking Effect FastSpeech 2: Fast and High-Quality End-to-End Text to Speech
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 275a9f07-8958-4fb7-8499-07f61f28cec4 · inbound
FADEL: Uncertainty-aware Fake Audio Detection with Evidential Deep Learning FastSpeech 2: Fast and High-Quality End-to-End Text to Speech
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4c31c211-b8d0-40f6-8ce2-44ce49450fd3 · inbound
Perceptual implications of automatic anonymization in pathological speech FastSpeech 2: Fast and High-Quality End-to-End Text to Speech
Reference 89
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation ea6bd5fc-2563-43c4-949b-1e27e10d4eeb · inbound
A Multi-Agent AI Framework for Immersive Audiobook Production through Spatial Audio and Neural Narration FastSpeech 2: Fast and High-Quality End-to-End Text to Speech
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e6c22446-75de-408e-9bdf-9c1dddaee454 · inbound
Score-Based Training for Energy-Based TTS Models FastSpeech 2: Fast and High-Quality End-to-End Text to Speech
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b7501c1f-dd39-4816-9344-aa9aacd8daad · inbound
LLM-based Generative Error Correction for Rare Words with Synthetic Data and Phonetic Context FastSpeech 2: Fast and High-Quality End-to-End Text to Speech
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d013a140-7618-4028-8f5d-28a9f72203de · inbound
MPE-TTS: Customized Emotion Zero-Shot Text-To-Speech Using Multi-Modal Prompt FastSpeech 2: Fast and High-Quality End-to-End Text to Speech
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bd0474c8-cad5-43b2-acc6-e2e05a9ca0d3 · inbound
CloneShield: A Framework for Universal Perturbation Against Zero-Shot Voice Cloning FastSpeech 2: Fast and High-Quality End-to-End Text to Speech
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7fc88f8e-e561-4c13-89a6-35e5a654f82f · inbound
SpeakStream: Streaming Text-to-Speech with Interleaved Data FastSpeech 2: Fast and High-Quality End-to-End Text to Speech
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bf85582c-7e86-4658-a5b5-f5e2cdf7c42d · inbound
RESOUND: Speech Reconstruction from Silent Videos via Acoustic-Semantic Decomposed Modeling FastSpeech 2: Fast and High-Quality End-to-End Text to Speech
Reference 2022
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fc3de9b6-9a65-4842-a3ef-4433dcddfc69 · inbound
Counterfactual Activation Editing for Post-hoc Prosody and Mispronunciation Correction in TTS Models FastSpeech 2: Fast and High-Quality End-to-End Text to Speech
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e0cd1bd0-62d1-4a9c-8e8c-c581427ab987 · inbound
UmbraTTS: Adapting Text-to-Speech to Environmental Contexts with Flow Matching FastSpeech 2: Fast and High-Quality End-to-End Text to Speech
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 762799b2-c25c-48d3-be60-9f847c2a8a4d · inbound
VS-Singer: Vision-Guided Stereo Singing Voice Synthesis with Consistency Schr\"odinger Bridge FastSpeech 2: Fast and High-Quality End-to-End Text to Speech
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d5e5367b-6e78-4ca4-b051-14f110120e45 · inbound
InstructTTSEval: Benchmarking Complex Natural-Language Instruction Following in Text-to-Speech Systems FastSpeech 2: Fast and High-Quality End-to-End Text to Speech
Reference 2021
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 51258cdc-1a04-47bc-b851-77e66cfb4e41 · inbound
SmoothSinger: A Conditional Diffusion Model for Singing Voice Synthesis with Multi-Resolution Architecture FastSpeech 2: Fast and High-Quality End-to-End Text to Speech
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b5600eec-68cb-4ea2-b308-f9cfa7d32d1f · inbound
JAM-Flow: Joint Audio-Motion Synthesis with Flow Matching FastSpeech 2: Fast and High-Quality End-to-End Text to Speech
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 8341b7aa-fa69-4608-9424-d0d851b6f797 · inbound
Collecting, Curating, and Annotating Good Quality Speech deepfake dataset for Famous Figures: Process and Challenges FastSpeech 2: Fast and High-Quality End-to-End Text to Speech
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5402b3f9-b82d-4aa5-aab3-8a802652788f · inbound
RepeaTTS: Towards Feature Discovery through Repeated Fine-Tuning FastSpeech 2: Fast and High-Quality End-to-End Text to Speech
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 628de9bd-5f27-43bd-ab8d-750b22b24240 · inbound
Quantize More, Lose Less: Autoregressive Generation from Residually Quantized Speech Representations FastSpeech 2: Fast and High-Quality End-to-End Text to Speech
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 133feca1-a02c-4ef2-a45a-904b8aa3e946 · inbound
Step-Audio 2 Technical Report FastSpeech 2: Fast and High-Quality End-to-End Text to Speech
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation dbec8d13-d99a-4ef8-bbec-77b4844aaab6 · inbound
Technical report: Impact of Duration Prediction on Speaker-specific TTS for Indian Languages FastSpeech 2: Fast and High-Quality End-to-End Text to Speech
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2b9a4059-25c8-49d5-9df7-ee5e831806cc · inbound
Generative Model Unlearning: A Survey through Target Events, Unlearning Operators, and Evaluation Protocols FastSpeech 2: Fast and High-Quality End-to-End Text to Speech
Reference 185
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 10183caa-b808-4807-92b6-767070d9fb8e · inbound
Adaptive Duration Model for Text Speech Alignment FastSpeech 2: Fast and High-Quality End-to-End Text to Speech
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 68004d0d-dddb-45cc-b8d5-53f7c625913f · inbound
Inference-time Scaling for Diffusion-based Audio Super-resolution FastSpeech 2: Fast and High-Quality End-to-End Text to Speech
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 747a176c-8695-4ef4-a090-5eede55c3179 · inbound
XEmoRAG: Cross-Lingual Emotion Transfer with Controllable Intensity Using Retrieval-Augmented Generation FastSpeech 2: Fast and High-Quality End-to-End Text to Speech
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7d0148a8-a446-4da9-8b35-6f0dd695f5e8 · inbound
SwiftF0: Fast and Accurate Monophonic Pitch Detection FastSpeech 2: Fast and High-Quality End-to-End Text to Speech
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1257031f-77b7-42fc-a4b6-f15c4431988c · inbound
MixedG2P-T5: G2P-free Speech Synthesis for Mixed-script texts using Speech Self-Supervised Learning and Language Model FastSpeech 2: Fast and High-Quality End-to-End Text to Speech
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e9e67638-e956-49a0-bc13-2c90fd26f737 · inbound
OLaPh: Optimal Language Phonemizer FastSpeech 2: Fast and High-Quality End-to-End Text to Speech
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 80ba89bc-1bef-46a9-97a6-e5657b0b5713 · inbound
AUDDT: A Unified Benchmark Toolkit for Audio and Speech Deepfake Detectors FastSpeech 2: Fast and High-Quality End-to-End Text to Speech
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7dac97be-17ba-447e-ad96-257d4e781f18 · inbound
WhisperVC: Decoupled Cross-Domain Alignment and Speech Generation for Low-Resource Whisper-to-Normal Conversion FastSpeech 2: Fast and High-Quality End-to-End Text to Speech
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e2bb892b-6864-48fd-8c99-9ea5d662fef8 · inbound
Intelligent Agents with Emotional Intelligence: Current Trends, Challenges, and Future Prospects FastSpeech 2: Fast and High-Quality End-to-End Text to Speech
Reference 140
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 795dc8f7-e99f-4731-a327-18aba2d1db48 · inbound
OmniCustom: Sync Audio-Video Customization Via Joint Audio-Video Generation Model FastSpeech 2: Fast and High-Quality End-to-End Text to Speech
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 759b864d-0c3b-408b-9987-b89f0e364d76 · inbound
MIDI-Informed Singing Accompaniment Generation in a Compositional Song Pipeline FastSpeech 2: Fast and High-Quality End-to-End Text to Speech
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 8d0d5933-a97a-415d-99c4-b9cfd8414a70 · inbound
Evolution Strategy-Based Calibration for Low-Bit Quantization of Speech Models FastSpeech 2: Fast and High-Quality End-to-End Text to Speech
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6c288869-8085-4a6a-8da6-d7d552cceac1 · inbound
AST: Adaptive, Seamless, and Training-Free Precise Speech Editing FastSpeech 2: Fast and High-Quality End-to-End Text to Speech
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 9bf430db-384e-425a-a788-69cf379c3eb3 · inbound
AST: Adaptive, Seamless, and Training-Free Precise Speech Editing FastSpeech 2: Fast and High-Quality End-to-End Text to Speech
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8c7c3b5d-04c4-40d9-9054-e0883607fd77 · inbound
Voice Mapping of Text-to-Speech Systems: A Metric-Based Approach for Voice Quality Assessment FastSpeech 2: Fast and High-Quality End-to-End Text to Speech
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 7d3b91d2-5017-4284-a5a8-6182bb24e22d · inbound
MelShield: Robust Mel-Domain Audio Watermarking for Provenance Attribution of AI Generated Synthesized Speech FastSpeech 2: Fast and High-Quality End-to-End Text to Speech
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 82fed65f-3184-4ec3-8c2e-3ca5adf453b1 · inbound
HapticLDM: A Diffusion Model for Text-to-Vibrotactile Generation FastSpeech 2: Fast and High-Quality End-to-End Text to Speech
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation d04f7cec-2598-4d24-b07d-362ee582d878 · inbound
N\"ushuVoice: Reviving the Voice of Endangered N\"ushu with Pitch-Aware Text-to-Speech FastSpeech 2: Fast and High-Quality End-to-End Text to Speech
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation abb0c836-e6d8-4c3e-b29c-9ba1b2ef9c9f · inbound
FineCombo-TTS: Collaborative and Precise Controllable Speech Synthesis Using Text Descriptions and Reference Speech FastSpeech 2: Fast and High-Quality End-to-End Text to Speech
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 8a76256e-c556-48ec-b6c9-e65bce9e280c · inbound
DisSpeech: Low-Resource Controllable Mandarin Stuttered Speech Synthesis for ASR Augmentation FastSpeech 2: Fast and High-Quality End-to-End Text to Speech
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation f4de4787-a61c-49b9-8988-98a03e38c516 · inbound
Learning to Evade: Adaptive Attacks on Audio Watermarking FastSpeech 2: Fast and High-Quality End-to-End Text to Speech
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 6f0c767e-c6a4-4b61-86fb-329d8473eeb5 · inbound
MeloDISinger: Melody-Aware & Duration-Preserving Singing Voice Editing with Audio Infilling FastSpeech 2: Fast and High-Quality End-to-End Text to Speech
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 401d0283-227c-44fd-924b-d3115d05a1fd · inbound
FlexiSLM: A Dynamic and Controllable Frame Rate Spoken Language Model FastSpeech 2: Fast and High-Quality End-to-End Text to Speech
Reference 160
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 6fa43e09-9a85-4bbf-9982-bcb5baabf37d · inbound
BlueMagpie-TTS: A Token-Efficient Tokenizer, Language Model, and TTS for Taiwanese-Accent Code-Switching Speech FastSpeech 2: Fast and High-Quality End-to-End Text to Speech
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 8b525661-3042-45e4-ad7a-d4be247ee97a · inbound
AutoSIFT: Automatic Style Sifting for Controllable Speech Generation with Arbitrary Style Infilling FastSpeech 2: Fast and High-Quality End-to-End Text to Speech
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a0d3e1bf-75f0-4f90-b74b-e46458c72f28 · inbound
StellarTTS: Sparse Temporal Embedding for Low-Latency and Robust Speech Synthesis FastSpeech 2: Fast and High-Quality End-to-End Text to Speech
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d0c683b4-48b1-4d44-8145-b7b42e8bd38a · inbound
Designed Vocalizations Dataset: Sound-Designed Human and Animal Voices for Non-human Voice Conversion FastSpeech 2: Fast and High-Quality End-to-End Text to Speech
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b241bb6a-636c-4942-a9b6-e30e605fe28d · inbound
Let Me Look at You: Advanced Facial Expression Modeling for Conversational Speech Synthesis FastSpeech 2: Fast and High-Quality End-to-End Text to Speech
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b272a6a6-74fa-42e1-a810-a34233ee9888 · inbound
Beyond One-Size-Fits-All: Personalized and Culturally Adaptive Emotional TTS via Interactive Optimization of Individual Emotion Perception Spaces FastSpeech 2: Fast and High-Quality End-to-End Text to Speech
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b3081ec9-7616-4677-9e7a-06de3c16e19b · inbound
SemBridge: Semantic Token Anchoring for Continuous-Latent Autoregressive Speech Generation FastSpeech 2: Fast and High-Quality End-to-End Text to Speech
Reference 119
Source-reported events for the cited work
Unavailable: canonical work link unavailable.