Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 22 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 40 inbound Pith citation observations for arXiv:2306.00814.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-16T00:37:04.982405Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-05T02:28:24.338817Z
0 of 0 outbound references displayed
External citation measurements
11
pith, observed 2026-08-05T02:28:24.338817Z
No outbound reference observations are available for this paper version.
Observation a45dbe55-9889-4076-9708-b7e7239e9345 · inbound
F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching Vocos: Closing the gap between time-domain and Fourier-based neural vocoders for high-quality audio synthesis
Reference 140
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 223fd8ff-ccae-411d-ad3f-1723c05de0e4 · inbound
ESTVocoder: An Excitation-Spectral-Transformed Neural Vocoder Conditioned on Mel Spectrogram Vocos: Closing the gap between time-domain and Fourier-based neural vocoders for high-quality audio synthesis
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ad3737eb-3bb6-48d8-bc61-6ccb7e8432de · inbound
A Neural Denoising Vocoder for Clean Waveform Generation from Noisy Mel-Spectrogram based on Amplitude and Phase Predictions Vocos: Closing the gap between time-domain and Fourier-based neural vocoders for high-quality audio synthesis
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 82e50c85-551c-4381-b35f-509eeb55e5c5 · inbound
TS3-Codec: Transformer-Based Simple Streaming Single Codec Vocos: Closing the gap between time-domain and Fourier-based neural vocoders for high-quality audio synthesis
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation acc77a20-f55a-4746-86b1-74de8fce3ff3 · inbound
GenSE: Generative Speech Enhancement via Language Models using Hierarchical Modeling Vocos: Closing the gap between time-domain and Fourier-based neural vocoders for high-quality audio synthesis
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7d5ca28d-17ae-426b-a5b3-3936b1de4e81 · inbound
Metis: A Foundation Speech Generation Model with Masked Generative Pre-training Vocos: Closing the gap between time-domain and Fourier-based neural vocoders for high-quality audio synthesis
Reference 80
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1098c571-dcdf-4bfb-a94c-32ebe44b8a97 · inbound
OZSpeech: One-step Zero-shot Speech Synthesis with Learned-Prior-Conditioned Flow Matching Vocos: Closing the gap between time-domain and Fourier-based neural vocoders for high-quality audio synthesis
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e3045518-e5d9-4e24-b1e7-6bfbecd78b09 · inbound
FlowTSE: Target Speaker Extraction with Flow Matching Vocos: Closing the gap between time-domain and Fourier-based neural vocoders for high-quality audio synthesis
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fa95a800-6595-417b-af9a-f94e5bea44ef · inbound
From Tens of Hours to Tens of Thousands: Scaling Back-Translation for Speech Recognition Vocos: Closing the gap between time-domain and Fourier-based neural vocoders for high-quality audio synthesis
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d35f41a3-464e-4842-b976-c54aa9b8cf8e · inbound
FlowSE: Efficient and High-Quality Speech Enhancement via Flow Matching Vocos: Closing the gap between time-domain and Fourier-based neural vocoders for high-quality audio synthesis
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5a4498b8-5f69-4669-bc9a-9b2c2b66c83d · inbound
Mel-McNet: A Mel-Scale Framework for Online Multichannel Speech Enhancement Vocos: Closing the gap between time-domain and Fourier-based neural vocoders for high-quality audio synthesis
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f7213156-5f1e-49bc-937f-5e62869b6ad6 · inbound
DS-Codec: Dual-Stage Training with Mirror-to-NonMirror Architecture Switching for Speech Codec Vocos: Closing the gap between time-domain and Fourier-based neural vocoders for high-quality audio synthesis
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eeacbfb5-38ea-4628-8ef8-d397f3701c09 · inbound
MagiCodec: Simple Masked Gaussian-Injected Codec for High-Fidelity Reconstruction and Generation Vocos: Closing the gap between time-domain and Fourier-based neural vocoders for high-quality audio synthesis
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b55dffe9-05b2-41cb-bde1-7f3dcca9ea0b · inbound
Efficient Speech Enhancement via Embeddings from Pre-trained Generative Audioencoders Vocos: Closing the gap between time-domain and Fourier-based neural vocoders for high-quality audio synthesis
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 513f7c56-5181-4066-ab7d-01fffc8e987d · inbound
SpeechRefiner: Towards Perceptual Quality Refinement for Front-End Algorithms Vocos: Closing the gap between time-domain and Fourier-based neural vocoders for high-quality audio synthesis
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 16507183-20b4-4104-8119-91ffd070184e · inbound
XY-Tokenizer: Mitigating the Semantic-Acoustic Conflict in Low-Bitrate Speech Codecs Vocos: Closing the gap between time-domain and Fourier-based neural vocoders for high-quality audio synthesis
Reference 2001
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 18874d1a-a21c-4683-8636-141d30755463 · inbound
Traceable TTS: Toward Watermark-Free TTS with Strong Traceability Vocos: Closing the gap between time-domain and Fourier-based neural vocoders for high-quality audio synthesis
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5b2c4a59-bc6d-40f7-8ffb-889dc66184bc · inbound
Autoregressive Speech Enhancement via Acoustic Tokens Vocos: Closing the gap between time-domain and Fourier-based neural vocoders for high-quality audio synthesis
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 676782b7-0ac9-4013-aa6c-2f897b0f37d4 · inbound
DMOSpeech 2: Reinforcement Learning for Duration Prediction in Metric-Optimized Speech Synthesis Vocos: Closing the gap between time-domain and Fourier-based neural vocoders for high-quality audio synthesis
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9eb0fb14-fa1f-4604-a4f8-8687aee6342d · inbound
HH-Codec: High Compression High-fidelity Discrete Neural Codec for Spoken Language Modeling Vocos: Closing the gap between time-domain and Fourier-based neural vocoders for high-quality audio synthesis
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 68a7a26e-2885-4c50-9e47-7653efdfdfa9 · inbound
SenSE: Semantic-Aware High-Fidelity Universal Speech Enhancement Vocos: Closing the gap between time-domain and Fourier-based neural vocoders for high-quality audio synthesis
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 62936d30-161f-4588-b923-eb3c1dbe8c65 · inbound
Two-Dimensional Quantization for Geometry-Aware Audio Coding Vocos: Closing the gap between time-domain and Fourier-based neural vocoders for high-quality audio synthesis
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 178320c8-2265-497d-8752-ec691fc40ac4 · inbound
Woosh: A Sound Effects Foundation Model Vocos: Closing the gap between time-domain and Fourier-based neural vocoders for high-quality audio synthesis
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation a0c32f32-2706-4912-86a0-4f01aa865198 · inbound
UniPASE: A Generative Model for Universal Speech Enhancement with High Fidelity and Low Hallucinations Vocos: Closing the gap between time-domain and Fourier-based neural vocoders for high-quality audio synthesis
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 9a992810-63b1-45e6-81b4-95ddc09f16ff · inbound
UniPASE: A Generative Model for Universal Speech Enhancement with High Fidelity and Low Hallucinations Vocos: Closing the gap between time-domain and Fourier-based neural vocoders for high-quality audio synthesis
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 311e960d-e47a-4882-9ac5-9415515192ac · inbound
WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling Vocos: Closing the gap between time-domain and Fourier-based neural vocoders for high-quality audio synthesis
Reference 72
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 648201a4-36d9-4773-a494-e62e2afccb2f · inbound
BareWave: Waveform-Native Flow-Matching Text-to-Speech Vocos: Closing the gap between time-domain and Fourier-based neural vocoders for high-quality audio synthesis
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 99123b19-0579-4663-b5b5-e39575b762bb · inbound
Afrispeech Semantics: Evaluating Audio Semantic Reasoning in Spoken Language Models Across Domains and Accents Vocos: Closing the gap between time-domain and Fourier-based neural vocoders for high-quality audio synthesis
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 6c09b5b2-5213-461a-bda2-ccf89aa39fb3 · inbound
SARA: A Dual-Stream VAE for High-Fidelity Speech Generation via Integrating Semantic and Acoustic Representations Vocos: Closing the gap between time-domain and Fourier-based neural vocoders for high-quality audio synthesis
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 8dec23c5-03d7-4633-a10e-3d92de0d2bad · inbound
PhASE-Flow: Phonetic-Conditioned Acoustic Flow Matching in SSL Representation Domain for Speech Enhancement Vocos: Closing the gap between time-domain and Fourier-based neural vocoders for high-quality audio synthesis
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 432d4ae4-8932-40ba-b173-96e958d3901b · inbound
NAC: Neural Action Codec for Vision-Language-Action Models Vocos: Closing the gap between time-domain and Fourier-based neural vocoders for high-quality audio synthesis
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 662649d3-2a3e-464e-aed4-30e1c9418328 · inbound
Synthesizing the Lombard Effect: Multi-Level Control of Speech Clarity and Vocal Effort in TTS Vocos: Closing the gap between time-domain and Fourier-based neural vocoders for high-quality audio synthesis
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation cd3727ec-abf2-4b56-94b8-836d6b00c7f6 · inbound
FlowTTS-GRPO: Online Reinforcement Learning with Multi-Objective Reward Optimization for Flow-Matching Based Text-to-Speech Vocos: Closing the gap between time-domain and Fourier-based neural vocoders for high-quality audio synthesis
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 4c3a6dcd-7094-4446-ad5e-af2b5c4292ab · inbound
FlowTTS-GRPO: Online Reinforcement Learning with Multi-Objective Reward Optimization for Flow-Matching Based Text-to-Speech Vocos: Closing the gap between time-domain and Fourier-based neural vocoders for high-quality audio synthesis
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3d2018f4-9fdd-48aa-9fc2-f2f9f6ddbafc · inbound
FlexiSLM: A Dynamic and Controllable Frame Rate Spoken Language Model Vocos: Closing the gap between time-domain and Fourier-based neural vocoders for high-quality audio synthesis
Reference 116
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 898e4653-2033-4961-a5d5-9e5126652721 · inbound
Unified Audio Intelligence Without Regressing on Text Intelligence Vocos: Closing the gap between time-domain and Fourier-based neural vocoders for high-quality audio synthesis
Reference 167
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 824fa72a-15bd-48fd-ae54-090373c1d9e3 · inbound
Unified Audio Intelligence Without Regressing on Text Intelligence Vocos: Closing the gap between time-domain and Fourier-based neural vocoders for high-quality audio synthesis
Reference 167
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cf042f65-49d9-4778-b6aa-65d59d4f3b99 · inbound
ZipL-Dialog: Memory-Efficient Long-Form Spoken Dialog Synthesis via Latent Flow Matching Vocos: Closing the gap between time-domain and Fourier-based neural vocoders for high-quality audio synthesis
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4922feff-a1cd-4b7e-97b4-377f7cebfe07 · inbound
Staged Depth-Pruning Distillation of a Flow-Matching Text-to-Speech Teacher: A Compact Hindi Speech Synthesizer Vocos: Closing the gap between time-domain and Fourier-based neural vocoders for high-quality audio synthesis
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bef0a40f-7795-4b00-ac2a-a536a6ff8af7 · inbound
Phoenix TTS: High-Fidelity Synthesis and Voice Conversion via Flow-Matching-Driven Speech Tokenization Vocos: Closing the gap between time-domain and Fourier-based neural vocoders for high-quality audio synthesis
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.