Pith. sign in

Paper Citation Record · LEDGER

Vocos: Closing the gap between time-domain and Fourier-based neural vocoders for high-quality audio synthesis

As of 22 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 40 inbound Pith citation observations for arXiv:2306.00814.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2306.00814 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 40 of 40 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 40 of 40 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T00:37:04.982405Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

11
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation a45dbe55-9889-4076-9708-b7e7239e9345 · inbound

F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching cites this paper.

F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching Vocos: Closing the gap between time-domain and Fourier-based neural vocoders for high-quality audio synthesis

Reference 140

Resolution
verified exact
arxiv_id, observed 2026-05-16T06:06:41.521748Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-16T06:06:41.348728Z digest=sha256:b075b134f67512ef8e2a8eee9dc1c8f5ac0d68224eceeebf9d40a12dc493371f

Observation 223fd8ff-ccae-411d-ad3f-1723c05de0e4 · inbound

ESTVocoder: An Excitation-Spectral-Transformed Neural Vocoder Conditioned on Mel Spectrogram cites this paper.

ESTVocoder: An Excitation-Spectral-Transformed Neural Vocoder Conditioned on Mel Spectrogram Vocos: Closing the gap between time-domain and Fourier-based neural vocoders for high-quality audio synthesis

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-12T18:47:51.609553Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:47:51.609553Z digest=sha256:081738fe9dd344ce2d2a5f0ad8259cc07de3ceb0c6edfd945d8ef70ec308f3af

Observation ad3737eb-3bb6-48d8-bc61-6ccb7e8432de · inbound

A Neural Denoising Vocoder for Clean Waveform Generation from Noisy Mel-Spectrogram based on Amplitude and Phase Predictions cites this paper.

A Neural Denoising Vocoder for Clean Waveform Generation from Noisy Mel-Spectrogram based on Amplitude and Phase Predictions Vocos: Closing the gap between time-domain and Fourier-based neural vocoders for high-quality audio synthesis

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-12T17:47:13.037199Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T17:47:13.037199Z digest=sha256:f399b47ea5f906c23b44f1ced45765ab88901bba6e3d4f822c179c041239b94e

Observation 82e50c85-551c-4381-b35f-509eeb55e5c5 · inbound

TS3-Codec: Transformer-Based Simple Streaming Single Codec cites this paper.

TS3-Codec: Transformer-Based Simple Streaming Single Codec Vocos: Closing the gap between time-domain and Fourier-based neural vocoders for high-quality audio synthesis

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-12T10:57:08.526943Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T10:57:08.526943Z digest=sha256:b226ad984c7bddc81c0488fe5f993caf6fbca52fbf0e8a40d260bf834d09533f

Observation acc77a20-f55a-4746-86b1-74de8fce3ff3 · inbound

GenSE: Generative Speech Enhancement via Language Models using Hierarchical Modeling cites this paper.

GenSE: Generative Speech Enhancement via Language Models using Hierarchical Modeling Vocos: Closing the gap between time-domain and Fourier-based neural vocoders for high-quality audio synthesis

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-09T10:37:03.701969Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T10:37:03.701969Z digest=sha256:e94d89a575cdc86bdda275966e218c44e2f84c07535654ca041aedab0d7bfb9e

Observation 7d5ca28d-17ae-426b-a5b3-3936b1de4e81 · inbound

Metis: A Foundation Speech Generation Model with Masked Generative Pre-training cites this paper.

Metis: A Foundation Speech Generation Model with Masked Generative Pre-training Vocos: Closing the gap between time-domain and Fourier-based neural vocoders for high-quality audio synthesis

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-09T05:54:18.731349Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T05:54:18.731349Z digest=sha256:ce2d2bcea9608e2aa8bcc7578a0f803a227bee7515648fc45419150a6dfa81c8

Observation 1098c571-dcdf-4bfb-a94c-32ebe44b8a97 · inbound

OZSpeech: One-step Zero-shot Speech Synthesis with Learned-Prior-Conditioned Flow Matching cites this paper.

OZSpeech: One-step Zero-shot Speech Synthesis with Learned-Prior-Conditioned Flow Matching Vocos: Closing the gap between time-domain and Fourier-based neural vocoders for high-quality audio synthesis

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-15T20:30:50.836022Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:30:50.836022Z digest=sha256:e14f2e7e45f613e2b9aaeba6a86b63561b086e5d2129e215e768123fd3440f81

Observation e3045518-e5d9-4e24-b1e7-6bfbecd78b09 · inbound

FlowTSE: Target Speaker Extraction with Flow Matching cites this paper.

FlowTSE: Target Speaker Extraction with Flow Matching Vocos: Closing the gap between time-domain and Fourier-based neural vocoders for high-quality audio synthesis

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T15:37:09.616746Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:37:09.616746Z digest=sha256:9ff20d341cd74a3885c23cdb37316aebe92f820c4fadccb9546de205677c1522

Observation fa95a800-6595-417b-af9a-f94e5bea44ef · inbound

From Tens of Hours to Tens of Thousands: Scaling Back-Translation for Speech Recognition cites this paper.

From Tens of Hours to Tens of Thousands: Scaling Back-Translation for Speech Recognition Vocos: Closing the gap between time-domain and Fourier-based neural vocoders for high-quality audio synthesis

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:57.198647Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:55:57.198647Z digest=sha256:887a3c341660a8246fe7cd9459b6c1399e432365531693858537a98fe5b63e34

Observation d35f41a3-464e-4842-b976-c54aa9b8cf8e · inbound

FlowSE: Efficient and High-Quality Speech Enhancement via Flow Matching cites this paper.

FlowSE: Efficient and High-Quality Speech Enhancement via Flow Matching Vocos: Closing the gap between time-domain and Fourier-based neural vocoders for high-quality audio synthesis

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T14:18:26.198033Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:18:26.198033Z digest=sha256:608849c6c7e87fbf35807edfa525106b7164cb60c7069be457406cecad63d450

Observation 5a4498b8-5f69-4669-bc9a-9b2c2b66c83d · inbound

Mel-McNet: A Mel-Scale Framework for Online Multichannel Speech Enhancement cites this paper.

Mel-McNet: A Mel-Scale Framework for Online Multichannel Speech Enhancement Vocos: Closing the gap between time-domain and Fourier-based neural vocoders for high-quality audio synthesis

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T14:15:18.545648Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:15:18.545648Z digest=sha256:a6fff5500830861f350800fa084d12bce38df107fe1bd07835b3ce39336c5dd7

Observation f7213156-5f1e-49bc-937f-5e62869b6ad6 · inbound

DS-Codec: Dual-Stage Training with Mirror-to-NonMirror Architecture Switching for Speech Codec cites this paper.

DS-Codec: Dual-Stage Training with Mirror-to-NonMirror Architecture Switching for Speech Codec Vocos: Closing the gap between time-domain and Fourier-based neural vocoders for high-quality audio synthesis

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T12:29:42.327056Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:29:42.327056Z digest=sha256:cd1e40d9691b5a24f8ed548c82f2b87283a45e2c566ea89f22175fbe2fdde9ce

Observation eeacbfb5-38ea-4628-8ef8-d397f3701c09 · inbound

MagiCodec: Simple Masked Gaussian-Injected Codec for High-Fidelity Reconstruction and Generation cites this paper.

MagiCodec: Simple Masked Gaussian-Injected Codec for High-Fidelity Reconstruction and Generation Vocos: Closing the gap between time-domain and Fourier-based neural vocoders for high-quality audio synthesis

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T12:12:29.406719Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:12:29.406719Z digest=sha256:953fafcb96845600ad7275b66e37083e93cdd3a86531a9da8741bd1313dff026

Observation b55dffe9-05b2-41cb-bde1-7f3dcca9ea0b · inbound

Efficient Speech Enhancement via Embeddings from Pre-trained Generative Audioencoders cites this paper.

Efficient Speech Enhancement via Embeddings from Pre-trained Generative Audioencoders Vocos: Closing the gap between time-domain and Fourier-based neural vocoders for high-quality audio synthesis

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:06.083636Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:06.083636Z digest=sha256:31eb45c1a2d0ca8ef2ecbf9c80d48538f7f60b2445c143bdfb053d74f6b40622

Observation 513f7c56-5181-4066-ab7d-01fffc8e987d · inbound

SpeechRefiner: Towards Perceptual Quality Refinement for Front-End Algorithms cites this paper.

SpeechRefiner: Towards Perceptual Quality Refinement for Front-End Algorithms Vocos: Closing the gap between time-domain and Fourier-based neural vocoders for high-quality audio synthesis

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T00:30:35.396683Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:30:35.396683Z digest=sha256:013307a884186ba8dbe0aa847cb8068b5c94c1c6491eea9bac9fe11e4428ef99

Observation 16507183-20b4-4104-8119-91ffd070184e · inbound

XY-Tokenizer: Mitigating the Semantic-Acoustic Conflict in Low-Bitrate Speech Codecs cites this paper.

XY-Tokenizer: Mitigating the Semantic-Acoustic Conflict in Low-Bitrate Speech Codecs Vocos: Closing the gap between time-domain and Fourier-based neural vocoders for high-quality audio synthesis

Reference 2001

Resolution
unresolved
no resolver link, observed 2026-08-06T21:51:35.862909Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:51:35.862909Z digest=sha256:63fe8615a32e39c21e90216848f388d70de092122801bb607b29c2f952dceebf

Observation 18874d1a-a21c-4683-8636-141d30755463 · inbound

Traceable TTS: Toward Watermark-Free TTS with Strong Traceability cites this paper.

Traceable TTS: Toward Watermark-Free TTS with Strong Traceability Vocos: Closing the gap between time-domain and Fourier-based neural vocoders for high-quality audio synthesis

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T20:06:18.832332Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:06:18.832332Z digest=sha256:d62c9f31341ad3796685eb421809ffc15133723cc1c41ef36cb36e7770663d8b

Observation 5b2c4a59-bc6d-40f7-8ffb-889dc66184bc · inbound

Autoregressive Speech Enhancement via Acoustic Tokens cites this paper.

Autoregressive Speech Enhancement via Acoustic Tokens Vocos: Closing the gap between time-domain and Fourier-based neural vocoders for high-quality audio synthesis

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-06T16:41:24.295671Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:41:24.295671Z digest=sha256:b22e238645bd1f7de3152ffab92c048c6087cb9f429c4cd5c06f1315760754e0

Observation 676782b7-0ac9-4013-aa6c-2f897b0f37d4 · inbound

DMOSpeech 2: Reinforcement Learning for Duration Prediction in Metric-Optimized Speech Synthesis cites this paper.

DMOSpeech 2: Reinforcement Learning for Duration Prediction in Metric-Optimized Speech Synthesis Vocos: Closing the gap between time-domain and Fourier-based neural vocoders for high-quality audio synthesis

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:22.435245Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:48:22.435245Z digest=sha256:7804347e938f116f010015cdff58018be78b306103dfc4cfab61cc8acdd54853

Observation 9eb0fb14-fa1f-4604-a4f8-8687aee6342d · inbound

HH-Codec: High Compression High-fidelity Discrete Neural Codec for Spoken Language Modeling cites this paper.

HH-Codec: High Compression High-fidelity Discrete Neural Codec for Spoken Language Modeling Vocos: Closing the gap between time-domain and Fourier-based neural vocoders for high-quality audio synthesis

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-15T18:11:55.375110Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:11:55.375110Z digest=sha256:682a4af98c7fa55036381242d44f22b1cf752092b1b3ffd7ce2ef629f3a54b95

Observation 68a7a26e-2885-4c50-9e47-7653efdfdfa9 · inbound

SenSE: Semantic-Aware High-Fidelity Universal Speech Enhancement cites this paper.

SenSE: Semantic-Aware High-Fidelity Universal Speech Enhancement Vocos: Closing the gap between time-domain and Fourier-based neural vocoders for high-quality audio synthesis

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-18T12:36:22.339938Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-18T12:35:19.023045Z digest=sha256:93f69a6d56fef309c95a10c537b1514260bb6c5ef812d9adc4aca02c474b689b

Observation 62936d30-161f-4588-b923-eb3c1dbe8c65 · inbound

Two-Dimensional Quantization for Geometry-Aware Audio Coding cites this paper.

Two-Dimensional Quantization for Geometry-Aware Audio Coding Vocos: Closing the gap between time-domain and Fourier-based neural vocoders for high-quality audio synthesis

Reference 63

Resolution
verified exact
arxiv_id, observed 2026-05-21T18:20:29.314335Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-21T18:16:51.486807Z digest=sha256:d46866117d0b94783a5890f3678f35e10bd4f173ea40097ff7be0be13df2e682

Observation 178320c8-2265-497d-8752-ec691fc40ac4 · inbound

Woosh: A Sound Effects Foundation Model cites this paper.

Woosh: A Sound Effects Foundation Model Vocos: Closing the gap between time-domain and Fourier-based neural vocoders for high-quality audio synthesis

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-13T20:53:15.813464Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-13T20:51:08.144573Z digest=sha256:2cbbc2effe14ac7854886df8cb73061e84175a6389d98360c83691cb84558f5d

Observation a0c32f32-2706-4912-86a0-4f01aa865198 · inbound

UniPASE: A Generative Model for Universal Speech Enhancement with High Fidelity and Low Hallucinations cites this paper.

UniPASE: A Generative Model for Universal Speech Enhancement with High Fidelity and Low Hallucinations Vocos: Closing the gap between time-domain and Fourier-based neural vocoders for high-quality audio synthesis

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-10T10:19:19.976408Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-10T10:18:16.972414Z digest=sha256:08a259c3a28611a69a6c516841af5a82b78416b13e4c532317f51cae0cf0a13a

Observation 9a992810-63b1-45e6-81b4-95ddc09f16ff · inbound

UniPASE: A Generative Model for Universal Speech Enhancement with High Fidelity and Low Hallucinations cites this paper.

UniPASE: A Generative Model for Universal Speech Enhancement with High Fidelity and Low Hallucinations Vocos: Closing the gap between time-domain and Fourier-based neural vocoders for high-quality audio synthesis

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-02T16:16:37.406485Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T16:16:37.406485Z digest=sha256:96fafdbab1cfed787fb9dad7a4c410c660ea38301650dc91d08529b44050b26b

Observation 311e960d-e47a-4882-9ac5-9415515192ac · inbound

WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling cites this paper.

WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling Vocos: Closing the gap between time-domain and Fourier-based neural vocoders for high-quality audio synthesis

Reference 72

Resolution
verified exact
arxiv_id, observed 2026-07-02T05:16:39.843462Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-28T08:18:42.002083Z digest=sha256:723347f3c57fbb38be6b4da8af0fd94021a7ad66ea07ac68f678fa02532cdab5

Observation 648201a4-36d9-4773-a494-e62e2afccb2f · inbound

BareWave: Waveform-Native Flow-Matching Text-to-Speech cites this paper.

BareWave: Waveform-Native Flow-Matching Text-to-Speech Vocos: Closing the gap between time-domain and Fourier-based neural vocoders for high-quality audio synthesis

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-07-03T03:37:35.201518Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-27T15:14:03.681833Z digest=sha256:d1ea967fd0bf47ed5e5f81cbef21133dd669b716063e5782ab18f302f49fbc85

Observation 99123b19-0579-4663-b5b5-e39575b762bb · inbound

Afrispeech Semantics: Evaluating Audio Semantic Reasoning in Spoken Language Models Across Domains and Accents cites this paper.

Afrispeech Semantics: Evaluating Audio Semantic Reasoning in Spoken Language Models Across Domains and Accents Vocos: Closing the gap between time-domain and Fourier-based neural vocoders for high-quality audio synthesis

Reference 59

Resolution
verified exact
arxiv_id, observed 2026-06-30T22:15:05.635697Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-06-30T22:11:44.891731Z digest=sha256:eec52bcd512f8f012444fe9361b50657a8c8061d78bcdc9d4a41e5dfefefe3c6

Observation 6c09b5b2-5213-461a-bda2-ccf89aa39fb3 · inbound

SARA: A Dual-Stream VAE for High-Fidelity Speech Generation via Integrating Semantic and Acoustic Representations cites this paper.

SARA: A Dual-Stream VAE for High-Fidelity Speech Generation via Integrating Semantic and Acoustic Representations Vocos: Closing the gap between time-domain and Fourier-based neural vocoders for high-quality audio synthesis

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-07-03T12:58:08.358364Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-27T08:38:51.371500Z digest=sha256:b61ad408edb90658cc64fa89c9fd98ef266473fcb4a2d5d7e03512f9e0497825

Observation 8dec23c5-03d7-4633-a10e-3d92de0d2bad · inbound

PhASE-Flow: Phonetic-Conditioned Acoustic Flow Matching in SSL Representation Domain for Speech Enhancement cites this paper.

PhASE-Flow: Phonetic-Conditioned Acoustic Flow Matching in SSL Representation Domain for Speech Enhancement Vocos: Closing the gap between time-domain and Fourier-based neural vocoders for high-quality audio synthesis

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-07-03T23:09:00.912005Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-26T23:05:08.894347Z digest=sha256:07685253444a663fa3b299fe42cb1bd6d4d101a338429341981b54cee26a47a8

Observation 432d4ae4-8932-40ba-b173-96e958d3901b · inbound

NAC: Neural Action Codec for Vision-Language-Action Models cites this paper.

NAC: Neural Action Codec for Vision-Language-Action Models Vocos: Closing the gap between time-domain and Fourier-based neural vocoders for high-quality audio synthesis

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-07-04T06:59:37.890291Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-26T13:59:53.484305Z digest=sha256:780b93fe54fe98f5a9e0fa61cfcb2493cc0c3ee404486e05dee2de124cc56dc2

Observation 662649d3-2a3e-464e-aed4-30e1c9418328 · inbound

Synthesizing the Lombard Effect: Multi-Level Control of Speech Clarity and Vocal Effort in TTS cites this paper.

Synthesizing the Lombard Effect: Multi-Level Control of Speech Clarity and Vocal Effort in TTS Vocos: Closing the gap between time-domain and Fourier-based neural vocoders for high-quality audio synthesis

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-07-04T12:29:51.743004Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-26T06:49:45.405554Z digest=sha256:af14b6e7eaa1a6e6bd5d4716756c1bb65cb3490189fdda7636d7f2529d54eff1

Observation cd3727ec-abf2-4b56-94b8-836d6b00c7f6 · inbound

FlowTTS-GRPO: Online Reinforcement Learning with Multi-Objective Reward Optimization for Flow-Matching Based Text-to-Speech cites this paper.

FlowTTS-GRPO: Online Reinforcement Learning with Multi-Objective Reward Optimization for Flow-Matching Based Text-to-Speech Vocos: Closing the gap between time-domain and Fourier-based neural vocoders for high-quality audio synthesis

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-07-04T12:19:49.654721Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-26T07:02:36.499424Z digest=sha256:aa2a9eb6d91f335e13e2456bcf30f6f5994717c59e2930c61b1b69e1149939b4

Observation 4c3a6dcd-7094-4446-ad5e-af2b5c4292ab · inbound

FlowTTS-GRPO: Online Reinforcement Learning with Multi-Objective Reward Optimization for Flow-Matching Based Text-to-Speech cites this paper.

FlowTTS-GRPO: Online Reinforcement Learning with Multi-Objective Reward Optimization for Flow-Matching Based Text-to-Speech Vocos: Closing the gap between time-domain and Fourier-based neural vocoders for high-quality audio synthesis

Reference 36

Resolution
unresolved
no resolver link, observed 2026-07-12T12:44:20.831164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T12:44:20.831164Z digest=sha256:95415f96558e4aeae66ce22a320ad6cb4db820fe36e01ca635a2aaa0c9a87f88

Observation 3d2018f4-9fdd-48aa-9fc2-f2f9f6ddbafc · inbound

FlexiSLM: A Dynamic and Controllable Frame Rate Spoken Language Model cites this paper.

FlexiSLM: A Dynamic and Controllable Frame Rate Spoken Language Model Vocos: Closing the gap between time-domain and Fourier-based neural vocoders for high-quality audio synthesis

Reference 116

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T11:55:42.189973Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-07-01T03:50:26.873406Z digest=sha256:5107f08c059bd290dd3f608ada77eb53a020f15040d5f623e2d568cccae40c99

Observation 898e4653-2033-4961-a5d5-9e5126652721 · inbound

Unified Audio Intelligence Without Regressing on Text Intelligence cites this paper.

Unified Audio Intelligence Without Regressing on Text Intelligence Vocos: Closing the gap between time-domain and Fourier-based neural vocoders for high-quality audio synthesis

Reference 167

Resolution
metadata mismatch
local_arxiv, observed 2026-07-08T00:04:22.674104Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-07-07T23:59:38.702609Z digest=sha256:0dae801fcdba78f471db834b04dd2b129d941502b0bfc8efa768a6b5070ee302

Observation 824fa72a-15bd-48fd-ae54-090373c1d9e3 · inbound

Unified Audio Intelligence Without Regressing on Text Intelligence cites this paper.

Unified Audio Intelligence Without Regressing on Text Intelligence Vocos: Closing the gap between time-domain and Fourier-based neural vocoders for high-quality audio synthesis

Reference 167

Resolution
unresolved
no resolver link, observed 2026-07-11T07:46:49.059192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T07:46:49.059192Z digest=sha256:0cfa54cd57da6a725eefa25d5a2bec04fec507f5cc69ebacf322377f5a435772

Observation cf042f65-49d9-4778-b6aa-65d59d4f3b99 · inbound

ZipL-Dialog: Memory-Efficient Long-Form Spoken Dialog Synthesis via Latent Flow Matching cites this paper.

ZipL-Dialog: Memory-Efficient Long-Form Spoken Dialog Synthesis via Latent Flow Matching Vocos: Closing the gap between time-domain and Fourier-based neural vocoders for high-quality audio synthesis

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-02T06:31:32.258567Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:31:32.258567Z digest=sha256:5ffb0eac244489a95816bbf9f867ea201fbde30421c3b0fd5c9e27d8c6dc8bb5

Observation 4922feff-a1cd-4b7e-97b4-377f7cebfe07 · inbound

Staged Depth-Pruning Distillation of a Flow-Matching Text-to-Speech Teacher: A Compact Hindi Speech Synthesizer cites this paper.

Staged Depth-Pruning Distillation of a Flow-Matching Text-to-Speech Teacher: A Compact Hindi Speech Synthesizer Vocos: Closing the gap between time-domain and Fourier-based neural vocoders for high-quality audio synthesis

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-01T18:46:14.038134Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T18:46:14.038134Z digest=sha256:67b7c1d63d9a286c3f63d55bb87784cefe94221fe15ff5c62f9e417f45c63d4b

Observation bef0a40f-7795-4b00-ac2a-a536a6ff8af7 · inbound

Phoenix TTS: High-Fidelity Synthesis and Voice Conversion via Flow-Matching-Driven Speech Tokenization cites this paper.

Phoenix TTS: High-Fidelity Synthesis and Voice Conversion via Flow-Matching-Driven Speech Tokenization Vocos: Closing the gap between time-domain and Fourier-based neural vocoders for high-quality audio synthesis

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-16T00:37:04.982405Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:37:04.982405Z digest=sha256:b05c3e2cc442977375292c269f00c47cbe4da64d1fddcfe70ad1dee5a3bca7d5