Pith. sign in

Paper Citation Record · LEDGER

Phoenix TTS: High-Fidelity Synthesis and Voice Conversion via Flow-Matching-Driven Speech Tokenization

As of 20 August 2026, this Paper Citation Record lists 41 of 41 outbound references and 0 inbound Pith citation observations for arXiv:2608.11737.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.11737 v1

Coverage vector

measured 41 of 41 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-16T00:37:05.011278Z

measured 41 of 41 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

41 of 41 outbound references displayed

  • verified exact0
  • verified fuzzy5
  • unresolved36
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 6299754f-6941-43b3-9c65-1897d925b67e · outbound

This paper cites Seed-TTS: A Family of High-Quality Versatile Speech Generation Models.

Phoenix TTS: High-Fidelity Synthesis and Voice Conversion via Flow-Matching-Driven Speech Tokenization Seed-TTS: A Family of High-Quality Versatile Speech Generation Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-16T00:37:04.855832Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:37:04.855832Z digest=sha256:54b054eff292dd9ed8afbce7e5dbb381ef0f8420d44728edbc16942d16fa23d9

Observation 24f6b3a3-7e86-4b8a-bf33-72248c77c095 · outbound

This paper cites XTTS: a Massively Multilingual Zero-Shot Text-to-Speech Model.

Phoenix TTS: High-Fidelity Synthesis and Voice Conversion via Flow-Matching-Driven Speech Tokenization XTTS: a Massively Multilingual Zero-Shot Text-to-Speech Model

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-16T00:37:04.860178Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:37:04.860178Z digest=sha256:034d6c2ced2ee212d47880b098b6737e7658cdb2b81e9246a18114cb8034c88c

Observation ef19d6b0-298f-446a-838c-2426a92e094b · outbound

This paper cites GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio.

Phoenix TTS: High-Fidelity Synthesis and Voice Conversion via Flow-Matching-Driven Speech Tokenization GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-16T00:37:04.864550Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:37:04.864550Z digest=sha256:46ad08d92436cc44bef79166c234c5df5698fd1c8b16c9c009601303b1a52f8c

Observation c7db203f-2dcf-4328-b109-506bf25fb32d · outbound

This paper cites Ds-codec: Dual-stage training with mirror-to-nonmirror architecture switching for speech codec.

Phoenix TTS: High-Fidelity Synthesis and Voice Conversion via Flow-Matching-Driven Speech Tokenization Ds-codec: Dual-stage training with mirror-to-nonmirror architecture switching for speech codec

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:37:05.631902Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T00:37:04.868619Z digest=sha256:9bfc2d25475b2fb51b52aab2cbaa6a2dcfc3eea89b0d1a05fcc3ce4e51738c37

Observation 775d7f71-93cd-4c9b-aacd-a5b1b1044671 · outbound

This paper cites SARA: A Dual-Stream VAE for High-Fidelity Speech Generation via Integrating Semantic and Acoustic Representations.

Phoenix TTS: High-Fidelity Synthesis and Voice Conversion via Flow-Matching-Driven Speech Tokenization SARA: A Dual-Stream VAE for High-Fidelity Speech Generation via Integrating Semantic and Acoustic Representations

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-16T00:37:04.872459Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:37:04.872459Z digest=sha256:20885a08cc421590de4041798cb58c3f80ed881240c513578e719d57e64a25dd

Observation 26e56797-04fd-4565-8829-026f3014147c · outbound

This paper cites Wavlm: Large-scale self-supervised pre- training for full stack speech processing.IEEE Journal of Selected Topics in Signal Processing, 16(6):1505–1518, 2022.

Phoenix TTS: High-Fidelity Synthesis and Voice Conversion via Flow-Matching-Driven Speech Tokenization Wavlm: Large-scale self-supervised pre- training for full stack speech processing.IEEE Journal of Selected Topics in Signal Processing, 16(6):1505–1518, 2022

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:37:05.616681Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T00:37:04.876165Z digest=sha256:ae48740ffc2e7849c76be6275cd79e7a68bcbcea6363966a9b19c25511e693f4

Observation d0725086-5c4b-4806-88ac-ad3c29b57b46 · outbound

This paper cites Neural codec language models are zero-shot text to speech synthesizers.IEEE Transactions on Audio, Speech and Language Processing, 33:705–718, 2025.

Phoenix TTS: High-Fidelity Synthesis and Voice Conversion via Flow-Matching-Driven Speech Tokenization Neural codec language models are zero-shot text to speech synthesizers.IEEE Transactions on Audio, Speech and Language Processing, 33:705–718, 2025

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-16T00:37:04.880497Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:37:04.880497Z digest=sha256:cae870d4e9ed62c289e81e663b30f90d9a79fb73b9aaf6e91a3ba366676752fc

Observation f3a9f138-c9ae-41a3-b09a-77a8b5b633de · outbound

This paper cites F5-tts: A fairytaler that fakes fluent and faithful speech with flow matching.

Phoenix TTS: High-Fidelity Synthesis and Voice Conversion via Flow-Matching-Driven Speech Tokenization F5-tts: A fairytaler that fakes fluent and faithful speech with flow matching

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-16T00:37:04.884261Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:37:04.884261Z digest=sha256:c4ad2ff66976158583b459af35db5cd0d019e7e9016d5cb60572e4249567b44c

Observation 97897886-1e82-4a23-8a54-9682a6b4b788 · outbound

This paper cites W2v-bert: Combining contrastive learning and masked language modeling for self-supervised speech pre-training.

Phoenix TTS: High-Fidelity Synthesis and Voice Conversion via Flow-Matching-Driven Speech Tokenization W2v-bert: Combining contrastive learning and masked language modeling for self-supervised speech pre-training

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-16T00:37:04.888080Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:37:04.888080Z digest=sha256:930b40ff7b88ac29cb3b42970fccf27c0b28cfaab66c861dd21b86d59f9c3812

Observation cf54ab6f-ceb9-43f5-8639-0c71bc5ee143 · outbound

This paper cites CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens.

Phoenix TTS: High-Fidelity Synthesis and Voice Conversion via Flow-Matching-Driven Speech Tokenization CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-16T00:37:04.891909Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:37:04.891909Z digest=sha256:503470e9291be7c4f4ff36d787e4f25701e94cbb951467f94251de93010eecd1

Observation 4f0f3479-6dfa-4191-9acf-269cb5071608 · outbound

This paper cites CosyVoice 3: Towards In-the-wild Speech Generation via Scaling-up and Post-training.

Phoenix TTS: High-Fidelity Synthesis and Voice Conversion via Flow-Matching-Driven Speech Tokenization CosyVoice 3: Towards In-the-wild Speech Generation via Scaling-up and Post-training

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-16T00:37:04.896325Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:37:04.896325Z digest=sha256:9a2d35c846415ce946c60131dcfec4fdee55b94747f5c8e78f39b41427052cad

Observation dff943c9-eaae-4729-be9c-2d189e7ef9cc · outbound

This paper cites CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models.

Phoenix TTS: High-Fidelity Synthesis and Voice Conversion via Flow-Matching-Driven Speech Tokenization CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-16T00:37:04.900278Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:37:04.900278Z digest=sha256:5e1b45c1bff078eaf02cb6671a5116ab9f669792dac0d81603c561cdae516683

Observation 47121a70-ff15-4021-a5c8-1efac8aa7538 · outbound

This paper cites High Fidelity Neural Audio Compression.

Phoenix TTS: High-Fidelity Synthesis and Voice Conversion via Flow-Matching-Driven Speech Tokenization High Fidelity Neural Audio Compression

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-16T00:37:04.903997Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:37:04.903997Z digest=sha256:167a4edeece9fc7653f414c6435d132bce22607e92a9b20fb9399b47f6fd1f20

Observation cde3c9fd-c867-4eae-8885-5230e107130d · outbound

This paper cites E2 tts: Embarrassingly easy fully non-autoregressive zero-shot tts.

Phoenix TTS: High-Fidelity Synthesis and Voice Conversion via Flow-Matching-Driven Speech Tokenization E2 tts: Embarrassingly easy fully non-autoregressive zero-shot tts

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-16T00:37:04.907580Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:37:04.907580Z digest=sha256:578e812dc8d8d7eb7d799b838ea63714dbe623268258073be2a617f30e9af14b

Observation ee9a78e7-0c36-4084-b44b-9198b2f3fd84 · outbound

This paper cites FunASR: A Fundamental End-to-End Speech Recognition Toolkit.

Phoenix TTS: High-Fidelity Synthesis and Voice Conversion via Flow-Matching-Driven Speech Tokenization FunASR: A Fundamental End-to-End Speech Recognition Toolkit

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-16T00:37:04.911143Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:37:04.911143Z digest=sha256:f6ddb2c96f89061117e4e76daa9929e12f496b4f918acd66c70a29d0f472b0ca

Observation 848bc857-ef5c-43a4-a6b0-e6b640ba068f · outbound

This paper cites Conformer: Convolution-augmented Transformer for Speech Recognition.

Phoenix TTS: High-Fidelity Synthesis and Voice Conversion via Flow-Matching-Driven Speech Tokenization Conformer: Convolution-augmented Transformer for Speech Recognition

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-16T00:37:04.915043Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:37:04.915043Z digest=sha256:7ca33282f0607ac8459d804b81b7043b217376b4040d4ab088ace44907cd8dec

Observation d68c1bcb-43ad-440c-b9ef-a1e1f4946ff5 · outbound

This paper cites FireRedTTS: A Foundation Text-To-Speech Framework for Industry-Level Generative Speech Applications.

Phoenix TTS: High-Fidelity Synthesis and Voice Conversion via Flow-Matching-Driven Speech Tokenization FireRedTTS: A Foundation Text-To-Speech Framework for Industry-Level Generative Speech Applications

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-16T00:37:04.919126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:37:04.919126Z digest=sha256:d39472f453dc8c35cf43a9d775ca2a409e2cc4bb5874f37ee63b10f2922a594e

Observation f76cc8bc-1f3a-4b96-927c-a71265d63cf7 · outbound

This paper cites Didispeech: A large scale mandarin speech corpus.

Phoenix TTS: High-Fidelity Synthesis and Voice Conversion via Flow-Matching-Driven Speech Tokenization Didispeech: A large scale mandarin speech corpus

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-16T00:37:04.923020Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:37:04.923020Z digest=sha256:c934fe342b56c768a180674bcc8b11b55352d8510228c1eb8eb63ede256f93c3

Observation 2f2cd431-891c-4a07-ad23-0bb3de3965c2 · outbound

This paper cites Emilia: An extensive, multilingual, and diverse speech dataset for large-scale speech generation.

Phoenix TTS: High-Fidelity Synthesis and Voice Conversion via Flow-Matching-Driven Speech Tokenization Emilia: An extensive, multilingual, and diverse speech dataset for large-scale speech generation

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-16T00:37:04.926907Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:37:04.926907Z digest=sha256:4608693e8b51edc017bf775c22d4cd23ba16d3db962de11e567e631068fdc4c5

Observation faf71c01-bf65-4aa7-b665-7701c3e67bc2 · outbound

This paper cites Hubert: Self-supervised speech representation learning by masked prediction of hidden units.IEEE/ACM transactions on audio, speech, and language processing, 29:3451–3460, 2021.

Phoenix TTS: High-Fidelity Synthesis and Voice Conversion via Flow-Matching-Driven Speech Tokenization Hubert: Self-supervised speech representation learning by masked prediction of hidden units.IEEE/ACM transactions on audio, speech, and language processing, 29:3451–3460, 2021

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:37:05.549427Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T00:37:04.931008Z digest=sha256:10599a77a0f143ff06e28092b469e08d1ab4553dd723c91c783f0b30270576bd

Observation 3657d78c-4938-47de-9af2-cfa2f82cc64d · outbound

This paper cites Ditar: Diffusion transformer autoregressive modeling for speech generation.arXiv preprint arXiv:2502.03930, 2025.

Phoenix TTS: High-Fidelity Synthesis and Voice Conversion via Flow-Matching-Driven Speech Tokenization Ditar: Diffusion transformer autoregressive modeling for speech generation.arXiv preprint arXiv:2502.03930, 2025

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-16T00:37:04.935124Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:37:04.935124Z digest=sha256:825c5fd5bdee69fd354be50bd23d2974e54d53c7752ad125681a543adfc0fee5

Observation 38fea359-7336-4eca-8aec-8ec2fa84dc1d · outbound

This paper cites MegaTTS 3: Sparse Alignment Enhanced Latent Diffusion Transformer for Zero-Shot Speech Synthesis.

Phoenix TTS: High-Fidelity Synthesis and Voice Conversion via Flow-Matching-Driven Speech Tokenization MegaTTS 3: Sparse Alignment Enhanced Latent Diffusion Transformer for Zero-Shot Speech Synthesis

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-16T00:37:04.938708Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:37:04.938708Z digest=sha256:d88b855a9fd168fded0ba8af240fff031f7ebe919a6c5fc7d298b63efad04ca8

Observation 7eb41eb7-2020-4432-89de-a3686b7633d0 · outbound

This paper cites Libriheavy: A 50,000 hours asr corpus with punctuation casing and context.

Phoenix TTS: High-Fidelity Synthesis and Voice Conversion via Flow-Matching-Driven Speech Tokenization Libriheavy: A 50,000 hours asr corpus with punctuation casing and context

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:37:05.534863Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T00:37:04.942649Z digest=sha256:95ece6b9f75be2df4b0563133ffaf9749a8a575d74bee2b8c499b3b9e5b44afc

Observation 201224c3-6514-4c19-a9a5-45de04100265 · outbound

This paper cites High- fidelity audio compression with improved rvqgan.Advances in Neural Information Processing Systems, 36:27980–27993, 2023.

Phoenix TTS: High-Fidelity Synthesis and Voice Conversion via Flow-Matching-Driven Speech Tokenization High- fidelity audio compression with improved rvqgan.Advances in Neural Information Processing Systems, 36:27980–27993, 2023

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T00:37:05.521990Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-16T00:37:04.946072Z digest=sha256:33d0955bc5941294022d34843ad171492e20bcd2b03940fc6c3baee508e80751

Observation 0f71aab3-e39d-4989-b3f4-31a948676316 · outbound

This paper cites DiTTo-TTS: Diffusion Transformers for Scalable Text-to-Speech without Domain-Specific Factors.

Phoenix TTS: High-Fidelity Synthesis and Voice Conversion via Flow-Matching-Driven Speech Tokenization DiTTo-TTS: Diffusion Transformers for Scalable Text-to-Speech without Domain-Specific Factors

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-16T00:37:04.949485Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:37:04.949485Z digest=sha256:c5178a0e11e7d946c1a827d6a9731011fd06503ad92ec3f1a98c94efab202289

Observation 6e64899a-7835-46d9-bb90-30dbda7a87e4 · outbound

This paper cites Zero-shot Voice Conversion with Diffusion Transformers.

Phoenix TTS: High-Fidelity Synthesis and Voice Conversion via Flow-Matching-Driven Speech Tokenization Zero-shot Voice Conversion with Diffusion Transformers

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-16T00:37:04.953206Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:37:04.953206Z digest=sha256:c852d9aaba5965def2d1172c6a07efb49eebdaf38cc0de34dc3b59fc927b077e

Observation 4b115cb4-af6c-4869-9f42-bc6710c32c45 · outbound

This paper cites Autoregressive Diffusion Transformer for Text-to-Speech Synthesis.

Phoenix TTS: High-Fidelity Synthesis and Voice Conversion via Flow-Matching-Driven Speech Tokenization Autoregressive Diffusion Transformer for Text-to-Speech Synthesis

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-16T00:37:04.957117Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:37:04.957117Z digest=sha256:9df1ed6562f4deebc049204ef5429a553cf55935be7fad2a3b6fb6f51c9eb926

Observation ea442264-10cd-41a2-98d6-e8cf3df62d6f · outbound

This paper cites WenetSpeech4TTS: A 12,800-hour Mandarin TTS Corpus for Large Speech Generation Model Benchmark.

Phoenix TTS: High-Fidelity Synthesis and Voice Conversion via Flow-Matching-Driven Speech Tokenization WenetSpeech4TTS: A 12,800-hour Mandarin TTS Corpus for Large Speech Generation Model Benchmark

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-16T00:37:04.960658Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:37:04.960658Z digest=sha256:2907e8f84399ea1da14ca81789ef07e85c5ad8579cf75eb6f38a1f3d2d9d7b6a

Observation c0a8bac5-196d-4068-8128-52d64fa6ed3c · outbound

This paper cites Librispeech-pc: Benchmark for evaluation of punctuation and capitaliza- tion capabilities of end-to-end asr models.

Phoenix TTS: High-Fidelity Synthesis and Voice Conversion via Flow-Matching-Driven Speech Tokenization Librispeech-pc: Benchmark for evaluation of punctuation and capitaliza- tion capabilities of end-to-end asr models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-16T00:37:04.964166Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:37:04.964166Z digest=sha256:c96deae5a828a2f4d27177a5c76a656032d0b5b34739cf9ec7e5f96168badfde

Observation 90ac5d2e-5115-4339-aa2e-12d74942f7f7 · outbound

This paper cites Autoregressive speech synthesis without vector quantization.

Phoenix TTS: High-Fidelity Synthesis and Voice Conversion via Flow-Matching-Driven Speech Tokenization Autoregressive speech synthesis without vector quantization

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-16T00:37:04.967650Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:37:04.967650Z digest=sha256:754c559d0631c3b6258a29058b0f4ee1e514ee84ea1fa47e71a4ff62da57f8dd

Observation 3caaa9a8-a762-44b4-abb2-2e4f28dba4d6 · outbound

This paper cites Librispeech: an asr corpus based on public domain audio books.

Phoenix TTS: High-Fidelity Synthesis and Voice Conversion via Flow-Matching-Driven Speech Tokenization Librispeech: an asr corpus based on public domain audio books

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-16T00:37:04.971209Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:37:04.971209Z digest=sha256:e448d0545ac3d8188dd191fe3d3c65fc3cdbb3f068070611e8350daed7c2b082

Observation a9d67925-d54f-4a3d-99e7-02fc5d60a9ff · outbound

This paper cites Scalable diffusion models with transformers.

Phoenix TTS: High-Fidelity Synthesis and Voice Conversion via Flow-Matching-Driven Speech Tokenization Scalable diffusion models with transformers

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-16T00:37:04.975118Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:37:04.975118Z digest=sha256:f6fd4f75c5ead37af812d3b50e0e1a2ef9f0b2985b26047ab4725e298e43e305

Observation dbfe518d-49fe-447c-aead-9e343fd9bef5 · outbound

This paper cites Robust speech recognition via large-scale weak supervision.

Phoenix TTS: High-Fidelity Synthesis and Voice Conversion via Flow-Matching-Driven Speech Tokenization Robust speech recognition via large-scale weak supervision

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-16T00:37:04.978650Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:37:04.978650Z digest=sha256:65de46f72e950a41c183bd22212f1d7d9d68715a58763d13d569a8c19af527b6

Observation bef0a40f-7795-4b00-ac2a-a536a6ff8af7 · outbound

This paper cites Vocos: Closing the gap between time-domain and Fourier-based neural vocoders for high-quality audio synthesis.

Phoenix TTS: High-Fidelity Synthesis and Voice Conversion via Flow-Matching-Driven Speech Tokenization Vocos: Closing the gap between time-domain and Fourier-based neural vocoders for high-quality audio synthesis

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-16T00:37:04.982405Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:37:04.982405Z digest=sha256:b05c3e2cc442977375292c269f00c47cbe4da64d1fddcfe70ad1dee5a3bca7d5

Observation a1c13e3b-44d7-4ed7-934c-2568ce3b2d8d · outbound

This paper cites Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens.

Phoenix TTS: High-Fidelity Synthesis and Voice Conversion via Flow-Matching-Driven Speech Tokenization Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-16T00:37:04.986772Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:37:04.986772Z digest=sha256:2e6066e5d44dd779f20c4d1b64b275a1e9ef79cea7bef82b8d54457ca777ae07

Observation ae40bb28-6d73-43a9-a72c-fa9cccee1085 · outbound

This paper cites MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer.

Phoenix TTS: High-Fidelity Synthesis and Voice Conversion via Flow-Matching-Driven Speech Tokenization MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-16T00:37:04.990844Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:37:04.990844Z digest=sha256:aa6b5915af7015dce49809228d5f459d365f71b7f8f22823b4871931010ed3cd

Observation 81311671-a9d0-4270-9e88-5af927fd25c1 · outbound

This paper cites BigCodec: Pushing the Limits of Low-Bitrate Neural Speech Codec.

Phoenix TTS: High-Fidelity Synthesis and Voice Conversion via Flow-Matching-Driven Speech Tokenization BigCodec: Pushing the Limits of Low-Bitrate Neural Speech Codec

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-16T00:37:04.994189Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:37:04.994189Z digest=sha256:c80e63513f6ca92daa16a9baca0c189a022bb2f3447a4e774864af80a4937140

Observation c8019c36-794d-4b8e-9aed-58afb3dd4281 · outbound

This paper cites SpeechTokenizer: Unified Speech Tokenizer for Speech Large Language Models.

Phoenix TTS: High-Fidelity Synthesis and Voice Conversion via Flow-Matching-Driven Speech Tokenization SpeechTokenizer: Unified Speech Tokenizer for Speech Large Language Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-16T00:37:04.998310Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:37:04.998310Z digest=sha256:a9253640111ded519aa9246d5b04ac9583c26afb165fe4803928340ed24dbde6

Observation 3e209853-2f7b-4c21-9c50-9e32064c4973 · outbound

This paper cites X-VC: Zero-shot Streaming Voice Conversion in Codec Space.

Phoenix TTS: High-Fidelity Synthesis and Voice Conversion via Flow-Matching-Driven Speech Tokenization X-VC: Zero-shot Streaming Voice Conversion in Codec Space

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-16T00:37:05.002453Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:37:05.002453Z digest=sha256:259221662c5ddedbb70d16e61f3956b87f0a8037bf43972f5aa23e60ec440525

Observation 17a4ed74-2ab2-4784-b52c-030b5cb0a95a · outbound

This paper cites Indextts2: A breakthrough in emotionally expressive and duration-controlled auto-regressive zero-shot text-to-speech.

Phoenix TTS: High-Fidelity Synthesis and Voice Conversion via Flow-Matching-Driven Speech Tokenization Indextts2: A breakthrough in emotionally expressive and duration-controlled auto-regressive zero-shot text-to-speech

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-16T00:37:05.007627Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:37:05.007627Z digest=sha256:eec35b64cf3835441a3f9a524420940881bdea07bc67d4540f016afe179671af

Observation e29db32b-1c62-41d3-8dc8-db4e941dadf6 · outbound

This paper cites V oxcpm: Tokenizer-free tts for context-aware speech generation and true-to-life voice cloning.arXiv preprint arXiv:2509.24650, 2025.

Phoenix TTS: High-Fidelity Synthesis and Voice Conversion via Flow-Matching-Driven Speech Tokenization V oxcpm: Tokenizer-free tts for context-aware speech generation and true-to-life voice cloning.arXiv preprint arXiv:2509.24650, 2025

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-16T00:37:05.011278Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:37:05.011278Z digest=sha256:d318d20a0460a4c137a0242d6b63faeba26b1b99ca73233d8d2036dbec5d9216

Pith citing papers

No inbound Pith citation observations are available.