Pith. sign in

Paper Citation Record · LEDGER

IndexTTS: An Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System

As of 10 August 2026, this Paper Citation Record lists 25 of 25 outbound references and 31 inbound Pith citation observations for arXiv:2502.05512.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.05512 v1

Coverage vector

measured 25 of 25 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-08T19:07:00.588636Z

measured 56 of 56 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 31 of 31 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-10T04:20:41.984640Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

25 of 25 outbound references displayed

  • verified exact0
  • verified fuzzy1
  • unresolved23
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

1
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 6530d5cc-7b61-4af1-ad91-7f92bfd767cd · outbound

This paper cites IndexTTS: An Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System.

IndexTTS: An Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System IndexTTS: An Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-08T19:07:00.481427Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:07:00.481427Z digest=sha256:906502fa1e273df9d797916993d5ae880d01beae4ffe9e70e113f548bdd726a3

Observation a670a59a-cc8b-4b05-946f-0b0e678666c0 · outbound

This paper cites [BT], prompt text, text, [ET], [BA], prompt audio, audio, [EA].

IndexTTS: An Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System [BT], prompt text, text, [ET], [BA], prompt audio, audio, [EA]

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T19:07:00.919750Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T19:07:00.486856Z digest=sha256:1e86469267f57150dfab9a22b44c4c856f3ef3ba962750eb35c9dbcce43ee831

Observation f09f3234-2912-49b8-9f20-b1e6f601f349 · outbound

This paper cites Dataset All training data was collected from the internet, with an initial 120,000 hours of raw audio.

IndexTTS: An Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System Dataset All training data was collected from the internet, with an initial 120,000 hours of raw audio

Reference 3

Resolution
malformed identifier
raw_fallback, observed 2026-08-08T19:07:00.905617Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-08T19:07:00.491564Z digest=sha256:1a459cb8e8b946fee1628a9932966cf02e3b2bc193a4f19048e86ecc063a4fd6

Observation 4b07bdf7-89e2-4fbc-bd3f-5ea2dc2abbf2 · outbound

This paper cites XTTS: a Massively Multilingual Zero-Shot Text-to-Speech Model.

IndexTTS: An Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System XTTS: a Massively Multilingual Zero-Shot Text-to-Speech Model

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-08T19:07:00.496354Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:07:00.496354Z digest=sha256:e8e30e278ddf88535faba187e9bdd99673e4fe8c0eef71c8ca5ee579371a4ddd

Observation 56fb5224-3e15-411f-9401-d2b9b29b2629 · outbound

This paper cites Fish-Speech: Leveraging Large Language Models for Advanced Multilingual Text-to-Speech Synthesis.

IndexTTS: An Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System Fish-Speech: Leveraging Large Language Models for Advanced Multilingual Text-to-Speech Synthesis

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-08T19:07:00.501225Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:07:00.501225Z digest=sha256:c33f78e43b3873c8ba14ce005376d57c6856500b7c89bf13d20139cce6b212d8

Observation 99cdb013-4c37-4c63-aff5-ef9260ebb7f3 · outbound

This paper cites CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens.

IndexTTS: An Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-08T19:07:00.505811Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:07:00.505811Z digest=sha256:9d7d9802075b7471b20eef8876d6abeeb91023d157498c0d2ea5d2accee3965b

Observation 0fadcd10-c646-4158-be36-c0dddd35c6bb · outbound

This paper cites FireRedTTS: A Foundation Text-To-Speech Framework for Industry-Level Generative Speech Applications.

IndexTTS: An Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System FireRedTTS: A Foundation Text-To-Speech Framework for Industry-Level Generative Speech Applications

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-08T19:07:00.510661Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:07:00.510661Z digest=sha256:db9c2b2df8ed1368d96b0f3a34759249fbe51ed9882048a542ec80d99beed2ed

Observation 5ad645aa-8a29-424e-865c-dd703843400a · outbound

This paper cites F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching.

IndexTTS: An Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-08T19:07:00.515212Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:07:00.515212Z digest=sha256:83a6bf83ea55143bb8031999e7a2f95c97e72899870a63b06fb4c7eaf66f8771

Observation 9ca92841-67a1-4390-9ec0-f00df24db369 · outbound

This paper cites Mega-TTS 2: Boosting Prompting Mechanisms for Zero-Shot Speech Synthesis.

IndexTTS: An Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System Mega-TTS 2: Boosting Prompting Mechanisms for Zero-Shot Speech Synthesis

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-08T19:07:00.520391Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:07:00.520391Z digest=sha256:a452c03c9902e7fe5cc9b41dc096df210300f9a818c047911718fa0944978cc5

Observation f8bede2c-be27-4dc2-85dd-36125b0de7dd · outbound

This paper cites Yourtts: Towards zero-shot multi-speaker tts and zero-shot voice conversion for everyone,.

IndexTTS: An Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System Yourtts: Towards zero-shot multi-speaker tts and zero-shot voice conversion for everyone,

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-08T19:07:00.524666Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:07:00.524666Z digest=sha256:8529cea46f34ad19b274baa3757171aee1db9d283c96ed4617dd0517af90fa6f

Observation 89e29809-4d7c-40d5-a481-5b6ae6b1667b · outbound

This paper cites Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers.

IndexTTS: An Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-08T19:07:00.528669Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:07:00.528669Z digest=sha256:36b90d478cf39e172a0592d16bd56bcdb0c0682495d474345a35e870e504167c

Observation 9e1ef766-eea4-4b45-86f0-abf236a47554 · outbound

This paper cites Seed-TTS: A Family of High-Quality Versatile Speech Generation Models.

IndexTTS: An Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System Seed-TTS: A Family of High-Quality Versatile Speech Generation Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-08T19:07:00.532973Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:07:00.532973Z digest=sha256:9899c9e3ac7c60107896bf36a7ee9a442461ad6323fb040a5ba2e4058cd6b01b

Observation 54a85830-836c-4f55-ac96-7fdf24b2aad4 · outbound

This paper cites Hifi-gan: Generative adversarial net- works for efficient and high fidelity speech synthesis,.

IndexTTS: An Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System Hifi-gan: Generative adversarial net- works for efficient and high fidelity speech synthesis,

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-08T19:07:00.537933Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:07:00.537933Z digest=sha256:7224296473b3b77a27c1e587de85f5b7f3bd46adf5236d4fa09976b717d031f5

Observation ecf8f860-9809-4c0e-8dfd-c0805f94c1a2 · outbound

This paper cites Better speech synthesis through scaling.

IndexTTS: An Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System Better speech synthesis through scaling

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-08T19:07:00.542381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:07:00.542381Z digest=sha256:eb299556c9e0315049300223cd1162dabbeab7e6bd5262cb541378648358ac2e

Observation 98c78b57-5ca5-488d-8648-a662f0e598d4 · outbound

This paper cites Neural discrete represen- tation learning,.

IndexTTS: An Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System Neural discrete represen- tation learning,

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-08T19:07:00.546792Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:07:00.546792Z digest=sha256:ceb978f97ab36c121440909dd685879d277a7307a247362893f129e5b3303e78

Observation 21ab49b3-9225-4430-bf1b-39e4d32a42d1 · outbound

This paper cites Finite Scalar Quantization: VQ-VAE Made Simple.

IndexTTS: An Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System Finite Scalar Quantization: VQ-VAE Made Simple

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-08T19:07:00.550792Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:07:00.550792Z digest=sha256:380d8ff3b9a18c7513d29fe0aaed2a1e36521735e4ab9e895ea7aab8f5e1cbaa

Observation 579bf133-2af6-4ed5-8767-e6b373bc5b1c · outbound

This paper cites BigVGAN: A Universal Neural Vocoder with Large-Scale Training.

IndexTTS: An Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System BigVGAN: A Universal Neural Vocoder with Large-Scale Training

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-08T19:07:00.555072Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:07:00.555072Z digest=sha256:fae71989e1da50bdd80645b7d466e22affb86d8de0853662823659af1c48f0e8

Observation 8b9e2605-4eb6-4c05-9646-c89d1cd2044c · outbound

This paper cites Matcha-tts: A fast tts architecture with conditional flow match- ing,.

IndexTTS: An Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System Matcha-tts: A fast tts architecture with conditional flow match- ing,

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-08T19:07:00.559444Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:07:00.559444Z digest=sha256:6acefd90d2e5bd85c5514dd3b7df285eef7a1a9b6f919a78a5ddf6d0dbe85834

Observation 906e5a61-7860-47a1-bd52-349f57a0e54e · outbound

This paper cites Demucs: Deep Extractor for Music Sources with extra unlabeled data remixed.

IndexTTS: An Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System Demucs: Deep Extractor for Music Sources with extra unlabeled data remixed

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-08T19:07:00.563570Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:07:00.563570Z digest=sha256:9541a791af46049f833c704dca32d758c7375ac722f5023bfe8b7f90f3cb49eb

Observation 4b29838f-8228-4c82-b65f-c4478b6af569 · outbound

This paper cites Lib- rispeech: an asr corpus based on public domain audio books,.

IndexTTS: An Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System Lib- rispeech: an asr corpus based on public domain audio books,

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-08T19:07:00.567811Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:07:00.567811Z digest=sha256:a5703d02e35f65cee201c08b8678a8047a2755526b1dbf96515c15358ddf1959

Observation 51618b6c-1f70-47c8-9318-c6eeb7266fcc · outbound

This paper cites Aishell-1: An open- source mandarin speech corpus and a speech recognition base- line,.

IndexTTS: An Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System Aishell-1: An open- source mandarin speech corpus and a speech recognition base- line,

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-08T19:07:00.571915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:07:00.571915Z digest=sha256:b00230c729acf8ef09ba68dd62acc71ff8e51e7cb21ac3a5766390a357b71da6

Observation 278b1ffd-223a-4ed5-b5d0-e5c554b3a460 · outbound

This paper cites Common Voice: A Massively-Multilingual Speech Corpus.

IndexTTS: An Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System Common Voice: A Massively-Multilingual Speech Corpus

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-08T19:07:00.575954Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:07:00.575954Z digest=sha256:707ac7709dba63acce94f03616fef14738ab5192254e470e832a2e47e31c368c

Observation 97b06b7c-850a-4ebe-9f0e-e9a21b7871f5 · outbound

This paper cites Paraformer: Fast and Accurate Parallel Transformer for Non-autoregressive End-to-End Speech Recognition.

IndexTTS: An Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System Paraformer: Fast and Accurate Parallel Transformer for Non-autoregressive End-to-End Speech Recognition

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-08T19:07:00.580468Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:07:00.580468Z digest=sha256:b155138efe42d41b3b2aa69787ee3a04030655baca9a243cc0040dfeb00bbdd1

Observation 27a95564-2160-45b0-8af3-8ffe9a76837a · outbound

This paper cites Robust speech recognition via large-scale weak supervision,.

IndexTTS: An Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System Robust speech recognition via large-scale weak supervision,

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-08T19:07:00.584577Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:07:00.584577Z digest=sha256:8c140ad17574babd1d6d4d2c51e914b4f4c31e9088120de60b5c9ddd06c26889

Observation 56c59340-6fa5-4555-87f6-1aabf4fc68a6 · outbound

This paper cites CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models.

IndexTTS: An Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-08T19:07:00.588636Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:07:00.588636Z digest=sha256:c020abe5b357bbee55cd81d36e02a3c77fef97ca6f4df08b4b7c580592668b6c

Pith citing papers

Observation 6530d5cc-7b61-4af1-ad91-7f92bfd767cd · inbound

IndexTTS: An Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System cites this paper.

IndexTTS: An Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System IndexTTS: An Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-08T19:07:00.481427Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:07:00.481427Z digest=sha256:906502fa1e273df9d797916993d5ae880d01beae4ffe9e70e113f548bdd726a3

Observation cceffa41-b9b1-4038-bf68-672cd9a46730 · inbound

UniTTS: An end-to-end TTS system without decoupling of acoustic and semantic information cites this paper.

UniTTS: An end-to-end TTS system without decoupling of acoustic and semantic information IndexTTS: An Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T14:53:14.979998Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:53:14.979998Z digest=sha256:9962c718a1976ad4796c4dcfa1358c2a836941406ad633866c4636028abd5370

Observation ca5794d7-3d49-4f58-bef4-4742bc912a27 · inbound

CosyVoice 3: Towards In-the-wild Speech Generation via Scaling-up and Post-training cites this paper.

CosyVoice 3: Towards In-the-wild Speech Generation via Scaling-up and Post-training IndexTTS: An Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-16T05:27:25.569998Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T05:27:25.425188Z digest=sha256:ac0cf021c8143bacf99845deee7de5e4356dd2b9bfaab3487acdaa68ad57288f

Observation 03ab161a-d36d-4a93-9e79-0c4f7662f124 · inbound

CloneShield: A Framework for Universal Perturbation Against Zero-Shot Voice Cloning cites this paper.

CloneShield: A Framework for Universal Perturbation Against Zero-Shot Voice Cloning IndexTTS: An Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T14:25:01.135750Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:25:01.135750Z digest=sha256:e6c8400099906cefae9f1e021a1fe14eac1e213a8e26049e34bfa0bd2922637b

Observation 743b8f20-f8f9-484d-9729-16dcc0f97530 · inbound

IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech cites this paper.

IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech IndexTTS: An Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T23:21:52.901086Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:21:52.901086Z digest=sha256:caec5bd070f827bb85ef9b34ef469660a214928728623d6ea177ea4d623fa6b3

Observation 6df3652a-8560-48d3-9cff-ff5f41003c0b · inbound

ZipVoice-Dialog: Non-Autoregressive Spoken Dialogue Generation with Flow Matching cites this paper.

ZipVoice-Dialog: Non-Autoregressive Spoken Dialogue Generation with Flow Matching IndexTTS: An Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-19T04:32:03.696915Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-19T04:29:41.285194Z digest=sha256:339b02faaec634e48c21b8fec85b186296987fb9e2cf2f578c3918645884647d

Observation b005b13d-8d31-46e1-aaf2-2a23300bdd76 · inbound

Robust Residual Finite Scalar Quantization for Neural Compression cites this paper.

Robust Residual Finite Scalar Quantization for Neural Compression IndexTTS: An Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-05T18:21:55.877250Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T18:21:55.877250Z digest=sha256:35aae2e17d79a88a3ffd41b64af8e354c5c814884e89028d0773b5410f800bd1

Observation f6b11d61-fc62-4741-9e45-9c2fd67d14da · inbound

FireRedTTS-2: Towards Long Conversational Speech Generation for Podcast and Chatbot cites this paper.

FireRedTTS-2: Towards Long Conversational Speech Generation for Podcast and Chatbot IndexTTS: An Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-05T12:00:45.299724Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:00:45.299724Z digest=sha256:6530c2daeab77bc709f9f5ff22cef870844a9f66b85c9225824a4610752515ba

Observation 0ee0c70b-a7db-4912-9cc0-47f9039061f6 · inbound

DiTReducio: A Training-Free Acceleration for DiT-Based TTS via Progressive Calibration cites this paper.

DiTReducio: A Training-Free Acceleration for DiT-Based TTS via Progressive Calibration IndexTTS: An Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-04T19:11:59.502401Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T19:11:59.502401Z digest=sha256:d38f42b6303a9b7066163ff4365806a04036b1ff8efcdd4652e45c9084381d45

Observation c0d2a60b-8c49-41a8-b9ca-d0d7f9072168 · inbound

EchoFake: A Replay-Aware Dataset for Practical Speech Deepfake Detection cites this paper.

EchoFake: A Replay-Aware Dataset for Practical Speech Deepfake Detection IndexTTS: An Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-18T05:02:23.325753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-18T05:01:57.044869Z digest=sha256:c067bad69e579ac95841a6f5c348f1d48e2d723339fe9ed34b95a5479104cdb2

Observation a6cc9382-8537-4204-9c75-ad97e33a7f16 · inbound

OmniVoice: Towards Omnilingual Zero-Shot Text-to-Speech with Diffusion Language Models cites this paper.

OmniVoice: Towards Omnilingual Zero-Shot Text-to-Speech with Diffusion Language Models IndexTTS: An Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-13T23:03:24.762622Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-13T23:00:18.720371Z digest=sha256:99a81653ac3939cd70e5044fbac62ff062e658530bc9e13e88bea7c54be37f2a

Observation 9b8586ec-f497-4039-87a1-671dcfc07bf8 · inbound

AT-ADD: All-Type Audio Deepfake Detection Challenge Evaluation Plan cites this paper.

AT-ADD: All-Type Audio Deepfake Detection Challenge Evaluation Plan IndexTTS: An Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-11T06:21:00.716219Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T17:40:10.657590Z digest=sha256:42f52d9085db8904711f6c95c26c0c727bcf71c59db9559e8753e1f34fabe8c4

Observation 7f9524ac-1a88-4870-b2e1-c130847c7aeb · inbound

WAND: Windowed Attention and Knowledge Distillation for Efficient Autoregressive Text-to-Speech Models cites this paper.

WAND: Windowed Attention and Knowledge Distillation for Efficient Autoregressive Text-to-Speech Models IndexTTS: An Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-15T10:29:55.951508Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-15T10:29:51.313835Z digest=sha256:1579e493b451ad1321c6c269162d424678e747860c316b0fa2be8f84ffb632c4

Observation 066c6c75-1c36-4075-ab14-f14e9810b178 · inbound

Interactive ASR: Towards Human-Like Interaction and Semantic Coherence Evaluation for Agentic Speech Recognition cites this paper.

Interactive ASR: Towards Human-Like Interaction and Semantic Coherence Evaluation for Agentic Speech Recognition IndexTTS: An Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-11T00:41:07.162050Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T18:22:08.670559Z digest=sha256:c9abc3e640103a0614c0d663cfad4c9ba01607f7a6eb64bd89e28b14257001ba

Observation 32632292-6130-4474-9201-86b204be205b · inbound

ActorMind: Emulating Human Actor Reasoning for Speech Role-Playing cites this paper.

ActorMind: Emulating Human Actor Reasoning for Speech Role-Playing IndexTTS: An Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-11T09:21:02.483890Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-10T16:03:15.572657Z digest=sha256:5d0a103a461f938bd83e08bac844480a1a01b743715b54cd7fa4ba95cec3ec67

Observation 9253335f-202f-4fdd-ba32-02f22861ae7f · inbound

AST: Adaptive, Seamless, and Training-Free Precise Speech Editing cites this paper.

AST: Adaptive, Seamless, and Training-Free Precise Speech Editing IndexTTS: An Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-10T08:02:25.232303Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T07:59:34.622096Z digest=sha256:936bba9b0c4105f69f1ecd3546baec7669cf7be46d233049e6d87a3e08faae71

Observation 6ede54ab-6ac8-4d72-911d-eb83bc9a9429 · inbound

AST: Adaptive, Seamless, and Training-Free Precise Speech Editing cites this paper.

AST: Adaptive, Seamless, and Training-Free Precise Speech Editing IndexTTS: An Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-04T05:29:37.697055Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T05:29:37.697055Z digest=sha256:27ff19fd1f6bda52e24a9922936b468cf34b508f340ef9d56bc7203d835c0ac9

Observation 5c0fc56d-4fe0-47e3-a0b1-fb609c5fe0c9 · inbound

MINT-Bench: A Comprehensive Multilingual Benchmark for Instruction-Following Text-to-Speech cites this paper.

MINT-Bench: A Comprehensive Multilingual Benchmark for Instruction-Following Text-to-Speech IndexTTS: An Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-10T04:04:47.235530Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T04:03:38.919545Z digest=sha256:fa7c91f5162897c68f43bccf04beda6cb31bb4a151ec5456105ac40436ea93cb

Observation a21a2f5c-3be3-4b7a-b59b-38e9beef61e2 · inbound

RoboKA: KAN Informed Multimodal Learning for RoboCall Surveillance System cites this paper.

RoboKA: KAN Informed Multimodal Learning for RoboCall Surveillance System IndexTTS: An Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-11T15:26:10.281890Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-09T19:59:00.255786Z digest=sha256:c9e4865296a1e38590da7b2641f6c6d1c2350c3f8d67a0ee15a57be35cc11317

Observation 7f83fb1f-d3f9-4bf7-a807-81baadf1f855 · inbound

SemaVoice: Semantic-Aware Continuous Autoregressive Speech Synthesis cites this paper.

SemaVoice: Semantic-Aware Continuous Autoregressive Speech Synthesis IndexTTS: An Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-19T19:02:43.600587Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-19T18:58:15.288299Z digest=sha256:66169f54cb62699db0547c4e1d8393e862128bfc8dd4321b492ef1e4ccb15a29

Observation 3cc031fc-1a1c-4bbb-be1f-c66b39d9a931 · inbound

AgentSteerTTS: A Multi-Agent Closed-Loop Framework for Composite-Instruction Text-to-Speech cites this paper.

AgentSteerTTS: A Multi-Agent Closed-Loop Framework for Composite-Instruction Text-to-Speech IndexTTS: An Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System

Reference 108

Resolution
verified exact
arxiv_id, observed 2026-05-20T21:19:03.270673Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-20T21:14:58.814362Z digest=sha256:68039abafdc8a0f2992fb9324ea2ecd974fb1b34bb1fef08fa1be815603edfdd

Observation ac48d7d0-7235-4016-ac85-5d345ec6882a · inbound

PilotTTS: A Disciplined Modular Recipe for Competitive Speech Synthesis cites this paper.

PilotTTS: A Disciplined Modular Recipe for Competitive Speech Synthesis IndexTTS: An Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-06-29T16:23:39.890163Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-29T15:51:21.519785Z digest=sha256:77bfa279600506abacf83c1dd0a0fd1cfc4336793d642413afd44650b2844934

Observation c8a2368c-09d1-4a4a-85c6-16870403ee90 · inbound

Towards Human-Like Interactive Speech Recognition With Agentic Correction and Semantic Evaluation cites this paper.

Towards Human-Like Interactive Speech Recognition With Agentic Correction and Semantic Evaluation IndexTTS: An Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-06-29T07:33:13.683613Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-29T07:30:11.718647Z digest=sha256:43bf7d17c3951c6399c6c06f5c525e64cb469cd772835ed71b3ba78b3897363d

Observation 09e111f7-6c4a-4bf2-a6a8-1ced99766cb2 · inbound

UniVocal: Unified Speech-Singing Code-Switching Synthesis cites this paper.

UniVocal: Unified Speech-Singing Code-Switching Synthesis IndexTTS: An Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System

Reference 65

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T00:46:24.601926Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-06-28T13:17:13.510587Z digest=sha256:c9cac3172861a91053b586e9da47977d5c482cae2ff1c90b4215dc987e2650bc

Observation 8380a461-bd96-48c6-baaa-409dd38bda18 · inbound

Read What You Hear: Reference-Free Hypotheses Evaluation with Acoustic Discrepancy cites this paper.

Read What You Hear: Reference-Free Hypotheses Evaluation with Acoustic Discrepancy IndexTTS: An Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-07-02T11:06:53.393587Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-28T04:37:26.452974Z digest=sha256:7388fd6b5940317de29132bf1fbad7e2e1efd459cdfad7e812465b8216875b7a

Observation f4b6d5e9-3a04-4460-bec4-761851ae30b0 · inbound

Joycent: Diffusion-based Accent TTS without Accented Phone Prediction cites this paper.

Joycent: Diffusion-based Accent TTS without Accented Phone Prediction IndexTTS: An Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-07-03T18:08:46.949342Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-27T03:15:41.454335Z digest=sha256:38eb6ecb55c95bb969934a9fee497d9e6999103440d5583ef1d22673fab2f842

Observation 0e620d67-2f50-448b-8493-788856d9f13d · inbound

UniSAE: Unified Speech Attribute Editing on Speaker, Emotion and Low-Level Content via Discrete Phonetic Posteriorgram Modelling cites this paper.

UniSAE: Unified Speech Attribute Editing on Speaker, Emotion and Low-Level Content via Discrete Phonetic Posteriorgram Modelling IndexTTS: An Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-07-01T11:45:46.194761Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-07-01T04:05:27.343684Z digest=sha256:9ed8611e22194ca4348d83808209204b0a57c66c6045849443852fc4a65bddba

Observation 472c3733-0573-4870-bc53-b39520c86dba · inbound

FlexiSLM: A Dynamic and Controllable Frame Rate Spoken Language Model cites this paper.

FlexiSLM: A Dynamic and Controllable Frame Rate Spoken Language Model IndexTTS: An Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System

Reference 137

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T11:45:47.283406Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-07-01T03:50:26.873406Z digest=sha256:ff5a0427fb8ed54d7d15e2bba167a2e6b7df0748ea699221a0b7aeb23603094d

Observation 82b3df18-4735-4230-91db-c7ca95c4d4ea · inbound

AutoSIFT: Automatic Style Sifting for Controllable Speech Generation with Arbitrary Style Infilling cites this paper.

AutoSIFT: Automatic Style Sifting for Controllable Speech Generation with Arbitrary Style Infilling IndexTTS: An Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-02T06:27:04.866415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:27:04.866415Z digest=sha256:2bbd3560161e0e2c3b180f1488f9f0917a67b9496ef5c13446fca3faceccfb72

Observation 148d8915-cd7c-42ea-b5af-27826fc4e2d2 · inbound

X-Translator: A Real-Time Multilingual Speaker-Aware Speech-to-Speech Translation System cites this paper.

X-Translator: A Real-Time Multilingual Speaker-Aware Speech-to-Speech Translation System IndexTTS: An Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-01T17:45:44.391170Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T17:45:44.391170Z digest=sha256:031a02cd8f03af029d6aa447ed005d2954cf7fc110534c4ca4c2506542166a69

Observation eb1187fd-25e5-46d6-aa79-62e7fea4410d · inbound

SemBridge: Semantic Token Anchoring for Continuous-Latent Autoregressive Speech Generation cites this paper.

SemBridge: Semantic Token Anchoring for Continuous-Latent Autoregressive Speech Generation IndexTTS: An Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-10T04:20:41.984640Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T04:20:41.984640Z digest=sha256:c7232e2f69318200943ff92b40ca0c4773359bb5fa6ec8c47467713c297e87a7