Pith. sign in

Paper Citation Record · LEDGER

FireRedTTS: A Foundation Text-To-Speech Framework for Industry-Level Generative Speech Applications

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 42 inbound Pith citation observations for arXiv:2409.03283.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2409.03283 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 42 of 42 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 42 of 42 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T13:26:19.124096Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T08:29:41.338585Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation d20dfc8c-8fda-41d8-81f3-2483d87ab40e · inbound

F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching cites this paper.

F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching FireRedTTS: A Foundation Text-To-Speech Framework for Industry-Level Generative Speech Applications

Reference 97

Resolution
verified exact
arxiv_id, observed 2026-05-16T06:06:41.552278Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-16T06:06:41.348728Z digest=sha256:3137d0d514650e3382f1c3b756d9b47733c8f023926c6598d94ce0c7df8f9de5

Observation e3754d63-1370-4fd4-8898-c1ddcb4dc393 · inbound

CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models cites this paper.

CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models FireRedTTS: A Foundation Text-To-Speech Framework for Industry-Level Generative Speech Applications

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-13T06:19:09.576355Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-13T06:19:09.507440Z digest=sha256:9545b188d40c48387f11f4f0a0d19f4d1b20c02ae6a781ec8379e8ccad2c0642

Observation ab9f82f4-1ab7-4ad4-82d3-3aab0706d44f · inbound

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement cites this paper.

Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement FireRedTTS: A Foundation Text-To-Speech Framework for Industry-Level Generative Speech Applications

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-08T13:26:19.124096Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T13:26:19.124096Z digest=sha256:a9c08c29edf08efd9d2d54d9927c4c6f66bbfa7c3d12d6da3b9b0e2a9330758e

Observation 6ba236b4-77cf-483f-80a7-d4709efe4c0d · inbound

Accelerating Autoregressive Speech Synthesis Inference With Speech Speculative Decoding cites this paper.

Accelerating Autoregressive Speech Synthesis Inference With Speech Speculative Decoding FireRedTTS: A Foundation Text-To-Speech Framework for Industry-Level Generative Speech Applications

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T15:23:15.859744Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:23:15.859744Z digest=sha256:dc836aa493366b035824060a6cae59411d5354eda16b3c9832f05dab79c178b2

Observation c380d880-2de5-43e2-b763-1b8e6964a994 · inbound

CosyVoice 3: Towards In-the-wild Speech Generation via Scaling-up and Post-training cites this paper.

CosyVoice 3: Towards In-the-wild Speech Generation via Scaling-up and Post-training FireRedTTS: A Foundation Text-To-Speech Framework for Industry-Level Generative Speech Applications

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-16T05:27:25.554102Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-16T05:27:25.425188Z digest=sha256:9ae6037d266d555afb884f4a484505c94e334f6353eef6c236668177fd7e01a4

Observation 235bff6a-1867-49b6-87de-bd2096abb91f · inbound

VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation cites this paper.

VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation FireRedTTS: A Foundation Text-To-Speech Framework for Industry-Level Generative Speech Applications

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-07T14:19:56.932091Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:19:56.932091Z digest=sha256:76fe5cb9a44bbaf712aee45d73747edaa5f3a64973de98dc55ed6f36459dfc9c

Observation 594d5d61-61ff-4221-becd-05815f436794 · inbound

Accelerating Diffusion-based Text-to-Speech Model Training with Dual Modality Alignment cites this paper.

Accelerating Diffusion-based Text-to-Speech Model Training with Dual Modality Alignment FireRedTTS: A Foundation Text-To-Speech Framework for Industry-Level Generative Speech Applications

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T14:15:55.201763Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:15:55.201763Z digest=sha256:655c616f5eeda352e40d235d52a3f72850283c0ebb09e9962847e912b893ffa7

Observation 7425d4dc-2ea1-4ce5-9d9b-c5eae321ef72 · inbound

AudioGenie: A Training-Free Multi-Agent Framework for Diverse Multimodality-to-Multiaudio Generation cites this paper.

AudioGenie: A Training-Free Multi-Agent Framework for Diverse Multimodality-to-Multiaudio Generation FireRedTTS: A Foundation Text-To-Speech Framework for Industry-Level Generative Speech Applications

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T13:21:54.561667Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:21:54.561667Z digest=sha256:6d1d2619fe93db550764f482b6a03f0cb8a4b227ffecc470c7634fee95317b1d

Observation 47614f9a-f3b0-4357-b723-32f725028c9f · inbound

ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching cites this paper.

ZipVoice: Fast and High-Quality Zero-Shot Text-to-Speech with Flow Matching FireRedTTS: A Foundation Text-To-Speech Framework for Industry-Level Generative Speech Applications

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T00:45:13.608161Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:45:13.608161Z digest=sha256:36c6d6f1d2afc4604561390217c650e0076286c198395e56a85cb4dc5eb6db08

Observation f04ef65d-a581-47f3-8108-7d8574e282b2 · inbound

IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech cites this paper.

IndexTTS2: A Breakthrough in Emotionally Expressive and Duration-Controlled Auto-Regressive Zero-Shot Text-to-Speech FireRedTTS: A Foundation Text-To-Speech Framework for Industry-Level Generative Speech Applications

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T23:21:54.069634Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:21:54.069634Z digest=sha256:5092944c05a2414e4b89df79fc8f5e2fd15245c3fafa40832146bf2aeae8ff21

Observation a4cb8f56-6738-412c-b666-4824b108b49c · inbound

JAM-Flow: Joint Audio-Motion Synthesis with Flow Matching cites this paper.

JAM-Flow: Joint Audio-Motion Synthesis with Flow Matching FireRedTTS: A Foundation Text-To-Speech Framework for Industry-Level Generative Speech Applications

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-05-22T00:50:50.989669Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-22T00:46:39.196042Z digest=sha256:8ef25bfc89eaa2df1bb7f3e3f8f78e334f3e57ecd196d9c58a77923aa06a9491

Observation 30f9b39f-def3-4812-afda-b754c9957879 · inbound

Speech Quality Assessment Model Based on Mixture of Experts: System-Level Performance Enhancement and Utterance-Level Challenge Analysis cites this paper.

Speech Quality Assessment Model Based on Mixture of Experts: System-Level Performance Enhancement and Utterance-Level Challenge Analysis FireRedTTS: A Foundation Text-To-Speech Framework for Industry-Level Generative Speech Applications

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T19:14:11.205847Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:14:11.205847Z digest=sha256:fe6729557f8ec4719321a56f6cb301f6a0e445343ce695d6737b5acb708b66c7

Observation d117a21e-b6f3-46f3-b34f-e7ac10819055 · inbound

ZipVoice-Dialog: Non-Autoregressive Spoken Dialogue Generation with Flow Matching cites this paper.

ZipVoice-Dialog: Non-Autoregressive Spoken Dialogue Generation with Flow Matching FireRedTTS: A Foundation Text-To-Speech Framework for Industry-Level Generative Speech Applications

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-19T04:32:03.685695Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-19T04:29:41.285194Z digest=sha256:2585b8ab08e5b120d3d568b981fde1870b5210b1e28aadcce710bb5695cb5539

Observation 5f720103-5c55-497f-8569-2b43f5e60094 · inbound

NonverbalTTS: A Public English Corpus of Text-Aligned Nonverbal Vocalizations with Emotion Annotations for Text-to-Speech cites this paper.

NonverbalTTS: A Public English Corpus of Text-Aligned Nonverbal Vocalizations with Emotion Annotations for Text-to-Speech FireRedTTS: A Foundation Text-To-Speech Framework for Industry-Level Generative Speech Applications

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T16:33:52.374067Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:33:52.374067Z digest=sha256:90a2e16a6d731b876f8a729535c465c9b282cfe237fa0184849a70d42091df4c

Observation bd611872-f44b-4e2d-aa23-8a946cd85e5e · inbound

SpeechFake: A Large-Scale Multilingual Speech Deepfake Dataset Incorporating Cutting-Edge Generation Methods cites this paper.

SpeechFake: A Large-Scale Multilingual Speech Deepfake Dataset Incorporating Cutting-Edge Generation Methods FireRedTTS: A Foundation Text-To-Speech Framework for Industry-Level Generative Speech Applications

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T12:49:17.719579Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T12:49:17.719579Z digest=sha256:9d0063ee74bceae5f021a815a3d14f72456320f07ccfffaedf65540194a6fb34

Observation c671d657-7cf7-46fc-a3ff-b7c02cdf8e9e · inbound

NVSpeech: An Integrated and Scalable Pipeline for Human-Like Speech Modeling with Paralinguistic Vocalizations cites this paper.

NVSpeech: An Integrated and Scalable Pipeline for Human-Like Speech Modeling with Paralinguistic Vocalizations FireRedTTS: A Foundation Text-To-Speech Framework for Industry-Level Generative Speech Applications

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T00:51:30.398007Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T00:51:30.398007Z digest=sha256:ae4246bafcf5c746b921a720eae6b69a58d53f8a862d1567d29bee856138990f

Observation 04207df1-0bfa-4b6d-9458-20207a08e057 · inbound

CLEAR: Continuous Latent Autoregressive Modeling for High-quality and Low-latency Speech Synthesis cites this paper.

CLEAR: Continuous Latent Autoregressive Modeling for High-quality and Low-latency Speech Synthesis FireRedTTS: A Foundation Text-To-Speech Framework for Industry-Level Generative Speech Applications

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-05T16:01:52.598217Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:01:52.598217Z digest=sha256:319f1497226caf3e5e46d9acda51df846593eb4a6a93c6bf65f606b0ac3ead9f

Observation cefde6a1-a184-456f-8e52-c5e09ce4e093 · inbound

FireRedTTS-2: Towards Long Conversational Speech Generation for Podcast and Chatbot cites this paper.

FireRedTTS-2: Towards Long Conversational Speech Generation for Podcast and Chatbot FireRedTTS: A Foundation Text-To-Speech Framework for Industry-Level Generative Speech Applications

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-05T12:00:45.274842Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:00:45.274842Z digest=sha256:44fd636b50f91486de53d43641011331306405d7dfb0cab092d59475a6722870

Observation dfafd71d-d7a0-4111-ac51-e2955428a5c4 · inbound

FireRedChat: A Pluggable, Full-Duplex Voice Interaction System with Cascaded and Semi-Cascaded Implementations cites this paper.

FireRedChat: A Pluggable, Full-Duplex Voice Interaction System with Cascaded and Semi-Cascaded Implementations FireRedTTS: A Foundation Text-To-Speech Framework for Industry-Level Generative Speech Applications

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-04T23:33:04.290230Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T23:33:04.290230Z digest=sha256:303f15342d1a7ac4d02f78eb8f1c6f5d2eb8b7d0ec04c786cb4e828d2daa9e6e

Observation d28099c9-99dc-49a7-8377-68848241ab74 · inbound

Ovi: Twin Backbone Cross-Modal Fusion for Audio-Video Generation cites this paper.

Ovi: Twin Backbone Cross-Modal Fusion for Audio-Video Generation FireRedTTS: A Foundation Text-To-Speech Framework for Industry-Level Generative Speech Applications

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T01:03:15.223871Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-17T01:03:15.183360Z digest=sha256:f7485c329af9855e2b0cd3f2577e5d7bc57c23ebf6dec85f6b8ca58b2abba3f6

Observation e26d98a2-a154-49b9-9b21-89d5298ec760 · inbound

UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models cites this paper.

UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models FireRedTTS: A Foundation Text-To-Speech Framework for Industry-Level Generative Speech Applications

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-04T11:29:32.431418Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T11:29:32.431418Z digest=sha256:744f1a1bfc89353f9e273508a4ab48948eadba02a311843c2ff566bcda5899d4

Observation 823a16f6-e458-4820-adef-e45758a0ca7b · inbound

EchoFake: A Replay-Aware Dataset for Practical Speech Deepfake Detection cites this paper.

EchoFake: A Replay-Aware Dataset for Practical Speech Deepfake Detection FireRedTTS: A Foundation Text-To-Speech Framework for Industry-Level Generative Speech Applications

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-18T05:02:23.383902Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-18T05:01:57.044869Z digest=sha256:1a43c277d6715d44336488fa3a0b0e1c13b70b15ee47c57e42ab37880e021f3e

Observation 77153097-b6dd-4302-877f-27f9e3db101b · inbound

OmniVoice: Towards Omnilingual Zero-Shot Text-to-Speech with Diffusion Language Models cites this paper.

OmniVoice: Towards Omnilingual Zero-Shot Text-to-Speech with Diffusion Language Models FireRedTTS: A Foundation Text-To-Speech Framework for Industry-Level Generative Speech Applications

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-13T23:03:24.891256Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-13T23:00:18.720371Z digest=sha256:b53cdd16d282f810f02ac548c48a3cdbfd9ffb46c22c8072e2aafa96b21a0639

Observation dff7dc9f-38a9-4243-a83e-443a16023c49 · inbound

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey cites this paper.

Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey FireRedTTS: A Foundation Text-To-Speech Framework for Industry-Level Generative Speech Applications

Reference 87

Resolution
verified exact
arxiv_id, observed 2026-05-11T08:30:57.154210Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T16:36:33.264166Z digest=sha256:086a48e95759986ae30160eee15918d366b27a191df296cfbf4a06177dc3f77c

Observation 46d95095-df79-4898-95cf-9d0b808f6a00 · inbound

AffectCodec: Emotion-Preserving Neural Speech Codec for Expressive Speech Modeling cites this paper.

AffectCodec: Emotion-Preserving Neural Speech Codec for Expressive Speech Modeling FireRedTTS: A Foundation Text-To-Speech Framework for Industry-Level Generative Speech Applications

Reference 51

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T01:07:00.265442Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-13T01:04:54.506749Z digest=sha256:967ee6f7a09e3f78e453a560d6f7757a0e9720dc87406e2665b565649b4f6469

Observation be6446a3-77f2-498c-818d-b1dbf5582866 · inbound

SemaVoice: Semantic-Aware Continuous Autoregressive Speech Synthesis cites this paper.

SemaVoice: Semantic-Aware Continuous Autoregressive Speech Synthesis FireRedTTS: A Foundation Text-To-Speech Framework for Industry-Level Generative Speech Applications

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-19T19:02:43.589090Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-19T18:58:15.288299Z digest=sha256:681ff3942275fdc476ea6301504ecce39548ba467c6b440642548f07530021ca

Observation ee013815-f54d-44c8-92bb-fac7cfed4656 · inbound

SwanVoice: Expressive Long-Form Zero-Shot Speech Synthesis for Both Monologue and Dialogue cites this paper.

SwanVoice: Expressive Long-Form Zero-Shot Speech Synthesis for Both Monologue and Dialogue FireRedTTS: A Foundation Text-To-Speech Framework for Industry-Level Generative Speech Applications

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-07-01T20:26:13.104940Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-28T21:05:54.061395Z digest=sha256:6cacd91af2d79a1f35719882e27079e9ffffab30041d097c025edf6df84b1564

Observation 5554fe46-4731-44ee-8b2e-3cec4b45bde0 · inbound

UniVocal: Unified Speech-Singing Code-Switching Synthesis cites this paper.

UniVocal: Unified Speech-Singing Code-Switching Synthesis FireRedTTS: A Foundation Text-To-Speech Framework for Industry-Level Generative Speech Applications

Reference 45

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T00:46:24.607854Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-28T13:17:13.510587Z digest=sha256:6fcc5b94856b8398446bdfc848a26200ee7412ec18829eb651c6ba9d8b47c882

Observation 4476e729-e5e4-4ced-aa61-c1d5762c8250 · inbound

WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling cites this paper.

WavTTS: Towards High-Quality Zero-Shot TTS via Direct Raw Waveform Modeling FireRedTTS: A Foundation Text-To-Speech Framework for Industry-Level Generative Speech Applications

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-07-02T05:16:39.768674Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-28T08:18:42.002083Z digest=sha256:f61a4ca01a6a8c02fee86269c1256fdd1ab0b0a1ef1d7cb01fce16b2378a2aa5

Observation d6c9f23d-4232-4d1b-b5be-cbde1fa35f93 · inbound

VoxCPM2 Technical Report cites this paper.

VoxCPM2 Technical Report FireRedTTS: A Foundation Text-To-Speech Framework for Industry-Level Generative Speech Applications

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-07-02T19:47:19.733226Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-27T21:18:22.911332Z digest=sha256:e9fb8f72f09599ce21980efaee2af482d84a34e8b66ab1396915269a3545feb7

Observation 45724a0f-9436-45d4-983c-cf0fbc79f786 · inbound

TLDR: Compressing Audio Tokens for Efficient Autoregressive Text-to-Speech cites this paper.

TLDR: Compressing Audio Tokens for Efficient Autoregressive Text-to-Speech FireRedTTS: A Foundation Text-To-Speech Framework for Industry-Level Generative Speech Applications

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-07-03T03:07:35.850892Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-27T15:27:01.142747Z digest=sha256:834ef3e51890721557a06a9bf197db86825f75507f278221fdf849b2523fdb2d

Observation 0f3a401c-6ffa-45d9-8b66-5b3376effd5b · inbound

HoliDubber: Holistic Video Dubbing for Complex Acoustic Scenes via Text-Guided Audio Synthesis cites this paper.

HoliDubber: Holistic Video Dubbing for Complex Acoustic Scenes via Text-Guided Audio Synthesis FireRedTTS: A Foundation Text-To-Speech Framework for Industry-Level Generative Speech Applications

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-07-03T03:27:35.634796Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-27T15:20:14.761337Z digest=sha256:2935f480b15c7e548b44d9a853a8f0e3565758da5fe11bb4c818e4ff319bcf70

Observation b4dfecfa-5f66-46ce-9ec4-c8d33a0bfd67 · inbound

FlashTTS: Fast Streaming TTS with MTP Acceleration and X-pred Mean Flow Distillation cites this paper.

FlashTTS: Fast Streaming TTS with MTP Acceleration and X-pred Mean Flow Distillation FireRedTTS: A Foundation Text-To-Speech Framework for Industry-Level Generative Speech Applications

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-07-03T03:27:36.170307Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-27T15:15:10.060770Z digest=sha256:0e2d8f117a85e126290e0556117afe45e6511666e06ffa7ce5fa5267e2b6a3af

Observation 58c08025-067f-41e3-96ba-0ad7c52fcac6 · inbound

End-to-End Training for Discrete Token LLM based TTS System cites this paper.

End-to-End Training for Discrete Token LLM based TTS System FireRedTTS: A Foundation Text-To-Speech Framework for Industry-Level Generative Speech Applications

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-07-03T03:27:34.892211Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-27T15:22:06.893507Z digest=sha256:9ec88d676794e9e22a4f438569c80d574ed48309439274fc11e201aa41cb10f0

Observation 24829c5b-27c9-47d4-acbd-828dfe98d711 · inbound

One-Step Token-to-Waveform Generation with MeanFlow in Latent Space cites this paper.

One-Step Token-to-Waveform Generation with MeanFlow in Latent Space FireRedTTS: A Foundation Text-To-Speech Framework for Industry-Level Generative Speech Applications

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-07-03T23:19:03.036123Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-26T22:45:53.073871Z digest=sha256:7ae3681874113ff92af631db15174b68bb2c48b0cd7e93e4fc7396ec59b0a4d4

Observation 0479e9c8-ef06-4c56-97af-b5a27e608693 · inbound

Streaming T5-based Text-to-Speech Synthesis with Limited Lookahead cites this paper.

Streaming T5-based Text-to-Speech Synthesis with Limited Lookahead FireRedTTS: A Foundation Text-To-Speech Framework for Industry-Level Generative Speech Applications

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-07-04T08:29:41.340017Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-26T11:42:44.035902Z digest=sha256:93ec4bac1e096fbd9c9d0c7aae37178191631c5b545ba34e4ae00fb635b41261

Observation d9b06688-7d9e-4526-84ae-34abc4897273 · inbound

FlexiSLM: A Dynamic and Controllable Frame Rate Spoken Language Model cites this paper.

FlexiSLM: A Dynamic and Controllable Frame Rate Spoken Language Model FireRedTTS: A Foundation Text-To-Speech Framework for Industry-Level Generative Speech Applications

Reference 191

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T11:45:47.144382Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-07-01T03:50:26.873406Z digest=sha256:4e0066b0a5eb44302846d69f7f57fd903ce68f02ef0df5cfd53ad50a81fc9863

Observation a0d0e77a-334f-45cf-a8b6-7686f25d5dba · inbound

Enhancing Flow Matching with A Unified Guidance Framework for Efficient and Robust Speech Synthesis cites this paper.

Enhancing Flow Matching with A Unified Guidance Framework for Efficient and Robust Speech Synthesis FireRedTTS: A Foundation Text-To-Speech Framework for Industry-Level Generative Speech Applications

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-07-02T06:56:44.308754Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-02T06:36:42.174254Z digest=sha256:ba2f72bafa122ebf7158e6e3c801c879a7fe388db68f72f32983c76c13beffe9

Observation 7916e613-da85-4173-8696-8df411191f06 · inbound

ReGen: Hierarchical Multi-Prompt Representation Generation for Efficient Waveform Diffusion Models cites this paper.

ReGen: Hierarchical Multi-Prompt Representation Generation for Efficient Waveform Diffusion Models FireRedTTS: A Foundation Text-To-Speech Framework for Industry-Level Generative Speech Applications

Reference 102

Resolution
unresolved
no resolver link, observed 2026-07-13T05:10:26.667731Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-13T05:10:26.667731Z digest=sha256:b1ef0bd34201d0de4a4d94b45447fbd9ea744b46d261401bc69d55c0a2c05cea

Observation 1c9468fe-a430-472d-8256-dce62ab51a78 · inbound

StellarTTS: Sparse Temporal Embedding for Low-Latency and Robust Speech Synthesis cites this paper.

StellarTTS: Sparse Temporal Embedding for Low-Latency and Robust Speech Synthesis FireRedTTS: A Foundation Text-To-Speech Framework for Industry-Level Generative Speech Applications

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-01T11:33:14.254397Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:33:14.254397Z digest=sha256:af28804bcf143f47114395a048c41c02f9e14a57c90bb1a5f18229785e2e7ac9

Observation c1eae712-ec60-4c69-95c0-911f42e4aa10 · inbound

Faster IndexTTS-2: Accelerating and Streaming Autoregressive Zero-Shot Text-to-Speech Synthesis on GPUs cites this paper.

Faster IndexTTS-2: Accelerating and Streaming Autoregressive Zero-Shot Text-to-Speech Synthesis on GPUs FireRedTTS: A Foundation Text-To-Speech Framework for Industry-Level Generative Speech Applications

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-01T08:40:27.699814Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:40:27.699814Z digest=sha256:fd3824206e2aa5d753b186e9bbba84ad0176b135634c01ca7e5f672c522d909a

Observation 680b1adf-f1f3-48ed-9db1-03bf0ed52b91 · inbound

Stable Autoregressive Speech Generation with Low-Frame-Rate High-Dimensional Continuous Tokens cites this paper.

Stable Autoregressive Speech Generation with Low-Frame-Rate High-Dimensional Continuous Tokens FireRedTTS: A Foundation Text-To-Speech Framework for Industry-Level Generative Speech Applications

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-03T08:35:49.194945Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T08:35:49.194945Z digest=sha256:272f4634247ba345015608bb147b87a423430c847c31407210b1cc0e5176cfa7