Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-01T08:40:31.629008Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 35 of 35 outbound references and 0 inbound Pith citation observations for arXiv:2607.21042.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-01T08:40:31.629008Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
35 of 35 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation fec16d61-d81e-443d-aed7-23f47a919d30 · outbound
Faster IndexTTS-2: Accelerating and Streaming Autoregressive Zero-Shot Text-to-Speech Synthesis on GPUs IndexTTS2: A breakthrough in emotionally expressive and duration- controlled auto-regressive zero-shot text-to-speech,
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d6f1d8cf-1a05-46ed-b690-81d6ebccf83d · outbound
Faster IndexTTS-2: Accelerating and Streaming Autoregressive Zero-Shot Text-to-Speech Synthesis on GPUs CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b4ee5dff-afe5-4298-a849-0a991a309525 · outbound
Faster IndexTTS-2: Accelerating and Streaming Autoregressive Zero-Shot Text-to-Speech Synthesis on GPUs CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 96e8430a-ea47-465d-8339-2ee17845339c · outbound
Faster IndexTTS-2: Accelerating and Streaming Autoregressive Zero-Shot Text-to-Speech Synthesis on GPUs Fish-Speech: Leveraging Large Language Models for Advanced Multilingual Text-to-Speech Synthesis
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 69ce9690-b3f7-48be-9344-2ff340b19901 · outbound
Faster IndexTTS-2: Accelerating and Streaming Autoregressive Zero-Shot Text-to-Speech Synthesis on GPUs Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 982ae0b7-0e29-4b7e-909f-bbe192749aeb · outbound
Faster IndexTTS-2: Accelerating and Streaming Autoregressive Zero-Shot Text-to-Speech Synthesis on GPUs Qwen3-TTS Technical Report
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c1eae712-ec60-4c69-95c0-911f42e4aa10 · outbound
Faster IndexTTS-2: Accelerating and Streaming Autoregressive Zero-Shot Text-to-Speech Synthesis on GPUs FireRedTTS: A Foundation Text-To-Speech Framework for Industry-Level Generative Speech Applications
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 62330e2b-7265-45bf-b130-cffe8ce10383 · outbound
Faster IndexTTS-2: Accelerating and Streaming Autoregressive Zero-Shot Text-to-Speech Synthesis on GPUs Streaming T5-based Text-to-Speech Synthesis with Limited Lookahead
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bf11fb27-ee48-4b88-b5fd-f345fae17195 · outbound
Faster IndexTTS-2: Accelerating and Streaming Autoregressive Zero-Shot Text-to-Speech Synthesis on GPUs F5-TTS: A fairytaler that fakes fluent and faithful speech with flow matching,
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 66841eb7-e957-4f41-93a6-585e617657eb · outbound
Faster IndexTTS-2: Accelerating and Streaming Autoregressive Zero-Shot Text-to-Speech Synthesis on GPUs MaskGCT: Zero-shot text-to-speech with masked generative codec transformer,
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 38630a00-faf5-4be8-9e30-678341d467ab · outbound
Faster IndexTTS-2: Accelerating and Streaming Autoregressive Zero-Shot Text-to-Speech Synthesis on GPUs E2 TTS: Embarrassingly easy fully non-autoregressive zero-shot TTS,
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9403fed2-bd57-4861-8c89-584338ba551b · outbound
Faster IndexTTS-2: Accelerating and Streaming Autoregressive Zero-Shot Text-to-Speech Synthesis on GPUs ZipV oice: Fast and high-quality zero-shot text-to-speech with flow matching,
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 31622478-221e-469f-b918-289070be5762 · outbound
Faster IndexTTS-2: Accelerating and Streaming Autoregressive Zero-Shot Text-to-Speech Synthesis on GPUs InstantSpeech: Instant synchronous text- to-speech synthesis for LLM-driven voice chatbots,
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 19192bbc-4e90-4097-92f5-3afb94e74293 · outbound
Faster IndexTTS-2: Accelerating and Streaming Autoregressive Zero-Shot Text-to-Speech Synthesis on GPUs NVIDIA TensorRT,
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 86f99114-cb22-410d-baca-080180d3b210 · outbound
Faster IndexTTS-2: Accelerating and Streaming Autoregressive Zero-Shot Text-to-Speech Synthesis on GPUs TensorRT-LLM,
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 24e9842b-7211-4936-bba4-543e5ce34603 · outbound
Faster IndexTTS-2: Accelerating and Streaming Autoregressive Zero-Shot Text-to-Speech Synthesis on GPUs Efficient memory management for large language model serving with PagedAttention,
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2a0c6fb1-3d97-4155-97db-0d5d453e3ac6 · outbound
Faster IndexTTS-2: Accelerating and Streaming Autoregressive Zero-Shot Text-to-Speech Synthesis on GPUs SGLang: Efficient execution of structured language model programs,
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 98505226-bd62-453b-bae7-f3951934fe91 · outbound
Faster IndexTTS-2: Accelerating and Streaming Autoregressive Zero-Shot Text-to-Speech Synthesis on GPUs ONNX Runtime,
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bf5dbdbf-4d1e-4b8f-8f10-633bd713f2d9 · outbound
Faster IndexTTS-2: Accelerating and Streaming Autoregressive Zero-Shot Text-to-Speech Synthesis on GPUs Language models are unsupervised multitask learners,
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 21d5285d-f0ab-4207-b46e-2eeb66e0ea26 · outbound
Faster IndexTTS-2: Accelerating and Streaming Autoregressive Zero-Shot Text-to-Speech Synthesis on GPUs Conformer: Convolution-augmented Transformer for Speech Recognition
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 12e67d0c-c9a7-4959-8eae-1defe7b0c0eb · outbound
Faster IndexTTS-2: Accelerating and Streaming Autoregressive Zero-Shot Text-to-Speech Synthesis on GPUs Flow matching for generative modeling,
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 178b5e08-1039-4ddc-adf0-a0e355e094af · outbound
Faster IndexTTS-2: Accelerating and Streaming Autoregressive Zero-Shot Text-to-Speech Synthesis on GPUs Scalable diffusion models with transformers,
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f122820f-8cdf-4ba4-877c-474326dcfa7b · outbound
Faster IndexTTS-2: Accelerating and Streaming Autoregressive Zero-Shot Text-to-Speech Synthesis on GPUs Classifier-Free Diffusion Guidance
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e5d0f3a8-6f14-4794-95e9-61cf4c3f3ecd · outbound
Faster IndexTTS-2: Accelerating and Streaming Autoregressive Zero-Shot Text-to-Speech Synthesis on GPUs BigVGAN: A Universal Neural Vocoder with Large-Scale Training
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5644a396-7193-45a7-9609-94d1141bbfeb · outbound
Faster IndexTTS-2: Accelerating and Streaming Autoregressive Zero-Shot Text-to-Speech Synthesis on GPUs W2v-BERT: Combining contrastive learning and masked lan- guage modeling for self-supervised speech pre-training,
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b32c882a-d706-4c04-87e0-27129d8f2570 · outbound
Faster IndexTTS-2: Accelerating and Streaming Autoregressive Zero-Shot Text-to-Speech Synthesis on GPUs Neural discrete representation learning,
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e823c9e9-3761-41c9-8ae0-0e6045ecd682 · outbound
Faster IndexTTS-2: Accelerating and Streaming Autoregressive Zero-Shot Text-to-Speech Synthesis on GPUs CAM++: A fast and efficient network for speaker verification using context-aware masking,
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation adba63e7-9b82-4b95-9037-b5d133917ecd · outbound
Faster IndexTTS-2: Accelerating and Streaming Autoregressive Zero-Shot Text-to-Speech Synthesis on GPUs Seed-TTS: A Family of High-Quality Versatile Speech Generation Models
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 93436678-4bc4-41f6-9dae-44af003e61e7 · outbound
Faster IndexTTS-2: Accelerating and Streaming Autoregressive Zero-Shot Text-to-Speech Synthesis on GPUs Common voice: A massively-multilingual speech corpus,
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8f384ed3-a057-4f77-a7bd-93e8e741b587 · outbound
Faster IndexTTS-2: Accelerating and Streaming Autoregressive Zero-Shot Text-to-Speech Synthesis on GPUs DiDiSpeech: A large scale mandarin speech corpus,
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d4c71c9f-d584-46f2-b0b7-c67e4afd8fab · outbound
Faster IndexTTS-2: Accelerating and Streaming Autoregressive Zero-Shot Text-to-Speech Synthesis on GPUs Robust speech recognition via large-scale weak super- vision,
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 67d7bdf6-6f35-4641-a6fc-4d39b8665843 · outbound
Faster IndexTTS-2: Accelerating and Streaming Autoregressive Zero-Shot Text-to-Speech Synthesis on GPUs FunASR: A fundamental end-to-end speech recognition toolkit,
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a80f1a77-2a32-4105-a6a5-d30f411bdd55 · outbound
Faster IndexTTS-2: Accelerating and Streaming Autoregressive Zero-Shot Text-to-Speech Synthesis on GPUs ECAPA-TDNN: emphasized channel attention, propagation and aggregation in TDNN based speaker verification,
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c09dd944-95cd-4d36-9d19-c2fe9cdbafea · outbound
Faster IndexTTS-2: Accelerating and Streaming Autoregressive Zero-Shot Text-to-Speech Synthesis on GPUs WavLM: Large-scale self-supervised pre- training for full stack speech processing,
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f26d5304-4e82-4d12-b4ec-f0eca8ea1b43 · outbound
Faster IndexTTS-2: Accelerating and Streaming Autoregressive Zero-Shot Text-to-Speech Synthesis on GPUs UTMOS: UTokyo-SaruLab system for V oiceMOS chal- lenge 2022,
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.