Pith. sign in

Paper Citation Record · LEDGER

Faster IndexTTS-2: Accelerating and Streaming Autoregressive Zero-Shot Text-to-Speech Synthesis on GPUs

As of 7 August 2026, this Paper Citation Record lists 35 of 35 outbound references and 0 inbound Pith citation observations for arXiv:2607.21042.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.21042 v1

Coverage vector

measured 35 of 35 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-01T08:40:31.629008Z

measured 35 of 35 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

35 of 35 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved35
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation fec16d61-d81e-443d-aed7-23f47a919d30 · outbound

This paper cites IndexTTS2: A breakthrough in emotionally expressive and duration- controlled auto-regressive zero-shot text-to-speech,.

Faster IndexTTS-2: Accelerating and Streaming Autoregressive Zero-Shot Text-to-Speech Synthesis on GPUs IndexTTS2: A breakthrough in emotionally expressive and duration- controlled auto-regressive zero-shot text-to-speech,

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-01T08:40:27.242144Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:40:27.242144Z digest=sha256:756c934dd4c7b7daca2b041d93b73cca56e90905b90845b3dfcfd0fc20d936af

Observation d6f1d8cf-1a05-46ed-b690-81d6ebccf83d · outbound

This paper cites CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens.

Faster IndexTTS-2: Accelerating and Streaming Autoregressive Zero-Shot Text-to-Speech Synthesis on GPUs CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-01T08:40:27.323776Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:40:27.323776Z digest=sha256:e37cb63b67640d0a7ea8f9b7685bb0753989ff8c1e20a3512975a455df0023e8

Observation b4ee5dff-afe5-4298-a849-0a991a309525 · outbound

This paper cites CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models.

Faster IndexTTS-2: Accelerating and Streaming Autoregressive Zero-Shot Text-to-Speech Synthesis on GPUs CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-01T08:40:27.396520Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:40:27.396520Z digest=sha256:25428a379fa1d16204d18cf24ec06e1b9694446f5b214f584909577f56cd8925

Observation 96e8430a-ea47-465d-8339-2ee17845339c · outbound

This paper cites Fish-Speech: Leveraging Large Language Models for Advanced Multilingual Text-to-Speech Synthesis.

Faster IndexTTS-2: Accelerating and Streaming Autoregressive Zero-Shot Text-to-Speech Synthesis on GPUs Fish-Speech: Leveraging Large Language Models for Advanced Multilingual Text-to-Speech Synthesis

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-01T08:40:27.509915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:40:27.509915Z digest=sha256:a7fd5eec80770e90161024228d3a2c48226751dc2a968e12a075f20493094091

Observation 69ce9690-b3f7-48be-9344-2ff340b19901 · outbound

This paper cites Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens.

Faster IndexTTS-2: Accelerating and Streaming Autoregressive Zero-Shot Text-to-Speech Synthesis on GPUs Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-01T08:40:27.564412Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:40:27.564412Z digest=sha256:cf635b5dedd6e39864e53a61c9d18f9041217972fb72b78f3c439a3c30ff3238

Observation 982ae0b7-0e29-4b7e-909f-bbe192749aeb · outbound

This paper cites Qwen3-TTS Technical Report.

Faster IndexTTS-2: Accelerating and Streaming Autoregressive Zero-Shot Text-to-Speech Synthesis on GPUs Qwen3-TTS Technical Report

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-01T08:40:27.601726Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:40:27.601726Z digest=sha256:baff00544ee3334e024322a75af01c07f97d9280d67d10be94e74f8702fb7953

Observation c1eae712-ec60-4c69-95c0-911f42e4aa10 · outbound

This paper cites FireRedTTS: A Foundation Text-To-Speech Framework for Industry-Level Generative Speech Applications.

Faster IndexTTS-2: Accelerating and Streaming Autoregressive Zero-Shot Text-to-Speech Synthesis on GPUs FireRedTTS: A Foundation Text-To-Speech Framework for Industry-Level Generative Speech Applications

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-01T08:40:27.699814Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:40:27.699814Z digest=sha256:fd3824206e2aa5d753b186e9bbba84ad0176b135634c01ca7e5f672c522d909a

Observation 62330e2b-7265-45bf-b130-cffe8ce10383 · outbound

This paper cites Streaming T5-based Text-to-Speech Synthesis with Limited Lookahead.

Faster IndexTTS-2: Accelerating and Streaming Autoregressive Zero-Shot Text-to-Speech Synthesis on GPUs Streaming T5-based Text-to-Speech Synthesis with Limited Lookahead

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-01T08:40:27.846912Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:40:27.846912Z digest=sha256:cf35d45982c520106e7a3331ce58fa0e8bb2f0880142603a4e5b5375f180ed28

Observation bf11fb27-ee48-4b88-b5fd-f345fae17195 · outbound

This paper cites F5-TTS: A fairytaler that fakes fluent and faithful speech with flow matching,.

Faster IndexTTS-2: Accelerating and Streaming Autoregressive Zero-Shot Text-to-Speech Synthesis on GPUs F5-TTS: A fairytaler that fakes fluent and faithful speech with flow matching,

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-01T08:40:28.013513Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:40:28.013513Z digest=sha256:70c9d2b0abc3b9d20869a2a66bdfd7f0e76fa5c941f5c17559b2977a91ef5f2e

Observation 66841eb7-e957-4f41-93a6-585e617657eb · outbound

This paper cites MaskGCT: Zero-shot text-to-speech with masked generative codec transformer,.

Faster IndexTTS-2: Accelerating and Streaming Autoregressive Zero-Shot Text-to-Speech Synthesis on GPUs MaskGCT: Zero-shot text-to-speech with masked generative codec transformer,

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-01T08:40:28.181140Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:40:28.181140Z digest=sha256:a2bd710f2f0548489e6a7b85227d074011f1aa56ad31820650a65ca45d19e192

Observation 38630a00-faf5-4be8-9e30-678341d467ab · outbound

This paper cites E2 TTS: Embarrassingly easy fully non-autoregressive zero-shot TTS,.

Faster IndexTTS-2: Accelerating and Streaming Autoregressive Zero-Shot Text-to-Speech Synthesis on GPUs E2 TTS: Embarrassingly easy fully non-autoregressive zero-shot TTS,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-01T08:40:28.346583Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:40:28.346583Z digest=sha256:44234d93e3cda88993bc3c9743d1b1ca8374b42f5567607fc2ca75784656e43d

Observation 9403fed2-bd57-4861-8c89-584338ba551b · outbound

This paper cites ZipV oice: Fast and high-quality zero-shot text-to-speech with flow matching,.

Faster IndexTTS-2: Accelerating and Streaming Autoregressive Zero-Shot Text-to-Speech Synthesis on GPUs ZipV oice: Fast and high-quality zero-shot text-to-speech with flow matching,

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-01T08:40:28.514609Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:40:28.514609Z digest=sha256:4f68cf1cb914634caa336df313a9e9c4e20519197ddb69032d521cb07d291a6c

Observation 31622478-221e-469f-b918-289070be5762 · outbound

This paper cites InstantSpeech: Instant synchronous text- to-speech synthesis for LLM-driven voice chatbots,.

Faster IndexTTS-2: Accelerating and Streaming Autoregressive Zero-Shot Text-to-Speech Synthesis on GPUs InstantSpeech: Instant synchronous text- to-speech synthesis for LLM-driven voice chatbots,

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-01T08:40:28.656148Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:40:28.656148Z digest=sha256:be65790a32ad33e7f6aefa75b05a5ea9d13c12669799adc76d9dacfb331d0b7e

Observation 19192bbc-4e90-4097-92f5-3afb94e74293 · outbound

This paper cites NVIDIA TensorRT,.

Faster IndexTTS-2: Accelerating and Streaming Autoregressive Zero-Shot Text-to-Speech Synthesis on GPUs NVIDIA TensorRT,

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-01T08:40:28.824143Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:40:28.824143Z digest=sha256:f811af86a49f85501f5a482c4951aefb4e687b8b62e6c567df2b09297c5bb2a4

Observation 86f99114-cb22-410d-baca-080180d3b210 · outbound

This paper cites TensorRT-LLM,.

Faster IndexTTS-2: Accelerating and Streaming Autoregressive Zero-Shot Text-to-Speech Synthesis on GPUs TensorRT-LLM,

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-01T08:40:28.954933Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:40:28.954933Z digest=sha256:120b68773cb087774a67ca784716c6809ef5548ef33161edebde89e62dea29aa

Observation 24e9842b-7211-4936-bba4-543e5ce34603 · outbound

This paper cites Efficient memory management for large language model serving with PagedAttention,.

Faster IndexTTS-2: Accelerating and Streaming Autoregressive Zero-Shot Text-to-Speech Synthesis on GPUs Efficient memory management for large language model serving with PagedAttention,

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-01T08:40:29.041783Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:40:29.041783Z digest=sha256:855a40ae2997a59939c9ab5bad42c54e2e0b4eaa4ca474b28bf929cba5c88ad8

Observation 2a0c6fb1-3d97-4155-97db-0d5d453e3ac6 · outbound

This paper cites SGLang: Efficient execution of structured language model programs,.

Faster IndexTTS-2: Accelerating and Streaming Autoregressive Zero-Shot Text-to-Speech Synthesis on GPUs SGLang: Efficient execution of structured language model programs,

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-01T08:40:29.152701Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:40:29.152701Z digest=sha256:6e8fe35afe3a8d1bf585b0f45c7ca4404dbd5aa3dcd7c02f8413c4202d3a5627

Observation 98505226-bd62-453b-bae7-f3951934fe91 · outbound

This paper cites ONNX Runtime,.

Faster IndexTTS-2: Accelerating and Streaming Autoregressive Zero-Shot Text-to-Speech Synthesis on GPUs ONNX Runtime,

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-01T08:40:29.228782Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:40:29.228782Z digest=sha256:9698444aa58eaf974a9c3ecfd564b296c9eed2b0b15f2f8a87a8857399baa53e

Observation bf5dbdbf-4d1e-4b8f-8f10-633bd713f2d9 · outbound

This paper cites Language models are unsupervised multitask learners,.

Faster IndexTTS-2: Accelerating and Streaming Autoregressive Zero-Shot Text-to-Speech Synthesis on GPUs Language models are unsupervised multitask learners,

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-01T08:40:29.342098Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:40:29.342098Z digest=sha256:c5bbc0de8ed3e5e8aa3e0fd24dd0ac34817590c48ab3cb1f00c5e3ab12d9c5a5

Observation 21d5285d-f0ab-4207-b46e-2eeb66e0ea26 · outbound

This paper cites Conformer: Convolution-augmented Transformer for Speech Recognition.

Faster IndexTTS-2: Accelerating and Streaming Autoregressive Zero-Shot Text-to-Speech Synthesis on GPUs Conformer: Convolution-augmented Transformer for Speech Recognition

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-01T08:40:29.429613Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:40:29.429613Z digest=sha256:4de1b4bdb4e1e28d1ec183523bf68ff7d71b9bb65a8cf781652b57f36d588b5b

Observation 12e67d0c-c9a7-4959-8eae-1defe7b0c0eb · outbound

This paper cites Flow matching for generative modeling,.

Faster IndexTTS-2: Accelerating and Streaming Autoregressive Zero-Shot Text-to-Speech Synthesis on GPUs Flow matching for generative modeling,

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-01T08:40:29.553448Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:40:29.553448Z digest=sha256:c23642a36127c911aec5b79367616ecb9cf838d2aa3d81ce4d848a682af50f9e

Observation 178b5e08-1039-4ddc-adf0-a0e355e094af · outbound

This paper cites Scalable diffusion models with transformers,.

Faster IndexTTS-2: Accelerating and Streaming Autoregressive Zero-Shot Text-to-Speech Synthesis on GPUs Scalable diffusion models with transformers,

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-01T08:40:29.648884Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:40:29.648884Z digest=sha256:4f591c8bc2a4d14040820a8c0e990fe6428d190438d524a041837af43702e25c

Observation f122820f-8cdf-4ba4-877c-474326dcfa7b · outbound

This paper cites Classifier-Free Diffusion Guidance.

Faster IndexTTS-2: Accelerating and Streaming Autoregressive Zero-Shot Text-to-Speech Synthesis on GPUs Classifier-Free Diffusion Guidance

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-01T08:40:29.739573Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:40:29.739573Z digest=sha256:dc8761992cde010875176053eb5083e08aaa5359292cdc854da61795bf60eb67

Observation e5d0f3a8-6f14-4794-95e9-61cf4c3f3ecd · outbound

This paper cites BigVGAN: A Universal Neural Vocoder with Large-Scale Training.

Faster IndexTTS-2: Accelerating and Streaming Autoregressive Zero-Shot Text-to-Speech Synthesis on GPUs BigVGAN: A Universal Neural Vocoder with Large-Scale Training

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-01T08:40:29.851867Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:40:29.851867Z digest=sha256:d499a9d5faeadd0e7bed84bc466e1fd82a830fa243a95a5e3475e5312d282295

Observation 5644a396-7193-45a7-9609-94d1141bbfeb · outbound

This paper cites W2v-BERT: Combining contrastive learning and masked lan- guage modeling for self-supervised speech pre-training,.

Faster IndexTTS-2: Accelerating and Streaming Autoregressive Zero-Shot Text-to-Speech Synthesis on GPUs W2v-BERT: Combining contrastive learning and masked lan- guage modeling for self-supervised speech pre-training,

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-01T08:40:29.993395Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:40:29.993395Z digest=sha256:037ecf24bd03360983ccac1903194299fc5da6fa90f8bd33e74689ee4f6de0eb

Observation b32c882a-d706-4c04-87e0-27129d8f2570 · outbound

This paper cites Neural discrete representation learning,.

Faster IndexTTS-2: Accelerating and Streaming Autoregressive Zero-Shot Text-to-Speech Synthesis on GPUs Neural discrete representation learning,

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-01T08:40:30.127935Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:40:30.127935Z digest=sha256:3606b40b2db5810f8b21ad1c3b934fec3258768ce9ea2bb9bb2d6f40fd152e4d

Observation e823c9e9-3761-41c9-8ae0-0e6045ecd682 · outbound

This paper cites CAM++: A fast and efficient network for speaker verification using context-aware masking,.

Faster IndexTTS-2: Accelerating and Streaming Autoregressive Zero-Shot Text-to-Speech Synthesis on GPUs CAM++: A fast and efficient network for speaker verification using context-aware masking,

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-01T08:40:30.329231Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:40:30.329231Z digest=sha256:72768e338a00ca7c3520a7d4bacdd0525a17edd1e5bfad7421bb3431792dd30b

Observation adba63e7-9b82-4b95-9037-b5d133917ecd · outbound

This paper cites Seed-TTS: A Family of High-Quality Versatile Speech Generation Models.

Faster IndexTTS-2: Accelerating and Streaming Autoregressive Zero-Shot Text-to-Speech Synthesis on GPUs Seed-TTS: A Family of High-Quality Versatile Speech Generation Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-01T08:40:30.520790Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:40:30.520790Z digest=sha256:a08b8d7e14f39b732ebb1ce2eeb3d5dfa893101162397bed1c216a47e07382e2

Observation 93436678-4bc4-41f6-9dae-44af003e61e7 · outbound

This paper cites Common voice: A massively-multilingual speech corpus,.

Faster IndexTTS-2: Accelerating and Streaming Autoregressive Zero-Shot Text-to-Speech Synthesis on GPUs Common voice: A massively-multilingual speech corpus,

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-01T08:40:30.659984Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:40:30.659984Z digest=sha256:3b59a6136e39bf159e35b2d5c455bc6d89d5488268364454c3ae79f0b538008a

Observation 8f384ed3-a057-4f77-a7bd-93e8e741b587 · outbound

This paper cites DiDiSpeech: A large scale mandarin speech corpus,.

Faster IndexTTS-2: Accelerating and Streaming Autoregressive Zero-Shot Text-to-Speech Synthesis on GPUs DiDiSpeech: A large scale mandarin speech corpus,

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-01T08:40:30.811354Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:40:30.811354Z digest=sha256:58d064af04e708f8895d18ca6874560e2e14075c864a3a822d601aebd073a9fd

Observation d4c71c9f-d584-46f2-b0b7-c67e4afd8fab · outbound

This paper cites Robust speech recognition via large-scale weak super- vision,.

Faster IndexTTS-2: Accelerating and Streaming Autoregressive Zero-Shot Text-to-Speech Synthesis on GPUs Robust speech recognition via large-scale weak super- vision,

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-01T08:40:30.981779Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:40:30.981779Z digest=sha256:2fc04ce186411c5114ffb2ae2e308f6bbc37305d2d30f395a1acb74f3d37e495

Observation 67d7bdf6-6f35-4641-a6fc-4d39b8665843 · outbound

This paper cites FunASR: A fundamental end-to-end speech recognition toolkit,.

Faster IndexTTS-2: Accelerating and Streaming Autoregressive Zero-Shot Text-to-Speech Synthesis on GPUs FunASR: A fundamental end-to-end speech recognition toolkit,

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-01T08:40:31.162278Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:40:31.162278Z digest=sha256:6245d3b0fdc78aa6b4ea223ba12ed63cb8e6396d58dc433e3a5fbf553ec66607

Observation a80f1a77-2a32-4105-a6a5-d30f411bdd55 · outbound

This paper cites ECAPA-TDNN: emphasized channel attention, propagation and aggregation in TDNN based speaker verification,.

Faster IndexTTS-2: Accelerating and Streaming Autoregressive Zero-Shot Text-to-Speech Synthesis on GPUs ECAPA-TDNN: emphasized channel attention, propagation and aggregation in TDNN based speaker verification,

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-01T08:40:31.274726Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:40:31.274726Z digest=sha256:76c0b40ababe0c539070744a0f85079599d6015b89cc26486461371fe841316f

Observation c09dd944-95cd-4d36-9d19-c2fe9cdbafea · outbound

This paper cites WavLM: Large-scale self-supervised pre- training for full stack speech processing,.

Faster IndexTTS-2: Accelerating and Streaming Autoregressive Zero-Shot Text-to-Speech Synthesis on GPUs WavLM: Large-scale self-supervised pre- training for full stack speech processing,

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-01T08:40:31.417593Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:40:31.417593Z digest=sha256:027747cad652b97d3d638075c031376ea7bcc7bac64fea2eb2e6aa70013a0eea

Observation f26d5304-4e82-4d12-b4ec-f0eca8ea1b43 · outbound

This paper cites UTMOS: UTokyo-SaruLab system for V oiceMOS chal- lenge 2022,.

Faster IndexTTS-2: Accelerating and Streaming Autoregressive Zero-Shot Text-to-Speech Synthesis on GPUs UTMOS: UTokyo-SaruLab system for V oiceMOS chal- lenge 2022,

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-01T08:40:31.629008Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T08:40:31.629008Z digest=sha256:1fb1ff5ac3ba6edb95ec11316413b31ad34eeaaf5ae9b12ca74c73b1c99c21a1

Pith citing papers

No inbound Pith citation observations are available.