Pith. sign in

Paper Citation Record · LEDGER

TTS-1 Technical Report

As of 13 August 2026, this Paper Citation Record lists 48 of 48 outbound references and 4 inbound Pith citation observations for arXiv:2507.21138.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.21138 v1

Coverage vector

measured 48 of 48 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T15:02:03.077732Z

measured 52 of 52 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-03T06:36:31.523597Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T23:19:02.942476Z

Reference resolution

48 of 48 outbound references displayed

  • verified exact0
  • verified fuzzy17
  • unresolved31
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 1116d9ff-1bb6-45fe-8c3b-2d690506a98d · outbound

This paper cites The People's Speech: A Large-Scale Diverse English Speech Recognition Dataset for Commercial Usage.

TTS-1 Technical Report The People's Speech: A Large-Scale Diverse English Speech Recognition Dataset for Commercial Usage

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T15:02:02.830714Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:02:02.830714Z digest=sha256:789aaae286380cf1771b09c2a3a032c40e4416122eb7b9e906aeb7bb66ac33ce

Observation b6165c63-f5b3-490c-908c-7f7aa6a4e881 · outbound

This paper cites Yodas: Youtube-oriented dataset for audio and speech.

TTS-1 Technical Report Yodas: Youtube-oriented dataset for audio and speech

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:02:04.059235Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T15:02:02.837634Z digest=sha256:9b37419d61f55fb8537a4c3d1a248680436b3c483de62e25f0ce548189e14e42

Observation 9a8e03bc-ef3b-494a-8385-5600549cb4dc · outbound

This paper cites Emilia: An extensive, multilingual, and diverse speech dataset for large-scale speech generation.

TTS-1 Technical Report Emilia: An extensive, multilingual, and diverse speech dataset for large-scale speech generation

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:02:04.040769Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T15:02:02.843161Z digest=sha256:2fbda9592dbc9caebc17bb5075fc3de64bfd7d7358c5328fca90e4d8cab6cbe3

Observation a5392520-ba95-430c-a8b1-d442e5564146 · outbound

This paper cites Styletts 2: Towards human-level text-to-speech through style diffusion and adversarial training with large speech language models.

TTS-1 Technical Report Styletts 2: Towards human-level text-to-speech through style diffusion and adversarial training with large speech language models

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:02:04.025179Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T15:02:02.847985Z digest=sha256:6d31100a10471c9c894f487dd6c24778d31de8246d62cc1afaeed401e0d9acc9

Observation d993e778-bc6a-4974-a47b-8ab4e75e5f8b · outbound

This paper cites Fastspeech: Fast, robust and controllable text to speech.

TTS-1 Technical Report Fastspeech: Fast, robust and controllable text to speech

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:02:04.008207Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T15:02:02.854081Z digest=sha256:dd9e911ac1a906d3db5996c0fc2b5fb1a1dfb336d5670412238509dd6c4598b6

Observation 448ecfc7-dea8-42b5-b8e6-c5ee136542be · outbound

This paper cites Better speech synthesis through scaling.

TTS-1 Technical Report Better speech synthesis through scaling

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T15:02:02.858858Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:02:02.858858Z digest=sha256:49ec2ed78d05ac636c459de04ea5e6e16050d04a0f44993b128129e3839cc88e

Observation e9573b28-862d-43f8-928b-4878a58d2629 · outbound

This paper cites Conditional variational autoencoder with adversarial learning for end-to-end text-to-speech.

TTS-1 Technical Report Conditional variational autoencoder with adversarial learning for end-to-end text-to-speech

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:02:03.988145Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T15:02:02.865173Z digest=sha256:fe9647b1f0b779387c6508dfdf8dea3855d1b67568bfcbd857925baaeb7bf4d2

Observation e204a9e3-3816-48fe-84cf-4e4f73ff3a0c · outbound

This paper cites V oicebox: Text-guided multilingual universal speech generation at scale.

TTS-1 Technical Report V oicebox: Text-guided multilingual universal speech generation at scale

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:02:03.965541Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T15:02:02.870328Z digest=sha256:6fa1b96af7398ef2f17966984a62cb85d1feb8aa45d087c3c689b582038d6ea1

Observation 3fb837f9-85f5-43b5-add2-e03a99597d36 · outbound

This paper cites Language models are unsupervised multitask learners.

TTS-1 Technical Report Language models are unsupervised multitask learners

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T15:02:02.875785Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:02:02.875785Z digest=sha256:96b18fe1d429439a4f8deaae94fa63e44ca126e61d2e61d90f402cffa423fc28

Observation 7c2d9832-c784-4272-b9a4-311689e7b0d2 · outbound

This paper cites Training Compute-Optimal Large Language Models.

TTS-1 Technical Report Training Compute-Optimal Large Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T15:02:02.880909Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:02:02.880909Z digest=sha256:646c6c61ca450597dd32244599ddfec1511d7738b00272bdcd5ca995b758bad6

Observation d96c6f96-d5ae-4aa6-b75b-63251abf5781 · outbound

This paper cites Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers.

TTS-1 Technical Report Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T15:02:02.886470Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:02:02.886470Z digest=sha256:374cb3560ca879b621a32eaa7bdd651d2bbfd9e7075566d72f65ad17658e77b1

Observation bc5e8aef-029b-4609-8889-5fef279eb141 · outbound

This paper cites MiniMax-Speech: Intrinsic Zero-Shot Text-to-Speech with a Learnable Speaker Encoder.

TTS-1 Technical Report MiniMax-Speech: Intrinsic Zero-Shot Text-to-Speech with a Learnable Speaker Encoder

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T15:02:02.891456Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:02:02.891456Z digest=sha256:0731444f79f68b5940efb90a88b9cb51a86a25ccd8ed5cc67c2ea39f73eaf496

Observation aed9d272-a437-42ff-bec8-c77046dd3603 · outbound

This paper cites Seed-TTS: A Family of High-Quality Versatile Speech Generation Models.

TTS-1 Technical Report Seed-TTS: A Family of High-Quality Versatile Speech Generation Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T15:02:02.897698Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:02:02.897698Z digest=sha256:0a1a4ac155c01de91d81711de271db78c5cc72e3c3a7abee64a79fb0aa828ec9

Observation a3c777f1-cd98-4493-aae4-b99d5ea637dd · outbound

This paper cites CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens.

TTS-1 Technical Report CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T15:02:02.904024Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:02:02.904024Z digest=sha256:23c7ddb7e2c495e69c7cb54eb3d62e34e2165dbc008df969a25d3b7a9436b5bd

Observation a2cd3a78-d85f-44a1-9c4b-f1c80281eac1 · outbound

This paper cites CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models.

TTS-1 Technical Report CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T15:02:02.908687Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:02:02.908687Z digest=sha256:7ff3e3e3cd4cfa05d4b502f989d871587a30288d71496d64b86ca829250016e6

Observation bd31abf3-cdc3-4957-9bde-e4ca9251251f · outbound

This paper cites Redpajama: an open dataset for training large language models.

TTS-1 Technical Report Redpajama: an open dataset for training large language models

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:02:03.926487Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T15:02:02.913946Z digest=sha256:23345f0e137330e5055c9395f99625c37e7ca0f525bd3a6aa7504aea91746615

Observation 669d28f8-b554-4228-a0ee-fdb73850def1 · outbound

This paper cites Open instruction generalist (oig) dataset.

TTS-1 Technical Report Open instruction generalist (oig) dataset

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:02:03.903451Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T15:02:02.919362Z digest=sha256:2d1f9bb076c6a454ece13799e1484867e331df34b0b2298b9e4e88036cd6ec9b

Observation 1f6b538f-24df-4cef-b876-74eb4f3f71dc · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

TTS-1 Technical Report DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T15:02:02.924574Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:02:02.924574Z digest=sha256:3d13023e2a33f24607d5488ebe6d07580b6f5e1269d26a54892e87f1b396f44f

Observation 5a853971-6149-46e6-9863-4098efe1267e · outbound

This paper cites Llasa: Scaling Train-Time and Inference-Time Compute for Llama-based Speech Synthesis.

TTS-1 Technical Report Llasa: Scaling Train-Time and Inference-Time Compute for Llama-based Speech Synthesis

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T15:02:02.930108Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:02:02.930108Z digest=sha256:69c90959ce13ccba5947883ee7baafa65f9d3ea9bb4e1fdce5706ca43dfd3f68

Observation 61e20d6f-5552-4c0c-a9a8-b4a2d805e76e · outbound

This paper cites Wavlm: Large-scale self-supervised pre-training for full stack speech processing.

TTS-1 Technical Report Wavlm: Large-scale self-supervised pre-training for full stack speech processing

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:02:03.884270Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T15:02:02.935030Z digest=sha256:a2871c53c35174f94d9315b8365204b593cd25626f6cc1ad58f8ed52800922c6

Observation befd186d-6c49-493d-8b30-f788cea54faa · outbound

This paper cites Dnsmos p.

TTS-1 Technical Report Dnsmos p

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:02:03.865444Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T15:02:02.940134Z digest=sha256:2fa9332573bcb08ce03962d9be79b3ba340c5cfaa1aa61437c6013a0904daa99

Observation 3d900436-35b3-4f17-b713-eba6b6f78fa3 · outbound

This paper cites Lora: Low-rank adaptation of large language models.

TTS-1 Technical Report Lora: Low-rank adaptation of large language models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T15:02:02.944597Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:02:02.944597Z digest=sha256:ab2ae1291f185c7ef7f8e5bb9415dd1f8a9166a888b6e3110cdc0026c4d215d9

Observation bdf5b87d-9598-4cab-88a2-5c88c38d010a · outbound

This paper cites Deep residual learning for image recognition.

TTS-1 Technical Report Deep residual learning for image recognition

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T15:02:02.949788Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:02:02.949788Z digest=sha256:189aa12c639aa6d7803a093fa8f21dc1e7e07dd8959cf8235dff990d0c69416a

Observation 089d6aab-8a5f-4074-bdbd-071b1dd9a10c · outbound

This paper cites The Llama 3 Herd of Models.

TTS-1 Technical Report The Llama 3 Herd of Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T15:02:02.954136Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:02:02.954136Z digest=sha256:b71e10d7a9a2dfc451b51fe2c192628fdd8cb9c67fb94a81144bf1f035735696

Observation 6e452096-9f3d-4855-8a96-101daea28de3 · outbound

This paper cites Matrix multiplication background user’s guide.

TTS-1 Technical Report Matrix multiplication background user’s guide

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:02:03.825369Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T15:02:02.959872Z digest=sha256:f80f12d9b690b275eecc998f70160ad3462b1f6195e7847a54fcbaf649bf8603

Observation 63165417-09a3-4464-b43e-d01040b7e49f · outbound

This paper cites Initializing new word embeddings for pretrained language models.

TTS-1 Technical Report Initializing new word embeddings for pretrained language models

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:02:03.808174Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T15:02:02.964733Z digest=sha256:220ce50ded64703101753313c6ec4bea4f611c120a118baad9671355a3935119

Observation dee61c16-a693-413d-baec-a8d71bc82094 · outbound

This paper cites BigCodec: Pushing the Limits of Low-Bitrate Neural Speech Codec.

TTS-1 Technical Report BigCodec: Pushing the Limits of Low-Bitrate Neural Speech Codec

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T15:02:02.970136Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:02:02.970136Z digest=sha256:f86a0110cdff9d1d9daee3c9b4b29da6728a055cf86b34cf9b496300fc40cf9e

Observation e7372153-22da-4005-9a06-a28caf869736 · outbound

This paper cites Hifi-gan: Generative adversarial networks for efficient and high fidelity speech synthesis.

TTS-1 Technical Report Hifi-gan: Generative adversarial networks for efficient and high fidelity speech synthesis

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:02:03.788491Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T15:02:02.974412Z digest=sha256:b14adcb8c6fa9c1056c4786edbedecc48d8cfbc8bd9e067a403c845461157df1

Observation 331a3e81-ee3b-4418-9bf2-aac95671b668 · outbound

This paper cites High Fidelity Neural Audio Compression.

TTS-1 Technical Report High Fidelity Neural Audio Compression

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T15:02:02.980832Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:02:02.980832Z digest=sha256:ed40b63bec2e7abd4e3209501d048a73139fa99abfc3741363e397c55069e459

Observation dfac856f-3877-4bf0-9bdb-d925ad897b44 · outbound

This paper cites Adding Instructions during Pretraining: Effective Way of Controlling Toxicity in Language Models.

TTS-1 Technical Report Adding Instructions during Pretraining: Effective Way of Controlling Toxicity in Language Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T15:02:02.985709Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:02:02.985709Z digest=sha256:e0c598a49b900e8905be6d0b3fda801053f95fb7485471e47d692a759f9585e0

Observation b2d0e952-b052-4dc2-8a25-3c9367e1cfe1 · outbound

This paper cites A neural probabilistic language model.

TTS-1 Technical Report A neural probabilistic language model

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T15:02:02.990541Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:02:02.990541Z digest=sha256:6b0a1139fe9a1f6c4afd12dbc11b4560fbc5e88995d96e8c650b9ac02264bf61

Observation 9133858a-f8b9-420b-8e82-19916efa00be · outbound

This paper cites FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning.

TTS-1 Technical Report FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T15:02:02.997072Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:02:02.997072Z digest=sha256:12a26bbe8d549deb34f9fba011662d70977fa414f53bccae09d7d2285459279a

Observation 29c6c949-fa70-4f7c-8811-4deab52cb8f1 · outbound

This paper cites Adam: A Method for Stochastic Optimization.

TTS-1 Technical Report Adam: A Method for Stochastic Optimization

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T15:02:03.001891Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:02:03.001891Z digest=sha256:53af3be25a5e5e3727421e80c0beca60e305396a60d25d997502f8b1d0acecc3

Observation dc53bb10-0b74-41a7-93a0-bac436025d36 · outbound

This paper cites PyTorch Distributed: Experiences on Accelerating Data Parallel Training.

TTS-1 Technical Report PyTorch Distributed: Experiences on Accelerating Data Parallel Training

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T15:02:03.006360Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:02:03.006360Z digest=sha256:2f5f7b7227d4b4138097d701ec39e45454211a98b9a688e83321ea780fd6a125

Observation d962fbbe-3828-4961-9347-cf0e52fb0db0 · outbound

This paper cites PyTorch FSDP: Experiences on Scaling Fully Sharded Data Parallel.

TTS-1 Technical Report PyTorch FSDP: Experiences on Scaling Fully Sharded Data Parallel

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T15:02:03.011140Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:02:03.011140Z digest=sha256:5a6690a2e24e6e5a4a88905051de438c242f91c2f0f8def45e7e3754d13e68e5

Observation e27dba3b-4a19-4c67-a1b4-2ed803652cb5 · outbound

This paper cites Parler-tts.

TTS-1 Technical Report Parler-tts

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:02:03.755019Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T15:02:03.016065Z digest=sha256:828f6e5c14449f25ca93e7fbceec1332141cfd9be4ec3015c8c7c1a54b309da5

Observation 8a5c05da-c5d0-4796-bf6d-9303960c483b · outbound

This paper cites Zero: Memory optimizations toward training trillion parameter models.

TTS-1 Technical Report Zero: Memory optimizations toward training trillion parameter models

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:02:03.737535Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T15:02:03.021572Z digest=sha256:c17589828485b28d2666fbfa20ff30f52679c5fc26e6d360315841c0238fe28d

Observation 52ded5a3-2265-49cd-b401-61758dbf0392 · outbound

This paper cites CosyVoice 3: Towards In-the-wild Speech Generation via Scaling-up and Post-training.

TTS-1 Technical Report CosyVoice 3: Towards In-the-wild Speech Generation via Scaling-up and Post-training

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T15:02:03.026897Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:02:03.026897Z digest=sha256:2da67aee502bae3875072ad5a2ca865fddbacffcf4ca2ef1a59ac8e3e9ae1ac4

Observation 3640e72a-bbf6-47a0-89f5-b238dca3c766 · outbound

This paper cites Enhancing Zero-shot Text-to-Speech Synthesis with Human Feedback.

TTS-1 Technical Report Enhancing Zero-shot Text-to-Speech Synthesis with Human Feedback

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T15:02:03.032073Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:02:03.032073Z digest=sha256:5d9a8c4607908eb54113a3d560261ae1564a6a53d94087643161e08f5e31d9d8

Observation 4a88a8cd-6997-471a-8f21-bc960cf587e3 · outbound

This paper cites DeepSeek-V3 Technical Report.

TTS-1 Technical Report DeepSeek-V3 Technical Report

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T15:02:03.037268Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:02:03.037268Z digest=sha256:38784bc3b6d1ec2e61c09ff4a344a9c8a9a9dbfe3d2f66ed40d744ecfc5f28a6

Observation 63c76d66-c01b-4edd-9a8a-2caebbdb197f · outbound

This paper cites DeepSeek LLM: Scaling Open-Source Language Models with Longtermism.

TTS-1 Technical Report DeepSeek LLM: Scaling Open-Source Language Models with Longtermism

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T15:02:03.042360Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:02:03.042360Z digest=sha256:141758bc8cdab978b0904a3cd9f0408b7c49b99a89425fc25f64b19285eed0c2

Observation 4ca58588-1f1f-4943-8194-60f8dbc61c15 · outbound

This paper cites Direct preference optimization: Your language model is secretly a reward model.

TTS-1 Technical Report Direct preference optimization: Your language model is secretly a reward model

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T15:02:03.047248Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:02:03.047248Z digest=sha256:5512494d6d313ae811e438cbf439b90c3e73ac69195ad382e614aff3691fa6b1

Observation f720f069-b85f-48e8-afe8-ae422da1397e · outbound

This paper cites Understanding R1-Zero-Like Training: A Critical Perspective.

TTS-1 Technical Report Understanding R1-Zero-Like Training: A Critical Perspective

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T15:02:03.052195Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:02:03.052195Z digest=sha256:26110101ec69a3ddcf4b5d32307b83d623a98b92a15f329b774014afe5c5155f

Observation b9b904ce-8b94-46ad-9c3d-99590258e76e · outbound

This paper cites Robust speech recognition via large-scale weak supervision.

TTS-1 Technical Report Robust speech recognition via large-scale weak supervision

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T15:02:03.057567Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:02:03.057567Z digest=sha256:a2e04d5d671dfe89efccc885909c0c282c9358a7fefd8bdb36c94117a8736aa0

Observation 6d2518ba-0c5b-4908-ace2-7d94f7171699 · outbound

This paper cites HuggingFace's Transformers: State-of-the-art Natural Language Processing.

TTS-1 Technical Report HuggingFace's Transformers: State-of-the-art Natural Language Processing

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T15:02:03.063141Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:02:03.063141Z digest=sha256:09f310aacdba099e339b94b7992ca091024d89ddc834d930b52ea21a5819bc30

Observation bd80dfc2-0d2c-400a-818f-ebd963c3eeab · outbound

This paper cites PyTorch: An Imperative Style, High-Performance Deep Learning Library.

TTS-1 Technical Report PyTorch: An Imperative Style, High-Performance Deep Learning Library

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T15:02:03.068483Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:02:03.068483Z digest=sha256:fc50d7db02d9b8e592c6463f76a2a993be2c4e6c9251b4b51cd8d556d0ec75f4

Observation a606eda4-cfd8-44ee-85bf-7b99585a2b9b · outbound

This paper cites Pytorch lightning.

TTS-1 Technical Report Pytorch lightning

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:02:03.690672Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T15:02:03.073406Z digest=sha256:c6f4f06974f5ae1550c4c47459be4ad7e769bb21e22c4dc8917a6a44ea53e50d

Observation c88b760d-9f09-4e4e-a31c-fb9f6ae7cc7b · outbound

This paper cites Efficient memory management for large language model serving with pagedattention.

TTS-1 Technical Report Efficient memory management for large language model serving with pagedattention

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:02:03.674403Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T15:02:03.077732Z digest=sha256:34134a15cc41b0e293348d379879df3db415606765cd563c2ad3bfc7b8a860b7

Pith citing papers

Observation 64fed19a-4d58-49a6-867f-8977988efecd · inbound

Evaluating and Rewarding LALMs for Expressive Role-Play TTS via Mean Continuation Log-Probability cites this paper.

Evaluating and Rewarding LALMs for Expressive Role-Play TTS via Mean Continuation Log-Probability TTS-1 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-03T06:36:31.523597Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:36:31.523597Z digest=sha256:2474b2ab69bf0140efe544a13ea122474cbf7382225815bc79a69de4bf173382

Observation b76085d3-b265-4c38-beb9-aa07f37b5812 · inbound

JaiTTS: A Thai Voice Cloning Model cites this paper.

JaiTTS: A Thai Voice Cloning Model TTS-1 Technical Report

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-12T10:21:28.850669Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-05-07T06:34:45.357695Z digest=sha256:4d0937dac0d6f795f3b5dd129154029867bb0a61617238f750a6e76ce44be0d1

Observation cca52b6c-b63b-414c-9517-cb5c8bc46c2a · inbound

JaiTTS: A Thai Voice Cloning Model cites this paper.

JaiTTS: A Thai Voice Cloning Model TTS-1 Technical Report

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-11T22:16:26.706638Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-05-08T03:09:24.541657Z digest=sha256:4d337ecfa49bab8f6ae84708958d27598b8bb434b75a70cc164bf311463a42c8

Observation cc3f5c45-0029-4f23-a0c6-f49bfa10d585 · inbound

Reliable Neural-Codec Text-to-Speech by ASR Self-Verification and Distillation: Near-Zero Catastrophic Failures Across Models and Codecs cites this paper.

Reliable Neural-Codec Text-to-Speech by ASR Self-Verification and Distillation: Near-Zero Catastrophic Failures Across Models and Codecs TTS-1 Technical Report

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-07-03T23:19:02.945574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-06-26T22:46:47.786929Z digest=sha256:afc2ba51cbabafe4654c7303c39b11937d4b714acf8d248cd6b7fb1b77ec8837