Pith. sign in

Paper Citation Record · LEDGER

TTS-1 Technical Report

As of 8 August 2026, this Paper Citation Record lists 48 of 48 outbound references and 4 inbound Pith citation observations for arXiv:2507.21138.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.21138 v1

Coverage vector

measured 48 of 48 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T15:02:03.077732Z

measured 52 of 52 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-03T06:36:31.523597Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T23:19:02.942476Z

Reference resolution

48 of 48 outbound references displayed

  • verified exact0
  • verified fuzzy17
  • unresolved31
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 1116d9ff-1bb6-45fe-8c3b-2d690506a98d · outbound

This paper cites The People's Speech: A Large-Scale Diverse English Speech Recognition Dataset for Commercial Usage.

TTS-1 Technical Report The People's Speech: A Large-Scale Diverse English Speech Recognition Dataset for Commercial Usage

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T15:02:02.830714Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:02:02.830714Z digest=sha256:1a8d88e6f2ba59e3b755dc051b4f93aa18594223ef42ebd27c90bb3658ac04c1

Observation b6165c63-f5b3-490c-908c-7f7aa6a4e881 · outbound

This paper cites Yodas: Youtube-oriented dataset for audio and speech.

TTS-1 Technical Report Yodas: Youtube-oriented dataset for audio and speech

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:02:04.059235Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:02:02.837634Z digest=sha256:0399468830204ebf6498a08cf2fe76f2bec187fc572df6da38f0bc649182e5e9

Observation 9a8e03bc-ef3b-494a-8385-5600549cb4dc · outbound

This paper cites Emilia: An extensive, multilingual, and diverse speech dataset for large-scale speech generation.

TTS-1 Technical Report Emilia: An extensive, multilingual, and diverse speech dataset for large-scale speech generation

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:02:04.040769Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:02:02.843161Z digest=sha256:d43ddf1a9d552314dec22466c92a0b9d06589e8924d0a11a05409281a9bc2b2b

Observation a5392520-ba95-430c-a8b1-d442e5564146 · outbound

This paper cites Styletts 2: Towards human-level text-to-speech through style diffusion and adversarial training with large speech language models.

TTS-1 Technical Report Styletts 2: Towards human-level text-to-speech through style diffusion and adversarial training with large speech language models

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:02:04.025179Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:02:02.847985Z digest=sha256:8480b2bae26107bb86cbe97885d5ce36c98153150d0933eb9cadbd08e2e11333

Observation d993e778-bc6a-4974-a47b-8ab4e75e5f8b · outbound

This paper cites Fastspeech: Fast, robust and controllable text to speech.

TTS-1 Technical Report Fastspeech: Fast, robust and controllable text to speech

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:02:04.008207Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:02:02.854081Z digest=sha256:ec4a4410152e5a4e1004847b31bf5d4c9dbb022fdbf042b3f3a16705d7f2907e

Observation 448ecfc7-dea8-42b5-b8e6-c5ee136542be · outbound

This paper cites Better speech synthesis through scaling.

TTS-1 Technical Report Better speech synthesis through scaling

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T15:02:02.858858Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:02:02.858858Z digest=sha256:58267e25bc9154cd4355fbea73eeea02c4c55165c425536efda639ed83ca7289

Observation e9573b28-862d-43f8-928b-4878a58d2629 · outbound

This paper cites Conditional variational autoencoder with adversarial learning for end-to-end text-to-speech.

TTS-1 Technical Report Conditional variational autoencoder with adversarial learning for end-to-end text-to-speech

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:02:03.988145Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:02:02.865173Z digest=sha256:73cdf7cef3270177fc433c1bb7a5d1e1ca2cb37d1783304b0770f406efc625a4

Observation e204a9e3-3816-48fe-84cf-4e4f73ff3a0c · outbound

This paper cites V oicebox: Text-guided multilingual universal speech generation at scale.

TTS-1 Technical Report V oicebox: Text-guided multilingual universal speech generation at scale

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:02:03.965541Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:02:02.870328Z digest=sha256:93b870030efa3cb47247497ce075f3ab2f8067dfcef491767ea75dabf231af59

Observation 3fb837f9-85f5-43b5-add2-e03a99597d36 · outbound

This paper cites Language models are unsupervised multitask learners.

TTS-1 Technical Report Language models are unsupervised multitask learners

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T15:02:02.875785Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:02:02.875785Z digest=sha256:2155cd0fcb63642066465719106ff2d2e8ef329ac30bf0faf887132eef7912e0

Observation 7c2d9832-c784-4272-b9a4-311689e7b0d2 · outbound

This paper cites Training Compute-Optimal Large Language Models.

TTS-1 Technical Report Training Compute-Optimal Large Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T15:02:02.880909Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:02:02.880909Z digest=sha256:a55d97e6bb91f3b1c66eb9ea3e2af8afaf5a63840409616f5d1740be7a48bf6b

Observation d96c6f96-d5ae-4aa6-b75b-63251abf5781 · outbound

This paper cites Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers.

TTS-1 Technical Report Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T15:02:02.886470Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:02:02.886470Z digest=sha256:0d02622237a5352917dbcb5cef956218537e7258068a96ee61103748c9dc7753

Observation bc5e8aef-029b-4609-8889-5fef279eb141 · outbound

This paper cites MiniMax-Speech: Intrinsic Zero-Shot Text-to-Speech with a Learnable Speaker Encoder.

TTS-1 Technical Report MiniMax-Speech: Intrinsic Zero-Shot Text-to-Speech with a Learnable Speaker Encoder

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T15:02:02.891456Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:02:02.891456Z digest=sha256:24da5d6e84b6ef71741fa396f84ef5ea8511fc21657ceab8c5109cef3f1b3dc6

Observation aed9d272-a437-42ff-bec8-c77046dd3603 · outbound

This paper cites Seed-TTS: A Family of High-Quality Versatile Speech Generation Models.

TTS-1 Technical Report Seed-TTS: A Family of High-Quality Versatile Speech Generation Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T15:02:02.897698Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:02:02.897698Z digest=sha256:46689f4aac1dd2fccd0ccfb489563dfd1849ba74feeeb881d706c7b8f048ed61

Observation a3c777f1-cd98-4493-aae4-b99d5ea637dd · outbound

This paper cites CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens.

TTS-1 Technical Report CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T15:02:02.904024Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:02:02.904024Z digest=sha256:a09d3e78c907e80e4f0a69747283b7ead5f22cbf3fc886155b4569e3537ee39b

Observation a2cd3a78-d85f-44a1-9c4b-f1c80281eac1 · outbound

This paper cites CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models.

TTS-1 Technical Report CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T15:02:02.908687Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:02:02.908687Z digest=sha256:011fba5027ad8e1ede8b582757eba29b12a15cb08f58c25228f5c5e644ed2e5d

Observation bd31abf3-cdc3-4957-9bde-e4ca9251251f · outbound

This paper cites Redpajama: an open dataset for training large language models.

TTS-1 Technical Report Redpajama: an open dataset for training large language models

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:02:03.926487Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:02:02.913946Z digest=sha256:7d54189a44f1f69a216156387f87549a2b13523d0d9c57af010a40ff7ab640f2

Observation 669d28f8-b554-4228-a0ee-fdb73850def1 · outbound

This paper cites Open instruction generalist (oig) dataset.

TTS-1 Technical Report Open instruction generalist (oig) dataset

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:02:03.903451Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:02:02.919362Z digest=sha256:354b7020864b3a7e5be2974d08c03f680723e7e146f51f9f9426eef62296b467

Observation 1f6b538f-24df-4cef-b876-74eb4f3f71dc · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

TTS-1 Technical Report DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T15:02:02.924574Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:02:02.924574Z digest=sha256:3b57abe3d86b0dd37073809b282955ba0f1398ef901fa591b965a4519dc1b22f

Observation 5a853971-6149-46e6-9863-4098efe1267e · outbound

This paper cites Llasa: Scaling Train-Time and Inference-Time Compute for Llama-based Speech Synthesis.

TTS-1 Technical Report Llasa: Scaling Train-Time and Inference-Time Compute for Llama-based Speech Synthesis

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T15:02:02.930108Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:02:02.930108Z digest=sha256:4804d9e425d399eca0662f5cf5ef67edf340acc0a07dc3d3209f07557af1ef57

Observation 61e20d6f-5552-4c0c-a9a8-b4a2d805e76e · outbound

This paper cites Wavlm: Large-scale self-supervised pre-training for full stack speech processing.

TTS-1 Technical Report Wavlm: Large-scale self-supervised pre-training for full stack speech processing

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:02:03.884270Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:02:02.935030Z digest=sha256:5b21410f340348afaeb8cee5e6327bd1f1d2156d9e5dad98719d886bac6ab049

Observation befd186d-6c49-493d-8b30-f788cea54faa · outbound

This paper cites Dnsmos p.

TTS-1 Technical Report Dnsmos p

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:02:03.865444Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:02:02.940134Z digest=sha256:927014b4586477b4e5b1573fc33dd3b72ad71034c1c57e91e94d6ecb4a2c8ec7

Observation 3d900436-35b3-4f17-b713-eba6b6f78fa3 · outbound

This paper cites Lora: Low-rank adaptation of large language models.

TTS-1 Technical Report Lora: Low-rank adaptation of large language models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T15:02:02.944597Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:02:02.944597Z digest=sha256:993eba5390a5c892aa779bc509851b86796d5af1de4a64eac4a01c63b2049c8e

Observation bdf5b87d-9598-4cab-88a2-5c88c38d010a · outbound

This paper cites Deep residual learning for image recognition.

TTS-1 Technical Report Deep residual learning for image recognition

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T15:02:02.949788Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:02:02.949788Z digest=sha256:8bc685d3ba3006301b04ffd06efca26e6f25f4873d5ffc39edde6249e15d138d

Observation 089d6aab-8a5f-4074-bdbd-071b1dd9a10c · outbound

This paper cites The Llama 3 Herd of Models.

TTS-1 Technical Report The Llama 3 Herd of Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T15:02:02.954136Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:02:02.954136Z digest=sha256:fa39b087eef264c71e0b25c39d2c7856f5291fa28167fc4ba8d49529d71cf31b

Observation 6e452096-9f3d-4855-8a96-101daea28de3 · outbound

This paper cites Matrix multiplication background user’s guide.

TTS-1 Technical Report Matrix multiplication background user’s guide

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:02:03.825369Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:02:02.959872Z digest=sha256:9f6df9c04a359ba61ef67cc35da0b8db69e1ec7eaf77f99e4c23050e1caa8d90

Observation 63165417-09a3-4464-b43e-d01040b7e49f · outbound

This paper cites Initializing new word embeddings for pretrained language models.

TTS-1 Technical Report Initializing new word embeddings for pretrained language models

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:02:03.808174Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:02:02.964733Z digest=sha256:33c0e01bb6fbd41b7be442194b6a33f798119b31dfa5b4fa8f5d302c824d9ef8

Observation dee61c16-a693-413d-baec-a8d71bc82094 · outbound

This paper cites BigCodec: Pushing the Limits of Low-Bitrate Neural Speech Codec.

TTS-1 Technical Report BigCodec: Pushing the Limits of Low-Bitrate Neural Speech Codec

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T15:02:02.970136Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:02:02.970136Z digest=sha256:066e6f2004ed735767b73a2100234f3716e34b8931e4ac9af21349e46eed3cea

Observation e7372153-22da-4005-9a06-a28caf869736 · outbound

This paper cites Hifi-gan: Generative adversarial networks for efficient and high fidelity speech synthesis.

TTS-1 Technical Report Hifi-gan: Generative adversarial networks for efficient and high fidelity speech synthesis

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:02:03.788491Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:02:02.974412Z digest=sha256:33d371a1ca8cafd3c7523dcc734414b97c58eb11f529c8fa21e70a632ff122b8

Observation 331a3e81-ee3b-4418-9bf2-aac95671b668 · outbound

This paper cites High Fidelity Neural Audio Compression.

TTS-1 Technical Report High Fidelity Neural Audio Compression

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T15:02:02.980832Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:02:02.980832Z digest=sha256:b961b3975a60f6134189b2144e52fb2504d4e2d682fc09182440799c82710c01

Observation dfac856f-3877-4bf0-9bdb-d925ad897b44 · outbound

This paper cites Adding Instructions during Pretraining: Effective Way of Controlling Toxicity in Language Models.

TTS-1 Technical Report Adding Instructions during Pretraining: Effective Way of Controlling Toxicity in Language Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T15:02:02.985709Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:02:02.985709Z digest=sha256:b896997e91ee9c148a7e8d358f63fc2a59bf0143a36f55c60c4593281b27d570

Observation b2d0e952-b052-4dc2-8a25-3c9367e1cfe1 · outbound

This paper cites A neural probabilistic language model.

TTS-1 Technical Report A neural probabilistic language model

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T15:02:02.990541Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:02:02.990541Z digest=sha256:e203cbdd0f8aa54bcd10596d97a4e7ca1ff467bfc387f13eb0ef0cdccf3dda0a

Observation 9133858a-f8b9-420b-8e82-19916efa00be · outbound

This paper cites FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning.

TTS-1 Technical Report FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T15:02:02.997072Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:02:02.997072Z digest=sha256:00f9a7b31c4c2449f2095cde6448693c1c69d6927e76545ab135c95582092525

Observation 29c6c949-fa70-4f7c-8811-4deab52cb8f1 · outbound

This paper cites Adam: A Method for Stochastic Optimization.

TTS-1 Technical Report Adam: A Method for Stochastic Optimization

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T15:02:03.001891Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:02:03.001891Z digest=sha256:b60f1f4454c21564b12abbc8817cc57f24144bc85a84aa43a220da30c4ecbe0a

Observation dc53bb10-0b74-41a7-93a0-bac436025d36 · outbound

This paper cites PyTorch Distributed: Experiences on Accelerating Data Parallel Training.

TTS-1 Technical Report PyTorch Distributed: Experiences on Accelerating Data Parallel Training

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T15:02:03.006360Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:02:03.006360Z digest=sha256:2204988bf3f0aebaffc84fc863dd85e3934aee80c15a3d4bb35b134bf4fa55d1

Observation d962fbbe-3828-4961-9347-cf0e52fb0db0 · outbound

This paper cites PyTorch FSDP: Experiences on Scaling Fully Sharded Data Parallel.

TTS-1 Technical Report PyTorch FSDP: Experiences on Scaling Fully Sharded Data Parallel

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T15:02:03.011140Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:02:03.011140Z digest=sha256:9b50896f3629dd25798ad945e27d3854c7c0669c4b184435b77bfe109173be56

Observation e27dba3b-4a19-4c67-a1b4-2ed803652cb5 · outbound

This paper cites Parler-tts.

TTS-1 Technical Report Parler-tts

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:02:03.755019Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:02:03.016065Z digest=sha256:9c9dbb54aa7501fb4e5487ffbda063f0b52f16c0267c5c4e90aa6042737019f2

Observation 8a5c05da-c5d0-4796-bf6d-9303960c483b · outbound

This paper cites Zero: Memory optimizations toward training trillion parameter models.

TTS-1 Technical Report Zero: Memory optimizations toward training trillion parameter models

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:02:03.737535Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:02:03.021572Z digest=sha256:c7b663695f14cee5f92e6043350c94260cee584284792f1afd4600ffcaa4f4ca

Observation 52ded5a3-2265-49cd-b401-61758dbf0392 · outbound

This paper cites CosyVoice 3: Towards In-the-wild Speech Generation via Scaling-up and Post-training.

TTS-1 Technical Report CosyVoice 3: Towards In-the-wild Speech Generation via Scaling-up and Post-training

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T15:02:03.026897Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:02:03.026897Z digest=sha256:398b12c8d147ab16f301161b1cfaabe2489f68a40cb1a22e4dfd90d94364b0d4

Observation 3640e72a-bbf6-47a0-89f5-b238dca3c766 · outbound

This paper cites Enhancing Zero-shot Text-to-Speech Synthesis with Human Feedback.

TTS-1 Technical Report Enhancing Zero-shot Text-to-Speech Synthesis with Human Feedback

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T15:02:03.032073Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:02:03.032073Z digest=sha256:2dda4c03aa4a04824c33788fc9f9d88330645a6891ad5c9747f8f8e5a3cf7701

Observation 4a88a8cd-6997-471a-8f21-bc960cf587e3 · outbound

This paper cites DeepSeek-V3 Technical Report.

TTS-1 Technical Report DeepSeek-V3 Technical Report

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T15:02:03.037268Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:02:03.037268Z digest=sha256:5b5e90aae5821bfdb59a0450bf7a2999147667f51f3ca3f1d5fd4afb1c20f281

Observation 63c76d66-c01b-4edd-9a8a-2caebbdb197f · outbound

This paper cites DeepSeek LLM: Scaling Open-Source Language Models with Longtermism.

TTS-1 Technical Report DeepSeek LLM: Scaling Open-Source Language Models with Longtermism

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T15:02:03.042360Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:02:03.042360Z digest=sha256:e6e2bd7f2b129cba4b1c0e499b54b4da8baf7a3a0d592c4513d8ee9e5d8b4880

Observation 4ca58588-1f1f-4943-8194-60f8dbc61c15 · outbound

This paper cites Direct preference optimization: Your language model is secretly a reward model.

TTS-1 Technical Report Direct preference optimization: Your language model is secretly a reward model

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T15:02:03.047248Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:02:03.047248Z digest=sha256:4fc8b5f036d025877d1e9a0edad8457c62b73bf96dc1da676de9165f58208172

Observation f720f069-b85f-48e8-afe8-ae422da1397e · outbound

This paper cites Understanding R1-Zero-Like Training: A Critical Perspective.

TTS-1 Technical Report Understanding R1-Zero-Like Training: A Critical Perspective

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T15:02:03.052195Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:02:03.052195Z digest=sha256:afe1d30e67d1077e74bffcd3e083a99339dd98ef0881aed720adfe248749ca26

Observation b9b904ce-8b94-46ad-9c3d-99590258e76e · outbound

This paper cites Robust speech recognition via large-scale weak supervision.

TTS-1 Technical Report Robust speech recognition via large-scale weak supervision

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T15:02:03.057567Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:02:03.057567Z digest=sha256:424c27f07625d113186a0ed8e759354cfe7b864af720fddb6bc131802ca497a1

Observation 6d2518ba-0c5b-4908-ace2-7d94f7171699 · outbound

This paper cites HuggingFace's Transformers: State-of-the-art Natural Language Processing.

TTS-1 Technical Report HuggingFace's Transformers: State-of-the-art Natural Language Processing

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T15:02:03.063141Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:02:03.063141Z digest=sha256:949be851719ad5545481a4969ec8ca7d49ebe85d704f602386845e6c35293177

Observation bd80dfc2-0d2c-400a-818f-ebd963c3eeab · outbound

This paper cites PyTorch: An Imperative Style, High-Performance Deep Learning Library.

TTS-1 Technical Report PyTorch: An Imperative Style, High-Performance Deep Learning Library

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T15:02:03.068483Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:02:03.068483Z digest=sha256:e1faaa9527aecd913079307d5cf502f8fafb36da4f85b99cbeaf457ca232020a

Observation a606eda4-cfd8-44ee-85bf-7b99585a2b9b · outbound

This paper cites Pytorch lightning.

TTS-1 Technical Report Pytorch lightning

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:02:03.690672Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:02:03.073406Z digest=sha256:2925fb0aa8c4dcdff90bc0c1531d5fb07cc1b5351e2d5cca51ac5606be287bed

Observation c88b760d-9f09-4e4e-a31c-fb9f6ae7cc7b · outbound

This paper cites Efficient memory management for large language model serving with pagedattention.

TTS-1 Technical Report Efficient memory management for large language model serving with pagedattention

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:02:03.674403Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T15:02:03.077732Z digest=sha256:40fcfe0f9edc40a987a1bae41c2c91eef024ad66eb0dfcee0e3968e8ad536e4a

Pith citing papers

Observation 64fed19a-4d58-49a6-867f-8977988efecd · inbound

Evaluating and Rewarding LALMs for Expressive Role-Play TTS via Mean Continuation Log-Probability cites this paper.

Evaluating and Rewarding LALMs for Expressive Role-Play TTS via Mean Continuation Log-Probability TTS-1 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-03T06:36:31.523597Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T06:36:31.523597Z digest=sha256:c387b019e5eaae45ab8d7a77160e67c6024d8d71c5a1ab7999ac711c7f8a6d6f

Observation b76085d3-b265-4c38-beb9-aa07f37b5812 · inbound

JaiTTS: A Thai Voice Cloning Model cites this paper.

JaiTTS: A Thai Voice Cloning Model TTS-1 Technical Report

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-12T10:21:28.850669Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-07T06:34:45.357695Z digest=sha256:a47a219c9dccf9b4d156aeef9206ebaa3080f82e89e2b9109c4575acfd6b5076

Observation cca52b6c-b63b-414c-9517-cb5c8bc46c2a · inbound

JaiTTS: A Thai Voice Cloning Model cites this paper.

JaiTTS: A Thai Voice Cloning Model TTS-1 Technical Report

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-11T22:16:26.706638Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-08T03:09:24.541657Z digest=sha256:c81d576909b38b2c0324194df23bae4cd9c8fcf025f5809e10b5e1955e835b76

Observation cc3f5c45-0029-4f23-a0c6-f49bfa10d585 · inbound

Reliable Neural-Codec Text-to-Speech by ASR Self-Verification and Distillation: Near-Zero Catastrophic Failures Across Models and Codecs cites this paper.

Reliable Neural-Codec Text-to-Speech by ASR Self-Verification and Distillation: Near-Zero Catastrophic Failures Across Models and Codecs TTS-1 Technical Report

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-07-03T23:19:02.945574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T22:46:47.786929Z digest=sha256:fc1f9c200de9af5ac7bc6eca08d6532a818118b2e97efa79bef5440161379b92