Pith. sign in

Paper Citation Record · LEDGER

TouchTTS: An Embarrassingly Simple TTS Framework that Everyone Can Touch

As of 20 August 2026, this Paper Citation Record lists 39 of 39 outbound references and 8 inbound Pith citation observations for arXiv:2412.08237.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.08237 v2

Coverage vector

measured 39 of 39 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T18:08:22.468176Z

measured 47 of 47 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 8 of 8 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T15:22:08.946434Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T20:10:07.248893Z

Reference resolution

39 of 39 outbound references displayed

  • verified exact0
  • verified fuzzy11
  • unresolved28
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 9c0f2c5a-92cd-4bf5-a35f-dcc129a8f942 · outbound

This paper cites CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens.

TouchTTS: An Embarrassingly Simple TTS Framework that Everyone Can Touch CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T18:08:21.220313Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:08:21.220313Z digest=sha256:f24d22effac843b1b50a498b297e5986289a51f609fa41ceff8de3cecfb65076

Observation e5d6af9a-26b2-4955-be85-259db7942f6a · outbound

This paper cites Matcha-tts: A fast tts architecture with conditional flow matching.

TouchTTS: An Embarrassingly Simple TTS Framework that Everyone Can Touch Matcha-tts: A fast tts architecture with conditional flow matching

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:08:23.948746Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T18:08:21.283284Z digest=sha256:4a9f92aef3d78eca1dded79414a146585d3c0c8f27b925c6385343db23eb2bbd

Observation 7e53d2af-665a-4a8a-b9a9-a03dd1d7563e · outbound

This paper cites Qwen2 Technical Report.

TouchTTS: An Embarrassingly Simple TTS Framework that Everyone Can Touch Qwen2 Technical Report

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T18:08:21.290547Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:08:21.290547Z digest=sha256:32876fbd638dca5cd38446db456eb26c79746c3e1cd9b3a3619e1b73fe2dbb4a

Observation 00377421-e3be-43b9-814f-606104334678 · outbound

This paper cites Tensorrt-llm, https://github.com/nvidia/tensorrt-llm, 2024.

TouchTTS: An Embarrassingly Simple TTS Framework that Everyone Can Touch Tensorrt-llm, https://github.com/nvidia/tensorrt-llm, 2024

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:08:23.873058Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T18:08:21.296528Z digest=sha256:ca7cac22a2d272854664bb3ecd291b2e4fa992d54c10fb42ce3472a6fe3e8e9d

Observation b9ce65ba-d788-4f68-a602-44dc3328224e · outbound

This paper cites Gonzalez, Hao Zhang, and Ion Stoica.

TouchTTS: An Embarrassingly Simple TTS Framework that Everyone Can Touch Gonzalez, Hao Zhang, and Ion Stoica

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T18:08:21.302797Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:08:21.302797Z digest=sha256:1b50544e2377fd4d1f385ba3f5974a23e003db72c35450a8c15307509ba07408

Observation 76716d8a-b0b7-4683-b004-930eec04fc31 · outbound

This paper cites A Survey on In-context Learning.

TouchTTS: An Embarrassingly Simple TTS Framework that Everyone Can Touch A Survey on In-context Learning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T18:08:21.360766Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:08:21.360766Z digest=sha256:25981b39e8012e370c906d76069f94b9460a65a91ce90ab4c2fb99564c5a578e

Observation 14708426-a94d-4b89-b445-2ee36eab8058 · outbound

This paper cites Fish-Speech: Leveraging Large Language Models for Advanced Multilingual Text-to-Speech Synthesis.

TouchTTS: An Embarrassingly Simple TTS Framework that Everyone Can Touch Fish-Speech: Leveraging Large Language Models for Advanced Multilingual Text-to-Speech Synthesis

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T18:08:21.514869Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:08:21.514869Z digest=sha256:ef4a0a06c1d7cf14ade93dd63a492b2016876f453cbff5d8df8c45dd1a2be71c

Observation a36ebeec-3cd5-4aee-872c-8f34e3681db2 · outbound

This paper cites FireRedTTS: A Foundation Text-To-Speech Framework for Industry-Level Generative Speech Applications.

TouchTTS: An Embarrassingly Simple TTS Framework that Everyone Can Touch FireRedTTS: A Foundation Text-To-Speech Framework for Industry-Level Generative Speech Applications

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T18:08:21.521139Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:08:21.521139Z digest=sha256:dd8c7f61e7106110d0656db32583661960053768b757a493ca4466281fe0a8b3

Observation ecb87789-7171-4a96-8948-f07973122c76 · outbound

This paper cites Takin: A Cohort of Superior Quality Zero-shot Speech Generation Models.

TouchTTS: An Embarrassingly Simple TTS Framework that Everyone Can Touch Takin: A Cohort of Superior Quality Zero-shot Speech Generation Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T18:08:21.526764Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:08:21.526764Z digest=sha256:ec411b4356ae1d6498ecdd3f19a2efd96d4673aaf30149eff4af33012e30be19

Observation 7f4dffea-ff9c-4850-aeaf-e92507e95c58 · outbound

This paper cites Wenetspeech4tts: A 12,800-hour mandarin tts corpus for large speech generation model benchmark, 2024.

TouchTTS: An Embarrassingly Simple TTS Framework that Everyone Can Touch Wenetspeech4tts: A 12,800-hour mandarin tts corpus for large speech generation model benchmark, 2024

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:08:23.798628Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T18:08:21.533197Z digest=sha256:923dd6c36f3bc7e07d849a6bbdd05dda64b61b0bc17822773317ae840ee2c65b

Observation 4f15adfc-e4a4-470c-883a-e7813dc1edb5 · outbound

This paper cites Emilia: An Extensive, Multilingual, and Diverse Speech Dataset for Large-Scale Speech Generation.

TouchTTS: An Embarrassingly Simple TTS Framework that Everyone Can Touch Emilia: An Extensive, Multilingual, and Diverse Speech Dataset for Large-Scale Speech Generation

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T18:08:21.539179Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:08:21.539179Z digest=sha256:a703c8c6884c8449abb795e4988de47aa5aad344385dcf8198069b0378984e79

Observation 71f22068-6a51-4ae3-ab9e-631c43399a61 · outbound

This paper cites Autoprep: An automatic preprocessing framework for in-the-wild speech data.

TouchTTS: An Embarrassingly Simple TTS Framework that Everyone Can Touch Autoprep: An automatic preprocessing framework for in-the-wild speech data

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:08:23.781474Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T18:08:21.544007Z digest=sha256:90bd5172d3ff1a097bae6e6167eb5a939e4edb9a288c290f66213fc6e665e72d

Observation 9524d2cb-d989-4da6-9acf-a538ec5c9f44 · outbound

This paper cites Flow Matching for Generative Modeling.

TouchTTS: An Embarrassingly Simple TTS Framework that Everyone Can Touch Flow Matching for Generative Modeling

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T18:08:21.670486Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:08:21.670486Z digest=sha256:69e7a63c37ba2829045844c2492bbb4cb98673a298a863821abddd663c14596b

Observation 85701e04-db90-4a23-9fd6-ce4d017651e8 · outbound

This paper cites Dnsmos p.

TouchTTS: An Embarrassingly Simple TTS Framework that Everyone Can Touch Dnsmos p

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:08:23.733607Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T18:08:21.726313Z digest=sha256:6e74a7373863efcd5d94909f44c52bb7b680b1b71ee5afc26227961489566006

Observation 8538d50f-60f3-444b-afd5-da95ded1fda2 · outbound

This paper cites AudioPaLM: A Large Language Model That Can Speak and Listen.

TouchTTS: An Embarrassingly Simple TTS Framework that Everyone Can Touch AudioPaLM: A Large Language Model That Can Speak and Listen

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T18:08:21.732047Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:08:21.732047Z digest=sha256:4cd3b1d70e864383e9e5443d64accb746fe8ae964973693bb27023e1533b2596

Observation e7907a02-61d0-4b53-a45f-ecd6273793c0 · outbound

This paper cites Conformer-1: Robust ASR via Large-Scale Semisupervised Bootstrapping.

TouchTTS: An Embarrassingly Simple TTS Framework that Everyone Can Touch Conformer-1: Robust ASR via Large-Scale Semisupervised Bootstrapping

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T18:08:21.738126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:08:21.738126Z digest=sha256:2793f91994dcd91082cf51940fbe7c56382e5da55b34de9067edf8ffbb6d46c1

Observation 0fee4344-9b68-48e6-905d-4b0eaa6ce871 · outbound

This paper cites Anatomy of Industrial Scale Multilingual ASR.

TouchTTS: An Embarrassingly Simple TTS Framework that Everyone Can Touch Anatomy of Industrial Scale Multilingual ASR

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T18:08:21.743850Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:08:21.743850Z digest=sha256:a8df476b4f971dcf90d5f93adea91920dcb40c955de6ff0020fc4e2f07271ac1

Observation 6098f619-c285-4b20-bb0e-e7f1852d105a · outbound

This paper cites Robust speech recognition via large-scale weak supervision.

TouchTTS: An Embarrassingly Simple TTS Framework that Everyone Can Touch Robust speech recognition via large-scale weak supervision

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:08:23.589187Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T18:08:21.750379Z digest=sha256:ae97cea0ce471e2f92888372721294e1cdb01d4f7c76e1f375f55588aff48fc9

Observation f2025600-be32-418c-900e-0c687c575d17 · outbound

This paper cites Paraformer: Fast and Accurate Parallel Transformer for Non-autoregressive End-to-End Speech Recognition.

TouchTTS: An Embarrassingly Simple TTS Framework that Everyone Can Touch Paraformer: Fast and Accurate Parallel Transformer for Non-autoregressive End-to-End Speech Recognition

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T18:08:21.797820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:08:21.797820Z digest=sha256:a31104a8bf9aa4ee59104395ca8a4b84197ab4322e751512291e072369218407

Observation ff202718-2327-4c02-839a-e0ca95c00bef · outbound

This paper cites WeNet 2.0: More Productive End-to-End Speech Recognition Toolkit.

TouchTTS: An Embarrassingly Simple TTS Framework that Everyone Can Touch WeNet 2.0: More Productive End-to-End Speech Recognition Toolkit

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T18:08:21.882088Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:08:21.882088Z digest=sha256:8462018735eb10dc3bced20545d527dfe84b85d614881275d9dc86a0f0f299eb

Observation d6265433-5bf9-4479-acd3-589ade9fe126 · outbound

This paper cites Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers.

TouchTTS: An Embarrassingly Simple TTS Framework that Everyone Can Touch Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T18:08:21.919160Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:08:21.919160Z digest=sha256:aeec6458c0411ee98e2788bc830542fb22f65836c11569995b639eaec0fd9fa7

Observation 2e1f8654-d813-465f-92f4-eba1ea976522 · outbound

This paper cites Conditional variational autoencoder with adversar- ial learning for end-to-end text-to-speech.

TouchTTS: An Embarrassingly Simple TTS Framework that Everyone Can Touch Conditional variational autoencoder with adversar- ial learning for end-to-end text-to-speech

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:08:23.522689Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T18:08:21.927770Z digest=sha256:f7bfb944861c12f8881898790b6649fc5415cb0b157c43b08c94c63474325da4

Observation 9b60a389-9c89-4723-9c16-42610d48c005 · outbound

This paper cites FastSpeech: Fast, robust and controllable text to speech.

TouchTTS: An Embarrassingly Simple TTS Framework that Everyone Can Touch FastSpeech: Fast, robust and controllable text to speech

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:08:23.505291Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T18:08:21.933675Z digest=sha256:b2254acd5b9472ea401b782db0ffd2e60d5907af2efdd0c09df6b9f9776cdfde

Observation 013ba208-de80-4ee5-81bd-29eda3a8bcd1 · outbound

This paper cites an unresolved cited work.

TouchTTS: An Embarrassingly Simple TTS Framework that Everyone Can Touch Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-08-11T18:08:23.407059Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T18:08:21.938510Z digest=sha256:b5c22d0cbec68d84171bb6766a6dc3e9502a809f0e8af594de50de88e43ca95f

Observation 74e0850d-acf8-4e64-b6e7-a75d99f1b7ad · outbound

This paper cites Flow- TTS: A non-autoregressive network for text to speech based on flow.

TouchTTS: An Embarrassingly Simple TTS Framework that Everyone Can Touch Flow- TTS: A non-autoregressive network for text to speech based on flow

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:08:23.271402Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T18:08:21.943925Z digest=sha256:b165b645feafd7e64944ee7a5476d9ab7f82e7d77062e9c2333d2b112ea65438

Observation e0f7b5e2-e811-47fb-8381-d9442528c73a · outbound

This paper cites An Embarrassingly Simple Approach for LLM with Strong ASR Capacity.

TouchTTS: An Embarrassingly Simple TTS Framework that Everyone Can Touch An Embarrassingly Simple Approach for LLM with Strong ASR Capacity

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T18:08:22.019176Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:08:22.019176Z digest=sha256:3dcc616986154d03bae2fe42b5444a452e454f40ac97c07ccd15ee7b354d711e

Observation ee79a058-2e8c-4acf-9ec8-70da2858fa0a · outbound

This paper cites LauraGPT: Listen, Attend, Understand, and Regenerate Audio with GPT.

TouchTTS: An Embarrassingly Simple TTS Framework that Everyone Can Touch LauraGPT: Listen, Attend, Understand, and Regenerate Audio with GPT

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-11T18:08:22.096260Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:08:22.096260Z digest=sha256:07ada1087ae9e7c164e54f60b368cd67c8aa7c7cef443b55041adb1672049a2d

Observation 647a6ebd-3673-4ad8-b35c-6f63a7df5f33 · outbound

This paper cites Seed-ASR: Understanding Diverse Speech and Contexts with LLM-based Speech Recognition.

TouchTTS: An Embarrassingly Simple TTS Framework that Everyone Can Touch Seed-ASR: Understanding Diverse Speech and Contexts with LLM-based Speech Recognition

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-11T18:08:22.102985Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:08:22.102985Z digest=sha256:6aa55a869eefa1b8c32faa7d465e6b582108776976e4bcd93345147d33f2b68f

Observation 1ee258fb-61d1-43ac-bc9a-339dfb8e6b5c · outbound

This paper cites dMel: Speech Tokenization made Simple.

TouchTTS: An Embarrassingly Simple TTS Framework that Everyone Can Touch dMel: Speech Tokenization made Simple

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-11T18:08:22.110679Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:08:22.110679Z digest=sha256:ac5ef0a4114607981d83c03443f5aef7eb59e92e7f8e7ea642b9645a325a0dbb

Observation 07d337c4-65c0-46ef-a01c-3a652fbb48a8 · outbound

This paper cites Librispeech: an asr corpus based on public domain audio books.

TouchTTS: An Embarrassingly Simple TTS Framework that Everyone Can Touch Librispeech: an asr corpus based on public domain audio books

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-11T18:08:22.117560Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:08:22.117560Z digest=sha256:6e6074dbb94543d78359ea5d65ac82e5b55a7a29b77dbea8a95c1e35b68438f7

Observation fc953978-061d-4e64-af49-b7fe1ea0dfed · outbound

This paper cites Scaling Speech-Text Pre-training with Synthetic Interleaved Data.

TouchTTS: An Embarrassingly Simple TTS Framework that Everyone Can Touch Scaling Speech-Text Pre-training with Synthetic Interleaved Data

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-11T18:08:22.122969Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:08:22.122969Z digest=sha256:a576a7e425ffa75e2af35ad0c77f114d1f44e5fcc56d84bf5d9f9aa18b1ccab4

Observation bbd21626-3df2-46ed-b674-41b41737dec3 · outbound

This paper cites Spirit LM: Interleaved Spoken and Written Language Model.

TouchTTS: An Embarrassingly Simple TTS Framework that Everyone Can Touch Spirit LM: Interleaved Spoken and Written Language Model

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-11T18:08:22.225106Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:08:22.225106Z digest=sha256:c7945f9728275e808f5ae5bfb9eb56930c22dfd009c7122ebfab89a31e4254a1

Observation 3eec4792-0ce4-4306-95c5-512c23899f9f · outbound

This paper cites Seed-TTS: A Family of High-Quality Versatile Speech Generation Models.

TouchTTS: An Embarrassingly Simple TTS Framework that Everyone Can Touch Seed-TTS: A Family of High-Quality Versatile Speech Generation Models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-11T18:08:22.349166Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:08:22.349166Z digest=sha256:502221f95b8a94e1638b58c532a3e478f6dbd8ffcf001777e1d12a248d44fb83

Observation d8acfa9a-6753-4159-8d29-c7345c7d2d95 · outbound

This paper cites Didispeech: A large scale mandarin speech corpus.

TouchTTS: An Embarrassingly Simple TTS Framework that Everyone Can Touch Didispeech: A large scale mandarin speech corpus

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-11T18:08:22.357484Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:08:22.357484Z digest=sha256:57895a2b944abbf17f5498130fc4ba83a76e999c00c54e3478ca2a2b20d6a194

Observation 1ba52b68-e7f6-4832-9721-4ba82c72dc29 · outbound

This paper cites Common Voice: A Massively-Multilingual Speech Corpus.

TouchTTS: An Embarrassingly Simple TTS Framework that Everyone Can Touch Common Voice: A Massively-Multilingual Speech Corpus

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-11T18:08:22.362702Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:08:22.362702Z digest=sha256:3104096a054c70727ce6b846c9341c6b2d1f24ed3be9686ffd64e0e1bf8e331e

Observation a72a38ed-83fb-46f8-851e-bcfa8cea037e · outbound

This paper cites F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching.

TouchTTS: An Embarrassingly Simple TTS Framework that Everyone Can Touch F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-11T18:08:22.371436Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:08:22.371436Z digest=sha256:ca68d224e7f1018469595526eea26a585ae0d25c030ee2e8161216d2962e4cb2

Observation 7ac6d73c-4f62-49b0-a77c-5a5d04a92fbf · outbound

This paper cites MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer.

TouchTTS: An Embarrassingly Simple TTS Framework that Everyone Can Touch MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-11T18:08:22.378032Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T18:08:22.378032Z digest=sha256:e846a7141d8ca607f44d472864e6f0a4c8cd9ead4a4d3811e5d2ec0dc0953f7c

Observation a228842c-9982-4f5b-ad87-6d326cb7bbf1 · outbound

This paper cites Wespeaker: A research and production oriented speaker embedding learning toolkit.

TouchTTS: An Embarrassingly Simple TTS Framework that Everyone Can Touch Wespeaker: A research and production oriented speaker embedding learning toolkit

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:08:23.231266Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T18:08:22.435225Z digest=sha256:1f46a68054d616a36c23442743a67011cb15be9630480ec6580cf98ff56778d5

Observation 0707ff01-6773-4813-950f-d5ec9cff023c · outbound

This paper cites Speechcolab leaderboard, https://github.com/speechcolab/leaderboard, 2021.

TouchTTS: An Embarrassingly Simple TTS Framework that Everyone Can Touch Speechcolab leaderboard, https://github.com/speechcolab/leaderboard, 2021

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:08:23.031107Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-11T18:08:22.468176Z digest=sha256:26749ac3ebfe2879539e9f3648eeba1954038884a3ed274043ad157d07d0fc93

Pith citing papers

Observation b8573576-e615-4d8f-a016-8e8b90ae8429 · inbound

TouchASP: Elastic Automatic Speech Perception that Everyone Can Touch cites this paper.

TouchASP: Elastic Automatic Speech Perception that Everyone Can Touch TouchTTS: An Embarrassingly Simple TTS Framework that Everyone Can Touch

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T11:18:33.039702Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:18:33.039702Z digest=sha256:b61c214ebf7694a4730fe66c3a97b91319d41498228e4d72808cf36fc8430432

Observation 06bef987-0894-470b-9668-add4ed54f302 · inbound

Interleaved Speech-Text Language Models for Simple Streaming Text-to-Speech Synthesis cites this paper.

Interleaved Speech-Text Language Models for Simple Streaming Text-to-Speech Synthesis TouchTTS: An Embarrassingly Simple TTS Framework that Everyone Can Touch

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-11T10:50:49.009679Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T10:50:49.009679Z digest=sha256:f7ceb74e6781ed6cb40700b10a83a8ec47a6b4a9cbaa1ccee8d98afdce71d7cf

Observation 649e4fff-8935-4783-895b-cab52ba5cc8d · inbound

CosyVoice 3: Towards In-the-wild Speech Generation via Scaling-up and Post-training cites this paper.

CosyVoice 3: Towards In-the-wild Speech Generation via Scaling-up and Post-training TouchTTS: An Embarrassingly Simple TTS Framework that Everyone Can Touch

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-16T05:27:25.549297Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-16T05:27:25.425188Z digest=sha256:175ec2bda8f61c6eb94c090fe65c0441f2df62e080b09617f8dcd05948e63d81

Observation ae15d67d-1ce5-49f0-a8db-e26d0cd91dcd · inbound

Balalaika: Data-Centric, Prosody-Aware Annotation Pipeline for Russian Speech cites this paper.

Balalaika: Data-Centric, Prosody-Aware Annotation Pipeline for Russian Speech TouchTTS: An Embarrassingly Simple TTS Framework that Everyone Can Touch

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-06T16:30:09.691792Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:30:09.691792Z digest=sha256:ecdd6d6801d171ca0ef7adc296bd65e11e6cbbf91bc21052ab7d74bb8a57f2d9

Observation 4b5bb2d7-2828-4156-b580-53b0b7c92847 · inbound

DMOSpeech 2: Reinforcement Learning for Duration Prediction in Metric-Optimized Speech Synthesis cites this paper.

DMOSpeech 2: Reinforcement Learning for Duration Prediction in Metric-Optimized Speech Synthesis TouchTTS: An Embarrassingly Simple TTS Framework that Everyone Can Touch

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:20.978975Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:48:20.978975Z digest=sha256:8a27a7ea1ca914afaa8a34b9b6563266408c53937be856285601fd24efe2ba85

Observation 8fbcb36f-463e-406c-b347-f32bcbc11b1e · inbound

Borderless Long Speech Synthesis cites this paper.

Borderless Long Speech Synthesis TouchTTS: An Embarrassingly Simple TTS Framework that Everyone Can Touch

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-15T07:39:50.196070Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-15T07:39:45.386208Z digest=sha256:c5a855ee51d390f16a4ec4207e1ac341946728b9cf91f27cee6a9059f6d5040d

Observation d48e1a1e-4f92-43dc-9f48-dbc23b520d2e · inbound

Sarashina2.2-TTS: Tackling Kanji Polyphony in Japanese Speech Generation via Data Scaling and Targeted Data Synthesis cites this paper.

Sarashina2.2-TTS: Tackling Kanji Polyphony in Japanese Speech Generation via Data Scaling and Targeted Data Synthesis TouchTTS: An Embarrassingly Simple TTS Framework that Everyone Can Touch

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-07-04T20:10:07.250749Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-25T20:43:29.118351Z digest=sha256:5b89f52002e3fd2d8698740f27783fec5dd6f445ff59944dfee775c8a1dcf17f

Observation 849841b5-c2c5-445e-be47-e9310f0da1e6 · inbound

Experience-Calibrated Contrastive Decoding for Mitigating Hallucinations in LM-Based Text-to-Speech cites this paper.

Experience-Calibrated Contrastive Decoding for Mitigating Hallucinations in LM-Based Text-to-Speech TouchTTS: An Embarrassingly Simple TTS Framework that Everyone Can Touch

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-15T15:22:08.946434Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T15:22:08.946434Z digest=sha256:2e0b0ffc20ca1920ed358b7eed4dee569ffbfce21ac5a038c4b6dfcd101a6e08