Pith. sign in

Paper Citation Record · LEDGER

UniTTS: An end-to-end TTS system without decoupling of acoustic and semantic information

As of 8 August 2026, this Paper Citation Record lists 47 of 47 outbound references and 0 inbound Pith citation observations for arXiv:2505.17426.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.17426 v1

Coverage vector

measured 47 of 47 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:53:18.993310Z

measured 47 of 47 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

47 of 47 outbound references displayed

  • verified exact0
  • verified fuzzy15
  • unresolved31
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 9d33002b-6b3c-47b2-89c1-3fcdaa12fcef · outbound

This paper cites Seed-TTS: A Family of High-Quality Versatile Speech Generation Models.

UniTTS: An end-to-end TTS system without decoupling of acoustic and semantic information Seed-TTS: A Family of High-Quality Versatile Speech Generation Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T14:53:14.619002Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:53:14.619002Z digest=sha256:e560551bc847862899e0f694a255e5ed814a6bc687c9bcc911b8c28b70e0421b

Observation e4d26acb-124c-476a-b934-3f4e48d06b50 · outbound

This paper cites wav2vec 2.0: A framework for self-supervised learning of speech represen tations.

UniTTS: An end-to-end TTS system without decoupling of acoustic and semantic information wav2vec 2.0: A framework for self-supervised learning of speech represen tations

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:53:23.656153Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:53:14.670603Z digest=sha256:1345f4255984646e8c7ff176cddc8b507249dae7c7ee4452ef2bfb14c73ea43d

Observation 966c5132-20c0-440d-b084-c0e5db6586e4 · outbound

This paper cites F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching.

UniTTS: An end-to-end TTS system without decoupling of acoustic and semantic information F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T14:53:14.733869Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:53:14.733869Z digest=sha256:e46bd2e3e4750e15f1af7655a30bc093bca54534baba5187f5462ae41891e516

Observation ec869804-b9ce-409a-9c1f-486251b2b3c6 · outbound

This paper cites High Fidelity Neural Audio Compression.

UniTTS: An end-to-end TTS system without decoupling of acoustic and semantic information High Fidelity Neural Audio Compression

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T14:53:14.823243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:53:14.823243Z digest=sha256:b1d27dfd0e5e4aa78ed92992efb60ddb8233d75b8809a76c98e31bc2b596a689

Observation 213665fe-a09a-410c-8c8e-d40078715e2e · outbound

This paper cites Moshi: a speech-text foundation model for real-time dialogue.

UniTTS: An end-to-end TTS system without decoupling of acoustic and semantic information Moshi: a speech-text foundation model for real-time dialogue

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T14:53:14.907708Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:53:14.907708Z digest=sha256:5c54201dfa5cf1c332bb0fdcdcd85906d1183c72b34bd89e8bc7ba2f6d48848b

Observation cceffa41-b9b1-4038-bf68-672cd9a46730 · outbound

This paper cites IndexTTS: An Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System.

UniTTS: An end-to-end TTS system without decoupling of acoustic and semantic information IndexTTS: An Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T14:53:14.979998Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:53:14.979998Z digest=sha256:f2f45c612718c073de85dbf5f8827af7d00236f8b6cda27d4f5980a114ae0c5d

Observation e607820d-e8a4-430f-955c-79c2edb8e86d · outbound

This paper cites CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens.

UniTTS: An end-to-end TTS system without decoupling of acoustic and semantic information CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T14:53:15.064261Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:53:15.064261Z digest=sha256:163df763513704e489feb5e2c2c033a0c7fbbce728affa56af7d928ffc8942d6

Observation 0facc9cb-428e-4f1e-8901-17226756ef20 · outbound

This paper cites CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models.

UniTTS: An end-to-end TTS system without decoupling of acoustic and semantic information CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T14:53:15.130965Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:53:15.130965Z digest=sha256:4a16ed6e903df5c24dd8fff1a1a87ef286f7ba9511ba50d43d898c7791f222a3

Observation a113fd37-95e0-4b9c-8a1b-a3282570a878 · outbound

This paper cites Paraformer: Fast and Accurate Parallel Transformer for Non-autoregressive End-to-End Speech Recognition.

UniTTS: An end-to-end TTS system without decoupling of acoustic and semantic information Paraformer: Fast and Accurate Parallel Transformer for Non-autoregressive End-to-End Speech Recognition

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T14:53:15.212421Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:53:15.212421Z digest=sha256:93b3a656d66a8b153ef85f2f78771d81ce2a7c42cdd72d35f764e59ae9d46474

Observation 7b0deeda-b6f5-4529-8ee7-58b20615f6a1 · outbound

This paper cites The Llama 3 Herd of Models.

UniTTS: An end-to-end TTS system without decoupling of acoustic and semantic information The Llama 3 Herd of Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T14:53:15.302882Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:53:15.302882Z digest=sha256:3805e7be1c3a5e9f2ac3801964bca09a2657ffb72def13e5865931c563bc4fe9

Observation 2347bbc7-4675-4758-900f-fef985939400 · outbound

This paper cites Emilia: An extensive, multilingual, and diverse speech dataset for large-scale speech generation.

UniTTS: An end-to-end TTS system without decoupling of acoustic and semantic information Emilia: An extensive, multilingual, and diverse speech dataset for large-scale speech generation

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:53:23.450367Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:53:15.463966Z digest=sha256:074d3cd12ea10c2a1583c42e0fae2e551f6706c692982c61325300b84bb757ae

Observation a56d4796-c3f9-436d-91ed-da5e3a3e48f0 · outbound

This paper cites Hubert: Self-supervised spe ech representation learning by masked prediction of hidden units, 2021.

UniTTS: An end-to-end TTS system without decoupling of acoustic and semantic information Hubert: Self-supervised spe ech representation learning by masked prediction of hidden units, 2021

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:53:23.285384Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:53:15.590347Z digest=sha256:2d88fb560d24e771b212525a48ea632a3b2f25b1bab16c07a80ed5a109d52d83

Observation f9ef53f3-404f-46bb-8008-bc5d8b4d1de1 · outbound

This paper cites WavTokenizer: an Efficient Acoustic Discrete Codec Tokenizer for Audio Language Modeling.

UniTTS: An end-to-end TTS system without decoupling of acoustic and semantic information WavTokenizer: an Efficient Acoustic Discrete Codec Tokenizer for Audio Language Modeling

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T14:53:15.691624Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:53:15.691624Z digest=sha256:7dd6666b94b19ef763a5ce0fa60420a45d1a0c6259d7fed4491d06e366dd5d13

Observation fcdc4326-4c69-4601-a2ae-5b9241fa4864 · outbound

This paper cites Libriheavy: A 50,000 hours asr corpus with punctuation casing and con- text.

UniTTS: An end-to-end TTS system without decoupling of acoustic and semantic information Libriheavy: A 50,000 hours asr corpus with punctuation casing and con- text

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:53:23.159193Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:53:15.795893Z digest=sha256:bf4e5c1833cba3083ae57e644f736561d4dde63f4f7f9655f062dbe94668e657

Observation a3f97cc6-f5d9-4d2f-8b49-b022605b9663 · outbound

This paper cites Scaling Laws for Neural Language Models.

UniTTS: An end-to-end TTS system without decoupling of acoustic and semantic information Scaling Laws for Neural Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T14:53:15.934554Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:53:15.934554Z digest=sha256:797b7dd013cdc29f1b0508f1b63438e93038dd0d8e01edb7f9a90f19e5c265eb

Observation dbc7af43-37c4-44bd-94e4-40c2e31177c9 · outbound

This paper cites Hifi-gan: Generative adversarial networks for efficient and high fidelity speech synthesis.

UniTTS: An end-to-end TTS system without decoupling of acoustic and semantic information Hifi-gan: Generative adversarial networks for efficient and high fidelity speech synthesis

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:53:23.005927Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:53:16.069772Z digest=sha256:d49446f106c9e2d4adfb91ab9fe27130e5ebb5db35006597388860b2147599e6

Observation 15d8824e-45e5-40de-add5-0652964ca1f8 · outbound

This paper cites Single-Codec: Single-Codebook Speech Codec towards High-Performance Speech Generation.

UniTTS: An end-to-end TTS system without decoupling of acoustic and semantic information Single-Codec: Single-Codebook Speech Codec towards High-Performance Speech Generation

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T14:53:16.222193Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:53:16.222193Z digest=sha256:ec8517dbd7e6a0cbf236849afd35d91bd35285837747f58841e6d60f31f38f5d

Observation 1bafbe8f-77ba-41e9-bc48-584777613c6e · outbound

This paper cites Fish-Speech: Leveraging Large Language Models for Advanced Multilingual Text-to-Speech Synthesis.

UniTTS: An end-to-end TTS system without decoupling of acoustic and semantic information Fish-Speech: Leveraging Large Language Models for Advanced Multilingual Text-to-Speech Synthesis

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T14:53:16.316201Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:53:16.316201Z digest=sha256:eff238ff5cadac01ad17bc9f2704550c84362ae05db0fbdd07f5119a9e3c230b

Observation d7a1aebb-b7f9-4d63-b941-2dfb44717064 · outbound

This paper cites Unitok: A unified tokenizer for visual generati on and understanding.

UniTTS: An end-to-end TTS system without decoupling of acoustic and semantic information Unitok: A unified tokenizer for visual generati on and understanding

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T14:53:16.366904Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:53:16.366904Z digest=sha256:bc0c3e88b88bf30cbc8002d18133bbc1dbeacb88db96b5a72fa5df862ffc2726

Observation 30bb3cbe-ee07-4b8c-98c1-e66a3eca2135 · outbound

This paper cites WenetSpeech4TTS: A 12,800-hour Mandarin TTS Corpus for Large Speech Generation Model Benchmark.

UniTTS: An end-to-end TTS system without decoupling of acoustic and semantic information WenetSpeech4TTS: A 12,800-hour Mandarin TTS Corpus for Large Speech Generation Model Benchmark

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T14:53:16.449290Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:53:16.449290Z digest=sha256:a21146dfda7eed94a50566957a52aa05070f56dd933b85a6c25a13bf0c5429d1

Observation 10d591a9-3c72-4840-b415-418f7995e5cd · outbound

This paper cites Autoregressive Speech Synthesis without Vector Quantization.

UniTTS: An end-to-end TTS system without decoupling of acoustic and semantic information Autoregressive Speech Synthesis without Vector Quantization

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T14:53:16.516851Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:53:16.516851Z digest=sha256:f36c4ddf2d8815f793c2c060b5d25fed00a36fb8c8b109ce94e43cb903fa8daa

Observation e04f95dd-98d0-4488-a399-b92592437938 · outbound

This paper cites Finite Scalar Quantization: VQ-VAE Made Simple.

UniTTS: An end-to-end TTS system without decoupling of acoustic and semantic information Finite Scalar Quantization: VQ-VAE Made Simple

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T14:53:16.616864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:53:16.616864Z digest=sha256:b22559e6ffdde4298bdfda59924387cd0b99f774b2ed34d74dba29d9a94341ba

Observation 53fa332c-d27a-43da-966e-4b834a73cad8 · outbound

This paper cites Scaling Transformers for Low-Bitrate High-Quality Speech Coding.

UniTTS: An end-to-end TTS system without decoupling of acoustic and semantic information Scaling Transformers for Low-Bitrate High-Quality Speech Coding

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T14:53:16.774762Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:53:16.774762Z digest=sha256:e7d2b75c422b88b9f0bca415501914267f46d23a1828de8e78c1b3bb6a79acf6

Observation 72b749f9-c540-4490-8fa4-134eb682d2ac · outbound

This paper cites Loss-sensitive generative adversarial ne tworks on lipschitz densities, 2018.

UniTTS: An end-to-end TTS system without decoupling of acoustic and semantic information Loss-sensitive generative adversarial ne tworks on lipschitz densities, 2018

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:53:22.341747Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:53:16.862660Z digest=sha256:4f6a5b86d69e478aa7780a724d45240d8b526e1e4041ace4416f2f076f8dff0b

Observation 87751d8b-54bb-406f-aeb7-ba408ac83444 · outbound

This paper cites Robust speech recognition via large-scale weak supervision.

UniTTS: An end-to-end TTS system without decoupling of acoustic and semantic information Robust speech recognition via large-scale weak supervision

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:53:21.518013Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:53:16.955473Z digest=sha256:87ed877cb4f9a2310e82c7d303256d7a5866048f7bcc0765d793053828b52675

Observation dc727f44-2f28-4626-9890-34064b2961a3 · outbound

This paper cites Language models are unsupervised multitask learners.

UniTTS: An end-to-end TTS system without decoupling of acoustic and semantic information Language models are unsupervised multitask learners

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:53:21.364048Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:53:17.076214Z digest=sha256:68e410c07f7c5bcab454e834b251dad53666068fcd67244eaea10ef0a0f4255a

Observation f120fd47-36e9-42e9-ae14-53b979da2abb · outbound

This paper cites Direct preference optimization: Y our language model is secretly a reward model.

UniTTS: An end-to-end TTS system without decoupling of acoustic and semantic information Direct preference optimization: Y our language model is secretly a reward model

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:53:21.223798Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:53:17.171284Z digest=sha256:4d5ab8649ffb67d03f6d5d2b33003e54741bd87f82d1d22f5e5cb30324e6e7a6

Observation 9701ba54-0ecd-4c53-9aae-316cfd4e8577 · outbound

This paper cites Dnsmos p.

UniTTS: An end-to-end TTS system without decoupling of acoustic and semantic information Dnsmos p

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:53:21.106441Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:53:17.261661Z digest=sha256:96b5566034c5e99351f7f45d47cfaf4ea2c188576405f0123f425dc9c17cbfa1

Observation afef3ee6-190a-4445-9f43-6feb6b7343c1 · outbound

This paper cites Utmos: Utokyo-sarulab system for voicem os challenge 2022, 2022.

UniTTS: An end-to-end TTS system without decoupling of acoustic and semantic information Utmos: Utokyo-sarulab system for voicem os challenge 2022, 2022

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:53:20.984555Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:53:17.331644Z digest=sha256:004ee90fd554a4db311fed62b9c714a6fc79ae6e4af7f6d86b9f97cd40e37244

Observation 66a420ae-6782-43af-96c9-7844c3edf45c · outbound

This paper cites Neural discret e representation learning.

UniTTS: An end-to-end TTS system without decoupling of acoustic and semantic information Neural discret e representation learning

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:53:20.810423Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:53:17.426791Z digest=sha256:76c72d0ca29885398e7c3912812695733eff946be5ce7d5be9f2e452c89ca366

Observation f6fd71d8-af69-4bf9-956a-c795fc93f5f9 · outbound

This paper cites Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers.

UniTTS: An end-to-end TTS system without decoupling of acoustic and semantic information Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T14:53:17.522496Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:53:17.522496Z digest=sha256:5e8563e0b8fabb5bf83f4e545fcbc0f65a9537da838a3e742f4710a3a3b1d6c3

Observation 2a745d27-1788-4f6c-9bd8-47e1b4265e34 · outbound

This paper cites Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens.

UniTTS: An end-to-end TTS system without decoupling of acoustic and semantic information Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T14:53:17.592637Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:53:17.592637Z digest=sha256:0966bad455ce9593bdf4d3881e964ce4fe3c58d799d66ecc469a0bdae1372f09

Observation 6b94630e-26d1-41bc-9996-33324ce4bddc · outbound

This paper cites Convnext v2: Co-designing and scaling conv nets with masked autoencoders, 2023.

UniTTS: An end-to-end TTS system without decoupling of acoustic and semantic information Convnext v2: Co-designing and scaling conv nets with masked autoencoders, 2023

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:53:20.707290Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:53:17.651875Z digest=sha256:de27031525007a36807f4e7663a75f5e7eb2643ab17ea1e800d0569db600369a

Observation bdc4cb98-4edf-4f45-80c6-4ec8e460e120 · outbound

This paper cites BigCodec: Pushing the Limits of Low-Bitrate Neural Speech Codec.

UniTTS: An end-to-end TTS system without decoupling of acoustic and semantic information BigCodec: Pushing the Limits of Low-Bitrate Neural Speech Codec

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T14:53:17.728250Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:53:17.728250Z digest=sha256:f74d29a44c3f1e85f7e3ddc8807072501f832cdc8212a16a2be952b7ba6c5b65

Observation 0abc68c5-c879-4e25-ab80-8811ab6e65ca · outbound

This paper cites Qwen2.5 Technical Report.

UniTTS: An end-to-end TTS system without decoupling of acoustic and semantic information Qwen2.5 Technical Report

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T14:53:17.785623Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:53:17.785623Z digest=sha256:7501633d771b0385987e1e809e1170acec902c34dfd0afd535918bb33f5ab5aa

Observation c17515ea-574e-4cd3-9cdd-f3aa64e2267c · outbound

This paper cites HiFi-Codec: Group-residual Vector quantization for High Fidelity Audio Codec.

UniTTS: An end-to-end TTS system without decoupling of acoustic and semantic information HiFi-Codec: Group-residual Vector quantization for High Fidelity Audio Codec

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T14:53:17.856062Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:53:17.856062Z digest=sha256:16fb5aea940144715649c6e6acd41ce652c9581e2cfa268543696a928cc5dad2

Observation e3ba37b1-24a5-4a5c-a7e0-b44ff18477d7 · outbound

This paper cites Codec does matter: Explo ring the semantic shortcoming of codec for audio language model.

UniTTS: An end-to-end TTS system without decoupling of acoustic and semantic information Codec does matter: Explo ring the semantic shortcoming of codec for audio language model

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:53:20.545166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:53:17.908432Z digest=sha256:e1888fe43fbff364beb751ed430e41d0b7ff5c678ce253e1eddb1c0a07a2b1f0

Observation 7f6a92ec-8d1c-4ad1-af56-c56ba9d1e6b1 · outbound

This paper cites Llasa: Scaling Train-Time and Inference-Time Compute for Llama-based Speech Synthesis.

UniTTS: An end-to-end TTS system without decoupling of acoustic and semantic information Llasa: Scaling Train-Time and Inference-Time Compute for Llama-based Speech Synthesis

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T14:53:17.983678Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:53:17.983678Z digest=sha256:026d318ced41e3b451ea8d1b3d2eae94ade4fb446da930571451f3bf11764261

Observation e4532399-60e4-4eed-a94b-1a509428d199 · outbound

This paper cites Vector-quantized Image Modeling with Improved VQGAN.

UniTTS: An end-to-end TTS system without decoupling of acoustic and semantic information Vector-quantized Image Modeling with Improved VQGAN

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T14:53:18.048407Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:53:18.048407Z digest=sha256:2e7117e2798c1f35127e8ae7b3aec0ebfb14360a1c6fb925d940a02a68b80705

Observation 87603f81-26c6-43e7-9b68-d341b58d5cb1 · outbound

This paper cites Soundstream: An end-to-end neural audio codec.

UniTTS: An end-to-end TTS system without decoupling of acoustic and semantic information Soundstream: An end-to-end neural audio codec

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:53:20.421251Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:53:18.148272Z digest=sha256:fb4b44132e63dc0b04c8a37570f03b79467ac2960e92ff15461d1c8e520971f0

Observation 1af4cf71-c62d-4614-881a-7eb6c32833fc · outbound

This paper cites SpeechTokenizer: Unified Speech Tokenizer for Speech Large Language Models.

UniTTS: An end-to-end TTS system without decoupling of acoustic and semantic information SpeechTokenizer: Unified Speech Tokenizer for Speech Large Language Models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T14:53:18.271919Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:53:18.271919Z digest=sha256:3745c0a7313d4082874de05f861056e3ee5caec5fd75b6fb00c650c3fd1898b5

Observation dee18e11-8a7f-4bf7-bd7a-55144e5b916e · outbound

This paper cites Scaling the Codebook Size of VQGAN to 100,000 with a Utilization Rate of 99%.

UniTTS: An end-to-end TTS system without decoupling of acoustic and semantic information Scaling the Codebook Size of VQGAN to 100,000 with a Utilization Rate of 99%

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T14:53:18.403349Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:53:18.403349Z digest=sha256:f7c8d8ba52d026c41a8f6f8769db7f5b71089a5b1143fc6167d5fa990c446178

Observation bf398295-290a-4cd1-a408-de9d69fe6597 · outbound

This paper cites Autoregressive spe ech synthesis with next-distribution prediction.

UniTTS: An end-to-end TTS system without decoupling of acoustic and semantic information Autoregressive spe ech synthesis with next-distribution prediction

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T14:53:18.507278Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:53:18.507278Z digest=sha256:3331389858a58eb1ed5c297e46d5f73e5ef1f28d5509a34601b07c679a55efd7

Observation 40441d08-839b-4330-9113-0461d391477a · outbound

This paper cites an unresolved cited work.

UniTTS: An end-to-end TTS system without decoupling of acoustic and semantic information Unresolved cited work

Reference 44

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:53:20.293072Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:53:18.615503Z digest=sha256:7fb31c93972067ef2bcc5e4b67da884598ca85d0bf9e95b4591fd61c94e37477

Observation 409776ea-b68d-4b8a-874b-af7b89ec1c41 · outbound

This paper cites an unresolved cited work.

UniTTS: An end-to-end TTS system without decoupling of acoustic and semantic information Unresolved cited work

Reference 45

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:53:20.071455Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:53:18.743653Z digest=sha256:f1421e36584f1c05042e5d5c2612f1510fef4b38ce86fe4f1d28f024b974b05c

Observation 04b30b9b-3f9c-483b-ab74-e59c35936428 · outbound

This paper cites an unresolved cited work.

UniTTS: An end-to-end TTS system without decoupling of acoustic and semantic information Unresolved cited work

Reference 46

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:53:19.867514Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:53:18.852069Z digest=sha256:7c8a35e6ef1f8af863d809efbec6257ad5f5ee3cc467e0ac2d937ebfaf677539

Observation 173d7a8f-9474-45ea-9412-76d0b909fe51 · outbound

This paper cites 16 Table 13: Multi STFT Discriminator parameter settings of Di stilCodec.

UniTTS: An end-to-end TTS system without decoupling of acoustic and semantic information 16 Table 13: Multi STFT Discriminator parameter settings of Di stilCodec

Reference 47

Resolution
malformed identifier
raw_fallback, observed 2026-08-07T14:53:19.646311Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:53:18.993310Z digest=sha256:8a84112ea19b7bec1ae824047db9b8803b2e54031d9d73a0c00243deaad9e345

Pith citing papers

No inbound Pith citation observations are available.