Pith. sign in

Paper Citation Record · LEDGER

UniTTS: An end-to-end TTS system without decoupling of acoustic and semantic information

As of 18 August 2026, this Paper Citation Record lists 47 of 47 outbound references and 0 inbound Pith citation observations for arXiv:2505.17426.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.17426 v1

Coverage vector

measured 47 of 47 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:53:18.993310Z

measured 47 of 47 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

47 of 47 outbound references displayed

  • verified exact0
  • verified fuzzy15
  • unresolved31
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 9d33002b-6b3c-47b2-89c1-3fcdaa12fcef · outbound

This paper cites Seed-TTS: A Family of High-Quality Versatile Speech Generation Models.

UniTTS: An end-to-end TTS system without decoupling of acoustic and semantic information Seed-TTS: A Family of High-Quality Versatile Speech Generation Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T14:53:14.619002Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:53:14.619002Z digest=sha256:dfd5b0d2f2c0c2724a72e5bf2d6f84d3c5b60593fd97c6e3899b5b77aeec470f

Observation e4d26acb-124c-476a-b934-3f4e48d06b50 · outbound

This paper cites wav2vec 2.0: A framework for self-supervised learning of speech represen tations.

UniTTS: An end-to-end TTS system without decoupling of acoustic and semantic information wav2vec 2.0: A framework for self-supervised learning of speech represen tations

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:53:23.656153Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T14:53:14.670603Z digest=sha256:56c02b4f44740f88e4a2bc4e001c93c8d56e73176a7056a7814f3fdc017ee7b4

Observation 966c5132-20c0-440d-b084-c0e5db6586e4 · outbound

This paper cites F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching.

UniTTS: An end-to-end TTS system without decoupling of acoustic and semantic information F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T14:53:14.733869Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:53:14.733869Z digest=sha256:825d7cd0d7f87ff820a138303a48475026762d1b3bf6fb560dc0a2da55fe29ae

Observation ec869804-b9ce-409a-9c1f-486251b2b3c6 · outbound

This paper cites High Fidelity Neural Audio Compression.

UniTTS: An end-to-end TTS system without decoupling of acoustic and semantic information High Fidelity Neural Audio Compression

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T14:53:14.823243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:53:14.823243Z digest=sha256:f0791bcfdc1f5d20d5975540d4368f27956a9b9dfb033e05c6445ca8ab292970

Observation 213665fe-a09a-410c-8c8e-d40078715e2e · outbound

This paper cites Moshi: a speech-text foundation model for real-time dialogue.

UniTTS: An end-to-end TTS system without decoupling of acoustic and semantic information Moshi: a speech-text foundation model for real-time dialogue

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T14:53:14.907708Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:53:14.907708Z digest=sha256:3beee56543b13022f0cd56259babbb61ae04561c21072d13fa91fc4b32482033

Observation cceffa41-b9b1-4038-bf68-672cd9a46730 · outbound

This paper cites IndexTTS: An Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System.

UniTTS: An end-to-end TTS system without decoupling of acoustic and semantic information IndexTTS: An Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T14:53:14.979998Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:53:14.979998Z digest=sha256:938f7564a7107780b5e2fec96f52691a804daf10919be8067ecefb7b9e069d50

Observation e607820d-e8a4-430f-955c-79c2edb8e86d · outbound

This paper cites CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens.

UniTTS: An end-to-end TTS system without decoupling of acoustic and semantic information CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T14:53:15.064261Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:53:15.064261Z digest=sha256:e3db4776d4d2234725460a6826832ca0512f9cd99e441de44e54ac3164517759

Observation 0facc9cb-428e-4f1e-8901-17226756ef20 · outbound

This paper cites CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models.

UniTTS: An end-to-end TTS system without decoupling of acoustic and semantic information CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T14:53:15.130965Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:53:15.130965Z digest=sha256:beb9400ce31d01ed27722dd5223cb02dc537222981cab800e5570a7c33755ac4

Observation a113fd37-95e0-4b9c-8a1b-a3282570a878 · outbound

This paper cites Paraformer: Fast and Accurate Parallel Transformer for Non-autoregressive End-to-End Speech Recognition.

UniTTS: An end-to-end TTS system without decoupling of acoustic and semantic information Paraformer: Fast and Accurate Parallel Transformer for Non-autoregressive End-to-End Speech Recognition

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T14:53:15.212421Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:53:15.212421Z digest=sha256:c76ffc1884b74b87f46a3fec378505a4b8e951f07308c7887e8ee0963da2370a

Observation 7b0deeda-b6f5-4529-8ee7-58b20615f6a1 · outbound

This paper cites The Llama 3 Herd of Models.

UniTTS: An end-to-end TTS system without decoupling of acoustic and semantic information The Llama 3 Herd of Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T14:53:15.302882Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:53:15.302882Z digest=sha256:409aed46ada96be6e8b30e5b28c276147c0509e3467645ab193221bdbe4454a8

Observation 2347bbc7-4675-4758-900f-fef985939400 · outbound

This paper cites Emilia: An extensive, multilingual, and diverse speech dataset for large-scale speech generation.

UniTTS: An end-to-end TTS system without decoupling of acoustic and semantic information Emilia: An extensive, multilingual, and diverse speech dataset for large-scale speech generation

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:53:23.450367Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T14:53:15.463966Z digest=sha256:35e022dac71a181495e813ac08ae452c2bc45d43403653a6cc40ec505f3be598

Observation a56d4796-c3f9-436d-91ed-da5e3a3e48f0 · outbound

This paper cites Hubert: Self-supervised spe ech representation learning by masked prediction of hidden units, 2021.

UniTTS: An end-to-end TTS system without decoupling of acoustic and semantic information Hubert: Self-supervised spe ech representation learning by masked prediction of hidden units, 2021

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:53:23.285384Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T14:53:15.590347Z digest=sha256:f45a9cbe665a74c9dac355b1463835f66460349d70871f68457c1209b1ca79e9

Observation f9ef53f3-404f-46bb-8008-bc5d8b4d1de1 · outbound

This paper cites WavTokenizer: an Efficient Acoustic Discrete Codec Tokenizer for Audio Language Modeling.

UniTTS: An end-to-end TTS system without decoupling of acoustic and semantic information WavTokenizer: an Efficient Acoustic Discrete Codec Tokenizer for Audio Language Modeling

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T14:53:15.691624Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:53:15.691624Z digest=sha256:c513853242fceb41f4aaa718a710ce317f6102951bedb7996a46f37f1afb7df6

Observation fcdc4326-4c69-4601-a2ae-5b9241fa4864 · outbound

This paper cites Libriheavy: A 50,000 hours asr corpus with punctuation casing and con- text.

UniTTS: An end-to-end TTS system without decoupling of acoustic and semantic information Libriheavy: A 50,000 hours asr corpus with punctuation casing and con- text

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:53:23.159193Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T14:53:15.795893Z digest=sha256:0c6ecce5e5a92ba66e1885946957e461e29b119b3e317042abe4aa6f51c2ba53

Observation a3f97cc6-f5d9-4d2f-8b49-b022605b9663 · outbound

This paper cites Scaling Laws for Neural Language Models.

UniTTS: An end-to-end TTS system without decoupling of acoustic and semantic information Scaling Laws for Neural Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T14:53:15.934554Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:53:15.934554Z digest=sha256:48d4d09edfc21ede4eef9b4cc8a5128fe407899ca55c779c71ebfce08ce313f0

Observation dbc7af43-37c4-44bd-94e4-40c2e31177c9 · outbound

This paper cites Hifi-gan: Generative adversarial networks for efficient and high fidelity speech synthesis.

UniTTS: An end-to-end TTS system without decoupling of acoustic and semantic information Hifi-gan: Generative adversarial networks for efficient and high fidelity speech synthesis

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:53:23.005927Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T14:53:16.069772Z digest=sha256:dc59d7df139c93211c9e9e29d8a6b07e2c88145a3b38a64c523a44ffcc6f80d7

Observation 15d8824e-45e5-40de-add5-0652964ca1f8 · outbound

This paper cites Single-Codec: Single-Codebook Speech Codec towards High-Performance Speech Generation.

UniTTS: An end-to-end TTS system without decoupling of acoustic and semantic information Single-Codec: Single-Codebook Speech Codec towards High-Performance Speech Generation

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T14:53:16.222193Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:53:16.222193Z digest=sha256:045979444c43981b18883b2dcf1f08cc691d39be3986777083583092da2c17f2

Observation 1bafbe8f-77ba-41e9-bc48-584777613c6e · outbound

This paper cites Fish-Speech: Leveraging Large Language Models for Advanced Multilingual Text-to-Speech Synthesis.

UniTTS: An end-to-end TTS system without decoupling of acoustic and semantic information Fish-Speech: Leveraging Large Language Models for Advanced Multilingual Text-to-Speech Synthesis

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T14:53:16.316201Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:53:16.316201Z digest=sha256:c7ea6f82a80196e865f6c7729b1b1e57898cfaf34bef95b53146c398aea89648

Observation d7a1aebb-b7f9-4d63-b941-2dfb44717064 · outbound

This paper cites Unitok: A unified tokenizer for visual generati on and understanding.

UniTTS: An end-to-end TTS system without decoupling of acoustic and semantic information Unitok: A unified tokenizer for visual generati on and understanding

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T14:53:16.366904Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:53:16.366904Z digest=sha256:29a2fd8f0a7c62e80dfb8780403ba34ac469b022e00ffb252f3540d01848f131

Observation 30bb3cbe-ee07-4b8c-98c1-e66a3eca2135 · outbound

This paper cites WenetSpeech4TTS: A 12,800-hour Mandarin TTS Corpus for Large Speech Generation Model Benchmark.

UniTTS: An end-to-end TTS system without decoupling of acoustic and semantic information WenetSpeech4TTS: A 12,800-hour Mandarin TTS Corpus for Large Speech Generation Model Benchmark

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T14:53:16.449290Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:53:16.449290Z digest=sha256:9b7843a19f894c28a700ce52c8bf747aaecf4491180f85bff1c5e2c816b955ed

Observation 10d591a9-3c72-4840-b415-418f7995e5cd · outbound

This paper cites Autoregressive Speech Synthesis without Vector Quantization.

UniTTS: An end-to-end TTS system without decoupling of acoustic and semantic information Autoregressive Speech Synthesis without Vector Quantization

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T14:53:16.516851Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:53:16.516851Z digest=sha256:2ebe74242e903ebcdba8c1700ff65112af6b4889b57ccc5e74c2d024d6a982c7

Observation e04f95dd-98d0-4488-a399-b92592437938 · outbound

This paper cites Finite Scalar Quantization: VQ-VAE Made Simple.

UniTTS: An end-to-end TTS system without decoupling of acoustic and semantic information Finite Scalar Quantization: VQ-VAE Made Simple

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T14:53:16.616864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:53:16.616864Z digest=sha256:11e0cc46ba4ef72b326cd0aac68892f819913a6cc34a39685d5e0ad6021ec9fb

Observation 53fa332c-d27a-43da-966e-4b834a73cad8 · outbound

This paper cites Scaling Transformers for Low-Bitrate High-Quality Speech Coding.

UniTTS: An end-to-end TTS system without decoupling of acoustic and semantic information Scaling Transformers for Low-Bitrate High-Quality Speech Coding

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T14:53:16.774762Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:53:16.774762Z digest=sha256:54b8c2797e79f6078bb8f4c9bf47049bc56665db0232508cf89aea0c003bc00a

Observation 72b749f9-c540-4490-8fa4-134eb682d2ac · outbound

This paper cites Loss-sensitive generative adversarial ne tworks on lipschitz densities, 2018.

UniTTS: An end-to-end TTS system without decoupling of acoustic and semantic information Loss-sensitive generative adversarial ne tworks on lipschitz densities, 2018

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:53:22.341747Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T14:53:16.862660Z digest=sha256:46b0ba2b3c449a6d3f72b51c07d5f999bb209ecd161d1bf3fed99ea7fa047818

Observation 87751d8b-54bb-406f-aeb7-ba408ac83444 · outbound

This paper cites Robust speech recognition via large-scale weak supervision.

UniTTS: An end-to-end TTS system without decoupling of acoustic and semantic information Robust speech recognition via large-scale weak supervision

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:53:21.518013Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T14:53:16.955473Z digest=sha256:320c426abf8f31196c6225bd23475ec0942e9546f117a9faa7cc6817be82291a

Observation dc727f44-2f28-4626-9890-34064b2961a3 · outbound

This paper cites Language models are unsupervised multitask learners.

UniTTS: An end-to-end TTS system without decoupling of acoustic and semantic information Language models are unsupervised multitask learners

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:53:21.364048Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T14:53:17.076214Z digest=sha256:d184b7146c74b1554b3d108b44ef2687448e8c3336e4048a04398612684539f4

Observation f120fd47-36e9-42e9-ae14-53b979da2abb · outbound

This paper cites Direct preference optimization: Y our language model is secretly a reward model.

UniTTS: An end-to-end TTS system without decoupling of acoustic and semantic information Direct preference optimization: Y our language model is secretly a reward model

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:53:21.223798Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T14:53:17.171284Z digest=sha256:509b86adf24cfcd7f5cb50154498f1b85b575c2ed49ab0390004ae2b6f292afc

Observation 9701ba54-0ecd-4c53-9aae-316cfd4e8577 · outbound

This paper cites Dnsmos p.

UniTTS: An end-to-end TTS system without decoupling of acoustic and semantic information Dnsmos p

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:53:21.106441Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T14:53:17.261661Z digest=sha256:b9856259d02141a95f4477caf5d6e1d8a1afb71f8fb63501d2d76b08cc2bae06

Observation afef3ee6-190a-4445-9f43-6feb6b7343c1 · outbound

This paper cites Utmos: Utokyo-sarulab system for voicem os challenge 2022, 2022.

UniTTS: An end-to-end TTS system without decoupling of acoustic and semantic information Utmos: Utokyo-sarulab system for voicem os challenge 2022, 2022

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:53:20.984555Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T14:53:17.331644Z digest=sha256:b94be97df49e1e8e060efa6eee53e481c649bb8cba626cb500ad121165fadd77

Observation 66a420ae-6782-43af-96c9-7844c3edf45c · outbound

This paper cites Neural discret e representation learning.

UniTTS: An end-to-end TTS system without decoupling of acoustic and semantic information Neural discret e representation learning

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:53:20.810423Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T14:53:17.426791Z digest=sha256:14f5477aed13db84c7df0dbfb9b0994fbe599a6ad744118c072b4e85f44535a5

Observation f6fd71d8-af69-4bf9-956a-c795fc93f5f9 · outbound

This paper cites Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers.

UniTTS: An end-to-end TTS system without decoupling of acoustic and semantic information Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T14:53:17.522496Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:53:17.522496Z digest=sha256:b9d513ac970b47bd5299ccf78d28b4b0b0543f232fec3b4697976551f025a504

Observation 2a745d27-1788-4f6c-9bd8-47e1b4265e34 · outbound

This paper cites Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens.

UniTTS: An end-to-end TTS system without decoupling of acoustic and semantic information Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T14:53:17.592637Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:53:17.592637Z digest=sha256:58db9d189008a123949a3af69da78accd1cabdcf7097258364f703795c84fd28

Observation 6b94630e-26d1-41bc-9996-33324ce4bddc · outbound

This paper cites Convnext v2: Co-designing and scaling conv nets with masked autoencoders, 2023.

UniTTS: An end-to-end TTS system without decoupling of acoustic and semantic information Convnext v2: Co-designing and scaling conv nets with masked autoencoders, 2023

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:53:20.707290Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T14:53:17.651875Z digest=sha256:9a595aae479880399cc0d7c60bfa091c0ce75fbbea90f4ba756990676b711fc7

Observation bdc4cb98-4edf-4f45-80c6-4ec8e460e120 · outbound

This paper cites BigCodec: Pushing the Limits of Low-Bitrate Neural Speech Codec.

UniTTS: An end-to-end TTS system without decoupling of acoustic and semantic information BigCodec: Pushing the Limits of Low-Bitrate Neural Speech Codec

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T14:53:17.728250Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:53:17.728250Z digest=sha256:e52eeea42f357f179f9d4fce76ca9582de80a3f118e70ae5d953ed8f91e5d640

Observation 0abc68c5-c879-4e25-ab80-8811ab6e65ca · outbound

This paper cites Qwen2.5 Technical Report.

UniTTS: An end-to-end TTS system without decoupling of acoustic and semantic information Qwen2.5 Technical Report

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T14:53:17.785623Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:53:17.785623Z digest=sha256:5969d927d961e710c77f49a693d2a9e3f048fe2b5f25ff2b16599a101484a84c

Observation c17515ea-574e-4cd3-9cdd-f3aa64e2267c · outbound

This paper cites HiFi-Codec: Group-residual Vector quantization for High Fidelity Audio Codec.

UniTTS: An end-to-end TTS system without decoupling of acoustic and semantic information HiFi-Codec: Group-residual Vector quantization for High Fidelity Audio Codec

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T14:53:17.856062Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:53:17.856062Z digest=sha256:65be4f9a81ba1141960f7a8f64b8ff96995e8bfc539767ca08bc9beea8931f72

Observation e3ba37b1-24a5-4a5c-a7e0-b44ff18477d7 · outbound

This paper cites Codec does matter: Explo ring the semantic shortcoming of codec for audio language model.

UniTTS: An end-to-end TTS system without decoupling of acoustic and semantic information Codec does matter: Explo ring the semantic shortcoming of codec for audio language model

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:53:20.545166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T14:53:17.908432Z digest=sha256:afa0d7ee7a72106b0f720ddb64fb561cf5e9d212e3eebcd7252a6d1e23e14436

Observation 7f6a92ec-8d1c-4ad1-af56-c56ba9d1e6b1 · outbound

This paper cites Llasa: Scaling Train-Time and Inference-Time Compute for Llama-based Speech Synthesis.

UniTTS: An end-to-end TTS system without decoupling of acoustic and semantic information Llasa: Scaling Train-Time and Inference-Time Compute for Llama-based Speech Synthesis

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T14:53:17.983678Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:53:17.983678Z digest=sha256:b0935af18ea14a2e9b324e7962ecc093ec2f56c72153bb0c63b28eed863b1513

Observation e4532399-60e4-4eed-a94b-1a509428d199 · outbound

This paper cites Vector-quantized Image Modeling with Improved VQGAN.

UniTTS: An end-to-end TTS system without decoupling of acoustic and semantic information Vector-quantized Image Modeling with Improved VQGAN

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T14:53:18.048407Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:53:18.048407Z digest=sha256:bfde6ed1dbfd20d83696a2cd7c17b4c8d3d8e3dbe15e61e0cbf4f482cafdce09

Observation 87603f81-26c6-43e7-9b68-d341b58d5cb1 · outbound

This paper cites Soundstream: An end-to-end neural audio codec.

UniTTS: An end-to-end TTS system without decoupling of acoustic and semantic information Soundstream: An end-to-end neural audio codec

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:53:20.421251Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T14:53:18.148272Z digest=sha256:ae706151515d30a41af262f8df6113945eac3e18e1c91df40597821cc4832e5c

Observation 1af4cf71-c62d-4614-881a-7eb6c32833fc · outbound

This paper cites SpeechTokenizer: Unified Speech Tokenizer for Speech Large Language Models.

UniTTS: An end-to-end TTS system without decoupling of acoustic and semantic information SpeechTokenizer: Unified Speech Tokenizer for Speech Large Language Models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T14:53:18.271919Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:53:18.271919Z digest=sha256:cba7d0c6078a5e616a68486763b4b4b9c835de6889abf4123675aa88696953ae

Observation dee18e11-8a7f-4bf7-bd7a-55144e5b916e · outbound

This paper cites Scaling the Codebook Size of VQGAN to 100,000 with a Utilization Rate of 99%.

UniTTS: An end-to-end TTS system without decoupling of acoustic and semantic information Scaling the Codebook Size of VQGAN to 100,000 with a Utilization Rate of 99%

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T14:53:18.403349Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:53:18.403349Z digest=sha256:f431aa6e21a5768c35ac76f0c50e82168094f956277765ba071e23d08269be91

Observation bf398295-290a-4cd1-a408-de9d69fe6597 · outbound

This paper cites Autoregressive spe ech synthesis with next-distribution prediction.

UniTTS: An end-to-end TTS system without decoupling of acoustic and semantic information Autoregressive spe ech synthesis with next-distribution prediction

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T14:53:18.507278Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:53:18.507278Z digest=sha256:e29b54ef7c72a8b7b7b4c856387f304de28177e0a870660ef03498d6737c6591

Observation 40441d08-839b-4330-9113-0461d391477a · outbound

This paper cites an unresolved cited work.

UniTTS: An end-to-end TTS system without decoupling of acoustic and semantic information Unresolved cited work

Reference 44

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:53:20.293072Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T14:53:18.615503Z digest=sha256:e7941f1f7ea6b9cd0d65f3c4591cc78a00cedc8632fe9b6072859cfac632ccf3

Observation 409776ea-b68d-4b8a-874b-af7b89ec1c41 · outbound

This paper cites an unresolved cited work.

UniTTS: An end-to-end TTS system without decoupling of acoustic and semantic information Unresolved cited work

Reference 45

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:53:20.071455Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T14:53:18.743653Z digest=sha256:b665a990801f529780082df68707854a4aa09630a883c624dbdc6493819e4ae2

Observation 04b30b9b-3f9c-483b-ab74-e59c35936428 · outbound

This paper cites an unresolved cited work.

UniTTS: An end-to-end TTS system without decoupling of acoustic and semantic information Unresolved cited work

Reference 46

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:53:19.867514Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T14:53:18.852069Z digest=sha256:a979705323a777fe5da24da4ba1ee3e69ba1cdc429900315d1940b7eaab72740

Observation 173d7a8f-9474-45ea-9412-76d0b909fe51 · outbound

This paper cites 16 Table 13: Multi STFT Discriminator parameter settings of Di stilCodec.

UniTTS: An end-to-end TTS system without decoupling of acoustic and semantic information 16 Table 13: Multi STFT Discriminator parameter settings of Di stilCodec

Reference 47

Resolution
malformed identifier
raw_fallback, observed 2026-08-07T14:53:19.646311Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T14:53:18.993310Z digest=sha256:3a5760a8fc26c713530b4c4aa977d2d6eae033d1bbfa760b48cb21561e19cf0b

Pith citing papers

No inbound Pith citation observations are available.