Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T14:53:18.993310Z
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 47 of 47 outbound references and 0 inbound Pith citation observations for arXiv:2505.17426.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T14:53:18.993310Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
47 of 47 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 9d33002b-6b3c-47b2-89c1-3fcdaa12fcef · outbound
UniTTS: An end-to-end TTS system without decoupling of acoustic and semantic information Seed-TTS: A Family of High-Quality Versatile Speech Generation Models
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e4d26acb-124c-476a-b934-3f4e48d06b50 · outbound
UniTTS: An end-to-end TTS system without decoupling of acoustic and semantic information wav2vec 2.0: A framework for self-supervised learning of speech represen tations
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 966c5132-20c0-440d-b084-c0e5db6586e4 · outbound
UniTTS: An end-to-end TTS system without decoupling of acoustic and semantic information F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ec869804-b9ce-409a-9c1f-486251b2b3c6 · outbound
UniTTS: An end-to-end TTS system without decoupling of acoustic and semantic information High Fidelity Neural Audio Compression
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 213665fe-a09a-410c-8c8e-d40078715e2e · outbound
UniTTS: An end-to-end TTS system without decoupling of acoustic and semantic information Moshi: a speech-text foundation model for real-time dialogue
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cceffa41-b9b1-4038-bf68-672cd9a46730 · outbound
UniTTS: An end-to-end TTS system without decoupling of acoustic and semantic information IndexTTS: An Industrial-Level Controllable and Efficient Zero-Shot Text-To-Speech System
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e607820d-e8a4-430f-955c-79c2edb8e86d · outbound
UniTTS: An end-to-end TTS system without decoupling of acoustic and semantic information CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0facc9cb-428e-4f1e-8901-17226756ef20 · outbound
UniTTS: An end-to-end TTS system without decoupling of acoustic and semantic information CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a113fd37-95e0-4b9c-8a1b-a3282570a878 · outbound
UniTTS: An end-to-end TTS system without decoupling of acoustic and semantic information Paraformer: Fast and Accurate Parallel Transformer for Non-autoregressive End-to-End Speech Recognition
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7b0deeda-b6f5-4529-8ee7-58b20615f6a1 · outbound
UniTTS: An end-to-end TTS system without decoupling of acoustic and semantic information The Llama 3 Herd of Models
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2347bbc7-4675-4758-900f-fef985939400 · outbound
UniTTS: An end-to-end TTS system without decoupling of acoustic and semantic information Emilia: An extensive, multilingual, and diverse speech dataset for large-scale speech generation
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation a56d4796-c3f9-436d-91ed-da5e3a3e48f0 · outbound
UniTTS: An end-to-end TTS system without decoupling of acoustic and semantic information Hubert: Self-supervised spe ech representation learning by masked prediction of hidden units, 2021
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation f9ef53f3-404f-46bb-8008-bc5d8b4d1de1 · outbound
UniTTS: An end-to-end TTS system without decoupling of acoustic and semantic information WavTokenizer: an Efficient Acoustic Discrete Codec Tokenizer for Audio Language Modeling
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fcdc4326-4c69-4601-a2ae-5b9241fa4864 · outbound
UniTTS: An end-to-end TTS system without decoupling of acoustic and semantic information Libriheavy: A 50,000 hours asr corpus with punctuation casing and con- text
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation a3f97cc6-f5d9-4d2f-8b49-b022605b9663 · outbound
UniTTS: An end-to-end TTS system without decoupling of acoustic and semantic information Scaling Laws for Neural Language Models
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dbc7af43-37c4-44bd-94e4-40c2e31177c9 · outbound
UniTTS: An end-to-end TTS system without decoupling of acoustic and semantic information Hifi-gan: Generative adversarial networks for efficient and high fidelity speech synthesis
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 15d8824e-45e5-40de-add5-0652964ca1f8 · outbound
UniTTS: An end-to-end TTS system without decoupling of acoustic and semantic information Single-Codec: Single-Codebook Speech Codec towards High-Performance Speech Generation
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1bafbe8f-77ba-41e9-bc48-584777613c6e · outbound
UniTTS: An end-to-end TTS system without decoupling of acoustic and semantic information Fish-Speech: Leveraging Large Language Models for Advanced Multilingual Text-to-Speech Synthesis
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d7a1aebb-b7f9-4d63-b941-2dfb44717064 · outbound
UniTTS: An end-to-end TTS system without decoupling of acoustic and semantic information Unitok: A unified tokenizer for visual generati on and understanding
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 30bb3cbe-ee07-4b8c-98c1-e66a3eca2135 · outbound
UniTTS: An end-to-end TTS system without decoupling of acoustic and semantic information WenetSpeech4TTS: A 12,800-hour Mandarin TTS Corpus for Large Speech Generation Model Benchmark
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 10d591a9-3c72-4840-b415-418f7995e5cd · outbound
UniTTS: An end-to-end TTS system without decoupling of acoustic and semantic information Autoregressive Speech Synthesis without Vector Quantization
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e04f95dd-98d0-4488-a399-b92592437938 · outbound
UniTTS: An end-to-end TTS system without decoupling of acoustic and semantic information Finite Scalar Quantization: VQ-VAE Made Simple
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 53fa332c-d27a-43da-966e-4b834a73cad8 · outbound
UniTTS: An end-to-end TTS system without decoupling of acoustic and semantic information Scaling Transformers for Low-Bitrate High-Quality Speech Coding
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 72b749f9-c540-4490-8fa4-134eb682d2ac · outbound
UniTTS: An end-to-end TTS system without decoupling of acoustic and semantic information Loss-sensitive generative adversarial ne tworks on lipschitz densities, 2018
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 87751d8b-54bb-406f-aeb7-ba408ac83444 · outbound
UniTTS: An end-to-end TTS system without decoupling of acoustic and semantic information Robust speech recognition via large-scale weak supervision
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation dc727f44-2f28-4626-9890-34064b2961a3 · outbound
UniTTS: An end-to-end TTS system without decoupling of acoustic and semantic information Language models are unsupervised multitask learners
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation f120fd47-36e9-42e9-ae14-53b979da2abb · outbound
UniTTS: An end-to-end TTS system without decoupling of acoustic and semantic information Direct preference optimization: Y our language model is secretly a reward model
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 9701ba54-0ecd-4c53-9aae-316cfd4e8577 · outbound
UniTTS: An end-to-end TTS system without decoupling of acoustic and semantic information Dnsmos p
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation afef3ee6-190a-4445-9f43-6feb6b7343c1 · outbound
UniTTS: An end-to-end TTS system without decoupling of acoustic and semantic information Utmos: Utokyo-sarulab system for voicem os challenge 2022, 2022
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 66a420ae-6782-43af-96c9-7844c3edf45c · outbound
UniTTS: An end-to-end TTS system without decoupling of acoustic and semantic information Neural discret e representation learning
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation f6fd71d8-af69-4bf9-956a-c795fc93f5f9 · outbound
UniTTS: An end-to-end TTS system without decoupling of acoustic and semantic information Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2a745d27-1788-4f6c-9bd8-47e1b4265e34 · outbound
UniTTS: An end-to-end TTS system without decoupling of acoustic and semantic information Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6b94630e-26d1-41bc-9996-33324ce4bddc · outbound
UniTTS: An end-to-end TTS system without decoupling of acoustic and semantic information Convnext v2: Co-designing and scaling conv nets with masked autoencoders, 2023
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation bdc4cb98-4edf-4f45-80c6-4ec8e460e120 · outbound
UniTTS: An end-to-end TTS system without decoupling of acoustic and semantic information BigCodec: Pushing the Limits of Low-Bitrate Neural Speech Codec
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0abc68c5-c879-4e25-ab80-8811ab6e65ca · outbound
UniTTS: An end-to-end TTS system without decoupling of acoustic and semantic information Qwen2.5 Technical Report
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c17515ea-574e-4cd3-9cdd-f3aa64e2267c · outbound
UniTTS: An end-to-end TTS system without decoupling of acoustic and semantic information HiFi-Codec: Group-residual Vector quantization for High Fidelity Audio Codec
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e3ba37b1-24a5-4a5c-a7e0-b44ff18477d7 · outbound
UniTTS: An end-to-end TTS system without decoupling of acoustic and semantic information Codec does matter: Explo ring the semantic shortcoming of codec for audio language model
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 7f6a92ec-8d1c-4ad1-af56-c56ba9d1e6b1 · outbound
UniTTS: An end-to-end TTS system without decoupling of acoustic and semantic information Llasa: Scaling Train-Time and Inference-Time Compute for Llama-based Speech Synthesis
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e4532399-60e4-4eed-a94b-1a509428d199 · outbound
UniTTS: An end-to-end TTS system without decoupling of acoustic and semantic information Vector-quantized Image Modeling with Improved VQGAN
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 87603f81-26c6-43e7-9b68-d341b58d5cb1 · outbound
UniTTS: An end-to-end TTS system without decoupling of acoustic and semantic information Soundstream: An end-to-end neural audio codec
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 1af4cf71-c62d-4614-881a-7eb6c32833fc · outbound
UniTTS: An end-to-end TTS system without decoupling of acoustic and semantic information SpeechTokenizer: Unified Speech Tokenizer for Speech Large Language Models
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dee18e11-8a7f-4bf7-bd7a-55144e5b916e · outbound
UniTTS: An end-to-end TTS system without decoupling of acoustic and semantic information Scaling the Codebook Size of VQGAN to 100,000 with a Utilization Rate of 99%
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bf398295-290a-4cd1-a408-de9d69fe6597 · outbound
UniTTS: An end-to-end TTS system without decoupling of acoustic and semantic information Autoregressive spe ech synthesis with next-distribution prediction
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 40441d08-839b-4330-9113-0461d391477a · outbound
UniTTS: An end-to-end TTS system without decoupling of acoustic and semantic information Unresolved cited work
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 409776ea-b68d-4b8a-874b-af7b89ec1c41 · outbound
UniTTS: An end-to-end TTS system without decoupling of acoustic and semantic information Unresolved cited work
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 04b30b9b-3f9c-483b-ab74-e59c35936428 · outbound
UniTTS: An end-to-end TTS system without decoupling of acoustic and semantic information Unresolved cited work
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 173d7a8f-9474-45ea-9412-76d0b909fe51 · outbound
UniTTS: An end-to-end TTS system without decoupling of acoustic and semantic information 16 Table 13: Multi STFT Discriminator parameter settings of Di stilCodec
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
No inbound Pith citation observations are available.