Pith. sign in

Paper Citation Record · LEDGER

CLEAR: Continuous Latent Autoregressive Modeling for High-quality and Low-latency Speech Synthesis

As of 18 August 2026, this Paper Citation Record lists 69 of 69 outbound references and 3 inbound Pith citation observations for arXiv:2508.19098.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.19098 v1

Coverage vector

measured 69 of 69 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T16:01:52.889397Z

measured 72 of 72 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-27T21:18:22.911332Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T19:47:19.713205Z

Reference resolution

69 of 69 outbound references displayed

  • verified exact0
  • verified fuzzy33
  • unresolved34
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 958bcb74-b0d2-4b6d-b750-76a12a654377 · outbound

This paper cites Seed-TTS: A Family of High-Quality Versatile Speech Generation Models.

CLEAR: Continuous Latent Autoregressive Modeling for High-quality and Low-latency Speech Synthesis Seed-TTS: A Family of High-Quality Versatile Speech Generation Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-05T16:01:52.500071Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:01:52.500071Z digest=sha256:24fff7d0dccd7fee95c0cf5c57d7a7aaa1cbd2e95c773195da84f560c92dfb77

Observation ce96466a-3f3c-498a-b4c6-699d1b217b7a · outbound

This paper cites SpeechT5: Unified-Modal Encoder-Decoder Pre-Training for Spoken Language Processing, May 2022.

CLEAR: Continuous Latent Autoregressive Modeling for High-quality and Low-latency Speech Synthesis SpeechT5: Unified-Modal Encoder-Decoder Pre-Training for Spoken Language Processing, May 2022

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:01:54.886110Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T16:01:52.506197Z digest=sha256:2adcaf77e4a36f4be1efd934c3520a1e2ee6f9d8b35c881f0ec3846102826aca

Observation ee274c21-6231-4378-b2c4-d7d16bd2c174 · outbound

This paper cites Rethinking lossy compression: The rate-distortion-perception tradeoff.

CLEAR: Continuous Latent Autoregressive Modeling for High-quality and Low-latency Speech Synthesis Rethinking lossy compression: The rate-distortion-perception tradeoff

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:01:54.863775Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T16:01:52.511486Z digest=sha256:e7feaa63b116911360a3c61045879f43219529238459b5d94bf11a595e46aaa3

Observation 9568581b-d05c-427d-bcda-b47528a0d1a1 · outbound

This paper cites Audiolm: a language modeling approach to audio generation.

CLEAR: Continuous Latent Autoregressive Modeling for High-quality and Low-latency Speech Synthesis Audiolm: a language modeling approach to audio generation

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:01:54.843458Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T16:01:52.517582Z digest=sha256:277efa8d8be72392b3e0d941a13679c350f0d456ea85b7a4bc3641f5c0a04aa8

Observation 90aa4c5a-2bd8-46eb-8908-16a6a917b2c7 · outbound

This paper cites Deep Compression Autoencoder for Efficient High-Resolution Diffusion Models, April 2025.

CLEAR: Continuous Latent Autoregressive Modeling for High-quality and Low-latency Speech Synthesis Deep Compression Autoencoder for Efficient High-Resolution Diffusion Models, April 2025

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:01:54.823149Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T16:01:52.523257Z digest=sha256:30f61b018d1a16ca98d3a96f0f2d2e4fe797a0b69cc9239e350edbd308b96239

Observation 19a81d71-38a8-4e58-8a48-58b6b09987ee · outbound

This paper cites Wavlm: Large-scale self-supervised pre- training for full stack speech processing.

CLEAR: Continuous Latent Autoregressive Modeling for High-quality and Low-latency Speech Synthesis Wavlm: Large-scale self-supervised pre- training for full stack speech processing

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:01:54.802618Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T16:01:52.528372Z digest=sha256:fca566598ffb41588226b82508b4aa1d43c458d552be92e16b8578f1a82b4b1e

Observation 53ce4747-da48-485a-b0f5-1a133c8439b7 · outbound

This paper cites VALL-E 2: Neural Codec Language Models are Human Parity Zero-Shot Text to Speech Synthesizers.

CLEAR: Continuous Latent Autoregressive Modeling for High-quality and Low-latency Speech Synthesis VALL-E 2: Neural Codec Language Models are Human Parity Zero-Shot Text to Speech Synthesizers

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-05T16:01:52.534536Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:01:52.534536Z digest=sha256:9d17d0d42ba893e2e7550df10d8d0ebdb09b179f52c74dad7162028b7d24a376

Observation 745110e8-993f-4339-8a4b-51e03547cfbb · outbound

This paper cites F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching.

CLEAR: Continuous Latent Autoregressive Modeling for High-quality and Low-latency Speech Synthesis F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-05T16:01:52.541441Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:01:52.541441Z digest=sha256:a79fcddaed9223b436ae7381aa986193025f85441a9b423392ee50e01eb7406a

Observation 57fdbef0-4a32-4f96-8ff3-dfb48569668b · outbound

This paper cites High Fidelity Neural Audio Compression, October 2022.

CLEAR: Continuous Latent Autoregressive Modeling for High-quality and Low-latency Speech Synthesis High Fidelity Neural Audio Compression, October 2022

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:01:54.782035Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T16:01:52.547263Z digest=sha256:c71a00db251ecd0a2fe0ed929eb998445efc95b69572f233614d3b1df55b3535

Observation 14557648-9203-4b77-94aa-6dde161d30c6 · outbound

This paper cites CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens.

CLEAR: Continuous Latent Autoregressive Modeling for High-quality and Low-latency Speech Synthesis CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-05T16:01:52.553461Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:01:52.553461Z digest=sha256:86a96c8f4b661bae92523aacc27f24570bfe7cc0d19733d5503d4002875926b1

Observation 5adda8a6-57e9-4477-82d1-a3a07bf6a619 · outbound

This paper cites CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models.

CLEAR: Continuous Latent Autoregressive Modeling for High-quality and Low-latency Speech Synthesis CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-05T16:01:52.558566Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:01:52.558566Z digest=sha256:997183020fb4eb86ec79ea1dffc5d06fbc01d7a5bac4562d28de908bed8e0dcb

Observation b85f4d6d-3bfa-4a61-a653-8b689b6c30b2 · outbound

This paper cites E2 TTS: Embarrassingly Easy Fully Non-Autoregressive Zero-Shot TTS.

CLEAR: Continuous Latent Autoregressive Modeling for High-quality and Low-latency Speech Synthesis E2 TTS: Embarrassingly Easy Fully Non-Autoregressive Zero-Shot TTS

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-05T16:01:52.564223Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:01:52.564223Z digest=sha256:2f4058046449310d34468b474966a751cf86cfb184e1b4412dce4e714a9e0321

Observation 67b6aa19-bfba-4ecf-9d2b-7792931a7eb6 · outbound

This paper cites Taming Transformers for High-Resolution Image Synthesis.

CLEAR: Continuous Latent Autoregressive Modeling for High-quality and Low-latency Speech Synthesis Taming Transformers for High-Resolution Image Synthesis

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:01:54.766551Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T16:01:52.570589Z digest=sha256:c3cc44299a912ef519da45709cc88f598e9c8d33986243ab241f1d541f782498

Observation ae6eafc0-1ccf-4c0c-b577-b0739db50a6e · outbound

This paper cites Scaling Rectified Flow Transformers for High-Resolution Image Synthesis.

CLEAR: Continuous Latent Autoregressive Modeling for High-quality and Low-latency Speech Synthesis Scaling Rectified Flow Transformers for High-Resolution Image Synthesis

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-05T16:01:52.575279Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:01:52.575279Z digest=sha256:c816c3fa0d8092bd8bc2e0eb0612de50636bdc77647e8a13900aeef055142128

Observation fb7fc861-e8fa-4c79-a6fd-53f57b73239f · outbound

This paper cites Fast Timing-Conditioned Latent Audio Diffusion.

CLEAR: Continuous Latent Autoregressive Modeling for High-quality and Low-latency Speech Synthesis Fast Timing-Conditioned Latent Audio Diffusion

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-05T16:01:52.581625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:01:52.581625Z digest=sha256:a89fb9dd93179792e166c0864e25a21b0722b951f6ec06cfd7b03ba23f5a2774

Observation 44e1c1ad-4a1d-4d7f-8636-d23a5b4e962c · outbound

This paper cites E3 TTS: Easy End-to-End Diffusion-Based Text To Speech.

CLEAR: Continuous Latent Autoregressive Modeling for High-quality and Low-latency Speech Synthesis E3 TTS: Easy End-to-End Diffusion-Based Text To Speech

Reference 16

Resolution
malformed identifier
no resolver link, observed 2026-08-05T16:01:52.587056Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:01:52.587056Z digest=sha256:b9a96e540c4478c54e7bafe17efa2bc538c245290b36816ada6c710c65355846

Observation b0c4c34e-f724-4f59-97a1-ce13369336bd · outbound

This paper cites Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio.

CLEAR: Continuous Latent Autoregressive Modeling for High-quality and Low-latency Speech Synthesis Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:01:54.749600Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T16:01:52.592009Z digest=sha256:607b7067c5ca82ef909f45dd622a3040529dd85f8d39cbc6e59ec901b9db15d4

Observation 04207df1-0bfa-4b6d-9458-20207a08e057 · outbound

This paper cites FireRedTTS: A Foundation Text-To-Speech Framework for Industry-Level Generative Speech Applications.

CLEAR: Continuous Latent Autoregressive Modeling for High-quality and Low-latency Speech Synthesis FireRedTTS: A Foundation Text-To-Speech Framework for Industry-Level Generative Speech Applications

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-05T16:01:52.598217Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:01:52.598217Z digest=sha256:6afd6587a5c32a6c44eef4dfeca9dc0d39550177c10c0dadf2c6aea09d144d07

Observation bb6b0037-dbd3-4354-9c35-d3fcd46dd6fd · outbound

This paper cites VALL-E R: Robust and Efficient Zero-Shot Text-to-Speech Synthesis via Monotonic Alignment.

CLEAR: Continuous Latent Autoregressive Modeling for High-quality and Low-latency Speech Synthesis VALL-E R: Robust and Efficient Zero-Shot Text-to-Speech Synthesis via Monotonic Alignment

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-05T16:01:52.604143Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:01:52.604143Z digest=sha256:77c7f01e082efc07c0b335c2bae3075d88c1a112a9ea050440a5bd66bc1cb461

Observation e65ad875-0b89-4e36-a11c-da80d888d4e1 · outbound

This paper cites Classifier-Free Diffusion Guidance.

CLEAR: Continuous Latent Autoregressive Modeling for High-quality and Low-latency Speech Synthesis Classifier-Free Diffusion Guidance

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-05T16:01:52.611342Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:01:52.611342Z digest=sha256:0ad17b3ed2854436907a0adb4b9f994fa0f6e81ea57caa01eeb5c54812b8f0b1

Observation aef40a6e-b847-4b40-8e4c-da3022bfe9f7 · outbound

This paper cites Straightening out the straight- through estimator: Overcoming optimization challenges in vector quantized networks.

CLEAR: Continuous Latent Autoregressive Modeling for High-quality and Low-latency Speech Synthesis Straightening out the straight- through estimator: Overcoming optimization challenges in vector quantized networks

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:01:54.730366Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T16:01:52.617729Z digest=sha256:bab3f5ebb3662e53a6c8ee3fb605825ec91951eb52383a5973a890e73fc447e1

Observation 7833edb7-8922-4c02-aa15-c5cf3f7ac512 · outbound

This paper cites Ditar: Diffusion transformer au- toregressive modeling for speech generation, 2025.

CLEAR: Continuous Latent Autoregressive Modeling for High-quality and Low-latency Speech Synthesis Ditar: Diffusion transformer au- toregressive modeling for speech generation, 2025

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:01:54.711234Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T16:01:52.622999Z digest=sha256:d1770f3927093e2580447ca92820687087cefac5cbd5d1eea35d42d6e04fe222

Observation 2e5b129c-c9c6-4923-bcd8-3039bb2510ee · outbound

This paper cites Mega-TTS 2: Boosting Prompting Mechanisms for Zero-Shot Speech Synthesis, April 2024.

CLEAR: Continuous Latent Autoregressive Modeling for High-quality and Low-latency Speech Synthesis Mega-TTS 2: Boosting Prompting Mechanisms for Zero-Shot Speech Synthesis, April 2024

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:01:54.692761Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T16:01:52.628628Z digest=sha256:cc3e1c946318702832ef3143b007801567e785d6ea530fbc2f18c8548fbb0376

Observation 31454201-0a0c-44ca-837f-36efe3eea5f1 · outbound

This paper cites NaturalSpeech 3: Zero-Shot Speech Synthesis with Factorized Codec and Diffusion Models, April 2024.

CLEAR: Continuous Latent Autoregressive Modeling for High-quality and Low-latency Speech Synthesis NaturalSpeech 3: Zero-Shot Speech Synthesis with Factorized Codec and Diffusion Models, April 2024

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:01:54.672734Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T16:01:52.634749Z digest=sha256:6d9dc8844d3ebb2a7dd8430de7a3ef5bb993e16a48471eeaa70ff65766f59836

Observation 84602fb7-7bf0-4493-9b41-12650abc3872 · outbound

This paper cites Libriheavy: a 50,000 hours asr corpus with punctuation casing and context,.

CLEAR: Continuous Latent Autoregressive Modeling for High-quality and Low-latency Speech Synthesis Libriheavy: a 50,000 hours asr corpus with punctuation casing and context,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:01:54.654066Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T16:01:52.639373Z digest=sha256:941700e02f3e36b178fe523b23823b608ba94eda28a8707a4fdf5da4f60cd2ff

Observation fa7fcdce-99c5-42f5-a527-58038e72b5a1 · outbound

This paper cites Analyzing and Improving the Training Dynamics of Diffusion Models.

CLEAR: Continuous Latent Autoregressive Modeling for High-quality and Low-latency Speech Synthesis Analyzing and Improving the Training Dynamics of Diffusion Models

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:01:54.633340Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T16:01:52.650093Z digest=sha256:25c4e45b63376e3a30f9bf97d9fa8aa2ec8e0d47f817c61fa7434525b36f66d9

Observation bc03816e-8704-4a65-a3bf-6d7772f6b3be · outbound

This paper cites CLaM-TTS: Improving Neural Codec Language Model for Zero-Shot Text-to-Speech.

CLEAR: Continuous Latent Autoregressive Modeling for High-quality and Low-latency Speech Synthesis CLaM-TTS: Improving Neural Codec Language Model for Zero-Shot Text-to-Speech

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-05T16:01:52.654964Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:01:52.654964Z digest=sha256:5a21365d47640a6a778ee36b4748bf4a829692bda07f47c81b716fca8e4ac9b7

Observation 37e0749a-329f-4bde-968a-37b2e22944bd · outbound

This paper cites High-Fidelity Audio Compression with Improved RVQGAN.

CLEAR: Continuous Latent Autoregressive Modeling for High-quality and Low-latency Speech Synthesis High-Fidelity Audio Compression with Improved RVQGAN

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:01:54.610995Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T16:01:52.662364Z digest=sha256:d85debc662bdfb44e7159825e25eb92702942cacc0e8158d27eeb32958970d8a

Observation 5cfa107d-b223-498f-870b-eec2eb05cec5 · outbound

This paper cites BASE TTS: Lessons from building a billion- parameter Text-to-Speech model on 100K hours of data, February 2024.

CLEAR: Continuous Latent Autoregressive Modeling for High-quality and Low-latency Speech Synthesis BASE TTS: Lessons from building a billion- parameter Text-to-Speech model on 100K hours of data, February 2024

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:01:54.586357Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T16:01:52.668210Z digest=sha256:a97efbd6b8a1ed9b1357ae891841476d488859e210987934b3bd5a1045dba375

Observation d5de89ad-5244-4d9f-98e4-a6e6c7f4439d · outbound

This paper cites V oicebox: Text-guided multilin- gual universal speech generation at scale.

CLEAR: Continuous Latent Autoregressive Modeling for High-quality and Low-latency Speech Synthesis V oicebox: Text-guided multilin- gual universal speech generation at scale

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:01:54.567050Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T16:01:52.673697Z digest=sha256:0335309d56a6d4777d37891a9bded925a60d1116b42de8287321fc9a873f83f8

Observation eebed74a-9958-470e-95d0-23b257fd04be · outbound

This paper cites REPA-E: Unlocking V AE for End-to-End Tuning with Latent Diffusion Transformers, April 2025.

CLEAR: Continuous Latent Autoregressive Modeling for High-quality and Low-latency Speech Synthesis REPA-E: Unlocking V AE for End-to-End Tuning with Latent Diffusion Transformers, April 2025

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:01:54.549525Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T16:01:52.678990Z digest=sha256:0f665d28630a2c7348409e2ca3f6f12c19a2428aeafda04f6d1e69cac3b3cc67

Observation 91a25cc2-7bd8-42d5-8329-6c8aba46c0a8 · outbound

This paper cites Neural Speech Synthesis with Transformer Network.

CLEAR: Continuous Latent Autoregressive Modeling for High-quality and Low-latency Speech Synthesis Neural Speech Synthesis with Transformer Network

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-05T16:01:52.684377Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:01:52.684377Z digest=sha256:5f1d0bbff38d172d214edb3acaf356d17b586e8554efd4f2e7d5cc7e95342483

Observation 6e8c5c59-9594-4e39-b972-1cd22596e5c6 · outbound

This paper cites Autoregressive Image Generation without Vector Quantization.

CLEAR: Continuous Latent Autoregressive Modeling for High-quality and Low-latency Speech Synthesis Autoregressive Image Generation without Vector Quantization

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-05T16:01:52.689870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:01:52.689870Z digest=sha256:4ff1c2502b4c71fe1ff7c07e417ec9b09c55eee5213c0686ef695a93a6dd7447

Observation 54dc638d-e939-4cb4-8af7-14894dfa04f4 · outbound

This paper cites Flow Straight and Fast: Learning to Generate and Transfer Data with Rectified Flow.

CLEAR: Continuous Latent Autoregressive Modeling for High-quality and Low-latency Speech Synthesis Flow Straight and Fast: Learning to Generate and Transfer Data with Rectified Flow

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-05T16:01:52.695773Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:01:52.695773Z digest=sha256:afbec2d83e9e1bb35e0ce0f58cbd95e7b8ffaff069a17cca02f9e1d6c2ac2fa2

Observation 9fbfafa0-f539-453a-95f2-15b3161ffe4c · outbound

This paper cites Autoregressive Diffusion Transformer for Text-to-Speech Synthesis.

CLEAR: Continuous Latent Autoregressive Modeling for High-quality and Low-latency Speech Synthesis Autoregressive Diffusion Transformer for Text-to-Speech Synthesis

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-05T16:01:52.701539Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:01:52.701539Z digest=sha256:23032c57822f1cdd9188dbbce80ac920691bdff35d082e80de47c8d54e2d98b2

Observation 439fbd8b-bb4f-46e5-95be-fc4997f507ad · outbound

This paper cites LibriSpeech-PC: Benchmark for Evaluation of Punctuation and Capitalization Capabilities of End-to-End ASR Models.

CLEAR: Continuous Latent Autoregressive Modeling for High-quality and Low-latency Speech Synthesis LibriSpeech-PC: Benchmark for Evaluation of Punctuation and Capitalization Capabilities of End-to-End ASR Models

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:01:54.516431Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T16:01:52.707422Z digest=sha256:1ad4eb94572b9b8f095832c76b2127c874f2b3736372d9836781fd6eebf2d7ab

Observation 95e71bba-b9f8-4aa8-9666-b8f109020006 · outbound

This paper cites Autoregressive Speech Synthesis without Vector Quantization.

CLEAR: Continuous Latent Autoregressive Modeling for High-quality and Low-latency Speech Synthesis Autoregressive Speech Synthesis without Vector Quantization

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-05T16:01:52.712652Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:01:52.712652Z digest=sha256:131e18b04818fd8a6ee7d6ac883e116de4ff540a9bc3e999c5323c8600c8b60f

Observation 40e3b997-26db-4c2a-bd81-b5ea98b4355a · outbound

This paper cites Finite Scalar Quantization: VQ-V AE Made Simple, October 2023.

CLEAR: Continuous Latent Autoregressive Modeling for High-quality and Low-latency Speech Synthesis Finite Scalar Quantization: VQ-V AE Made Simple, October 2023

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:01:54.492317Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T16:01:52.717776Z digest=sha256:b06acbeff9b8afb676d55c8751963ed766245ac5bf7df2bd7a020aefb0a57a29

Observation ca2f0a00-cd01-453c-ab18-6e562a1fbf90 · outbound

This paper cites HALL-E: Hierarchical Neural Codec Language Model for Minute-Long Zero-Shot Text-to-Speech Synthesis.

CLEAR: Continuous Latent Autoregressive Modeling for High-quality and Low-latency Speech Synthesis HALL-E: Hierarchical Neural Codec Language Model for Minute-Long Zero-Shot Text-to-Speech Synthesis

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-05T16:01:52.722764Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:01:52.722764Z digest=sha256:3f9ed0d5788a441eb6e56eeea9dced399f48d0e53ac12f529c28077666f42ba1

Observation 1c9ac590-e942-4621-8e3c-34b488390f90 · outbound

This paper cites Librispeech: An ASR corpus based on public domain audio books.

CLEAR: Continuous Latent Autoregressive Modeling for High-quality and Low-latency Speech Synthesis Librispeech: An ASR corpus based on public domain audio books

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:01:54.470831Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T16:01:52.728945Z digest=sha256:265666ac09cec9b48ce79ba48aef2dd919a0678acad7a0802e66552ca7514c5a

Observation 25b5268c-6ecf-4fe3-9b77-07cb518ff5bd · outbound

This paper cites Robust speech recognition via large-scale weak supervision.

CLEAR: Continuous Latent Autoregressive Modeling for High-quality and Low-latency Speech Synthesis Robust speech recognition via large-scale weak supervision

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-05T16:01:52.736142Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:01:52.736142Z digest=sha256:4b49c5f81226fbf6530d1d34cd68f03a9c0b3b1d22ca25164010337120a84713

Observation 5e2ca104-69e1-476a-8e25-097231b26321 · outbound

This paper cites Revisiting over-smoothness in text to speech.

CLEAR: Continuous Latent Autoregressive Modeling for High-quality and Low-latency Speech Synthesis Revisiting over-smoothness in text to speech

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:01:54.435875Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T16:01:52.743452Z digest=sha256:06ebf8174c5dab526577fce56e865b7e5c7d6d88522a3697e0064284ceadd83d

Observation 7d78640a-6f08-41fe-8d3a-5d48bcd6aad6 · outbound

This paper cites UTMOS: UTokyo-SaruLab System for VoiceMOS Challenge 2022.

CLEAR: Continuous Latent Autoregressive Modeling for High-quality and Low-latency Speech Synthesis UTMOS: UTokyo-SaruLab System for VoiceMOS Challenge 2022

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-05T16:01:52.748760Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:01:52.748760Z digest=sha256:888127fd43bec7ba2e29a40138e8f5700c7c9ad54c5f428bda5146ca028229c7

Observation 7183e728-81f3-4342-807a-9d614121e605 · outbound

This paper cites AutoClip: Adaptive Gradient Clipping for Source Separation Networks, July 2020.

CLEAR: Continuous Latent Autoregressive Modeling for High-quality and Low-latency Speech Synthesis AutoClip: Adaptive Gradient Clipping for Source Separation Networks, July 2020

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:01:54.418343Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T16:01:52.756183Z digest=sha256:491c129166ebe453fde944b59aee0c67c44cc6225ebd5fec2179e3d395a5d9c7

Observation f5943c9b-07cd-4172-98a2-2b6298578e77 · outbound

This paper cites GLU Variants Improve Transformer.

CLEAR: Continuous Latent Autoregressive Modeling for High-quality and Low-latency Speech Synthesis GLU Variants Improve Transformer

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-05T16:01:52.761504Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:01:52.761504Z digest=sha256:b2930d2a3d2c28d8180bb541b4f20b1acb28a57eafd75ea36d88a64f3cfeed81

Observation 3c67adb7-6ae8-4540-9c83-0d94bb0ab2bd · outbound

This paper cites NaturalSpeech 2: Latent Diffusion Models are Natural and Zero-Shot Speech and Singing Synthesizers, May 2023.

CLEAR: Continuous Latent Autoregressive Modeling for High-quality and Low-latency Speech Synthesis NaturalSpeech 2: Latent Diffusion Models are Natural and Zero-Shot Speech and Singing Synthesizers, May 2023

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:01:54.396325Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T16:01:52.767480Z digest=sha256:7d54a60b623794898d824ec62546837431d05c0ec320566049aaeaf764e03067

Observation aeb04ce9-c2aa-4a98-9784-484353023d73 · outbound

This paper cites Real-time single image and video super-resolution using an efficient sub-pixel convolutional neural network.

CLEAR: Continuous Latent Autoregressive Modeling for High-quality and Low-latency Speech Synthesis Real-time single image and video super-resolution using an efficient sub-pixel convolutional neural network

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-05T16:01:52.774022Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:01:52.774022Z digest=sha256:d9bea8e42ff813e557d11ca129e44cb91f53bfe283b63adb84f16b94629938da

Observation f7350a1b-82de-4554-bdc7-c4b14864f0d5 · outbound

This paper cites Ella-v: Stable neural codec language modeling with alignment-guided sequence reordering.

CLEAR: Continuous Latent Autoregressive Modeling for High-quality and Low-latency Speech Synthesis Ella-v: Stable neural codec language modeling with alignment-guided sequence reordering

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:01:54.359518Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T16:01:52.778679Z digest=sha256:095f221744af2770d0684717df938595e5a287ab59812acc814a99774b4c4309

Observation 8d0864e8-9d61-4847-b4ae-8007d02a4b4a · outbound

This paper cites Steinmetz, Jordi Pons, Santiago Pascual, and Joan Serrà.

CLEAR: Continuous Latent Autoregressive Modeling for High-quality and Low-latency Speech Synthesis Steinmetz, Jordi Pons, Santiago Pascual, and Joan Serrà

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:01:54.339544Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T16:01:52.785227Z digest=sha256:5de1a0b8b2a257ee850727b54774693b7e2829d6d77dc8aa16d33da798186ef7

Observation 88ac9b82-424e-4284-90e6-7e3474890267 · outbound

This paper cites RoFormer: Enhanced Transformer with Rotary Position Embedding.

CLEAR: Continuous Latent Autoregressive Modeling for High-quality and Low-latency Speech Synthesis RoFormer: Enhanced Transformer with Rotary Position Embedding

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-05T16:01:52.792348Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:01:52.792348Z digest=sha256:60152b6dd7aa65581f844c8d8844adae7ed92b6ca3bc9f50bfdc605b915d1a10

Observation 7e22d374-289f-4f3a-9e3a-fd63b7c1ac92 · outbound

This paper cites NaturalSpeech: End-to-End Text-to-Speech Synthesis With Human-Level Quality.

CLEAR: Continuous Latent Autoregressive Modeling for High-quality and Low-latency Speech Synthesis NaturalSpeech: End-to-End Text-to-Speech Synthesis With Human-Level Quality

Reference 51

Resolution
metadata mismatch
raw_fallback, observed 2026-08-05T16:01:53.540373Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T16:01:52.797356Z digest=sha256:1dbc110b4acc21cff226b29b803be05db772a83a75a37c1523b6b1dd504a6f3c

Observation 1d9c93ff-5827-44e5-a4de-57432d2981a0 · outbound

This paper cites Continuous Speech Synthesis using per-token Latent Diffusion, October 2024.

CLEAR: Continuous Latent Autoregressive Modeling for High-quality and Low-latency Speech Synthesis Continuous Speech Synthesis using per-token Latent Diffusion, October 2024

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:01:54.309528Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T16:01:52.802678Z digest=sha256:9365d7f16da08f6415f19c7236937c9983f31fe09a34535e168d461bb8251a24

Observation 6e0e7da8-4df0-4875-9f8c-592ba1c12b98 · outbound

This paper cites Neural Discrete Representation Learning.

CLEAR: Continuous Latent Autoregressive Modeling for High-quality and Low-latency Speech Synthesis Neural Discrete Representation Learning

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:01:54.284600Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T16:01:52.808377Z digest=sha256:680b912818105fc407252ac633ae5e11a076c00747e9bcbd8e63b82a78d7cdf2

Observation e7a870c0-8ae9-4cb2-b7d2-6361397113fe · outbound

This paper cites Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers.

CLEAR: Continuous Latent Autoregressive Modeling for High-quality and Low-latency Speech Synthesis Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-05T16:01:52.813060Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:01:52.813060Z digest=sha256:bf1b853c60d27e6cba96d888b12979b8b535b6d1eebfaa0512646073a75a02c4

Observation c5b0d8c7-9cc3-4359-8b41-ce886d522b74 · outbound

This paper cites FELLE: Autoregressive Speech Synthesis with Token-Wise Coarse-to-Fine Flow Matching.

CLEAR: Continuous Latent Autoregressive Modeling for High-quality and Low-latency Speech Synthesis FELLE: Autoregressive Speech Synthesis with Token-Wise Coarse-to-Fine Flow Matching

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-05T16:01:52.818331Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:01:52.818331Z digest=sha256:5f564ef8f8a2f89035ef9ad8f12ee296777206ee599ee34d0857eed0b93fd8d1

Observation 61f5c358-f70e-48f2-9779-a8badee88a3c · outbound

This paper cites Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens.

CLEAR: Continuous Latent Autoregressive Modeling for High-quality and Low-latency Speech Synthesis Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-05T16:01:52.823558Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:01:52.823558Z digest=sha256:f3eee3726941c43948d0a516a192350e884d23d8235d8f69c93481a0afe0dabf

Observation 1d998edd-d252-4ed5-a46d-6df9e126adaa · outbound

This paper cites MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer, October 2024.

CLEAR: Continuous Latent Autoregressive Modeling for High-quality and Low-latency Speech Synthesis MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer, October 2024

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:01:54.263650Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T16:01:52.829746Z digest=sha256:0d9f9be2bc2e26fcbbdee882da9c2c3f8539353a31465e8245020e23d6b485fc

Observation 536b2ec0-62b9-481b-b8ee-ab4ccc1042b9 · outbound

This paper cites Towards audio language modeling -- an overview.

CLEAR: Continuous Latent Autoregressive Modeling for High-quality and Low-latency Speech Synthesis Towards audio language modeling -- an overview

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-05T16:01:52.834220Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:01:52.834220Z digest=sha256:44b3cbbb49d72e4325944b7a71545c100acf290e247d6bf16aead577c3fbc722

Observation 981dec4e-3a18-4f0b-a20a-d12e616454d4 · outbound

This paper cites RALL-E: Robust Codec Language Modeling with Chain-of-Thought Prompting for Text-to-Speech Synthesis.

CLEAR: Continuous Latent Autoregressive Modeling for High-quality and Low-latency Speech Synthesis RALL-E: Robust Codec Language Modeling with Chain-of-Thought Prompting for Text-to-Speech Synthesis

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-05T16:01:52.839342Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:01:52.839342Z digest=sha256:bd532d410a531062f8cb8cea11cbb1aeb33505e1f77a5eda2326a82b57cf56e0

Observation fa13fcd6-7014-4990-9911-08281a879493 · outbound

This paper cites On Layer Normalization in the Transformer Architecture.

CLEAR: Continuous Latent Autoregressive Modeling for High-quality and Low-latency Speech Synthesis On Layer Normalization in the Transformer Architecture

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-05T16:01:52.844345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:01:52.844345Z digest=sha256:6fcc6e7273c911920ecd687156663c244c8656264c518c108c0c4d61cfbce0ee

Observation 02ee09ec-e051-4671-805d-6f91e1e6141b · outbound

This paper cites Reconstruction vs.

CLEAR: Continuous Latent Autoregressive Modeling for High-quality and Low-latency Speech Synthesis Reconstruction vs

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:01:54.245081Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T16:01:52.849095Z digest=sha256:ba86280ec19e60608a61123af62dee3acf3edb37acb1616ff50858eb767762f9

Observation cc1ff55e-0c36-4212-9fcd-d9cfb7916350 · outbound

This paper cites Llasa: Scaling Train-Time and Inference-Time Compute for Llama-based Speech Synthesis.

CLEAR: Continuous Latent Autoregressive Modeling for High-quality and Low-latency Speech Synthesis Llasa: Scaling Train-Time and Inference-Time Compute for Llama-based Speech Synthesis

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-05T16:01:52.854277Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:01:52.854277Z digest=sha256:c80879617f4ccbff07273aac0801c0dc55dc6c0706ee3c4c39d17a2cddc56756

Observation ffefd789-2cff-4087-abec-ee7e034b4dd6 · outbound

This paper cites Representation Alignment for Generation: Training Diffusion Transformers Is Easier Than You Think, February 2025.

CLEAR: Continuous Latent Autoregressive Modeling for High-quality and Low-latency Speech Synthesis Representation Alignment for Generation: Training Diffusion Transformers Is Easier Than You Think, February 2025

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:01:54.219454Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T16:01:52.859410Z digest=sha256:3904c44c9f07f2ad8fcbb341620ac9c350eaf281831f2164876e2158c9bc64bb

Observation 0fb202ed-c747-45d3-918b-8fe81290ee9f · outbound

This paper cites Soundstream: An end-to-end neural audio codec.

CLEAR: Continuous Latent Autoregressive Modeling for High-quality and Low-latency Speech Synthesis Soundstream: An end-to-end neural audio codec

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-05T16:01:52.864521Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:01:52.864521Z digest=sha256:2d3adbb7ca27e5ccdd9667905d51840127b5800320ab3b40c8bedb1173068428

Observation e9ac53ff-9952-42b0-9262-3634970b2ed3 · outbound

This paper cites Weiss, Ye Jia, Zhifeng Chen, and Yonghui Wu.

CLEAR: Continuous Latent Autoregressive Modeling for High-quality and Low-latency Speech Synthesis Weiss, Ye Jia, Zhifeng Chen, and Yonghui Wu

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:01:54.188022Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T16:01:52.871677Z digest=sha256:766eb665adbc1e375f6a59c51fc928b31fc0b7c2aa302f53826976441cc188be

Observation 449586ca-366b-4904-8afd-99a0efc17c80 · outbound

This paper cites Root Mean Square Layer Normalization.

CLEAR: Continuous Latent Autoregressive Modeling for High-quality and Low-latency Speech Synthesis Root Mean Square Layer Normalization

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-05T16:01:52.877438Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:01:52.877438Z digest=sha256:ace12fb5b5c1d4270d217d0f7438d6d9f6e04e54cf854aabc85b9b0a762bdf57

Observation d9c9843c-e0b0-4d37-9727-f0540ccd7d32 · outbound

This paper cites Speak Foreign Languages with Your Own Voice: Cross-Lingual Neural Codec Language Modeling.

CLEAR: Continuous Latent Autoregressive Modeling for High-quality and Low-latency Speech Synthesis Speak Foreign Languages with Your Own Voice: Cross-Lingual Neural Codec Language Modeling

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-05T16:01:52.883671Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:01:52.883671Z digest=sha256:1e5d2b33fe9c7aaacf73f7947672a3858b9dfd17cc6b4e4f8d18009f3f22bf5d

Observation 3c846336-a41f-4239-85a9-dc8976d02ace · outbound

This paper cites ,𝑁𝑆𝐵,𝐶!,𝑁 + UpsamplingBlock Channel toSpace ChannelDuplicating𝐵,𝐶.

CLEAR: Continuous Latent Autoregressive Modeling for High-quality and Low-latency Speech Synthesis ,𝑁𝑆𝐵,𝐶!,𝑁 + UpsamplingBlock Channel toSpace ChannelDuplicating𝐵,𝐶

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:01:54.168792Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-05T16:01:52.889397Z digest=sha256:98bb9d9dfce4699dfed6fff666caee229ddbf9d793a1455955bd45c638b71404

Observation c35e1bbd-d098-48b1-b4de-3d045c067d50 · outbound

This paper cites Libriheavy: a 50,000 hours ASR corpus with punctuation casing and context.

CLEAR: Continuous Latent Autoregressive Modeling for High-quality and Low-latency Speech Synthesis Libriheavy: a 50,000 hours ASR corpus with punctuation casing and context

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-05T16:01:52.645000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:01:52.645000Z digest=sha256:800960dc323eb51f444f466082d07a21f0a742ab2844eeb8e75ee2d06b8a14e4

Pith citing papers

Observation f877a92c-605b-4346-bb9a-935b28c1d987 · inbound

On the Distillation Loss Functions of Speech VAE for Unified Reconstruction, Understanding, and Generation cites this paper.

On the Distillation Loss Functions of Speech VAE for Unified Reconstruction, Understanding, and Generation CLEAR: Continuous Latent Autoregressive Modeling for High-quality and Low-latency Speech Synthesis

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-11T10:25:59.528620Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T15:30:58.664375Z digest=sha256:97e79a5bad2707466abdba551603a31d88e5ee94f403240b07a03ebdb9f3db3e

Observation c6001f44-543d-4894-8ba1-62ff1c34b900 · inbound

SemaVoice: Semantic-Aware Continuous Autoregressive Speech Synthesis cites this paper.

SemaVoice: Semantic-Aware Continuous Autoregressive Speech Synthesis CLEAR: Continuous Latent Autoregressive Modeling for High-quality and Low-latency Speech Synthesis

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-19T19:02:43.569018Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-19T18:58:15.288299Z digest=sha256:efafddc4fcd66720307fd132dcb1eb31c3e6b28880be501560bc7f9aa8a4556c

Observation 6947956a-a27e-4039-9221-c3793b937b2b · inbound

VoxCPM2 Technical Report cites this paper.

VoxCPM2 Technical Report CLEAR: Continuous Latent Autoregressive Modeling for High-quality and Low-latency Speech Synthesis

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-07-02T19:47:19.714861Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-27T21:18:22.911332Z digest=sha256:13355f6dbf3987736da4b183c6ff36a7b8ea1a64c684fdccacdd06200490ee70