Pith. sign in

Paper Citation Record · LEDGER

CLEAR: Continuous Latent Autoregressive Modeling for High-quality and Low-latency Speech Synthesis

As of 8 August 2026, this Paper Citation Record lists 69 of 69 outbound references and 3 inbound Pith citation observations for arXiv:2508.19098.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.19098 v1

Coverage vector

measured 69 of 69 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T16:01:52.889397Z

measured 72 of 72 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-27T21:18:22.911332Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T19:47:19.713205Z

Reference resolution

69 of 69 outbound references displayed

  • verified exact0
  • verified fuzzy33
  • unresolved34
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 958bcb74-b0d2-4b6d-b750-76a12a654377 · outbound

This paper cites Seed-TTS: A Family of High-Quality Versatile Speech Generation Models.

CLEAR: Continuous Latent Autoregressive Modeling for High-quality and Low-latency Speech Synthesis Seed-TTS: A Family of High-Quality Versatile Speech Generation Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-05T16:01:52.500071Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:01:52.500071Z digest=sha256:a5f5256b5a21624b784c8f800aa872767787e04a4aa639a71ade7feea487ab8b

Observation ce96466a-3f3c-498a-b4c6-699d1b217b7a · outbound

This paper cites SpeechT5: Unified-Modal Encoder-Decoder Pre-Training for Spoken Language Processing, May 2022.

CLEAR: Continuous Latent Autoregressive Modeling for High-quality and Low-latency Speech Synthesis SpeechT5: Unified-Modal Encoder-Decoder Pre-Training for Spoken Language Processing, May 2022

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:01:54.886110Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T16:01:52.506197Z digest=sha256:cfdd92f78d373943b20fd85000458850effee52ed1adb896a2e242a71ca69447

Observation ee274c21-6231-4378-b2c4-d7d16bd2c174 · outbound

This paper cites Rethinking lossy compression: The rate-distortion-perception tradeoff.

CLEAR: Continuous Latent Autoregressive Modeling for High-quality and Low-latency Speech Synthesis Rethinking lossy compression: The rate-distortion-perception tradeoff

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:01:54.863775Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T16:01:52.511486Z digest=sha256:fcea0eb241679f42a10b1d286ce021000e2e5a84cec9f1a476addfa3785d08c2

Observation 9568581b-d05c-427d-bcda-b47528a0d1a1 · outbound

This paper cites Audiolm: a language modeling approach to audio generation.

CLEAR: Continuous Latent Autoregressive Modeling for High-quality and Low-latency Speech Synthesis Audiolm: a language modeling approach to audio generation

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:01:54.843458Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T16:01:52.517582Z digest=sha256:68da0fded67b5d850f4eedbe6dd5f5241083d6e7c52726ac1ff33362dfc04491

Observation 90aa4c5a-2bd8-46eb-8908-16a6a917b2c7 · outbound

This paper cites Deep Compression Autoencoder for Efficient High-Resolution Diffusion Models, April 2025.

CLEAR: Continuous Latent Autoregressive Modeling for High-quality and Low-latency Speech Synthesis Deep Compression Autoencoder for Efficient High-Resolution Diffusion Models, April 2025

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:01:54.823149Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T16:01:52.523257Z digest=sha256:d683332154ff8c39d594e72820f749d1602fb4313d8b755da1175c07681282ba

Observation 19a81d71-38a8-4e58-8a48-58b6b09987ee · outbound

This paper cites Wavlm: Large-scale self-supervised pre- training for full stack speech processing.

CLEAR: Continuous Latent Autoregressive Modeling for High-quality and Low-latency Speech Synthesis Wavlm: Large-scale self-supervised pre- training for full stack speech processing

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:01:54.802618Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T16:01:52.528372Z digest=sha256:716c553a7e25bf4837fc9a09a84619102536b4d60134c34d9277eadfc1f92330

Observation 53ce4747-da48-485a-b0f5-1a133c8439b7 · outbound

This paper cites VALL-E 2: Neural Codec Language Models are Human Parity Zero-Shot Text to Speech Synthesizers.

CLEAR: Continuous Latent Autoregressive Modeling for High-quality and Low-latency Speech Synthesis VALL-E 2: Neural Codec Language Models are Human Parity Zero-Shot Text to Speech Synthesizers

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-05T16:01:52.534536Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:01:52.534536Z digest=sha256:40dbf83fcac2b753fa876e901b26f58b177398995de92432a7638c3cb978dd47

Observation 745110e8-993f-4339-8a4b-51e03547cfbb · outbound

This paper cites F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching.

CLEAR: Continuous Latent Autoregressive Modeling for High-quality and Low-latency Speech Synthesis F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-05T16:01:52.541441Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:01:52.541441Z digest=sha256:1249ac878bf90a9a344351a8de76718ba6a2248011e19adb34b475ea73cb4419

Observation 57fdbef0-4a32-4f96-8ff3-dfb48569668b · outbound

This paper cites High Fidelity Neural Audio Compression, October 2022.

CLEAR: Continuous Latent Autoregressive Modeling for High-quality and Low-latency Speech Synthesis High Fidelity Neural Audio Compression, October 2022

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:01:54.782035Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T16:01:52.547263Z digest=sha256:fc1f80b8cbd82805064b537bdd754a0466689250a5bfafb48c155a533ff833f3

Observation 14557648-9203-4b77-94aa-6dde161d30c6 · outbound

This paper cites CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens.

CLEAR: Continuous Latent Autoregressive Modeling for High-quality and Low-latency Speech Synthesis CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-05T16:01:52.553461Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:01:52.553461Z digest=sha256:a18acfa17ff356259c677474350d46c733e4d9dca69f96260c22a023fa96e8c1

Observation 5adda8a6-57e9-4477-82d1-a3a07bf6a619 · outbound

This paper cites CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models.

CLEAR: Continuous Latent Autoregressive Modeling for High-quality and Low-latency Speech Synthesis CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-05T16:01:52.558566Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:01:52.558566Z digest=sha256:2a33bdbf5e3e6e311e3b022b17e8a178c7e5040a3df0bba2203cbb965fe457b3

Observation b85f4d6d-3bfa-4a61-a653-8b689b6c30b2 · outbound

This paper cites E2 TTS: Embarrassingly Easy Fully Non-Autoregressive Zero-Shot TTS.

CLEAR: Continuous Latent Autoregressive Modeling for High-quality and Low-latency Speech Synthesis E2 TTS: Embarrassingly Easy Fully Non-Autoregressive Zero-Shot TTS

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-05T16:01:52.564223Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:01:52.564223Z digest=sha256:bdb1cb46007555bbd2b8b01f0201b28d48d57be8b754b2346ffc1fa7e219f102

Observation 67b6aa19-bfba-4ecf-9d2b-7792931a7eb6 · outbound

This paper cites Taming Transformers for High-Resolution Image Synthesis.

CLEAR: Continuous Latent Autoregressive Modeling for High-quality and Low-latency Speech Synthesis Taming Transformers for High-Resolution Image Synthesis

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:01:54.766551Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T16:01:52.570589Z digest=sha256:f4ee1872c00a2b2f942cdcf7cef2f2e394dac5f7595f20ce0391ab26b0042357

Observation ae6eafc0-1ccf-4c0c-b577-b0739db50a6e · outbound

This paper cites Scaling Rectified Flow Transformers for High-Resolution Image Synthesis.

CLEAR: Continuous Latent Autoregressive Modeling for High-quality and Low-latency Speech Synthesis Scaling Rectified Flow Transformers for High-Resolution Image Synthesis

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-05T16:01:52.575279Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:01:52.575279Z digest=sha256:8be13b86e5b9ac1214d2fee43f2599ea1437076206c6f62600bc558f4f5148e8

Observation fb7fc861-e8fa-4c79-a6fd-53f57b73239f · outbound

This paper cites Fast Timing-Conditioned Latent Audio Diffusion.

CLEAR: Continuous Latent Autoregressive Modeling for High-quality and Low-latency Speech Synthesis Fast Timing-Conditioned Latent Audio Diffusion

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-05T16:01:52.581625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:01:52.581625Z digest=sha256:a7d3838e88bf7ad7164ccd38b1fd18673d7c7c631c6bafec9b5ff83435a52e70

Observation 44e1c1ad-4a1d-4d7f-8636-d23a5b4e962c · outbound

This paper cites E3 TTS: Easy End-to-End Diffusion-Based Text To Speech.

CLEAR: Continuous Latent Autoregressive Modeling for High-quality and Low-latency Speech Synthesis E3 TTS: Easy End-to-End Diffusion-Based Text To Speech

Reference 16

Resolution
malformed identifier
no resolver link, observed 2026-08-05T16:01:52.587056Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:01:52.587056Z digest=sha256:5a0558ed219e3ae058bf14100e1930ee1f8f819d56c5b4f6dda08d674bcdc394

Observation b0c4c34e-f724-4f59-97a1-ce13369336bd · outbound

This paper cites Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio.

CLEAR: Continuous Latent Autoregressive Modeling for High-quality and Low-latency Speech Synthesis Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:01:54.749600Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T16:01:52.592009Z digest=sha256:bb56d10e8d9dacaff67f56b10b667511ab43a4b9679e3168940e50411eaaeca1

Observation 04207df1-0bfa-4b6d-9458-20207a08e057 · outbound

This paper cites FireRedTTS: A Foundation Text-To-Speech Framework for Industry-Level Generative Speech Applications.

CLEAR: Continuous Latent Autoregressive Modeling for High-quality and Low-latency Speech Synthesis FireRedTTS: A Foundation Text-To-Speech Framework for Industry-Level Generative Speech Applications

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-05T16:01:52.598217Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:01:52.598217Z digest=sha256:319f1497226caf3e5e46d9acda51df846593eb4a6a93c6bf65f606b0ac3ead9f

Observation bb6b0037-dbd3-4354-9c35-d3fcd46dd6fd · outbound

This paper cites VALL-E R: Robust and Efficient Zero-Shot Text-to-Speech Synthesis via Monotonic Alignment.

CLEAR: Continuous Latent Autoregressive Modeling for High-quality and Low-latency Speech Synthesis VALL-E R: Robust and Efficient Zero-Shot Text-to-Speech Synthesis via Monotonic Alignment

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-05T16:01:52.604143Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:01:52.604143Z digest=sha256:99efb5b4d7c5444f14e97b5cb27ce3bae145b864a9d0da9eb12b31d83e6ad76c

Observation e65ad875-0b89-4e36-a11c-da80d888d4e1 · outbound

This paper cites Classifier-Free Diffusion Guidance.

CLEAR: Continuous Latent Autoregressive Modeling for High-quality and Low-latency Speech Synthesis Classifier-Free Diffusion Guidance

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-05T16:01:52.611342Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:01:52.611342Z digest=sha256:8fb57f130dfab6bc5f2887724b2c2468ea56765bf99473e82fc7a655ab22bc5b

Observation aef40a6e-b847-4b40-8e4c-da3022bfe9f7 · outbound

This paper cites Straightening out the straight- through estimator: Overcoming optimization challenges in vector quantized networks.

CLEAR: Continuous Latent Autoregressive Modeling for High-quality and Low-latency Speech Synthesis Straightening out the straight- through estimator: Overcoming optimization challenges in vector quantized networks

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:01:54.730366Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T16:01:52.617729Z digest=sha256:a6ac48b0ac59b19befe8427da6365e5884a8c43b2bedb106afeaeb635425b4a1

Observation 7833edb7-8922-4c02-aa15-c5cf3f7ac512 · outbound

This paper cites Ditar: Diffusion transformer au- toregressive modeling for speech generation, 2025.

CLEAR: Continuous Latent Autoregressive Modeling for High-quality and Low-latency Speech Synthesis Ditar: Diffusion transformer au- toregressive modeling for speech generation, 2025

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:01:54.711234Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T16:01:52.622999Z digest=sha256:2dcaf84bece8ca329f18c7717cdc0a1bcbc63d2f679bfb8c260dd116dc040b7a

Observation 2e5b129c-c9c6-4923-bcd8-3039bb2510ee · outbound

This paper cites Mega-TTS 2: Boosting Prompting Mechanisms for Zero-Shot Speech Synthesis, April 2024.

CLEAR: Continuous Latent Autoregressive Modeling for High-quality and Low-latency Speech Synthesis Mega-TTS 2: Boosting Prompting Mechanisms for Zero-Shot Speech Synthesis, April 2024

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:01:54.692761Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T16:01:52.628628Z digest=sha256:27230325eac3eaabc9fa28852188f9800dc8ac4c1f09e0014216dc52d9ad7ff3

Observation 31454201-0a0c-44ca-837f-36efe3eea5f1 · outbound

This paper cites NaturalSpeech 3: Zero-Shot Speech Synthesis with Factorized Codec and Diffusion Models, April 2024.

CLEAR: Continuous Latent Autoregressive Modeling for High-quality and Low-latency Speech Synthesis NaturalSpeech 3: Zero-Shot Speech Synthesis with Factorized Codec and Diffusion Models, April 2024

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:01:54.672734Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T16:01:52.634749Z digest=sha256:4e4173bc1faeb4b4d1d437a8ef42262ecd93d2921befa7d8624704049a843dc7

Observation 84602fb7-7bf0-4493-9b41-12650abc3872 · outbound

This paper cites Libriheavy: a 50,000 hours asr corpus with punctuation casing and context,.

CLEAR: Continuous Latent Autoregressive Modeling for High-quality and Low-latency Speech Synthesis Libriheavy: a 50,000 hours asr corpus with punctuation casing and context,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:01:54.654066Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T16:01:52.639373Z digest=sha256:493204985240e70bd16feeed6b4cfd3143ba9c29208ba9e1971d363f965e1894

Observation fa7fcdce-99c5-42f5-a527-58038e72b5a1 · outbound

This paper cites Analyzing and Improving the Training Dynamics of Diffusion Models.

CLEAR: Continuous Latent Autoregressive Modeling for High-quality and Low-latency Speech Synthesis Analyzing and Improving the Training Dynamics of Diffusion Models

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:01:54.633340Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T16:01:52.650093Z digest=sha256:7a9d113fb35039c41a10b24e83b60d6f4a7c0958e7eb5318b0bbb6b9463f2382

Observation bc03816e-8704-4a65-a3bf-6d7772f6b3be · outbound

This paper cites CLaM-TTS: Improving Neural Codec Language Model for Zero-Shot Text-to-Speech.

CLEAR: Continuous Latent Autoregressive Modeling for High-quality and Low-latency Speech Synthesis CLaM-TTS: Improving Neural Codec Language Model for Zero-Shot Text-to-Speech

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-05T16:01:52.654964Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:01:52.654964Z digest=sha256:87ef036b304c2ad8c6f59d5d89fc0eebb3f542ec70e1950e65dfcb52cde176c1

Observation 37e0749a-329f-4bde-968a-37b2e22944bd · outbound

This paper cites High-Fidelity Audio Compression with Improved RVQGAN.

CLEAR: Continuous Latent Autoregressive Modeling for High-quality and Low-latency Speech Synthesis High-Fidelity Audio Compression with Improved RVQGAN

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:01:54.610995Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T16:01:52.662364Z digest=sha256:64b81fb5975c565de172f3b5b69536b5f7ecc4fbbee7d4c9d0e153728850ee53

Observation 5cfa107d-b223-498f-870b-eec2eb05cec5 · outbound

This paper cites BASE TTS: Lessons from building a billion- parameter Text-to-Speech model on 100K hours of data, February 2024.

CLEAR: Continuous Latent Autoregressive Modeling for High-quality and Low-latency Speech Synthesis BASE TTS: Lessons from building a billion- parameter Text-to-Speech model on 100K hours of data, February 2024

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:01:54.586357Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T16:01:52.668210Z digest=sha256:026bffa2494d2b07165c6bc41bdf2f5453b698096943c41d9a4af9c6cd98ee7e

Observation d5de89ad-5244-4d9f-98e4-a6e6c7f4439d · outbound

This paper cites V oicebox: Text-guided multilin- gual universal speech generation at scale.

CLEAR: Continuous Latent Autoregressive Modeling for High-quality and Low-latency Speech Synthesis V oicebox: Text-guided multilin- gual universal speech generation at scale

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:01:54.567050Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T16:01:52.673697Z digest=sha256:67f25500380c5a05195d15416691a63b8dde8145a73dbd38d0dc63768f71cf00

Observation eebed74a-9958-470e-95d0-23b257fd04be · outbound

This paper cites REPA-E: Unlocking V AE for End-to-End Tuning with Latent Diffusion Transformers, April 2025.

CLEAR: Continuous Latent Autoregressive Modeling for High-quality and Low-latency Speech Synthesis REPA-E: Unlocking V AE for End-to-End Tuning with Latent Diffusion Transformers, April 2025

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:01:54.549525Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T16:01:52.678990Z digest=sha256:8b20b5023ddd471d689a57f939309cf424bd6ac6541a0517137830374d46652a

Observation 91a25cc2-7bd8-42d5-8329-6c8aba46c0a8 · outbound

This paper cites Neural Speech Synthesis with Transformer Network.

CLEAR: Continuous Latent Autoregressive Modeling for High-quality and Low-latency Speech Synthesis Neural Speech Synthesis with Transformer Network

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-05T16:01:52.684377Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:01:52.684377Z digest=sha256:cf3adbe79ac0abbfce7db82714d52a987ae3fd1c47789fa5907d59122acde31b

Observation 6e8c5c59-9594-4e39-b972-1cd22596e5c6 · outbound

This paper cites Autoregressive Image Generation without Vector Quantization.

CLEAR: Continuous Latent Autoregressive Modeling for High-quality and Low-latency Speech Synthesis Autoregressive Image Generation without Vector Quantization

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-05T16:01:52.689870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:01:52.689870Z digest=sha256:61a8aed0d0b6ec932231ed86833a726aad52dfa7f0231e13bf2b93b4e5e7fb2b

Observation 54dc638d-e939-4cb4-8af7-14894dfa04f4 · outbound

This paper cites Flow Straight and Fast: Learning to Generate and Transfer Data with Rectified Flow.

CLEAR: Continuous Latent Autoregressive Modeling for High-quality and Low-latency Speech Synthesis Flow Straight and Fast: Learning to Generate and Transfer Data with Rectified Flow

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-05T16:01:52.695773Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:01:52.695773Z digest=sha256:b85213d1c02343b614466748b8ac99e944c7397f382538c820ed1827f9cbc834

Observation 9fbfafa0-f539-453a-95f2-15b3161ffe4c · outbound

This paper cites Autoregressive Diffusion Transformer for Text-to-Speech Synthesis.

CLEAR: Continuous Latent Autoregressive Modeling for High-quality and Low-latency Speech Synthesis Autoregressive Diffusion Transformer for Text-to-Speech Synthesis

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-05T16:01:52.701539Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:01:52.701539Z digest=sha256:62915ea48f319c19dfddd73eac1ba73ded288dd1e4facc7b4836d211ad30f3d3

Observation 439fbd8b-bb4f-46e5-95be-fc4997f507ad · outbound

This paper cites LibriSpeech-PC: Benchmark for Evaluation of Punctuation and Capitalization Capabilities of End-to-End ASR Models.

CLEAR: Continuous Latent Autoregressive Modeling for High-quality and Low-latency Speech Synthesis LibriSpeech-PC: Benchmark for Evaluation of Punctuation and Capitalization Capabilities of End-to-End ASR Models

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:01:54.516431Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T16:01:52.707422Z digest=sha256:8927c5a4facb96aa2c92e8f7fa919b66ddf18fd6a617af534fe589c9084b7be9

Observation 95e71bba-b9f8-4aa8-9666-b8f109020006 · outbound

This paper cites Autoregressive Speech Synthesis without Vector Quantization.

CLEAR: Continuous Latent Autoregressive Modeling for High-quality and Low-latency Speech Synthesis Autoregressive Speech Synthesis without Vector Quantization

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-05T16:01:52.712652Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:01:52.712652Z digest=sha256:06e7ffb5b710834909cb78bb73186e75968c1236ca639b396ec0549f1f8e17be

Observation 40e3b997-26db-4c2a-bd81-b5ea98b4355a · outbound

This paper cites Finite Scalar Quantization: VQ-V AE Made Simple, October 2023.

CLEAR: Continuous Latent Autoregressive Modeling for High-quality and Low-latency Speech Synthesis Finite Scalar Quantization: VQ-V AE Made Simple, October 2023

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:01:54.492317Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T16:01:52.717776Z digest=sha256:e104ef8959b333f9d0e226bc0900ea2cefa234f6c68e2478855b478133e46671

Observation ca2f0a00-cd01-453c-ab18-6e562a1fbf90 · outbound

This paper cites HALL-E: Hierarchical Neural Codec Language Model for Minute-Long Zero-Shot Text-to-Speech Synthesis.

CLEAR: Continuous Latent Autoregressive Modeling for High-quality and Low-latency Speech Synthesis HALL-E: Hierarchical Neural Codec Language Model for Minute-Long Zero-Shot Text-to-Speech Synthesis

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-05T16:01:52.722764Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:01:52.722764Z digest=sha256:996983b277adeeb35d5cd409f0010e8520a700cd965045c802337eceb422483b

Observation 1c9ac590-e942-4621-8e3c-34b488390f90 · outbound

This paper cites Librispeech: An ASR corpus based on public domain audio books.

CLEAR: Continuous Latent Autoregressive Modeling for High-quality and Low-latency Speech Synthesis Librispeech: An ASR corpus based on public domain audio books

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:01:54.470831Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T16:01:52.728945Z digest=sha256:96ce92da1f0a3301c400b4f45f77d091d3e3220a893ddceb8230a4add0aacf54

Observation 25b5268c-6ecf-4fe3-9b77-07cb518ff5bd · outbound

This paper cites Robust speech recognition via large-scale weak supervision.

CLEAR: Continuous Latent Autoregressive Modeling for High-quality and Low-latency Speech Synthesis Robust speech recognition via large-scale weak supervision

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-05T16:01:52.736142Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:01:52.736142Z digest=sha256:27c7a700a8793c859a40c9481599ef3663fcb96ef4bf0a7ce45e6ae886016ebe

Observation 5e2ca104-69e1-476a-8e25-097231b26321 · outbound

This paper cites Revisiting over-smoothness in text to speech.

CLEAR: Continuous Latent Autoregressive Modeling for High-quality and Low-latency Speech Synthesis Revisiting over-smoothness in text to speech

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:01:54.435875Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T16:01:52.743452Z digest=sha256:5b735439bef691f84be0c9ac9b33c99a739e48ef64959555ccc8b9b4b791d335

Observation 7d78640a-6f08-41fe-8d3a-5d48bcd6aad6 · outbound

This paper cites UTMOS: UTokyo-SaruLab System for VoiceMOS Challenge 2022.

CLEAR: Continuous Latent Autoregressive Modeling for High-quality and Low-latency Speech Synthesis UTMOS: UTokyo-SaruLab System for VoiceMOS Challenge 2022

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-05T16:01:52.748760Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:01:52.748760Z digest=sha256:e9bda5871805040a6d3ce7d867da71239b73841f02f9bc0e4a82624b1c5833c5

Observation 7183e728-81f3-4342-807a-9d614121e605 · outbound

This paper cites AutoClip: Adaptive Gradient Clipping for Source Separation Networks, July 2020.

CLEAR: Continuous Latent Autoregressive Modeling for High-quality and Low-latency Speech Synthesis AutoClip: Adaptive Gradient Clipping for Source Separation Networks, July 2020

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:01:54.418343Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T16:01:52.756183Z digest=sha256:a6092f4432a360c0a61234465c0742f4df6ccfff7f0630874c04620a2fe99c85

Observation f5943c9b-07cd-4172-98a2-2b6298578e77 · outbound

This paper cites GLU Variants Improve Transformer.

CLEAR: Continuous Latent Autoregressive Modeling for High-quality and Low-latency Speech Synthesis GLU Variants Improve Transformer

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-05T16:01:52.761504Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:01:52.761504Z digest=sha256:9e2f90f3b8c987a709134807be8c9c74e75a0dbe0553278e506510d1438ec582

Observation 3c67adb7-6ae8-4540-9c83-0d94bb0ab2bd · outbound

This paper cites NaturalSpeech 2: Latent Diffusion Models are Natural and Zero-Shot Speech and Singing Synthesizers, May 2023.

CLEAR: Continuous Latent Autoregressive Modeling for High-quality and Low-latency Speech Synthesis NaturalSpeech 2: Latent Diffusion Models are Natural and Zero-Shot Speech and Singing Synthesizers, May 2023

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:01:54.396325Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T16:01:52.767480Z digest=sha256:35ae169c60544dfd2e9f8997ccfa45a3c1b1c7a520ccc93b21a3ba8cc4da4c14

Observation aeb04ce9-c2aa-4a98-9784-484353023d73 · outbound

This paper cites Real-time single image and video super-resolution using an efficient sub-pixel convolutional neural network.

CLEAR: Continuous Latent Autoregressive Modeling for High-quality and Low-latency Speech Synthesis Real-time single image and video super-resolution using an efficient sub-pixel convolutional neural network

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-05T16:01:52.774022Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:01:52.774022Z digest=sha256:0a8483d456590e88888a722e086048c62b0b4dcd36575d99ad93b2409e7e0cb2

Observation f7350a1b-82de-4554-bdc7-c4b14864f0d5 · outbound

This paper cites Ella-v: Stable neural codec language modeling with alignment-guided sequence reordering.

CLEAR: Continuous Latent Autoregressive Modeling for High-quality and Low-latency Speech Synthesis Ella-v: Stable neural codec language modeling with alignment-guided sequence reordering

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:01:54.359518Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T16:01:52.778679Z digest=sha256:339646d5f239d174a2df9463b4095c9bdfa17b213a9e6166fd92263f01ae0d3f

Observation 8d0864e8-9d61-4847-b4ae-8007d02a4b4a · outbound

This paper cites Steinmetz, Jordi Pons, Santiago Pascual, and Joan Serrà.

CLEAR: Continuous Latent Autoregressive Modeling for High-quality and Low-latency Speech Synthesis Steinmetz, Jordi Pons, Santiago Pascual, and Joan Serrà

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:01:54.339544Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T16:01:52.785227Z digest=sha256:bb680363269f5513fc059d67ff32650fec2eceb8ad946c79510577dae3eb8ca1

Observation 88ac9b82-424e-4284-90e6-7e3474890267 · outbound

This paper cites RoFormer: Enhanced Transformer with Rotary Position Embedding.

CLEAR: Continuous Latent Autoregressive Modeling for High-quality and Low-latency Speech Synthesis RoFormer: Enhanced Transformer with Rotary Position Embedding

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-05T16:01:52.792348Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:01:52.792348Z digest=sha256:09efddf393df586c6970e7493a12038b1fb4b11e37550e8c5d63bfb9afe00be3

Observation 7e22d374-289f-4f3a-9e3a-fd63b7c1ac92 · outbound

This paper cites NaturalSpeech: End-to-End Text-to-Speech Synthesis With Human-Level Quality.

CLEAR: Continuous Latent Autoregressive Modeling for High-quality and Low-latency Speech Synthesis NaturalSpeech: End-to-End Text-to-Speech Synthesis With Human-Level Quality

Reference 51

Resolution
metadata mismatch
raw_fallback, observed 2026-08-05T16:01:53.540373Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T16:01:52.797356Z digest=sha256:8b4ab5a6cad0220e53a1df96d1e515635eb731a45a68dfede3a37bc2c7410d4b

Observation 1d9c93ff-5827-44e5-a4de-57432d2981a0 · outbound

This paper cites Continuous Speech Synthesis using per-token Latent Diffusion, October 2024.

CLEAR: Continuous Latent Autoregressive Modeling for High-quality and Low-latency Speech Synthesis Continuous Speech Synthesis using per-token Latent Diffusion, October 2024

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:01:54.309528Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T16:01:52.802678Z digest=sha256:fda0704746e776469b3e06f2bed639dca33567f13c306a198c129838517fb91d

Observation 6e0e7da8-4df0-4875-9f8c-592ba1c12b98 · outbound

This paper cites Neural Discrete Representation Learning.

CLEAR: Continuous Latent Autoregressive Modeling for High-quality and Low-latency Speech Synthesis Neural Discrete Representation Learning

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:01:54.284600Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T16:01:52.808377Z digest=sha256:c3b9e0bfc041477939d03f5731b7302f4d3f7125c8b94aa678b151ee3e17c946

Observation e7a870c0-8ae9-4cb2-b7d2-6361397113fe · outbound

This paper cites Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers.

CLEAR: Continuous Latent Autoregressive Modeling for High-quality and Low-latency Speech Synthesis Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-05T16:01:52.813060Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:01:52.813060Z digest=sha256:643a830f74ff0156f2d540c7a3416baa733e081ea0d4cab34a93f51fb14fc13c

Observation c5b0d8c7-9cc3-4359-8b41-ce886d522b74 · outbound

This paper cites FELLE: Autoregressive Speech Synthesis with Token-Wise Coarse-to-Fine Flow Matching.

CLEAR: Continuous Latent Autoregressive Modeling for High-quality and Low-latency Speech Synthesis FELLE: Autoregressive Speech Synthesis with Token-Wise Coarse-to-Fine Flow Matching

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-05T16:01:52.818331Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:01:52.818331Z digest=sha256:eb4a344531f47bbe2fd96524e1e23cbd8eccd646dbb6e14e9bd1d870726bddfa

Observation 61f5c358-f70e-48f2-9779-a8badee88a3c · outbound

This paper cites Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens.

CLEAR: Continuous Latent Autoregressive Modeling for High-quality and Low-latency Speech Synthesis Spark-TTS: An Efficient LLM-Based Text-to-Speech Model with Single-Stream Decoupled Speech Tokens

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-05T16:01:52.823558Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:01:52.823558Z digest=sha256:ea6776a5294c8af323cff9261ee9861728cb38f8524262f033ba39602d83b296

Observation 1d998edd-d252-4ed5-a46d-6df9e126adaa · outbound

This paper cites MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer, October 2024.

CLEAR: Continuous Latent Autoregressive Modeling for High-quality and Low-latency Speech Synthesis MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer, October 2024

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:01:54.263650Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T16:01:52.829746Z digest=sha256:82ffdd4193075bf44fde13e669c2e164147ca704d990f2a4e3cb39839a55c628

Observation 536b2ec0-62b9-481b-b8ee-ab4ccc1042b9 · outbound

This paper cites Towards audio language modeling -- an overview.

CLEAR: Continuous Latent Autoregressive Modeling for High-quality and Low-latency Speech Synthesis Towards audio language modeling -- an overview

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-05T16:01:52.834220Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:01:52.834220Z digest=sha256:cfa497ef0ea74ad5ce306ac0ce3a04f2d760a198df0f6e489e338adc6429300d

Observation 981dec4e-3a18-4f0b-a20a-d12e616454d4 · outbound

This paper cites RALL-E: Robust Codec Language Modeling with Chain-of-Thought Prompting for Text-to-Speech Synthesis.

CLEAR: Continuous Latent Autoregressive Modeling for High-quality and Low-latency Speech Synthesis RALL-E: Robust Codec Language Modeling with Chain-of-Thought Prompting for Text-to-Speech Synthesis

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-05T16:01:52.839342Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:01:52.839342Z digest=sha256:cafe9e01a9b0ee57dc6349add4d691fa9df720df10d332ea3724f3f11429f0f9

Observation fa13fcd6-7014-4990-9911-08281a879493 · outbound

This paper cites On Layer Normalization in the Transformer Architecture.

CLEAR: Continuous Latent Autoregressive Modeling for High-quality and Low-latency Speech Synthesis On Layer Normalization in the Transformer Architecture

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-05T16:01:52.844345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:01:52.844345Z digest=sha256:af4cdd6e2640f133c788ada68904112e6aa379380276eb6eaa7ca40f3ac9ee7f

Observation 02ee09ec-e051-4671-805d-6f91e1e6141b · outbound

This paper cites Reconstruction vs.

CLEAR: Continuous Latent Autoregressive Modeling for High-quality and Low-latency Speech Synthesis Reconstruction vs

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:01:54.245081Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T16:01:52.849095Z digest=sha256:b6855343faab3c56fa05a602608c9e477ed2208329c122932a367f94004bb40f

Observation cc1ff55e-0c36-4212-9fcd-d9cfb7916350 · outbound

This paper cites Llasa: Scaling Train-Time and Inference-Time Compute for Llama-based Speech Synthesis.

CLEAR: Continuous Latent Autoregressive Modeling for High-quality and Low-latency Speech Synthesis Llasa: Scaling Train-Time and Inference-Time Compute for Llama-based Speech Synthesis

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-05T16:01:52.854277Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:01:52.854277Z digest=sha256:bff55f42e3a01f9bc86db4640ee0dfdf2f4d62fcd7136f82baec152112f8a3de

Observation ffefd789-2cff-4087-abec-ee7e034b4dd6 · outbound

This paper cites Representation Alignment for Generation: Training Diffusion Transformers Is Easier Than You Think, February 2025.

CLEAR: Continuous Latent Autoregressive Modeling for High-quality and Low-latency Speech Synthesis Representation Alignment for Generation: Training Diffusion Transformers Is Easier Than You Think, February 2025

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:01:54.219454Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T16:01:52.859410Z digest=sha256:3f9c115d667c6e16968eb1046e5ea8178c91493d93351112a0a8c1815ec41898

Observation 0fb202ed-c747-45d3-918b-8fe81290ee9f · outbound

This paper cites Soundstream: An end-to-end neural audio codec.

CLEAR: Continuous Latent Autoregressive Modeling for High-quality and Low-latency Speech Synthesis Soundstream: An end-to-end neural audio codec

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-05T16:01:52.864521Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:01:52.864521Z digest=sha256:2c2fa35df92c24d090008e4bbe2afffe8dbab740718b6cd9698a8bbc78209548

Observation e9ac53ff-9952-42b0-9262-3634970b2ed3 · outbound

This paper cites Weiss, Ye Jia, Zhifeng Chen, and Yonghui Wu.

CLEAR: Continuous Latent Autoregressive Modeling for High-quality and Low-latency Speech Synthesis Weiss, Ye Jia, Zhifeng Chen, and Yonghui Wu

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:01:54.188022Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T16:01:52.871677Z digest=sha256:6591bbe48becedc228478a24742bcd2f011a86224c2448c7d4ceac60291dc473

Observation 449586ca-366b-4904-8afd-99a0efc17c80 · outbound

This paper cites Root Mean Square Layer Normalization.

CLEAR: Continuous Latent Autoregressive Modeling for High-quality and Low-latency Speech Synthesis Root Mean Square Layer Normalization

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-05T16:01:52.877438Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:01:52.877438Z digest=sha256:a6cb87d2e523152f79e8eaa6ca8817f6a5dc2b37d0c7ed0e320f148deed9b308

Observation d9c9843c-e0b0-4d37-9727-f0540ccd7d32 · outbound

This paper cites Speak Foreign Languages with Your Own Voice: Cross-Lingual Neural Codec Language Modeling.

CLEAR: Continuous Latent Autoregressive Modeling for High-quality and Low-latency Speech Synthesis Speak Foreign Languages with Your Own Voice: Cross-Lingual Neural Codec Language Modeling

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-05T16:01:52.883671Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:01:52.883671Z digest=sha256:b572590c6a84beb69f574604c922f28a7e3bfb91a0afec01a2b8cf9775f4accd

Observation 3c846336-a41f-4239-85a9-dc8976d02ace · outbound

This paper cites ,𝑁𝑆𝐵,𝐶!,𝑁 + UpsamplingBlock Channel toSpace ChannelDuplicating𝐵,𝐶.

CLEAR: Continuous Latent Autoregressive Modeling for High-quality and Low-latency Speech Synthesis ,𝑁𝑆𝐵,𝐶!,𝑁 + UpsamplingBlock Channel toSpace ChannelDuplicating𝐵,𝐶

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T16:01:54.168792Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T16:01:52.889397Z digest=sha256:8b8a695daa66ba9a8db443c5c2aedb08ce67953d0eecb39a3698c9c6df696d8f

Observation c35e1bbd-d098-48b1-b4de-3d045c067d50 · outbound

This paper cites Libriheavy: a 50,000 hours ASR corpus with punctuation casing and context.

CLEAR: Continuous Latent Autoregressive Modeling for High-quality and Low-latency Speech Synthesis Libriheavy: a 50,000 hours ASR corpus with punctuation casing and context

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-05T16:01:52.645000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:01:52.645000Z digest=sha256:bf5b35de738b6debb23fdf32047d04dea652b3957809f9ef4e04e972e674f0f9

Pith citing papers

Observation f877a92c-605b-4346-bb9a-935b28c1d987 · inbound

On the Distillation Loss Functions of Speech VAE for Unified Reconstruction, Understanding, and Generation cites this paper.

On the Distillation Loss Functions of Speech VAE for Unified Reconstruction, Understanding, and Generation CLEAR: Continuous Latent Autoregressive Modeling for High-quality and Low-latency Speech Synthesis

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-11T10:25:59.528620Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T15:30:58.664375Z digest=sha256:0555e8a8d5f09be8bc8a32ec0c9074bdab784a73bd2ece782c38726feee65937

Observation c6001f44-543d-4894-8ba1-62ff1c34b900 · inbound

SemaVoice: Semantic-Aware Continuous Autoregressive Speech Synthesis cites this paper.

SemaVoice: Semantic-Aware Continuous Autoregressive Speech Synthesis CLEAR: Continuous Latent Autoregressive Modeling for High-quality and Low-latency Speech Synthesis

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-19T19:02:43.569018Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-19T18:58:15.288299Z digest=sha256:3a9e5100dc8c7aff8e55b29f95714c0ea9f680d2132e0db2988e37d05b0eb4e2

Observation 6947956a-a27e-4039-9221-c3793b937b2b · inbound

VoxCPM2 Technical Report cites this paper.

VoxCPM2 Technical Report CLEAR: Continuous Latent Autoregressive Modeling for High-quality and Low-latency Speech Synthesis

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-07-02T19:47:19.714861Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-27T21:18:22.911332Z digest=sha256:53b5a770ba24a674c418e821330c8e0b62abb851d44c3e9fa89ff21e8365b943