Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T14:22:31.869618Z
Paper Citation Record · LEDGER
As of 9 August 2026, this Paper Citation Record lists 35 of 35 outbound references and 2 inbound Pith citation observations for arXiv:2505.19206.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T14:22:31.869618Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-05T15:52:44.054140Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-02T17:07:12.867079Z
35 of 35 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 5c797376-1075-4f7a-b805-ed8753374b7c · outbound
SpeakStream: Streaming Text-to-Speech with Interleaved Data Audi- olm: a language modeling approach to audio generation,
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 2a49efeb-c160-4df2-b2ed-2eb489baef6e · outbound
SpeakStream: Streaming Text-to-Speech with Interleaved Data Qwen2.5-Omni Technical Report
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bc980c59-7ca6-4cfc-9ad1-f31144368a47 · outbound
SpeakStream: Streaming Text-to-Speech with Interleaved Data Spirit-lm: Interleaved spoken and written language model,
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c1bc2262-c193-473e-821f-5069f1f8eb6d · outbound
SpeakStream: Streaming Text-to-Speech with Interleaved Data MMAU: A Massive Multi-Task Audio Understanding and Reasoning Benchmark
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0b239529-5d94-462f-9b54-9ad986574e10 · outbound
SpeakStream: Streaming Text-to-Speech with Interleaved Data Zero-Shot Text-to-Speech from Continuous Text Streams
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8928467d-c253-4fba-ac9c-676f4539d528 · outbound
SpeakStream: Streaming Text-to-Speech with Interleaved Data Speak while you think: Streaming speech synthesis during text generation,
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation b18d51a4-4449-433f-b0d8-7f587de7c478 · outbound
SpeakStream: Streaming Text-to-Speech with Interleaved Data Interleaved Speech-Text Language Models for Simple Streaming Text-to-Speech Synthesis
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c136ad76-0c3f-424d-97f2-a7621681b0e8 · outbound
SpeakStream: Streaming Text-to-Speech with Interleaved Data dMel: Speech Tokenization made Simple
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4d648051-cf1a-46e1-9b8e-67188104c6c8 · outbound
SpeakStream: Streaming Text-to-Speech with Interleaved Data A 3T: Alignment-aware acoustic and text pretraining for speech synthesis and editing,
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 9ffb4b08-22c6-47bd-ad27-01f64133f11b · outbound
SpeakStream: Streaming Text-to-Speech with Interleaved Data Yourtts: Towards zero-shot multi-speaker tts and zero-shot voice conversion for everyone,
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 204d923a-49c9-443c-ad21-c47bea9e7cbb · outbound
SpeakStream: Streaming Text-to-Speech with Interleaved Data CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 69c1a5bb-5c73-4bd8-8a8f-4cba12c0e441 · outbound
SpeakStream: Streaming Text-to-Speech with Interleaved Data E3 tts: Easy end-to- end diffusion-based text to speech,
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 7fc88f8e-e561-4c13-89a6-35e5a654f82f · outbound
SpeakStream: Streaming Text-to-Speech with Interleaved Data FastSpeech 2: Fast and High-Quality End-to-End Text to Speech
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6fa8f085-6f5f-4e23-8147-693b3da14f4e · outbound
SpeakStream: Streaming Text-to-Speech with Interleaved Data Tacotron: Towards End-to-End Speech Synthesis
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 30b3b16d-7403-4add-a2d3-9a216fec6c90 · outbound
SpeakStream: Streaming Text-to-Speech with Interleaved Data (2024) Text-to-speech guide
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 1efec5e5-675f-4cf5-8bed-cd3d766ca544 · outbound
SpeakStream: Streaming Text-to-Speech with Interleaved Data Streamspeech: Low-latency neural architecture for high-quality on-device speech synthesis,
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 6209cf8c-2fd9-47ab-b891-3aee83ad4a78 · outbound
SpeakStream: Streaming Text-to-Speech with Interleaved Data MLX: Efficient and flexible machine learning on apple silicon,
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e261da5d-e116-4fe1-b3ab-5a9103f733a5 · outbound
SpeakStream: Streaming Text-to-Speech with Interleaved Data VALL-T: Decoder-Only Generative Transducer for Robust and Decoding-Controllable Text-to-Speech
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 20ab1739-357b-457d-91d3-b16a217f4b12 · outbound
SpeakStream: Streaming Text-to-Speech with Interleaved Data BERT: A Review of Applications in Natural Language Processing and Understanding
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d9bcda6d-b7a2-4e38-90e2-dbf56765bdbf · outbound
SpeakStream: Streaming Text-to-Speech with Interleaved Data Non-Attentive Tacotron: Robust and Controllable Neural TTS Synthesis Including Unsupervised Duration Modeling
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 264a593b-6af6-4e94-a1a6-e63767c50fec · outbound
SpeakStream: Streaming Text-to-Speech with Interleaved Data LiveSpeech: Low-Latency Zero-shot Text-to-Speech via Autoregressive Modeling of Audio Discrete Codes
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b70e1eb8-b561-42d4-bcc2-511129d73975 · outbound
SpeakStream: Streaming Text-to-Speech with Interleaved Data Transduce and speak: Neural transducer for text-to-speech with semantic token predic- tion,
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation fcf3c1f7-9d52-4c8b-b2ef-cf855d75e533 · outbound
SpeakStream: Streaming Text-to-Speech with Interleaved Data Parallel wavegan: A fast waveform generation model based on generative adversarial networks with multi-resolution spectrogram,
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5fc08e52-2a4b-4dfc-88ac-3e787debc790 · outbound
SpeakStream: Streaming Text-to-Speech with Interleaved Data Hifi-gan: Generative adversarial networks for efficient and high fidelity speech synthesis,
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bf912353-3894-4d2b-86e5-1a360d45569b · outbound
SpeakStream: Streaming Text-to-Speech with Interleaved Data Bigvgan: A universal neural vocoder with large-scale training,
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 939b6549-48d9-4fc6-b00e-96a60271dc9c · outbound
SpeakStream: Streaming Text-to-Speech with Interleaved Data V ocos: Closing the gap between time-domain and fourier- based neural vocoders for high-quality audio synthesis,
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 027f9df9-6925-4d0d-941b-fc979cc4d52a · outbound
SpeakStream: Streaming Text-to-Speech with Interleaved Data Non-causal to causal ssl-supported transfer learning: Towards a high-performance low-latency speech vocoder,
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 64393298-7ab5-4cbc-933d-fc9615b7d845 · outbound
SpeakStream: Streaming Text-to-Speech with Interleaved Data Coqui TTS: A deep learning toolkit for Text-to-Speech, battle-tested in research and production,
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 71b7ac5a-074f-4634-8442-ef81b3ea1d29 · outbound
SpeakStream: Streaming Text-to-Speech with Interleaved Data The lj speech dataset,
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3fd2acba-0c0e-4856-a3d9-9b89e5eb8477 · outbound
SpeakStream: Streaming Text-to-Speech with Interleaved Data Whisperx: Time-accurate speech transcription of long-form audio,
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation ad6a3e0d-0db1-425d-96dd-3d850d4678a6 · outbound
SpeakStream: Streaming Text-to-Speech with Interleaved Data Robust speech recognition via large-scale weak super- vision,
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 44a264d3-1988-48c0-bc65-f38876f91a97 · outbound
SpeakStream: Streaming Text-to-Speech with Interleaved Data Libritts-r: A restored multi-speaker text-to-speech corpus,
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation f1fbf252-58e1-4706-91c0-32a9c352fb57 · outbound
SpeakStream: Streaming Text-to-Speech with Interleaved Data Librispeech: an asr corpus based on public domain audio books,
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 841334ad-98d8-4578-a9aa-8572b38c2b0b · outbound
SpeakStream: Streaming Text-to-Speech with Interleaved Data Moshi: a speech-text foundation model for real-time dialogue
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d2eae670-95c0-493b-8556-5d1d38b0af9b · outbound
SpeakStream: Streaming Text-to-Speech with Interleaved Data Available: https://www.coqui.ai
Reference 2021
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 5fea5dc5-c540-4237-bc1d-d1f003ccf490 · inbound
ChipChat: Low-Latency Cascaded Conversational Agent in MLX SpeakStream: Streaming Text-to-Speech with Interleaved Data
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6227ce77-3ec2-40de-9470-32c1250b84ee · inbound
Don't Pause: Streaming Video-Language Synchrony for Online Video Understanding SpeakStream: Streaming Text-to-Speech with Interleaved Data
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.