Pith. sign in

Paper Citation Record · LEDGER

SpeakStream: Streaming Text-to-Speech with Interleaved Data

As of 9 August 2026, this Paper Citation Record lists 35 of 35 outbound references and 2 inbound Pith citation observations for arXiv:2505.19206.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.19206 v1

Coverage vector

measured 35 of 35 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:22:31.869618Z

measured 37 of 37 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T15:52:44.054140Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T17:07:12.867079Z

Reference resolution

35 of 35 outbound references displayed

  • verified exact0
  • verified fuzzy14
  • unresolved21
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 5c797376-1075-4f7a-b805-ed8753374b7c · outbound

This paper cites Audi- olm: a language modeling approach to audio generation,.

SpeakStream: Streaming Text-to-Speech with Interleaved Data Audi- olm: a language modeling approach to audio generation,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:22:33.551223Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:22:30.401774Z digest=sha256:d44d129736227efe5d43d4796a8f66712def3482bbc78382aa2d330bd4ea30e4

Observation 2a49efeb-c160-4df2-b2ed-2eb489baef6e · outbound

This paper cites Qwen2.5-Omni Technical Report.

SpeakStream: Streaming Text-to-Speech with Interleaved Data Qwen2.5-Omni Technical Report

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T14:22:30.544240Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:22:30.544240Z digest=sha256:fb6e0d77f30a31bd04bfe1bfe65ddddfd043a9d136ecb4688a2d9ae65fca349d

Observation bc980c59-7ca6-4cfc-9ad1-f31144368a47 · outbound

This paper cites Spirit-lm: Interleaved spoken and written language model,.

SpeakStream: Streaming Text-to-Speech with Interleaved Data Spirit-lm: Interleaved spoken and written language model,

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T14:22:30.595293Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:22:30.595293Z digest=sha256:eece483a4405d159f114efc41b439b00884cb97cf6154bc601f1a8ef73df3309

Observation c1bc2262-c193-473e-821f-5069f1f8eb6d · outbound

This paper cites MMAU: A Massive Multi-Task Audio Understanding and Reasoning Benchmark.

SpeakStream: Streaming Text-to-Speech with Interleaved Data MMAU: A Massive Multi-Task Audio Understanding and Reasoning Benchmark

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T14:22:30.673231Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:22:30.673231Z digest=sha256:7c3bc499c521c24dd42690e9a0cf0b6554b6db5ade3347714848641c911bb092

Observation 0b239529-5d94-462f-9b54-9ad986574e10 · outbound

This paper cites Zero-Shot Text-to-Speech from Continuous Text Streams.

SpeakStream: Streaming Text-to-Speech with Interleaved Data Zero-Shot Text-to-Speech from Continuous Text Streams

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T14:22:30.726922Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:22:30.726922Z digest=sha256:b8851db5f934c8aca80ca9b803bd480d6e3f1cb80e1feb761331ac4dff0e91ae

Observation 8928467d-c253-4fba-ac9c-676f4539d528 · outbound

This paper cites Speak while you think: Streaming speech synthesis during text generation,.

SpeakStream: Streaming Text-to-Speech with Interleaved Data Speak while you think: Streaming speech synthesis during text generation,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:22:33.477254Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:22:30.780606Z digest=sha256:ad0344377038c9f09058e380929bb7c46cc1dee7978565bc1190a274bf28183d

Observation b18d51a4-4449-433f-b0d8-7f587de7c478 · outbound

This paper cites Interleaved Speech-Text Language Models for Simple Streaming Text-to-Speech Synthesis.

SpeakStream: Streaming Text-to-Speech with Interleaved Data Interleaved Speech-Text Language Models for Simple Streaming Text-to-Speech Synthesis

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T14:22:30.842860Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:22:30.842860Z digest=sha256:d7a1a379ef7feac43cae9f645fe0159e21dc2216a27f670eb32caa8aa6962f42

Observation c136ad76-0c3f-424d-97f2-a7621681b0e8 · outbound

This paper cites dMel: Speech Tokenization made Simple.

SpeakStream: Streaming Text-to-Speech with Interleaved Data dMel: Speech Tokenization made Simple

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T14:22:30.915758Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:22:30.915758Z digest=sha256:431e4b64d46c73139b0721846db349dbfd29df62f0fba78283edb7526d295dd8

Observation 4d648051-cf1a-46e1-9b8e-67188104c6c8 · outbound

This paper cites A 3T: Alignment-aware acoustic and text pretraining for speech synthesis and editing,.

SpeakStream: Streaming Text-to-Speech with Interleaved Data A 3T: Alignment-aware acoustic and text pretraining for speech synthesis and editing,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:22:33.434512Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:22:31.004924Z digest=sha256:d8ffb5e2738c425a2ccba36c8936f4cf0fc6ddc40b926005702ae8b30319e650

Observation 9ffb4b08-22c6-47bd-ad27-01f64133f11b · outbound

This paper cites Yourtts: Towards zero-shot multi-speaker tts and zero-shot voice conversion for everyone,.

SpeakStream: Streaming Text-to-Speech with Interleaved Data Yourtts: Towards zero-shot multi-speaker tts and zero-shot voice conversion for everyone,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T14:22:31.029594Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:22:31.029594Z digest=sha256:798b64aec8b45fabaa842123e8968b5d5678a5723b1cfe7285b8d5c0a3674522

Observation 204d923a-49c9-443c-ad21-c47bea9e7cbb · outbound

This paper cites CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens.

SpeakStream: Streaming Text-to-Speech with Interleaved Data CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T14:22:31.087888Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:22:31.087888Z digest=sha256:ddee176f1e5b39e22131ae2a2e3efb64bac3e163d8af2477f7e4476218a1930e

Observation 69c1a5bb-5c73-4bd8-8a8f-4cba12c0e441 · outbound

This paper cites E3 tts: Easy end-to- end diffusion-based text to speech,.

SpeakStream: Streaming Text-to-Speech with Interleaved Data E3 tts: Easy end-to- end diffusion-based text to speech,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:22:33.343081Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:22:31.132497Z digest=sha256:5212b20c1b84bcc97e0115c60d55d95abeb0fe41dfeab9fddd43cd2e99a0f62a

Observation 7fc88f8e-e561-4c13-89a6-35e5a654f82f · outbound

This paper cites FastSpeech 2: Fast and High-Quality End-to-End Text to Speech.

SpeakStream: Streaming Text-to-Speech with Interleaved Data FastSpeech 2: Fast and High-Quality End-to-End Text to Speech

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T14:22:31.169600Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:22:31.169600Z digest=sha256:817923dcd75bb7c1669706fdfc0658c0772d1e9028c347b699d9de4f2540e39f

Observation 6fa8f085-6f5f-4e23-8147-693b3da14f4e · outbound

This paper cites Tacotron: Towards End-to-End Speech Synthesis.

SpeakStream: Streaming Text-to-Speech with Interleaved Data Tacotron: Towards End-to-End Speech Synthesis

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T14:22:31.221168Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:22:31.221168Z digest=sha256:0ecb07f7aa6523b56ba581447cf809b6980040a8b9fb73404d62b770dd2a833d

Observation 30b3b16d-7403-4add-a2d3-9a216fec6c90 · outbound

This paper cites (2024) Text-to-speech guide.

SpeakStream: Streaming Text-to-Speech with Interleaved Data (2024) Text-to-speech guide

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:22:33.304761Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:22:31.249785Z digest=sha256:71f85bf935c0004ef585e0b3c12c0cfd414ea38a1bf60931dd34b654cd583f2f

Observation 1efec5e5-675f-4cf5-8bed-cd3d766ca544 · outbound

This paper cites Streamspeech: Low-latency neural architecture for high-quality on-device speech synthesis,.

SpeakStream: Streaming Text-to-Speech with Interleaved Data Streamspeech: Low-latency neural architecture for high-quality on-device speech synthesis,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:22:33.252460Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:22:31.282325Z digest=sha256:e5e070615270a2f4d409b8d499a5eb1fe249e69a80825f0fc3f2a0fd85b36a6f

Observation 6209cf8c-2fd9-47ab-b891-3aee83ad4a78 · outbound

This paper cites MLX: Efficient and flexible machine learning on apple silicon,.

SpeakStream: Streaming Text-to-Speech with Interleaved Data MLX: Efficient and flexible machine learning on apple silicon,

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T14:22:31.327328Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:22:31.327328Z digest=sha256:b159536ac5128dfd79e8488f4ad8d92552e61daedc7c61df8abe8047897e4a66

Observation e261da5d-e116-4fe1-b3ab-5a9103f733a5 · outbound

This paper cites VALL-T: Decoder-Only Generative Transducer for Robust and Decoding-Controllable Text-to-Speech.

SpeakStream: Streaming Text-to-Speech with Interleaved Data VALL-T: Decoder-Only Generative Transducer for Robust and Decoding-Controllable Text-to-Speech

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T14:22:31.416374Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:22:31.416374Z digest=sha256:64711a48fb517de81dbeb38494396795022bc36d74e309ce07c379e2c833261c

Observation 20ab1739-357b-457d-91d3-b16a217f4b12 · outbound

This paper cites BERT: A Review of Applications in Natural Language Processing and Understanding.

SpeakStream: Streaming Text-to-Speech with Interleaved Data BERT: A Review of Applications in Natural Language Processing and Understanding

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T14:22:31.501147Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:22:31.501147Z digest=sha256:0b84c7dda7ba02da8ee4b785c0dad4dbd2871b58d36bcc7057b62061178df270

Observation d9bcda6d-b7a2-4e38-90e2-dbf56765bdbf · outbound

This paper cites Non-Attentive Tacotron: Robust and Controllable Neural TTS Synthesis Including Unsupervised Duration Modeling.

SpeakStream: Streaming Text-to-Speech with Interleaved Data Non-Attentive Tacotron: Robust and Controllable Neural TTS Synthesis Including Unsupervised Duration Modeling

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T14:22:31.555872Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:22:31.555872Z digest=sha256:ee1a3e2e07f3d7401adce8d6ff96f4f30bfa559f4b79e37c8f35c06a47e7339b

Observation 264a593b-6af6-4e94-a1a6-e63767c50fec · outbound

This paper cites LiveSpeech: Low-Latency Zero-shot Text-to-Speech via Autoregressive Modeling of Audio Discrete Codes.

SpeakStream: Streaming Text-to-Speech with Interleaved Data LiveSpeech: Low-Latency Zero-shot Text-to-Speech via Autoregressive Modeling of Audio Discrete Codes

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T14:22:31.639465Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:22:31.639465Z digest=sha256:326b2ae53ed319375dacec23228beac1025ee6de2193f1f022193132e6b1fe78

Observation b70e1eb8-b561-42d4-bcc2-511129d73975 · outbound

This paper cites Transduce and speak: Neural transducer for text-to-speech with semantic token predic- tion,.

SpeakStream: Streaming Text-to-Speech with Interleaved Data Transduce and speak: Neural transducer for text-to-speech with semantic token predic- tion,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:22:33.190432Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:22:31.674377Z digest=sha256:9608c738294f379187731a64894512ce14acb55bc40a3b510f79d264268e7d63

Observation fcf3c1f7-9d52-4c8b-b2ef-cf855d75e533 · outbound

This paper cites Parallel wavegan: A fast waveform generation model based on generative adversarial networks with multi-resolution spectrogram,.

SpeakStream: Streaming Text-to-Speech with Interleaved Data Parallel wavegan: A fast waveform generation model based on generative adversarial networks with multi-resolution spectrogram,

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T14:22:31.715642Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:22:31.715642Z digest=sha256:f23059c00e0d2f8ea22c9938f69d55ec75dd900704e305cbe82b3778922cd354

Observation 5fc08e52-2a4b-4dfc-88ac-3e787debc790 · outbound

This paper cites Hifi-gan: Generative adversarial networks for efficient and high fidelity speech synthesis,.

SpeakStream: Streaming Text-to-Speech with Interleaved Data Hifi-gan: Generative adversarial networks for efficient and high fidelity speech synthesis,

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T14:22:31.733397Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:22:31.733397Z digest=sha256:3349b93a475de3950d560618f8b1939b46d999665abe8033f911f5d5f0ab72cd

Observation bf912353-3894-4d2b-86e5-1a360d45569b · outbound

This paper cites Bigvgan: A universal neural vocoder with large-scale training,.

SpeakStream: Streaming Text-to-Speech with Interleaved Data Bigvgan: A universal neural vocoder with large-scale training,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:22:33.075120Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:22:31.746547Z digest=sha256:5583fd5f30bbe9859d026fbbeb53f8cfe99c1ae5ddc7e6dc624cde783f4ee8e2

Observation 939b6549-48d9-4fc6-b00e-96a60271dc9c · outbound

This paper cites V ocos: Closing the gap between time-domain and fourier- based neural vocoders for high-quality audio synthesis,.

SpeakStream: Streaming Text-to-Speech with Interleaved Data V ocos: Closing the gap between time-domain and fourier- based neural vocoders for high-quality audio synthesis,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:22:33.020489Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:22:31.759386Z digest=sha256:e6162312413b9156da386a6707330e9bf4dfe2b6bd69d5fc32539afc442f44b1

Observation 027f9df9-6925-4d0d-941b-fc979cc4d52a · outbound

This paper cites Non-causal to causal ssl-supported transfer learning: Towards a high-performance low-latency speech vocoder,.

SpeakStream: Streaming Text-to-Speech with Interleaved Data Non-causal to causal ssl-supported transfer learning: Towards a high-performance low-latency speech vocoder,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:22:32.945758Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:22:31.778855Z digest=sha256:7cb40c52421008a6840b8f08a7e670d458952bd39b2f92145a8c97aac9d897f2

Observation 64393298-7ab5-4cbc-933d-fc9615b7d845 · outbound

This paper cites Coqui TTS: A deep learning toolkit for Text-to-Speech, battle-tested in research and production,.

SpeakStream: Streaming Text-to-Speech with Interleaved Data Coqui TTS: A deep learning toolkit for Text-to-Speech, battle-tested in research and production,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:22:32.906779Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:22:31.787676Z digest=sha256:bfd5db5fb74f1be0f669548c6d9dc900799d774d6d6cd432d4f96ea8b2a80a61

Observation 71b7ac5a-074f-4634-8442-ef81b3ea1d29 · outbound

This paper cites The lj speech dataset,.

SpeakStream: Streaming Text-to-Speech with Interleaved Data The lj speech dataset,

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T14:22:31.807757Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:22:31.807757Z digest=sha256:3d6b54f2bfe30a8146cdf6c6aacaf304d5ba405ac09ef352cb805e3a490541be

Observation 3fd2acba-0c0e-4856-a3d9-9b89e5eb8477 · outbound

This paper cites Whisperx: Time-accurate speech transcription of long-form audio,.

SpeakStream: Streaming Text-to-Speech with Interleaved Data Whisperx: Time-accurate speech transcription of long-form audio,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:22:32.766258Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:22:31.816394Z digest=sha256:58bc38eb572eecd565c5bda414af144a63cf72a7393a79658e49ac02c98073dd

Observation ad6a3e0d-0db1-425d-96dd-3d850d4678a6 · outbound

This paper cites Robust speech recognition via large-scale weak super- vision,.

SpeakStream: Streaming Text-to-Speech with Interleaved Data Robust speech recognition via large-scale weak super- vision,

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T14:22:31.823985Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:22:31.823985Z digest=sha256:1397e4d93f7867bc42cfec7a3c1d15f7192f9e31bbfdf2d28d67eab48cb829cc

Observation 44a264d3-1988-48c0-bc65-f38876f91a97 · outbound

This paper cites Libritts-r: A restored multi-speaker text-to-speech corpus,.

SpeakStream: Streaming Text-to-Speech with Interleaved Data Libritts-r: A restored multi-speaker text-to-speech corpus,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:22:32.674771Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:22:31.845205Z digest=sha256:b31630969d3e8d4e697258e468d8b992c23663d4911c268d8e52ed1d91eac111

Observation f1fbf252-58e1-4706-91c0-32a9c352fb57 · outbound

This paper cites Librispeech: an asr corpus based on public domain audio books,.

SpeakStream: Streaming Text-to-Speech with Interleaved Data Librispeech: an asr corpus based on public domain audio books,

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T14:22:31.853069Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:22:31.853069Z digest=sha256:515bcda716836159a9d360e47a37fb099024b6e3054470a73645f615a8e7237e

Observation 841334ad-98d8-4578-a9aa-8572b38c2b0b · outbound

This paper cites Moshi: a speech-text foundation model for real-time dialogue.

SpeakStream: Streaming Text-to-Speech with Interleaved Data Moshi: a speech-text foundation model for real-time dialogue

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T14:22:31.869618Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:22:31.869618Z digest=sha256:86b04f72eb70f9b436d9a8e949a725118220d5fa8c4465092e1bb3ba095f0126

Observation d2eae670-95c0-493b-8556-5d1d38b0af9b · outbound

This paper cites Available: https://www.coqui.ai.

SpeakStream: Streaming Text-to-Speech with Interleaved Data Available: https://www.coqui.ai

Reference 2021

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:22:32.851116Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:22:31.800042Z digest=sha256:2a45e81d30bcb961e43730000ab474437001d5ae73d60ced5689254c8d028f76

Pith citing papers

Observation 5fea5dc5-c540-4237-bc1d-d1f003ccf490 · inbound

ChipChat: Low-Latency Cascaded Conversational Agent in MLX cites this paper.

ChipChat: Low-Latency Cascaded Conversational Agent in MLX SpeakStream: Streaming Text-to-Speech with Interleaved Data

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-05T15:52:44.054140Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:52:44.054140Z digest=sha256:7b6061dc79c41e02d8785aa40de5cf10100a530bf77e40a087af597bcaff1e66

Observation 6227ce77-3ec2-40de-9470-32c1250b84ee · inbound

Don't Pause: Streaming Video-Language Synchrony for Online Video Understanding cites this paper.

Don't Pause: Streaming Video-Language Synchrony for Online Video Understanding SpeakStream: Streaming Text-to-Speech with Interleaved Data

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-07-02T17:07:12.868675Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-27T22:11:01.690237Z digest=sha256:878d6ddfb1862ded8e61798d80370d8de878f4ac283deddea4c7178c587e3e05