Pith. sign in

Paper Citation Record · LEDGER

SpeakStream: Streaming Text-to-Speech with Interleaved Data

As of 18 August 2026, this Paper Citation Record lists 35 of 35 outbound references and 2 inbound Pith citation observations for arXiv:2505.19206.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.19206 v1

Coverage vector

measured 35 of 35 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:22:31.869618Z

measured 37 of 37 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T15:52:44.054140Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T17:07:12.867079Z

Reference resolution

35 of 35 outbound references displayed

  • verified exact0
  • verified fuzzy14
  • unresolved21
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 5c797376-1075-4f7a-b805-ed8753374b7c · outbound

This paper cites Audi- olm: a language modeling approach to audio generation,.

SpeakStream: Streaming Text-to-Speech with Interleaved Data Audi- olm: a language modeling approach to audio generation,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:22:33.551223Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T14:22:30.401774Z digest=sha256:80f9d4d09903345bb7d706b84197d68a10d8544ff40e25579c8ff4c3f2e7ba92

Observation 2a49efeb-c160-4df2-b2ed-2eb489baef6e · outbound

This paper cites Qwen2.5-Omni Technical Report.

SpeakStream: Streaming Text-to-Speech with Interleaved Data Qwen2.5-Omni Technical Report

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T14:22:30.544240Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:22:30.544240Z digest=sha256:a3464dd308fcfc375d6a89d848d6697f6091b2720409ce7f28f3fecfe39803e4

Observation bc980c59-7ca6-4cfc-9ad1-f31144368a47 · outbound

This paper cites Spirit-lm: Interleaved spoken and written language model,.

SpeakStream: Streaming Text-to-Speech with Interleaved Data Spirit-lm: Interleaved spoken and written language model,

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T14:22:30.595293Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:22:30.595293Z digest=sha256:ca58d2c8ab97b9e05b33485de0eeb3b7e39c901deacc6a459ee539afe090793a

Observation c1bc2262-c193-473e-821f-5069f1f8eb6d · outbound

This paper cites MMAU: A Massive Multi-Task Audio Understanding and Reasoning Benchmark.

SpeakStream: Streaming Text-to-Speech with Interleaved Data MMAU: A Massive Multi-Task Audio Understanding and Reasoning Benchmark

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T14:22:30.673231Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:22:30.673231Z digest=sha256:413b79aea133f2d364d870ecf011fe37fa96e0489d75fdb350b18584b10de7d2

Observation 0b239529-5d94-462f-9b54-9ad986574e10 · outbound

This paper cites Zero-Shot Text-to-Speech from Continuous Text Streams.

SpeakStream: Streaming Text-to-Speech with Interleaved Data Zero-Shot Text-to-Speech from Continuous Text Streams

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T14:22:30.726922Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:22:30.726922Z digest=sha256:933c18160ec7f07457fc983ed8a716f6fda836e94c038a3015d59dbb5cdfd40a

Observation 8928467d-c253-4fba-ac9c-676f4539d528 · outbound

This paper cites Speak while you think: Streaming speech synthesis during text generation,.

SpeakStream: Streaming Text-to-Speech with Interleaved Data Speak while you think: Streaming speech synthesis during text generation,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:22:33.477254Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T14:22:30.780606Z digest=sha256:d121e9a3c56ded25f85d3f8f46df1d0ea16815f85ac661a08d2570f0f2907d06

Observation b18d51a4-4449-433f-b0d8-7f587de7c478 · outbound

This paper cites Interleaved Speech-Text Language Models for Simple Streaming Text-to-Speech Synthesis.

SpeakStream: Streaming Text-to-Speech with Interleaved Data Interleaved Speech-Text Language Models for Simple Streaming Text-to-Speech Synthesis

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T14:22:30.842860Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:22:30.842860Z digest=sha256:3df7de5090bc4bd708701e1259c73a302aa1a3e6d2c0fed81413aea8b9077ac2

Observation c136ad76-0c3f-424d-97f2-a7621681b0e8 · outbound

This paper cites dMel: Speech Tokenization made Simple.

SpeakStream: Streaming Text-to-Speech with Interleaved Data dMel: Speech Tokenization made Simple

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T14:22:30.915758Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:22:30.915758Z digest=sha256:6e42ba9013aafa0bdf747079794a64294f61547c982a2604353e6d87fab09d04

Observation 4d648051-cf1a-46e1-9b8e-67188104c6c8 · outbound

This paper cites A 3T: Alignment-aware acoustic and text pretraining for speech synthesis and editing,.

SpeakStream: Streaming Text-to-Speech with Interleaved Data A 3T: Alignment-aware acoustic and text pretraining for speech synthesis and editing,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:22:33.434512Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T14:22:31.004924Z digest=sha256:cbb3e32260130c8b3fe685af5b8a60159beaf65e345c8f70691e54b35c32a1c7

Observation 9ffb4b08-22c6-47bd-ad27-01f64133f11b · outbound

This paper cites Yourtts: Towards zero-shot multi-speaker tts and zero-shot voice conversion for everyone,.

SpeakStream: Streaming Text-to-Speech with Interleaved Data Yourtts: Towards zero-shot multi-speaker tts and zero-shot voice conversion for everyone,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T14:22:31.029594Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:22:31.029594Z digest=sha256:be4bc17419277473074dfd4aa1cd77a10c794448cfc8a7a1c607e86ede0b391d

Observation 204d923a-49c9-443c-ad21-c47bea9e7cbb · outbound

This paper cites CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens.

SpeakStream: Streaming Text-to-Speech with Interleaved Data CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T14:22:31.087888Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:22:31.087888Z digest=sha256:91e771d3ca31c08736deb8a65de5d72d8a4d2c206ae4bc49c917049e8c25c288

Observation 69c1a5bb-5c73-4bd8-8a8f-4cba12c0e441 · outbound

This paper cites E3 tts: Easy end-to- end diffusion-based text to speech,.

SpeakStream: Streaming Text-to-Speech with Interleaved Data E3 tts: Easy end-to- end diffusion-based text to speech,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:22:33.343081Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T14:22:31.132497Z digest=sha256:b23ece9668c0c682c64f8382be42a61a173aedd4d1330cadf45363331469aaa6

Observation 7fc88f8e-e561-4c13-89a6-35e5a654f82f · outbound

This paper cites FastSpeech 2: Fast and High-Quality End-to-End Text to Speech.

SpeakStream: Streaming Text-to-Speech with Interleaved Data FastSpeech 2: Fast and High-Quality End-to-End Text to Speech

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T14:22:31.169600Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:22:31.169600Z digest=sha256:cd2d4e24846d7f8fec521236ff004b56a9de9bcf83256cc42ee78fa20a1721ff

Observation 6fa8f085-6f5f-4e23-8147-693b3da14f4e · outbound

This paper cites Tacotron: Towards End-to-End Speech Synthesis.

SpeakStream: Streaming Text-to-Speech with Interleaved Data Tacotron: Towards End-to-End Speech Synthesis

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T14:22:31.221168Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:22:31.221168Z digest=sha256:814dad70a30246d92ff2ab8641af61310dfa63d59b7bea4a70b52afc779dbf23

Observation 30b3b16d-7403-4add-a2d3-9a216fec6c90 · outbound

This paper cites (2024) Text-to-speech guide.

SpeakStream: Streaming Text-to-Speech with Interleaved Data (2024) Text-to-speech guide

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:22:33.304761Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T14:22:31.249785Z digest=sha256:d081efb6ba65309081a1e4d97b5e7f8d1b05d32e1bb342d155871f11c7ab9e39

Observation 1efec5e5-675f-4cf5-8bed-cd3d766ca544 · outbound

This paper cites Streamspeech: Low-latency neural architecture for high-quality on-device speech synthesis,.

SpeakStream: Streaming Text-to-Speech with Interleaved Data Streamspeech: Low-latency neural architecture for high-quality on-device speech synthesis,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:22:33.252460Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T14:22:31.282325Z digest=sha256:977eebdfe3acb9734b97f3172618e827db2137078c85d9f8571c01a8e6a3f984

Observation 6209cf8c-2fd9-47ab-b891-3aee83ad4a78 · outbound

This paper cites MLX: Efficient and flexible machine learning on apple silicon,.

SpeakStream: Streaming Text-to-Speech with Interleaved Data MLX: Efficient and flexible machine learning on apple silicon,

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T14:22:31.327328Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:22:31.327328Z digest=sha256:22b7993abbc5a8efbe5862e91fa38721f227a1edcec3d1edad24abc3158ac40b

Observation e261da5d-e116-4fe1-b3ab-5a9103f733a5 · outbound

This paper cites VALL-T: Decoder-Only Generative Transducer for Robust and Decoding-Controllable Text-to-Speech.

SpeakStream: Streaming Text-to-Speech with Interleaved Data VALL-T: Decoder-Only Generative Transducer for Robust and Decoding-Controllable Text-to-Speech

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T14:22:31.416374Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:22:31.416374Z digest=sha256:819fe5e5ffd50982e008787a9abb69182dfb561dba22b93f94eebd4e00db6a0f

Observation 20ab1739-357b-457d-91d3-b16a217f4b12 · outbound

This paper cites BERT: A Review of Applications in Natural Language Processing and Understanding.

SpeakStream: Streaming Text-to-Speech with Interleaved Data BERT: A Review of Applications in Natural Language Processing and Understanding

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T14:22:31.501147Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:22:31.501147Z digest=sha256:4dd5ccb115916e809f9837448c3df6716fd300c8fcb3006dc0da1e55b63f1dd3

Observation d9bcda6d-b7a2-4e38-90e2-dbf56765bdbf · outbound

This paper cites Non-Attentive Tacotron: Robust and Controllable Neural TTS Synthesis Including Unsupervised Duration Modeling.

SpeakStream: Streaming Text-to-Speech with Interleaved Data Non-Attentive Tacotron: Robust and Controllable Neural TTS Synthesis Including Unsupervised Duration Modeling

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T14:22:31.555872Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:22:31.555872Z digest=sha256:b9802a05f7f8cc5543233611de118435f795626f4610acc09cf3c715f3e3a8ce

Observation 264a593b-6af6-4e94-a1a6-e63767c50fec · outbound

This paper cites LiveSpeech: Low-Latency Zero-shot Text-to-Speech via Autoregressive Modeling of Audio Discrete Codes.

SpeakStream: Streaming Text-to-Speech with Interleaved Data LiveSpeech: Low-Latency Zero-shot Text-to-Speech via Autoregressive Modeling of Audio Discrete Codes

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T14:22:31.639465Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:22:31.639465Z digest=sha256:e7f768faf793f9e1712bd070a035cef266f0acad3b7b26559a71f1936e93012c

Observation b70e1eb8-b561-42d4-bcc2-511129d73975 · outbound

This paper cites Transduce and speak: Neural transducer for text-to-speech with semantic token predic- tion,.

SpeakStream: Streaming Text-to-Speech with Interleaved Data Transduce and speak: Neural transducer for text-to-speech with semantic token predic- tion,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:22:33.190432Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T14:22:31.674377Z digest=sha256:786f1bce6efef998173e29e38d396974ffcaf962f251329e95fe95e655300d56

Observation fcf3c1f7-9d52-4c8b-b2ef-cf855d75e533 · outbound

This paper cites Parallel wavegan: A fast waveform generation model based on generative adversarial networks with multi-resolution spectrogram,.

SpeakStream: Streaming Text-to-Speech with Interleaved Data Parallel wavegan: A fast waveform generation model based on generative adversarial networks with multi-resolution spectrogram,

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T14:22:31.715642Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:22:31.715642Z digest=sha256:c70e832df1b75dfed6531f7260cea852ddd1f081cd3ac182d1c0e4ade398c1e7

Observation 5fc08e52-2a4b-4dfc-88ac-3e787debc790 · outbound

This paper cites Hifi-gan: Generative adversarial networks for efficient and high fidelity speech synthesis,.

SpeakStream: Streaming Text-to-Speech with Interleaved Data Hifi-gan: Generative adversarial networks for efficient and high fidelity speech synthesis,

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T14:22:31.733397Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:22:31.733397Z digest=sha256:b83bbddd3d6151663b6913e1b53eb2dcc83bbec7ae70e9f28d76ffab555a5925

Observation bf912353-3894-4d2b-86e5-1a360d45569b · outbound

This paper cites Bigvgan: A universal neural vocoder with large-scale training,.

SpeakStream: Streaming Text-to-Speech with Interleaved Data Bigvgan: A universal neural vocoder with large-scale training,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:22:33.075120Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T14:22:31.746547Z digest=sha256:f6f4bbd0f72cc3d9516cba501f57812512a1fb7bfb15f702e082522718c92928

Observation 939b6549-48d9-4fc6-b00e-96a60271dc9c · outbound

This paper cites V ocos: Closing the gap between time-domain and fourier- based neural vocoders for high-quality audio synthesis,.

SpeakStream: Streaming Text-to-Speech with Interleaved Data V ocos: Closing the gap between time-domain and fourier- based neural vocoders for high-quality audio synthesis,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:22:33.020489Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T14:22:31.759386Z digest=sha256:ad0a567d7a90922bd5bf5bbf596d8584ef1593ddf2bbf5d3d757e5b6f80ade85

Observation 027f9df9-6925-4d0d-941b-fc979cc4d52a · outbound

This paper cites Non-causal to causal ssl-supported transfer learning: Towards a high-performance low-latency speech vocoder,.

SpeakStream: Streaming Text-to-Speech with Interleaved Data Non-causal to causal ssl-supported transfer learning: Towards a high-performance low-latency speech vocoder,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:22:32.945758Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T14:22:31.778855Z digest=sha256:86f62a44adacb2e3a091d325a3396a09aec0aaed589b0d176c74563df7735068

Observation 64393298-7ab5-4cbc-933d-fc9615b7d845 · outbound

This paper cites Coqui TTS: A deep learning toolkit for Text-to-Speech, battle-tested in research and production,.

SpeakStream: Streaming Text-to-Speech with Interleaved Data Coqui TTS: A deep learning toolkit for Text-to-Speech, battle-tested in research and production,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:22:32.906779Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T14:22:31.787676Z digest=sha256:4922308a1115dbab478b89d271a90fa8a68ef62748625babe251cf96deb7ca4c

Observation 71b7ac5a-074f-4634-8442-ef81b3ea1d29 · outbound

This paper cites The lj speech dataset,.

SpeakStream: Streaming Text-to-Speech with Interleaved Data The lj speech dataset,

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T14:22:31.807757Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:22:31.807757Z digest=sha256:cc33b41190df0e62fc3d8913177008f3d20008fd43b38a39ce1123d4ff073662

Observation 3fd2acba-0c0e-4856-a3d9-9b89e5eb8477 · outbound

This paper cites Whisperx: Time-accurate speech transcription of long-form audio,.

SpeakStream: Streaming Text-to-Speech with Interleaved Data Whisperx: Time-accurate speech transcription of long-form audio,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:22:32.766258Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T14:22:31.816394Z digest=sha256:95930915f741ce7a8ebe20977d002e16fcd190bb6716621ca5aa26c9f493b2da

Observation ad6a3e0d-0db1-425d-96dd-3d850d4678a6 · outbound

This paper cites Robust speech recognition via large-scale weak super- vision,.

SpeakStream: Streaming Text-to-Speech with Interleaved Data Robust speech recognition via large-scale weak super- vision,

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T14:22:31.823985Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:22:31.823985Z digest=sha256:17e9df955e048f441108b89804a6d11cfd8ccbe8b236229765e602d197f49303

Observation 44a264d3-1988-48c0-bc65-f38876f91a97 · outbound

This paper cites Libritts-r: A restored multi-speaker text-to-speech corpus,.

SpeakStream: Streaming Text-to-Speech with Interleaved Data Libritts-r: A restored multi-speaker text-to-speech corpus,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:22:32.674771Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T14:22:31.845205Z digest=sha256:3cbb3c0884e7e9bebef8675dc53a89a73a1419c2f1a8a42649fd67e77189a6c2

Observation f1fbf252-58e1-4706-91c0-32a9c352fb57 · outbound

This paper cites Librispeech: an asr corpus based on public domain audio books,.

SpeakStream: Streaming Text-to-Speech with Interleaved Data Librispeech: an asr corpus based on public domain audio books,

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T14:22:31.853069Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:22:31.853069Z digest=sha256:3736ef1ee9c1c59f30b87992e15cbc2e40c07264f9d49d83ca6d3f4fdbff3137

Observation 841334ad-98d8-4578-a9aa-8572b38c2b0b · outbound

This paper cites Moshi: a speech-text foundation model for real-time dialogue.

SpeakStream: Streaming Text-to-Speech with Interleaved Data Moshi: a speech-text foundation model for real-time dialogue

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T14:22:31.869618Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:22:31.869618Z digest=sha256:0b164f6598993680ab80bcb5799bd9a42d4c9fd7180c13e6afa6b647995b6f55

Observation d2eae670-95c0-493b-8556-5d1d38b0af9b · outbound

This paper cites Available: https://www.coqui.ai.

SpeakStream: Streaming Text-to-Speech with Interleaved Data Available: https://www.coqui.ai

Reference 2021

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:22:32.851116Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T14:22:31.800042Z digest=sha256:faf1458411e6eb91204b5c3fc60dce5605386553c6adccec305957ec72232c37

Pith citing papers

Observation 5fea5dc5-c540-4237-bc1d-d1f003ccf490 · inbound

ChipChat: Low-Latency Cascaded Conversational Agent in MLX cites this paper.

ChipChat: Low-Latency Cascaded Conversational Agent in MLX SpeakStream: Streaming Text-to-Speech with Interleaved Data

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-05T15:52:44.054140Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:52:44.054140Z digest=sha256:6d44b38ffb4763a7315edcdc32a7d044756d0f1a555a0866f032aed8c6fa1f20

Observation 6227ce77-3ec2-40de-9470-32c1250b84ee · inbound

Don't Pause: Streaming Video-Language Synchrony for Online Video Understanding cites this paper.

Don't Pause: Streaming Video-Language Synchrony for Online Video Understanding SpeakStream: Streaming Text-to-Speech with Interleaved Data

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-07-02T17:07:12.868675Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-27T22:11:01.690237Z digest=sha256:bb5a028f375f8fbf8611196a958c449f59bbdf68accbd26edf1553784780852f