Pith. sign in

Paper Citation Record · LEDGER

StreamMel: Real-Time Zero-shot Text-to-Speech via Interleaved Continuous Autoregressive Modeling

As of 7 August 2026, this Paper Citation Record lists 35 of 35 outbound references and 0 inbound Pith citation observations for arXiv:2506.12570.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.12570 v1

Coverage vector

measured 35 of 35 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T00:52:04.191022Z

measured 35 of 35 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

35 of 35 outbound references displayed

  • verified exact0
  • verified fuzzy23
  • unresolved11
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation bf49560b-b22b-4a55-bc8b-a52da9d87aef · outbound

This paper cites GPT-4 Technical Report.

StreamMel: Real-Time Zero-shot Text-to-Speech via Interleaved Continuous Autoregressive Modeling GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T00:52:01.571286Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:52:01.571286Z digest=sha256:3bb2eaf8912687760a63f89a4e7e94213b9fc494b59e7b121e22dd03a11a9400

Observation bca8a2f3-7387-45e2-9b6f-e2b86e6e457e · outbound

This paper cites Audiolm: A language modeling approach to audio generation,.

StreamMel: Real-Time Zero-shot Text-to-Speech via Interleaved Continuous Autoregressive Modeling Audiolm: A language modeling approach to audio generation,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:52:07.859576Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:52:01.613360Z digest=sha256:fd0778f335922a3a12f6ca0502881ab837e952a76c4832de2690d475a47db366

Observation f2786bbc-8f4c-4d59-89ae-3d62f6eb06f4 · outbound

This paper cites Flow matching for generative modeling,.

StreamMel: Real-Time Zero-shot Text-to-Speech via Interleaved Continuous Autoregressive Modeling Flow matching for generative modeling,

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T00:52:01.712370Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:52:01.712370Z digest=sha256:aecfc77de93415481b7ad3e6666b3236398d14fadc6877a0f080fbc46b9deacd

Observation 09327ca5-7b3b-4ede-a6c2-08250ab64b36 · outbound

This paper cites Denoising diffusion probabilistic models,.

StreamMel: Real-Time Zero-shot Text-to-Speech via Interleaved Continuous Autoregressive Modeling Denoising diffusion probabilistic models,

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T00:52:01.782086Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:52:01.782086Z digest=sha256:9723e25ea977885755c82dd11bd302b444e9e54ab5c482fd90ba6996e353d7ca

Observation b570ac72-c91a-4522-928a-e9ceca604e5f · outbound

This paper cites Neural codec language models are zero-shot text to speech synthesizers,.

StreamMel: Real-Time Zero-shot Text-to-Speech via Interleaved Continuous Autoregressive Modeling Neural codec language models are zero-shot text to speech synthesizers,

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T00:52:01.901321Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:52:01.901321Z digest=sha256:fdf425e7541a1c507966dcd967ae15cfc22656c2d2bc7edf95c6f256dd835b1e

Observation 18b91231-a431-4427-b361-ebbc14d04113 · outbound

This paper cites V ALL-E 2: Neural codec language models are human parity zero-shot text to speech synthesizers,.

StreamMel: Real-Time Zero-shot Text-to-Speech via Interleaved Continuous Autoregressive Modeling V ALL-E 2: Neural codec language models are human parity zero-shot text to speech synthesizers,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:52:07.616899Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:52:01.985281Z digest=sha256:d2ab6df6f8682b6bb92e3e0d29f7daa17da87b59f0f12d2ab08baff5e52e59a0

Observation 15a4defa-8156-4b3a-902d-49cd85587af8 · outbound

This paper cites CosyV oice: A scalable multilingual zero-shot text-to-speech synthesizer based on supervised semantic to- kens,.

StreamMel: Real-Time Zero-shot Text-to-Speech via Interleaved Continuous Autoregressive Modeling CosyV oice: A scalable multilingual zero-shot text-to-speech synthesizer based on supervised semantic to- kens,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:52:07.417375Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:52:02.050554Z digest=sha256:7a21ddb38491ffc73d758b27a5441e7fb35ad57927a86ee1ef16c7468983b41a

Observation b9089ddd-2bd8-46f9-94c8-a6a0b10a4ec5 · outbound

This paper cites Seed-TTS: A family of high- quality versatile speech generation models,.

StreamMel: Real-Time Zero-shot Text-to-Speech via Interleaved Continuous Autoregressive Modeling Seed-TTS: A family of high- quality versatile speech generation models,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:52:07.317431Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:52:02.180105Z digest=sha256:9801ffbaf7e9c3284626cc13680a13f06d92fedc40e9ab95197025e8adec02e9

Observation 8bfc3e62-250f-40b5-b08d-a9fb8af1b6eb · outbound

This paper cites CosyVoice 3: Towards In-the-wild Speech Generation via Scaling-up and Post-training.

StreamMel: Real-Time Zero-shot Text-to-Speech via Interleaved Continuous Autoregressive Modeling CosyVoice 3: Towards In-the-wild Speech Generation via Scaling-up and Post-training

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T00:52:02.203070Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:52:02.203070Z digest=sha256:d279f4c24807a2bd83d9ce77a77c3d9a037f005d6ae761a600fc802a4f000e4f

Observation 7a4cf7c7-35da-4e8f-be7a-fe8607cfd06e · outbound

This paper cites Pseudo-autoregressive neural codec language models for efficient zero-shot text-to-speech synthesis,.

StreamMel: Real-Time Zero-shot Text-to-Speech via Interleaved Continuous Autoregressive Modeling Pseudo-autoregressive neural codec language models for efficient zero-shot text-to-speech synthesis,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:52:07.129504Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:52:02.264484Z digest=sha256:d00c19af6de0646dd839122aa8ff056ae095fd09ef5036b28c89af6dd0330c0a

Observation 604d4ff9-0d40-492c-a316-7f8b8ab3404e · outbound

This paper cites Autoregressive speech synthesis without vector quantization,.

StreamMel: Real-Time Zero-shot Text-to-Speech via Interleaved Continuous Autoregressive Modeling Autoregressive speech synthesis without vector quantization,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:52:06.990494Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:52:02.305103Z digest=sha256:620bdc8994da0cdeaef54f986b3f2a49a498361f9c7b9f5816e549974ec3df50

Observation d6f696a4-cffe-4f11-ae11-587f4073529b · outbound

This paper cites Cosyvoice 2: Scalable streaming speech synthesis with large language models,.

StreamMel: Real-Time Zero-shot Text-to-Speech via Interleaved Continuous Autoregressive Modeling Cosyvoice 2: Scalable streaming speech synthesis with large language models,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:52:06.879628Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:52:02.351297Z digest=sha256:c778e2002b03b3a64e1ed70a5b8bd5637bc0c68a8a30444db2c24d1d14de7ae4

Observation 412ef6ee-0c2f-4a1f-b7b2-7d5c4a68b520 · outbound

This paper cites Interleaved speech-text language models are simple streaming text to speech synthesizers,.

StreamMel: Real-Time Zero-shot Text-to-Speech via Interleaved Continuous Autoregressive Modeling Interleaved speech-text language models are simple streaming text to speech synthesizers,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:52:06.747229Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:52:02.392778Z digest=sha256:a534f4b9dd17f6c1ca5b8ae676688965cd8382a3d8945c6bd4979c5878890206

Observation ad5cbb77-9e8e-4f73-ad1a-344020d4bdf4 · outbound

This paper cites Zero-Shot Streaming Text to Speech Synthesis with Transducer and Auto-Regressive Modeling.

StreamMel: Real-Time Zero-shot Text-to-Speech via Interleaved Continuous Autoregressive Modeling Zero-Shot Streaming Text to Speech Synthesis with Transducer and Auto-Regressive Modeling

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T00:52:02.492729Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:52:02.492729Z digest=sha256:47fcedbf7e6c1f9423f6acf11d9a9f0ef226f76c41904341b92929119a5369cd

Observation 254d6ae7-804d-420e-bdcc-a5b0db8c718d · outbound

This paper cites Syncspeech: Low-latency and efficient dual-stream text-to-speech based on temporal masked transformer,.

StreamMel: Real-Time Zero-shot Text-to-Speech via Interleaved Continuous Autoregressive Modeling Syncspeech: Low-latency and efficient dual-stream text-to-speech based on temporal masked transformer,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:52:06.604284Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:52:02.548264Z digest=sha256:a4e94b6aabe19206a75dcc296366003d812e32cd4396a699b7e3c548ee91272b

Observation 8d3e0bf1-4ff6-4c04-8ce3-36988ce854e1 · outbound

This paper cites Discrete audio representation as an alternative to mel-spectrograms for speaker and speech recognition,.

StreamMel: Real-Time Zero-shot Text-to-Speech via Interleaved Continuous Autoregressive Modeling Discrete audio representation as an alternative to mel-spectrograms for speaker and speech recognition,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:52:06.522611Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:52:02.596251Z digest=sha256:3af8e64d09235c11caa585295900eecee00e34a33a3e22b8049c3056e46901f3

Observation 72d6443a-abea-45aa-87b7-dbf0dd6e7ea1 · outbound

This paper cites Autoregressive diffusion transformer for text-to-speech synthesis,.

StreamMel: Real-Time Zero-shot Text-to-Speech via Interleaved Continuous Autoregressive Modeling Autoregressive diffusion transformer for text-to-speech synthesis,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:52:06.344259Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:52:02.665547Z digest=sha256:f3cb7850f0621e430f40b2766241ab241d867efa0ae7a52510e33f09a12d68df

Observation 7b78f69d-f61f-4bc9-87f5-6da2c7246e5b · outbound

This paper cites Autoregressive Image Generation without Vector Quantization.

StreamMel: Real-Time Zero-shot Text-to-Speech via Interleaved Continuous Autoregressive Modeling Autoregressive Image Generation without Vector Quantization

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T00:52:02.719163Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:52:02.719163Z digest=sha256:a5b6d8b8bbc597d5491acc497ea8c286ed760b7b6817c6b761f8393180a8be2f

Observation 67d72c9a-6b5f-4763-8b44-4b5490c36bdc · outbound

This paper cites FELLE: autoregressive speech synthesis with token-wise coarse-to-fine flow matching,.

StreamMel: Real-Time Zero-shot Text-to-Speech via Interleaved Continuous Autoregressive Modeling FELLE: autoregressive speech synthesis with token-wise coarse-to-fine flow matching,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:52:06.224882Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:52:02.782039Z digest=sha256:33f8f82639c560902cdb20d9e3211b49d1c813c00359621679a31fada1ce62db

Observation 4c47da4f-c6a1-4481-95c1-2ef253ecaca8 · outbound

This paper cites Ditar: Diffusion transformer autoregressive modeling for speech generation,.

StreamMel: Real-Time Zero-shot Text-to-Speech via Interleaved Continuous Autoregressive Modeling Ditar: Diffusion transformer autoregressive modeling for speech generation,

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T00:52:02.873718Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:52:02.873718Z digest=sha256:23c307c957e48dc2750c79a64a78cbfc29b173c6e26bc6bd5f686be9514ebc99

Observation 36ccc01f-48dc-49d0-91d7-a45a98b8938b · outbound

This paper cites Auto-encoding variational bayes,.

StreamMel: Real-Time Zero-shot Text-to-Speech via Interleaved Continuous Autoregressive Modeling Auto-encoding variational bayes,

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T00:52:02.939608Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:52:02.939608Z digest=sha256:aaf603c32ded42e774da84865a2cdffb9428f35a0c3e59a7794993aa09084b0f

Observation 7082f2ce-2d21-48db-abde-1422f8f5d842 · outbound

This paper cites Tacotron: Towards end- to-end speech synthesis,.

StreamMel: Real-Time Zero-shot Text-to-Speech via Interleaved Continuous Autoregressive Modeling Tacotron: Towards end- to-end speech synthesis,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:52:06.061420Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:52:03.055590Z digest=sha256:a525f6f8653c1e4507978806dddfc154165797833d83ba91c29cb07b74a5e53d

Observation 99733192-e98e-40c2-9b61-e859f987b52d · outbound

This paper cites Librispeech: an ASR corpus based on public domain audio books,.

StreamMel: Real-Time Zero-shot Text-to-Speech via Interleaved Continuous Autoregressive Modeling Librispeech: an ASR corpus based on public domain audio books,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:52:05.895076Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:52:03.132846Z digest=sha256:9a1b5b77ed076aa08ea9f05cc12a9729fd2b20ebb3674aab2cf66d866d2fe319

Observation 1e041d1c-da61-40fa-b08a-35a0d788755a · outbound

This paper cites V ALL-E R: robust and efficient zero-shot text-to-speech synthesis via monotonic alignment,.

StreamMel: Real-Time Zero-shot Text-to-Speech via Interleaved Continuous Autoregressive Modeling V ALL-E R: robust and efficient zero-shot text-to-speech synthesis via monotonic alignment,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:52:05.762832Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:52:03.223595Z digest=sha256:6288036d695e2f6dbabe852f887a8f452623fd8cce0619685b24fce6427a0564

Observation 6d419976-a6ed-4cf3-9dd7-97606372ab49 · outbound

This paper cites F5-TTS: A fairytaler that fakes fluent and faithful speech with flow matching,.

StreamMel: Real-Time Zero-shot Text-to-Speech via Interleaved Continuous Autoregressive Modeling F5-TTS: A fairytaler that fakes fluent and faithful speech with flow matching,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:52:05.621674Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:52:03.285128Z digest=sha256:459b851c399d7b67775386e9daadd4064f52fcee03917e7557fea2a004abb52b

Observation f6a18580-b41c-42af-b442-e71d1614e80c · outbound

This paper cites MaskGCT: Zero-shot text-to-speech with masked generative codec transformer,.

StreamMel: Real-Time Zero-shot Text-to-Speech via Interleaved Continuous Autoregressive Modeling MaskGCT: Zero-shot text-to-speech with masked generative codec transformer,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:52:05.442194Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:52:03.385295Z digest=sha256:cf6e0bcfd35d0203c777ac1a1a25bae26e1d2f659e742dd8543948f6a20b121f

Observation 7afb6a5c-1d16-46c5-9b28-d3f2c1bd98c9 · outbound

This paper cites Ramp: Retrieval-augmented mos prediction via confidence-based dynamic weighting,.

StreamMel: Real-Time Zero-shot Text-to-Speech via Interleaved Continuous Autoregressive Modeling Ramp: Retrieval-augmented mos prediction via confidence-based dynamic weighting,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:52:05.327200Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:52:03.475605Z digest=sha256:5c42649fa93885e0021e3cbfbef14d112a4173f1d212b78b3adff8b7c8e8d550

Observation 462c7a9a-7137-421b-829c-2e2aaa8cb3ec · outbound

This paper cites Ramp+: Retrieval-augmented mos prediction with prior knowledge integration,.

StreamMel: Real-Time Zero-shot Text-to-Speech via Interleaved Continuous Autoregressive Modeling Ramp+: Retrieval-augmented mos prediction with prior knowledge integration,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:52:05.232775Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:52:03.590269Z digest=sha256:eee196dde51eea6e6337a084efa7b96d155b6802bbc2fa822ecb69b2745f8147

Observation 7d1701fd-bfef-49fa-abcc-e7495a5d0181 · outbound

This paper cites Conformer: Convolution-augmented Transformer for Speech Recognition.

StreamMel: Real-Time Zero-shot Text-to-Speech via Interleaved Continuous Autoregressive Modeling Conformer: Convolution-augmented Transformer for Speech Recognition

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T00:52:03.732161Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:52:03.732161Z digest=sha256:241b59bc5c36ad917a58586181b9250613fb913bb241901139c64d81a4b8035c

Observation 204c52a3-173f-43c0-a320-2e4a4bc4181b · outbound

This paper cites HuBERT: Self-supervised speech representation learning by masked prediction of hidden units,.

StreamMel: Real-Time Zero-shot Text-to-Speech via Interleaved Continuous Autoregressive Modeling HuBERT: Self-supervised speech representation learning by masked prediction of hidden units,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:52:05.107680Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:52:03.807837Z digest=sha256:37d3ad1d314857a112a995b9de42e21af7da082a73d96bf3861c4fe4d7a16f2d

Observation ef7dbf33-7411-4e07-87b9-bc4450a22762 · outbound

This paper cites Robust Speech Recognition via Large-Scale Weak Supervision.

StreamMel: Real-Time Zero-shot Text-to-Speech via Interleaved Continuous Autoregressive Modeling Robust Speech Recognition via Large-Scale Weak Supervision

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T00:52:03.933734Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:52:03.933734Z digest=sha256:dc588dfc03b45bee732af14550f951b5f6d8756a391cbf92fbcbbe66e48f4c9f

Observation 50132868-6da2-4174-9ba8-8a8d3b9a67df · outbound

This paper cites WavLM: Large-scale self-supervised pre-training for full stack speech processing,.

StreamMel: Real-Time Zero-shot Text-to-Speech via Interleaved Continuous Autoregressive Modeling WavLM: Large-scale self-supervised pre-training for full stack speech processing,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:52:05.018828Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:52:04.039029Z digest=sha256:27d184100742c4f109e581b8932c3cea7383cbd0f2932f4b4b051fb4a5dc571f

Observation 63c65484-ae4f-4ce4-80a1-aa28b396b6c3 · outbound

This paper cites Libriheavy: a 50,000 hours ASR corpus with punctuation casing and context,.

StreamMel: Real-Time Zero-shot Text-to-Speech via Interleaved Continuous Autoregressive Modeling Libriheavy: a 50,000 hours ASR corpus with punctuation casing and context,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:52:04.863237Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:52:04.109157Z digest=sha256:9d6cf7b4679ad8a7b5d42598d3655495134fa8609271936c1d20fc7b5fa7105f

Observation 68ed7cca-35c7-4bf2-bf63-e61243bf22fc · outbound

This paper cites Emilia: An extensive, multilingual, and diverse speech dataset for large-scale speech generation,.

StreamMel: Real-Time Zero-shot Text-to-Speech via Interleaved Continuous Autoregressive Modeling Emilia: An extensive, multilingual, and diverse speech dataset for large-scale speech generation,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:52:04.682115Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:52:04.191022Z digest=sha256:42a3dbe7167ef970543a3f56f1b1091ffda6662397e38fb8653db010c3c4c5e0

Observation 2f204fc6-c3d9-4824-bc3b-a67581bde4cc · outbound

This paper cites Interleaved Speech-Text Language Models for Simple Streaming Text-to-Speech Synthesis.

StreamMel: Real-Time Zero-shot Text-to-Speech via Interleaved Continuous Autoregressive Modeling Interleaved Speech-Text Language Models for Simple Streaming Text-to-Speech Synthesis

Reference 2024

Resolution
metadata mismatch
local_arxiv, observed 2026-08-07T00:52:04.503082Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T00:52:02.441261Z digest=sha256:d15717dc1e5b1b81ee56d7ad2e57daab43b0bac6196e1c87650ef8b11c70ca63

Pith citing papers

No inbound Pith citation observations are available.