Pith. sign in

Paper Citation Record · LEDGER

StreamMel: Real-Time Zero-shot Text-to-Speech via Interleaved Continuous Autoregressive Modeling

As of 21 August 2026, this Paper Citation Record lists 35 of 35 outbound references and 0 inbound Pith citation observations for arXiv:2506.12570.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.12570 v1

Coverage vector

measured 35 of 35 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T00:52:04.191022Z

measured 35 of 35 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

35 of 35 outbound references displayed

  • verified exact0
  • verified fuzzy23
  • unresolved11
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation bf49560b-b22b-4a55-bc8b-a52da9d87aef · outbound

This paper cites GPT-4 Technical Report.

StreamMel: Real-Time Zero-shot Text-to-Speech via Interleaved Continuous Autoregressive Modeling GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T00:52:01.571286Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:52:01.571286Z digest=sha256:7afed365c4347a30a8b0751da3ec31c46d825675283e13993af8029fb13e9799

Observation bca8a2f3-7387-45e2-9b6f-e2b86e6e457e · outbound

This paper cites Audiolm: A language modeling approach to audio generation,.

StreamMel: Real-Time Zero-shot Text-to-Speech via Interleaved Continuous Autoregressive Modeling Audiolm: A language modeling approach to audio generation,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:52:07.859576Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T00:52:01.613360Z digest=sha256:912c6a97059e266949becf9a01b9457e31c1e661811a1e2e6a321f213b2840f9

Observation f2786bbc-8f4c-4d59-89ae-3d62f6eb06f4 · outbound

This paper cites Flow matching for generative modeling,.

StreamMel: Real-Time Zero-shot Text-to-Speech via Interleaved Continuous Autoregressive Modeling Flow matching for generative modeling,

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T00:52:01.712370Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:52:01.712370Z digest=sha256:44e0c7074f802b290a690e8edaca8e3e0e0b1c5d332012434a639effbe7b7d30

Observation 09327ca5-7b3b-4ede-a6c2-08250ab64b36 · outbound

This paper cites Denoising diffusion probabilistic models,.

StreamMel: Real-Time Zero-shot Text-to-Speech via Interleaved Continuous Autoregressive Modeling Denoising diffusion probabilistic models,

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T00:52:01.782086Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:52:01.782086Z digest=sha256:dbf2f8185e97b82eef321df75cc682c73ac939f3ed1b741ed5fbc4faba39a60a

Observation b570ac72-c91a-4522-928a-e9ceca604e5f · outbound

This paper cites Neural codec language models are zero-shot text to speech synthesizers,.

StreamMel: Real-Time Zero-shot Text-to-Speech via Interleaved Continuous Autoregressive Modeling Neural codec language models are zero-shot text to speech synthesizers,

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T00:52:01.901321Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:52:01.901321Z digest=sha256:a1a043d38468dabb72491c657fe2faa91c94644b09ffb4338cfc9e5646d818dd

Observation 18b91231-a431-4427-b361-ebbc14d04113 · outbound

This paper cites V ALL-E 2: Neural codec language models are human parity zero-shot text to speech synthesizers,.

StreamMel: Real-Time Zero-shot Text-to-Speech via Interleaved Continuous Autoregressive Modeling V ALL-E 2: Neural codec language models are human parity zero-shot text to speech synthesizers,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:52:07.616899Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T00:52:01.985281Z digest=sha256:157372bf39e4cbbad7564af40707bc38eccafd0079438330316047316129c556

Observation 15a4defa-8156-4b3a-902d-49cd85587af8 · outbound

This paper cites CosyV oice: A scalable multilingual zero-shot text-to-speech synthesizer based on supervised semantic to- kens,.

StreamMel: Real-Time Zero-shot Text-to-Speech via Interleaved Continuous Autoregressive Modeling CosyV oice: A scalable multilingual zero-shot text-to-speech synthesizer based on supervised semantic to- kens,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:52:07.417375Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T00:52:02.050554Z digest=sha256:ae5bbc3197ca25b30b71b511fdb8427f0d7b94f846914685637056ce429cf7e9

Observation b9089ddd-2bd8-46f9-94c8-a6a0b10a4ec5 · outbound

This paper cites Seed-TTS: A family of high- quality versatile speech generation models,.

StreamMel: Real-Time Zero-shot Text-to-Speech via Interleaved Continuous Autoregressive Modeling Seed-TTS: A family of high- quality versatile speech generation models,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:52:07.317431Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T00:52:02.180105Z digest=sha256:354420c8171816ecaac8f346590bd59ed772cdc8079067a91a29614f83d772f7

Observation 8bfc3e62-250f-40b5-b08d-a9fb8af1b6eb · outbound

This paper cites CosyVoice 3: Towards In-the-wild Speech Generation via Scaling-up and Post-training.

StreamMel: Real-Time Zero-shot Text-to-Speech via Interleaved Continuous Autoregressive Modeling CosyVoice 3: Towards In-the-wild Speech Generation via Scaling-up and Post-training

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T00:52:02.203070Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:52:02.203070Z digest=sha256:bd7f740d069976ad9c1ddc12da7d7db5cc07b939056ddd5a08b562a595d630b7

Observation 7a4cf7c7-35da-4e8f-be7a-fe8607cfd06e · outbound

This paper cites Pseudo-autoregressive neural codec language models for efficient zero-shot text-to-speech synthesis,.

StreamMel: Real-Time Zero-shot Text-to-Speech via Interleaved Continuous Autoregressive Modeling Pseudo-autoregressive neural codec language models for efficient zero-shot text-to-speech synthesis,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:52:07.129504Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T00:52:02.264484Z digest=sha256:5ac62c228a94abb876b741fe6f41a04c96db60d93252791de44d0b1d0cc169c3

Observation 604d4ff9-0d40-492c-a316-7f8b8ab3404e · outbound

This paper cites Autoregressive speech synthesis without vector quantization,.

StreamMel: Real-Time Zero-shot Text-to-Speech via Interleaved Continuous Autoregressive Modeling Autoregressive speech synthesis without vector quantization,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:52:06.990494Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T00:52:02.305103Z digest=sha256:4e9f8ce7ac83d86a206a9af46b7311742ed4cf1e3a2f276835573672ce82b518

Observation d6f696a4-cffe-4f11-ae11-587f4073529b · outbound

This paper cites Cosyvoice 2: Scalable streaming speech synthesis with large language models,.

StreamMel: Real-Time Zero-shot Text-to-Speech via Interleaved Continuous Autoregressive Modeling Cosyvoice 2: Scalable streaming speech synthesis with large language models,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:52:06.879628Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T00:52:02.351297Z digest=sha256:00cba50c5074d0edb2214268729fb9a6335eb0506773f7d94f98bad84127ccc8

Observation 412ef6ee-0c2f-4a1f-b7b2-7d5c4a68b520 · outbound

This paper cites Interleaved speech-text language models are simple streaming text to speech synthesizers,.

StreamMel: Real-Time Zero-shot Text-to-Speech via Interleaved Continuous Autoregressive Modeling Interleaved speech-text language models are simple streaming text to speech synthesizers,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:52:06.747229Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T00:52:02.392778Z digest=sha256:80e611622081df34f616285c54125989759b91e9f0011b8dba3a202e5be097b7

Observation ad5cbb77-9e8e-4f73-ad1a-344020d4bdf4 · outbound

This paper cites Zero-Shot Streaming Text to Speech Synthesis with Transducer and Auto-Regressive Modeling.

StreamMel: Real-Time Zero-shot Text-to-Speech via Interleaved Continuous Autoregressive Modeling Zero-Shot Streaming Text to Speech Synthesis with Transducer and Auto-Regressive Modeling

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T00:52:02.492729Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:52:02.492729Z digest=sha256:a4c661632ccbdfc482d5862cfa6d121db5092620501c00c796cf305034a5787c

Observation 254d6ae7-804d-420e-bdcc-a5b0db8c718d · outbound

This paper cites Syncspeech: Low-latency and efficient dual-stream text-to-speech based on temporal masked transformer,.

StreamMel: Real-Time Zero-shot Text-to-Speech via Interleaved Continuous Autoregressive Modeling Syncspeech: Low-latency and efficient dual-stream text-to-speech based on temporal masked transformer,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:52:06.604284Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T00:52:02.548264Z digest=sha256:9a9c535578e0f0069c2facf5b7142552daa2020d7c0b69b3003253f14dde1ce3

Observation 8d3e0bf1-4ff6-4c04-8ce3-36988ce854e1 · outbound

This paper cites Discrete audio representation as an alternative to mel-spectrograms for speaker and speech recognition,.

StreamMel: Real-Time Zero-shot Text-to-Speech via Interleaved Continuous Autoregressive Modeling Discrete audio representation as an alternative to mel-spectrograms for speaker and speech recognition,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:52:06.522611Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T00:52:02.596251Z digest=sha256:1498f29cfce3ab7f5a873b0e14933281bb7b81299d704495cded0a364990deb1

Observation 72d6443a-abea-45aa-87b7-dbf0dd6e7ea1 · outbound

This paper cites Autoregressive diffusion transformer for text-to-speech synthesis,.

StreamMel: Real-Time Zero-shot Text-to-Speech via Interleaved Continuous Autoregressive Modeling Autoregressive diffusion transformer for text-to-speech synthesis,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:52:06.344259Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T00:52:02.665547Z digest=sha256:8c09fc695a0921c4019cf046aa8a8b10448e350d74f580871619c3262dc4b6e8

Observation 7b78f69d-f61f-4bc9-87f5-6da2c7246e5b · outbound

This paper cites Autoregressive Image Generation without Vector Quantization.

StreamMel: Real-Time Zero-shot Text-to-Speech via Interleaved Continuous Autoregressive Modeling Autoregressive Image Generation without Vector Quantization

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T00:52:02.719163Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:52:02.719163Z digest=sha256:ef7fefe30031aedb421ef2f1cf847c643527af664867a232c0e6049dd8b8b635

Observation 67d72c9a-6b5f-4763-8b44-4b5490c36bdc · outbound

This paper cites FELLE: autoregressive speech synthesis with token-wise coarse-to-fine flow matching,.

StreamMel: Real-Time Zero-shot Text-to-Speech via Interleaved Continuous Autoregressive Modeling FELLE: autoregressive speech synthesis with token-wise coarse-to-fine flow matching,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:52:06.224882Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T00:52:02.782039Z digest=sha256:d949abc94baee734e58675f151d4f0f3329f2e474c5fe08d697b33cf79dc24b4

Observation 4c47da4f-c6a1-4481-95c1-2ef253ecaca8 · outbound

This paper cites Ditar: Diffusion transformer autoregressive modeling for speech generation,.

StreamMel: Real-Time Zero-shot Text-to-Speech via Interleaved Continuous Autoregressive Modeling Ditar: Diffusion transformer autoregressive modeling for speech generation,

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T00:52:02.873718Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:52:02.873718Z digest=sha256:6df0320460c9eaa47487990abfb51890672b2b8df7b4337e9491fe39c62ad0fb

Observation 36ccc01f-48dc-49d0-91d7-a45a98b8938b · outbound

This paper cites Auto-encoding variational bayes,.

StreamMel: Real-Time Zero-shot Text-to-Speech via Interleaved Continuous Autoregressive Modeling Auto-encoding variational bayes,

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T00:52:02.939608Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:52:02.939608Z digest=sha256:7815bf468db55a45ae25289f15f26ffb92ca506928b36d73447f1ea554fdd7a1

Observation 7082f2ce-2d21-48db-abde-1422f8f5d842 · outbound

This paper cites Tacotron: Towards end- to-end speech synthesis,.

StreamMel: Real-Time Zero-shot Text-to-Speech via Interleaved Continuous Autoregressive Modeling Tacotron: Towards end- to-end speech synthesis,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:52:06.061420Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T00:52:03.055590Z digest=sha256:a7111906b688046f09eb3a10e784f288029585195d194d0e22e39457be7e5f2d

Observation 99733192-e98e-40c2-9b61-e859f987b52d · outbound

This paper cites Librispeech: an ASR corpus based on public domain audio books,.

StreamMel: Real-Time Zero-shot Text-to-Speech via Interleaved Continuous Autoregressive Modeling Librispeech: an ASR corpus based on public domain audio books,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:52:05.895076Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T00:52:03.132846Z digest=sha256:2deca03ac0ae50749557f908f12cf77d31cdbd751040dd48e2da0b1b1d1d94dc

Observation 1e041d1c-da61-40fa-b08a-35a0d788755a · outbound

This paper cites V ALL-E R: robust and efficient zero-shot text-to-speech synthesis via monotonic alignment,.

StreamMel: Real-Time Zero-shot Text-to-Speech via Interleaved Continuous Autoregressive Modeling V ALL-E R: robust and efficient zero-shot text-to-speech synthesis via monotonic alignment,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:52:05.762832Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T00:52:03.223595Z digest=sha256:4b6508b98689aa3be42a2998a8532a1fed68c8fe54df82fa403e73e88f67dbc3

Observation 6d419976-a6ed-4cf3-9dd7-97606372ab49 · outbound

This paper cites F5-TTS: A fairytaler that fakes fluent and faithful speech with flow matching,.

StreamMel: Real-Time Zero-shot Text-to-Speech via Interleaved Continuous Autoregressive Modeling F5-TTS: A fairytaler that fakes fluent and faithful speech with flow matching,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:52:05.621674Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T00:52:03.285128Z digest=sha256:93ceee841dc49677d9324aa4d1e0f674eb5afd397d3e6be8b904f5ea8dfabe27

Observation f6a18580-b41c-42af-b442-e71d1614e80c · outbound

This paper cites MaskGCT: Zero-shot text-to-speech with masked generative codec transformer,.

StreamMel: Real-Time Zero-shot Text-to-Speech via Interleaved Continuous Autoregressive Modeling MaskGCT: Zero-shot text-to-speech with masked generative codec transformer,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:52:05.442194Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T00:52:03.385295Z digest=sha256:bfbe1dfed3af823a5f95691478b06cb5f508ea6b9c524a6bd50c6a87de49b26d

Observation 7afb6a5c-1d16-46c5-9b28-d3f2c1bd98c9 · outbound

This paper cites Ramp: Retrieval-augmented mos prediction via confidence-based dynamic weighting,.

StreamMel: Real-Time Zero-shot Text-to-Speech via Interleaved Continuous Autoregressive Modeling Ramp: Retrieval-augmented mos prediction via confidence-based dynamic weighting,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:52:05.327200Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T00:52:03.475605Z digest=sha256:7722edd14ad3141e042dcb16f7184a602e8eeaa0f73fe832c4c0ad4e991ecaca

Observation 462c7a9a-7137-421b-829c-2e2aaa8cb3ec · outbound

This paper cites Ramp+: Retrieval-augmented mos prediction with prior knowledge integration,.

StreamMel: Real-Time Zero-shot Text-to-Speech via Interleaved Continuous Autoregressive Modeling Ramp+: Retrieval-augmented mos prediction with prior knowledge integration,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:52:05.232775Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T00:52:03.590269Z digest=sha256:75e4277422418aab98661388f916ee7f340fefd5f2a85e2a1c2bfe49253dddd4

Observation 7d1701fd-bfef-49fa-abcc-e7495a5d0181 · outbound

This paper cites Conformer: Convolution-augmented Transformer for Speech Recognition.

StreamMel: Real-Time Zero-shot Text-to-Speech via Interleaved Continuous Autoregressive Modeling Conformer: Convolution-augmented Transformer for Speech Recognition

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T00:52:03.732161Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:52:03.732161Z digest=sha256:486f38acf1d8e1a808b0f52a8bba78754db61998356b1b1eb7ba079715b326db

Observation 204c52a3-173f-43c0-a320-2e4a4bc4181b · outbound

This paper cites HuBERT: Self-supervised speech representation learning by masked prediction of hidden units,.

StreamMel: Real-Time Zero-shot Text-to-Speech via Interleaved Continuous Autoregressive Modeling HuBERT: Self-supervised speech representation learning by masked prediction of hidden units,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:52:05.107680Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T00:52:03.807837Z digest=sha256:165149ef5468e614bf9369b7c8e55d18e79c9f04c4bf5d0c1066ec69a99ecb78

Observation ef7dbf33-7411-4e07-87b9-bc4450a22762 · outbound

This paper cites Robust Speech Recognition via Large-Scale Weak Supervision.

StreamMel: Real-Time Zero-shot Text-to-Speech via Interleaved Continuous Autoregressive Modeling Robust Speech Recognition via Large-Scale Weak Supervision

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T00:52:03.933734Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:52:03.933734Z digest=sha256:adfdaa7577b1e9dcea545b1cda855ffe391d5fb6ac2043aeac684cc450e4665a

Observation 50132868-6da2-4174-9ba8-8a8d3b9a67df · outbound

This paper cites WavLM: Large-scale self-supervised pre-training for full stack speech processing,.

StreamMel: Real-Time Zero-shot Text-to-Speech via Interleaved Continuous Autoregressive Modeling WavLM: Large-scale self-supervised pre-training for full stack speech processing,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:52:05.018828Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T00:52:04.039029Z digest=sha256:e3ff8ad705c5c68cb093dc40bd2d9e1f60d58318d3132777cc85dc7aaa0f0d4b

Observation 63c65484-ae4f-4ce4-80a1-aa28b396b6c3 · outbound

This paper cites Libriheavy: a 50,000 hours ASR corpus with punctuation casing and context,.

StreamMel: Real-Time Zero-shot Text-to-Speech via Interleaved Continuous Autoregressive Modeling Libriheavy: a 50,000 hours ASR corpus with punctuation casing and context,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:52:04.863237Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T00:52:04.109157Z digest=sha256:72f778fe238bef001c5a0d0eb4727a357d154ddcf84f869ddab30a8631f00a8d

Observation 68ed7cca-35c7-4bf2-bf63-e61243bf22fc · outbound

This paper cites Emilia: An extensive, multilingual, and diverse speech dataset for large-scale speech generation,.

StreamMel: Real-Time Zero-shot Text-to-Speech via Interleaved Continuous Autoregressive Modeling Emilia: An extensive, multilingual, and diverse speech dataset for large-scale speech generation,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:52:04.682115Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T00:52:04.191022Z digest=sha256:1183cdce737dece17fcca78a153efb3bbac645be31e3947408a9a755d606f187

Observation 2f204fc6-c3d9-4824-bc3b-a67581bde4cc · outbound

This paper cites Interleaved Speech-Text Language Models for Simple Streaming Text-to-Speech Synthesis.

StreamMel: Real-Time Zero-shot Text-to-Speech via Interleaved Continuous Autoregressive Modeling Interleaved Speech-Text Language Models for Simple Streaming Text-to-Speech Synthesis

Reference 2024

Resolution
metadata mismatch
local_arxiv, observed 2026-08-07T00:52:04.503082Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T00:52:02.441261Z digest=sha256:d31749f12d263073244732870191e1e8f8524f446662fce081696736b16bd809

Pith citing papers

No inbound Pith citation observations are available.