Pith. sign in

Paper Citation Record · LEDGER

Zero-Shot Streaming Text to Speech Synthesis with Transducer and Auto-Regressive Modeling

As of 15 August 2026, this Paper Citation Record lists 36 of 36 outbound references and 3 inbound Pith citation observations for arXiv:2505.19669.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.19669 v2

Coverage vector

measured 36 of 36 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:15:56.850541Z

measured 39 of 39 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:15:54.536075Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-06T11:22:28.361623Z

Reference resolution

36 of 36 outbound references displayed

  • verified exact2
  • verified fuzzy11
  • unresolved22
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 3db29749-9a2b-4027-bf1f-044ab8826a08 · outbound

This paper cites <eos> <bos> 𝑦! <eos> 𝑦#<bos> 𝑦.

Zero-Shot Streaming Text to Speech Synthesis with Transducer and Auto-Regressive Modeling <eos> <bos> 𝑦! <eos> 𝑦#<bos> 𝑦

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:16:00.735947Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T14:15:54.302540Z digest=sha256:ab02e16ebc640cc1f90d8dc52eeecca59295b396b266472bf2e734cdc61bdcf7

Observation 1ca17757-54e9-4855-a3da-83635fb401c5 · outbound

This paper cites Specifically, it uses a Transducer to convert the text into a sequence of se- mantic token in real time.

Zero-Shot Streaming Text to Speech Synthesis with Transducer and Auto-Regressive Modeling Specifically, it uses a Transducer to convert the text into a sequence of se- mantic token in real time

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:16:00.561077Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T14:15:54.368789Z digest=sha256:355183391ffd04af29a708ecc2f72177b86dc603e55a7f2dbfef6cf77d991e7a

Observation abd7dcf7-bf08-4280-a845-e12a5cd2ab21 · outbound

This paper cites Delete ⟨Bos⟩ Mechanism.

Zero-Shot Streaming Text to Speech Synthesis with Transducer and Auto-Regressive Modeling Delete ⟨Bos⟩ Mechanism

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:16:00.399535Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T14:15:54.457220Z digest=sha256:23ad61c3280aeb90b50863cc0ee643895194dfc8609abc7fb5e6b1283c373fc0

Observation b157e099-92a6-4c5f-86fa-8feddde23ab5 · outbound

This paper cites Zero-Shot Streaming Text to Speech Synthesis with Transducer and Auto-Regressive Modeling.

Zero-Shot Streaming Text to Speech Synthesis with Transducer and Auto-Regressive Modeling Zero-Shot Streaming Text to Speech Synthesis with Transducer and Auto-Regressive Modeling

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T14:15:54.536075Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:15:54.536075Z digest=sha256:e475f6d91e27ba1320d7414ceeca3a50714ef8d32fa618b26f835048fa60a18f

Observation b7359fb2-2924-46d9-ac34-68d8b76debfa · outbound

This paper cites 𝑚# 𝑚$ 𝑚% 𝑚& 𝑚'…… 𝑚! 𝑚% 𝑚.

Zero-Shot Streaming Text to Speech Synthesis with Transducer and Auto-Regressive Modeling 𝑚# 𝑚$ 𝑚% 𝑚& 𝑚'…… 𝑚! 𝑚% 𝑚

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:16:00.102532Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T14:15:54.612074Z digest=sha256:c313a377ab0c06e244f0169275d736f5f7d111a0ad458da6743362c09ed730bd

Observation 10fe7834-ee62-4958-a01f-13f37865905d · outbound

This paper cites Training Datasets We train SMLLE on the LibriSpeech dataset [27].

Zero-Shot Streaming Text to Speech Synthesis with Transducer and Auto-Regressive Modeling Training Datasets We train SMLLE on the LibriSpeech dataset [27]

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:59.849141Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T14:15:54.679034Z digest=sha256:48099a0104c39bf847a1d1be78394251417a0d5f4f8f58a38ffbcbe10bb0acff

Observation 637d4ad4-31c2-4e52-a91e-db1f8143081f · outbound

This paper cites SMLLE-R5.

Zero-Shot Streaming Text to Speech Synthesis with Transducer and Auto-Regressive Modeling SMLLE-R5

Reference 7

Resolution
malformed identifier
raw_fallback, observed 2026-08-07T14:15:59.537632Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T14:15:54.734281Z digest=sha256:b15be273c301358ec367c5e623e443c950b432f5e9d156c4b05d7e5f5d4a92b7

Observation 660d2a3d-5bc9-4dc2-b223-71d37dbf1aaa · outbound

This paper cites It uses a Transducer model to convert text into semantic tokens in real time and reconstructs them into mel-spectrograms frame by frame using an AR model.

Zero-Shot Streaming Text to Speech Synthesis with Transducer and Auto-Regressive Modeling It uses a Transducer model to convert text into semantic tokens in real time and reconstructs them into mel-spectrograms frame by frame using an AR model

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:59.284368Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T14:15:54.800671Z digest=sha256:c9a88cbfc169b8a86cdb6f17117093a724e9fbb0c872332e0d5e95cffe840464

Observation e61516d7-dcd4-4687-81c3-8ab1b49cc20a · outbound

This paper cites GPT-4 Technical Report.

Zero-Shot Streaming Text to Speech Synthesis with Transducer and Auto-Regressive Modeling GPT-4 Technical Report

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T14:15:54.839595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:15:54.839595Z digest=sha256:e86af6ca529ef11bab2dd57b2e08fa2107a5f728439588e0323fe863f3847cb9

Observation 10855297-8798-4096-83be-7c47490509e8 · outbound

This paper cites The Llama 3 Herd of Models.

Zero-Shot Streaming Text to Speech Synthesis with Transducer and Auto-Regressive Modeling The Llama 3 Herd of Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T14:15:54.910526Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:15:54.910526Z digest=sha256:007c507471d291f03fcd5288cce1ed1b109cd296b8196b604ceda020c999f85e

Observation e3a2c76e-35c3-4c52-be70-1536443c2ff4 · outbound

This paper cites Zero-shot text-to-image generation,.

Zero-Shot Streaming Text to Speech Synthesis with Transducer and Auto-Regressive Modeling Zero-shot text-to-image generation,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T14:15:54.998181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:15:54.998181Z digest=sha256:11298c17335e2b5549b3436b7ce6a199ce3fe529660b2ac4f47af94f9b386bea

Observation 91e8bcbc-9c99-4d09-9708-81e495670817 · outbound

This paper cites Learning transferable visual models from natural language supervision,.

Zero-Shot Streaming Text to Speech Synthesis with Transducer and Auto-Regressive Modeling Learning transferable visual models from natural language supervision,

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T14:15:55.105958Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:15:55.105958Z digest=sha256:d99344ebb11863a16d5223aa9456ac49567aba5e993bfcbcbfbfc832c450efc5

Observation 8688be3c-dae8-4999-810d-adf59919d480 · outbound

This paper cites Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers.

Zero-Shot Streaming Text to Speech Synthesis with Transducer and Auto-Regressive Modeling Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T14:15:55.172894Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:15:55.172894Z digest=sha256:5e233368fed5721e497168fb2b89a8f90a1dbf8847c8a4ec576a7fd2fc4b1f31

Observation fab4a743-7c6c-42bb-a1ab-0ef3ae3e293a · outbound

This paper cites Speak Foreign Languages with Your Own Voice: Cross-Lingual Neural Codec Language Modeling.

Zero-Shot Streaming Text to Speech Synthesis with Transducer and Auto-Regressive Modeling Speak Foreign Languages with Your Own Voice: Cross-Lingual Neural Codec Language Modeling

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T14:15:55.227783Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:15:55.227783Z digest=sha256:88e096cdf0775ec7d62f55eb1a063b6ed76c3f2175c8d3ca467d5b1f899a8b9c

Observation d5d276c2-d1f9-4b0d-b676-8d33a54b07dd · outbound

This paper cites Autoregressive Speech Synthesis without Vector Quantization.

Zero-Shot Streaming Text to Speech Synthesis with Transducer and Auto-Regressive Modeling Autoregressive Speech Synthesis without Vector Quantization

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T14:15:55.347789Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:15:55.347789Z digest=sha256:2fa15aac5c93a6ceab6c59d1f2a0750eec937be855697724814fd545a8d73552

Observation ba626b19-461e-4a04-8d57-da15f014d083 · outbound

This paper cites Moshi: a speech-text foundation model for real-time dialogue.

Zero-Shot Streaming Text to Speech Synthesis with Transducer and Auto-Regressive Modeling Moshi: a speech-text foundation model for real-time dialogue

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T14:15:55.424547Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:15:55.424547Z digest=sha256:34ca096b7b113a937e6ee0561ba56a03d0210cd589c6ef723cf878e9471cce11

Observation a56efe66-560a-4033-8703-cdd811a4327f · outbound

This paper cites Mini-Omni: Language Models Can Hear, Talk While Thinking in Streaming.

Zero-Shot Streaming Text to Speech Synthesis with Transducer and Auto-Regressive Modeling Mini-Omni: Language Models Can Hear, Talk While Thinking in Streaming

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T14:15:55.510988Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:15:55.510988Z digest=sha256:7bea6b4a948f569869f71eaefd4240a92155cfceaf047fa59349a928188c6444

Observation 132089e7-bcf9-4ff1-ada2-857a733acffe · outbound

This paper cites LLaMA-Omni: Seamless Speech Interaction with Large Language Models.

Zero-Shot Streaming Text to Speech Synthesis with Transducer and Auto-Regressive Modeling LLaMA-Omni: Seamless Speech Interaction with Large Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T14:15:55.598942Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:15:55.598942Z digest=sha256:1740ebb030c56cdf15ed90cb031341b8f12919896ca63b159cf6e66df51e7ac4

Observation 6bc85ce0-590c-4bb0-9971-49c256c354fc · outbound

This paper cites SALMONN: Towards Generic Hearing Abilities for Large Language Models.

Zero-Shot Streaming Text to Speech Synthesis with Transducer and Auto-Regressive Modeling SALMONN: Towards Generic Hearing Abilities for Large Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T14:15:55.630686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:15:55.630686Z digest=sha256:d823825872b8bcd7ab9ef9373cde4cdd03bd2dcef4305baeb7dd5a45da758100

Observation b5c68240-e06e-4395-a491-797cb4e11261 · outbound

This paper cites Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models.

Zero-Shot Streaming Text to Speech Synthesis with Transducer and Auto-Regressive Modeling Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T14:15:55.684869Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:15:55.684869Z digest=sha256:8f853d2bf18a6fc521f1ea3ad0549bb28ad8ce530c8c2499e6df8676af5e9391

Observation 0f9d2bf0-68e8-41e4-8fa5-9059d48c1758 · outbound

This paper cites WavLLM: Towards Robust and Adaptive Speech Large Language Model.

Zero-Shot Streaming Text to Speech Synthesis with Transducer and Auto-Regressive Modeling WavLLM: Towards Robust and Adaptive Speech Large Language Model

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T14:15:55.743399Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:15:55.743399Z digest=sha256:35917946de2962df70e796c7653e4140ff4f64490ce84d6de40e48508a5b591a

Observation 943ecc01-d9c2-4724-a43d-1676c5147316 · outbound

This paper cites LiveSpeech: Low-Latency Zero-shot Text-to-Speech via Autoregressive Modeling of Audio Discrete Codes.

Zero-Shot Streaming Text to Speech Synthesis with Transducer and Auto-Regressive Modeling LiveSpeech: Low-Latency Zero-shot Text-to-Speech via Autoregressive Modeling of Audio Discrete Codes

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T14:15:55.806875Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:15:55.806875Z digest=sha256:b4964eb8cd7f2a1c349a1acf109978d9fe0afbc29c480bd8572a04374bb7bc95

Observation f0f38e56-6acc-4499-aa22-b0ff256ed87f · outbound

This paper cites Speech-t: Transducer for text to speech and beyond,.

Zero-Shot Streaming Text to Speech Synthesis with Transducer and Auto-Regressive Modeling Speech-t: Transducer for text to speech and beyond,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:59.070271Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T14:15:55.885997Z digest=sha256:54911e7a8f47a47fdbaf2c43068d84524ca4654779ab254124122a2a9483b097

Observation d76bc8fa-d7af-4976-923f-a96854527548 · outbound

This paper cites Transduce and speak: Neural transducer for text-to-speech with semantic to- ken prediction,.

Zero-Shot Streaming Text to Speech Synthesis with Transducer and Auto-Regressive Modeling Transduce and speak: Neural transducer for text-to-speech with semantic to- ken prediction,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:58.838464Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T14:15:55.955500Z digest=sha256:38bbab0a17c23e4abbe74cbbfaa77d20724e6f45e3eff2f66e4c434f72a7c6bd

Observation d8ca4da6-6379-4673-9cd4-4329b7ea5d42 · outbound

This paper cites High Fidelity Text-to-Speech Via Discrete Tokens Using Token Transducer and Group Masked Language Model.

Zero-Shot Streaming Text to Speech Synthesis with Transducer and Auto-Regressive Modeling High Fidelity Text-to-Speech Via Discrete Tokens Using Token Transducer and Group Masked Language Model

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-08-07T14:15:57.452738Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T14:15:56.035770Z digest=sha256:6c6bd58888af3e3bca2b4b3577040332d3a96689514b7026bb22bf8e6e788115

Observation a5d1a7e1-0a42-4b77-9bae-a43bb0475ee4 · outbound

This paper cites TTS-Transducer: End-to-End Speech Synthesis with Neural Transducer.

Zero-Shot Streaming Text to Speech Synthesis with Transducer and Auto-Regressive Modeling TTS-Transducer: End-to-End Speech Synthesis with Neural Transducer

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-08-07T14:15:57.238904Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T14:15:56.168758Z digest=sha256:9012a7db5b1d7dfa8e49300240d8cd7f84d4dd60e6491be7b5710506256a6a2c

Observation 042c7359-715b-468c-98a5-c2a1e3645329 · outbound

This paper cites SpeechTokenizer: Unified Speech Tokenizer for Speech Large Language Models.

Zero-Shot Streaming Text to Speech Synthesis with Transducer and Auto-Regressive Modeling SpeechTokenizer: Unified Speech Tokenizer for Speech Large Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T14:15:56.248113Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:15:56.248113Z digest=sha256:363391e489b5f3e134a417ca836eb5ec53f34048da61fd1eb29d0de299950868

Observation b835ee00-fb8d-4e72-8092-748c120d4af8 · outbound

This paper cites ELLA-V: Stable Neural Codec Language Modeling with Alignment-guided Sequence Reordering.

Zero-Shot Streaming Text to Speech Synthesis with Transducer and Auto-Regressive Modeling ELLA-V: Stable Neural Codec Language Modeling with Alignment-guided Sequence Reordering

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T14:15:56.310389Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:15:56.310389Z digest=sha256:02093492f36e08c172c170ba84693d160190756bfc62f9cb6a7e2754d0df49a7

Observation 2ae8e0e9-d909-44bc-b465-2e2e7fd95af7 · outbound

This paper cites VALL-E R: Robust and Efficient Zero-Shot Text-to-Speech Synthesis via Monotonic Alignment.

Zero-Shot Streaming Text to Speech Synthesis with Transducer and Auto-Regressive Modeling VALL-E R: Robust and Efficient Zero-Shot Text-to-Speech Synthesis via Monotonic Alignment

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T14:15:56.366851Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:15:56.366851Z digest=sha256:3a1f128d6a6a2884eccdf506ff576a1cacdb112507568cef7c6ca88f3d5ee254

Observation fb0d1b7d-887d-45c5-a7ff-2646d3a864b2 · outbound

This paper cites RALL-E: Robust Codec Language Modeling with Chain-of-Thought Prompting for Text-to-Speech Synthesis.

Zero-Shot Streaming Text to Speech Synthesis with Transducer and Auto-Regressive Modeling RALL-E: Robust Codec Language Modeling with Chain-of-Thought Prompting for Text-to-Speech Synthesis

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T14:15:56.438913Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:15:56.438913Z digest=sha256:8472368609fee22243cef3b4d5de4404082d77c11f5614f996f3f9dfdb507c82

Observation 9181c0a4-cfee-4a30-b9b8-fe8424ceb737 · outbound

This paper cites CLaM-TTS: Improving neural codec language model for zero-shot text-to-speech,.

Zero-Shot Streaming Text to Speech Synthesis with Transducer and Auto-Regressive Modeling CLaM-TTS: Improving neural codec language model for zero-shot text-to-speech,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:58.542504Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T14:15:56.518067Z digest=sha256:c1975010eed67ae6d794e5e43798cf57d5b81b7b037c63d9ef1e9dac97728e40

Observation f6446f42-6a4d-4da2-8133-f9a06b7a46a7 · outbound

This paper cites VALL-E 2: Neural Codec Language Models are Human Parity Zero-Shot Text to Speech Synthesizers.

Zero-Shot Streaming Text to Speech Synthesis with Transducer and Auto-Regressive Modeling VALL-E 2: Neural Codec Language Models are Human Parity Zero-Shot Text to Speech Synthesizers

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T14:15:56.602314Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:15:56.602314Z digest=sha256:d00c77d4b149460c10f78a43808f3d7790634e4e6f696c9374fd631b9f142fe2

Observation 01897a29-0ba8-4177-8598-5b7d749b6ec3 · outbound

This paper cites V oicebox: Text-guided multilingual universal speech gen- eration at scale,.

Zero-Shot Streaming Text to Speech Synthesis with Transducer and Auto-Regressive Modeling V oicebox: Text-guided multilingual universal speech gen- eration at scale,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:58.207909Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T14:15:56.695232Z digest=sha256:82edde830402001218ae8c36f925f168c8e70e0e37e698819d079920cee2f7ff

Observation 187165c8-e0e0-453a-9c32-094d0596e4c6 · outbound

This paper cites Lib- rispeech: An ASR corpus based on public domain audio books,.

Zero-Shot Streaming Text to Speech Synthesis with Transducer and Auto-Regressive Modeling Lib- rispeech: An ASR corpus based on public domain audio books,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:58.035412Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T14:15:56.756041Z digest=sha256:19d41d5f97e88e40dfcffd4b59bba36e48dbfd99d67b6c43b325eb622b8e6e1f

Observation 4f863876-8cb1-4ca8-a1ed-7fe38cb5bdd7 · outbound

This paper cites LibriTTS: A Corpus Derived from LibriSpeech for Text-to-Speech.

Zero-Shot Streaming Text to Speech Synthesis with Transducer and Auto-Regressive Modeling LibriTTS: A Corpus Derived from LibriSpeech for Text-to-Speech

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T14:15:56.797055Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:15:56.797055Z digest=sha256:b08b8ddf30fe48b4f5e0fba5aa6a4af78d3cbdc0f27ce25da9d1fb7c889770cf

Observation 041daa36-4dc0-4883-aa0a-bc187e0c0483 · outbound

This paper cites Zero-Shot Text-to-Speech from Continuous Text Streams.

Zero-Shot Streaming Text to Speech Synthesis with Transducer and Auto-Regressive Modeling Zero-Shot Text-to-Speech from Continuous Text Streams

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T14:15:56.850541Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:15:56.850541Z digest=sha256:6781d969a294a6fdbbc56794c9ca470dba64730eb6ecf1c1a9a556b982381b48

Pith citing papers

Observation b157e099-92a6-4c5f-86fa-8feddde23ab5 · inbound

Zero-Shot Streaming Text to Speech Synthesis with Transducer and Auto-Regressive Modeling cites this paper.

Zero-Shot Streaming Text to Speech Synthesis with Transducer and Auto-Regressive Modeling Zero-Shot Streaming Text to Speech Synthesis with Transducer and Auto-Regressive Modeling

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T14:15:54.536075Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:15:54.536075Z digest=sha256:e475f6d91e27ba1320d7414ceeca3a50714ef8d32fa618b26f835048fa60a18f

Observation ad5cbb77-9e8e-4f73-ad1a-344020d4bdf4 · inbound

StreamMel: Real-Time Zero-shot Text-to-Speech via Interleaved Continuous Autoregressive Modeling cites this paper.

StreamMel: Real-Time Zero-shot Text-to-Speech via Interleaved Continuous Autoregressive Modeling Zero-Shot Streaming Text to Speech Synthesis with Transducer and Auto-Regressive Modeling

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T00:52:02.492729Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:52:02.492729Z digest=sha256:7ea50103244c364f3922ecadd5968332b766e154b049c4afd683df69c7d16f9e

Observation 4706dce0-0b29-472a-80c5-38080d13a8c1 · inbound

Next Tokens Denoising for Speech Synthesis cites this paper.

Next Tokens Denoising for Speech Synthesis Zero-Shot Streaming Text to Speech Synthesis with Transducer and Auto-Regressive Modeling

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-08-06T11:22:28.411634Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T11:22:27.182862Z digest=sha256:c29da22da713869db14ce73bf230dedee42a5d200479ff27eb9c937df783b6e3