Pith. sign in

Paper Citation Record · LEDGER

Zero-Shot Streaming Text to Speech Synthesis with Transducer and Auto-Regressive Modeling

As of 8 August 2026, this Paper Citation Record lists 36 of 36 outbound references and 3 inbound Pith citation observations for arXiv:2505.19669.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.19669 v2

Coverage vector

measured 36 of 36 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:15:56.850541Z

measured 39 of 39 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:15:54.536075Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-06T11:22:28.361623Z

Reference resolution

36 of 36 outbound references displayed

  • verified exact2
  • verified fuzzy11
  • unresolved22
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 3db29749-9a2b-4027-bf1f-044ab8826a08 · outbound

This paper cites <eos> <bos> 𝑦! <eos> 𝑦#<bos> 𝑦.

Zero-Shot Streaming Text to Speech Synthesis with Transducer and Auto-Regressive Modeling <eos> <bos> 𝑦! <eos> 𝑦#<bos> 𝑦

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:16:00.735947Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:15:54.302540Z digest=sha256:9d29ca2be363a8cf707edb9c3d94bf08ed67a6a490d9f288ae829a2b9e9a73f6

Observation 1ca17757-54e9-4855-a3da-83635fb401c5 · outbound

This paper cites Specifically, it uses a Transducer to convert the text into a sequence of se- mantic token in real time.

Zero-Shot Streaming Text to Speech Synthesis with Transducer and Auto-Regressive Modeling Specifically, it uses a Transducer to convert the text into a sequence of se- mantic token in real time

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:16:00.561077Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:15:54.368789Z digest=sha256:eabffa8404556ecb24916e406a128303d072682d1c4559e5566eca026d80dd92

Observation abd7dcf7-bf08-4280-a845-e12a5cd2ab21 · outbound

This paper cites Delete ⟨Bos⟩ Mechanism.

Zero-Shot Streaming Text to Speech Synthesis with Transducer and Auto-Regressive Modeling Delete ⟨Bos⟩ Mechanism

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:16:00.399535Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:15:54.457220Z digest=sha256:b1b9567c80565cd722c0b53fc418ca1efe966ec2cbdddcf0678f19cc2ba7e1dc

Observation b157e099-92a6-4c5f-86fa-8feddde23ab5 · outbound

This paper cites Zero-Shot Streaming Text to Speech Synthesis with Transducer and Auto-Regressive Modeling.

Zero-Shot Streaming Text to Speech Synthesis with Transducer and Auto-Regressive Modeling Zero-Shot Streaming Text to Speech Synthesis with Transducer and Auto-Regressive Modeling

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T14:15:54.536075Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:15:54.536075Z digest=sha256:bd5ece38c13b255793d0366a1f7d9c6fc0e798eda7fef2a92b12655180dbb4cd

Observation b7359fb2-2924-46d9-ac34-68d8b76debfa · outbound

This paper cites 𝑚# 𝑚$ 𝑚% 𝑚& 𝑚'…… 𝑚! 𝑚% 𝑚.

Zero-Shot Streaming Text to Speech Synthesis with Transducer and Auto-Regressive Modeling 𝑚# 𝑚$ 𝑚% 𝑚& 𝑚'…… 𝑚! 𝑚% 𝑚

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:16:00.102532Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:15:54.612074Z digest=sha256:8507ea407618974b2151bd2803ed874eea3da74707bd1eb11ac334dc237f6212

Observation 10fe7834-ee62-4958-a01f-13f37865905d · outbound

This paper cites Training Datasets We train SMLLE on the LibriSpeech dataset [27].

Zero-Shot Streaming Text to Speech Synthesis with Transducer and Auto-Regressive Modeling Training Datasets We train SMLLE on the LibriSpeech dataset [27]

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:59.849141Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:15:54.679034Z digest=sha256:ae23d2d429b0685dcf9a11630eee48d107af13b799c002debc485788e4d789df

Observation 637d4ad4-31c2-4e52-a91e-db1f8143081f · outbound

This paper cites SMLLE-R5.

Zero-Shot Streaming Text to Speech Synthesis with Transducer and Auto-Regressive Modeling SMLLE-R5

Reference 7

Resolution
malformed identifier
raw_fallback, observed 2026-08-07T14:15:59.537632Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:15:54.734281Z digest=sha256:40bb766c07eea4c3ef7566862900ce095a09689f1b7efae5fba4d1dbd28658eb

Observation 660d2a3d-5bc9-4dc2-b223-71d37dbf1aaa · outbound

This paper cites It uses a Transducer model to convert text into semantic tokens in real time and reconstructs them into mel-spectrograms frame by frame using an AR model.

Zero-Shot Streaming Text to Speech Synthesis with Transducer and Auto-Regressive Modeling It uses a Transducer model to convert text into semantic tokens in real time and reconstructs them into mel-spectrograms frame by frame using an AR model

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:59.284368Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:15:54.800671Z digest=sha256:bb72db48ef0476e362baac9f485e2437bc616b982966f7511db2dbf214fcd813

Observation e61516d7-dcd4-4687-81c3-8ab1b49cc20a · outbound

This paper cites GPT-4 Technical Report.

Zero-Shot Streaming Text to Speech Synthesis with Transducer and Auto-Regressive Modeling GPT-4 Technical Report

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T14:15:54.839595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:15:54.839595Z digest=sha256:b244483e46273451019093048e3d7d628dd3d269aee682418f1787303da204d5

Observation 10855297-8798-4096-83be-7c47490509e8 · outbound

This paper cites The Llama 3 Herd of Models.

Zero-Shot Streaming Text to Speech Synthesis with Transducer and Auto-Regressive Modeling The Llama 3 Herd of Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T14:15:54.910526Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:15:54.910526Z digest=sha256:60224bf17b0d71b0cfdfead155bde78b4105232f40b9fb5b6a5c5838ce252908

Observation e3a2c76e-35c3-4c52-be70-1536443c2ff4 · outbound

This paper cites Zero-shot text-to-image generation,.

Zero-Shot Streaming Text to Speech Synthesis with Transducer and Auto-Regressive Modeling Zero-shot text-to-image generation,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T14:15:54.998181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:15:54.998181Z digest=sha256:4cd1f283c08307662af75e5c3e0875faf8d0fe0e89afaf084ffbf28d78629390

Observation 91e8bcbc-9c99-4d09-9708-81e495670817 · outbound

This paper cites Learning transferable visual models from natural language supervision,.

Zero-Shot Streaming Text to Speech Synthesis with Transducer and Auto-Regressive Modeling Learning transferable visual models from natural language supervision,

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T14:15:55.105958Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:15:55.105958Z digest=sha256:4c5adc506cf80b4c977b814011c1f35f0c79c57f5aa402e2de57f005dac2986c

Observation 8688be3c-dae8-4999-810d-adf59919d480 · outbound

This paper cites Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers.

Zero-Shot Streaming Text to Speech Synthesis with Transducer and Auto-Regressive Modeling Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T14:15:55.172894Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:15:55.172894Z digest=sha256:9a78841abaf95ac311c985abbdeaca4cb68c78d4f5a4e25c9bc205ae6efb867a

Observation fab4a743-7c6c-42bb-a1ab-0ef3ae3e293a · outbound

This paper cites Speak Foreign Languages with Your Own Voice: Cross-Lingual Neural Codec Language Modeling.

Zero-Shot Streaming Text to Speech Synthesis with Transducer and Auto-Regressive Modeling Speak Foreign Languages with Your Own Voice: Cross-Lingual Neural Codec Language Modeling

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T14:15:55.227783Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:15:55.227783Z digest=sha256:a6b223e0d9be57afd2fde00f7a2acbe6c13310d4ba66a80f6cdbe8e81f9f0fa6

Observation d5d276c2-d1f9-4b0d-b676-8d33a54b07dd · outbound

This paper cites Autoregressive Speech Synthesis without Vector Quantization.

Zero-Shot Streaming Text to Speech Synthesis with Transducer and Auto-Regressive Modeling Autoregressive Speech Synthesis without Vector Quantization

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T14:15:55.347789Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:15:55.347789Z digest=sha256:fb0e6b6f8430599cf947d8e89cabe72e158f1dbc6b0c4f2b84d92b28cc7adf09

Observation ba626b19-461e-4a04-8d57-da15f014d083 · outbound

This paper cites Moshi: a speech-text foundation model for real-time dialogue.

Zero-Shot Streaming Text to Speech Synthesis with Transducer and Auto-Regressive Modeling Moshi: a speech-text foundation model for real-time dialogue

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T14:15:55.424547Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:15:55.424547Z digest=sha256:4a6373f19064403117354e0f2adf0231e8d8f5f84ca3e3fa56fded37bff67794

Observation a56efe66-560a-4033-8703-cdd811a4327f · outbound

This paper cites Mini-Omni: Language Models Can Hear, Talk While Thinking in Streaming.

Zero-Shot Streaming Text to Speech Synthesis with Transducer and Auto-Regressive Modeling Mini-Omni: Language Models Can Hear, Talk While Thinking in Streaming

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T14:15:55.510988Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:15:55.510988Z digest=sha256:5a99b1aab52d5914137f3e3944e9f6423bff9b328f99dd8888aa27cb15132e39

Observation 132089e7-bcf9-4ff1-ada2-857a733acffe · outbound

This paper cites LLaMA-Omni: Seamless Speech Interaction with Large Language Models.

Zero-Shot Streaming Text to Speech Synthesis with Transducer and Auto-Regressive Modeling LLaMA-Omni: Seamless Speech Interaction with Large Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T14:15:55.598942Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:15:55.598942Z digest=sha256:00d7f034970d82e7caaac6d772dd933ff4ac00f600e873b9ed9ff53d5d9d577b

Observation 6bc85ce0-590c-4bb0-9971-49c256c354fc · outbound

This paper cites SALMONN: Towards Generic Hearing Abilities for Large Language Models.

Zero-Shot Streaming Text to Speech Synthesis with Transducer and Auto-Regressive Modeling SALMONN: Towards Generic Hearing Abilities for Large Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T14:15:55.630686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:15:55.630686Z digest=sha256:cb88b9a523b50c34b4562c104d3508ee3a2faa4221da1f02d4ff5e3e126b8b03

Observation b5c68240-e06e-4395-a491-797cb4e11261 · outbound

This paper cites Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models.

Zero-Shot Streaming Text to Speech Synthesis with Transducer and Auto-Regressive Modeling Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T14:15:55.684869Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:15:55.684869Z digest=sha256:850bac2ce59bae417875d90fcd1fe7d055665ee2fd23c1cee7e40420312c2421

Observation 0f9d2bf0-68e8-41e4-8fa5-9059d48c1758 · outbound

This paper cites WavLLM: Towards Robust and Adaptive Speech Large Language Model.

Zero-Shot Streaming Text to Speech Synthesis with Transducer and Auto-Regressive Modeling WavLLM: Towards Robust and Adaptive Speech Large Language Model

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T14:15:55.743399Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:15:55.743399Z digest=sha256:af50d82721d75b2a6b422da9fc2020c1a034416eb8a5dec8863502b77cf624ba

Observation 943ecc01-d9c2-4724-a43d-1676c5147316 · outbound

This paper cites LiveSpeech: Low-Latency Zero-shot Text-to-Speech via Autoregressive Modeling of Audio Discrete Codes.

Zero-Shot Streaming Text to Speech Synthesis with Transducer and Auto-Regressive Modeling LiveSpeech: Low-Latency Zero-shot Text-to-Speech via Autoregressive Modeling of Audio Discrete Codes

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T14:15:55.806875Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:15:55.806875Z digest=sha256:a28b3870b76a9e7b9b91ec987c1102d34bd2c35fac96900cfc213c42df66ed80

Observation f0f38e56-6acc-4499-aa22-b0ff256ed87f · outbound

This paper cites Speech-t: Transducer for text to speech and beyond,.

Zero-Shot Streaming Text to Speech Synthesis with Transducer and Auto-Regressive Modeling Speech-t: Transducer for text to speech and beyond,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:59.070271Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:15:55.885997Z digest=sha256:134ede58877a896ad9a0d4c531dd548272a65d4079ee6ff7cc59207993bd312d

Observation d76bc8fa-d7af-4976-923f-a96854527548 · outbound

This paper cites Transduce and speak: Neural transducer for text-to-speech with semantic to- ken prediction,.

Zero-Shot Streaming Text to Speech Synthesis with Transducer and Auto-Regressive Modeling Transduce and speak: Neural transducer for text-to-speech with semantic to- ken prediction,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:58.838464Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:15:55.955500Z digest=sha256:d099842fabbae92d85203368e07acede4182b5480287c8ec8744788bb73805ff

Observation d8ca4da6-6379-4673-9cd4-4329b7ea5d42 · outbound

This paper cites High Fidelity Text-to-Speech Via Discrete Tokens Using Token Transducer and Group Masked Language Model.

Zero-Shot Streaming Text to Speech Synthesis with Transducer and Auto-Regressive Modeling High Fidelity Text-to-Speech Via Discrete Tokens Using Token Transducer and Group Masked Language Model

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-08-07T14:15:57.452738Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:15:56.035770Z digest=sha256:3dcdd8cca7f1b25a8daa42c132aaa9fa7ca94f9e53d1ba852f9472f64f6e6928

Observation a5d1a7e1-0a42-4b77-9bae-a43bb0475ee4 · outbound

This paper cites TTS-Transducer: End-to-End Speech Synthesis with Neural Transducer.

Zero-Shot Streaming Text to Speech Synthesis with Transducer and Auto-Regressive Modeling TTS-Transducer: End-to-End Speech Synthesis with Neural Transducer

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-08-07T14:15:57.238904Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:15:56.168758Z digest=sha256:0b173d1189a9cf0f60309f58f15eb68c3bbd511a9cbdfda19600d3898214c4d2

Observation 042c7359-715b-468c-98a5-c2a1e3645329 · outbound

This paper cites SpeechTokenizer: Unified Speech Tokenizer for Speech Large Language Models.

Zero-Shot Streaming Text to Speech Synthesis with Transducer and Auto-Regressive Modeling SpeechTokenizer: Unified Speech Tokenizer for Speech Large Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T14:15:56.248113Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:15:56.248113Z digest=sha256:7356261964f3ceb07f678e9631da74fc963c32b86ac20d233972d2ba2853e355

Observation b835ee00-fb8d-4e72-8092-748c120d4af8 · outbound

This paper cites ELLA-V: Stable Neural Codec Language Modeling with Alignment-guided Sequence Reordering.

Zero-Shot Streaming Text to Speech Synthesis with Transducer and Auto-Regressive Modeling ELLA-V: Stable Neural Codec Language Modeling with Alignment-guided Sequence Reordering

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T14:15:56.310389Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:15:56.310389Z digest=sha256:2c457508729aa25ae9413623c19c458dd62b65dab964d6760b6be5ea89f3d808

Observation 2ae8e0e9-d909-44bc-b465-2e2e7fd95af7 · outbound

This paper cites VALL-E R: Robust and Efficient Zero-Shot Text-to-Speech Synthesis via Monotonic Alignment.

Zero-Shot Streaming Text to Speech Synthesis with Transducer and Auto-Regressive Modeling VALL-E R: Robust and Efficient Zero-Shot Text-to-Speech Synthesis via Monotonic Alignment

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T14:15:56.366851Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:15:56.366851Z digest=sha256:7fd1042f65726cd4c789cd3ae69253c57783d3fc54a2bbb9d3dc655eed7bd430

Observation fb0d1b7d-887d-45c5-a7ff-2646d3a864b2 · outbound

This paper cites RALL-E: Robust Codec Language Modeling with Chain-of-Thought Prompting for Text-to-Speech Synthesis.

Zero-Shot Streaming Text to Speech Synthesis with Transducer and Auto-Regressive Modeling RALL-E: Robust Codec Language Modeling with Chain-of-Thought Prompting for Text-to-Speech Synthesis

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T14:15:56.438913Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:15:56.438913Z digest=sha256:bee6e8ebf5ef89899493ea0b6293f58ff13e2ef6ca2ba283c7ea8de451671a94

Observation 9181c0a4-cfee-4a30-b9b8-fe8424ceb737 · outbound

This paper cites CLaM-TTS: Improving neural codec language model for zero-shot text-to-speech,.

Zero-Shot Streaming Text to Speech Synthesis with Transducer and Auto-Regressive Modeling CLaM-TTS: Improving neural codec language model for zero-shot text-to-speech,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:58.542504Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:15:56.518067Z digest=sha256:cadb590dbd89410502fb4d28f5c47bf765ae284b6792d50b335a543a84fd5cec

Observation f6446f42-6a4d-4da2-8133-f9a06b7a46a7 · outbound

This paper cites VALL-E 2: Neural Codec Language Models are Human Parity Zero-Shot Text to Speech Synthesizers.

Zero-Shot Streaming Text to Speech Synthesis with Transducer and Auto-Regressive Modeling VALL-E 2: Neural Codec Language Models are Human Parity Zero-Shot Text to Speech Synthesizers

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T14:15:56.602314Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:15:56.602314Z digest=sha256:f241046236a74ae59229720fbf61da13bbf43d1e114e10788e8724dd333ff13c

Observation 01897a29-0ba8-4177-8598-5b7d749b6ec3 · outbound

This paper cites V oicebox: Text-guided multilingual universal speech gen- eration at scale,.

Zero-Shot Streaming Text to Speech Synthesis with Transducer and Auto-Regressive Modeling V oicebox: Text-guided multilingual universal speech gen- eration at scale,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:58.207909Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:15:56.695232Z digest=sha256:601a7765340b4d950e163a5b6bb956873179d6d676f9c0670fea2d6dc96c5359

Observation 187165c8-e0e0-453a-9c32-094d0596e4c6 · outbound

This paper cites Lib- rispeech: An ASR corpus based on public domain audio books,.

Zero-Shot Streaming Text to Speech Synthesis with Transducer and Auto-Regressive Modeling Lib- rispeech: An ASR corpus based on public domain audio books,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:15:58.035412Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T14:15:56.756041Z digest=sha256:b919711a3571a6aa97c5197fa78070caf46ccf0e9b9615d3dafbbfaee9920913

Observation 4f863876-8cb1-4ca8-a1ed-7fe38cb5bdd7 · outbound

This paper cites LibriTTS: A Corpus Derived from LibriSpeech for Text-to-Speech.

Zero-Shot Streaming Text to Speech Synthesis with Transducer and Auto-Regressive Modeling LibriTTS: A Corpus Derived from LibriSpeech for Text-to-Speech

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T14:15:56.797055Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:15:56.797055Z digest=sha256:d90618a6654175b5b63549510a56db25c7b029227648fa35f9c7af68813d3b35

Observation 041daa36-4dc0-4883-aa0a-bc187e0c0483 · outbound

This paper cites Zero-Shot Text-to-Speech from Continuous Text Streams.

Zero-Shot Streaming Text to Speech Synthesis with Transducer and Auto-Regressive Modeling Zero-Shot Text-to-Speech from Continuous Text Streams

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T14:15:56.850541Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:15:56.850541Z digest=sha256:1be412ca7c7a5840411975019bdf5ecac6946b393f31371ff77b31af89e5a227

Pith citing papers

Observation b157e099-92a6-4c5f-86fa-8feddde23ab5 · inbound

Zero-Shot Streaming Text to Speech Synthesis with Transducer and Auto-Regressive Modeling cites this paper.

Zero-Shot Streaming Text to Speech Synthesis with Transducer and Auto-Regressive Modeling Zero-Shot Streaming Text to Speech Synthesis with Transducer and Auto-Regressive Modeling

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T14:15:54.536075Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:15:54.536075Z digest=sha256:bd5ece38c13b255793d0366a1f7d9c6fc0e798eda7fef2a92b12655180dbb4cd

Observation ad5cbb77-9e8e-4f73-ad1a-344020d4bdf4 · inbound

StreamMel: Real-Time Zero-shot Text-to-Speech via Interleaved Continuous Autoregressive Modeling cites this paper.

StreamMel: Real-Time Zero-shot Text-to-Speech via Interleaved Continuous Autoregressive Modeling Zero-Shot Streaming Text to Speech Synthesis with Transducer and Auto-Regressive Modeling

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T00:52:02.492729Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:52:02.492729Z digest=sha256:42db6069532cc53ae24d08b6476d7e9f3fc4d4cb8b0553e281e32c1061b67f92

Observation 4706dce0-0b29-472a-80c5-38080d13a8c1 · inbound

Next Tokens Denoising for Speech Synthesis cites this paper.

Next Tokens Denoising for Speech Synthesis Zero-Shot Streaming Text to Speech Synthesis with Transducer and Auto-Regressive Modeling

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-08-06T11:22:28.411634Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T11:22:27.182862Z digest=sha256:701e548d9e74ace2d166a68e748111c8defb197e045bee3fb2ed20518348c4ad