Pith. sign in

Paper Citation Record · LEDGER

VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation

As of 18 August 2026, this Paper Citation Record lists 80 of 80 outbound references and 1 inbound Pith citation observation for arXiv:2505.19462.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.19462 v2

Coverage vector

measured 80 of 80 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:19:58.490777Z

measured 81 of 81 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-12T23:22:40.218077Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

80 of 80 outbound references displayed

  • verified exact2
  • verified fuzzy29
  • unresolved49
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 1d9a321e-26ae-4b4f-8f8e-d74b54358d5a · outbound

This paper cites Soundstream: An end-to-end neural audio codec.IEEE/ACM Transactions on Audio, Speech, and Language Processing, 30:495–507, 2021.

VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation Soundstream: An end-to-end neural audio codec.IEEE/ACM Transactions on Audio, Speech, and Language Processing, 30:495–507, 2021

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T14:19:49.841291Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:19:49.841291Z digest=sha256:5ff9091f7d8804f2d436f6cfca86bd32d88eeec7f01e5c05fb9422e4e1773f68

Observation 70b9dbd1-e4e9-4bdf-a20d-6b0e93c64b32 · outbound

This paper cites High Fidelity Neural Audio Compression.

VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation High Fidelity Neural Audio Compression

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T14:19:49.901237Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:19:49.901237Z digest=sha256:b4f6fda88a7451f6e272255e85d0ab41d648407584f08cfc45d57fe07d0ed8e1

Observation 2460dc9a-5101-4a73-b0ba-b16d0f199337 · outbound

This paper cites F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching.

VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T14:19:50.066833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:19:50.066833Z digest=sha256:aaa7a433c2237cd1d14970564c4f38799d11223c3b21a7e01488b8bf0a8741b3

Observation c05b4db4-dc42-462e-93ba-1d2c9cada271 · outbound

This paper cites Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N.

VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T14:19:50.173827Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:19:50.173827Z digest=sha256:c46b7862fcb7825184269ecba801b6adefb94e21666d1f039bb2bb6bdfe759af

Observation 413b2487-d923-4f77-921b-443342a17858 · outbound

This paper cites Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers.

VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T14:19:50.254673Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:19:50.254673Z digest=sha256:c3267345b41baf918e9b8ae0e28d2e0d3175f14e515c5a000ed7bd7064e55f3b

Observation e854fbbc-f772-4500-8022-aa9bc2b23ef7 · outbound

This paper cites Vall-e 2: Neural codec language models are human parity zero-shot text to speech synthesizers, 2024.

VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation Vall-e 2: Neural codec language models are human parity zero-shot text to speech synthesizers, 2024

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:20:05.424679Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T14:19:50.374743Z digest=sha256:683834f8a872a7f3ebed90cb6cd9c5815e2800e0d74bfcef8f63d82886f2a2f2

Observation 9918ab20-a36c-4e57-8ef1-e28ca562c615 · outbound

This paper cites V oicecraft: Zero-shot speech editing and text-to-speech in the wild.

VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation V oicecraft: Zero-shot speech editing and text-to-speech in the wild

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:20:05.056341Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T14:19:50.481272Z digest=sha256:44d8cb34cb3a53e5cfcd578578413b0baea478a6809752de0658a7e0122da15a

Observation 6404e8c6-308b-48ab-b2ff-92543293ef9a · outbound

This paper cites On generative spoken language modeling from raw audio.Transactions of the Association for Computational Linguistics, 9:1336–1354, 2021.

VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation On generative spoken language modeling from raw audio.Transactions of the Association for Computational Linguistics, 9:1336–1354, 2021

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:20:04.656485Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T14:19:50.603358Z digest=sha256:bd61a224a1bc398b6f474dcc3a0ca2f623326317916d245bd3fee55c744ba5d6

Observation 09cc527d-f973-43eb-9ce5-a5c6bd7309e5 · outbound

This paper cites Audiolm: A language modeling approach to audio generation.IEEE/ACM Transactions on Audio, Speech, and Language Processing, 31:2523–2533, 2022.

VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation Audiolm: A language modeling approach to audio generation.IEEE/ACM Transactions on Audio, Speech, and Language Processing, 31:2523–2533, 2022

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:20:04.287314Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T14:19:50.727860Z digest=sha256:09d8fa053ebed5a198710412e9e019e1d08d1b15a426ead3e70c8a83f2a34db0

Observation ebbb6bb1-5a6c-4a0b-8b3c-05f49645f4d4 · outbound

This paper cites AudioGen: Textually Guided Audio Generation.

VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation AudioGen: Textually Guided Audio Generation

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T14:19:50.833721Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:19:50.833721Z digest=sha256:92a7d6cc4d196e242e05fda9fd17dfe9ef7b02479c1a3a12dce85cbc0a55c360

Observation 986297ee-6df1-432f-954f-d02096739bf5 · outbound

This paper cites Speak, read and prompt: High-fidelity text-to-speech with minimal supervision.Transactions of the Association for Computational Linguistics, 11:1703–1718, 2023.

VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation Speak, read and prompt: High-fidelity text-to-speech with minimal supervision.Transactions of the Association for Computational Linguistics, 11:1703–1718, 2023

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:20:03.955411Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T14:19:50.893333Z digest=sha256:cc7ade73145514297e6f8707b5efbd42cc4c3e20008608ddf0c8c85eeb6eb55d

Observation bb56d6d4-287a-4e76-8944-dd5c35448e75 · outbound

This paper cites SoundStorm: Efficient Parallel Audio Generation.

VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation SoundStorm: Efficient Parallel Audio Generation

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T14:19:50.973655Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:19:50.973655Z digest=sha256:450d018fdde2eda927925874dc3180099dc7fe0ec18f4aa400de21c45f0fb148

Observation c44eb73f-b23e-479f-822f-8d99bf5d7fe1 · outbound

This paper cites Voicebox: Text-Guided Multilingual Universal Speech Generation at Scale.

VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation Voicebox: Text-Guided Multilingual Universal Speech Generation at Scale

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T14:19:51.036153Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:19:51.036153Z digest=sha256:b8185ab98fbfff6c5461e6f3f756d362657d6560da41e9896bb693067f419088

Observation 03333264-800a-4980-89ee-16fd2c942e25 · outbound

This paper cites MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer.

VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T14:19:51.135367Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:19:51.135367Z digest=sha256:97cf97acf481be00501d0bfbf6594e0a85c4b22d34fa310dc5326d9f0fd9398f

Observation 3a1a758a-6b3e-4758-a03d-e2c1796f49f5 · outbound

This paper cites Robust and Unbounded Length Generalization in Autoregressive Transformer-Based Text-to-Speech.

VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation Robust and Unbounded Length Generalization in Autoregressive Transformer-Based Text-to-Speech

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-08-07T14:19:58.963053Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T14:19:51.283267Z digest=sha256:ad43428fe0b3dc3b352cb7e624178c172eb244b3381b282cb29f00f48551af35

Observation 27763106-efe7-4f7b-834b-59b89b87e191 · outbound

This paper cites CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens.

VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T14:19:51.454188Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:19:51.454188Z digest=sha256:eec671a2fd2f94f5113a7ce3fe46a591d6a3c012b49a940378a5df781388f932

Observation 9382fd7f-03e0-478b-9daa-2c1b9d74e65d · outbound

This paper cites CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models.

VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T14:19:51.609437Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:19:51.609437Z digest=sha256:5369c054d768b8cbb2aba0efc278eda911eb560f7ff38f0549a4724cb7d4ef3a

Observation 3b032d13-40de-4694-b4d7-3fd94ebbce1d · outbound

This paper cites Llasa: Scaling train-time and inference-time compute for llama-based speech synthesis.

VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation Llasa: Scaling train-time and inference-time compute for llama-based speech synthesis

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:20:03.667748Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T14:19:51.705529Z digest=sha256:31559f77e80f273ed5b5aa466a6524629da4a82d0ecde336be561acbd5eba972

Observation b84f76b2-5175-4208-9515-21c1bc861091 · outbound

This paper cites Ella-v: Stable neural codec language modeling with alignment-guided sequence reordering.

VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation Ella-v: Stable neural codec language modeling with alignment-guided sequence reordering

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:20:03.396435Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T14:19:51.826778Z digest=sha256:7ed44ac46b093347ba7394f36470c6d41c6a8b378f225aa319de284d84db2561

Observation c099402e-4fee-4705-b271-6f0c51e6f718 · outbound

This paper cites VALL-E R: Robust and Efficient Zero-Shot Text-to-Speech Synthesis via Monotonic Alignment.

VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation VALL-E R: Robust and Efficient Zero-Shot Text-to-Speech Synthesis via Monotonic Alignment

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T14:19:51.952647Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:19:51.952647Z digest=sha256:7d205ae54b173e1a9ccc7a2388746de20504d044ec6a7b8ce94320b8e7611e0f

Observation 474ee284-a658-4782-b777-5fa639db753e · outbound

This paper cites Vall-t: Decoder-only generative transducer for robust and decoding- controllable text-to-speech.

VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation Vall-t: Decoder-only generative transducer for robust and decoding- controllable text-to-speech

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:20:03.193256Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T14:19:52.092794Z digest=sha256:7847c4bcdda7721dc07814dd9887951b702a1e512308adcfccff7d73811773f7

Observation 669cdead-4f97-4776-88a7-4186167f048e · outbound

This paper cites Attention- constrained inference for robust decoder-only text-to-speech.2024 IEEE Spoken Language Technology Workshop (SLT), pages 630–637, 2024.

VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation Attention- constrained inference for robust decoder-only text-to-speech.2024 IEEE Spoken Language Technology Workshop (SLT), pages 630–637, 2024

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:20:02.969811Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T14:19:52.174655Z digest=sha256:d97df8d2cef4132ac1649acf0f0840e88772754b3adf31b23fbc8c6689cb235f

Observation 60758394-101e-43fd-8c43-d555c12e201d · outbound

This paper cites Improving Robustness of LLM-based Speech Synthesis by Learning Monotonic Alignment.

VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation Improving Robustness of LLM-based Speech Synthesis by Learning Monotonic Alignment

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T14:19:52.306333Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:19:52.306333Z digest=sha256:610d77b6e41713f42f5247e44589bd919fbc54fd2787a80a93791053ea354461

Observation 85484e04-365d-408c-a020-e216707c56b9 · outbound

This paper cites Speaking from Coarse to Fine: Improving Neural Codec Language Model via Multi-Scale Speech Coding and Generation.

VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation Speaking from Coarse to Fine: Improving Neural Codec Language Model via Multi-Scale Speech Coding and Generation

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T14:19:52.424117Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:19:52.424117Z digest=sha256:c1f25537ea476167145b4309d49ce4a2553dba8ce24c6e05e2c754bd29ea9def

Observation 32e5db4f-f25a-4307-9d78-607b7fadd358 · outbound

This paper cites Mega-TTS 2: Boosting prompting mechanisms for zero-shot speech synthesis.

VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation Mega-TTS 2: Boosting prompting mechanisms for zero-shot speech synthesis

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:20:02.826188Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T14:19:52.548815Z digest=sha256:a91ac2f04fa0d91d76188b4af30b2f6a15ea001b38a660c717c224ed6b5f9a0d

Observation ed023195-40ab-4345-b43f-4eda5a5f4a1a · outbound

This paper cites SSR-Speech: Towards Stable, Safe and Robust Zero-shot Text-based Speech Editing and Synthesis.

VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation SSR-Speech: Towards Stable, Safe and Robust Zero-shot Text-based Speech Editing and Synthesis

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T14:19:52.702338Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:19:52.702338Z digest=sha256:1354bba98d1d01cd94e6f7f8804a9b2a6a48229d0d2ab3a3a05b9524e3e5f1b9

Observation 873d0146-bc69-4a97-8ac2-5e2a681475df · outbound

This paper cites RALL-E: Robust Codec Language Modeling with Chain-of-Thought Prompting for Text-to-Speech Synthesis.

VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation RALL-E: Robust Codec Language Modeling with Chain-of-Thought Prompting for Text-to-Speech Synthesis

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T14:19:52.899498Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:19:52.899498Z digest=sha256:ccf06913ca4e42cfe4445a0f348f76c8da46d1883d9e246148b80cdc641865c0

Observation 35f44ca9-e9d1-47d2-974d-62fe1ca1d699 · outbound

This paper cites Enhancing Zero-shot Text-to-Speech Synthesis with Human Feedback.

VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation Enhancing Zero-shot Text-to-Speech Synthesis with Human Feedback

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T14:19:53.043874Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:19:53.043874Z digest=sha256:30f964190fe634ec726c8d9127334a6f8a31d856b75d04cc927e7c4a1acc8b94

Observation 4ff82b7b-3459-4cce-a038-ca8c65736997 · outbound

This paper cites Robust Zero-Shot Text-to-Speech Synthesis with Reverse Inference Optimization.

VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation Robust Zero-Shot Text-to-Speech Synthesis with Reverse Inference Optimization

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T14:19:53.169474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:19:53.169474Z digest=sha256:f8f58ac69dfc519881a82242ca1b5af869186b4332ff24ffa4f6f9d1dad060f9

Observation cb710981-e29d-40bd-8c39-1305b93c9d47 · outbound

This paper cites Desta, Roy Fejgin, Rafael Valle, and Jason Li.

VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation Desta, Roy Fejgin, Rafael Valle, and Jason Li

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:20:02.632742Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T14:19:53.311490Z digest=sha256:0632ce207e4af8584baf98c39ef4a049479ad8f06c586473e41a77b5a8eaf93a

Observation 7e1872e8-7f8e-468d-bf75-a5acd1a15bc1 · outbound

This paper cites Speak Foreign Languages with Your Own Voice: Cross-Lingual Neural Codec Language Modeling.

VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation Speak Foreign Languages with Your Own Voice: Cross-Lingual Neural Codec Language Modeling

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T14:19:53.462342Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:19:53.462342Z digest=sha256:53d64621d42027d9771b6abb5a9dda0c48918531d12c0740fc77720adaf0cfea

Observation 012daf19-969a-4d96-a9d1-43fba4969a32 · outbound

This paper cites Generative Pre-trained Speech Language Model with Efficient Hierarchical Transformer.

VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation Generative Pre-trained Speech Language Model with Efficient Hierarchical Transformer

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T14:19:53.626191Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:19:53.626191Z digest=sha256:f34f8214eb25ba49613d2e5ab355ce72b1231c8bd7f290afafbf7f65186a7cb5

Observation dc3048c9-080b-4860-b527-1fcdc0f87fc5 · outbound

This paper cites VioLA: Unified Codec Language Models for Speech Recognition, Synthesis, and Translation.

VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation VioLA: Unified Codec Language Models for Speech Recognition, Synthesis, and Translation

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T14:19:53.771128Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:19:53.771128Z digest=sha256:49b0782a0f07e95b4492dfadbe7563a708baba4ffb44c8483197791ec9618a67

Observation 6279db10-5903-4d71-adf5-822e051483e3 · outbound

This paper cites LauraGPT: Listen, Attend, Understand, and Regenerate Audio with GPT.

VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation LauraGPT: Listen, Attend, Understand, and Regenerate Audio with GPT

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T14:19:53.892566Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:19:53.892566Z digest=sha256:020a6a78c23268ab03f5ec4e79f0a9cfd1bf4df9d760ff4d99d440ceb89f8a9a

Observation f41a6ef4-5484-4196-b1e3-a003a49b9111 · outbound

This paper cites an unresolved cited work.

VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation Unresolved cited work

Reference 35

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:20:02.462834Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T14:19:54.016530Z digest=sha256:0d7313e5ad5d0f01b9ec404a46caaa6a515b50a2b0965d2ac77b6c00f866f079

Observation d7ab772c-e025-41bb-a22a-24ce39553164 · outbound

This paper cites SpeechComposer: Unifying Multiple Speech Tasks with Prompt Composition.

VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation SpeechComposer: Unifying Multiple Speech Tasks with Prompt Composition

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T14:19:54.204415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:19:54.204415Z digest=sha256:4bccf30ab66d034c0a054a481b9f80b4c44996ad97579e2648b7fb160a66c19e

Observation 84f6cce3-b598-495d-8db5-7f4ce98de632 · outbound

This paper cites Metis: A foundation speech generation model with masked generative pre-training.

VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation Metis: A foundation speech generation model with masked generative pre-training

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:20:02.247472Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T14:19:54.327352Z digest=sha256:9b5376f462eb48476247e0d389b7a856dd6559d7e4001af43cd8e2c0048f2e0c

Observation 556d9d7d-cac4-4f08-b791-226968ee62bb · outbound

This paper cites Prompttts: Controllable text-to-speech with text descriptions.ICASSP 2023 - 2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pages 1–5, 2022.

VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation Prompttts: Controllable text-to-speech with text descriptions.ICASSP 2023 - 2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pages 1–5, 2022

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:20:02.047892Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T14:19:54.422588Z digest=sha256:f4c069a51642bdc4d4526ecfe19363cbadcef12412b241dedbf352cd91797a71

Observation fc53d6c9-e5c1-4e46-8f80-0122508deb35 · outbound

This paper cites InstructTTS: Modelling Expressive TTS in Discrete Latent Space with Natural Language Style Prompt.

VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation InstructTTS: Modelling Expressive TTS in Discrete Latent Space with Natural Language Style Prompt

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T14:19:54.567317Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:19:54.567317Z digest=sha256:63f9ad65250028f76e30f4c998ef29a48620fefcfde109c17ccca33ba223f300

Observation 2ceec292-ba82-4ff9-8de3-dfdc53dd723a · outbound

This paper cites PromptStyle: Controllable Style Transfer for Text-to-Speech with Natural Language Descriptions.

VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation PromptStyle: Controllable Style Transfer for Text-to-Speech with Natural Language Descriptions

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T14:19:54.715981Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:19:54.715981Z digest=sha256:41841b547626366704bc3e95f72ee275299210bb5be9efa3a53902c2a1016ff8

Observation a3ac4ea4-86e6-4c5e-976e-593653ca213a · outbound

This paper cites TextrolSpeech: A Text Style Control Speech Corpus With Codec Language Text-to-Speech Models.

VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation TextrolSpeech: A Text Style Control Speech Corpus With Codec Language Text-to-Speech Models

Reference 41

Resolution
verified exact
local_arxiv, observed 2026-08-07T14:19:58.766245Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T14:19:54.931851Z digest=sha256:0204c368eb8703343556a0bfd28e8d0cdf5307511460c4cea2d898045f872845

Observation 1a8ba9e7-f22b-493f-b6e8-46229dd6c244 · outbound

This paper cites PromptTTS 2: Describing and Generating Voices with Text Prompt.

VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation PromptTTS 2: Describing and Generating Voices with Text Prompt

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T14:19:55.063985Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:19:55.063985Z digest=sha256:c8ab6f889aba5e032b85b24125dc763b432d602f1b171972df96dddcecb0a8c6

Observation a0230389-f536-413d-b453-4a4d8395e27d · outbound

This paper cites Natural language guidance of high-fidelity text-to-speech with synthetic annotations.

VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation Natural language guidance of high-fidelity text-to-speech with synthetic annotations

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T14:19:55.191205Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:19:55.191205Z digest=sha256:28a61ae79bfd1a49f2630b3ca1549d6fc79a259c782d90d61b15fc5f9bdfb5ea

Observation 44ff2b8f-9d24-4841-84b3-62414e426ed8 · outbound

This paper cites Scaling rich style-prompted text-to-speech datasets.

VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation Scaling rich style-prompted text-to-speech datasets

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:20:01.854731Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T14:19:55.310667Z digest=sha256:356907f840f6edbfd9780558e86746c5fa8f40a4f9f1f8740711e6640245384a

Observation 61f0d21b-5e1d-40cf-91ef-18199528b189 · outbound

This paper cites Moshi: a speech-text foundation model for real-time dialogue.

VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation Moshi: a speech-text foundation model for real-time dialogue

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T14:19:55.465301Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:19:55.465301Z digest=sha256:eeed101a1ae420cad36b12061f5901d7a7cc6d14a8126981ef668efac433041f

Observation 91447976-ac9f-4a61-9ca5-c61009dce3b5 · outbound

This paper cites Language Model Can Listen While Speaking.

VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation Language Model Can Listen While Speaking

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T14:19:55.581577Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:19:55.581577Z digest=sha256:a73b58e79b5e0744846990ccb4a10cafd0e6f3258c4a932774aac33fd3dda058

Observation 9ca4f5b7-4a33-40d1-aa74-c22a6e2656c5 · outbound

This paper cites LLaMA-Omni: Seamless Speech Interaction with Large Language Models.

VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation LLaMA-Omni: Seamless Speech Interaction with Large Language Models

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T14:19:55.644063Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:19:55.644063Z digest=sha256:325c4db4888a4788bc8f1a9c228e4e75d63e353b714bbbb1cf5be1072258c28e

Observation 0c878640-8d68-4e7b-97c0-52e1752594f0 · outbound

This paper cites SALMONN-omni: A Codec-free LLM for Full-duplex Speech Understanding and Generation.

VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation SALMONN-omni: A Codec-free LLM for Full-duplex Speech Understanding and Generation

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T14:19:55.723227Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:19:55.723227Z digest=sha256:ddf176faef3cead109d8f4f6d63cca4f02c1e4740c5a61ebffab9b3f775a3cf3

Observation 832e6d0f-0bf6-40fd-bae9-48da9426e755 · outbound

This paper cites Enabling Real-Time Conversations with Minimal Training Costs.

VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation Enabling Real-Time Conversations with Minimal Training Costs

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T14:19:55.833609Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:19:55.833609Z digest=sha256:04b9f830c4ea72e82861a45571c747ae935a4b7fc6638cc083caabf1339dc156

Observation b1889e61-ab6e-4d2e-92cd-ac6d3819c761 · outbound

This paper cites IntrinsicVoice: Empowering LLMs with Intrinsic Real-time Voice Interaction Abilities.

VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation IntrinsicVoice: Empowering LLMs with Intrinsic Real-time Voice Interaction Abilities

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T14:19:55.952954Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:19:55.952954Z digest=sha256:de76d7b1128ed6ff4d5b37fab446822eef3a488d45e5928d7c9427bbae18e7b1

Observation c1dad2e6-6205-48be-adc3-86345098b25a · outbound

This paper cites Llm-enhanced dialogue management for full-duplex spoken dialogue systems.

VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation Llm-enhanced dialogue management for full-duplex spoken dialogue systems

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:20:01.644408Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T14:19:56.035454Z digest=sha256:508aca7c849dd58cb65705f9bec92656372f1495a701f0fd59987f4404715740

Observation 5c38d1d3-bd8b-427f-bff7-00677543bc84 · outbound

This paper cites BASE TTS: Lessons from building a billion-parameter Text-to-Speech model on 100K hours of data.

VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation BASE TTS: Lessons from building a billion-parameter Text-to-Speech model on 100K hours of data

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T14:19:56.106020Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:19:56.106020Z digest=sha256:f837ec0aacf7ec5a2a9a8c9e45e2b6ddb58922e6ae29aecf223e357ffa97386f

Observation 992085bb-1e13-4e77-a897-c0c0e2e2511d · outbound

This paper cites HALL-E: Hierarchical Neural Codec Language Model for Minute-Long Zero-Shot Text-to-Speech Synthesis.

VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation HALL-E: Hierarchical Neural Codec Language Model for Minute-Long Zero-Shot Text-to-Speech Synthesis

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T14:19:56.181491Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:19:56.181491Z digest=sha256:b249a4ce9441b37e54bcfb32967c3bdaabf1aa5e0c44c696397c596c5bd1f916

Observation d5407115-8935-4156-98a6-2901d3393e33 · outbound

This paper cites NaturalSpeech 2: Latent Diffusion Models are Natural and Zero-Shot Speech and Singing Synthesizers.

VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation NaturalSpeech 2: Latent Diffusion Models are Natural and Zero-Shot Speech and Singing Synthesizers

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T14:19:56.312162Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:19:56.312162Z digest=sha256:f70e7a0566398b0a29ad891fcbc43777465e76c069ecaf7671a37bc8c2777220

Observation 556b0628-1098-4b0f-8d76-a2ba4395e922 · outbound

This paper cites Shih, Rohan Badlani, João Felipe Santos, Evelina Bakhturina, Mikyas T.

VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation Shih, Rohan Badlani, João Felipe Santos, Evelina Bakhturina, Mikyas T

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:20:01.415277Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T14:19:56.381071Z digest=sha256:89b34d782f2382946e76cb889cbe0ad13a8eb56968f7b4fb6e8055eb1efbe54e

Observation 37fd8c4d-8a2b-4ab1-87ab-b21dca27f849 · outbound

This paper cites Ditto-tts: Diffusion transformers for scalable text-to-speech without domain-specific factors.

VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation Ditto-tts: Diffusion transformers for scalable text-to-speech without domain-specific factors

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:20:01.204599Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T14:19:56.451386Z digest=sha256:c69fa9e3d8948d34b093f6a33b7c0a5c5162bcb62cbfcab2ef26506da1a8a62b

Observation 433d7614-33db-4942-975d-0ae475faa212 · outbound

This paper cites Peebles and Saining Xie.

VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation Peebles and Saining Xie

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T14:19:56.583065Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:19:56.583065Z digest=sha256:427dafdab509c10edea9875316af003380e4d35cdb7db4745216d0c046841d62

Observation 22457582-96a8-4eea-9a32-f5ecc442de3a · outbound

This paper cites Dmospeech: Direct metric optimization via distilled diffusion model in zero-shot speech synthesis.

VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation Dmospeech: Direct metric optimization via distilled diffusion model in zero-shot speech synthesis

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:20:00.990413Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T14:19:56.683947Z digest=sha256:acfd711016b7c161ce0ab2ee6cbdea4f86c36109997ba3d121b36fdf56932078

Observation 1e2b3cef-62ad-48af-b47a-1ddb4dcb323e · outbound

This paper cites E2 tts: Embarrassingly easy fully non-autoregressive zero-shot tts.2024 IEEE Spoken Language Technology Workshop (SLT), pages 682–689, 2024.

VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation E2 tts: Embarrassingly easy fully non-autoregressive zero-shot tts.2024 IEEE Spoken Language Technology Workshop (SLT), pages 682–689, 2024

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:20:00.801253Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T14:19:56.779894Z digest=sha256:521e0d29cc21a8ed9bd2618e8f5a59a452502eaad5924ec7fe68748337182ca8

Observation a77996d5-2d3b-4528-9b14-7bd57feddcde · outbound

This paper cites Audiobox: Unified Audio Generation with Natural Language Prompts.

VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation Audiobox: Unified Audio Generation with Natural Language Prompts

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T14:19:56.842801Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:19:56.842801Z digest=sha256:2c4e935c987e70ce7fd66f1ea16e7c376b0a26459045cc8492c069a0850d34ff

Observation 235bff6a-1867-49b6-87de-bd2096abb91f · outbound

This paper cites FireRedTTS: A Foundation Text-To-Speech Framework for Industry-Level Generative Speech Applications.

VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation FireRedTTS: A Foundation Text-To-Speech Framework for Industry-Level Generative Speech Applications

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-07T14:19:56.932091Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:19:56.932091Z digest=sha256:9578f85cd4fd73de6546a17613dbd195cb37d3b9254b1c47be295c33e51fe6f9

Observation 62d1ea48-ddeb-422a-afc6-9cd1cf666d36 · outbound

This paper cites Simple and Controllable Music Generation.

VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation Simple and Controllable Music Generation

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-07T14:19:57.022040Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:19:57.022040Z digest=sha256:ddd59e84c9b1f30a46ac4c16b6ce6f25f59e634274db88a113ef6866be65f6e5

Observation b5d23f41-e7c4-4aef-9075-65944122b465 · outbound

This paper cites Neural machine translation by jointly learning to align and translate, 2016.

VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation Neural machine translation by jointly learning to align and translate, 2016

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-07T14:19:57.089308Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:19:57.089308Z digest=sha256:1670e81a630d9856605c32d3c4c160b66238fa25e2c5319ae8b311139f76e4e4

Observation d3c1883d-8489-4d10-998e-13991ac28e25 · outbound

This paper cites Roformer: Enhanced transformer with rotary position embedding, 2023.

VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation Roformer: Enhanced transformer with rotary position embedding, 2023

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:20:00.668512Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T14:19:57.181904Z digest=sha256:b2d83ee6657efecaf9d474fc8424b4dd131e68128b10c20ac4df4261fb977892

Observation deee603a-0ce1-4795-8030-32fc1e0db0d1 · outbound

This paper cites Gomez, Lukasz Kaiser, and Illia Polosukhin.

VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation Gomez, Lukasz Kaiser, and Illia Polosukhin

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-07T14:19:57.278975Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:19:57.278975Z digest=sha256:d4e40a9f54498a924c24520d50874e2a7a3a958d1afb6c600dbe711b9020f360

Observation 0df4e2a8-decb-4308-9f13-f80961a67480 · outbound

This paper cites an unresolved cited work.

VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation Unresolved cited work

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-07T14:19:57.355945Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:19:57.355945Z digest=sha256:4bd1eb9308ac8497ad7c7c4607fcb881c93759c9aca46d2ef03598b0bb0aca2b

Observation a4750667-7041-41f0-84fe-797038ed46c6 · outbound

This paper cites Smith, and Mike Lewis.

VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation Smith, and Mike Lewis

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-07T14:19:57.434211Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:19:57.434211Z digest=sha256:04678cc07628acbcb18367b05aee6e72607dba5dd3d14efdb2950abda8a77929

Observation 032496ba-6547-4544-80ef-b8c1d4fbca3b · outbound

This paper cites Emilia: An extensive, multilingual, and diverse speech dataset for large-scale speech generation.2024 IEEE Spoken Language Technology Workshop (SLT), pages 885–890, 2024.

VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation Emilia: An extensive, multilingual, and diverse speech dataset for large-scale speech generation.2024 IEEE Spoken Language Technology Workshop (SLT), pages 885–890, 2024

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:20:00.546516Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T14:19:57.537109Z digest=sha256:b790d3af1fa65d737fdfde8a13daa943e35bbda007f62f772729b37d1379dd7e

Observation 99fa98a6-c588-438f-b2d3-69705bb847a2 · outbound

This paper cites an unresolved cited work.

VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation Unresolved cited work

Reference 69

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:20:00.387300Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T14:19:57.594796Z digest=sha256:2f8ba4f8f0564cf97bc26445511fcbc80f13653c9c8e826bdd36c542ba70d15b

Observation 0b8773f0-7299-4522-ab03-186b7534942e · outbound

This paper cites eSpeak NG: Speech synthesiser.https://github.com/espeak-ng/espeak-ng.

VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation eSpeak NG: Speech synthesiser.https://github.com/espeak-ng/espeak-ng

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:20:00.232493Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T14:19:57.695586Z digest=sha256:f22e61be42428f10bec40838f86419ec0e52ce3e6246d1b02f5d272c916b6297

Observation 8d541a03-a710-477c-a4db-f2ce42905748 · outbound

This paper cites Seed-TTS: A Family of High-Quality Versatile Speech Generation Models.

VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation Seed-TTS: A Family of High-Quality Versatile Speech Generation Models

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-07T14:19:57.780080Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:19:57.780080Z digest=sha256:602a32647fed25af54902ee26ea19eeedc9f71a697cda0da007bda2f421d7a6d

Observation fc9d4cbb-a552-4058-8c24-680bc8651cad · outbound

This paper cites Zipformer: A faster and better encoder for automatic speech recognition.

VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation Zipformer: A faster and better encoder for automatic speech recognition

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:20:00.042837Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T14:19:57.837727Z digest=sha256:5e051abf7f90f808660e8d608f8f916768e340846a343ebdbb78ec188101b54b

Observation cc47ff8e-9be7-4e0f-aaca-e4557fceaf5d · outbound

This paper cites Spark-tts: An efficient llm-based text-to-speech model with single-stream decoupled speech tokens.

VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation Spark-tts: An efficient llm-based text-to-speech model with single-stream decoupled speech tokens

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:19:59.926601Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T14:19:57.919375Z digest=sha256:a7d8f4eab7519fed11516ec3e57ddfa8ed5b352ecc97266446e0ce26445f8f1c

Observation b8179d0d-46ab-486a-a30d-a4ba1206f51a · outbound

This paper cites Utmos: Utokyo-sarulab system for voicemos challenge 2022, 2022.

VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation Utmos: Utokyo-sarulab system for voicemos challenge 2022, 2022

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:19:59.759952Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T14:19:58.017967Z digest=sha256:5225d1c92c4a15d4121d8fe7af6afe63347ef4c377b10286f37c32ffde606170

Observation 7ffd8915-540a-457e-ad2a-feebe127b0e5 · outbound

This paper cites Robust Speech Recognition via Large-Scale Weak Supervision.

VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation Robust Speech Recognition via Large-Scale Weak Supervision

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-07T14:19:58.084628Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:19:58.084628Z digest=sha256:b1d774b4e8a3055ed213c91ae082727d842614ace3014c3d1edba79c64c661c2

Observation 9774d91c-d74b-419e-80da-0d25f97b1e5a · outbound

This paper cites Wavlm: Large-scale self-supervised pre-training for full stack speech processing.IEEE Journal of Selected Topics in Signal Processing, 16:1505–1518, 2021.

VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation Wavlm: Large-scale self-supervised pre-training for full stack speech processing.IEEE Journal of Selected Topics in Signal Processing, 16:1505–1518, 2021

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:19:59.605196Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T14:19:58.176788Z digest=sha256:3b6275d21cb07ad29538cf7d4bb077752cdc565e603518deefb50a2ef861e23e

Observation fcdb7826-d3bc-4038-8c87-1cdf16295d18 · outbound

This paper cites Speech quality assessment.

VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation Speech quality assessment

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:19:59.455875Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T14:19:58.249319Z digest=sha256:6d5d9cc892109fcbf206ea6716c87e2bf9598236743deaa4ede423c04b1e8591

Observation 8d91192c-2adb-4c99-bb69-82eb63c45c81 · outbound

This paper cites Llm.int8(): 8-bit matrix multiplication for transformers at scale, 2022.

VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation Llm.int8(): 8-bit matrix multiplication for transformers at scale, 2022

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-07T14:19:58.330976Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:19:58.330976Z digest=sha256:83db3d839b05b7655856de50fd82e382a416ffc99904a417afd3c53d01c89c0f

Observation bcfe4542-9f26-42fc-a7fd-831a2e9c5aa1 · outbound

This paper cites Fast and high-quality auto- regressive speech synthesis via speculative decoding.

VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation Fast and high-quality auto- regressive speech synthesis via speculative decoding

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:19:59.295479Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T14:19:58.406514Z digest=sha256:e948f22309a4387eba0809eb2a534314c1a294521329167c75c43956d1eef4f0

Observation 0fa78f88-7c67-4229-9457-1550633ce4dd · outbound

This paper cites natural” in comparative naturalness task with “similar.

VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation natural” in comparative naturalness task with “similar

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:19:59.140627Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T14:19:58.490777Z digest=sha256:c79b76c2b6aa572e93b37dc0690a533ea7a82dcecda06c17fb6b0cd6afabfce8

Pith citing papers

Observation 4a70ccd9-9a71-43f5-a190-b1e16e566dd0 · inbound

Bridging the Stability-Expressivity Gap: Synthetic Data Scaling and Preference Alignment for Low-Resource Spoken Language Models cites this paper.

Bridging the Stability-Expressivity Gap: Synthetic Data Scaling and Preference Alignment for Low-Resource Spoken Language Models VoiceStar: Robust Zero-Shot Autoregressive TTS with Duration Control and Extrapolation

Reference 21

Resolution
unresolved
no resolver link, observed 2026-07-12T23:22:40.218077Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T23:22:40.218077Z digest=sha256:6c25ddc14d96c00bffa2118f58d11a7bdd9dabae9363303398e7f322cf5f6814