Pith. sign in

Paper Citation Record · LEDGER

Robust and Efficient Autoregressive Speech Synthesis with Dynamic Chunk-wise Prediction Policy

As of 19 August 2026, this Paper Citation Record lists 50 of 50 outbound references and 2 inbound Pith citation observations for arXiv:2506.22023.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.22023 v1

Coverage vector

measured 50 of 50 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T22:19:03.875806Z

measured 52 of 52 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-27T01:05:20.432379Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T20:58:57.478873Z

Reference resolution

50 of 50 outbound references displayed

  • verified exact2
  • verified fuzzy18
  • unresolved30
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation a5e68496-bb31-450b-84ca-0952e784f0c2 · outbound

This paper cites A Survey of Large Language Models.

Robust and Efficient Autoregressive Speech Synthesis with Dynamic Chunk-wise Prediction Policy A Survey of Large Language Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T22:18:58.451990Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:18:58.451990Z digest=sha256:8e91bff3a1883a06358cc090ea6b6c7b7ad38866bc20505c86f385c54b4aa004

Observation cb1ccaf6-217f-4166-bf0c-6b91003ee642 · outbound

This paper cites Codec-SUPERB: An In-Depth Analysis of Sound Codec Models.

Robust and Efficient Autoregressive Speech Synthesis with Dynamic Chunk-wise Prediction Policy Codec-SUPERB: An In-Depth Analysis of Sound Codec Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T22:18:58.537491Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:18:58.537491Z digest=sha256:7af1540531a9f4c551d4707c6683147abbb86516a9ff8147dd75fd45881a4713

Observation bd8c8cb0-03ac-476c-8e5b-b72ad47d999e · outbound

This paper cites Recent advances in discrete speech tokens: A review,.

Robust and Efficient Autoregressive Speech Synthesis with Dynamic Chunk-wise Prediction Policy Recent advances in discrete speech tokens: A review,

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T22:18:58.693492Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:18:58.693492Z digest=sha256:b8bf9d2cf0194d0fcfeab0667eafd8b09c4be008f229206047494c525cedc297

Observation 8ef43034-e09c-47b0-b03f-f0e7494a5942 · outbound

This paper cites On The Landscape of Spoken Language Models: A Comprehensive Survey.

Robust and Efficient Autoregressive Speech Synthesis with Dynamic Chunk-wise Prediction Policy On The Landscape of Spoken Language Models: A Comprehensive Survey

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T22:18:58.895945Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:18:58.895945Z digest=sha256:1292eaadde22372012f5793b2e0120ac767572e867e71f8e5eb14d2e1477120a

Observation 170c9bb8-d83e-4618-bde3-c75ad123e85e · outbound

This paper cites Neural Discrete Representation Learning,.

Robust and Efficient Autoregressive Speech Synthesis with Dynamic Chunk-wise Prediction Policy Neural Discrete Representation Learning,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:19:09.288020Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T22:18:59.021254Z digest=sha256:d6f371dd285283d4c0b0606e0e1b3a452140df4270bae0c2475e56247d5c8ff3

Observation 7d465b53-a010-46e1-83f7-f162dcbfc1dd · outbound

This paper cites Fast and high-quality auto-regressive speech synthesis via speculative decoding,.

Robust and Efficient Autoregressive Speech Synthesis with Dynamic Chunk-wise Prediction Policy Fast and high-quality auto-regressive speech synthesis via speculative decoding,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:19:09.091961Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T22:18:59.166926Z digest=sha256:83c5a986a5a833c5b445d436f79552051127001ba7154c463e240e8bf51b66fb

Observation e7cfcf5d-e68c-4595-8ce2-625a9860e24b · outbound

This paper cites Accelerating Codec-based Speech Synthesis with Multi-Token Prediction and Speculative Decoding.

Robust and Efficient Autoregressive Speech Synthesis with Dynamic Chunk-wise Prediction Policy Accelerating Codec-based Speech Synthesis with Multi-Token Prediction and Speculative Decoding

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-08-06T22:19:04.696229Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T22:18:59.328921Z digest=sha256:72b3a37100d366c7c93e79cd424f366fdfbca2c975c1ef44807111a8374283b1

Observation 7f734fbc-4e28-473c-b2b5-189bfff8215b · outbound

This paper cites VocalNet: Speech LLM with Multi-Token Prediction for Faster and High-Quality Generation.

Robust and Efficient Autoregressive Speech Synthesis with Dynamic Chunk-wise Prediction Policy VocalNet: Speech LLM with Multi-Token Prediction for Faster and High-Quality Generation

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T22:18:59.507675Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:18:59.507675Z digest=sha256:8b3e313cf48278f2519329a930cb99f64399dd78c5ea62848d156fa665d4f772

Observation 4b95c2a2-7701-426e-9b56-1e2a926f9b7d · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Robust and Efficient Autoregressive Speech Synthesis with Dynamic Chunk-wise Prediction Policy DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T22:18:59.608837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:18:59.608837Z digest=sha256:255c74b5daac87c042da818bea5af4d155a34dd9d8708e0e672d3fd6138b864b

Observation ac48b1a8-3cfd-458a-8666-e40d88a486a5 · outbound

This paper cites SFT Memorizes, RL Generalizes: A Comparative Study of Foundation Model Post-training.

Robust and Efficient Autoregressive Speech Synthesis with Dynamic Chunk-wise Prediction Policy SFT Memorizes, RL Generalizes: A Comparative Study of Foundation Model Post-training

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T22:18:59.735703Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:18:59.735703Z digest=sha256:b285f0c02627683053914009e0d0de5113a8765a9ddc54bb667adaf506ffef3c

Observation 0eaed673-25f5-44ed-9d4a-9084025b75a2 · outbound

This paper cites Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?.

Robust and Efficient Autoregressive Speech Synthesis with Dynamic Chunk-wise Prediction Policy Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T22:18:59.877417Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:18:59.877417Z digest=sha256:e64721ce968a0912098191a6f52ee53d3ce541252aa17bcb505e88e8210ece86

Observation 10172a57-2cf3-4b8f-a1e9-81406f67e2dd · outbound

This paper cites LibriTTS: A Corpus Derived from LibriSpeech for Text-to- Speech,.

Robust and Efficient Autoregressive Speech Synthesis with Dynamic Chunk-wise Prediction Policy LibriTTS: A Corpus Derived from LibriSpeech for Text-to- Speech,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:19:08.859990Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T22:19:00.015587Z digest=sha256:9283d2a7c6768d646f8db67e52ef76a3247dd3638859a21a9140e1f64f600ed2

Observation b04355ef-2ce0-4264-8d3c-142da83fbb8a · outbound

This paper cites HuBERT: Self-Supervised Speech Representation Learning by Masked Prediction of Hidden Units,.

Robust and Efficient Autoregressive Speech Synthesis with Dynamic Chunk-wise Prediction Policy HuBERT: Self-Supervised Speech Representation Learning by Masked Prediction of Hidden Units,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:19:08.546115Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T22:19:00.124151Z digest=sha256:189e4a875196f73cae75eef48b109d483f6d44c30e0455ac259da98c7c95124b

Observation f64dcb36-58b7-422b-83b9-42bbada1cba2 · outbound

This paper cites CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens.

Robust and Efficient Autoregressive Speech Synthesis with Dynamic Chunk-wise Prediction Policy CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T22:19:00.278755Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:19:00.278755Z digest=sha256:30eae9890bc70132e2db178cb13d26d8d049f1d77b58a467a51c4be352cc7d0e

Observation cca8f652-0bf6-4cff-9abb-116298515508 · outbound

This paper cites Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers,.

Robust and Efficient Autoregressive Speech Synthesis with Dynamic Chunk-wise Prediction Policy Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:19:08.260615Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T22:19:00.426911Z digest=sha256:eb646079ef657dc5dc3f367f9af403bd8040435e3fb3d74f4c622194b555ce44

Observation cbe6412f-31f3-4fdc-b33f-731c6b255e13 · outbound

This paper cites V oiceCraft: Zero-Shot Speech Editing and Text-to-Speech in the Wild,.

Robust and Efficient Autoregressive Speech Synthesis with Dynamic Chunk-wise Prediction Policy V oiceCraft: Zero-Shot Speech Editing and Text-to-Speech in the Wild,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:19:07.967672Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T22:19:00.547714Z digest=sha256:50759d73f8cd68bcfbebdbbce931b306cc920b13bc121d901f126fb651ddc9e5

Observation 3fff0af9-2b2b-4938-aa93-41cff71acb69 · outbound

This paper cites Seed-TTS: A Family of High-Quality Versatile Speech Generation Models.

Robust and Efficient Autoregressive Speech Synthesis with Dynamic Chunk-wise Prediction Policy Seed-TTS: A Family of High-Quality Versatile Speech Generation Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T22:19:00.703156Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:19:00.703156Z digest=sha256:c17f3803da68d3fee6ac8b132be8060b063af4504b069befb9124d93cd59fb2a

Observation 45eb768c-4d27-4724-8018-1a3d215099e3 · outbound

This paper cites BASE TTS: Lessons from building a billion-parameter Text-to-Speech model on 100K hours of data.

Robust and Efficient Autoregressive Speech Synthesis with Dynamic Chunk-wise Prediction Policy BASE TTS: Lessons from building a billion-parameter Text-to-Speech model on 100K hours of data

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T22:19:00.848288Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:19:00.848288Z digest=sha256:fd083f9095b8f11f88570a68f49e0f2a0323401d215aa65b4122155edc39ce32

Observation 0f544de7-5d73-42a3-aa1b-fde17836f6c3 · outbound

This paper cites ELLA-V: Stable Neural Codec Language Modelling with Alignment-Guided Sequence Reordering,.

Robust and Efficient Autoregressive Speech Synthesis with Dynamic Chunk-wise Prediction Policy ELLA-V: Stable Neural Codec Language Modelling with Alignment-Guided Sequence Reordering,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:19:07.735392Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T22:19:00.908080Z digest=sha256:390719526b2d0f8d366ebe027015703205cacf08e836dc50b9576edd5709a04a

Observation da940102-74c5-4465-8666-46c1934a7c50 · outbound

This paper cites VALL-E R: Robust and Efficient Zero-Shot Text-to-Speech Synthesis via Monotonic Alignment.

Robust and Efficient Autoregressive Speech Synthesis with Dynamic Chunk-wise Prediction Policy VALL-E R: Robust and Efficient Zero-Shot Text-to-Speech Synthesis via Monotonic Alignment

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T22:19:01.038507Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:19:01.038507Z digest=sha256:a17a1c92e5d4a53780628255cfc0fbab570c6e43d6b6ed3697750c6f7d68f51a

Observation 8b915efd-50a3-4683-9d04-669e86c5673f · outbound

This paper cites V ALL-T: Decoder-only generative transducer for robust and decoding-controllable text-to-speech,.

Robust and Efficient Autoregressive Speech Synthesis with Dynamic Chunk-wise Prediction Policy V ALL-T: Decoder-only generative transducer for robust and decoding-controllable text-to-speech,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:19:07.508375Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T22:19:01.123238Z digest=sha256:0b3e5cae3adde1933d5ec8e3acd7a4d34f87f15c63e2b1ce6fda5b51bb233d80

Observation 0b967284-e98e-45fc-b7af-9a6177e2998a · outbound

This paper cites RALL-E: Robust Codec Language Modeling with Chain-of-Thought Prompting for Text-to-Speech Synthesis.

Robust and Efficient Autoregressive Speech Synthesis with Dynamic Chunk-wise Prediction Policy RALL-E: Robust Codec Language Modeling with Chain-of-Thought Prompting for Text-to-Speech Synthesis

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T22:19:01.200105Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:19:01.200105Z digest=sha256:321c2a0bd4746190d803b59886e01b77a09c5d062126ebcc72b8cb9f0d84e63a

Observation 2c9ca762-3128-47f8-95e8-b181b5323ef8 · outbound

This paper cites SNAC: Multi-Scale Neural Audio Codec.

Robust and Efficient Autoregressive Speech Synthesis with Dynamic Chunk-wise Prediction Policy SNAC: Multi-Scale Neural Audio Codec

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T22:19:01.314755Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:19:01.314755Z digest=sha256:b89e5c3643a5af5190bed5cc6694a7529120fd78ea11b8ae3067529094ccb66e

Observation 9ff7cc4e-5338-474b-a2b5-c427656fce26 · outbound

This paper cites Speaking from Coarse to Fine: Improving Neural Codec Language Model via Multi-Scale Speech Coding and Generation.

Robust and Efficient Autoregressive Speech Synthesis with Dynamic Chunk-wise Prediction Policy Speaking from Coarse to Fine: Improving Neural Codec Language Model via Multi-Scale Speech Coding and Generation

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-08-06T22:19:04.337592Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T22:19:01.383749Z digest=sha256:ffa2d8ba23602ed1bd6311f661fba7cc12542125f3b8d85e47026590233c022f

Observation 453300fe-b9e7-4766-bb5e-e29390984124 · outbound

This paper cites UniAudio 1.5: Large Language Model-Driven Audio Codec is A Few-Shot Audio Task Learner,.

Robust and Efficient Autoregressive Speech Synthesis with Dynamic Chunk-wise Prediction Policy UniAudio 1.5: Large Language Model-Driven Audio Codec is A Few-Shot Audio Task Learner,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:19:07.260505Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T22:19:01.499041Z digest=sha256:23cdfb8ab02c5d9bf62481d7180918176296fca044bd366a0e166b0713e143b3

Observation 623308b0-5410-4c93-921d-76604959d8cc · outbound

This paper cites SpeechTokenizer: Unified Speech Tokenizer for Speech Language Models,.

Robust and Efficient Autoregressive Speech Synthesis with Dynamic Chunk-wise Prediction Policy SpeechTokenizer: Unified Speech Tokenizer for Speech Language Models,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:19:07.030523Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T22:19:01.620066Z digest=sha256:1ea5f2261ee39811d49a53c54239ffac876c7197be80e54f060fc859397a968b

Observation 541ac633-3efc-459c-82c6-140d69c2f425 · outbound

This paper cites Moshi: a speech-text foundation model for real-time dialogue.

Robust and Efficient Autoregressive Speech Synthesis with Dynamic Chunk-wise Prediction Policy Moshi: a speech-text foundation model for real-time dialogue

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T22:19:01.760430Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:19:01.760430Z digest=sha256:c2986dd0cc1ccfd1d8ebf5fa0d073ee5740f2cf67120edb8c1bd26e2203df0d0

Observation f31aca70-c936-420a-b30f-4a1e6c6ec694 · outbound

This paper cites AudioLM: A Language Modeling Approach to Audio Generation,.

Robust and Efficient Autoregressive Speech Synthesis with Dynamic Chunk-wise Prediction Policy AudioLM: A Language Modeling Approach to Audio Generation,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:19:06.814846Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T22:19:01.801928Z digest=sha256:acadb391c9361de46bc7c4dec9ba1074fd937483871bab185cfd72eee3afaaa7

Observation 87d7d26a-7d45-4b9e-a157-954651f6f045 · outbound

This paper cites Speak, Read and Prompt: High-Fidelity Text-to- Speech with Minimal Supervision,.

Robust and Efficient Autoregressive Speech Synthesis with Dynamic Chunk-wise Prediction Policy Speak, Read and Prompt: High-Fidelity Text-to- Speech with Minimal Supervision,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:19:06.546180Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T22:19:01.863234Z digest=sha256:8e5e4a81fc0efdb3f2072da0154a50810b669f6604c7d4df3357ba7c3e695d2e

Observation 212b35dc-657a-48f6-ae02-a3db385f2a66 · outbound

This paper cites MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer,.

Robust and Efficient Autoregressive Speech Synthesis with Dynamic Chunk-wise Prediction Policy MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:19:06.308846Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T22:19:01.965874Z digest=sha256:001923045ed0d433973816a16ee17b4290281e5295c6474ab229e3502bef66f9

Observation 021b628d-e59d-41e5-9d24-abc2e64e2af3 · outbound

This paper cites Medusa: Simple LLM Inference Acceleration Framework with Multiple Decoding Heads.

Robust and Efficient Autoregressive Speech Synthesis with Dynamic Chunk-wise Prediction Policy Medusa: Simple LLM Inference Acceleration Framework with Multiple Decoding Heads

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T22:19:02.032915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:19:02.032915Z digest=sha256:35502272ba5cc95ff05b0b0ebbd10af11ad387e6666bdd404346b98809890865

Observation 36d5de7f-a8ec-4f5a-bd80-308858dd3cb8 · outbound

This paper cites Better & Faster Large Language Models via Multi-token Prediction.

Robust and Efficient Autoregressive Speech Synthesis with Dynamic Chunk-wise Prediction Policy Better & Faster Large Language Models via Multi-token Prediction

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T22:19:02.090093Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:19:02.090093Z digest=sha256:f59d2e2a54d0a0fb757f406789b552f1fc734174983de0606ae3fdfd450e9c95

Observation 01db0498-5f06-4f96-ac2c-35d8cbbe18d4 · outbound

This paper cites DeepSeek-V3 Technical Report.

Robust and Efficient Autoregressive Speech Synthesis with Dynamic Chunk-wise Prediction Policy DeepSeek-V3 Technical Report

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T22:19:02.210858Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:19:02.210858Z digest=sha256:be8cb7a2907017dc2ddc07989675dfb263fb073b1da0e66430e19aeaec6c2802

Observation 034f4c16-c96a-4c01-8782-c94c93b10491 · outbound

This paper cites Training language models to follow instructions with human feedback,.

Robust and Efficient Autoregressive Speech Synthesis with Dynamic Chunk-wise Prediction Policy Training language models to follow instructions with human feedback,

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T22:19:02.317947Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:19:02.317947Z digest=sha256:a63ccb95b4822ae3929fa8fdb9aa008034211ec09a5a23f0890dd123c174cd1d

Observation b1f83a33-32f4-4d46-ad4d-407412ceb971 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Robust and Efficient Autoregressive Speech Synthesis with Dynamic Chunk-wise Prediction Policy Proximal Policy Optimization Algorithms

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T22:19:02.460098Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:19:02.460098Z digest=sha256:1b93040e4635ff15ffaeebc74a23c31d08667a5113cf6b4fee0a906e3a10a197

Observation 8b2ec461-051b-48dd-a584-65794bf75d01 · outbound

This paper cites Direct preference optimization: Your language model is secretly a reward model,.

Robust and Efficient Autoregressive Speech Synthesis with Dynamic Chunk-wise Prediction Policy Direct preference optimization: Your language model is secretly a reward model,

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T22:19:02.578605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:19:02.578605Z digest=sha256:055951961f78e222949cd04819f38271e6ff9864bff0e00144cbd167a42ba5e4

Observation ed5e40cf-f30d-403d-b151-12f36ac6cea7 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Robust and Efficient Autoregressive Speech Synthesis with Dynamic Chunk-wise Prediction Policy DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T22:19:02.652495Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:19:02.652495Z digest=sha256:c72f3ca70fe659247c005b613178baedbfd31e1cde285969150b1a0dc792aa09

Observation b38bdcb0-3913-450b-8c11-8b7e27459dd7 · outbound

This paper cites Align-SLM: Textless Spoken Language Models with Reinforcement Learning from AI Feedback.

Robust and Efficient Autoregressive Speech Synthesis with Dynamic Chunk-wise Prediction Policy Align-SLM: Textless Spoken Language Models with Reinforcement Learning from AI Feedback

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T22:19:02.738861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:19:02.738861Z digest=sha256:a36c2879f8072e6b74c2c3f277a92a2b7c02c0cbb0dc3e2d5a9aba349bfa2d69

Observation 76cb5b0b-a16d-4a52-87f3-b884024da7cf · outbound

This paper cites Robust Zero-Shot Text-to-Speech Synthesis with Reverse Inference Optimization.

Robust and Efficient Autoregressive Speech Synthesis with Dynamic Chunk-wise Prediction Policy Robust Zero-Shot Text-to-Speech Synthesis with Reverse Inference Optimization

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T22:19:02.830612Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:19:02.830612Z digest=sha256:b95f19ab7360ddd1eb33a1bd63f989efb794d3571ca745edad3e7dc50dd50f56

Observation aff5959b-423b-4aed-a646-6c0e663d169a · outbound

This paper cites Enhancing Zero-shot Text-to-Speech Synthesis with Human Feedback.

Robust and Efficient Autoregressive Speech Synthesis with Dynamic Chunk-wise Prediction Policy Enhancing Zero-shot Text-to-Speech Synthesis with Human Feedback

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T22:19:02.888700Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:19:02.888700Z digest=sha256:bae9f7e6304fe29d1c5e709b33ec47f5086584b79dc6d7ed56f7f10fdd4dff3a

Observation 576192fd-545d-423f-a2fd-742d69e6552d · outbound

This paper cites SpeechAlign: Aligning speech generation to human preferences,.

Robust and Efficient Autoregressive Speech Synthesis with Dynamic Chunk-wise Prediction Policy SpeechAlign: Aligning speech generation to human preferences,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:19:06.010259Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T22:19:02.963953Z digest=sha256:c8d8b930aca1af37fbec74662923ec8c3801ee745a758da5710b40777aa35e9b

Observation 71a4a6ff-5a9b-4326-9b1f-18cdfee7b83c · outbound

This paper cites Fine-grained preference optimization improves zero-shot text-to-speech,.

Robust and Efficient Autoregressive Speech Synthesis with Dynamic Chunk-wise Prediction Policy Fine-grained preference optimization improves zero-shot text-to-speech,

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T22:19:03.066762Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:19:03.066762Z digest=sha256:30dda5fad924571dce7a3e0babddd1f9d1ed386a0c93121117f8e858caa5743f

Observation 9a995078-6c79-4198-9ada-73e1eedc2b32 · outbound

This paper cites VALL-E 2: Neural Codec Language Models are Human Parity Zero-Shot Text to Speech Synthesizers.

Robust and Efficient Autoregressive Speech Synthesis with Dynamic Chunk-wise Prediction Policy VALL-E 2: Neural Codec Language Models are Human Parity Zero-Shot Text to Speech Synthesizers

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T22:19:03.126226Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:19:03.126226Z digest=sha256:81205e5ddbb264486deea8e91c4ec05e26f244a4b9e1b24b4c600aea7c6060e5

Observation f9935a8d-bf66-4312-a1fc-ca1e4ab6f111 · outbound

This paper cites UniCATS: A Unified Context-Aware Text-to-Speech Framework with Contextual VQ-Diffusion and V ocoding,.

Robust and Efficient Autoregressive Speech Synthesis with Dynamic Chunk-wise Prediction Policy UniCATS: A Unified Context-Aware Text-to-Speech Framework with Contextual VQ-Diffusion and V ocoding,

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:19:05.793544Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T22:19:03.250253Z digest=sha256:00d1aa418c6a17e624b2788178ec3cb368bc9a661409a84ee079d3bd4982de52

Observation e51de082-c27f-4ccf-9cef-17c1e6e69518 · outbound

This paper cites LSCodec: Low-Bitrate and Speaker-Decoupled Discrete Speech Codec.

Robust and Efficient Autoregressive Speech Synthesis with Dynamic Chunk-wise Prediction Policy LSCodec: Low-Bitrate and Speaker-Decoupled Discrete Speech Codec

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T22:19:03.356798Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:19:03.356798Z digest=sha256:c7ca3e311e5e1c69af2200ff420dbd115685c99b7bd74f1431b9b67af87eca21

Observation 65daf919-b05e-4438-8033-4f3a98cb4407 · outbound

This paper cites WavLM: Large-Scale Self-Supervised Pre-Training for Full Stack Speech Processing,.

Robust and Efficient Autoregressive Speech Synthesis with Dynamic Chunk-wise Prediction Policy WavLM: Large-Scale Self-Supervised Pre-Training for Full Stack Speech Processing,

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:19:05.611424Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T22:19:03.471288Z digest=sha256:3d56f16ae5401d24499bb148378b8545df75840b8e10d421b78db26bb0b30e83

Observation d8c83b8d-7981-4eae-a62d-7fefc3f3a776 · outbound

This paper cites Sigmoid-weighted linear units for neural network function approximation in reinforcement learning,.

Robust and Efficient Autoregressive Speech Synthesis with Dynamic Chunk-wise Prediction Policy Sigmoid-weighted linear units for neural network function approximation in reinforcement learning,

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-06T22:19:03.560664Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:19:03.560664Z digest=sha256:6f7b9d226e5d448def22d27ba49b93a46460230ac01b541e15353744f3690cc2

Observation d2d0d98b-d896-40cd-99c4-b34e37f99f92 · outbound

This paper cites UTMOS: UTokyo- SaruLab System for V oiceMOS Challenge 2022,.

Robust and Efficient Autoregressive Speech Synthesis with Dynamic Chunk-wise Prediction Policy UTMOS: UTokyo- SaruLab System for V oiceMOS Challenge 2022,

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:19:05.438167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T22:19:03.693568Z digest=sha256:949eb3b8b27d2881c4aeaf1d2e626b8727c95b669ff6ef1d7cc9b7f0835d0c20

Observation 75ad788d-482b-47b3-9388-04dc6c7cf032 · outbound

This paper cites Robust Speech Recognition via Large-Scale Weak Supervision,.

Robust and Efficient Autoregressive Speech Synthesis with Dynamic Chunk-wise Prediction Policy Robust Speech Recognition via Large-Scale Weak Supervision,

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:19:05.082403Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-06T22:19:03.772752Z digest=sha256:92cc879590cb14f10ea0c292bda6399f72b0cbc3ab3dd85fa936d20438e24260

Observation 5835c07c-950d-4118-957f-594a85fa128b · outbound

This paper cites Zipformer: A faster and better encoder for automatic speech recognition.

Robust and Efficient Autoregressive Speech Synthesis with Dynamic Chunk-wise Prediction Policy Zipformer: A faster and better encoder for automatic speech recognition

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T22:19:03.875806Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:19:03.875806Z digest=sha256:065998077f6654622563cf5ba40dd0614c974bc97f52ced140476e4b37dcc8d5

Pith citing papers

Observation 03bae24a-489b-4704-8727-8f4f7b6df872 · inbound

From Static Inference to Dynamic Interaction: A Survey of Streaming Large Language Models cites this paper.

From Static Inference to Dynamic Interaction: A Survey of Streaming Large Language Models Robust and Efficient Autoregressive Speech Synthesis with Dynamic Chunk-wise Prediction Policy

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-15T16:20:09.991805Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-15T16:16:15.819622Z digest=sha256:1d3544ea3dbe7b653b25b4b9cc3d2c212248b5196d9f87e587899edda43b6aa3

Observation e5d064c5-efdc-4ec2-aab7-be973da6cf34 · inbound

ASTRA: A Scalable Next-Generation ATCO Training Simulator with Autonomous Simpilots cites this paper.

ASTRA: A Scalable Next-Generation ATCO Training Simulator with Autonomous Simpilots Robust and Efficient Autoregressive Speech Synthesis with Dynamic Chunk-wise Prediction Policy

Reference 55

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T20:58:57.480674Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-06-27T01:05:20.432379Z digest=sha256:17d99e1a264cbd1101e8c146bb53cdef02dddc77aae022cedfc3091392b0ca67