Pith. sign in

Paper Citation Record · LEDGER

Robust and Efficient Autoregressive Speech Synthesis with Dynamic Chunk-wise Prediction Policy

As of 7 August 2026, this Paper Citation Record lists 50 of 50 outbound references and 2 inbound Pith citation observations for arXiv:2506.22023.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.22023 v1

Coverage vector

measured 50 of 50 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T22:19:03.875806Z

measured 52 of 52 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-27T01:05:20.432379Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T20:58:57.478873Z

Reference resolution

50 of 50 outbound references displayed

  • verified exact2
  • verified fuzzy18
  • unresolved30
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation a5e68496-bb31-450b-84ca-0952e784f0c2 · outbound

This paper cites A Survey of Large Language Models.

Robust and Efficient Autoregressive Speech Synthesis with Dynamic Chunk-wise Prediction Policy A Survey of Large Language Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T22:18:58.451990Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:18:58.451990Z digest=sha256:a243f50293fe9d5928aaf649b80d539c5ad1df44065aea8ed4678e336d274206

Observation cb1ccaf6-217f-4166-bf0c-6b91003ee642 · outbound

This paper cites Codec-SUPERB: An In-Depth Analysis of Sound Codec Models.

Robust and Efficient Autoregressive Speech Synthesis with Dynamic Chunk-wise Prediction Policy Codec-SUPERB: An In-Depth Analysis of Sound Codec Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T22:18:58.537491Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:18:58.537491Z digest=sha256:cdb864da7c5389dce94237b2b4eb150d210eeb39ba5ff2c348399a4f21fac9ab

Observation bd8c8cb0-03ac-476c-8e5b-b72ad47d999e · outbound

This paper cites Recent advances in discrete speech tokens: A review,.

Robust and Efficient Autoregressive Speech Synthesis with Dynamic Chunk-wise Prediction Policy Recent advances in discrete speech tokens: A review,

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T22:18:58.693492Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:18:58.693492Z digest=sha256:01dece907b3bb1334ca7159f7812920604bee66942020fb6bce08f0676da685b

Observation 8ef43034-e09c-47b0-b03f-f0e7494a5942 · outbound

This paper cites On The Landscape of Spoken Language Models: A Comprehensive Survey.

Robust and Efficient Autoregressive Speech Synthesis with Dynamic Chunk-wise Prediction Policy On The Landscape of Spoken Language Models: A Comprehensive Survey

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T22:18:58.895945Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:18:58.895945Z digest=sha256:2825e5dc1f632958f4393c6b2e799a3951d75de7426d8c7b2f6237cf17c8905b

Observation 170c9bb8-d83e-4618-bde3-c75ad123e85e · outbound

This paper cites Neural Discrete Representation Learning,.

Robust and Efficient Autoregressive Speech Synthesis with Dynamic Chunk-wise Prediction Policy Neural Discrete Representation Learning,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:19:09.288020Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:18:59.021254Z digest=sha256:528c0e049cd99fff2b64afc0f5c2f980723cc08403e03f3c8d040d1f86429e2c

Observation 7d465b53-a010-46e1-83f7-f162dcbfc1dd · outbound

This paper cites Fast and high-quality auto-regressive speech synthesis via speculative decoding,.

Robust and Efficient Autoregressive Speech Synthesis with Dynamic Chunk-wise Prediction Policy Fast and high-quality auto-regressive speech synthesis via speculative decoding,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:19:09.091961Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:18:59.166926Z digest=sha256:e2cf43eaf789c2d66bab4d9340480c70180d8c56a9c6cd39c0e6bd1a3205c9d1

Observation e7cfcf5d-e68c-4595-8ce2-625a9860e24b · outbound

This paper cites Accelerating Codec-based Speech Synthesis with Multi-Token Prediction and Speculative Decoding.

Robust and Efficient Autoregressive Speech Synthesis with Dynamic Chunk-wise Prediction Policy Accelerating Codec-based Speech Synthesis with Multi-Token Prediction and Speculative Decoding

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-08-06T22:19:04.696229Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:18:59.328921Z digest=sha256:41a12799cd44a62ad4a25448f768c76aa2f9b243fbb9f4891327e257996176e0

Observation 7f734fbc-4e28-473c-b2b5-189bfff8215b · outbound

This paper cites VocalNet: Speech LLM with Multi-Token Prediction for Faster and High-Quality Generation.

Robust and Efficient Autoregressive Speech Synthesis with Dynamic Chunk-wise Prediction Policy VocalNet: Speech LLM with Multi-Token Prediction for Faster and High-Quality Generation

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T22:18:59.507675Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:18:59.507675Z digest=sha256:6ee0d143a275f8942f18fb17db0cdf6f8f61308a3fdb56d78121204eef8f1920

Observation 4b95c2a2-7701-426e-9b56-1e2a926f9b7d · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Robust and Efficient Autoregressive Speech Synthesis with Dynamic Chunk-wise Prediction Policy DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T22:18:59.608837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:18:59.608837Z digest=sha256:fba3944fec98e5d5548b7a1c4456fc70154675607d1b071dd88cd7490cc825d3

Observation ac48b1a8-3cfd-458a-8666-e40d88a486a5 · outbound

This paper cites SFT Memorizes, RL Generalizes: A Comparative Study of Foundation Model Post-training.

Robust and Efficient Autoregressive Speech Synthesis with Dynamic Chunk-wise Prediction Policy SFT Memorizes, RL Generalizes: A Comparative Study of Foundation Model Post-training

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T22:18:59.735703Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:18:59.735703Z digest=sha256:e6c35ce88edc9f3ec844c46a2ce107a067fe43d04afa3735d45d468cd326df7c

Observation 0eaed673-25f5-44ed-9d4a-9084025b75a2 · outbound

This paper cites Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?.

Robust and Efficient Autoregressive Speech Synthesis with Dynamic Chunk-wise Prediction Policy Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T22:18:59.877417Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:18:59.877417Z digest=sha256:67e53210cd2b1c17c411cdc1cf746456bd807cee105f3b60fa7c9c6a3e280adc

Observation 10172a57-2cf3-4b8f-a1e9-81406f67e2dd · outbound

This paper cites LibriTTS: A Corpus Derived from LibriSpeech for Text-to- Speech,.

Robust and Efficient Autoregressive Speech Synthesis with Dynamic Chunk-wise Prediction Policy LibriTTS: A Corpus Derived from LibriSpeech for Text-to- Speech,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:19:08.859990Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:19:00.015587Z digest=sha256:5f86bbad8df48b3f886d0ea561d135b5f46dc8c35e156cd0046c4446a2b40b6f

Observation b04355ef-2ce0-4264-8d3c-142da83fbb8a · outbound

This paper cites HuBERT: Self-Supervised Speech Representation Learning by Masked Prediction of Hidden Units,.

Robust and Efficient Autoregressive Speech Synthesis with Dynamic Chunk-wise Prediction Policy HuBERT: Self-Supervised Speech Representation Learning by Masked Prediction of Hidden Units,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:19:08.546115Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:19:00.124151Z digest=sha256:856d2197f01a8039d5b1064eea39de80ec5d12750049d4ffc2b336bea6f816a7

Observation f64dcb36-58b7-422b-83b9-42bbada1cba2 · outbound

This paper cites CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens.

Robust and Efficient Autoregressive Speech Synthesis with Dynamic Chunk-wise Prediction Policy CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T22:19:00.278755Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:19:00.278755Z digest=sha256:a52d0a9a07dc32fb28e482150aa9be35fbf9b88e1cd8ec18b88d83830a225aec

Observation cca8f652-0bf6-4cff-9abb-116298515508 · outbound

This paper cites Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers,.

Robust and Efficient Autoregressive Speech Synthesis with Dynamic Chunk-wise Prediction Policy Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:19:08.260615Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:19:00.426911Z digest=sha256:93549c65f05a78dbe4514edc3927de75cd784c446ecb9c82467f9e1c16924465

Observation cbe6412f-31f3-4fdc-b33f-731c6b255e13 · outbound

This paper cites V oiceCraft: Zero-Shot Speech Editing and Text-to-Speech in the Wild,.

Robust and Efficient Autoregressive Speech Synthesis with Dynamic Chunk-wise Prediction Policy V oiceCraft: Zero-Shot Speech Editing and Text-to-Speech in the Wild,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:19:07.967672Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:19:00.547714Z digest=sha256:d78fe3427b0945852f15f51aabecee434acb9154e153bdf4e866509d96ab3448

Observation 3fff0af9-2b2b-4938-aa93-41cff71acb69 · outbound

This paper cites Seed-TTS: A Family of High-Quality Versatile Speech Generation Models.

Robust and Efficient Autoregressive Speech Synthesis with Dynamic Chunk-wise Prediction Policy Seed-TTS: A Family of High-Quality Versatile Speech Generation Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T22:19:00.703156Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:19:00.703156Z digest=sha256:b090d2a988b8cca15817f3cc0460b1a404a27f680607a61965ae8cff5925bb94

Observation 45eb768c-4d27-4724-8018-1a3d215099e3 · outbound

This paper cites BASE TTS: Lessons from building a billion-parameter Text-to-Speech model on 100K hours of data.

Robust and Efficient Autoregressive Speech Synthesis with Dynamic Chunk-wise Prediction Policy BASE TTS: Lessons from building a billion-parameter Text-to-Speech model on 100K hours of data

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T22:19:00.848288Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:19:00.848288Z digest=sha256:08a863807fd922d38ed9eb2aac3e15ddfb04421f573f8aef08d9fbf6654a5adc

Observation 0f544de7-5d73-42a3-aa1b-fde17836f6c3 · outbound

This paper cites ELLA-V: Stable Neural Codec Language Modelling with Alignment-Guided Sequence Reordering,.

Robust and Efficient Autoregressive Speech Synthesis with Dynamic Chunk-wise Prediction Policy ELLA-V: Stable Neural Codec Language Modelling with Alignment-Guided Sequence Reordering,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:19:07.735392Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:19:00.908080Z digest=sha256:96679f45eb8b4ea87d2f7b0d2f1327a78e2659959c20099036740ca9337162b6

Observation da940102-74c5-4465-8666-46c1934a7c50 · outbound

This paper cites VALL-E R: Robust and Efficient Zero-Shot Text-to-Speech Synthesis via Monotonic Alignment.

Robust and Efficient Autoregressive Speech Synthesis with Dynamic Chunk-wise Prediction Policy VALL-E R: Robust and Efficient Zero-Shot Text-to-Speech Synthesis via Monotonic Alignment

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T22:19:01.038507Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:19:01.038507Z digest=sha256:653ab295c9bf99b001aa173387e62b62ecbef976841da577f46548f95a74951a

Observation 8b915efd-50a3-4683-9d04-669e86c5673f · outbound

This paper cites V ALL-T: Decoder-only generative transducer for robust and decoding-controllable text-to-speech,.

Robust and Efficient Autoregressive Speech Synthesis with Dynamic Chunk-wise Prediction Policy V ALL-T: Decoder-only generative transducer for robust and decoding-controllable text-to-speech,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:19:07.508375Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:19:01.123238Z digest=sha256:c5175bbb1874fb0bb99fa5fcafa885aab08680b079ad7020a6e388333f7ce005

Observation 0b967284-e98e-45fc-b7af-9a6177e2998a · outbound

This paper cites RALL-E: Robust Codec Language Modeling with Chain-of-Thought Prompting for Text-to-Speech Synthesis.

Robust and Efficient Autoregressive Speech Synthesis with Dynamic Chunk-wise Prediction Policy RALL-E: Robust Codec Language Modeling with Chain-of-Thought Prompting for Text-to-Speech Synthesis

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T22:19:01.200105Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:19:01.200105Z digest=sha256:b7c528debf41ea70e3c9f11ea247a5c2b6ec3743ae2dcf03777de85027d2ac7b

Observation 2c9ca762-3128-47f8-95e8-b181b5323ef8 · outbound

This paper cites SNAC: Multi-Scale Neural Audio Codec.

Robust and Efficient Autoregressive Speech Synthesis with Dynamic Chunk-wise Prediction Policy SNAC: Multi-Scale Neural Audio Codec

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T22:19:01.314755Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:19:01.314755Z digest=sha256:246372da102d0c3255c128d1776aaf94deac13661aed0ea836a11e62c6177ea4

Observation 9ff7cc4e-5338-474b-a2b5-c427656fce26 · outbound

This paper cites Speaking from Coarse to Fine: Improving Neural Codec Language Model via Multi-Scale Speech Coding and Generation.

Robust and Efficient Autoregressive Speech Synthesis with Dynamic Chunk-wise Prediction Policy Speaking from Coarse to Fine: Improving Neural Codec Language Model via Multi-Scale Speech Coding and Generation

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-08-06T22:19:04.337592Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:19:01.383749Z digest=sha256:b1a7962f5a884d471f95de2c3162104f0dd639b62bd05931b8aaadfe5c7ff8f4

Observation 453300fe-b9e7-4766-bb5e-e29390984124 · outbound

This paper cites UniAudio 1.5: Large Language Model-Driven Audio Codec is A Few-Shot Audio Task Learner,.

Robust and Efficient Autoregressive Speech Synthesis with Dynamic Chunk-wise Prediction Policy UniAudio 1.5: Large Language Model-Driven Audio Codec is A Few-Shot Audio Task Learner,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:19:07.260505Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:19:01.499041Z digest=sha256:3446dac835c97ddb7a02b20a10ed2b4a5425baabf812f5c377a2f1c5d61635a3

Observation 623308b0-5410-4c93-921d-76604959d8cc · outbound

This paper cites SpeechTokenizer: Unified Speech Tokenizer for Speech Language Models,.

Robust and Efficient Autoregressive Speech Synthesis with Dynamic Chunk-wise Prediction Policy SpeechTokenizer: Unified Speech Tokenizer for Speech Language Models,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:19:07.030523Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:19:01.620066Z digest=sha256:918dc484e3a41365a4921654aee0aa8f16e6e7e50e6cdb81faf8c27c7d46deca

Observation 541ac633-3efc-459c-82c6-140d69c2f425 · outbound

This paper cites Moshi: a speech-text foundation model for real-time dialogue.

Robust and Efficient Autoregressive Speech Synthesis with Dynamic Chunk-wise Prediction Policy Moshi: a speech-text foundation model for real-time dialogue

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T22:19:01.760430Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:19:01.760430Z digest=sha256:3a1c4b76d08016b170aead1939b7f5ce46174e18368d485d6f73b659da15d7cc

Observation f31aca70-c936-420a-b30f-4a1e6c6ec694 · outbound

This paper cites AudioLM: A Language Modeling Approach to Audio Generation,.

Robust and Efficient Autoregressive Speech Synthesis with Dynamic Chunk-wise Prediction Policy AudioLM: A Language Modeling Approach to Audio Generation,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:19:06.814846Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:19:01.801928Z digest=sha256:f00751f2f4f3b8ea7d6f230124a988583fb092b1fdbb45977a7f66b78de227c5

Observation 87d7d26a-7d45-4b9e-a157-954651f6f045 · outbound

This paper cites Speak, Read and Prompt: High-Fidelity Text-to- Speech with Minimal Supervision,.

Robust and Efficient Autoregressive Speech Synthesis with Dynamic Chunk-wise Prediction Policy Speak, Read and Prompt: High-Fidelity Text-to- Speech with Minimal Supervision,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:19:06.546180Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:19:01.863234Z digest=sha256:b4e3379a533fa3f6f83c0956938ae4e9f692e1669f4aa6e791b7b7f1b31d228d

Observation 212b35dc-657a-48f6-ae02-a3db385f2a66 · outbound

This paper cites MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer,.

Robust and Efficient Autoregressive Speech Synthesis with Dynamic Chunk-wise Prediction Policy MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:19:06.308846Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:19:01.965874Z digest=sha256:7b165799b1a62fd1e982261cad360f6c00e799742fe2cbb5c85e55db3ab3519b

Observation 021b628d-e59d-41e5-9d24-abc2e64e2af3 · outbound

This paper cites Medusa: Simple LLM Inference Acceleration Framework with Multiple Decoding Heads.

Robust and Efficient Autoregressive Speech Synthesis with Dynamic Chunk-wise Prediction Policy Medusa: Simple LLM Inference Acceleration Framework with Multiple Decoding Heads

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T22:19:02.032915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:19:02.032915Z digest=sha256:066ab13335580fd6c93c4118ad6ca4a02dabfb0c8ef63326bd35468618dcf6a8

Observation 36d5de7f-a8ec-4f5a-bd80-308858dd3cb8 · outbound

This paper cites Better & Faster Large Language Models via Multi-token Prediction.

Robust and Efficient Autoregressive Speech Synthesis with Dynamic Chunk-wise Prediction Policy Better & Faster Large Language Models via Multi-token Prediction

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T22:19:02.090093Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:19:02.090093Z digest=sha256:fc3f0b7366550aca8c4044a85f972f721d552d2428a445e8a93d72997243abab

Observation 01db0498-5f06-4f96-ac2c-35d8cbbe18d4 · outbound

This paper cites DeepSeek-V3 Technical Report.

Robust and Efficient Autoregressive Speech Synthesis with Dynamic Chunk-wise Prediction Policy DeepSeek-V3 Technical Report

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T22:19:02.210858Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:19:02.210858Z digest=sha256:ddbfd1d3b677a2652f41c495706ba36e43e27c614ca81696754cb9af8f387435

Observation 034f4c16-c96a-4c01-8782-c94c93b10491 · outbound

This paper cites Training language models to follow instructions with human feedback,.

Robust and Efficient Autoregressive Speech Synthesis with Dynamic Chunk-wise Prediction Policy Training language models to follow instructions with human feedback,

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T22:19:02.317947Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:19:02.317947Z digest=sha256:91da4fad9b738d55fc16a78eac4d6d68484278934e6720c16cdd2bd6cde12d6d

Observation b1f83a33-32f4-4d46-ad4d-407412ceb971 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Robust and Efficient Autoregressive Speech Synthesis with Dynamic Chunk-wise Prediction Policy Proximal Policy Optimization Algorithms

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T22:19:02.460098Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:19:02.460098Z digest=sha256:862094ccb76ab4ccd5190b7478848dac567b59dae3c3ac8170ea4c331a3a6faa

Observation 8b2ec461-051b-48dd-a584-65794bf75d01 · outbound

This paper cites Direct preference optimization: Your language model is secretly a reward model,.

Robust and Efficient Autoregressive Speech Synthesis with Dynamic Chunk-wise Prediction Policy Direct preference optimization: Your language model is secretly a reward model,

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T22:19:02.578605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:19:02.578605Z digest=sha256:aeb222dfe65e157ac50bddea3d9a45643708613e055abcb957e93e9a8e9efbfd

Observation ed5e40cf-f30d-403d-b151-12f36ac6cea7 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Robust and Efficient Autoregressive Speech Synthesis with Dynamic Chunk-wise Prediction Policy DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T22:19:02.652495Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:19:02.652495Z digest=sha256:fb8a2432b534f311ddbb384f734b443bdd8f4ad30c0e03ebfcd92aafc6b79571

Observation b38bdcb0-3913-450b-8c11-8b7e27459dd7 · outbound

This paper cites Align-SLM: Textless Spoken Language Models with Reinforcement Learning from AI Feedback.

Robust and Efficient Autoregressive Speech Synthesis with Dynamic Chunk-wise Prediction Policy Align-SLM: Textless Spoken Language Models with Reinforcement Learning from AI Feedback

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T22:19:02.738861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:19:02.738861Z digest=sha256:78b69fbd5d7b9f806dea666539ddca2045fddb0eeba3308b82853fd045c46ddd

Observation 76cb5b0b-a16d-4a52-87f3-b884024da7cf · outbound

This paper cites Robust Zero-Shot Text-to-Speech Synthesis with Reverse Inference Optimization.

Robust and Efficient Autoregressive Speech Synthesis with Dynamic Chunk-wise Prediction Policy Robust Zero-Shot Text-to-Speech Synthesis with Reverse Inference Optimization

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T22:19:02.830612Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:19:02.830612Z digest=sha256:98336c11f57666d6aad97ce865946accd49e72083faee78be7a03e96b0afe57c

Observation aff5959b-423b-4aed-a646-6c0e663d169a · outbound

This paper cites Enhancing Zero-shot Text-to-Speech Synthesis with Human Feedback.

Robust and Efficient Autoregressive Speech Synthesis with Dynamic Chunk-wise Prediction Policy Enhancing Zero-shot Text-to-Speech Synthesis with Human Feedback

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T22:19:02.888700Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:19:02.888700Z digest=sha256:b1fccb4a7d234bc4f39873eb357390dbd6a23c4fc494b59bb67ca27a98e9bc06

Observation 576192fd-545d-423f-a2fd-742d69e6552d · outbound

This paper cites SpeechAlign: Aligning speech generation to human preferences,.

Robust and Efficient Autoregressive Speech Synthesis with Dynamic Chunk-wise Prediction Policy SpeechAlign: Aligning speech generation to human preferences,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:19:06.010259Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:19:02.963953Z digest=sha256:0f878ff54675e8758da2e01d7af22de29c3b5eae2561d8a271780861f48c5297

Observation 71a4a6ff-5a9b-4326-9b1f-18cdfee7b83c · outbound

This paper cites Fine-grained preference optimization improves zero-shot text-to-speech,.

Robust and Efficient Autoregressive Speech Synthesis with Dynamic Chunk-wise Prediction Policy Fine-grained preference optimization improves zero-shot text-to-speech,

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T22:19:03.066762Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:19:03.066762Z digest=sha256:124626773473e57dfb03f8d848a2d2f24a0671b0179ce3edda3102c23c29dbb4

Observation 9a995078-6c79-4198-9ada-73e1eedc2b32 · outbound

This paper cites VALL-E 2: Neural Codec Language Models are Human Parity Zero-Shot Text to Speech Synthesizers.

Robust and Efficient Autoregressive Speech Synthesis with Dynamic Chunk-wise Prediction Policy VALL-E 2: Neural Codec Language Models are Human Parity Zero-Shot Text to Speech Synthesizers

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T22:19:03.126226Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:19:03.126226Z digest=sha256:dae9a79154879134bf7910aa6c2d2a4ad4e6f4edae091a19a802c76fc3eaee50

Observation f9935a8d-bf66-4312-a1fc-ca1e4ab6f111 · outbound

This paper cites UniCATS: A Unified Context-Aware Text-to-Speech Framework with Contextual VQ-Diffusion and V ocoding,.

Robust and Efficient Autoregressive Speech Synthesis with Dynamic Chunk-wise Prediction Policy UniCATS: A Unified Context-Aware Text-to-Speech Framework with Contextual VQ-Diffusion and V ocoding,

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:19:05.793544Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:19:03.250253Z digest=sha256:3ece8cc2ca28751fce3a71fe4b0404e0ced8fe0fb3becad1c9242693efd28086

Observation e51de082-c27f-4ccf-9cef-17c1e6e69518 · outbound

This paper cites LSCodec: Low-Bitrate and Speaker-Decoupled Discrete Speech Codec.

Robust and Efficient Autoregressive Speech Synthesis with Dynamic Chunk-wise Prediction Policy LSCodec: Low-Bitrate and Speaker-Decoupled Discrete Speech Codec

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T22:19:03.356798Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:19:03.356798Z digest=sha256:5f81b13d85ab25ca5331fcc0c0bd66022c1c70e6d41037f6faeda6549acff647

Observation 65daf919-b05e-4438-8033-4f3a98cb4407 · outbound

This paper cites WavLM: Large-Scale Self-Supervised Pre-Training for Full Stack Speech Processing,.

Robust and Efficient Autoregressive Speech Synthesis with Dynamic Chunk-wise Prediction Policy WavLM: Large-Scale Self-Supervised Pre-Training for Full Stack Speech Processing,

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:19:05.611424Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:19:03.471288Z digest=sha256:6fc18e0b6726d0b5f3595213c73809aa40785745b888cd5bf80c099627cd7c3a

Observation d8c83b8d-7981-4eae-a62d-7fefc3f3a776 · outbound

This paper cites Sigmoid-weighted linear units for neural network function approximation in reinforcement learning,.

Robust and Efficient Autoregressive Speech Synthesis with Dynamic Chunk-wise Prediction Policy Sigmoid-weighted linear units for neural network function approximation in reinforcement learning,

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-06T22:19:03.560664Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:19:03.560664Z digest=sha256:8b7c53ece27da8b5e72c3317dbab62a60302aace8a601818d7bdf5467843d442

Observation d2d0d98b-d896-40cd-99c4-b34e37f99f92 · outbound

This paper cites UTMOS: UTokyo- SaruLab System for V oiceMOS Challenge 2022,.

Robust and Efficient Autoregressive Speech Synthesis with Dynamic Chunk-wise Prediction Policy UTMOS: UTokyo- SaruLab System for V oiceMOS Challenge 2022,

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:19:05.438167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:19:03.693568Z digest=sha256:a7f7af1921aec0fdef5d19c90199b493ce18ff1c904fb810eeb0e892d24a9615

Observation 75ad788d-482b-47b3-9388-04dc6c7cf032 · outbound

This paper cites Robust Speech Recognition via Large-Scale Weak Supervision,.

Robust and Efficient Autoregressive Speech Synthesis with Dynamic Chunk-wise Prediction Policy Robust Speech Recognition via Large-Scale Weak Supervision,

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:19:05.082403Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T22:19:03.772752Z digest=sha256:c04a3c1742f4177c5ff8a68fb0ea2c7494958ab8ba6d5cde026e8ebcca4255e4

Observation 5835c07c-950d-4118-957f-594a85fa128b · outbound

This paper cites Zipformer: A faster and better encoder for automatic speech recognition.

Robust and Efficient Autoregressive Speech Synthesis with Dynamic Chunk-wise Prediction Policy Zipformer: A faster and better encoder for automatic speech recognition

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T22:19:03.875806Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:19:03.875806Z digest=sha256:372c436920a2a1c087af4e0f49e20e843b0d274dd1c8087aa497aa16a689a142

Pith citing papers

Observation 03bae24a-489b-4704-8727-8f4f7b6df872 · inbound

From Static Inference to Dynamic Interaction: A Survey of Streaming Large Language Models cites this paper.

From Static Inference to Dynamic Interaction: A Survey of Streaming Large Language Models Robust and Efficient Autoregressive Speech Synthesis with Dynamic Chunk-wise Prediction Policy

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-15T16:20:09.991805Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T16:16:15.819622Z digest=sha256:f6834bff937dd3a835a172869a4c5253be283a6302a83d92245df104faf70e4a

Observation e5d064c5-efdc-4ec2-aab7-be973da6cf34 · inbound

ASTRA: A Scalable Next-Generation ATCO Training Simulator with Autonomous Simpilots cites this paper.

ASTRA: A Scalable Next-Generation ATCO Training Simulator with Autonomous Simpilots Robust and Efficient Autoregressive Speech Synthesis with Dynamic Chunk-wise Prediction Policy

Reference 55

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T20:58:57.480674Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-27T01:05:20.432379Z digest=sha256:ef6cdcd4e78ac07978d4c10ba4c7d00fdc766698fc89b20844c459d2eb5b6399