Pith. sign in

Paper Citation Record · LEDGER

Quantize More, Lose Less: Autoregressive Generation from Residually Quantized Speech Representations

As of 18 August 2026, this Paper Citation Record lists 18 of 18 outbound references and 1 inbound Pith citation observation for arXiv:2507.12197.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.12197 v1

Coverage vector

measured 18 of 18 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T16:55:50.720088Z

measured 19 of 19 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T04:20:42.550446Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-08T04:20:43.349222Z

Reference resolution

18 of 18 outbound references displayed

  • verified exact0
  • verified fuzzy1
  • unresolved17
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 971e9b2c-bccb-40d3-ac1b-bd6a2a5cb146 · outbound

This paper cites Seed-TTS: A Family of High-Quality Versatile Speech Generation Models.

Quantize More, Lose Less: Autoregressive Generation from Residually Quantized Speech Representations Seed-TTS: A Family of High-Quality Versatile Speech Generation Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T16:55:49.452658Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:55:49.452658Z digest=sha256:64a953f9229f1b612fe5b68bc93949fc3170dec013f15dcefd71330d33d161bf

Observation 90cb3af2-4b64-47b7-972a-918d54202787 · outbound

This paper cites High Fidelity Neural Audio Compression.

Quantize More, Lose Less: Autoregressive Generation from Residually Quantized Speech Representations High Fidelity Neural Audio Compression

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T16:55:49.747483Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:55:49.747483Z digest=sha256:e799834ddcf426c5cb463baf15392b53e9ab52536602cd9ea5877e8d2b49698b

Observation 70fe884f-f9ee-406b-9b21-b6aaa865037e · outbound

This paper cites Moshi: a speech-text foundation model for real-time dialogue.

Quantize More, Lose Less: Autoregressive Generation from Residually Quantized Speech Representations Moshi: a speech-text foundation model for real-time dialogue

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T16:55:49.826784Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:55:49.826784Z digest=sha256:6bc82786136b04028c738a08bfd37d9ce4196a0504b27a5c2054e4ef8929e790

Observation 8ee08fab-3606-49ca-a030-696e0f56c49b · outbound

This paper cites CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens.

Quantize More, Lose Less: Autoregressive Generation from Residually Quantized Speech Representations CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T16:55:49.922280Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:55:49.922280Z digest=sha256:9ec516f3a2240e0b548a7735befa00e802e3da1b96e41e642a7ae9e32a9c4840

Observation 06c0925f-3d35-4bb7-9e45-55b0626476e6 · outbound

This paper cites Speak, Read and Prompt: High-Fidelity Text-to-Speech with Minimal Supervision.

Quantize More, Lose Less: Autoregressive Generation from Residually Quantized Speech Representations Speak, Read and Prompt: High-Fidelity Text-to-Speech with Minimal Supervision

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T16:55:49.981725Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:55:49.981725Z digest=sha256:882b5a191953e0613d5d7f1ea9b7dc5b794253853f639734ce4a4ccc1bd3d267

Observation 5cf50fe2-f983-4810-9806-b75e9749fe36 · outbound

This paper cites BASE TTS: Lessons from building a billion-parameter Text-to-Speech model on 100K hours of data.

Quantize More, Lose Less: Autoregressive Generation from Residually Quantized Speech Representations BASE TTS: Lessons from building a billion-parameter Text-to-Speech model on 100K hours of data

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T16:55:50.078904Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:55:50.078904Z digest=sha256:b8153d6fc2ed63ecac4e9c717bc2b6bb65b00a8fca0a7f8fb41b05d5131add0e

Observation 87337e4f-1806-45e6-bfb4-bede7ddf54e3 · outbound

This paper cites StyleTTS 2: Towards Human-Level Text-to-Speech through Style Diffusion and Adversarial Training with Large Speech Language Models.

Quantize More, Lose Less: Autoregressive Generation from Residually Quantized Speech Representations StyleTTS 2: Towards Human-Level Text-to-Speech through Style Diffusion and Adversarial Training with Large Speech Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T16:55:50.153909Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:55:50.153909Z digest=sha256:487eddcf4eee62265a59193ba8f0aa248b703c68518231adafcd7fd551c4d0ca

Observation 02ec6650-2b1e-4979-8b8e-003742c21c5d · outbound

This paper cites Scaling Transformers for Low-Bitrate High-Quality Speech Coding.

Quantize More, Lose Less: Autoregressive Generation from Residually Quantized Speech Representations Scaling Transformers for Low-Bitrate High-Quality Speech Coding

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T16:55:50.237226Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:55:50.237226Z digest=sha256:01f738f851d3040312969fb6ee2e41b38e18831c824cc7035c7f9d0af180bf29

Observation 628de9bd-5f27-43bd-ab8d-750b22b24240 · outbound

This paper cites FastSpeech 2: Fast and High-Quality End-to-End Text to Speech.

Quantize More, Lose Less: Autoregressive Generation from Residually Quantized Speech Representations FastSpeech 2: Fast and High-Quality End-to-End Text to Speech

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T16:55:50.287364Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:55:50.287364Z digest=sha256:c26f23a7d923ee93ca80f50cf574d775bf67024444dc361574cc5f35d771869e

Observation 506e6954-1a24-43c5-aaa2-a2c3c992b00e · outbound

This paper cites UniAudio: An Audio Foundation Model Toward Universal Audio Generation.

Quantize More, Lose Less: Autoregressive Generation from Residually Quantized Speech Representations UniAudio: An Audio Foundation Model Toward Universal Audio Generation

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T16:55:50.479314Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:55:50.479314Z digest=sha256:e8054d581269dc75ee6896742d730d85cdd28d263c25860baf1b51c0c67a98b6

Observation 8df2d612-4795-4727-945d-4dbb76992aac · outbound

This paper cites Multi-band melgan: Faster waveform generation for high-quality text-to-speech.

Quantize More, Lose Less: Autoregressive Generation from Residually Quantized Speech Representations Multi-band melgan: Faster waveform generation for high-quality text-to-speech

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:55:51.136842Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T16:55:50.545372Z digest=sha256:0ae88d027c777002e637c9de2f80c9831b1349cb68a08667bebcd07df521da3c

Observation f0330872-1376-4cc9-92cb-e1e56c0a857c · outbound

This paper cites Llasa: Scaling Train-Time and Inference-Time Compute for Llama-based Speech Synthesis.

Quantize More, Lose Less: Autoregressive Generation from Residually Quantized Speech Representations Llasa: Scaling Train-Time and Inference-Time Compute for Llama-based Speech Synthesis

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T16:55:50.665549Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:55:50.665549Z digest=sha256:04a27ed2462cf8e3e9613dc512771d1f4c2faeb4fdb15971e3aaf60b62a4a164

Observation 5a717487-dfee-48e8-8cad-7c7e4e17f5a7 · outbound

This paper cites MiniMax-Speech: Intrinsic Zero-Shot Text-to-Speech with a Learnable Speaker Encoder.

Quantize More, Lose Less: Autoregressive Generation from Residually Quantized Speech Representations MiniMax-Speech: Intrinsic Zero-Shot Text-to-Speech with a Learnable Speaker Encoder

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T16:55:50.720088Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:55:50.720088Z digest=sha256:d9ce08f048f60eeeb62a46a6d2437006de715cc988cc16e8e20e87b6886d9ec7

Observation 45171cba-3a9e-437e-8d8f-363ea2e819ae · outbound

This paper cites HiFi-GAN: Generative Adversarial Networks for Efficient and High Fidelity Speech Synthesis.

Quantize More, Lose Less: Autoregressive Generation from Residually Quantized Speech Representations HiFi-GAN: Generative Adversarial Networks for Efficient and High Fidelity Speech Synthesis

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-06T16:55:50.034113Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:55:50.034113Z digest=sha256:550af08855ba9b6b1290489409081fb7d8d60aacbf4a79fbaba0c26034671c72

Observation affebd7b-7615-4ba4-9882-6e541db2e17b · outbound

This paper cites doi: 10.1109/jstsp.2022.3188113.

Quantize More, Lose Less: Autoregressive Generation from Residually Quantized Speech Representations doi: 10.1109/jstsp.2022.3188113

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-06T16:55:49.676482Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:55:49.676482Z digest=sha256:919e3ba486a29beffbc0ee3ecd466887395c397c7989c90cc94677a24228b41a

Observation 799f044e-036b-4f20-a99e-2041da5c1500 · outbound

This paper cites XTTS: a Massively Multilingual Zero-Shot Text-to-Speech Model.

Quantize More, Lose Less: Autoregressive Generation from Residually Quantized Speech Representations XTTS: a Massively Multilingual Zero-Shot Text-to-Speech Model

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-06T16:55:49.583767Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:55:49.583767Z digest=sha256:edbd1771717dd05f959c38c9ff1ab026ba8f35dff69d9a760a0824048de2f62a

Observation edcddf70-53f9-47ca-b41a-b5fb3e02e6dd · outbound

This paper cites Better speech synthesis through scaling.

Quantize More, Lose Less: Autoregressive Generation from Residually Quantized Speech Representations Better speech synthesis through scaling

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-06T16:55:49.512702Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:55:49.512702Z digest=sha256:d79f7701d0617e08f0ed1966fcf8154c8222ae2ecae672dd6a56d8bdb3e37ade

Observation a32f064e-3d61-4988-8f59-e4bee014f2fa · outbound

This paper cites MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer.

Quantize More, Lose Less: Autoregressive Generation from Residually Quantized Speech Representations MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-06T16:55:50.368580Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:55:50.368580Z digest=sha256:e37bde89f1e39024245d947f814ba2603d46dda02a265b458f0c2dae36b75b8c

Pith citing papers

Observation a434ea2d-a5f3-4c61-b77b-df25f926ee55 · inbound

MeloCodec: Harnessing Melodic Priors for High-Fidelity Singing Voice Representation cites this paper.

MeloCodec: Harnessing Melodic Priors for High-Fidelity Singing Voice Representation Quantize More, Lose Less: Autoregressive Generation from Residually Quantized Speech Representations

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-08-08T04:20:43.354150Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-08T04:20:42.550446Z digest=sha256:131af2de5c2a368680371299f54590e80f5580f44bc13d415f21e433317b8ee1