Pith. sign in

Paper Citation Record · LEDGER

ReLMCodec: Designing Predictable Speech Tokens from Pre-Quantization Phoneme Structure

As of 16 August 2026, this Paper Citation Record lists 26 of 26 outbound references and 0 inbound Pith citation observations for arXiv:2608.08286.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.08286 v1

Coverage vector

measured 26 of 26 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T00:15:23.055549Z

measured 26 of 26 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

26 of 26 outbound references displayed

  • verified exact6
  • verified fuzzy1
  • unresolved15
  • parse uncertain0
  • malformed identifier3
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation b750918b-8b95-4e55-9042-cf6776c0b7fb · outbound

This paper cites Moshi: a speech-text foundation model for real-time dialogue.

ReLMCodec: Designing Predictable Speech Tokens from Pre-Quantization Phoneme Structure Moshi: a speech-text foundation model for real-time dialogue

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-12T00:15:22.966458Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:15:22.966458Z digest=sha256:f9307cc382234305dd79f62b13a8746895c90a989ba281ae7138aeb2b7b9ad49

Observation 2cdbb5ac-acb0-4bf6-b5e8-1379f62bb23d · outbound

This paper cites an unresolved cited work.

ReLMCodec: Designing Predictable Speech Tokens from Pre-Quantization Phoneme Structure Unresolved cited work

Reference 6

Resolution
verified exact
raw_fallback, observed 2026-08-12T00:15:23.716930Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T00:15:22.969847Z digest=sha256:5698459a7fb0bf2d08ffaabf9752c29d30178d557628ad97e600ec7bf884df0b

Observation 4cafe02c-b60c-4800-9652-81b9419f94f3 · outbound

This paper cites CosyVoice 3: Towards In-the-wild Speech Generation via Scaling-up and Post-training.

ReLMCodec: Designing Predictable Speech Tokens from Pre-Quantization Phoneme Structure CosyVoice 3: Towards In-the-wild Speech Generation via Scaling-up and Post-training

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-12T00:15:22.972684Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:15:22.972684Z digest=sha256:4cea067be448124e1aa8a68fe883d944f1601be901a961f1759cfc854fde0377

Observation 85648047-3cb3-4e35-b1c2-735661c22afe · outbound

This paper cites 16 YiweiGuo,ZhihanLi,ChenpengDu,HankunWang,XieChen, and Kai Yu.

ReLMCodec: Designing Predictable Speech Tokens from Pre-Quantization Phoneme Structure 16 YiweiGuo,ZhihanLi,ChenpengDu,HankunWang,XieChen, and Kai Yu

Reference 8

Resolution
verified exact
doi, observed 2026-08-12T00:15:23.116937Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T00:15:22.976967Z digest=sha256:7880b660d07b48cf8d18c2f0b7d94ac9ffaa95ae65463de622df3ee3b810524c

Observation c7876959-d0f1-4515-8286-e7a222a7e710 · outbound

This paper cites NadavHar-Tuv,OrTal,andYossiAdi.

ReLMCodec: Designing Predictable Speech Tokens from Pre-Quantization Phoneme Structure NadavHar-Tuv,OrTal,andYossiAdi

Reference 9

Resolution
verified exact
doi, observed 2026-08-12T00:15:23.106902Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T00:15:22.980061Z digest=sha256:eddbced0a05eb2b35428c57d5b374af94660ff6f5d58a015ee2ff97cbad51df2

Observation 78c4e1da-8186-4199-81db-0b9f26c668a6 · outbound

This paper cites Harry Julian, Rachel Beeson, Lohith Konathala, Johanna Ulin, and Jiameng Gao.

ReLMCodec: Designing Predictable Speech Tokens from Pre-Quantization Phoneme Structure Harry Julian, Rachel Beeson, Lohith Konathala, Johanna Ulin, and Jiameng Gao

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-12T00:15:22.986532Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:15:22.986532Z digest=sha256:80fe49faf24cc89ebfd8739726fc43b97c2df6354a2e79057f63b69120e1ac11

Observation e7566782-808e-4cb6-ab93-18919a4410d6 · outbound

This paper cites Finite Scalar Quantization Enables Redundant and Transmission-Robust Neural Audio Compression at Low Bit-rates.

ReLMCodec: Designing Predictable Speech Tokens from Pre-Quantization Phoneme Structure Finite Scalar Quantization Enables Redundant and Transmission-Robust Neural Audio Compression at Low Bit-rates

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-12T00:15:22.989838Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:15:22.989838Z digest=sha256:177dd656000b1106dfca3ccf607dda053742ac6ff2ebff462f7068e6ae5cf45b

Observation 3791d20b-f13b-4c85-addc-9828bfc49cd0 · outbound

This paper cites Montreal forced aligner: Trainable text-speech alignment using Kaldi.

ReLMCodec: Designing Predictable Speech Tokens from Pre-Quantization Phoneme Structure Montreal forced aligner: Trainable text-speech alignment using Kaldi

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T00:15:23.772910Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T00:15:22.994783Z digest=sha256:2417e0c4c8854980a98bdb623b560d26189dd55a7df944e406d87341a32b94fa

Observation e61731eb-39d7-4959-a2b6-697f3f267a64 · outbound

This paper cites Vassil Panayotov, Guoguo Chen, Daniel Povey, and Sanjeev Khudanpur.

ReLMCodec: Designing Predictable Speech Tokens from Pre-Quantization Phoneme Structure Vassil Panayotov, Guoguo Chen, Daniel Povey, and Sanjeev Khudanpur

Reference 15

Resolution
verified exact
doi, observed 2026-08-12T00:15:23.082455Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T00:15:23.013594Z digest=sha256:613e31ee5a7be550d84c36a3c15b3d64d9b8523bf795c1910844a4f493a7804c

Observation 78820abc-3b26-4e8f-b799-c9ffb95c07a7 · outbound

This paper cites Robust Speech Recognition via Large-Scale Weak Supervision.

ReLMCodec: Designing Predictable Speech Tokens from Pre-Quantization Phoneme Structure Robust Speech Recognition via Large-Scale Weak Supervision

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-12T00:15:23.020983Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:15:23.020983Z digest=sha256:a27eeaf98510d2496a32d5c3a89277d9d903e370acf6e2da3b1811fee2ff340e

Observation 1e4b5ed5-0c19-49f4-8d7f-5b9810faf0ec · outbound

This paper cites UTMOS: UTokyo-SaruLab System for VoiceMOS Challenge 2022.

ReLMCodec: Designing Predictable Speech Tokens from Pre-Quantization Phoneme Structure UTMOS: UTokyo-SaruLab System for VoiceMOS Challenge 2022

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-12T00:15:23.024338Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:15:23.024338Z digest=sha256:180786c8b002707ccc3196e0f436b83d8c2aeea2d6551a4ee394df6d5b16ba74

Observation 0e13cfd5-72e7-48c3-8426-d5203d46a295 · outbound

This paper cites MagiCodec: Simple Masked Gaussian-Injected Codec for High-Fidelity Reconstruction and Generation.

ReLMCodec: Designing Predictable Speech Tokens from Pre-Quantization Phoneme Structure MagiCodec: Simple Masked Gaussian-Injected Codec for High-Fidelity Reconstruction and Generation

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-12T00:15:23.032346Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:15:23.032346Z digest=sha256:15a3b0ac23139a97142bb0a99c60d244fccb354157837a4cada9aa66ab5c80c1

Observation 8f600229-0b91-43ba-9765-7dbfcefa954f · outbound

This paper cites Content is What Remains: Invariant Speech Tokenization from Parallel Utterances.

ReLMCodec: Designing Predictable Speech Tokens from Pre-Quantization Phoneme Structure Content is What Remains: Invariant Speech Tokenization from Parallel Utterances

Reference 21

Resolution
metadata mismatch
local_arxiv, observed 2026-08-12T00:15:23.478725Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T00:15:23.039683Z digest=sha256:3ec8e4e8cdce435fef2ba58878beb6c15dd0e12e651fbb3e50a70d43199b0dad

Observation 19fe2a0b-5d0b-430a-9885-8f3e678956f9 · outbound

This paper cites Yuancheng Wang, Haoyue Zhan, Liwei Liu, Ruihong Zeng, Haotian Guo, Jiachen Zheng, Qiang Zhang, Xueyao Zhang, Shunsi Zhang, and Zhizheng Wu.

ReLMCodec: Designing Predictable Speech Tokens from Pre-Quantization Phoneme Structure Yuancheng Wang, Haoyue Zhan, Liwei Liu, Ruihong Zeng, Haotian Guo, Jiachen Zheng, Qiang Zhang, Xueyao Zhang, Shunsi Zhang, and Zhizheng Wu

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-12T00:15:23.042977Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:15:23.042977Z digest=sha256:443477b9680fb697d706ea2939968caaa960ee9c5b4f7dd10e66177f1137ab42

Observation d910d2a4-a97d-46fd-b6af-db581bfa13d3 · outbound

This paper cites Detai Xin, Xu Tan, Shinnosuke Takamichi, and Hiroshi Saruwatari.

ReLMCodec: Designing Predictable Speech Tokens from Pre-Quantization Phoneme Structure Detai Xin, Xu Tan, Shinnosuke Takamichi, and Hiroshi Saruwatari

Reference 23

Resolution
malformed identifier
no resolver link, observed 2026-08-12T00:15:23.046190Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:15:23.046190Z digest=sha256:adca892b0baafdda96f12119c254826e908a895b46e7d37cb2420a1d15247f71

Observation de98056c-99a6-4a4a-ae4c-b4ef6728cb0d · outbound

This paper cites An Yang et al.

ReLMCodec: Designing Predictable Speech Tokens from Pre-Quantization Phoneme Structure An Yang et al

Reference 24

Resolution
verified exact
doi, observed 2026-08-12T00:15:23.320953Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T00:15:23.049313Z digest=sha256:d5dd01356b4ff71d845db96aec33ac9cc50628035323dad78bf584b1e44a6c16

Observation 51606635-d0c9-4acd-8ae2-23c75d0c566c · outbound

This paper cites HiFi-Codec: Group-residual Vector quantization for High Fidelity Audio Codec.

ReLMCodec: Designing Predictable Speech Tokens from Pre-Quantization Phoneme Structure HiFi-Codec: Group-residual Vector quantization for High Fidelity Audio Codec

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-12T00:15:23.052303Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:15:23.052303Z digest=sha256:52c642b55e90218f3a144654087007b544e35bd796371e60468c35995bf3fea3

Observation 1ddde720-ecad-422b-8964-00b9c9ef1c70 · outbound

This paper cites Xin Zhang, Dong Zhang, Shimin Li, Yaqian Zhou, and Xipeng Qiu.

ReLMCodec: Designing Predictable Speech Tokens from Pre-Quantization Phoneme Structure Xin Zhang, Dong Zhang, Shimin Li, Yaqian Zhou, and Xipeng Qiu

Reference 26

Resolution
malformed identifier
no resolver link, observed 2026-08-12T00:15:23.055549Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:15:23.055549Z digest=sha256:c249653188a1142e373c3e2ca4aefddd83af779f5dec280c24b567633cafcbc7

Observation 3c3ffca4-fce9-462f-916b-a4fb3d0967ea · outbound

This paper cites an unresolved cited work.

ReLMCodec: Designing Predictable Speech Tokens from Pre-Quantization Phoneme Structure Unresolved cited work

Reference 2015

Resolution
verified exact
raw_fallback, observed 2026-08-12T00:15:23.598627Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T00:15:23.017011Z digest=sha256:6173691807b1fed598013713032f53d1cd3be98f23463732bdc491542201d7ca

Observation e3db7876-5287-4fe0-ac6d-b954b0f0a347 · outbound

This paper cites Pooneh Mousavi, Jarod Duret, Salah Zaiem, Luca Della Libera, Artem Ploujnikov, Cem Subakan, and Mirco Ravanelli.

ReLMCodec: Designing Predictable Speech Tokens from Pre-Quantization Phoneme Structure Pooneh Mousavi, Jarod Duret, Salah Zaiem, Luca Della Libera, Artem Ploujnikov, Cem Subakan, and Mirco Ravanelli

Reference 2017

Resolution
malformed identifier
no resolver link, observed 2026-08-12T00:15:22.998055Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:15:22.998055Z digest=sha256:3919ab25064b1e79df4a542968ebb854f58ce66ecaa5a80b9a872ac5000a896d

Observation f6f5a47b-d744-4265-a5a9-2e296c29b069 · outbound

This paper cites Qwen3-TTS Technical Report.

ReLMCodec: Designing Predictable Speech Tokens from Pre-Quantization Phoneme Structure Qwen3-TTS Technical Report

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-12T00:15:22.983028Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:15:22.983028Z digest=sha256:c707eca840e53f02fdbc4f1d6df0dda4e718f8f9798298b0e7922b2903753058

Observation b45f9ab1-6ed2-4ca9-bedf-ac2352ae781a · outbound

This paper cites High Fidelity Neural Audio Compression.

ReLMCodec: Designing Predictable Speech Tokens from Pre-Quantization Phoneme Structure High Fidelity Neural Audio Compression

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-12T00:15:22.963072Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:15:22.963072Z digest=sha256:f71bebe09574a28fa019927281d19c92e8f526dd9909c7a6a46e01fa9b151c51

Observation cb07ca61-e82c-45c3-a0a8-9ccbe848488a · outbound

This paper cites Seamless: Multilingual Expressive and Streaming Speech Translation.

ReLMCodec: Designing Predictable Speech Tokens from Pre-Quantization Phoneme Structure Seamless: Multilingual Expressive and Streaming Speech Translation

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-12T00:15:23.027676Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:15:23.027676Z digest=sha256:b1778b359b0850d4559c6e9cb41a1dffedd2ad48e722751b42f6e20053d7b6cb

Observation 7e2eb48c-4064-4630-a7c3-89ab44b59d4b · outbound

This paper cites VALL-E 2: Neural Codec Language Models are Human Parity Zero-Shot Text to Speech Synthesizers.

ReLMCodec: Designing Predictable Speech Tokens from Pre-Quantization Phoneme Structure VALL-E 2: Neural Codec Language Models are Human Parity Zero-Shot Text to Speech Synthesizers

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-12T00:15:22.951195Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:15:22.951195Z digest=sha256:505558cbee781ce659539f1f19d7a55eadf0f51efbd2769de6a065ee5348890c

Observation f7c4e0b9-93a1-44fb-bb2a-67184bdb22dc · outbound

This paper cites Yushen Chen, Kai Hu, Long Zhou, Shulin Feng, Xusheng Yang, Hangting Chen, and Xie Chen.

ReLMCodec: Designing Predictable Speech Tokens from Pre-Quantization Phoneme Structure Yushen Chen, Kai Hu, Long Zhou, Shulin Feng, Xusheng Yang, Hangting Chen, and Xie Chen

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-12T00:15:22.955426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:15:22.955426Z digest=sha256:a6a12aa44d1949aea4b355f28a4d5e68c5a7f197ac18087228166f6d6d83872e

Observation ff2a4f74-3c11-4560-a2aa-5a0385ce0fb6 · outbound

This paper cites On the Distillation Loss Functions of Speech VAE for Unified Reconstruction, Understanding, and Generation.

ReLMCodec: Designing Predictable Speech Tokens from Pre-Quantization Phoneme Structure On the Distillation Loss Functions of Speech VAE for Unified Reconstruction, Understanding, and Generation

Reference 2026

Resolution
unresolved
no resolver link, observed 2026-08-12T00:15:22.959556Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:15:22.959556Z digest=sha256:c66009489af9dc319bbf53f7c0cd348fc62d0a1bf60fc1bc41afe51c426fc80f

Pith citing papers

No inbound Pith citation observations are available.