Pith. sign in

Paper Citation Record · LEDGER

HH-Codec: High Compression High-fidelity Discrete Neural Codec for Spoken Language Modeling

As of 17 August 2026, this Paper Citation Record lists 25 of 25 outbound references and 2 inbound Pith citation observations for arXiv:2507.18897.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.18897 v1

Coverage vector

measured 25 of 25 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T18:11:55.394112Z

measured 27 of 27 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-28T05:20:49.952030Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

25 of 25 outbound references displayed

  • verified exact0
  • verified fuzzy3
  • unresolved21
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 462a63bc-6a3d-4e3e-bf9d-c9d5f4b96dc3 · outbound

This paper cites GPT-4 Technical Report.

HH-Codec: High Compression High-fidelity Discrete Neural Codec for Spoken Language Modeling GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-15T18:11:55.297544Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:11:55.297544Z digest=sha256:a8962d274108b99d7f81302bf9b4ff646f6df54fcef0356f688f6c9034260b04

Observation 23e1e775-e1b9-4f7d-a881-f28be003e22b · outbound

This paper cites Overview of the evs codec architecture.

HH-Codec: High Compression High-fidelity Discrete Neural Codec for Spoken Language Modeling Overview of the evs codec architecture

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:11:55.778299Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T18:11:55.327955Z digest=sha256:413d85c88169b7bde144d95ff9743c5f66613ddba6dec3f6e50342de8423d3ac

Observation fe1a2dd0-d4b6-4767-882d-e715fcaf1d60 · outbound

This paper cites Textless Speech Emotion Conversion using Discrete and Decomposed Representations.

HH-Codec: High Compression High-fidelity Discrete Neural Codec for Spoken Language Modeling Textless Speech Emotion Conversion using Discrete and Decomposed Representations

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T18:11:55.335336Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:11:55.335336Z digest=sha256:599a3a02ad00850414522e71dbaafffd86daff33fe21cfe5e95a375e15483aca

Observation 20bdf8b8-6bc0-486e-8ee6-fbf87c9fa810 · outbound

This paper cites Multi-scale sub- band constant-q transform discriminator for high-fidelity vocoder.

HH-Codec: High Compression High-fidelity Discrete Neural Codec for Spoken Language Modeling Multi-scale sub- band constant-q transform discriminator for high-fidelity vocoder

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:11:55.766887Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T18:11:55.343115Z digest=sha256:ea74d6de65916d33adac4aa10c79d0d292af5bb51bb30d751c2c2ea4ec9b10f5

Observation aea01c95-f0f8-475a-b9e9-fd2a39d1bcb2 · outbound

This paper cites Emilia: An extensive, multi- lingual, and diverse speech dataset for large-scale speech generation.

HH-Codec: High Compression High-fidelity Discrete Neural Codec for Spoken Language Modeling Emilia: An extensive, multi- lingual, and diverse speech dataset for large-scale speech generation

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-15T18:11:55.346729Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:11:55.346729Z digest=sha256:57212047c75896f88f33a4e029b2602e6a167c65547ac6744002c7b3f9e2bdec

Observation 69485992-1638-4842-b044-579ad86e708e · outbound

This paper cites WavTokenizer: an Efficient Acoustic Discrete Codec Tokenizer for Audio Language Modeling.

HH-Codec: High Compression High-fidelity Discrete Neural Codec for Spoken Language Modeling WavTokenizer: an Efficient Acoustic Discrete Codec Tokenizer for Audio Language Modeling

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-15T18:11:55.350325Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:11:55.350325Z digest=sha256:35bc461d48e9c11a3a2858cb83836042f765006827346271b5e3cb72f2971916

Observation bb1a8488-8f32-4b5f-9aca-b31bc52845b5 · outbound

This paper cites BigVGAN: A Universal Neural Vocoder with Large-Scale Training.

HH-Codec: High Compression High-fidelity Discrete Neural Codec for Spoken Language Modeling BigVGAN: A Universal Neural Vocoder with Large-Scale Training

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-15T18:11:55.353862Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:11:55.353862Z digest=sha256:b373715f1d93e73e65766bb74b9f866d895ec0a821b4876665dcb30b050d9cfd

Observation edbdc2fb-8313-446b-aea5-3ed0bb6e2846 · outbound

This paper cites Single-Codec: Single-Codebook Speech Codec towards High-Performance Speech Generation.

HH-Codec: High Compression High-fidelity Discrete Neural Codec for Spoken Language Modeling Single-Codec: Single-Codebook Speech Codec towards High-Performance Speech Generation

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-15T18:11:55.357367Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:11:55.357367Z digest=sha256:c172834de50d346c44e899169f2ed2478fdf5650c7e7720d0abc862768139377

Observation 48a87d0d-3442-44be-9bd9-331e8abd7710 · outbound

This paper cites Finite Scalar Quantization: VQ-VAE Made Simple.

HH-Codec: High Compression High-fidelity Discrete Neural Codec for Spoken Language Modeling Finite Scalar Quantization: VQ-VAE Made Simple

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-15T18:11:55.360754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:11:55.360754Z digest=sha256:890257ac6fd37ff0d569dda44bb098b041bfbf92ae22b70ca1b85dd499051fa7

Observation 37aa6348-e795-4132-9b4d-342faa23e802 · outbound

This paper cites Enhanced Direct Speech-to-Speech Translation Using Self-supervised Pre-training and Data Augmentation.

HH-Codec: High Compression High-fidelity Discrete Neural Codec for Spoken Language Modeling Enhanced Direct Speech-to-Speech Translation Using Self-supervised Pre-training and Data Augmentation

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-15T18:11:55.364167Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:11:55.364167Z digest=sha256:d6b9132605635b8a65afa963a39722852174249c2e92b3f246952c28fce7664b

Observation 9779287d-7d2f-4980-a85e-2d13418a2679 · outbound

This paper cites UTMOS: UTokyo-SaruLab System for VoiceMOS Challenge 2022.

HH-Codec: High Compression High-fidelity Discrete Neural Codec for Spoken Language Modeling UTMOS: UTokyo-SaruLab System for VoiceMOS Challenge 2022

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-15T18:11:55.367705Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:11:55.367705Z digest=sha256:59f3e8c316d16068e15a40b26ff9bc22595f9f462585f8e157e1368d9d8170c5

Observation cc83373a-0ebe-4efa-9419-006e75ff1fbc · outbound

This paper cites Neural Machine Translation of Rare Words with Subword Units.

HH-Codec: High Compression High-fidelity Discrete Neural Codec for Spoken Language Modeling Neural Machine Translation of Rare Words with Subword Units

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-15T18:11:55.371713Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:11:55.371713Z digest=sha256:8c604fffe667e3e683bffc23503fce66dac4ba860cc9d924f1defc02f72497b5

Observation 9eb0fb14-fa1f-4604-a4f8-8687aee6342d · outbound

This paper cites Vocos: Closing the gap between time-domain and Fourier-based neural vocoders for high-quality audio synthesis.

HH-Codec: High Compression High-fidelity Discrete Neural Codec for Spoken Language Modeling Vocos: Closing the gap between time-domain and Fourier-based neural vocoders for high-quality audio synthesis

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-15T18:11:55.375110Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:11:55.375110Z digest=sha256:a0b0377788e9fa1aa27fc288064092fab195f10367912de044011910d0cf7bd5

Observation 410dfe13-3645-4307-afc3-b8fc8a6f6117 · outbound

This paper cites H., Hendriks, R.

HH-Codec: High Compression High-fidelity Discrete Neural Codec for Spoken Language Modeling H., Hendriks, R

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T18:11:55.746789Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T18:11:55.378902Z digest=sha256:43cac400ae6a86c070a8bd5862c72597e358698aa3c1c0abf322e04f5b01aff0

Observation bfaf7900-0d04-4e9e-a88a-0bfc24356ab2 · outbound

This paper cites LibriTTS: A Corpus Derived from LibriSpeech for Text-to-Speech.

HH-Codec: High Compression High-fidelity Discrete Neural Codec for Spoken Language Modeling LibriTTS: A Corpus Derived from LibriSpeech for Text-to-Speech

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-15T18:11:55.389989Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:11:55.389989Z digest=sha256:a9eb2302a9d27707737535a4a2d7cea7615fff9f844a7595674cf7eecd45062e

Observation ecaa0243-654e-4c06-8a15-9b880b43ace4 · outbound

This paper cites Addressing representa- tion collapse in vector quantized models with one linear layer.

HH-Codec: High Compression High-fidelity Discrete Neural Codec for Spoken Language Modeling Addressing representa- tion collapse in vector quantized models with one linear layer

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-15T18:11:55.394112Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:11:55.394112Z digest=sha256:f8d2c77630fa7a61fa6be1721355272c040041daefc1af48839d7800c831a3a3

Observation 8d92e18f-e846-483e-bbbd-016068783d3e · outbound

This paper cites SEANet: A Multi-modal Speech Enhancement Network.

HH-Codec: High Compression High-fidelity Discrete Neural Codec for Spoken Language Modeling SEANet: A Multi-modal Speech Enhancement Network

Reference 2010

Resolution
unresolved
no resolver link, observed 2026-08-15T18:11:55.382457Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:11:55.382457Z digest=sha256:97979dbe7e409a19965d6443cf9c75deb2e6f54d345f71857092866b7660b7a6

Observation e3cd9dc7-a8cc-4d21-84b1-b77f000f398c · outbound

This paper cites CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models.

HH-Codec: High Compression High-fidelity Discrete Neural Codec for Spoken Language Modeling CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models

Reference 2015

Resolution
unresolved
no resolver link, observed 2026-08-15T18:11:55.331365Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:11:55.331365Z digest=sha256:0d4086a87fd9ce82cef0bbd339c36d2be8bc496d3a8b6442ad295254e7d0f42c

Observation 9f8e55ff-669e-4306-99a9-5fc2d3ffc1f0 · outbound

This paper cites Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers.

HH-Codec: High Compression High-fidelity Discrete Neural Codec for Spoken Language Modeling Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers

Reference 2016

Resolution
unresolved
no resolver link, observed 2026-08-15T18:11:55.386364Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:11:55.386364Z digest=sha256:dbf042f31310f5d8f1710bfec3dba1a5ca067c38b75ef644437bbe7c577f552e

Observation 784dd7a5-97e1-4b6f-8cf6-d14127b47a93 · outbound

This paper cites Virtual Class Enhanced Discriminative Embedding Learning.

HH-Codec: High Compression High-fidelity Discrete Neural Codec for Spoken Language Modeling Virtual Class Enhanced Discriminative Embedding Learning

Reference 2018

Resolution
metadata mismatch
local_arxiv, observed 2026-08-15T18:11:55.697983Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T18:11:55.315040Z digest=sha256:6bb31209b9b3cd661ae5cb685d4970ae827cd3c2c0db53237fc87ccd4893bcca

Observation 623e8020-74cd-4149-ac46-cf8d4d9fafba · outbound

This paper cites Estimating or Propagating Gradients Through Stochastic Neurons for Conditional Computation.

HH-Codec: High Compression High-fidelity Discrete Neural Codec for Spoken Language Modeling Estimating or Propagating Gradients Through Stochastic Neurons for Conditional Computation

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-15T18:11:55.306285Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:11:55.306285Z digest=sha256:a36593e65522b6fc3671dc63dad6d81abf1265f6ab2927b425c35176477ac7b1

Observation 2d85cbd5-7560-4c54-8f46-8d607f89fb94 · outbound

This paper cites Restructuring Vector Quantization with the Rotation Trick.

HH-Codec: High Compression High-fidelity Discrete Neural Codec for Spoken Language Modeling Restructuring Vector Quantization with the Rotation Trick

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-15T18:11:55.339064Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:11:55.339064Z digest=sha256:544dcdd8feb224e8bac2fba40f312eddd87917cac879420d28fdf0520bdc2205

Observation f94f1c66-4099-4ddc-a52c-ce56e9e9c916 · outbound

This paper cites Moshi: a speech-text foundation model for real-time dialogue.

HH-Codec: High Compression High-fidelity Discrete Neural Codec for Spoken Language Modeling Moshi: a speech-text foundation model for real-time dialogue

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-15T18:11:55.323965Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:11:55.323965Z digest=sha256:cb96f40d7b3ddb7dc97a6e2a92813041d241837a807382597a87da548d966cd7

Observation 0a16ce14-c7d5-4a2c-bd3a-3876edc30654 · outbound

This paper cites Seed-TTS: A Family of High-Quality Versatile Speech Generation Models.

HH-Codec: High Compression High-fidelity Discrete Neural Codec for Spoken Language Modeling Seed-TTS: A Family of High-Quality Versatile Speech Generation Models

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-15T18:11:55.301847Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:11:55.301847Z digest=sha256:3891230ad174934a3b521aaa652f61f722bd19a03b72cadf0bb37935bebcbdd9

Observation a54e8fc2-36dc-4a05-aa11-07d2d9aa33c8 · outbound

This paper cites Vision Transformers Need Registers.

HH-Codec: High Compression High-fidelity Discrete Neural Codec for Spoken Language Modeling Vision Transformers Need Registers

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-15T18:11:55.318713Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T18:11:55.318713Z digest=sha256:ef30c02d209ad264bf36279034116124a84eea8fbe4c970f4d971aae030cf3da

Pith citing papers

Observation ec4ecc8f-4c80-40da-9c5e-df10243578cd · inbound

Listen, Pause, and Reason: Toward Perception-Grounded Hybrid Reasoning for Audio Understanding cites this paper.

Listen, Pause, and Reason: Toward Perception-Grounded Hybrid Reasoning for Audio Understanding HH-Codec: High Compression High-fidelity Discrete Neural Codec for Spoken Language Modeling

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-10T09:28:38.891667Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-10T09:24:25.616750Z digest=sha256:9e303e343064cc81d6804f19da88955835bda3a52bcf82a9d77c75e39dd105ec

Observation 4fd98e9a-bfcf-4e56-bc5f-8fea8212fb17 · inbound

CleanCodec: Efficient and Robust Speech Tokenization via Perceptually Guided Encoding cites this paper.

CleanCodec: Efficient and Robust Speech Tokenization via Perceptually Guided Encoding HH-Codec: High Compression High-fidelity Discrete Neural Codec for Spoken Language Modeling

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-06-28T05:21:39.776420Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-06-28T05:20:49.952030Z digest=sha256:ebf2eea43287f09606b6beab91d9c0e04c9604c46406a635b6bcfbcb6a59dca6