Pith. sign in

Paper Citation Record · LEDGER

vq-wav2vec: Self-Supervised Learning of Discrete Speech Representations

As of 20 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 19 inbound Pith citation observations for arXiv:1910.05453.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
1910.05453 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 19 of 19 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 19 of 19 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T16:53:27.032602Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T03:27:34.881381Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation bd5da3c8-f4b7-40d2-acf0-ff3087252564 · inbound

VPBSD:Vessel-Pattern-Based Semi-Supervised Distillation for Efficient 3D Microscopic Cerebrovascular Segmentation cites this paper.

VPBSD:Vessel-Pattern-Based Semi-Supervised Distillation for Efficient 3D Microscopic Cerebrovascular Segmentation vq-wav2vec: Self-Supervised Learning of Discrete Speech Representations

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-12T20:35:51.368719Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:35:51.368719Z digest=sha256:69e65bbe0555923cc15830be99a6e2f9b9b96370c73ea4937b9614abc88e7159

Observation 1ed81267-acbe-47a6-b11a-49bbf046eafd · inbound

ESTVocoder: An Excitation-Spectral-Transformed Neural Vocoder Conditioned on Mel Spectrogram cites this paper.

ESTVocoder: An Excitation-Spectral-Transformed Neural Vocoder Conditioned on Mel Spectrogram vq-wav2vec: Self-Supervised Learning of Discrete Speech Representations

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-12T18:47:51.315555Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T18:47:51.315555Z digest=sha256:08b2980cb18b54918a5d6948872c46768fc8aa09456fab6958799605b41c32ca

Observation 058cc7c1-7830-4a7c-9327-49db50e7c999 · inbound

Comparative Analysis of Audio Feature Extraction for Real-Time Talking Portrait Synthesis cites this paper.

Comparative Analysis of Audio Feature Extraction for Real-Time Talking Portrait Synthesis vq-wav2vec: Self-Supervised Learning of Discrete Speech Representations

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-12T16:47:09.012586Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:47:09.012586Z digest=sha256:02d27b73d9bf7e8ca5ba73067c0552c88404e7406d25d230c8a5b719d895ce8b

Observation f66d3ebe-02a5-4a06-9377-416c6743b044 · inbound

The SVASR System for Text-dependent Speaker Verification (TdSV) AAIC Challenge 2024 cites this paper.

The SVASR System for Text-dependent Speaker Verification (TdSV) AAIC Challenge 2024 vq-wav2vec: Self-Supervised Learning of Discrete Speech Representations

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-12T13:27:06.782418Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T13:27:06.782418Z digest=sha256:67132eb13da52bb9b1992e87019bb60d578973dea5cbc38f3dfea5e491b513ab

Observation c9d80b7b-e771-4944-bb35-9a1a7226f957 · inbound

Efficient Video Face Enhancement with Enhanced Spatial-Temporal Consistency cites this paper.

Efficient Video Face Enhancement with Enhanced Spatial-Temporal Consistency vq-wav2vec: Self-Supervised Learning of Discrete Speech Representations

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T13:09:28.731472Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:09:28.731472Z digest=sha256:90d20e629fcf5cc84f35e4cd443acbc1978071ba724c25e66c79676d16870272

Observation 6cb689c4-cabb-413c-86e2-2fc7fd27fe21 · inbound

Scalable Image Tokenization with Index Backpropagation Quantization cites this paper.

Scalable Image Tokenization with Index Backpropagation Quantization vq-wav2vec: Self-Supervised Learning of Discrete Speech Representations

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T23:15:46.089540Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:15:46.089540Z digest=sha256:26b724551bb4c4057673d5832dbee083082c294bfaf9fdbb1192584622b3ed40

Observation f73d5b1f-11ef-4532-a430-e588e16558cd · inbound

Bilevel Joint Unsupervised and Supervised Training for Automatic Speech Recognition cites this paper.

Bilevel Joint Unsupervised and Supervised Training for Automatic Speech Recognition vq-wav2vec: Self-Supervised Learning of Discrete Speech Representations

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T17:51:21.274491Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:51:21.274491Z digest=sha256:54d1b3f5955cce8a8a628dca688b4d67522e8b315768ea7d298435368c81358a

Observation 3eafa8b5-a3f6-48ba-b8ff-647d49e61366 · inbound

Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey cites this paper.

Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey vq-wav2vec: Self-Supervised Learning of Discrete Speech Representations

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T14:59:01.130923Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:59:01.130923Z digest=sha256:dd39a3d4c85dc3e9d09a74b6ca9d4931e50b0008adf999745a6674950b163083

Observation 3d1b2d55-e44f-4080-9d8a-81bdb4d79a59 · inbound

Multimodal Medical Code Tokenizer cites this paper.

Multimodal Medical Code Tokenizer vq-wav2vec: Self-Supervised Learning of Discrete Speech Representations

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-09T00:42:56.928928Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T00:42:56.928928Z digest=sha256:839e5b54b8dc9c0752a8b5f2c5d01ac015615998cfbc8bdd03adebb300c3c35d

Observation 3d42f6fe-bbc3-447c-92a1-b6a758faf5cf · inbound

Heterogeneous Self-Supervised Acoustic Pre-Training with Local Constraints cites this paper.

Heterogeneous Self-Supervised Acoustic Pre-Training with Local Constraints vq-wav2vec: Self-Supervised Learning of Discrete Speech Representations

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T16:53:27.032602Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:53:27.032602Z digest=sha256:646b2b64426bf23e719652d40b270ef13bdc20b02431daee219b0f7161144bea

Observation f358c909-a9a3-40c7-8130-9d129917742f · inbound

Bitrate-Controlled Diffusion for Disentangling Motion and Content in Video cites this paper.

Bitrate-Controlled Diffusion for Disentangling Motion and Content in Video vq-wav2vec: Self-Supervised Learning of Discrete Speech Representations

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-04T20:44:59.929335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:44:59.929335Z digest=sha256:f2a482d8a443975650f0577639d49942cb17043a820b3fd4b3ddb14b26c6435f

Observation 25d505d1-4f49-424e-9ff5-4b9809c63378 · inbound

StableToken: A Noise-Robust Semantic Speech Tokenizer for Resilient SpeechLLMs cites this paper.

StableToken: A Noise-Robust Semantic Speech Tokenizer for Resilient SpeechLLMs vq-wav2vec: Self-Supervised Learning of Discrete Speech Representations

Reference 5

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T13:01:24.421342Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-18T12:57:04.450462Z digest=sha256:540cde2222f1a65f477039fe7bb0802f47e05585684f67802981c36b685ed3f5

Observation cf743998-8796-4234-9694-97d682b945bc · inbound

Incomplete Multi-View Multi-Label Classification via Shared Codebook and Fused-Teacher Self-Distillation cites this paper.

Incomplete Multi-View Multi-Label Classification via Shared Codebook and Fused-Teacher Self-Distillation vq-wav2vec: Self-Supervised Learning of Discrete Speech Representations

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T16:53:00.174441Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-13T16:50:19.940458Z digest=sha256:48397cfd28ba623e5f0f23d5f781df3bfce900072366a6b5617e4e1ad9705238

Observation 4857b5f0-a324-44d1-9439-9d915dabba4f · inbound

PairAlign: A Framework for Sequence Tokenization via Self-Alignment with Applications to Audio Tokenization cites this paper.

PairAlign: A Framework for Sequence Tokenization via Self-Alignment with Applications to Audio Tokenization vq-wav2vec: Self-Supervised Learning of Discrete Speech Representations

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-11T19:16:08.887715Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-08T12:25:52.847432Z digest=sha256:bdf2aa39487bc9cd282c067cb7245a57397a71d99f518e732d4acaa4aa8f15b1

Observation a62efb3b-79da-4838-af2e-43f700c455bc · inbound

PairAlign: A Framework for Sequence Tokenization via Self-Alignment with Applications to Audio Tokenization cites this paper.

PairAlign: A Framework for Sequence Tokenization via Self-Alignment with Applications to Audio Tokenization vq-wav2vec: Self-Supervised Learning of Discrete Speech Representations

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-06-30T23:15:07.879196Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-30T23:14:32.494076Z digest=sha256:d289277a6e1f06a7d5bffea1d56c97083b9141b2a539a5f57969d4a0e348138d

Observation 2082989e-05b8-4318-81fd-1a4cd4ee0a7a · inbound

Multilingual and Multimodal LLMs in the Wild: Building for Low-Resource Languages cites this paper.

Multilingual and Multimodal LLMs in the Wild: Building for Low-Resource Languages vq-wav2vec: Self-Supervised Learning of Discrete Speech Representations

Reference 235

Resolution
verified exact
arxiv_id, observed 2026-05-20T14:38:21.749867Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-20T14:33:36.100966Z digest=sha256:e6078f69ff4d7b799e2a1586c21b1e4db810d3b0f746e77156fe829d25fd1a59

Observation 76f3b6db-b6af-4849-8009-e46213ca13ca · inbound

How Optimality Structures Sparse Dictionaries: A Theory for Understanding SAE Representations cites this paper.

How Optimality Structures Sparse Dictionaries: A Theory for Understanding SAE Representations vq-wav2vec: Self-Supervised Learning of Discrete Speech Representations

Reference 272

Resolution
verified exact
arxiv_id, observed 2026-07-02T01:46:26.206836Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-06-28T11:36:45.020411Z digest=sha256:dfcb791dac7523f4ce384f49df938dd72fd48939033f5de20bb58f77b914e16b

Observation db3eab54-383c-4847-9ac9-28123ace8316 · inbound

End-to-End Training for Discrete Token LLM based TTS System cites this paper.

End-to-End Training for Discrete Token LLM based TTS System vq-wav2vec: Self-Supervised Learning of Discrete Speech Representations

Reference 24

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T03:27:34.882863Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-27T15:22:06.893507Z digest=sha256:7f77bdb4fc886e974c996355d43378d7482a5660207a5bea26ff966d31910201

Observation 0753e8d9-374c-4dd6-8a37-673cfd273461 · inbound

Seeds Before Objectives: Rethinking Evaluation for Low-Resource Garhwali ASR cites this paper.

Seeds Before Objectives: Rethinking Evaluation for Low-Resource Garhwali ASR vq-wav2vec: Self-Supervised Learning of Discrete Speech Representations

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-12T19:47:23.643442Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:47:23.643442Z digest=sha256:b661ae1a4d7583fdd632580b73269f7f4c35c896fedd3af5a887a19b14dd9c54