Pith. sign in

Paper Citation Record · LEDGER

Representing Speech Through Autoregressive Prediction of Cochlear Tokens

As of 20 August 2026, this Paper Citation Record lists 65 of 65 outbound references and 1 inbound Pith citation observation for arXiv:2508.11598.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.11598 v1

Coverage vector

measured 65 of 65 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T19:54:58.755219Z

measured 66 of 66 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T19:54:58.525812Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T19:54:59.403728Z

Reference resolution

65 of 65 outbound references displayed

  • verified exact4
  • verified fuzzy35
  • unresolved24
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c87c66bf-5d6a-4851-91b1-d73aff9c5ccf · outbound

This paper cites Representing Speech Through Autoregressive Prediction of Cochlear Tokens.

Representing Speech Through Autoregressive Prediction of Cochlear Tokens Representing Speech Through Autoregressive Prediction of Cochlear Tokens

Reference 1

Resolution
metadata mismatch
local_arxiv, observed 2026-08-05T19:54:59.463829Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T19:54:58.525812Z digest=sha256:73161a30f53ff0508dfdec08579a609160db6e6ab8490e6ef7882f2af16f70f9

Observation e70517cc-c426-43e4-b519-7ec9eba9316d · outbound

This paper cites water” and “river.

Representing Speech Through Autoregressive Prediction of Cochlear Tokens water” and “river

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:55:03.623324Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T19:54:58.530279Z digest=sha256:583e25ced8a5acd1fdb52ff02d7a7cbc2d7e19a70f713d8e5f3cb60d07a637cf

Observation 1dbc760f-0472-4349-bb5c-34037b0f7723 · outbound

This paper cites er” was often confused with “r.

Representing Speech Through Autoregressive Prediction of Cochlear Tokens er” was often confused with “r

Reference 3

Resolution
malformed identifier
raw_fallback, observed 2026-08-05T19:55:03.613327Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T19:54:58.534959Z digest=sha256:08cd3fdff0cfe2f9bd04a4d9b9224da97b5ab476978121b5034659ef2aaff087

Observation 8418ae8c-cc49-47e6-8870-8f8d089400a8 · outbound

This paper cites Transformation Imitation.

Representing Speech Through Autoregressive Prediction of Cochlear Tokens Transformation Imitation

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:55:03.603935Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T19:54:58.538930Z digest=sha256:8d0a833de38fc238dd74c8ad5b6fcffbcdde0b72e10de67fbc6beefab50c0932

Observation 87b8727c-cc16-4faa-a61c-702bb755647c · outbound

This paper cites acknowledges support from The K.

Representing Speech Through Autoregressive Prediction of Cochlear Tokens acknowledges support from The K

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:55:03.594277Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T19:54:58.542617Z digest=sha256:2fcb97a2573be6f2f1a6332fa90aa27cf50dfcec56da9c7f626e61fe15afdade

Observation d35525ca-2a30-43d0-9329-ae571f573311 · outbound

This paper cites SUPERB: Speech processing Universal PERformance Benchmark.

Representing Speech Through Autoregressive Prediction of Cochlear Tokens SUPERB: Speech processing Universal PERformance Benchmark

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-05T19:54:58.546041Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:54:58.546041Z digest=sha256:9ebece583c4bd5e6d1f1e16320815e8f8e27c0957ae088b1df4177285d2a6825

Observation 7257bd88-eb8c-483a-84c1-1faded21bf78 · outbound

This paper cites Codec-SUPERB @ SLT 2024: A lightweight benchmark for neural audio codec models.

Representing Speech Through Autoregressive Prediction of Cochlear Tokens Codec-SUPERB @ SLT 2024: A lightweight benchmark for neural audio codec models

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-08-05T19:54:59.341956Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T19:54:58.550290Z digest=sha256:926362a9ca004b7596ae5d94705a58b52ce81c2d9b55200abb9e85f83675b6cf

Observation ba6f1da7-8ec5-4af4-a722-1c40c19dfc80 · outbound

This paper cites A Review of Deep Learning Techniques for Speech Processing.

Representing Speech Through Autoregressive Prediction of Cochlear Tokens A Review of Deep Learning Techniques for Speech Processing

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-05T19:54:58.553897Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:54:58.553897Z digest=sha256:5832ecbcf9f734537195a90cc01d2505fb352076fa361c66f2dc5b617db399e3

Observation 86bff55e-0e93-4447-aa6c-b3938d7f5878 · outbound

This paper cites Soundstream: An end-to-end neural audio codec,.

Representing Speech Through Autoregressive Prediction of Cochlear Tokens Soundstream: An end-to-end neural audio codec,

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-05T19:54:58.558742Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:54:58.558742Z digest=sha256:c8dbab6eb4a38b0bee3ad0f5939b54a3efaa3a8801808e7f4b13eb10a1b4cc71

Observation 587e6140-14f7-4998-a625-0c5aa2904316 · outbound

This paper cites High Fidelity Neural Audio Compression.

Representing Speech Through Autoregressive Prediction of Cochlear Tokens High Fidelity Neural Audio Compression

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-05T19:54:58.561899Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:54:58.561899Z digest=sha256:5d49b5619623ba24a14a7de45f0fbb0dfff71655221ddec3057ebd7066a4c4cf

Observation f225f353-cfd7-4e63-b6bc-70213e002bb5 · outbound

This paper cites HiFi-Codec: Group-residual Vector quantization for High Fidelity Audio Codec.

Representing Speech Through Autoregressive Prediction of Cochlear Tokens HiFi-Codec: Group-residual Vector quantization for High Fidelity Audio Codec

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-05T19:54:58.565677Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:54:58.565677Z digest=sha256:f13e809273bdd51832f688da9747a1847e8e8ee221c89b8e1661220f119fa9fa

Observation 4075519f-d176-4b8a-bce3-882d355bb389 · outbound

This paper cites High-fidelity audio compression with improved rvqgan,.

Representing Speech Through Autoregressive Prediction of Cochlear Tokens High-fidelity audio compression with improved rvqgan,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:55:03.577698Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T19:54:58.569242Z digest=sha256:7f9280034b5106f08d923279b8b444647f7c93095951000eda85201b146fea7a

Observation b43b3d0a-1d76-4908-9c67-cf1874151514 · outbound

This paper cites Language-Codec: Bridging Discrete Codec Representations and Speech Language Models.

Representing Speech Through Autoregressive Prediction of Cochlear Tokens Language-Codec: Bridging Discrete Codec Representations and Speech Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-05T19:54:58.572361Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:54:58.572361Z digest=sha256:fabb01c8627e73f87c8266c0a2ae10580caa6bf03f71db2a4a5d2a9cc9daea50

Observation 3f67f68b-b935-425c-86b6-f24317671d3e · outbound

This paper cites Speechtok- enizer: Unified speech tokenizer for speech language models,.

Representing Speech Through Autoregressive Prediction of Cochlear Tokens Speechtok- enizer: Unified speech tokenizer for speech language models,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:55:03.566990Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T19:54:58.575845Z digest=sha256:5efa50f1f6e986513d3ee994527b92904522436c1223bbbd436d4cd9879bffcc

Observation 61224e10-6883-4e89-a553-dd944a8b051c · outbound

This paper cites CLaM-TTS: Improving Neural Codec Language Model for Zero-Shot Text-to-Speech.

Representing Speech Through Autoregressive Prediction of Cochlear Tokens CLaM-TTS: Improving Neural Codec Language Model for Zero-Shot Text-to-Speech

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-05T19:54:58.579073Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:54:58.579073Z digest=sha256:ec154f35ffde02817ed8c78afb7c5b619824bdaa78239df81f0d6c71af8d92a7

Observation 1886f8f4-88fc-4818-9a5b-9a573165cfbb · outbound

This paper cites Spectral Codecs: Improving Non-Autoregressive Speech Synthesis with Spectrogram-Based Audio Codecs.

Representing Speech Through Autoregressive Prediction of Cochlear Tokens Spectral Codecs: Improving Non-Autoregressive Speech Synthesis with Spectrogram-Based Audio Codecs

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-05T19:54:58.582582Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:54:58.582582Z digest=sha256:66cacb6046ce029befa3c1ea7d2293b22cc5f4090359d2b83567b0f09c8ff4c2

Observation 7f87b3e5-0f62-484c-a2bf-6ee9e158b49b · outbound

This paper cites Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers.

Representing Speech Through Autoregressive Prediction of Cochlear Tokens Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-05T19:54:58.586186Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:54:58.586186Z digest=sha256:9763636faeacaaefeabc47143043cfd25e7c5f573beef550e83bfdaf743e5a2e

Observation e7744f64-4ab0-4588-a3cc-55690244c93e · outbound

This paper cites Hierarchical organization of hu- man auditory cortex: evidence from acoustic invariance in the re- sponse to intelligible speech,.

Representing Speech Through Autoregressive Prediction of Cochlear Tokens Hierarchical organization of hu- man auditory cortex: evidence from acoustic invariance in the re- sponse to intelligible speech,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:55:03.462945Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T19:54:58.589508Z digest=sha256:45eba804ad1c11fd2886d8d5e082bbe2f60a4d4b0555a801bebbc90c4d3a395d

Observation 5d4cc089-2728-4852-b9a9-0427f40f18a8 · outbound

This paper cites Intonational speech prosody encoding in the human auditory cortex,.

Representing Speech Through Autoregressive Prediction of Cochlear Tokens Intonational speech prosody encoding in the human auditory cortex,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:55:03.286871Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T19:54:58.592847Z digest=sha256:b23ea841269bd3155d00930b613964c6583866e683e7918863841186adb28fe6

Observation 8464c6ed-1479-49e3-b937-6b7d393d173a · outbound

This paper cites Hubert: Self-supervised speech rep- resentation learning by masked prediction of hidden units,.

Representing Speech Through Autoregressive Prediction of Cochlear Tokens Hubert: Self-supervised speech rep- resentation learning by masked prediction of hidden units,

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-05T19:54:58.595978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:54:58.595978Z digest=sha256:8559b9c63386f264027567d063cc3b0b0c85e676f76a53b7ff9b72de12b576f8

Observation 9516fb69-478a-4798-bfa4-95e9d6cda6bf · outbound

This paper cites W2v-BERT: Combining Contrastive Learning and Masked Language Modeling for Self-Supervised Speech Pre-Training.

Representing Speech Through Autoregressive Prediction of Cochlear Tokens W2v-BERT: Combining Contrastive Learning and Masked Language Modeling for Self-Supervised Speech Pre-Training

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-05T19:54:58.599394Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:54:58.599394Z digest=sha256:706c4046887eb8adc27402dcd5e4f1b36c6e63e79293bd375e7363af98c54abb

Observation 43f76c7d-ebb6-4dc9-9aa6-e9fbefa6023f · outbound

This paper cites Wavlm: Large-scale self- supervised pre-training for full stack speech processing,.

Representing Speech Through Autoregressive Prediction of Cochlear Tokens Wavlm: Large-scale self- supervised pre-training for full stack speech processing,

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-05T19:54:58.602756Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:54:58.602756Z digest=sha256:77cc35a02f5819f9c211afa4829df40448a61dc9a8870d54a3520ec6e6e3befd

Observation ed71dc35-6842-4fef-9545-5fdea44e45e5 · outbound

This paper cites Blind phoneme segmentation with temporal prediction errors.

Representing Speech Through Autoregressive Prediction of Cochlear Tokens Blind phoneme segmentation with temporal prediction errors

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-08-05T19:54:59.172535Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T19:54:58.605940Z digest=sha256:33586f962e3b75e4b58c131d2207e1cd4176dcf5d1d96cde3e1a12d4b7385787

Observation b177be0e-b3b8-4219-b791-7006268441cc · outbound

This paper cites An Unsupervised Autoregressive Model for Speech Representation Learning.

Representing Speech Through Autoregressive Prediction of Cochlear Tokens An Unsupervised Autoregressive Model for Speech Representation Learning

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-05T19:54:58.609844Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:54:58.609844Z digest=sha256:6bd96bdf818e6a79bdb3f10f25223e5ee10d1ba25601383933feb4b6c0785ccc

Observation 73d39a96-2ea6-41e1-881d-0a986ab7417b · outbound

This paper cites Acquiring language from speech by learning to remember and predict,.

Representing Speech Through Autoregressive Prediction of Cochlear Tokens Acquiring language from speech by learning to remember and predict,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:55:03.113543Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T19:54:58.613387Z digest=sha256:4d0d361c15f3b97102dd1a5003dbe4b250212da1b8158295bcc428c84f4b48aa

Observation 04458adf-6e74-495f-ae62-26106e634722 · outbound

This paper cites Vector-Quantized Autoregressive Predictive Coding.

Representing Speech Through Autoregressive Prediction of Cochlear Tokens Vector-Quantized Autoregressive Predictive Coding

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-05T19:54:58.616637Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:54:58.616637Z digest=sha256:b4f2cc9f8b82a6ac7cf8c67567f5399e6532882900a29b3faeef5921ae04a051

Observation 4125ea9c-5623-4cdf-97da-333987beef75 · outbound

This paper cites Audio albert: A lite bert for self-supervised learning of audio representation,.

Representing Speech Through Autoregressive Prediction of Cochlear Tokens Audio albert: A lite bert for self-supervised learning of audio representation,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:55:02.985774Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T19:54:58.621129Z digest=sha256:5eb3c0fced8c773555774671293ef185186e5148ea7932e74ca91781875fa537

Observation 6774d475-e580-4647-9742-715858ad99a6 · outbound

This paper cites On generative spoken language modeling from raw au- dio,.

Representing Speech Through Autoregressive Prediction of Cochlear Tokens On generative spoken language modeling from raw au- dio,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:55:02.859177Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T19:54:58.624665Z digest=sha256:3e31e2ae3a518491993b8eb143d4ad9b1c63864ef267260e6dafb200386703a8

Observation c35d794d-6ba2-43a0-a662-565838525c7c · outbound

This paper cites Audiolm: a language modeling approach to audio gener- ation,.

Representing Speech Through Autoregressive Prediction of Cochlear Tokens Audiolm: a language modeling approach to audio gener- ation,

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-05T19:54:58.628627Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:54:58.628627Z digest=sha256:3764cde1f2d3271dcb17de0c48ee3164b7b8f6c27c08c5b69b5190ec31171d95

Observation 4a55b06c-d015-495f-a306-39f267076dbf · outbound

This paper cites Improving Textless Spoken Language Understanding with Discrete Units as Intermediate Target.

Representing Speech Through Autoregressive Prediction of Cochlear Tokens Improving Textless Spoken Language Understanding with Discrete Units as Intermediate Target

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-08-05T19:54:59.113274Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T19:54:58.632070Z digest=sha256:c5d0e61b6ecd9fb40a73b6a3ec3f13b7038b9434c53028702f828924db0a5053

Observation 4eb29282-33de-4d8d-b2ac-40aa568d4df9 · outbound

This paper cites BERT: Pre- training of Deep Bidirectional Transformers for Language Under- standing,.

Representing Speech Through Autoregressive Prediction of Cochlear Tokens BERT: Pre- training of Deep Bidirectional Transformers for Language Under- standing,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:55:02.733522Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T19:54:58.635519Z digest=sha256:542a1031d3fb13f48056bbf7176d5a8b10df92f1c00534398ea6e60d04145767

Observation 54f7547d-408e-47fd-b58f-3c2612a5e04e · outbound

This paper cites Representation Learning with Contrastive Predictive Coding.

Representing Speech Through Autoregressive Prediction of Cochlear Tokens Representation Learning with Contrastive Predictive Coding

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-05T19:54:58.639389Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:54:58.639389Z digest=sha256:3a1c0cf555f3e34df69d1bd362ec48c34db7ca9671867cc77d9ff16b473dc130

Observation 05670626-1df8-47d1-a877-7bccc75d583d · outbound

This paper cites Wav2vec 2.0: a framework for self-supervised learning of speech representa- tions,.

Representing Speech Through Autoregressive Prediction of Cochlear Tokens Wav2vec 2.0: a framework for self-supervised learning of speech representa- tions,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:55:02.595018Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T19:54:58.642906Z digest=sha256:17109d3c5cdcf9fa06cf4d75e0a8185bcc41574a2d92485ed708e8d6f4a56db8

Observation d0ea69db-a483-47df-a0f9-960dc3ab132b · outbound

This paper cites Contrastive learning of general-purpose audio representations,.

Representing Speech Through Autoregressive Prediction of Cochlear Tokens Contrastive learning of general-purpose audio representations,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:55:02.473345Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T19:54:58.647061Z digest=sha256:668fb7ad8a1e4cbcc9a54352684ad5616a374078605b93b81a8143b83a05508b

Observation c036ce17-2c66-4ac3-a5f0-bec0c25cf112 · outbound

This paper cites Unsupervised speech segmentation and variable rate repre- sentation learning using segmental contrastive predictive coding,.

Representing Speech Through Autoregressive Prediction of Cochlear Tokens Unsupervised speech segmentation and variable rate repre- sentation learning using segmental contrastive predictive coding,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:55:02.333859Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T19:54:58.650700Z digest=sha256:0cb381925770c537ee38dc36a17a2b092259d21b08b3dfe175a24deec31d3a89

Observation d8f38926-fe18-4bef-b40f-325195474c9d · outbound

This paper cites Self-normalization and noise- robustness in early auditory representations,.

Representing Speech Through Autoregressive Prediction of Cochlear Tokens Self-normalization and noise- robustness in early auditory representations,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:55:02.231236Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T19:54:58.653998Z digest=sha256:253c39f58bdd894162ba711c14760ce8d77f348b71d92d4949f2e2c1def13b5a

Observation a5178f83-656d-412d-a88b-be14c90d6f35 · outbound

This paper cites an unresolved cited work.

Representing Speech Through Autoregressive Prediction of Cochlear Tokens Unresolved cited work

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-05T19:54:58.657328Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:54:58.657328Z digest=sha256:42d6bea77d7b017c29046fb7d5a94edb50babf4777724b8e471952a3cfbd4eea

Observation 764197d2-608a-47ea-a155-1fae0d4a3019 · outbound

This paper cites Multiresolution spectrotem- poral analysis of complex sounds,.

Representing Speech Through Autoregressive Prediction of Cochlear Tokens Multiresolution spectrotem- poral analysis of complex sounds,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:55:02.100079Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T19:54:58.660417Z digest=sha256:a49da221486f8d2c41754bc1073cbf8af47fe24a14eb0a141bc860ca7b28bf57

Observation 7dfd53ae-2f9a-4016-9a9a-a5fa14fbd665 · outbound

This paper cites Model metamers reveal divergent invariances between biological and ar- tificial neural networks,.

Representing Speech Through Autoregressive Prediction of Cochlear Tokens Model metamers reveal divergent invariances between biological and ar- tificial neural networks,

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:55:01.976683Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T19:54:58.663600Z digest=sha256:dc0ad1e50afb59784654d12966ab0258141528917a2c02a857a621ac9d739b24

Observation 6a985c03-b404-43b8-903e-93fc23763d3c · outbound

This paper cites Derivation of auditory filter shapes from notched-noise data,.

Representing Speech Through Autoregressive Prediction of Cochlear Tokens Derivation of auditory filter shapes from notched-noise data,

Reference 40

Resolution
verified exact
raw_fallback, observed 2026-08-05T19:54:58.998024Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T19:54:58.666884Z digest=sha256:3701348b3a334df524b342ba9f4fbf31d143273558dc3d1cee008fcadb38e28f

Observation 71303311-fc97-4f31-aa77-67021373af4d · outbound

This paper cites Sound texture perception via statistics of the auditory periphery: evidence from sound syn- thesis,.

Representing Speech Through Autoregressive Prediction of Cochlear Tokens Sound texture perception via statistics of the auditory periphery: evidence from sound syn- thesis,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:55:01.847895Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T19:54:58.670026Z digest=sha256:1478b02805ecc6988d5be38c22e4814eb2bd67c79b761aca4762fa2f41fa7c00

Observation c0098db9-3dc5-4551-b07b-e52f6010aaf4 · outbound

This paper cites Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation.

Representing Speech Through Autoregressive Prediction of Cochlear Tokens Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-05T19:54:58.673271Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:54:58.673271Z digest=sha256:0af75676c931db33e7561cc2f871493c5735dcf83f7b9552243b180602e1b0d6

Observation 2b59ce30-34e4-413c-8261-07bcdeafbaef · outbound

This paper cites Lib- rispeech: an asr corpus based on public domain audio books,.

Representing Speech Through Autoregressive Prediction of Cochlear Tokens Lib- rispeech: an asr corpus based on public domain audio books,

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-05T19:54:58.677108Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:54:58.677108Z digest=sha256:228abad42108f2036d64b64a9318538739601701b6871996397fc9cc37b8bf5f

Observation 653e79ba-084b-461e-be72-dbbf8515a5e7 · outbound

This paper cites Improving language understanding by generative pre-training,.

Representing Speech Through Autoregressive Prediction of Cochlear Tokens Improving language understanding by generative pre-training,

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:55:01.721880Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T19:54:58.680937Z digest=sha256:bb3d1ec80a945e1a11d68c57ef56e70833d0350351f511079cf3cdb20e9bbfe3

Observation f5858c51-8ec3-42f6-899c-28e3d76aa871 · outbound

This paper cites Gaussian Error Linear Units (GELUs).

Representing Speech Through Autoregressive Prediction of Cochlear Tokens Gaussian Error Linear Units (GELUs)

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-05T19:54:58.684181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:54:58.684181Z digest=sha256:1311235334e26ea10ab5e9cf88f16291110a0f2735d26cb3c1d53bb7fa68cd79

Observation 1e70758c-5b26-475c-ab84-35265a244a88 · outbound

This paper cites Root Mean Square Layer Normalization.

Representing Speech Through Autoregressive Prediction of Cochlear Tokens Root Mean Square Layer Normalization

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-05T19:54:58.688240Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:54:58.688240Z digest=sha256:0dc83bafcfc02cb28b38bbf0dfb51c9cb501ac5a6afdd9f05dc73b94b2ffb6ee

Observation 3d165a1c-5b8b-4e06-a373-0f6b2f87e3a8 · outbound

This paper cites Libri-light: A benchmark for asr with limited or no supervision,.

Representing Speech Through Autoregressive Prediction of Cochlear Tokens Libri-light: A benchmark for asr with limited or no supervision,

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:55:01.611021Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T19:54:58.691750Z digest=sha256:3aec017b23eead1e9693ac42b9045efef9744931e01eac6b060943712bca6d9c

Observation e2c0c792-4af0-45a9-a891-87aa239c8340 · outbound

This paper cites Timit acoustic phonetic continuous speech cor- pus,.

Representing Speech Through Autoregressive Prediction of Cochlear Tokens Timit acoustic phonetic continuous speech cor- pus,

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:55:01.458253Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T19:54:58.695205Z digest=sha256:5ca7795e86fc8092c3dd787d3436c4481dd22918ea69db216ec8606a544d38d0

Observation 98d1da4f-be71-4a68-852d-3bc94bbb395d · outbound

This paper cites Speaker-independent phone recogni- tion using hidden markov models,.

Representing Speech Through Autoregressive Prediction of Cochlear Tokens Speaker-independent phone recogni- tion using hidden markov models,

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:55:01.279439Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T19:54:58.699082Z digest=sha256:19396cafa1143d30e74ae14ad688dae4b677c42f61c39fc4a0b296d68ca7a89d

Observation eb7042ec-1ac7-468a-b583-3db817e346b2 · outbound

This paper cites Scikit-learn: Machine Learning in Python,.

Representing Speech Through Autoregressive Prediction of Cochlear Tokens Scikit-learn: Machine Learning in Python,

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:55:01.140934Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T19:54:58.703299Z digest=sha256:c1b28001c654dcf204f8367f6330eb8f2a602f85b05f95dbe508510e370fe387

Observation a2d0db29-082a-4ab4-8bf5-da843b7777f7 · outbound

This paper cites The Zero Resource Speech Benchmark 2021: Metrics and baselines for unsupervised spoken language modeling.

Representing Speech Through Autoregressive Prediction of Cochlear Tokens The Zero Resource Speech Benchmark 2021: Metrics and baselines for unsupervised spoken language modeling

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-05T19:54:58.706814Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:54:58.706814Z digest=sha256:b67002b9b95b50d095dd169fc82ea212e4baf715d274f19631438df165b468e7

Observation 38b04a89-0a8f-4ecd-b6ce-a817b2096755 · outbound

This paper cites Over-reliance on english hinders cognitive science,.

Representing Speech Through Autoregressive Prediction of Cochlear Tokens Over-reliance on english hinders cognitive science,

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:55:01.002185Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T19:54:58.710255Z digest=sha256:3e28b9f3467191c578dafa5de3cb710c6fb5897b9cc98ba9dfc66cda29b6d70c

Observation 099f2fd0-5e9a-45d3-8cef-f1ff9189d8b0 · outbound

This paper cites To- wards inclusive automatic speech recognition,.

Representing Speech Through Autoregressive Prediction of Cochlear Tokens To- wards inclusive automatic speech recognition,

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:55:00.891669Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T19:54:58.713597Z digest=sha256:c9e708ce104016855569b1b067d2fe9e5b81b935474877d0aaa931efe4f91cfb

Observation 3aaeb462-9d71-4e5c-ba65-82b91e492f2a · outbound

This paper cites Say- cam: A large, longitudinal audiovisual dataset recorded from the infant’s perspective,.

Representing Speech Through Autoregressive Prediction of Cochlear Tokens Say- cam: A large, longitudinal audiovisual dataset recorded from the infant’s perspective,

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:55:00.778519Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T19:54:58.716758Z digest=sha256:9766ec81ae1ee718f63690d3287586e339d7e06393abf58f7ad368dde9593800

Observation 071b6908-1925-49bf-b58e-757b99b7b746 · outbound

This paper cites Call for Papers -- The BabyLM Challenge: Sample-efficient pretraining on a developmentally plausible corpus.

Representing Speech Through Autoregressive Prediction of Cochlear Tokens Call for Papers -- The BabyLM Challenge: Sample-efficient pretraining on a developmentally plausible corpus

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-05T19:54:58.720381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:54:58.720381Z digest=sha256:f53e5158cc2cd953efb81d96f3ee9cdfd8226d0168dd90f2843c595afe24f46c

Observation 19b671da-5252-4497-9d4b-a0caffead11c · outbound

This paper cites A task-optimized neural network repli- cates human auditory behavior, predicts brain responses, and re- veals a cortical processing hierarchy,.

Representing Speech Through Autoregressive Prediction of Cochlear Tokens A task-optimized neural network repli- cates human auditory behavior, predicts brain responses, and re- veals a cortical processing hierarchy,

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:55:00.644771Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T19:54:58.723735Z digest=sha256:d175050690b5939bd93a21f29eca5264d0db8055bc7cc3936ac8b88d191fee04

Observation d7f67bdb-df30-48c7-bf43-d124bd131507 · outbound

This paper cites Toward a realistic model of speech processing in the brain with self-supervised learning,.

Representing Speech Through Autoregressive Prediction of Cochlear Tokens Toward a realistic model of speech processing in the brain with self-supervised learning,

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:55:00.557865Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T19:54:58.727007Z digest=sha256:01ca39b27ac125dd1c07844138e6382eb916fd1b857ccfc9bae2f57cbf34cbe5

Observation 02a3ce6b-c9d5-4b31-b65d-c42660187c23 · outbound

This paper cites Dissecting neural computations of the human auditory pathway using deep neural networks for speech,.

Representing Speech Through Autoregressive Prediction of Cochlear Tokens Dissecting neural computations of the human auditory pathway using deep neural networks for speech,

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:55:00.429799Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T19:54:58.730190Z digest=sha256:5a641568936227f98a05cd30492adde3eee77b46fb0de44dce76ba9b3326866b

Observation aee78f11-cfe0-443e-8813-2da3653d8bce · outbound

This paper cites Many but not all deep neural network audio models capture brain responses and exhibit correspondence between model stages and brain regions,.

Representing Speech Through Autoregressive Prediction of Cochlear Tokens Many but not all deep neural network audio models capture brain responses and exhibit correspondence between model stages and brain regions,

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:55:00.297487Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T19:54:58.733488Z digest=sha256:9efa4ecaf88b6b594f5116f99df8918865c88d58e4422d6e14eb638649c44edb

Observation 91f4b4e3-66e6-4218-ac8c-0418f1508ccb · outbound

This paper cites Speech taskonomy: Which speech tasks are the most predictive of fmri brain activity?.

Representing Speech Through Autoregressive Prediction of Cochlear Tokens Speech taskonomy: Which speech tasks are the most predictive of fmri brain activity?

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:55:00.192596Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T19:54:58.736777Z digest=sha256:786f0b52a50fb464fe53e78801f0c8162a73fdbb7a61c4bc0b50ce043d3b66a4

Observation 3aff2a49-132d-4df5-803b-27c0a1a5079a · outbound

This paper cites Language in brains, minds, and machines,.

Representing Speech Through Autoregressive Prediction of Cochlear Tokens Language in brains, minds, and machines,

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:55:00.058048Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T19:54:58.741001Z digest=sha256:cccc94290ee87911c1464f49d0504be048c92965c0b6cfba25641a8d255592f5

Observation 0264a1d7-45ec-4d1d-a2b8-29e6e9f2387a · outbound

This paper cites Brain-tuned Speech Models Better Reflect Speech Processing Stages in the Brain.

Representing Speech Through Autoregressive Prediction of Cochlear Tokens Brain-tuned Speech Models Better Reflect Speech Processing Stages in the Brain

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-05T19:54:58.744160Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:54:58.744160Z digest=sha256:11b49fbdcc92564cd2f36bc5aeeb810b100d2430f0be536e8be912fc7cf7ce74

Observation c3849522-5e98-42e5-b262-13da0475a8e1 · outbound

This paper cites An algorithm for the machine calculation of complex fourier series,.

Representing Speech Through Autoregressive Prediction of Cochlear Tokens An algorithm for the machine calculation of complex fourier series,

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:54:59.936362Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T19:54:58.747668Z digest=sha256:fae085fe84670b5f196fbc49255aadf97f6a7e4cd82baafae85b3037d2eea46f

Observation 13352801-f597-4a72-8ac0-fc4301aedecb · outbound

This paper cites Lib- rispeech: An ASR corpus based on public domain audio books,.

Representing Speech Through Autoregressive Prediction of Cochlear Tokens Lib- rispeech: An ASR corpus based on public domain audio books,

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:54:59.764253Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T19:54:58.751677Z digest=sha256:19199f6ffd8a24323f4bcc9c1efaf0e084efe0639512ed024d6d6fab4b3c95ca

Observation ec17cb52-6a2a-46f9-a014-6dc96e011ce3 · outbound

This paper cites codebook usage.

Representing Speech Through Autoregressive Prediction of Cochlear Tokens codebook usage

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:54:59.574844Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T19:54:58.755219Z digest=sha256:145b51ac92c6d04053fcf27272b7fe20c2b39859e46a87af384d15b748e8b9ba

Pith citing papers

Observation c87c66bf-5d6a-4851-91b1-d73aff9c5ccf · inbound

Representing Speech Through Autoregressive Prediction of Cochlear Tokens cites this paper.

Representing Speech Through Autoregressive Prediction of Cochlear Tokens Representing Speech Through Autoregressive Prediction of Cochlear Tokens

Reference 1

Resolution
metadata mismatch
local_arxiv, observed 2026-08-05T19:54:59.463829Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T19:54:58.525812Z digest=sha256:73161a30f53ff0508dfdec08579a609160db6e6ab8490e6ef7882f2af16f70f9