Pith. sign in

Paper Citation Record · LEDGER

Representing Speech Through Autoregressive Prediction of Cochlear Tokens

As of 8 August 2026, this Paper Citation Record lists 65 of 65 outbound references and 1 inbound Pith citation observation for arXiv:2508.11598.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.11598 v1

Coverage vector

measured 65 of 65 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T19:54:58.755219Z

measured 66 of 66 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T19:54:58.525812Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T19:54:59.403728Z

Reference resolution

65 of 65 outbound references displayed

  • verified exact4
  • verified fuzzy35
  • unresolved24
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c87c66bf-5d6a-4851-91b1-d73aff9c5ccf · outbound

This paper cites Representing Speech Through Autoregressive Prediction of Cochlear Tokens.

Representing Speech Through Autoregressive Prediction of Cochlear Tokens Representing Speech Through Autoregressive Prediction of Cochlear Tokens

Reference 1

Resolution
metadata mismatch
local_arxiv, observed 2026-08-05T19:54:59.463829Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T19:54:58.525812Z digest=sha256:13ab2df04b857eb15d38575f100f4e6e5a23cca02a972ecec33b91d63eee790e

Observation e70517cc-c426-43e4-b519-7ec9eba9316d · outbound

This paper cites water” and “river.

Representing Speech Through Autoregressive Prediction of Cochlear Tokens water” and “river

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:55:03.623324Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T19:54:58.530279Z digest=sha256:8f63621053ccc3af3847ea8be7148c9639f80c3998368fd185d673cff5fba0dd

Observation 1dbc760f-0472-4349-bb5c-34037b0f7723 · outbound

This paper cites er” was often confused with “r.

Representing Speech Through Autoregressive Prediction of Cochlear Tokens er” was often confused with “r

Reference 3

Resolution
malformed identifier
raw_fallback, observed 2026-08-05T19:55:03.613327Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T19:54:58.534959Z digest=sha256:466e506262942cc879b0c50280fa0e77f01528aae32019e0adc54966b5e1052a

Observation 8418ae8c-cc49-47e6-8870-8f8d089400a8 · outbound

This paper cites Transformation Imitation.

Representing Speech Through Autoregressive Prediction of Cochlear Tokens Transformation Imitation

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:55:03.603935Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T19:54:58.538930Z digest=sha256:bee0d81e88d3f874d8530d3af0bcb67c609b40ea47993fa8831a5e310280069c

Observation 87b8727c-cc16-4faa-a61c-702bb755647c · outbound

This paper cites acknowledges support from The K.

Representing Speech Through Autoregressive Prediction of Cochlear Tokens acknowledges support from The K

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:55:03.594277Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T19:54:58.542617Z digest=sha256:d481094d2b5af805ffd675890f92fce4fa94feb0698de2088048af576ccb4488

Observation d35525ca-2a30-43d0-9329-ae571f573311 · outbound

This paper cites SUPERB: Speech processing Universal PERformance Benchmark.

Representing Speech Through Autoregressive Prediction of Cochlear Tokens SUPERB: Speech processing Universal PERformance Benchmark

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-05T19:54:58.546041Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:54:58.546041Z digest=sha256:1c55f33149e21ea563c656a51f73d4e487462332206490c8caa718441813988e

Observation 7257bd88-eb8c-483a-84c1-1faded21bf78 · outbound

This paper cites Codec-SUPERB @ SLT 2024: A lightweight benchmark for neural audio codec models.

Representing Speech Through Autoregressive Prediction of Cochlear Tokens Codec-SUPERB @ SLT 2024: A lightweight benchmark for neural audio codec models

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-08-05T19:54:59.341956Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T19:54:58.550290Z digest=sha256:71182080d87aaf9175a1dd396bed623c2428daa4ea6a76a0150173a0d1b91d55

Observation ba6f1da7-8ec5-4af4-a722-1c40c19dfc80 · outbound

This paper cites A Review of Deep Learning Techniques for Speech Processing.

Representing Speech Through Autoregressive Prediction of Cochlear Tokens A Review of Deep Learning Techniques for Speech Processing

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-05T19:54:58.553897Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:54:58.553897Z digest=sha256:844e710eeaf58aa618d46fcdcfb456bfd44e0c1953afc177519f25ea174719c2

Observation 86bff55e-0e93-4447-aa6c-b3938d7f5878 · outbound

This paper cites Soundstream: An end-to-end neural audio codec,.

Representing Speech Through Autoregressive Prediction of Cochlear Tokens Soundstream: An end-to-end neural audio codec,

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-05T19:54:58.558742Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:54:58.558742Z digest=sha256:2570edd82c94ceba1166e2687db554d2822a520af51750619be5a685441116e2

Observation 587e6140-14f7-4998-a625-0c5aa2904316 · outbound

This paper cites High Fidelity Neural Audio Compression.

Representing Speech Through Autoregressive Prediction of Cochlear Tokens High Fidelity Neural Audio Compression

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-05T19:54:58.561899Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:54:58.561899Z digest=sha256:7ab1d3761c6bf61db3d55f456ed2d870cf80ac384c4c2b5398907ad52e42d373

Observation f225f353-cfd7-4e63-b6bc-70213e002bb5 · outbound

This paper cites HiFi-Codec: Group-residual Vector quantization for High Fidelity Audio Codec.

Representing Speech Through Autoregressive Prediction of Cochlear Tokens HiFi-Codec: Group-residual Vector quantization for High Fidelity Audio Codec

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-05T19:54:58.565677Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:54:58.565677Z digest=sha256:9c5fdd0ecbde68e70ee1a3899819a48a9b0d5263adf5f68944bd68d72917173d

Observation 4075519f-d176-4b8a-bce3-882d355bb389 · outbound

This paper cites High-fidelity audio compression with improved rvqgan,.

Representing Speech Through Autoregressive Prediction of Cochlear Tokens High-fidelity audio compression with improved rvqgan,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:55:03.577698Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T19:54:58.569242Z digest=sha256:e13f3b8941552e2bc7e7ce3ba40400c0317831003cb5c3bd4bab69b952dbac46

Observation b43b3d0a-1d76-4908-9c67-cf1874151514 · outbound

This paper cites Language-Codec: Bridging Discrete Codec Representations and Speech Language Models.

Representing Speech Through Autoregressive Prediction of Cochlear Tokens Language-Codec: Bridging Discrete Codec Representations and Speech Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-05T19:54:58.572361Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:54:58.572361Z digest=sha256:3e5ddf6256977962317c9a62cd65bb790a3d74d41ba3eeafb9e735812e40231b

Observation 3f67f68b-b935-425c-86b6-f24317671d3e · outbound

This paper cites Speechtok- enizer: Unified speech tokenizer for speech language models,.

Representing Speech Through Autoregressive Prediction of Cochlear Tokens Speechtok- enizer: Unified speech tokenizer for speech language models,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:55:03.566990Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T19:54:58.575845Z digest=sha256:ab56784ecb1bc476766739427329949c0d9e9ee529357b15d5698e2c393c3aa0

Observation 61224e10-6883-4e89-a553-dd944a8b051c · outbound

This paper cites CLaM-TTS: Improving Neural Codec Language Model for Zero-Shot Text-to-Speech.

Representing Speech Through Autoregressive Prediction of Cochlear Tokens CLaM-TTS: Improving Neural Codec Language Model for Zero-Shot Text-to-Speech

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-05T19:54:58.579073Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:54:58.579073Z digest=sha256:3fc8bc5e120f9cb8c922778ff1a87173671c7324f9c9c882f289c16d9285b117

Observation 1886f8f4-88fc-4818-9a5b-9a573165cfbb · outbound

This paper cites Spectral Codecs: Improving Non-Autoregressive Speech Synthesis with Spectrogram-Based Audio Codecs.

Representing Speech Through Autoregressive Prediction of Cochlear Tokens Spectral Codecs: Improving Non-Autoregressive Speech Synthesis with Spectrogram-Based Audio Codecs

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-05T19:54:58.582582Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:54:58.582582Z digest=sha256:f0bb1116a21adcdfb1048b1206f2f6abb7bb001361e9c2804b7adaf9ec78bed1

Observation 7f87b3e5-0f62-484c-a2bf-6ee9e158b49b · outbound

This paper cites Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers.

Representing Speech Through Autoregressive Prediction of Cochlear Tokens Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-05T19:54:58.586186Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:54:58.586186Z digest=sha256:84cdd961182846e521683b29569ed1f6c1e8d781c59c0aa839106a2ed52be96b

Observation e7744f64-4ab0-4588-a3cc-55690244c93e · outbound

This paper cites Hierarchical organization of hu- man auditory cortex: evidence from acoustic invariance in the re- sponse to intelligible speech,.

Representing Speech Through Autoregressive Prediction of Cochlear Tokens Hierarchical organization of hu- man auditory cortex: evidence from acoustic invariance in the re- sponse to intelligible speech,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:55:03.462945Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T19:54:58.589508Z digest=sha256:fb003798037df4f44438aed4cd0dadcd8f8d296938df6798e126a337c8cfe56a

Observation 5d4cc089-2728-4852-b9a9-0427f40f18a8 · outbound

This paper cites Intonational speech prosody encoding in the human auditory cortex,.

Representing Speech Through Autoregressive Prediction of Cochlear Tokens Intonational speech prosody encoding in the human auditory cortex,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:55:03.286871Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T19:54:58.592847Z digest=sha256:067fd77f9c8301a7b90b2b835d3ea2f27e10fbc2482aff2dd9d7827881a84c48

Observation 8464c6ed-1479-49e3-b937-6b7d393d173a · outbound

This paper cites Hubert: Self-supervised speech rep- resentation learning by masked prediction of hidden units,.

Representing Speech Through Autoregressive Prediction of Cochlear Tokens Hubert: Self-supervised speech rep- resentation learning by masked prediction of hidden units,

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-05T19:54:58.595978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:54:58.595978Z digest=sha256:fa89201b913a20ab6b2d95cfdd1d95fecd4cefb8c70d276577c0ff6264b4f737

Observation 9516fb69-478a-4798-bfa4-95e9d6cda6bf · outbound

This paper cites W2v-BERT: Combining Contrastive Learning and Masked Language Modeling for Self-Supervised Speech Pre-Training.

Representing Speech Through Autoregressive Prediction of Cochlear Tokens W2v-BERT: Combining Contrastive Learning and Masked Language Modeling for Self-Supervised Speech Pre-Training

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-05T19:54:58.599394Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:54:58.599394Z digest=sha256:85b85b91e4ec98d937ca973c7956a655a52cfd097dee3a40e86e01a28dd3576a

Observation 43f76c7d-ebb6-4dc9-9aa6-e9fbefa6023f · outbound

This paper cites Wavlm: Large-scale self- supervised pre-training for full stack speech processing,.

Representing Speech Through Autoregressive Prediction of Cochlear Tokens Wavlm: Large-scale self- supervised pre-training for full stack speech processing,

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-05T19:54:58.602756Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:54:58.602756Z digest=sha256:ee495031155a05d1c7f951b653fdcdff9718d273fb18f24c5e1cf4e268e42fad

Observation ed71dc35-6842-4fef-9545-5fdea44e45e5 · outbound

This paper cites Blind phoneme segmentation with temporal prediction errors.

Representing Speech Through Autoregressive Prediction of Cochlear Tokens Blind phoneme segmentation with temporal prediction errors

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-08-05T19:54:59.172535Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T19:54:58.605940Z digest=sha256:8cb1ee2ac68e0c6c0a7f5b10bc84fde14779dea0b67d9a4f188092c1ba684f3f

Observation b177be0e-b3b8-4219-b791-7006268441cc · outbound

This paper cites An Unsupervised Autoregressive Model for Speech Representation Learning.

Representing Speech Through Autoregressive Prediction of Cochlear Tokens An Unsupervised Autoregressive Model for Speech Representation Learning

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-05T19:54:58.609844Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:54:58.609844Z digest=sha256:d4d88d1678b2e93f0a0b99ce9af2fcc135743a097a802459801e34daa8998aef

Observation 73d39a96-2ea6-41e1-881d-0a986ab7417b · outbound

This paper cites Acquiring language from speech by learning to remember and predict,.

Representing Speech Through Autoregressive Prediction of Cochlear Tokens Acquiring language from speech by learning to remember and predict,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:55:03.113543Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T19:54:58.613387Z digest=sha256:5a1303afcf78e4715770c786f06727f79772587a02c9b0e12d09125b75a7fb3b

Observation 04458adf-6e74-495f-ae62-26106e634722 · outbound

This paper cites Vector-Quantized Autoregressive Predictive Coding.

Representing Speech Through Autoregressive Prediction of Cochlear Tokens Vector-Quantized Autoregressive Predictive Coding

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-05T19:54:58.616637Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:54:58.616637Z digest=sha256:9d28b7f75a06519a32e5040044b7eaab5a9244689b35d26d0690bb446df48d65

Observation 4125ea9c-5623-4cdf-97da-333987beef75 · outbound

This paper cites Audio albert: A lite bert for self-supervised learning of audio representation,.

Representing Speech Through Autoregressive Prediction of Cochlear Tokens Audio albert: A lite bert for self-supervised learning of audio representation,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:55:02.985774Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T19:54:58.621129Z digest=sha256:4ded708b0b96c0d806a5e7704b050dcb49e1c8bf8ed2c4b9e9f3fbfe97838e26

Observation 6774d475-e580-4647-9742-715858ad99a6 · outbound

This paper cites On generative spoken language modeling from raw au- dio,.

Representing Speech Through Autoregressive Prediction of Cochlear Tokens On generative spoken language modeling from raw au- dio,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:55:02.859177Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T19:54:58.624665Z digest=sha256:985d10044015cf99a51a5359d3a5ffe6a0aca05fcac566e85ee2081336e90a61

Observation c35d794d-6ba2-43a0-a662-565838525c7c · outbound

This paper cites Audiolm: a language modeling approach to audio gener- ation,.

Representing Speech Through Autoregressive Prediction of Cochlear Tokens Audiolm: a language modeling approach to audio gener- ation,

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-05T19:54:58.628627Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:54:58.628627Z digest=sha256:e71ce58a0506d440aa3321bbff0d889f527b37262b9cfce1728f5e01f53c2896

Observation 4a55b06c-d015-495f-a306-39f267076dbf · outbound

This paper cites Improving Textless Spoken Language Understanding with Discrete Units as Intermediate Target.

Representing Speech Through Autoregressive Prediction of Cochlear Tokens Improving Textless Spoken Language Understanding with Discrete Units as Intermediate Target

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-08-05T19:54:59.113274Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T19:54:58.632070Z digest=sha256:f755ff65ec707d113cc83c75e770c748bcde9783d608ab20496acb638b36b9ea

Observation 4eb29282-33de-4d8d-b2ac-40aa568d4df9 · outbound

This paper cites BERT: Pre- training of Deep Bidirectional Transformers for Language Under- standing,.

Representing Speech Through Autoregressive Prediction of Cochlear Tokens BERT: Pre- training of Deep Bidirectional Transformers for Language Under- standing,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:55:02.733522Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T19:54:58.635519Z digest=sha256:12ed7e6c9daab269b8ed1d2fb6909a2cd96ec923c40415138a9535d3d4628fa7

Observation 54f7547d-408e-47fd-b58f-3c2612a5e04e · outbound

This paper cites Representation Learning with Contrastive Predictive Coding.

Representing Speech Through Autoregressive Prediction of Cochlear Tokens Representation Learning with Contrastive Predictive Coding

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-05T19:54:58.639389Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:54:58.639389Z digest=sha256:80580c6d114e61f8ef71356b1b3d53dfd4afb2c3491cca1d1e0dc2793811608b

Observation 05670626-1df8-47d1-a877-7bccc75d583d · outbound

This paper cites Wav2vec 2.0: a framework for self-supervised learning of speech representa- tions,.

Representing Speech Through Autoregressive Prediction of Cochlear Tokens Wav2vec 2.0: a framework for self-supervised learning of speech representa- tions,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:55:02.595018Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T19:54:58.642906Z digest=sha256:263de24d52564bcdd421121b4eb3dd63c274f8fd856750810dc61bbb6363ef2c

Observation d0ea69db-a483-47df-a0f9-960dc3ab132b · outbound

This paper cites Contrastive learning of general-purpose audio representations,.

Representing Speech Through Autoregressive Prediction of Cochlear Tokens Contrastive learning of general-purpose audio representations,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:55:02.473345Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T19:54:58.647061Z digest=sha256:ee321d2a8e2aa6d19f00d8524c180a5ff4fe24e886b312804685ab701a70d0b5

Observation c036ce17-2c66-4ac3-a5f0-bec0c25cf112 · outbound

This paper cites Unsupervised speech segmentation and variable rate repre- sentation learning using segmental contrastive predictive coding,.

Representing Speech Through Autoregressive Prediction of Cochlear Tokens Unsupervised speech segmentation and variable rate repre- sentation learning using segmental contrastive predictive coding,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:55:02.333859Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T19:54:58.650700Z digest=sha256:4b3809af18d5ddaf6514b27c34f376ee87f55f133f23dc331450ad3663556207

Observation d8f38926-fe18-4bef-b40f-325195474c9d · outbound

This paper cites Self-normalization and noise- robustness in early auditory representations,.

Representing Speech Through Autoregressive Prediction of Cochlear Tokens Self-normalization and noise- robustness in early auditory representations,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:55:02.231236Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T19:54:58.653998Z digest=sha256:c7989ca23a1d35bbfccaddc0807b9211f734958235b260f79430a4b9a440a42b

Observation a5178f83-656d-412d-a88b-be14c90d6f35 · outbound

This paper cites an unresolved cited work.

Representing Speech Through Autoregressive Prediction of Cochlear Tokens Unresolved cited work

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-05T19:54:58.657328Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:54:58.657328Z digest=sha256:689a903aeefa51e48961f01d5ac972e5bea493a009038a0a9b70835275fab7ae

Observation 764197d2-608a-47ea-a155-1fae0d4a3019 · outbound

This paper cites Multiresolution spectrotem- poral analysis of complex sounds,.

Representing Speech Through Autoregressive Prediction of Cochlear Tokens Multiresolution spectrotem- poral analysis of complex sounds,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:55:02.100079Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T19:54:58.660417Z digest=sha256:c2c7590a895eeabec61d21f0d349731756e61422a8f6dfff84015e6900c47c5e

Observation 7dfd53ae-2f9a-4016-9a9a-a5fa14fbd665 · outbound

This paper cites Model metamers reveal divergent invariances between biological and ar- tificial neural networks,.

Representing Speech Through Autoregressive Prediction of Cochlear Tokens Model metamers reveal divergent invariances between biological and ar- tificial neural networks,

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:55:01.976683Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T19:54:58.663600Z digest=sha256:e864c197564c61e19d3610de9f3f400d81b6a9cbe2f00caba2c976549b2a45e8

Observation 6a985c03-b404-43b8-903e-93fc23763d3c · outbound

This paper cites Derivation of auditory filter shapes from notched-noise data,.

Representing Speech Through Autoregressive Prediction of Cochlear Tokens Derivation of auditory filter shapes from notched-noise data,

Reference 40

Resolution
verified exact
raw_fallback, observed 2026-08-05T19:54:58.998024Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T19:54:58.666884Z digest=sha256:a9607eabd87ed9da429e8937f4dd515dff7187e8b0f399d600e2ab16b5de3338

Observation 71303311-fc97-4f31-aa77-67021373af4d · outbound

This paper cites Sound texture perception via statistics of the auditory periphery: evidence from sound syn- thesis,.

Representing Speech Through Autoregressive Prediction of Cochlear Tokens Sound texture perception via statistics of the auditory periphery: evidence from sound syn- thesis,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:55:01.847895Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T19:54:58.670026Z digest=sha256:5dd9cad21af67a2fc754c901de5a4f81e37a7579e756d1252d8d257fbfafe2f6

Observation c0098db9-3dc5-4551-b07b-e52f6010aaf4 · outbound

This paper cites Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation.

Representing Speech Through Autoregressive Prediction of Cochlear Tokens Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-05T19:54:58.673271Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:54:58.673271Z digest=sha256:3acb2f1e3caa6eac147abba9f396c3ee77ba9f699d3e302b9714f22b838f492e

Observation 2b59ce30-34e4-413c-8261-07bcdeafbaef · outbound

This paper cites Lib- rispeech: an asr corpus based on public domain audio books,.

Representing Speech Through Autoregressive Prediction of Cochlear Tokens Lib- rispeech: an asr corpus based on public domain audio books,

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-05T19:54:58.677108Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:54:58.677108Z digest=sha256:b79ddf061cf11fc98dd40124370c410fc9313224d5eb44b6d4eb8c421e5e4a7b

Observation 653e79ba-084b-461e-be72-dbbf8515a5e7 · outbound

This paper cites Improving language understanding by generative pre-training,.

Representing Speech Through Autoregressive Prediction of Cochlear Tokens Improving language understanding by generative pre-training,

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:55:01.721880Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T19:54:58.680937Z digest=sha256:da42263783b1746e3b5c677a08e193d4820d7d204644e95d0c82a856045579fa

Observation f5858c51-8ec3-42f6-899c-28e3d76aa871 · outbound

This paper cites Gaussian Error Linear Units (GELUs).

Representing Speech Through Autoregressive Prediction of Cochlear Tokens Gaussian Error Linear Units (GELUs)

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-05T19:54:58.684181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:54:58.684181Z digest=sha256:75d0d49ff0f6972bde8e50b683ba5da176266b681650a14ad51cb71379e04f82

Observation 1e70758c-5b26-475c-ab84-35265a244a88 · outbound

This paper cites Root Mean Square Layer Normalization.

Representing Speech Through Autoregressive Prediction of Cochlear Tokens Root Mean Square Layer Normalization

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-05T19:54:58.688240Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:54:58.688240Z digest=sha256:c9de211211df20aecc3fe7b77f86fb96eae2dab3134c27224a3f12a7ae375954

Observation 3d165a1c-5b8b-4e06-a373-0f6b2f87e3a8 · outbound

This paper cites Libri-light: A benchmark for asr with limited or no supervision,.

Representing Speech Through Autoregressive Prediction of Cochlear Tokens Libri-light: A benchmark for asr with limited or no supervision,

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:55:01.611021Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T19:54:58.691750Z digest=sha256:c35e941eb603b117e9cbd465300c23db0c8d5cb9f4a7d4db893fa105c38cb234

Observation e2c0c792-4af0-45a9-a891-87aa239c8340 · outbound

This paper cites Timit acoustic phonetic continuous speech cor- pus,.

Representing Speech Through Autoregressive Prediction of Cochlear Tokens Timit acoustic phonetic continuous speech cor- pus,

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:55:01.458253Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T19:54:58.695205Z digest=sha256:a052de1f41f22287c1f63c3bc49e3d181ece0fda4a78e953f57ccb9c87e46a53

Observation 98d1da4f-be71-4a68-852d-3bc94bbb395d · outbound

This paper cites Speaker-independent phone recogni- tion using hidden markov models,.

Representing Speech Through Autoregressive Prediction of Cochlear Tokens Speaker-independent phone recogni- tion using hidden markov models,

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:55:01.279439Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T19:54:58.699082Z digest=sha256:98d504922e17540dcd25ae34fcd2c984fa7dfd97d0b77cadfe181af45daf6c7a

Observation eb7042ec-1ac7-468a-b583-3db817e346b2 · outbound

This paper cites Scikit-learn: Machine Learning in Python,.

Representing Speech Through Autoregressive Prediction of Cochlear Tokens Scikit-learn: Machine Learning in Python,

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:55:01.140934Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T19:54:58.703299Z digest=sha256:33838650137de8ce5e355d230373896fd5637b6b72f8511df80f182a4d474c15

Observation a2d0db29-082a-4ab4-8bf5-da843b7777f7 · outbound

This paper cites The Zero Resource Speech Benchmark 2021: Metrics and baselines for unsupervised spoken language modeling.

Representing Speech Through Autoregressive Prediction of Cochlear Tokens The Zero Resource Speech Benchmark 2021: Metrics and baselines for unsupervised spoken language modeling

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-05T19:54:58.706814Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:54:58.706814Z digest=sha256:feb333a7b0badd660c7148f3cffd30fa8d8c0c4f9001bc1151b475409052f8d3

Observation 38b04a89-0a8f-4ecd-b6ce-a817b2096755 · outbound

This paper cites Over-reliance on english hinders cognitive science,.

Representing Speech Through Autoregressive Prediction of Cochlear Tokens Over-reliance on english hinders cognitive science,

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:55:01.002185Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T19:54:58.710255Z digest=sha256:987d55dceac5b324d9a3ae07da9f618f93304a895008cbc343da4abd12e2a4a5

Observation 099f2fd0-5e9a-45d3-8cef-f1ff9189d8b0 · outbound

This paper cites To- wards inclusive automatic speech recognition,.

Representing Speech Through Autoregressive Prediction of Cochlear Tokens To- wards inclusive automatic speech recognition,

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:55:00.891669Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T19:54:58.713597Z digest=sha256:6578c8f89c497b2d435010e7538f9d6cd4e8f5c8bb3a78c0bd44f0cc1ca63ccb

Observation 3aaeb462-9d71-4e5c-ba65-82b91e492f2a · outbound

This paper cites Say- cam: A large, longitudinal audiovisual dataset recorded from the infant’s perspective,.

Representing Speech Through Autoregressive Prediction of Cochlear Tokens Say- cam: A large, longitudinal audiovisual dataset recorded from the infant’s perspective,

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:55:00.778519Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T19:54:58.716758Z digest=sha256:79b37f2025771447b0cc87a0a9cf542658ee90248373a18f6ede3b6ebcd1d5bb

Observation 071b6908-1925-49bf-b58e-757b99b7b746 · outbound

This paper cites Call for Papers -- The BabyLM Challenge: Sample-efficient pretraining on a developmentally plausible corpus.

Representing Speech Through Autoregressive Prediction of Cochlear Tokens Call for Papers -- The BabyLM Challenge: Sample-efficient pretraining on a developmentally plausible corpus

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-05T19:54:58.720381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:54:58.720381Z digest=sha256:0a426e8a91028ce408e9ad9e3d62f27df1f982b283fb3a475466813b2a97cc8c

Observation 19b671da-5252-4497-9d4b-a0caffead11c · outbound

This paper cites A task-optimized neural network repli- cates human auditory behavior, predicts brain responses, and re- veals a cortical processing hierarchy,.

Representing Speech Through Autoregressive Prediction of Cochlear Tokens A task-optimized neural network repli- cates human auditory behavior, predicts brain responses, and re- veals a cortical processing hierarchy,

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:55:00.644771Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T19:54:58.723735Z digest=sha256:9d45319c18b3408820e21377c3223f6e36e0b4d6d3d3224409cdf135b6b2f6ca

Observation d7f67bdb-df30-48c7-bf43-d124bd131507 · outbound

This paper cites Toward a realistic model of speech processing in the brain with self-supervised learning,.

Representing Speech Through Autoregressive Prediction of Cochlear Tokens Toward a realistic model of speech processing in the brain with self-supervised learning,

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:55:00.557865Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T19:54:58.727007Z digest=sha256:a3bfc5b11061a6d432bdac3320945dd32d4263052f3226f3294a60c2c2e9cc10

Observation 02a3ce6b-c9d5-4b31-b65d-c42660187c23 · outbound

This paper cites Dissecting neural computations of the human auditory pathway using deep neural networks for speech,.

Representing Speech Through Autoregressive Prediction of Cochlear Tokens Dissecting neural computations of the human auditory pathway using deep neural networks for speech,

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:55:00.429799Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T19:54:58.730190Z digest=sha256:29c8246293f33f9f6520f22564fea0bbfeeabbde2d49c832f41c120190eae1e5

Observation aee78f11-cfe0-443e-8813-2da3653d8bce · outbound

This paper cites Many but not all deep neural network audio models capture brain responses and exhibit correspondence between model stages and brain regions,.

Representing Speech Through Autoregressive Prediction of Cochlear Tokens Many but not all deep neural network audio models capture brain responses and exhibit correspondence between model stages and brain regions,

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:55:00.297487Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T19:54:58.733488Z digest=sha256:d884af3dc3b325df81cb574ed93e802eda00d16fe516687b24f5a58bd18395e8

Observation 91f4b4e3-66e6-4218-ac8c-0418f1508ccb · outbound

This paper cites Speech taskonomy: Which speech tasks are the most predictive of fmri brain activity?.

Representing Speech Through Autoregressive Prediction of Cochlear Tokens Speech taskonomy: Which speech tasks are the most predictive of fmri brain activity?

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:55:00.192596Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T19:54:58.736777Z digest=sha256:3da797010cae2de0f7318f63504afed6409aa66e67e7eecb9cac70a2bf37a2c3

Observation 3aff2a49-132d-4df5-803b-27c0a1a5079a · outbound

This paper cites Language in brains, minds, and machines,.

Representing Speech Through Autoregressive Prediction of Cochlear Tokens Language in brains, minds, and machines,

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:55:00.058048Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T19:54:58.741001Z digest=sha256:f0cb849ee31e7687ea7c9776f6e52ff8d9c557331cc618977c25bfaed4a4d2b7

Observation 0264a1d7-45ec-4d1d-a2b8-29e6e9f2387a · outbound

This paper cites Brain-tuned Speech Models Better Reflect Speech Processing Stages in the Brain.

Representing Speech Through Autoregressive Prediction of Cochlear Tokens Brain-tuned Speech Models Better Reflect Speech Processing Stages in the Brain

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-05T19:54:58.744160Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:54:58.744160Z digest=sha256:763319343ff2c6f3c0452693b73ae11ce2b450bca21832869663fe8bfece98cc

Observation c3849522-5e98-42e5-b262-13da0475a8e1 · outbound

This paper cites An algorithm for the machine calculation of complex fourier series,.

Representing Speech Through Autoregressive Prediction of Cochlear Tokens An algorithm for the machine calculation of complex fourier series,

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:54:59.936362Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T19:54:58.747668Z digest=sha256:6308532e51071af15745aee42964e7ba2aeb089c1d1e69ab227e88678fa61058

Observation 13352801-f597-4a72-8ac0-fc4301aedecb · outbound

This paper cites Lib- rispeech: An ASR corpus based on public domain audio books,.

Representing Speech Through Autoregressive Prediction of Cochlear Tokens Lib- rispeech: An ASR corpus based on public domain audio books,

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:54:59.764253Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T19:54:58.751677Z digest=sha256:b86e2a64417929e68b7165fe6f6ad31868ec7f0c5830231dd3f774aff395c073

Observation ec17cb52-6a2a-46f9-a014-6dc96e011ce3 · outbound

This paper cites codebook usage.

Representing Speech Through Autoregressive Prediction of Cochlear Tokens codebook usage

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T19:54:59.574844Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T19:54:58.755219Z digest=sha256:b1cd22a26bffaa04e2328cc00ce382b431b65343288077e02ce40ec4fdb69fb9

Pith citing papers

Observation c87c66bf-5d6a-4851-91b1-d73aff9c5ccf · inbound

Representing Speech Through Autoregressive Prediction of Cochlear Tokens cites this paper.

Representing Speech Through Autoregressive Prediction of Cochlear Tokens Representing Speech Through Autoregressive Prediction of Cochlear Tokens

Reference 1

Resolution
metadata mismatch
local_arxiv, observed 2026-08-05T19:54:59.463829Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-05T19:54:58.525812Z digest=sha256:13ab2df04b857eb15d38575f100f4e6e5a23cca02a972ecec33b91d63eee790e