Pith. sign in

Paper Citation Record · LEDGER

PAST: Phonetic-Acoustic Speech Tokenizer

As of 10 August 2026, this Paper Citation Record lists 48 of 48 outbound references and 2 inbound Pith citation observations for arXiv:2505.14470.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.14470 v2

Coverage vector

measured 48 of 48 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:36:46.248344Z

measured 50 of 50 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:36:41.638006Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T12:48:12.260808Z

Reference resolution

48 of 48 outbound references displayed

  • verified exact0
  • verified fuzzy22
  • unresolved26
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 6d1d0618-b4ee-44a9-8e73-da79dbbe5c50 · outbound

This paper cites These models usually operate over acoustic tokens or phonetic speech tokens (also known as semantic tokens).

PAST: Phonetic-Acoustic Speech Tokenizer These models usually operate over acoustic tokens or phonetic speech tokens (also known as semantic tokens)

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:52.440433Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:36:41.113789Z digest=sha256:c65554f28c9d2398eae0aece9aa3febdae0842888205b444e4ea79cdde3cdec5

Observation 24be95c4-b49a-460b-b670-fa9d1c2ba676 · outbound

This paper cites an unresolved cited work.

PAST: Phonetic-Acoustic Speech Tokenizer Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:36:52.213664Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:36:41.232581Z digest=sha256:92627636846779c715c82bd66e429813bcdf2a8ddc2a3a3162344a2512b2b603

Observation 8e225647-b7d9-4597-9a81-c995d783fc67 · outbound

This paper cites an unresolved cited work.

PAST: Phonetic-Acoustic Speech Tokenizer Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:36:51.919616Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:36:41.364786Z digest=sha256:978733cc49ed361938937f8b10e99b66177065764cf75a47ee77889819add82c

Observation 2073e341-ac4e-420d-a8c5-69758a47bf76 · outbound

This paper cites an unresolved cited work.

PAST: Phonetic-Acoustic Speech Tokenizer Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:36:51.608579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:36:41.471072Z digest=sha256:dab4939e9b8087342fb4aeaea5ed16d4fb506ede6b6e6f0ee2d7c231286f4a4d

Observation 3bd6c800-6965-417d-a7ae-ccea8cbdb256 · outbound

This paper cites an unresolved cited work.

PAST: Phonetic-Acoustic Speech Tokenizer Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:36:51.301505Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:36:41.550800Z digest=sha256:8cf38d162015c24ed37177373c3a47ee70ea3fef5a658a211dc3947d5031b28e

Observation fd002efa-0613-4dbf-80fe-94327ad589b8 · outbound

This paper cites PAST: Phonetic-Acoustic Speech Tokenizer.

PAST: Phonetic-Acoustic Speech Tokenizer PAST: Phonetic-Acoustic Speech Tokenizer

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:41.638006Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:41.638006Z digest=sha256:99cc89b85649a50ee82425ca5f2405c87142820f713c644e170b929ca3df7d51

Observation 5a9ed943-e7c3-4768-ad3c-9a6b1bb4955c · outbound

This paper cites Problem Setup Our model is composed of three main components: Encoder, Quantizer, and Decoder.

PAST: Phonetic-Acoustic Speech Tokenizer Problem Setup Our model is composed of three main components: Encoder, Quantizer, and Decoder

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:51.009788Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:36:41.730614Z digest=sha256:d14381c0e9d61d2973bf2c2e80ce85ba3b4987d52a9482235af57c42dafdbfbe

Observation 51300dc2-8046-4718-b918-da187b26b91a · outbound

This paper cites Soundstream: An end-to-end neural audio codec,.

PAST: Phonetic-Acoustic Speech Tokenizer Soundstream: An end-to-end neural audio codec,

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:43.064970Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:43.064970Z digest=sha256:ee6f8e65b09b2d1cb4f133eeaf6f8faad7dc10a058ab2e3bc7e35c3427cae27a

Observation 92fd0231-a65c-4b49-8483-1dddd9a223bd · outbound

This paper cites Data We use all training subsets of LibriSpeech [29] and TIMIT [30] for our training set, yielding a total of965hours of raw audio.

PAST: Phonetic-Acoustic Speech Tokenizer Data We use all training subsets of LibriSpeech [29] and TIMIT [30] for our training set, yielding a total of965hours of raw audio

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:50.579406Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:36:41.861572Z digest=sha256:30f90c5649755d09ee2fea473b44469bc277cc7611b4f99388d739fa0ae73897

Observation a4d1ca86-875f-44c2-b182-6aeb909450aa · outbound

This paper cites Baseline Comparison We compare PAST with two baseline hybrid models, Speech- Tokenizer and X-Codec, on both reconstruction and phonetic information metrics.

PAST: Phonetic-Acoustic Speech Tokenizer Baseline Comparison We compare PAST with two baseline hybrid models, Speech- Tokenizer and X-Codec, on both reconstruction and phonetic information metrics

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:50.396232Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:36:41.954031Z digest=sha256:1b785f4ab0f7368174cfb0d87c7c30723b585623fa8a4d5fa6f758e2f0527836

Observation 6e8e1277-a058-4f74-b933-60c5cf533c7b · outbound

This paper cites an unresolved cited work.

PAST: Phonetic-Acoustic Speech Tokenizer Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:36:50.215266Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:36:42.029656Z digest=sha256:326b5392a68a921f4d424a598f06f7242ccd69a734d851121a8d930b94b605bb

Observation c69020ab-9cb6-4e28-ae47-4ad8073d2623 · outbound

This paper cites Generative spoken language model based on continuous word-sized audio tokens,.

PAST: Phonetic-Acoustic Speech Tokenizer Generative spoken language model based on continuous word-sized audio tokens,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:50.067667Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:36:42.157713Z digest=sha256:88417b37dc90b3f091fa1db6a683c112db560f3737c5fef79922e17376ac1226

Observation 91225a23-a38d-4679-a578-fe8950aea761 · outbound

This paper cites Textless acoustic model with self-supervised distillation for noise-robust expressive speech-to-speech transla- tion,.

PAST: Phonetic-Acoustic Speech Tokenizer Textless acoustic model with self-supervised distillation for noise-robust expressive speech-to-speech transla- tion,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:49.907843Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:36:42.293050Z digest=sha256:57a8e84218fa52c1a928ab4f6963cabe8f78a521d3c0fd4fbd7f22021b50c180

Observation 2b67d8cf-cf05-4e28-820a-7ef18474dfc2 · outbound

This paper cites Audiolm: a language modeling approach to au- dio generation,.

PAST: Phonetic-Acoustic Speech Tokenizer Audiolm: a language modeling approach to au- dio generation,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:49.760241Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:36:42.467128Z digest=sha256:98e56eb791c61f3b434d5cf3fc9374c3bef6c673efd1cd58ff3a35f38c1e1a75

Observation c14e6a54-8a6b-4e38-ac5d-5e1881a85c0b · outbound

This paper cites Text-free prosody-aware generative spoken language modeling,.

PAST: Phonetic-Acoustic Speech Tokenizer Text-free prosody-aware generative spoken language modeling,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:49.588890Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:36:42.590165Z digest=sha256:d481145dbe9d5a506296fde2ae643ec3ed36986095b4561f7c6792e5306ec78a

Observation 9ae9b94b-4759-4cd1-bf42-0a0da3dece75 · outbound

This paper cites On generative spoken language modeling from raw audio,.

PAST: Phonetic-Acoustic Speech Tokenizer On generative spoken language modeling from raw audio,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:49.409764Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:36:42.723486Z digest=sha256:716ab8337e011378e94e1b54d95790bd3793e3d11ed7d04fa3296f41d0eb46a7

Observation c37f05e4-a93e-4275-bdd8-d56991b118f6 · outbound

This paper cites Textually pretrained speech language mod- els,.

PAST: Phonetic-Acoustic Speech Tokenizer Textually pretrained speech language mod- els,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:49.263010Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:36:42.863861Z digest=sha256:2ee0d621c085138f8537a5ff4cf3cae9719af317d228f18ff6c51036a285e7e7

Observation 093fc44b-671b-450c-b501-926022514b33 · outbound

This paper cites High Fidelity Neural Audio Compression.

PAST: Phonetic-Acoustic Speech Tokenizer High Fidelity Neural Audio Compression

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:42.972992Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:42.972992Z digest=sha256:63deaba2f98a5c4f3c8ffaec34ee23b4ac531b178903e7911570d86028f96bbd

Observation 9ad809f5-5aec-45ed-9b29-cc3bc56237f8 · outbound

This paper cites Pyramidcodec: Hierarchical codec for long-form music generation in audio domain,.

PAST: Phonetic-Acoustic Speech Tokenizer Pyramidcodec: Hierarchical codec for long-form music generation in audio domain,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:48.210900Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:36:44.394744Z digest=sha256:c82053be0da36a067241540b3f5b6d47ce655890e94e1158241f24ba7f1bbefd

Observation 41a88ea8-e634-44c9-a273-a8d22177ed8b · outbound

This paper cites wav2vec 2.0: A framework for self-supervised learning of speech repre- sentations,.

PAST: Phonetic-Acoustic Speech Tokenizer wav2vec 2.0: A framework for self-supervised learning of speech repre- sentations,

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:43.217290Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:43.217290Z digest=sha256:6e798312fe3008340ceb177efd5b0369f6cfa0e51ca54388608552f7af1f66c8

Observation c6552382-832e-4c72-bdb3-c74f93108585 · outbound

This paper cites Hubert: Self-supervised speech represen- tation learning by masked prediction of hidden units,.

PAST: Phonetic-Acoustic Speech Tokenizer Hubert: Self-supervised speech represen- tation learning by masked prediction of hidden units,

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:43.335300Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:43.335300Z digest=sha256:104b581ab994a2f5534e50989ce28dc55de28f09ba4f9bad8212d20581af2563

Observation f4541b15-f6fc-48f2-a527-dd324dc16ca4 · outbound

This paper cites Analysing discrete self supervised speech representation for spoken language modeling,.

PAST: Phonetic-Acoustic Speech Tokenizer Analysing discrete self supervised speech representation for spoken language modeling,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:49.132860Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:36:43.431637Z digest=sha256:42114f72632ba1d574e33ff523c3593c0d1bf1c74428461492a49a0da0ed95d8

Observation fce9e7a6-b01d-4b23-badc-691f7bf72840 · outbound

This paper cites Hifi-gan: Generative adversarial net- works for efficient and high fidelity speech synthesis,.

PAST: Phonetic-Acoustic Speech Tokenizer Hifi-gan: Generative adversarial net- works for efficient and high fidelity speech synthesis,

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:43.497780Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:43.497780Z digest=sha256:98e53db3494ffc568a2684f2ad27e4e0207562bf9387e8b41b9c9f0ee8f536c8

Observation 5ea204e1-eb42-41a0-a676-61cebaea42d3 · outbound

This paper cites Speechtok- enizer: Unified speech tokenizer for speech language models,.

PAST: Phonetic-Acoustic Speech Tokenizer Speechtok- enizer: Unified speech tokenizer for speech language models,

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:43.576586Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:43.576586Z digest=sha256:5e2604e9ec846f978ea8f08d7bb74efd047b335323c5c719a1033837528a0585

Observation 65286334-7429-4af7-a48b-ad9146096970 · outbound

This paper cites Codec Does Matter: Exploring the Semantic Shortcoming of Codec for Audio Language Model.

PAST: Phonetic-Acoustic Speech Tokenizer Codec Does Matter: Exploring the Semantic Shortcoming of Codec for Audio Language Model

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:43.700479Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:43.700479Z digest=sha256:bdb41490f63846e85cb44892b8e8a317c0e105df6065a3444cf3951cb61ba9cd

Observation bbf55102-cca8-4a26-b575-518bf7a68b9d · outbound

This paper cites Moshi: a speech-text foundation model for real-time dialogue,.

PAST: Phonetic-Acoustic Speech Tokenizer Moshi: a speech-text foundation model for real-time dialogue,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:48.972558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:36:43.818074Z digest=sha256:be1c5d05b3485c45b681614f0bc68e019cf571278894609545cc99ff55866f4f

Observation 356517eb-c842-450f-85a4-0c0e666dcdab · outbound

This paper cites Wavlm: Large-scale self-supervised pre-training for full stack speech processing,.

PAST: Phonetic-Acoustic Speech Tokenizer Wavlm: Large-scale self-supervised pre-training for full stack speech processing,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:48.715687Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:36:43.938820Z digest=sha256:c41ee42112c9beefb6806cc8bfb4d0805fa9e9aeb3f6a6629a1d38024be6726a

Observation 467667de-8c01-4e10-ac1f-b910263e95df · outbound

This paper cites an unresolved cited work.

PAST: Phonetic-Acoustic Speech Tokenizer Unresolved cited work

Reference 28

Resolution
unresolved
raw_fallback, observed 2026-08-07T15:36:50.782575Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:36:41.796360Z digest=sha256:b544a3288a8a629b0d6590b793ff21483e200c2bde27f6eb948c79645034d408

Observation ae2a569d-ba4b-42a9-a956-bcf254094187 · outbound

This paper cites HiFi-Codec: Group-residual Vector quantization for High Fidelity Audio Codec.

PAST: Phonetic-Acoustic Speech Tokenizer HiFi-Codec: Group-residual Vector quantization for High Fidelity Audio Codec

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:44.119792Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:44.119792Z digest=sha256:e4aab43ff2b3855d8e4ded538dc2fe5b79350e647a1751de6863f63bfdb4c38b

Observation a7c73625-ae79-45d1-9199-36b2bc83944a · outbound

This paper cites High fidelity neural audio compression,.

PAST: Phonetic-Acoustic Speech Tokenizer High fidelity neural audio compression,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:48.445553Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:36:44.241497Z digest=sha256:b3d086bf202597cf012fa5d207ee936c3c6079a7c6e2a204cf31fe3e2f9dceec

Observation 722f6378-a58a-4deb-81f4-32f4f5f10418 · outbound

This paper cites Audiodec: An open-source streaming high- fidelity neural audio codec,.

PAST: Phonetic-Acoustic Speech Tokenizer Audiodec: An open-source streaming high- fidelity neural audio codec,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:47.971760Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:36:44.527302Z digest=sha256:8f58e6dde2d55106d97631bcdd0f7c7b7ec01f9a613731c0bcb67f027c82c531

Observation 3945ee03-0863-472a-9fb8-29c097240400 · outbound

This paper cites Funcodec: A funda- mental, reproducible and integrable open-source toolkit for neural speech codec,.

PAST: Phonetic-Acoustic Speech Tokenizer Funcodec: A funda- mental, reproducible and integrable open-source toolkit for neural speech codec,

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:44.656266Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:44.656266Z digest=sha256:261cea34f102597b0e0882ec4fdc54b2d6325d865dad7710320e7400e9b4aa95

Observation 5e8865d9-12f2-4ef8-872f-9f91f8d12791 · outbound

This paper cites Wavtokenizer: an efficient acoustic discrete codec tokenizer for audio language modeling,.

PAST: Phonetic-Acoustic Speech Tokenizer Wavtokenizer: an efficient acoustic discrete codec tokenizer for audio language modeling,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:47.689975Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:36:44.766303Z digest=sha256:ab986dca0c819ac79ea0501b26cfa903652c80a347c0cd0ab87dbb41197329db

Observation 35126764-2f60-4e9d-8e58-84d1b7e60eeb · outbound

This paper cites Scaling Speech-Text Pre-training with Synthetic Interleaved Data.

PAST: Phonetic-Acoustic Speech Tokenizer Scaling Speech-Text Pre-training with Synthetic Interleaved Data

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:44.886879Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:44.886879Z digest=sha256:e9ea692ed0217240f2d50ddfb37851209fa6174d7f14f31bbf131504579b1538

Observation 5a62e195-0f55-469f-a96c-04642d3a81fe · outbound

This paper cites Robust speech recognition via large-scale weak supervision,.

PAST: Phonetic-Acoustic Speech Tokenizer Robust speech recognition via large-scale weak supervision,

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:44.992730Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:44.992730Z digest=sha256:e246415ea6aa25558267231d7f78566c0902db7990f546aceb30377d2df0e160

Observation 9b62e96f-ebfd-4928-8e66-23090e9c5428 · outbound

This paper cites LAST: Language Model Aware Speech Tokenization.

PAST: Phonetic-Acoustic Speech Tokenizer LAST: Language Model Aware Speech Tokenization

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:45.112795Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:45.112795Z digest=sha256:1d8069ef43e6bb54e74b2009cee89baa08871052be36ca11214a67375c4ae444

Observation 2f90e1a3-4ff1-4f91-9f46-27076370343f · outbound

This paper cites NAST: Noise Aware Speech Tokenization for Speech Language Models.

PAST: Phonetic-Acoustic Speech Tokenizer NAST: Noise Aware Speech Tokenization for Speech Language Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:45.240767Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:45.240767Z digest=sha256:75b31b051119f5ef0707de9fb4279062fca7cc89d37ff4d4de0fdd41df1e7333

Observation 9620259b-327c-4042-8f3c-652536fdf4b1 · outbound

This paper cites A systematic compar- ison of phonetic aware techniques for speech enhancement,.

PAST: Phonetic-Acoustic Speech Tokenizer A systematic compar- ison of phonetic aware techniques for speech enhancement,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:47.409472Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:36:45.347084Z digest=sha256:9c8c59bef99be71d732978c594b5131696dd442b5b657736c28ef419a413e7e2

Observation cc04afe2-3599-4af3-8073-c2bfb689ab90 · outbound

This paper cites Con- nectionist temporal classification: labelling unsegmented se- quence data with recurrent neural networks,.

PAST: Phonetic-Acoustic Speech Tokenizer Con- nectionist temporal classification: labelling unsegmented se- quence data with recurrent neural networks,

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:45.445023Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:45.445023Z digest=sha256:799a6dd4a381f675cebbc70defdf5a40d453ced5a92c69033329d73fec106341

Observation ef3f6159-93f0-405c-8250-5d7ade4b97f2 · outbound

This paper cites Lib- rispeech: An asr corpus based on public domain audio books,.

PAST: Phonetic-Acoustic Speech Tokenizer Lib- rispeech: An asr corpus based on public domain audio books,

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:45.562599Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:45.562599Z digest=sha256:134e69aca2e20e10d994cc5b87faa98415a4b176fa6c85494e0eb6cf60946598

Observation 839a24ed-933d-4929-8277-66023614f503 · outbound

This paper cites Timit acoustic-phonetic continuous speech corpus,.

PAST: Phonetic-Acoustic Speech Tokenizer Timit acoustic-phonetic continuous speech corpus,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:47.162559Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:36:45.656172Z digest=sha256:712764a1248ee64c850eb2bb10481096774383c6867446fe9fd109561db45dcf

Observation 154029ea-b6bf-4d81-9854-c7ae7695950e · outbound

This paper cites Visqol v3: An open source production ready objec- tive speech and audio metric,.

PAST: Phonetic-Acoustic Speech Tokenizer Visqol v3: An open source production ready objec- tive speech and audio metric,

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:45.725481Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:45.725481Z digest=sha256:dbf74e040669c5483325487cbed7b5de9e36809ab896d7cc3b6b08883471590c

Observation bb65a8c7-94e0-44ae-9425-442ed6a4016d · outbound

This paper cites Perceptual eval- uation of speech quality (pesq)-a new method for speech quality assessment of telephone networks and codecs,.

PAST: Phonetic-Acoustic Speech Tokenizer Perceptual eval- uation of speech quality (pesq)-a new method for speech quality assessment of telephone networks and codecs,

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:45.788652Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:45.788652Z digest=sha256:7f5e1c43df4edac3b16a34b1d632dc22d0a30dd55bc394385fed3999c629c8f3

Observation 01192191-8782-4dca-8efc-a5b90cce4364 · outbound

This paper cites Evaluating speech features with the minimal- pair abx task: analysis of the classical mfc/plp pipeline,.

PAST: Phonetic-Acoustic Speech Tokenizer Evaluating speech features with the minimal- pair abx task: analysis of the classical mfc/plp pipeline,

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:46.884652Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:36:45.874985Z digest=sha256:1dde0844ce2fafd4c4266d0b5d2eabc7dcf698f7d59531ebc246ef70e6f812ea

Observation 1be950b0-83c3-403c-b17f-6ace2183e66e · outbound

This paper cites DASB - Discrete Audio and Speech Benchmark.

PAST: Phonetic-Acoustic Speech Tokenizer DASB - Discrete Audio and Speech Benchmark

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:45.931571Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:45.931571Z digest=sha256:ff93d808359cfefd6453bb4eef8c4e04e75f7c43cacffafae2c7656c12c8cc6d

Observation 7a95b9ff-de8a-497d-8d58-def212839c4e · outbound

This paper cites AudioGen: Textually Guided Audio Generation.

PAST: Phonetic-Acoustic Speech Tokenizer AudioGen: Textually Guided Audio Generation

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:46.052253Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:46.052253Z digest=sha256:c25c770fd9b65168769e4d3e1a996bb35f3d4f24114f29650700a25f9e78f1e9

Observation 1909cf34-ad72-475b-b8b3-427cf0d212b2 · outbound

This paper cites Simple and controllable music gen- eration,.

PAST: Phonetic-Acoustic Speech Tokenizer Simple and controllable music gen- eration,

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:46.626989Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:36:46.148093Z digest=sha256:b1213665d99adbddff654a4bbf767ac1eaad9cfc7adb20f5cbd5ee7464fbafa1

Observation f5e42fd6-873a-4829-94b9-d8e497fbe07f · outbound

This paper cites The zero resource speech benchmark 2021: Metrics and baselines for unsupervised spoken language model- ing,.

PAST: Phonetic-Acoustic Speech Tokenizer The zero resource speech benchmark 2021: Metrics and baselines for unsupervised spoken language model- ing,

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:36:46.529710Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T15:36:46.248344Z digest=sha256:771f1fc43df5d2029a2e589822c9b2b99fc3c663cf0fb1468e54d818e2ccc4d4

Pith citing papers

Observation fd002efa-0613-4dbf-80fe-94327ad589b8 · inbound

PAST: Phonetic-Acoustic Speech Tokenizer cites this paper.

PAST: Phonetic-Acoustic Speech Tokenizer PAST: Phonetic-Acoustic Speech Tokenizer

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T15:36:41.638006Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:36:41.638006Z digest=sha256:99cc89b85649a50ee82425ca5f2405c87142820f713c644e170b929ca3df7d51

Observation af391476-d310-4819-81f2-fea65bdf6141 · inbound

Benchmarking Neural Speech Compression from a Rate-Distortion Perspective cites this paper.

Benchmarking Neural Speech Compression from a Rate-Distortion Perspective PAST: Phonetic-Acoustic Speech Tokenizer

Reference 77

Resolution
verified exact
arxiv_id, observed 2026-07-03T12:48:12.262338Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-27T08:43:35.279033Z digest=sha256:0190e9b72503bb5b49f1b162e767b5c05fef2b52df411c3f1c93576480ed399e