Pith. sign in

Paper Citation Record · LEDGER

Content is What Remains: Invariant Speech Tokenization from Parallel Utterances

As of 16 August 2026, this Paper Citation Record lists 52 of 52 outbound references and 2 inbound Pith citation observations for arXiv:2607.19033.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.19033 v1

Coverage vector

measured 52 of 52 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T15:37:36.656380Z

measured 54 of 54 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T15:37:36.420218Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-12T00:15:23.475461Z

Reference resolution

52 of 52 outbound references displayed

  • verified exact3
  • verified fuzzy27
  • unresolved21
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation b2af3fbb-4cc0-42f9-ace0-8327fc1aeac3 · outbound

This paper cites Content is What Remains: Invariant Speech Tokenization from Parallel Utterances.

Content is What Remains: Invariant Speech Tokenization from Parallel Utterances Content is What Remains: Invariant Speech Tokenization from Parallel Utterances

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-15T15:37:36.420218Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:37:36.420218Z digest=sha256:877cd3ca45e75bfc52031f0f613dff2875d2e1ffbec1d16067039dfa28b52d9e

Observation 36122b67-fcf0-4c05-b518-26339874cf68 · outbound

This paper cites an unresolved cited work.

Content is What Remains: Invariant Speech Tokenization from Parallel Utterances Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-15T15:37:37.773073Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T15:37:36.426600Z digest=sha256:8c739ce486ce7fd57e2bc19d81984da2fda4f27adba246cab14b96f6ecb6794b

Observation 314f73bb-bf00-4cba-84c2-caa1eccad9cc · outbound

This paper cites Further we evaluate the downstream compres- sion benefit unlocked by the above.

Content is What Remains: Invariant Speech Tokenization from Parallel Utterances Further we evaluate the downstream compres- sion benefit unlocked by the above

Reference 3

Resolution
malformed identifier
raw_fallback, observed 2026-08-15T15:37:37.757774Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T15:37:36.431802Z digest=sha256:6fef89a7df6d283093920fee73b3b095c652ca69f2748b785133b004e62b1ce7

Observation e619767f-ffcb-42b7-92b3-ccf4b6312d00 · outbound

This paper cites an unresolved cited work.

Content is What Remains: Invariant Speech Tokenization from Parallel Utterances Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-15T15:37:37.742427Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T15:37:36.436685Z digest=sha256:5b157b7d11c576bdb9716596d62b6c7b22c80f64eaa57139584e96ea6c17fde9

Observation f6d5b190-beab-4257-8444-22f4875b90b3 · outbound

This paper cites The experiments were run manually and results were manually verified.

Content is What Remains: Invariant Speech Tokenization from Parallel Utterances The experiments were run manually and results were manually verified

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:37:37.726711Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T15:37:36.441464Z digest=sha256:3d3b77650efb530f2d81e22a3d92103807c632450f59fd62dd082d9b7440db51

Observation de499ba7-9b34-418c-9973-674a7badd1cc · outbound

This paper cites AudioLM: a language modeling approach to audio generation,.

Content is What Remains: Invariant Speech Tokenization from Parallel Utterances AudioLM: a language modeling approach to audio generation,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:37:37.710583Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T15:37:36.446270Z digest=sha256:ef323d97215cc7d7a2f9d8cc236c8aa2fcf9cc7a7dde2f3946a0e1c8cd29065e

Observation fc188f6a-54ef-425d-9b17-b5e03b654f03 · outbound

This paper cites SpeechGPT: Empowering large language models with intrinsic cross-modal conversational abilities,.

Content is What Remains: Invariant Speech Tokenization from Parallel Utterances SpeechGPT: Empowering large language models with intrinsic cross-modal conversational abilities,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:37:37.696744Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T15:37:36.450990Z digest=sha256:71675c35ea7cf7eebfb85960fc9c5cac83fe0af78a0b493c67483735e84c78e2

Observation 33894dc6-3747-4fd1-9516-1c6b297ca72e · outbound

This paper cites Moshi: a speech-text foundation model for real-time dialogue.

Content is What Remains: Invariant Speech Tokenization from Parallel Utterances Moshi: a speech-text foundation model for real-time dialogue

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T15:37:36.456476Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:37:36.456476Z digest=sha256:1df8fe081a5f9a7c3a359118e6b0a0a1d55eaa0f92da7fd2061d18ec1ee8808b

Observation 63a33f02-57d5-4c38-ba68-5d94e182fff3 · outbound

This paper cites VioLA: Unified Codec Language Models for Speech Recognition, Synthesis, and Translation.

Content is What Remains: Invariant Speech Tokenization from Parallel Utterances VioLA: Unified Codec Language Models for Speech Recognition, Synthesis, and Translation

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T15:37:36.461335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:37:36.461335Z digest=sha256:e04adfcab548ed111aba11afa1bb977e756f2b0ef706b16fd303a9827d6fe869

Observation 8ee4ab36-a958-40c1-a6f9-d0f6bd0ea221 · outbound

This paper cites SpeechTok- enizer: Unified speech tokenizer for speech language models,.

Content is What Remains: Invariant Speech Tokenization from Parallel Utterances SpeechTok- enizer: Unified speech tokenizer for speech language models,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:37:37.681063Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T15:37:36.466266Z digest=sha256:04ec51f599de0ed91eaa5d6f48d61f863f2ce40fc241d2f1e61dfc652545f1d3

Observation ef9bd97c-7591-4b86-a813-f90d20517ba1 · outbound

This paper cites AudioDec: An open-source streaming high-fidelity neural audio codec,.

Content is What Remains: Invariant Speech Tokenization from Parallel Utterances AudioDec: An open-source streaming high-fidelity neural audio codec,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:37:37.664857Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T15:37:36.470649Z digest=sha256:be97d280fdd9e8e09b395592b3b4465e22a2dda076dd8d124fe6c0abe40528dc

Observation 1cbd49f7-3425-4cf6-9c55-a96ff4db9109 · outbound

This paper cites Hubert: Self-supervised speech rep- resentation learning by masked prediction of hidden units,.

Content is What Remains: Invariant Speech Tokenization from Parallel Utterances Hubert: Self-supervised speech rep- resentation learning by masked prediction of hidden units,

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-15T15:37:36.474582Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:37:36.474582Z digest=sha256:641e33114a2a3173428a9c11365656f297f68233238393c1f2cf6d6ff6b17340

Observation 6e300fcc-ab13-44cf-9cad-b4d7168681a6 · outbound

This paper cites Wavlm: Large-scale self-supervised pre-training for full stack speech processing,.

Content is What Remains: Invariant Speech Tokenization from Parallel Utterances Wavlm: Large-scale self-supervised pre-training for full stack speech processing,

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-15T15:37:36.478614Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:37:36.478614Z digest=sha256:ccd7be3233ed55490b797aa4923ecb21341d12e4598627c1ba8155db901fc106

Observation 7799387e-c3bd-4bd1-9ae7-8f47588a0dcd · outbound

This paper cites ContentVec: An improved self-supervised speech representation by disentangling speakers,.

Content is What Remains: Invariant Speech Tokenization from Parallel Utterances ContentVec: An improved self-supervised speech representation by disentangling speakers,

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-15T15:37:36.486606Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:37:36.486606Z digest=sha256:4854367f3ff49a6e12c443974efc62d3dd19f0c10de933457b3eb0913988690d

Observation faba2e41-4dfc-4b35-a2dd-931154e42af9 · outbound

This paper cites Estimating the completeness of discrete speech units,.

Content is What Remains: Invariant Speech Tokenization from Parallel Utterances Estimating the completeness of discrete speech units,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:37:37.620815Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T15:37:36.490350Z digest=sha256:eb3946179fba3b73544b0c8781ab181d28c8085ef8c41b9fe311bbfe8bb89e04

Observation 1aff2f3f-d144-43b1-9ae7-26d7c0194291 · outbound

This paper cites Augmentation invariant discrete representation for generative spoken language modeling,.

Content is What Remains: Invariant Speech Tokenization from Parallel Utterances Augmentation invariant discrete representation for generative spoken language modeling,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:37:37.606133Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T15:37:36.494023Z digest=sha256:292d3b46eebc60eeb04ce8b081227ceca94fd585d4d56e1f3a5468a8ea289735

Observation 88d4a449-ca1f-4533-bce0-993379af8834 · outbound

This paper cites StableToken: A Noise-Robust Semantic Speech Tokenizer for Resilient SpeechLLMs.

Content is What Remains: Invariant Speech Tokenization from Parallel Utterances StableToken: A Noise-Robust Semantic Speech Tokenizer for Resilient SpeechLLMs

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-08-15T15:37:37.080968Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T15:37:36.531979Z digest=sha256:5f0c8dfcc85aaf7e0643c20fc0516a95bc76f9fc9977f1e2d97d835b84d8f3aa

Observation 788381a9-7117-4f0a-8075-db053f96d105 · outbound

This paper cites STAB: Speech Tokenizer Assessment Benchmark.

Content is What Remains: Invariant Speech Tokenization from Parallel Utterances STAB: Speech Tokenizer Assessment Benchmark

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-15T15:37:36.502980Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:37:36.502980Z digest=sha256:4ec121e2566741834c8fe59b1d171f81209c65efb5142f758ec8a3b4ebde8161

Observation 7ac573bf-c9f7-4578-9aff-11cb22ddf585 · outbound

This paper cites Dc- spin: A speaker-invariant speech tokenizer for spoken language models,.

Content is What Remains: Invariant Speech Tokenization from Parallel Utterances Dc- spin: A speaker-invariant speech tokenizer for spoken language models,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:37:37.577027Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T15:37:36.507781Z digest=sha256:ee0e92ea7cce3607bc2184569f7a515c9f8af20690dba771be923a076127293d

Observation bac9b4da-0f77-450a-a7af-a908838c7e47 · outbound

This paper cites Rethinking discrete speech representation tokens for accent generation,.

Content is What Remains: Invariant Speech Tokenization from Parallel Utterances Rethinking discrete speech representation tokens for accent generation,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:37:37.563488Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T15:37:36.512769Z digest=sha256:a3413e5378b00ca7c599d00b84506e13dc9b787ad9051c6a65b5ddb336b6748d

Observation a0a54f54-523e-406b-90c0-9dee637fb615 · outbound

This paper cites Robust Speech Recognition via Large-Scale Weak Supervision.

Content is What Remains: Invariant Speech Tokenization from Parallel Utterances Robust Speech Recognition via Large-Scale Weak Supervision

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-15T15:37:36.555809Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:37:36.555809Z digest=sha256:2be5b4fbba40e4e6c66589cc3d7d42bf2b45b10e72b53223a6d79a88a6c53324

Observation 86d3fbae-e589-4fea-8288-7939cd8d6701 · outbound

This paper cites Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement.

Content is What Remains: Invariant Speech Tokenization from Parallel Utterances Vevo: Controllable Zero-Shot Voice Imitation with Self-Supervised Disentanglement

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-15T15:37:36.522433Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:37:36.522433Z digest=sha256:525b7b77238b6ac773573e1f54e631a9fa1c9c57dbe2a9b87262e90360c5fd27

Observation 1f13b3ec-36bc-437f-a1e3-c157d76e754b · outbound

This paper cites NAST: Noise Aware Speech Tokenization for Speech Language Models.

Content is What Remains: Invariant Speech Tokenization from Parallel Utterances NAST: Noise Aware Speech Tokenization for Speech Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-15T15:37:36.527205Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:37:36.527205Z digest=sha256:5804a4e5eb22f540cf795a75ab4b82cd4d4c98440ce9d1d7035b974b7c1993d0

Observation 3c456030-231e-48df-b613-ed6a1f945cb7 · outbound

This paper cites The ryerson audio-visual database of emotional speech and song (ravdess): A dynamic, multimodal set of facial and vocal expressions in north american english,.

Content is What Remains: Invariant Speech Tokenization from Parallel Utterances The ryerson audio-visual database of emotional speech and song (ravdess): A dynamic, multimodal set of facial and vocal expressions in north american english,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:37:37.506664Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T15:37:36.573839Z digest=sha256:6b830fcf00f5b58fe4481e06bebadfc21807122c039d48ff659e7b27b0fb781b

Observation 82920521-3379-497c-9f60-eed4e20f0fba · outbound

This paper cites Dualcodec: A low-frame-rate, semantically-enhanced neural audio codec for speech generation,.

Content is What Remains: Invariant Speech Tokenization from Parallel Utterances Dualcodec: A low-frame-rate, semantically-enhanced neural audio codec for speech generation,

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-15T15:37:36.537592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:37:36.537592Z digest=sha256:9b086d655762a0c8e50bddc33a70a74c082e808aacd6beac89ae804568af744f

Observation ecb525e6-c774-4269-aad9-03ca55301624 · outbound

This paper cites XY-Tokenizer: Mitigating the Semantic-Acoustic Conflict in Low-Bitrate Speech Codecs.

Content is What Remains: Invariant Speech Tokenization from Parallel Utterances XY-Tokenizer: Mitigating the Semantic-Acoustic Conflict in Low-Bitrate Speech Codecs

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-15T15:37:36.541957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:37:36.541957Z digest=sha256:5a5bfaa2f7d53dd352d8cdcfb44edc6e0d98a93ce1063ddcd2cd063c09c0a93d

Observation 53c8d88c-fbf8-44e6-be24-86d970fdba9e · outbound

This paper cites Sac: Neural speech codec with semantic-acoustic dual-stream quantization,.

Content is What Remains: Invariant Speech Tokenization from Parallel Utterances Sac: Neural speech codec with semantic-acoustic dual-stream quantization,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:37:37.549691Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T15:37:36.546380Z digest=sha256:2353ee5004e6b5a132162112d49b6909a54a2830d23a5c8e9aa92f96d0ebec1e

Observation da88886d-b705-43c7-ab98-b0ecb3e04f3d · outbound

This paper cites The chains corpus: Characterizing individual speakers,.

Content is What Remains: Invariant Speech Tokenization from Parallel Utterances The chains corpus: Characterizing individual speakers,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:37:37.444811Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T15:37:36.591693Z digest=sha256:b4d670d27689f18bc5a0344d2f41358ddc743ec4f2516d85446d2525bc019de5

Observation 8fd06ec6-4823-49f3-a1ab-76f6db781d33 · outbound

This paper cites CSTR VCTK Corpus: English multi-speaker corpus for CSTR voice cloning toolkit (version 0.92),.

Content is What Remains: Invariant Speech Tokenization from Parallel Utterances CSTR VCTK Corpus: English multi-speaker corpus for CSTR voice cloning toolkit (version 0.92),

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-15T15:37:36.596071Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:37:36.596071Z digest=sha256:2474857608e6ee58251d405bdc938f681c20dcfed63d73418da74c57b5c81340

Observation 0513336b-4b7e-4198-a1d4-69a8f94e2445 · outbound

This paper cites W2v-bert: Combining contrastive learning and masked language modeling for self-supervised speech pre-training,.

Content is What Remains: Invariant Speech Tokenization from Parallel Utterances W2v-bert: Combining contrastive learning and masked language modeling for self-supervised speech pre-training,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:37:37.536085Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T15:37:36.560362Z digest=sha256:4ac025c5f170af8d7a24d7d17a739a5e9c0daee399d1a2b8c2694beb44294b2c

Observation 23478bcb-0e67-4a27-b311-d4e27796c0a7 · outbound

This paper cites Seen and unseen emo- tional style transfer for voice conversion with a new emotional speech dataset,.

Content is What Remains: Invariant Speech Tokenization from Parallel Utterances Seen and unseen emo- tional style transfer for voice conversion with a new emotional speech dataset,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:37:37.413903Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T15:37:36.604243Z digest=sha256:24843f15c7d42c30406fd28c0d7d048612ec9dbfcd8d4a9e49b9e27925d10925

Observation 36c02c41-c9f7-4d4e-91b8-9ff7888675b9 · outbound

This paper cites Kokoro-82m,.

Content is What Remains: Invariant Speech Tokenization from Parallel Utterances Kokoro-82m,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:37:37.521724Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T15:37:36.569307Z digest=sha256:cb85b5681ac14d165d5cba5666e256d01f6fe7553e223239833bfb4d9b3295f6

Observation 249fdd17-c16b-4c20-801b-f1fc4f7f76cc · outbound

This paper cites Darpa timit acoustic phonetic con- tinuous speech corpus cdrom,.

Content is What Remains: Invariant Speech Tokenization from Parallel Utterances Darpa timit acoustic phonetic con- tinuous speech corpus cdrom,

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-15T15:37:36.612873Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:37:36.612873Z digest=sha256:d2ebc424cf5e184382268c37ace9e82bf1e40316edbb5b53eb8b2d493c2b9c2d

Observation 202ec8d5-3725-495b-b14e-ff0d587c5921 · outbound

This paper cites Learning sound event classifiers from web audio with noisy labels,.

Content is What Remains: Invariant Speech Tokenization from Parallel Utterances Learning sound event classifiers from web audio with noisy labels,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:37:37.490573Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T15:37:36.578405Z digest=sha256:3dcfec7beec07e6932a67e84429867b2d4067f8a96d5366e4b09b7787058a018

Observation 5d8b0ac3-d261-4242-b2a6-0d84029a0fd8 · outbound

This paper cites Montreal forced aligner: Trainable text-speech align- ment using kaldi,.

Content is What Remains: Invariant Speech Tokenization from Parallel Utterances Montreal forced aligner: Trainable text-speech align- ment using kaldi,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:37:37.473348Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T15:37:36.582751Z digest=sha256:2a9528627fe1bdcc571f4fd07c938f3ffd33d69744c70c9b7fd55a4fa29029dd

Observation b4983d30-63a2-4b2c-9469-83fc06b98bbb · outbound

This paper cites The cmu arctic speech databases,.

Content is What Remains: Invariant Speech Tokenization from Parallel Utterances The cmu arctic speech databases,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:37:37.459318Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T15:37:36.587397Z digest=sha256:7f69305d12d49794acbabeca4bac40eb2516f7c770d3321feb54cede421de22c

Observation d5f5f318-0d7e-400b-a3b6-fd88c3fa8697 · outbound

This paper cites Con- nectionist temporal classification: Labelling unsegmented se- quence data with recurrent neural networks,.

Content is What Remains: Invariant Speech Tokenization from Parallel Utterances Con- nectionist temporal classification: Labelling unsegmented se- quence data with recurrent neural networks,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:37:37.362800Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T15:37:36.629857Z digest=sha256:406db313051f1664f611be05f9dfc9db8ebc06c485a243d509e0edfe0cb3ff4d

Observation 522fa691-bc23-42c5-a351-dc5f53aab19e · outbound

This paper cites Soft-dtw: a differentiable loss func- tion for time-series,.

Content is What Remains: Invariant Speech Tokenization from Parallel Utterances Soft-dtw: a differentiable loss func- tion for time-series,

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:37:37.345344Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T15:37:36.634292Z digest=sha256:3b0631c36fab09567b41528cb3c79094db2c733244bd90e1e2e2bf442ca8f325

Observation 5574bffc-00ab-4c3f-a271-54df18374705 · outbound

This paper cites Open-source multi-speaker corpora of the English accents in the British isles,.

Content is What Remains: Invariant Speech Tokenization from Parallel Utterances Open-source multi-speaker corpora of the English accents in the British isles,

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:37:37.429609Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T15:37:36.600245Z digest=sha256:95de6cd7e43090afe326e42235b182f32c65bd0f322014a4822ed47c77688fc8

Observation 82939289-67ef-4a9f-9fdf-fbf74550678d · outbound

This paper cites Mean teachers are better role models: Weight-averaged consistency targets improve semi- supervised deep learning results,.

Content is What Remains: Invariant Speech Tokenization from Parallel Utterances Mean teachers are better role models: Weight-averaged consistency targets improve semi- supervised deep learning results,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:37:37.312918Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T15:37:36.642925Z digest=sha256:dc920c3619db86cbb09064c6695602cd5ac91a76d35f25e2d03f93e91b79d835

Observation 329955bc-1630-41c4-b757-18b1d5e1a461 · outbound

This paper cites Learning Disentangled Speech Representations.

Content is What Remains: Invariant Speech Tokenization from Parallel Utterances Learning Disentangled Speech Representations

Reference 42

Resolution
verified exact
local_arxiv, observed 2026-08-15T15:37:36.892990Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T15:37:36.608319Z digest=sha256:e53ee2764fe2b70a2f42eb2466b529190a4c438de5f373b9e41af623741a2509

Observation 05f03d3a-5db6-4f46-992a-1b274dd7c4fd · outbound

This paper cites Evaluating speech features with the minimal-pair ABX task: Analysis of the classical MFC/PLP pipeline,.

Content is What Remains: Invariant Speech Tokenization from Parallel Utterances Evaluating speech features with the minimal-pair ABX task: Analysis of the classical MFC/PLP pipeline,

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:37:37.297591Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T15:37:36.651768Z digest=sha256:80ac42ded314351893af2d77c366a3e60ce1c680a5241fec4630bc737dc54db7

Observation d4590b98-b95a-4637-87c8-9ba486d5eae0 · outbound

This paper cites Lib- rispeech: an asr corpus based on public domain audio books,.

Content is What Remains: Invariant Speech Tokenization from Parallel Utterances Lib- rispeech: an asr corpus based on public domain audio books,

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-15T15:37:36.617067Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:37:36.617067Z digest=sha256:dc199be1c28837b1f58d34a7185a83ced9b0ab1cf9614bcac881ec7f36ed85f2

Observation 7e20d7c1-efd6-4333-8d97-b160364a448f · outbound

This paper cites Ted-lium 3: Twice as much data and corpus repartition for experiments on speaker adaptation,.

Content is What Remains: Invariant Speech Tokenization from Parallel Utterances Ted-lium 3: Twice as much data and corpus repartition for experiments on speaker adaptation,

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:37:37.378582Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T15:37:36.621152Z digest=sha256:3cc17f67722e7ed6082addad70a2df5e7c5b4ea7df119ad11b4c0cd6c97804b8

Observation eaf1e97c-3c83-40af-8879-2ae53a5176df · outbound

This paper cites Libri-light: A benchmark for asr with limited or no supervision,.

Content is What Remains: Invariant Speech Tokenization from Parallel Utterances Libri-light: A benchmark for asr with limited or no supervision,

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-15T15:37:36.625423Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:37:36.625423Z digest=sha256:07909a017012f8240f3f50a0094a11c5065eec5bc87e577307b591669487ef40

Observation 598fdfe8-c382-4126-aeac-c811a3627b47 · outbound

This paper cites Textless speech-to-speech translation on real data,.

Content is What Remains: Invariant Speech Tokenization from Parallel Utterances Textless speech-to-speech translation on real data,

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:37:37.329727Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T15:37:36.638543Z digest=sha256:fb6aad640fd2c303bf1fc5c1cd86e2758ee5672a81749aff7914ffea350e6d87

Observation 99d84a0e-d1f6-4749-aed6-5d6fce1da5a9 · outbound

This paper cites Efficient Self-supervised Learning with Contextualized Target Representations for Vision, Speech and Language.

Content is What Remains: Invariant Speech Tokenization from Parallel Utterances Efficient Self-supervised Learning with Contextualized Target Representations for Vision, Speech and Language

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-15T15:37:36.647187Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:37:36.647187Z digest=sha256:93d008fa04715a2d2b66cef8424c9761ed261c7734c61281ae4053c8ab4cfcd5

Observation c145d71b-3f7a-490b-9298-274e2256af8f · outbound

This paper cites X-vectors: Robust dnn embeddings for speaker recognition,.

Content is What Remains: Invariant Speech Tokenization from Parallel Utterances X-vectors: Robust dnn embeddings for speaker recognition,

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:37:37.281487Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T15:37:36.656380Z digest=sha256:96d32d0d3bd296e7f22e580aeab5e7d59f0dac0e436d021206535c806e985e59

Observation f0fcf3c6-851d-4378-8a71-0801f38f29b2 · outbound

This paper cites Available: https://aclanthology.org/2023.iwslt-1.

Content is What Remains: Invariant Speech Tokenization from Parallel Utterances Available: https://aclanthology.org/2023.iwslt-1

Reference 477

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T15:37:37.591389Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T15:37:36.498516Z digest=sha256:c8afbf9a7ef8607b82818a6e3f695d256e6a5e2e3eed2511032b696c76d6fbcd

Observation 623f462a-0fe4-4be0-9f4e-bf8f3c8b2a43 · outbound

This paper cites W2v-BERT: Combining Contrastive Learning and Masked Language Modeling for Self-Supervised Speech Pre-Training.

Content is What Remains: Invariant Speech Tokenization from Parallel Utterances W2v-BERT: Combining Contrastive Learning and Masked Language Modeling for Self-Supervised Speech Pre-Training

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-15T15:37:36.564661Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:37:36.564661Z digest=sha256:a7b0fe07c30af4b03406dd4d9195e25408a36a76e146e70ff0c103f029f7144a

Observation bd2e5f66-ea7e-40f2-b825-38d5ceec6c26 · outbound

This paper cites Available: https://arxiv.org/abs/2510.16841.

Content is What Remains: Invariant Speech Tokenization from Parallel Utterances Available: https://arxiv.org/abs/2510.16841

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-15T15:37:36.551536Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:37:36.551536Z digest=sha256:53c84ab50b1dd8f79f16637304ea2e0a37e7998ff5a6f10bc4be761b4ce8af22

Observation b7c1552a-d0ca-4218-83b3-dbcf5865334b · outbound

This paper cites Available: https://arxiv.org/abs/2601.19786.

Content is What Remains: Invariant Speech Tokenization from Parallel Utterances Available: https://arxiv.org/abs/2601.19786

Reference 2026

Resolution
verified exact
raw_fallback, observed 2026-08-15T15:37:37.184946Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T15:37:36.517410Z digest=sha256:5fa3867d7c9047d09404332fae4fe9aaf16f50c39efd57ec7c6117e57dcbc209

Pith citing papers

Observation b2af3fbb-4cc0-42f9-ace0-8327fc1aeac3 · inbound

Content is What Remains: Invariant Speech Tokenization from Parallel Utterances cites this paper.

Content is What Remains: Invariant Speech Tokenization from Parallel Utterances Content is What Remains: Invariant Speech Tokenization from Parallel Utterances

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-15T15:37:36.420218Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:37:36.420218Z digest=sha256:877cd3ca45e75bfc52031f0f613dff2875d2e1ffbec1d16067039dfa28b52d9e

Observation 8f600229-0b91-43ba-9765-7dbfcefa954f · inbound

ReLMCodec: Designing Predictable Speech Tokens from Pre-Quantization Phoneme Structure cites this paper.

ReLMCodec: Designing Predictable Speech Tokens from Pre-Quantization Phoneme Structure Content is What Remains: Invariant Speech Tokenization from Parallel Utterances

Reference 21

Resolution
metadata mismatch
local_arxiv, observed 2026-08-12T00:15:23.478725Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-12T00:15:23.039683Z digest=sha256:3ec8e4e8cdce435fef2ba58878beb6c15dd0e12e651fbb3e50a70d43199b0dad