Pith. sign in

Paper Citation Record · LEDGER

Exploring the Effect of Segmentation and Vocabulary Size on Speech Tokenization for Speech Language Models

As of 23 August 2026, this Paper Citation Record lists 38 of 38 outbound references and 1 inbound Pith citation observation for arXiv:2505.17446.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.17446 v2

Coverage vector

measured 38 of 38 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:50:02.022181Z

measured 39 of 39 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:49:57.859145Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-07T14:50:02.310312Z

Reference resolution

38 of 38 outbound references displayed

  • verified exact1
  • verified fuzzy30
  • unresolved7
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 8cf812e2-ced0-49ae-874a-a2004ad39b8e · outbound

This paper cites Exploring the Effect of Segmentation and Vocabulary Size on Speech Tokenization for Speech Language Models.

Exploring the Effect of Segmentation and Vocabulary Size on Speech Tokenization for Speech Language Models Exploring the Effect of Segmentation and Vocabulary Size on Speech Tokenization for Speech Language Models

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-08-07T14:50:02.409712Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T14:49:57.859145Z digest=sha256:b8a92b66363db5131b25f4938f3064236142e6a29cdb46e793a8718355307bfe

Observation 0c1fb68c-fbc5-45de-b901-2f473c3febeb · outbound

This paper cites Throughout this study, we used HuBERT [7] as an SSL model and extracted representations from the ninth layer.

Exploring the Effect of Segmentation and Vocabulary Size on Speech Tokenization for Speech Language Models Throughout this study, we used HuBERT [7] as an SSL model and extracted representations from the ninth layer

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:50:09.461214Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T14:49:57.887646Z digest=sha256:179b3f052f56a20f117cb6db4b7ae05cec1b57178789809f05fade96dfa746ff

Observation 139ffd34-fb6c-4203-bc97-72dd762a371d · outbound

This paper cites an unresolved cited work.

Exploring the Effect of Segmentation and Vocabulary Size on Speech Tokenization for Speech Language Models Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:50:09.330352Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T14:49:57.995598Z digest=sha256:d6e98ad88b9b8e4a9bade8d66baa14f552e1600f4f1fc829dd4da1805a30dae9

Observation fe7a207b-24d1-47c7-aa0f-b692ededf434 · outbound

This paper cites an unresolved cited work.

Exploring the Effect of Segmentation and Vocabulary Size on Speech Tokenization for Speech Language Models Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:50:09.205464Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T14:49:58.128986Z digest=sha256:5a614d4979bc6154a274b49c74d4806e97b1f0b9faabe7aac315678644286050

Observation 80c5be5a-d540-48db-a5fe-c97e88c5ee36 · outbound

This paper cites an unresolved cited work.

Exploring the Effect of Segmentation and Vocabulary Size on Speech Tokenization for Speech Language Models Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:50:09.076506Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T14:49:58.238918Z digest=sha256:c34e5e355698cbb982104490881976d3cd0b873d41b3a5b4f7a55cc871729191

Observation 6f133b16-38b6-4852-aea0-89f614884e69 · outbound

This paper cites Dataset As a training set for SLM, we used LibriSpeech [17], a 960-hour English audiobook corpus.

Exploring the Effect of Segmentation and Vocabulary Size on Speech Tokenization for Speech Language Models Dataset As a training set for SLM, we used LibriSpeech [17], a 960-hour English audiobook corpus

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:50:08.949260Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T14:49:58.300160Z digest=sha256:bad2744b76ebca93746cc224dd8fd5d58471a9090071d345642fcb8bd33e938b

Observation 0845f3d5-355d-4e8a-974b-cb9c7422d12a · outbound

This paper cites Figure 2 shows results on fixed boundary settings.

Exploring the Effect of Segmentation and Vocabulary Size on Speech Tokenization for Speech Language Models Figure 2 shows results on fixed boundary settings

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:50:08.846595Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T14:49:58.407640Z digest=sha256:f806f85ec2d94cb7d3c9f62043989c9d711dfc2da9fc7a0052fbbf3c289375e9

Observation c225dd39-0b87-4151-a99c-c7a4a2ac90da · outbound

This paper cites yonder".

Exploring the Effect of Segmentation and Vocabulary Size on Speech Tokenization for Speech Language Models yonder"

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:50:08.730832Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T14:49:58.481649Z digest=sha256:d235cd238985ee3e73be7cadc7acdb9c0369a9c2b694ba9a623f14b5acd1f14f

Observation 564c9789-a8e0-4a74-943a-2d439fbab684 · outbound

This paper cites We conducted mul- tiple speech tokenizations based on the combination of the fixed/variable segmentation and the cluster size.

Exploring the Effect of Segmentation and Vocabulary Size on Speech Tokenization for Speech Language Models We conducted mul- tiple speech tokenizations based on the combination of the fixed/variable segmentation and the cluster size

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:50:08.591058Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T14:49:58.570939Z digest=sha256:76af2fb569dcb584c527e06365077489129d6f485b367661c9dcf87ae52a17e7

Observation 3fd089f2-4175-4795-b3fc-07373eed3187 · outbound

This paper cites an unresolved cited work.

Exploring the Effect of Segmentation and Vocabulary Size on Speech Tokenization for Speech Language Models Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:50:08.423077Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T14:49:58.640221Z digest=sha256:7da25f8bab6eb904fb6ef8fe45c4f45c85652af7d1645e1133d2ea776c4c55e5

Observation 87f1e936-e923-42bf-a50c-0b65c1b67fcc · outbound

This paper cites On Generative Spoken Language Modeling from Raw Audio,.

Exploring the Effect of Segmentation and Vocabulary Size on Speech Tokenization for Speech Language Models On Generative Spoken Language Modeling from Raw Audio,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:50:08.193845Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T14:49:58.753979Z digest=sha256:1eab99ae242977e5a0ad7ab7efc25a3e25ec415a38922d2e662d28c5afbcf003

Observation 70af3af0-d481-4544-949c-607c1f9a2449 · outbound

This paper cites Textually Pretrained Speech Language Models,.

Exploring the Effect of Segmentation and Vocabulary Size on Speech Tokenization for Speech Language Models Textually Pretrained Speech Language Models,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:50:08.028899Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T14:49:58.883691Z digest=sha256:1bc0ed13253a5c1bbce834c95031ed7797b589db705f98fa99591b88159ca454

Observation 5a3393b4-1624-43bc-8954-f4f254b0e0ba · outbound

This paper cites Audiolm: A language modeling approach to audio generation,.

Exploring the Effect of Segmentation and Vocabulary Size on Speech Tokenization for Speech Language Models Audiolm: A language modeling approach to audio generation,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:50:07.884237Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T14:49:58.963630Z digest=sha256:4a312116da478eccaf48a2ac86a5f3993551e825e75b3754b36824919761c697

Observation 2a8b485e-ab94-43da-9934-9c0545480a8d · outbound

This paper cites WavLLM: Towards Robust and Adaptive Speech Large Language Model,.

Exploring the Effect of Segmentation and Vocabulary Size on Speech Tokenization for Speech Language Models WavLLM: Towards Robust and Adaptive Speech Large Language Model,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:50:07.713715Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T14:49:59.048288Z digest=sha256:7243dc4a17a18a840b75c669946da02b792887a15d8ebfab5b108da30f5572d4

Observation 3d90e8be-6ab2-4752-bf34-49cb88ce772e · outbound

This paper cites Representation Learning with Contrastive Predictive Coding,.

Exploring the Effect of Segmentation and Vocabulary Size on Speech Tokenization for Speech Language Models Representation Learning with Contrastive Predictive Coding,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:50:07.558127Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T14:49:59.162114Z digest=sha256:8f42921c639b3ca71e656dc239b3bf6fd54f9a8bff18932fce27a2265c7b146c

Observation 399a0d2d-8807-4c88-86db-a7dbe180c12f · outbound

This paper cites Wav2vec 2.0: A Framework for Self-Supervised Learning of Speech Represen- tations,.

Exploring the Effect of Segmentation and Vocabulary Size on Speech Tokenization for Speech Language Models Wav2vec 2.0: A Framework for Self-Supervised Learning of Speech Represen- tations,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:50:07.379926Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T14:49:59.229667Z digest=sha256:2b6cc048625b6c0dcc7927c0c5e5b7097dbb03e0793cf0b96afa84440b41d305

Observation a2b91708-4c6d-450d-be67-836a8f2da1a2 · outbound

This paper cites HuBERT: Self-Supervised Speech Rep- resentation Learning by Masked Prediction of Hidden Units,.

Exploring the Effect of Segmentation and Vocabulary Size on Speech Tokenization for Speech Language Models HuBERT: Self-Supervised Speech Rep- resentation Learning by Masked Prediction of Hidden Units,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:50:07.145029Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T14:49:59.368364Z digest=sha256:fd874a3328cf688449a7c253751bff9b437e48c244cd1b3ec86ec5c1b41444ae

Observation 5cb4db3e-8525-4a82-9f68-88280f53a06e · outbound

This paper cites The Zero Resource Speech Benchmark 2021: Metrics and baselines for unsupervised spoken language modeling,.

Exploring the Effect of Segmentation and Vocabulary Size on Speech Tokenization for Speech Language Models The Zero Resource Speech Benchmark 2021: Metrics and baselines for unsupervised spoken language modeling,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:50:06.928005Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T14:49:59.502483Z digest=sha256:aa2f5624f68f05c6a623216a911d19d7e466e85712d67a102099469f4dfd6f98

Observation 1bc6e752-38ab-46c8-b9ff-bbc5bdd885e5 · outbound

This paper cites Generative Spoken Dialogue Language Model- ing,.

Exploring the Effect of Segmentation and Vocabulary Size on Speech Tokenization for Speech Language Models Generative Spoken Dialogue Language Model- ing,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:50:06.699670Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T14:49:59.637090Z digest=sha256:9a034f91dd15320a10e602289f31250ddfd8a9acb288c640d013522e75015a23

Observation 0d1665fd-9f77-4ee3-8638-31813885e252 · outbound

This paper cites Direct Speech- to-Speech Translation With Discrete Units,.

Exploring the Effect of Segmentation and Vocabulary Size on Speech Tokenization for Speech Language Models Direct Speech- to-Speech Translation With Discrete Units,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:50:06.423491Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T14:49:59.792124Z digest=sha256:3c70753739a79a78656a509dd83edc947216646516800873b71550c2a9e4160f

Observation 1c98c5ae-2af5-432e-a24b-6fe9e72d7ae4 · outbound

This paper cites Text-free prosody-aware generative spoken language modeling,.

Exploring the Effect of Segmentation and Vocabulary Size on Speech Tokenization for Speech Language Models Text-free prosody-aware generative spoken language modeling,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:50:06.151214Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T14:49:59.930171Z digest=sha256:4be5291b0a237ced53b233276420d5559ed5cfafab476f894cba86f81e6e98c4

Observation 00f4e789-1c8c-404f-ace8-d8e74dbf0b7e · outbound

This paper cites Attention is all you need,.

Exploring the Effect of Segmentation and Vocabulary Size on Speech Tokenization for Speech Language Models Attention is all you need,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:50:05.922539Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T14:50:00.108485Z digest=sha256:a02a6d15fa5f1ad6f7a0c4b2f83ee3bd2e061d2b08f48d84b8a39919006bf209

Observation 2c6830ad-de9f-453a-9575-21f2fab6f352 · outbound

This paper cites Self-Supervised Speech Representations are More Phonetic than Semantic,.

Exploring the Effect of Segmentation and Vocabulary Size on Speech Tokenization for Speech Language Models Self-Supervised Speech Representations are More Phonetic than Semantic,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:50:05.697064Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T14:50:00.225350Z digest=sha256:647967b9da48dcd589e80dd3b415f1390f7a2259cff4d4e084134d9cfd5ecd61

Observation 3639a5ac-d4ab-4078-8b27-57f965f6b5a0 · outbound

This paper cites Generative Spoken Language Model based on continuous word-sized audio tokens,.

Exploring the Effect of Segmentation and Vocabulary Size on Speech Tokenization for Speech Language Models Generative Spoken Language Model based on continuous word-sized audio tokens,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:50:05.437143Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T14:50:00.329313Z digest=sha256:3bf1aeeb85103635111825de833057b65538d4cdd9b202791bb204430065676f

Observation 45e2ab50-0b6a-4966-8c70-02191b2dff11 · outbound

This paper cites SyllableLM: Learning Coarse Semantic Units for Speech Language Models.

Exploring the Effect of Segmentation and Vocabulary Size on Speech Tokenization for Speech Language Models SyllableLM: Learning Coarse Semantic Units for Speech Language Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T14:50:00.432842Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:50:00.432842Z digest=sha256:1aef101621647c60f1b55b9359212ee9a211445ce7103cc18e4bda59b41e97cb

Observation 4367e883-40a3-4b54-b078-4ae4544f7a68 · outbound

This paper cites Sylber: Syllabic Embedding Repre- sentation of Speech from Raw Audio,.

Exploring the Effect of Segmentation and Vocabulary Size on Speech Tokenization for Speech Language Models Sylber: Syllabic Embedding Repre- sentation of Speech from Raw Audio,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:50:05.138108Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T14:50:00.556666Z digest=sha256:11a856911084de410f13cb1d5c429e722bad9874332f4533795e8b36b7a8db9e

Observation 4635ecbd-ac80-4160-9225-3d93d1ffc5be · outbound

This paper cites Lib- rispeech: An ASR corpus based on public domain audio books,.

Exploring the Effect of Segmentation and Vocabulary Size on Speech Tokenization for Speech Language Models Lib- rispeech: An ASR corpus based on public domain audio books,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:50:04.894664Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T14:50:00.664857Z digest=sha256:1716ef7aeacde3febd6b3ff19893b7888222f94bc0fa6f68966cef3df500580e

Observation 3efbaa60-489f-409b-80d8-cfd65dd5e073 · outbound

This paper cites Libri-light: A benchmark for ASR with limited or no super- vision,.

Exploring the Effect of Segmentation and Vocabulary Size on Speech Tokenization for Speech Language Models Libri-light: A benchmark for ASR with limited or no super- vision,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:50:04.615415Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T14:50:00.802807Z digest=sha256:ec2d6c721d1f7ddd505354dc44bcc04107af7b1e2cc37f7947812caec730d63a

Observation 1e453aac-68e8-44d7-99a5-ce410e69dde7 · outbound

This paper cites OPT: Open Pre-trained Transformer Language Models.

Exploring the Effect of Segmentation and Vocabulary Size on Speech Tokenization for Speech Language Models OPT: Open Pre-trained Transformer Language Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T14:50:00.957373Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:50:00.957373Z digest=sha256:efe604b1b3d975f56d31f90a914a64c083c5d6dc7edbe556a07578b9b4cab396

Observation 7877ab0c-d921-437a-85b6-25eff4f21cbc · outbound

This paper cites ProsAudit, a prosodic benchmark for self-supervised speech models,.

Exploring the Effect of Segmentation and Vocabulary Size on Speech Tokenization for Speech Language Models ProsAudit, a prosodic benchmark for self-supervised speech models,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:50:04.320994Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T14:50:01.110052Z digest=sha256:a285c4ed077288b01bda1be8ad650c00c7207121c4cf20e3d5ef494594fe43aa

Observation 688f3b9f-cfa1-49ac-9e5f-3dba2828f275 · outbound

This paper cites A corpus and cloze evaluation for deeper understanding of commonsense stories,.

Exploring the Effect of Segmentation and Vocabulary Size on Speech Tokenization for Speech Language Models A corpus and cloze evaluation for deeper understanding of commonsense stories,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:50:04.110815Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T14:50:01.237648Z digest=sha256:c844924a57f891d9814e85aaeeb3afbb6d2d90ae7f96b38e89d15143cde804f8

Observation 5021f133-9058-4f65-ae6b-c35cee78d217 · outbound

This paper cites Praat: doing phonetics by com- puter [computer program]. version 6.4.27,.

Exploring the Effect of Segmentation and Vocabulary Size on Speech Tokenization for Speech Language Models Praat: doing phonetics by com- puter [computer program]. version 6.4.27,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:50:03.857528Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T14:50:01.363667Z digest=sha256:8fe92a6cec353ddbd14268786fb1a169ffb3145f21f0d699afebf03c4452eb05

Observation fc7b4e91-a9ec-48aa-a83e-67d0b3690784 · outbound

This paper cites Martinet, Elements of General Linguistics, ser.

Exploring the Effect of Segmentation and Vocabulary Size on Speech Tokenization for Speech Language Models Martinet, Elements of General Linguistics, ser

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:50:03.592334Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T14:50:01.513967Z digest=sha256:ddee085c6074870c7a179a91b53a2551c60774109b2038fdb1e2e2ca500101be

Observation 3805a6d2-15f9-4b42-9db8-6599ed20e1da · outbound

This paper cites Are Discrete Units Nec- essary for Spoken Language Modeling?.

Exploring the Effect of Segmentation and Vocabulary Size on Speech Tokenization for Speech Language Models Are Discrete Units Nec- essary for Spoken Language Modeling?

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:50:03.373953Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T14:50:01.653076Z digest=sha256:9a14c6634471b1cb12334ee1f393774d5ea8e48fd6e5f114990b3d7747494fa9

Observation 05e41d96-2af2-4730-a5b6-ac84e179efd6 · outbound

This paper cites Spirit LM: Interleaved Spoken and Written Language Model.

Exploring the Effect of Segmentation and Vocabulary Size on Speech Tokenization for Speech Language Models Spirit LM: Interleaved Spoken and Written Language Model

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T14:50:01.779226Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:50:01.779226Z digest=sha256:7af49dc1e8e5d57f81893846dad3eaf2d6860fc8a78fef9dfe6b886ce2781ea6

Observation 47015db1-6fa3-4ab6-99af-e34b107b2ff2 · outbound

This paper cites Multi- resolution hubert: Multi-resolution speech self-supervised learn- ing with masked unit prediction,.

Exploring the Effect of Segmentation and Vocabulary Size on Speech Tokenization for Speech Language Models Multi- resolution hubert: Multi-resolution speech self-supervised learn- ing with masked unit prediction,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:50:03.120389Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T14:50:01.870557Z digest=sha256:729550dce4941d77914d19780beebc528539fb8147c699daf75893a5ca4b6a09

Observation a782d3a6-28f9-42e3-ba2e-e69c11c1d4c7 · outbound

This paper cites Self-supervised contrastive learning for unsupervised phoneme segmentation,.

Exploring the Effect of Segmentation and Vocabulary Size on Speech Tokenization for Speech Language Models Self-supervised contrastive learning for unsupervised phoneme segmentation,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:50:02.865867Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T14:50:01.956551Z digest=sha256:43a7f457669b2ffdfb0c369ff924987c96bb6e6361979c07d9d0584b2ee19690

Observation abf54504-99ea-4786-be68-a93e630cec9a · outbound

This paper cites Unsupervised word segmentation using temporal gradient pseudo-labels,.

Exploring the Effect of Segmentation and Vocabulary Size on Speech Tokenization for Speech Language Models Unsupervised word segmentation using temporal gradient pseudo-labels,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:50:02.620676Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T14:50:02.022181Z digest=sha256:42811209c1d815113b1f92c22e8b932c3cf0835d99031dbc78b57a16505a4e75

Pith citing papers

Observation 8cf812e2-ced0-49ae-874a-a2004ad39b8e · inbound

Exploring the Effect of Segmentation and Vocabulary Size on Speech Tokenization for Speech Language Models cites this paper.

Exploring the Effect of Segmentation and Vocabulary Size on Speech Tokenization for Speech Language Models Exploring the Effect of Segmentation and Vocabulary Size on Speech Tokenization for Speech Language Models

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-08-07T14:50:02.409712Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-23T06:30:58.430688+00:00.

source=pdf_text observed=2026-08-07T14:49:57.859145Z digest=sha256:b8a92b66363db5131b25f4938f3064236142e6a29cdb46e793a8718355307bfe