Pith. sign in

Paper Citation Record · LEDGER

HuBERT: Self-Supervised Speech Representation Learning by Masked Prediction of Hidden Units

As of 21 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 34 inbound Pith citation observations for arXiv:2106.07447.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2106.07447 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 34 of 34 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 34 of 34 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T00:47:52.584754Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

25
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 063db37c-f8db-4c10-af47-3c256df3285b · inbound

Towards Speaker Identification with Minimal Dataset and Constrained Resources using 1D-Convolution Neural Network cites this paper.

Towards Speaker Identification with Minimal Dataset and Constrained Resources using 1D-Convolution Neural Network HuBERT: Self-Supervised Speech Representation Learning by Masked Prediction of Hidden Units

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-12T14:33:49.954272Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T14:33:49.954272Z digest=sha256:75da91a9467868e74e714f74662bbc6554596380d1ef5002dba6019e918fe6d7

Observation dc48ff5f-8d49-4f63-8941-bf4181efd1bc · inbound

WhiSPA: Semantically and Psychologically Aligned Whisper with Self-Supervised Contrastive and Student-Teacher Learning cites this paper.

WhiSPA: Semantically and Psychologically Aligned Whisper with Self-Supervised Contrastive and Student-Teacher Learning HuBERT: Self-Supervised Speech Representation Learning by Masked Prediction of Hidden Units

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-10T20:27:22.494754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T20:27:22.494754Z digest=sha256:c80a31b51ca3d406a9c9f906fed2a6d91ee8d8090422d769c0c0cdc17a31b907

Observation dd19769d-df76-4e6a-b750-8aebb223ab4f · inbound

Voila: Voice-Language Foundation Models for Real-Time Autonomous Interaction and Voice Role-Play cites this paper.

Voila: Voice-Language Foundation Models for Real-Time Autonomous Interaction and Voice Role-Play HuBERT: Self-Supervised Speech Representation Learning by Masked Prediction of Hidden Units

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-16T00:47:52.584754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:47:52.584754Z digest=sha256:9f519136ffc076a8fc636cf8d88f7e11b0dce0313e49b128f7dda5c8279828df

Observation 43644128-e3fd-42f2-83bc-2ddc70e9d148 · inbound

Survey of End-to-End Multi-Speaker Automatic Speech Recognition for Monaural Audio cites this paper.

Survey of End-to-End Multi-Speaker Automatic Speech Recognition for Monaural Audio HuBERT: Self-Supervised Speech Representation Learning by Masked Prediction of Hidden Units

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-15T21:03:18.019220Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:03:18.019220Z digest=sha256:cbbffbd01f95deb25f6028b28bbc0779a3be1f9711abda4e3b8c30088a65596d

Observation 3efc365d-704a-4714-8d27-54e70b726133 · inbound

EZ-VC: Easy Zero-shot Any-to-Any Voice Conversion cites this paper.

EZ-VC: Easy Zero-shot Any-to-Any Voice Conversion HuBERT: Self-Supervised Speech Representation Learning by Masked Prediction of Hidden Units

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-07T14:58:57.191969Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:58:57.191969Z digest=sha256:e31c09423040623bb80deaf89d6c0ffc12c23037fdc2be8834a66fa513555e43

Observation c86a60e8-64fa-429d-b203-491345f93955 · inbound

Multimodal Zero-Shot Framework for Deepfake Hate Speech Detection in Low-Resource Languages cites this paper.

Multimodal Zero-Shot Framework for Deepfake Hate Speech Detection in Low-Resource Languages HuBERT: Self-Supervised Speech Representation Learning by Masked Prediction of Hidden Units

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T05:20:04.453086Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:20:04.453086Z digest=sha256:e0c59d9ae1b42fb57b744f252ec19152950333fc65c4a12694a7c564f628a99a

Observation b907b8c8-5dd0-4009-a434-2028a84c0a45 · inbound

Benchmarking Foundation Speech and Language Models for Alzheimer's Disease and Related Dementia Detection from Spontaneous Speech cites this paper.

Benchmarking Foundation Speech and Language Models for Alzheimer's Disease and Related Dementia Detection from Spontaneous Speech HuBERT: Self-Supervised Speech Representation Learning by Masked Prediction of Hidden Units

Reference 3460

Resolution
unresolved
no resolver link, observed 2026-08-07T05:26:36.292242Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:26:36.292242Z digest=sha256:01c31d953089940b1f75b0655cc326529009253c7b0d16d88e33e52aca2e28bc

Observation 2757860d-2e1a-4049-800c-f969eeb485a8 · inbound

Zero-Shot Cognitive Impairment Detection from Speech Using AudioLLM cites this paper.

Zero-Shot Cognitive Impairment Detection from Speech Using AudioLLM HuBERT: Self-Supervised Speech Representation Learning by Masked Prediction of Hidden Units

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T23:41:49.569844Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:41:49.569844Z digest=sha256:e25ffa2a638c543afa14b4968a8d9cbd872599a922a3028518660ae61ccb87eb

Observation 762248e0-df10-4ea3-ac8f-a499ded69049 · inbound

XY-Tokenizer: Mitigating the Semantic-Acoustic Conflict in Low-Bitrate Speech Codecs cites this paper.

XY-Tokenizer: Mitigating the Semantic-Acoustic Conflict in Low-Bitrate Speech Codecs HuBERT: Self-Supervised Speech Representation Learning by Masked Prediction of Hidden Units

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T21:51:35.040172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:51:35.040172Z digest=sha256:f2716d704a524669c207d4bbaa4023229bce95ba650cdb5693a18e8dcdd40d41

Observation ca1dcac3-5898-4a56-8ae9-a51f4ae5a008 · inbound

Scaling Self-Supervised Representation Learning for Symbolic Piano Performance cites this paper.

Scaling Self-Supervised Representation Learning for Symbolic Piano Performance HuBERT: Self-Supervised Speech Representation Learning by Masked Prediction of Hidden Units

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T21:36:11.289042Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:36:11.289042Z digest=sha256:2f1a2e301cd44a80559e6863a599d21598c55bf0460494775f0a4d1cb454fb18

Observation 2f1a916e-e772-4f3d-8282-a493595b354f · inbound

Audio-JEPA: Joint-Embedding Predictive Architecture for Audio Representation Learning cites this paper.

Audio-JEPA: Joint-Embedding Predictive Architecture for Audio Representation Learning HuBERT: Self-Supervised Speech Representation Learning by Masked Prediction of Hidden Units

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T22:56:46.224028Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:56:46.224028Z digest=sha256:fecedd01f6d631369ba273e88ee40a215ddb2790c6e295648f347415a51c33bb

Observation ff90399d-1651-42e3-97fb-5e6f8bcc7037 · inbound

Self-supervised learning of speech representations with Dutch archival data cites this paper.

Self-supervised learning of speech representations with Dutch archival data HuBERT: Self-Supervised Speech Representation Learning by Masked Prediction of Hidden Units

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T19:50:22.355340Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:50:22.355340Z digest=sha256:2760d869d34c071203cf3db57f2d2e119a1059b8daefebada06bfdbdcddfb5a3

Observation e26faa88-32ee-45c7-8e30-07a804b5bb7d · inbound

Leveraging Context for Multimodal Fallacy Classification in Political Debates cites this paper.

Leveraging Context for Multimodal Fallacy Classification in Political Debates HuBERT: Self-Supervised Speech Representation Learning by Masked Prediction of Hidden Units

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T15:31:18.775855Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:31:18.775855Z digest=sha256:735f9aa8cd82d30668cd54886f252d21a1c6c591d7f1c79be85290f3f760baa6

Observation 986ec86d-5bf2-489e-a7fd-de4c8516c67d · inbound

Step-Audio 2 Technical Report cites this paper.

Step-Audio 2 Technical Report HuBERT: Self-Supervised Speech Representation Learning by Masked Prediction of Hidden Units

Reference 30

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T05:59:51.132301Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-16T05:59:50.900436Z digest=sha256:6a9f626d48680f7202ed09d2dfba96e2ab0920af599db86d8b4b3fefd5bc865a

Observation b7d3a70e-6d86-4d32-b58b-f91ce0aac5bf · inbound

A Concept-based approach to Voice Disorder Detection cites this paper.

A Concept-based approach to Voice Disorder Detection HuBERT: Self-Supervised Speech Representation Learning by Masked Prediction of Hidden Units

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T14:49:52.253608Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:49:52.253608Z digest=sha256:505d1bb9a600e9aa554786dfd7be92706f74de8ad69399fbb0f1dfa842720418

Observation d20f6562-f71c-4c16-8d5f-bdf6236aaf4f · inbound

Seeing is Believing: Emotion-Aware Audio-Visual Language Modeling for Expressive Speech Generation cites this paper.

Seeing is Believing: Emotion-Aware Audio-Visual Language Modeling for Expressive Speech Generation HuBERT: Self-Supervised Speech Representation Learning by Masked Prediction of Hidden Units

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-05T17:33:58.485256Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T17:33:58.485256Z digest=sha256:f194e2bcf27cbd91270b88922c3b76e23f947b12696b8b0fe6c1f5077d69f2b7

Observation c9954816-dbc8-43e3-9707-98f3ef9225ec · inbound

Entropy-based Coarse and Compressed Semantic Speech Representation Learning cites this paper.

Entropy-based Coarse and Compressed Semantic Speech Representation Learning HuBERT: Self-Supervised Speech Representation Learning by Masked Prediction of Hidden Units

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-05T13:36:04.730722Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T13:36:04.730722Z digest=sha256:114845c6a45cc07b5381e0d15154886d4b23f37378bd89415883aa9bd3960b63

Observation a0f2b216-8a76-4676-9c08-b3ca6222aa23 · inbound

Where Do Backdoors Live? A Component-Level Analysis of Backdoor Propagation in Speech Language Models cites this paper.

Where Do Backdoors Live? A Component-Level Analysis of Backdoor Propagation in Speech Language Models HuBERT: Self-Supervised Speech Representation Learning by Masked Prediction of Hidden Units

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-04T13:15:22.832912Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T13:15:22.832912Z digest=sha256:911b4d4e190082e6e3703cea00699d211b23bbe69cbf15480cafa967544b292e

Observation 14a60961-3961-41d0-a5f4-b6cea4c81c02 · inbound

A Two-Stage Dual-Modality Model for Facial Emotional Expression Recognition cites this paper.

A Two-Stage Dual-Modality Model for Facial Emotional Expression Recognition HuBERT: Self-Supervised Speech Representation Learning by Masked Prediction of Hidden Units

Reference 47

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T11:39:59.378823Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-15T11:35:38.330061Z digest=sha256:0ce60b74aa83c351d0a6067e3a81b76ab54b1c1e38bac75cd9906fb38a0650b8

Observation baacf565-a834-462f-bf1b-b9e2c8474f3a · inbound

findsylls: A Language-Agnostic Toolkit for Syllable-Level Speech Tokenization and Embedding cites this paper.

findsylls: A Language-Agnostic Toolkit for Syllable-Level Speech Tokenization and Embedding HuBERT: Self-Supervised Speech Representation Learning by Masked Prediction of Hidden Units

Reference 29

Resolution
unresolved
no resolver link, observed 2026-07-13T17:37:12.659819Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T17:37:12.659819Z digest=sha256:9cc4f88b94cf2ad5ffcda9650cd15b576711a0db8fb8827fbf23c9d174e7ee22

Observation f7728530-8560-4a6b-9fcb-4bc70517722d · inbound

Neural networks for Text-to-Speech evaluation cites this paper.

Neural networks for Text-to-Speech evaluation HuBERT: Self-Supervised Speech Representation Learning by Masked Prediction of Hidden Units

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T09:49:54.655165Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-15T09:46:26.884551Z digest=sha256:6af6e8cf421c4125861fe45296ff45adcbdb6e474e9642b0e85c36e0144a4f71

Observation b1c58a5a-15e7-47fa-849e-a00281d74a84 · inbound

Meow-Omni 1: A Multimodal Large Language Model for Feline Ethology cites this paper.

Meow-Omni 1: A Multimodal Large Language Model for Feline Ethology HuBERT: Self-Supervised Speech Representation Learning by Masked Prediction of Hidden Units

Reference 13

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T06:41:43.777162Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-12T04:02:23.847187Z digest=sha256:34dda9826388934874f7bd657d7d8a64d52d8b9b1c2bd00ce442e91e2db26a0f

Observation 590cc051-77a2-40c2-9753-b2bdf1d71bc4 · inbound

Fine-tuning language encoding models on slow fMRI improves prediction for fast ECoG cites this paper.

Fine-tuning language encoding models on slow fMRI improves prediction for fast ECoG HuBERT: Self-Supervised Speech Representation Learning by Masked Prediction of Hidden Units

Reference 6

Resolution
malformed identifier
arxiv_id, observed 2026-05-20T06:53:06.014414Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-20T06:48:52.078640Z digest=sha256:5fca23cc4e6e218fb70f387ab65ef9420a8972b941a5f7c7e6b312f370d6feff

Observation 27588fde-9bd5-4255-a614-0f6e73963047 · inbound

F3-Tokenizer: Taming Audio Autoencoder Latents for Understanding and Generation cites this paper.

F3-Tokenizer: Taming Audio Autoencoder Latents for Understanding and Generation HuBERT: Self-Supervised Speech Representation Learning by Masked Prediction of Hidden Units

Reference 5

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T15:47:06.020988Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-27T23:36:27.369551Z digest=sha256:88eb922db70aaaaffbb93386815c0fd23d949e2203cc3d9fda19b18d31f7a733

Observation fcef0e10-8117-47e2-959c-c7bf9d3907e5 · inbound

Pretrained self-supervised speech models can recognize unseen consonants cites this paper.

Pretrained self-supervised speech models can recognize unseen consonants HuBERT: Self-Supervised Speech Representation Learning by Masked Prediction of Hidden Units

Reference 34

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T10:07:56.083448Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-27T10:14:47.932613Z digest=sha256:345b9e896f93f04c3afa7b0df2ca00d75c0d023e2f134da8ce45a47b5611fb39

Observation 46238d6e-7817-465c-bc14-d6371a8c32b9 · inbound

Light-weight Pronunciation Assessment via Discrete Speech Token Surprisal cites this paper.

Light-weight Pronunciation Assessment via Discrete Speech Token Surprisal HuBERT: Self-Supervised Speech Representation Learning by Masked Prediction of Hidden Units

Reference 58

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T03:49:30.223989Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-26T17:37:25.043607Z digest=sha256:131d056e424e5a80c0ed0ca71785075a007c65a9d032469971f4057732d41793

Observation 4cba52db-200a-4e08-9c0a-e31d1b391a8c · inbound

Interleaved Speech Language Models Latently Work In Text cites this paper.

Interleaved Speech Language Models Latently Work In Text HuBERT: Self-Supervised Speech Representation Learning by Masked Prediction of Hidden Units

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-07-04T08:59:42.891310Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-26T10:41:19.777779Z digest=sha256:a34396318df71e6d640f2a8e22a8aff3aff1319c2a5bb7da614198032b2f05f8

Observation 12d732ad-24e4-49be-9020-97b89026ede6 · inbound

End-to-End Voice Intent Recognition for Spontaneous Human-Drone Interaction with Naive Users cites this paper.

End-to-End Voice Intent Recognition for Spontaneous Human-Drone Interaction with Naive Users HuBERT: Self-Supervised Speech Representation Learning by Masked Prediction of Hidden Units

Reference 30

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T07:29:38.221570Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-26T13:30:12.101045Z digest=sha256:ce430f75015469e5d79cef6f47cf0cb215ae27c643d04a265f7177ab6d46bcf9

Observation 05e6240c-cfc1-42e9-b23f-5ac9f84ef1b8 · inbound

Syntactic Belief Update as the Driver of Garden Path Processing Difficulty cites this paper.

Syntactic Belief Update as the Driver of Garden Path Processing Difficulty HuBERT: Self-Supervised Speech Representation Learning by Masked Prediction of Hidden Units

Reference 288

Resolution
metadata mismatch
arxiv_id, observed 2026-06-26T04:38:58.619776Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-06-26T04:38:01.183423Z digest=sha256:df2cae12451c88109126b5334f0cee3ee3f5341b546e8250f0b61cf398b42168

Observation 28de5576-2bb0-49ca-9446-7276fb253d47 · inbound

Speaker-Aware Temporal Aggregation Strategies on Segment Representations for Depression Detection in Dyadic Interaction: A Benchmark Study cites this paper.

Speaker-Aware Temporal Aggregation Strategies on Segment Representations for Depression Detection in Dyadic Interaction: A Benchmark Study HuBERT: Self-Supervised Speech Representation Learning by Masked Prediction of Hidden Units

Reference 7

Resolution
unresolved
no resolver link, observed 2026-07-12T06:14:42.321767Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T06:14:42.321767Z digest=sha256:f58b4f2b9374041e1a5b025d885f7d1fa89bbd6dc99bed77db4511426f6cdd9d

Observation 2e838a15-a1a9-403d-9d33-4fc1b0085c1f · inbound

Layer-wise Cross-Lingual Depression Detection from Speech: Analysis with Contrastive Alignment cites this paper.

Layer-wise Cross-Lingual Depression Detection from Speech: Analysis with Contrastive Alignment HuBERT: Self-Supervised Speech Representation Learning by Masked Prediction of Hidden Units

Reference 5

Resolution
unresolved
no resolver link, observed 2026-07-12T06:06:46.343211Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T06:06:46.343211Z digest=sha256:4dc598e09508884a161b9ac92250a5d02b4916e30a99fd51831c37198cfacde5

Observation 3df35c21-9c57-4186-97af-02c2123f6fe0 · inbound

Video2Reaction: Mapping Video to Audience Reaction Distribution in the Wild cites this paper.

Video2Reaction: Mapping Video to Audience Reaction Distribution in the Wild HuBERT: Self-Supervised Speech Representation Learning by Masked Prediction of Hidden Units

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-07-09T23:56:38.448201Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-09T23:51:55.422232Z digest=sha256:daa508e39921c62dfa1198429284a4567a0fc47a5fd76895d9e3a74aa67c351c

Observation 0c6ac7ef-ce87-48c6-acb2-7e2bfde25cf9 · inbound

DONDO: Open w2v-BERT Speech-Recognition Base Models for African Languages cites this paper.

DONDO: Open w2v-BERT Speech-Recognition Base Models for African Languages HuBERT: Self-Supervised Speech Representation Learning by Masked Prediction of Hidden Units

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-01T07:10:39.830211Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:10:39.830211Z digest=sha256:62e895d0df4f806339df05e3452ff7cb42b953681ab336fcce4a259b41dc1043

Observation d49c1e6b-732d-4b39-9099-322895f7ec7b · inbound

A Unifying Perspective on Audio Generative Modeling: Latent Representations and Modeling Strategies cites this paper.

A Unifying Perspective on Audio Generative Modeling: Latent Representations and Modeling Strategies HuBERT: Self-Supervised Speech Representation Learning by Masked Prediction of Hidden Units

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-14T04:32:47.948785Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T04:32:47.948785Z digest=sha256:05a6b391216122090336f4e927ca4706a302260bf0affea76577543bd31b8612