Pith. sign in

Paper Citation Record · LEDGER

XLS-R: Self-supervised Cross-lingual Speech Representation Learning at Scale

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 57 inbound Pith citation observations for arXiv:2111.09296.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2111.09296 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 57 of 57 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 57 of 57 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-09T04:35:12.037894Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 02381b35-f84c-445a-ae80-fce5332e9019 · inbound

Comprehensive Layer-wise Analysis of SSL Models for Audio Deepfake Detection cites this paper.

Comprehensive Layer-wise Analysis of SSL Models for Audio Deepfake Detection XLS-R: Self-supervised Cross-lingual Speech Representation Learning at Scale

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-09T04:35:12.037894Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T04:35:12.037894Z digest=sha256:48c8978aa06c7d63ce5dfa9c246c1835708226ef06d50ea308ae8debb20b27c1

Observation 8a519f52-c5d5-47fc-8e45-a2aaf4a6f663 · inbound

Evaluating Standard and Dialectal Frisian ASR: Multilingual Fine-tuning and Language Identification for Improved Low-resource Performance cites this paper.

Evaluating Standard and Dialectal Frisian ASR: Multilingual Fine-tuning and Language Identification for Improved Low-resource Performance XLS-R: Self-supervised Cross-lingual Speech Representation Learning at Scale

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-08T21:10:38.112255Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T21:10:38.112255Z digest=sha256:b84040cb97f62d50a07fbefb4b4487ee1fa65fa0c9dd3c517ebdc0c51dffaad8

Observation 7aaf16de-cdf7-47dd-9d8b-0e43c321bd87 · inbound

Enhancing Expressive Voice Conversion with Discrete Pitch-Conditioned Flow Matching Model cites this paper.

Enhancing Expressive Voice Conversion with Discrete Pitch-Conditioned Flow Matching Model XLS-R: Self-supervised Cross-lingual Speech Representation Learning at Scale

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-08T19:15:52.095800Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:15:52.095800Z digest=sha256:fe9e26088b61e9df82d29ffada4ca766f8ba68feae9050f81094f91fc7e97a0c

Observation fb63d79a-1d1c-4c4c-b71f-796f33f5281d · inbound

Recreating Neural Activity During Speech Production with Language and Speech Model Embeddings cites this paper.

Recreating Neural Activity During Speech Production with Language and Speech Model Embeddings XLS-R: Self-supervised Cross-lingual Speech Representation Learning at Scale

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T15:42:19.207531Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:42:19.207531Z digest=sha256:e459c6410569a57a8898873166b873abc976aaa6f9c2327eee8de1a39fe76fd2

Observation 5e4f22e8-cb45-4434-ba3e-6316c3a68fce · inbound

PSRB: A Comprehensive Benchmark for Evaluating Persian ASR Systems cites this paper.

PSRB: A Comprehensive Benchmark for Evaluating Persian ASR Systems XLS-R: Self-supervised Cross-lingual Speech Representation Learning at Scale

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T13:42:54.835853Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:42:54.835853Z digest=sha256:9100b062053267c95117fdad7970e2b77226ef10818dfb3476f2a070f1dd18da

Observation 36098707-1984-449d-af69-2d68b3e84bea · inbound

ZIPA: A family of efficient models for multilingual phone recognition cites this paper.

ZIPA: A family of efficient models for multilingual phone recognition XLS-R: Self-supervised Cross-lingual Speech Representation Learning at Scale

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-07T12:55:39.944247Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:55:39.944247Z digest=sha256:ed921d34854bc74571984bca9058248c6d08acfb2213ebde15de285b72f9c334

Observation 8c6ee2ef-5abc-41d7-a341-6ff54fb01f02 · inbound

Few-Shot Speech Deepfake Detection Adaptation with Gaussian Processes cites this paper.

Few-Shot Speech Deepfake Detection Adaptation with Gaussian Processes XLS-R: Self-supervised Cross-lingual Speech Representation Learning at Scale

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:29.628714Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:29.628714Z digest=sha256:0bd5c17a2f484111921e9ed80ba91dc326c5a8d1fb4711a665908591f21b76ae

Observation ad37c472-53a9-44e6-baed-4e1850d22a49 · inbound

Assortment of Attention Heads: Accelerating Federated PEFT with Head Pruning and Strategic Client Selection cites this paper.

Assortment of Attention Heads: Accelerating Federated PEFT with Head Pruning and Strategic Client Selection XLS-R: Self-supervised Cross-lingual Speech Representation Learning at Scale

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T12:05:51.843072Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:05:51.843072Z digest=sha256:f61923ac88c4223ff91afc40ac52192235a4735da58c1ea329661670bfe4b05e

Observation 38d25a94-b079-4b34-b82c-8025f1e96b67 · inbound

Rhythm Controllable and Efficient Zero-Shot Voice Conversion via Shortcut Flow Matching cites this paper.

Rhythm Controllable and Efficient Zero-Shot Voice Conversion via Shortcut Flow Matching XLS-R: Self-supervised Cross-lingual Speech Representation Learning at Scale

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:52.905967Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:57:52.905967Z digest=sha256:880785d26e6597ad9e6a655324362437b895e45978c02c27dc1436b36489881f

Observation df060ada-c158-40dd-a58e-c297ca8c74c8 · inbound

Towards Generalized Source Tracing for Codec-Based Deepfake Speech cites this paper.

Towards Generalized Source Tracing for Codec-Based Deepfake Speech XLS-R: Self-supervised Cross-lingual Speech Representation Learning at Scale

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T05:43:08.358475Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:43:08.358475Z digest=sha256:a563ff8c3001d9c5abe2bdd60628c7bd281e53ec47ab7b2b81b1e198e7a0a651

Observation 97d95895-7355-4b31-83b6-d50e6d76d3c4 · inbound

Joint ASR and Speaker Role Tagging with Serialized Output Training cites this paper.

Joint ASR and Speaker Role Tagging with Serialized Output Training XLS-R: Self-supervised Cross-lingual Speech Representation Learning at Scale

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T04:32:30.227311Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:32:30.227311Z digest=sha256:3c9aa87d0216402a65a203bfd25570ad7811ce3f3e677bae1b9303a8b86f4d11

Observation e892ec6f-7acc-49c8-a6fa-fde74123eb15 · inbound

From Sharpness to Better Generalization for Speech Deepfake Detection cites this paper.

From Sharpness to Better Generalization for Speech Deepfake Detection XLS-R: Self-supervised Cross-lingual Speech Representation Learning at Scale

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:06.077547Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:06.077547Z digest=sha256:c59bd376882efcb82f4b01e01dc16067968ee5147903875a83e624015b11b98e

Observation ff93de16-16f8-419b-9e1a-54ae05522f3d · inbound

Pushing the Performance of Synthetic Speech Detection with Kolmogorov-Arnold Networks and Self-Supervised Learning Models cites this paper.

Pushing the Performance of Synthetic Speech Detection with Kolmogorov-Arnold Networks and Self-Supervised Learning Models XLS-R: Self-supervised Cross-lingual Speech Representation Learning at Scale

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T00:22:26.715238Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:22:26.715238Z digest=sha256:7924a41cd42a1ab405b9a97048bc093549255ea7e4a63f2803cd282da4bbf5fc

Observation 4ba1af66-387e-4ddc-bbca-aa4840297d0d · inbound

Breaking the Transcription Bottleneck: Fine-tuning ASR Models for Extremely Low-Resource Fieldwork Languages cites this paper.

Breaking the Transcription Bottleneck: Fine-tuning ASR Models for Extremely Low-Resource Fieldwork Languages XLS-R: Self-supervised Cross-lingual Speech Representation Learning at Scale

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T23:35:28.171349Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:35:28.171349Z digest=sha256:287a6e57f35837e2be6b5dc06b22c63a21ea16a917f7cfa2550a499d45c06515

Observation bacce3da-fc5d-4cf5-bbbf-bfeef09b4d93 · inbound

Data Quality Issues in Multilingual Speech Datasets: The Need for Sociolinguistic Awareness and Proactive Language Planning cites this paper.

Data Quality Issues in Multilingual Speech Datasets: The Need for Sociolinguistic Awareness and Proactive Language Planning XLS-R: Self-supervised Cross-lingual Speech Representation Learning at Scale

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T23:34:49.009463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:34:49.009463Z digest=sha256:f016dd3c451898f2e4424060f4a88cb3ae1f27ca6443ec8be5bce4897a3ca455

Observation 045ac136-f4e3-4f8b-8dce-fd38e6707afd · inbound

Thinking Beyond Tokens: From Brain-Inspired Intelligence to Cognitive Foundations for Artificial General Intelligence and its Societal Impact cites this paper.

Thinking Beyond Tokens: From Brain-Inspired Intelligence to Cognitive Foundations for Artificial General Intelligence and its Societal Impact XLS-R: Self-supervised Cross-lingual Speech Representation Learning at Scale

Reference 159

Resolution
unresolved
no resolver link, observed 2026-08-06T21:07:11.272341Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:07:11.272341Z digest=sha256:a1e52c44e896cab7ba045b25c10514066a0652e77ad2d1043e1e81b34a7e7991

Observation dad648f9-9aa1-4934-b17b-0668ee6d121e · inbound

DiceHuBERT: Distilling HuBERT with a Self-Supervised Learning Objective cites this paper.

DiceHuBERT: Distilling HuBERT with a Self-Supervised Learning Objective XLS-R: Self-supervised Cross-lingual Speech Representation Learning at Scale

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T23:02:14.927874Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:02:14.927874Z digest=sha256:d88833d00e1c6db3b93da376cd6e19173d95973c2c6d02999be96cacbd66224d

Observation 845ae894-6d3c-49bf-8e4e-53e6d3f9e739 · inbound

Word stress in self-supervised speech models: A cross-linguistic comparison cites this paper.

Word stress in self-supervised speech models: A cross-linguistic comparison XLS-R: Self-supervised Cross-lingual Speech Representation Learning at Scale

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T19:44:26.885657Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:44:26.885657Z digest=sha256:e93f4a4cadb0c134f43498946fa4e98a3deaaa1d2e28bd089504a9a4a5b8d589

Observation 6910d54b-696a-4824-9f91-cfc860254baf · inbound

RepeaTTS: Towards Feature Discovery through Repeated Fine-Tuning cites this paper.

RepeaTTS: Towards Feature Discovery through Repeated Fine-Tuning XLS-R: Self-supervised Cross-lingual Speech Representation Learning at Scale

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T20:01:35.492150Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:01:35.492150Z digest=sha256:2b2e7e479a622ecd1a13a9d78a86a2a1ba6bc89479c2f483636f6c109008b9d7

Observation 1991a958-5633-48f9-ac24-634cc4e0fab1 · inbound

On Barriers to Archival Audio Processing cites this paper.

On Barriers to Archival Audio Processing XLS-R: Self-supervised Cross-lingual Speech Representation Learning at Scale

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T18:14:08.779769Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:14:08.779769Z digest=sha256:adbc2365af7a12a67210211c218bcdbbe271398d2aa2c86831dd7a0e61c9a64f

Observation f1c21fd1-9842-447b-8a1c-14db27d29faa · inbound

SpeechFake: A Large-Scale Multilingual Speech Deepfake Dataset Incorporating Cutting-Edge Generation Methods cites this paper.

SpeechFake: A Large-Scale Multilingual Speech Deepfake Dataset Incorporating Cutting-Edge Generation Methods XLS-R: Self-supervised Cross-lingual Speech Representation Learning at Scale

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T12:49:16.380954Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T12:49:16.380954Z digest=sha256:037c0413e469f981b69def35870c04ab9fd6f6e168b2c3f3658a2396d0197665

Observation ddfdcf68-6bbb-418c-92ca-b7e5ef60aa85 · inbound

CAM\~OES: A Comprehensive Automatic Speech Recognition Benchmark for European Portuguese cites this paper.

CAM\~OES: A Comprehensive Automatic Speech Recognition Benchmark for European Portuguese XLS-R: Self-supervised Cross-lingual Speech Representation Learning at Scale

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-05T15:39:21.516370Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:39:21.516370Z digest=sha256:144e3e0ad03413fca3c938e274340eb958752ca9c6ad66a6a79dc3962b2b1775

Observation b1e5f998-c1c8-4785-91e5-dd8b66b957c6 · inbound

Generalizable Audio Spoofing Detection using Non-Semantic Representations cites this paper.

Generalizable Audio Spoofing Detection using Non-Semantic Representations XLS-R: Self-supervised Cross-lingual Speech Representation Learning at Scale

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-05T13:55:18.277810Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:55:18.277810Z digest=sha256:0c3f01f8680b5b9af52ae73e5efec11d7528a43f389ba430e031c5aa58c9279c

Observation b423e264-8805-408e-8884-c868e8f5ad10 · inbound

Forensic Similarity for Speech Deepfakes cites this paper.

Forensic Similarity for Speech Deepfakes XLS-R: Self-supervised Cross-lingual Speech Representation Learning at Scale

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-18T10:31:14.902664Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-18T10:28:12.467606Z digest=sha256:15c4f5b650bed1c1c865509d9b0beab7130d19ea1eae67ae3f12a6b8970e0cd9

Observation 7ea9d4ba-0fda-47c3-ada1-5f2e3c64aae3 · inbound

SONAR: Spectral-Contrastive Audio Residuals for Generalizable Deepfake Detection cites this paper.

SONAR: Spectral-Contrastive Audio Residuals for Generalizable Deepfake Detection XLS-R: Self-supervised Cross-lingual Speech Representation Learning at Scale

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-03T20:05:29.992627Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:05:29.992627Z digest=sha256:03bc4268262b3848a5eef0dcd0653aeed5ed70dae5ec6e9f7c9e81a012ba238a

Observation 98b8ee19-d213-48f5-9aa8-f181363ba2b3 · inbound

A SUPERB-Style Benchmark of Self-Supervised Speech Models for Audio Deepfake Detection cites this paper.

A SUPERB-Style Benchmark of Self-Supervised Speech Models for Audio Deepfake Detection XLS-R: Self-supervised Cross-lingual Speech Representation Learning at Scale

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-15T17:26:22.179791Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-15T17:25:27.137423Z digest=sha256:577f8b83fecf411041646c1807f0254d355fae9842d232dd244729623f347dfa

Observation be6fd803-7145-4683-a19c-101dd423dd21 · inbound

Benchmarking Multilingual Speech Models on Pashto: Zero-Shot ASR, Script Failure, and Cross-Domain Evaluation cites this paper.

Benchmarking Multilingual Speech Models on Pashto: Zero-Shot ASR, Script Failure, and Cross-Domain Evaluation XLS-R: Self-supervised Cross-lingual Speech Representation Learning at Scale

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-10T22:35:48.700833Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T19:44:30.762851Z digest=sha256:d0eadb93d7d65b8851d588f17f6e0ca0bc4f4b1e6cb9e11c720c61ce48398514

Observation db07feee-6407-4425-8621-09c452858f2c · inbound

Giving Voice to the Constitution: Low-Resource Text-to-Speech for Quechua and Spanish Using a Bilingual Legal Corpus cites this paper.

Giving Voice to the Constitution: Low-Resource Text-to-Speech for Quechua and Spanish Using a Bilingual Legal Corpus XLS-R: Self-supervised Cross-lingual Speech Representation Learning at Scale

Reference 7

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T11:21:00.993257Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T15:01:20.902986Z digest=sha256:315a2064c6895c2bdd5122be2c7740874e5d96cbde9e0bfa0e8a42acbd64d98a

Observation dbb3df5e-639c-4e03-a9f6-4c0a3e50176d · inbound

Similarity Choice and Negative Scaling in Supervised Contrastive Learning for Deepfake Audio Detection cites this paper.

Similarity Choice and Negative Scaling in Supervised Contrastive Learning for Deepfake Audio Detection XLS-R: Self-supervised Cross-lingual Speech Representation Learning at Scale

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-12T08:46:25.862041Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-07T13:50:53.553261Z digest=sha256:40640d649673bb151c3ee0f1d7bdb99b16c8b5ea1783b2273ef41a1b2647268f

Observation e98638b2-ad8c-42d9-8f32-79e8ac4438e6 · inbound

Profiling the Voice: Speaker-Specific Phoneme Fingerprinting for Speech Deepfake Detection cites this paper.

Profiling the Voice: Speaker-Specific Phoneme Fingerprinting for Speech Deepfake Detection XLS-R: Self-supervised Cross-lingual Speech Representation Learning at Scale

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-20T01:22:55.817247Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-20T01:20:04.609575Z digest=sha256:9c53c110def4e7eac95a873c84b0ff9c59edf89baa1bde1c50e75d51b3f17826

Observation cce6bf73-5148-4a1d-b5a5-29cb6cbcb3d4 · inbound

MixFake: Benchmarking and Enhancing Audio Deepfake Detection in Diverse Real-world Mixed Audio cites this paper.

MixFake: Benchmarking and Enhancing Audio Deepfake Detection in Diverse Real-world Mixed Audio XLS-R: Self-supervised Cross-lingual Speech Representation Learning at Scale

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-25T03:15:17.023264Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-25T03:14:08.880072Z digest=sha256:289ac1b358f21c19255bb06c0199ca8d310641a34d87cc0114e69f8504014ab2

Observation ddaf9454-c96e-42c0-b5fe-c0ab99677897 · inbound

Escaping the Linearity Trap: Manifold Detours for Black-Box Adversarial Attacks on Singing Audio Deepfake Detection cites this paper.

Escaping the Linearity Trap: Manifold Detours for Black-Box Adversarial Attacks on Singing Audio Deepfake Detection XLS-R: Self-supervised Cross-lingual Speech Representation Learning at Scale

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-06-30T19:05:01.112227Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-30T18:58:31.318883Z digest=sha256:207246eaf702713b3188bc47300d66f08735d56c6812112bc57bc131dcfc631b

Observation 92c547c6-ca7c-4952-967d-fb9b26bfd232 · inbound

Dual-Branch Gated Fusion for Open-Set Audio Deepfake Source Tracing cites this paper.

Dual-Branch Gated Fusion for Open-Set Audio Deepfake Source Tracing XLS-R: Self-supervised Cross-lingual Speech Representation Learning at Scale

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-07-03T03:47:35.572305Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-27T14:45:45.359616Z digest=sha256:6f51316525fe2ea7f59880df69de68916c9b81e27f561f0f9a050ae6a4ecb020

Observation f4260727-f010-4480-8756-d5f2a3dee562 · inbound

Pretrained self-supervised speech models can recognize unseen consonants cites this paper.

Pretrained self-supervised speech models can recognize unseen consonants XLS-R: Self-supervised Cross-lingual Speech Representation Learning at Scale

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-07-03T10:07:56.070600Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-27T10:14:47.932613Z digest=sha256:723f89a0fd3e437d2369059b82bceb39558abab195f482870d945feadd14df2f

Observation cbe71276-4f07-4fdb-8176-4a33c86cc761 · inbound

SpAArSIST: Sparsified AASIST for Efficient and Reliable Anti-Spoofing cites this paper.

SpAArSIST: Sparsified AASIST for Efficient and Reliable Anti-Spoofing XLS-R: Self-supervised Cross-lingual Speech Representation Learning at Scale

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-07-03T12:58:08.668281Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-27T08:35:24.321197Z digest=sha256:f4dd1ce452b6c99b9936007bcfef397fc4d5a56e994ebd75e833c9a67a773242

Observation 18ed6e2d-fcd1-4e0b-9178-354240c9df74 · inbound

Responsible ASR: Overcoming Challenges of Foundational Models in Narrow-Band and Low-Resource Settings cites this paper.

Responsible ASR: Overcoming Challenges of Foundational Models in Narrow-Band and Low-Resource Settings XLS-R: Self-supervised Cross-lingual Speech Representation Learning at Scale

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-07-04T01:59:26.658181Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-26T20:03:28.491546Z digest=sha256:5476e35d865b22fb7d7520599b77e92edc734fadcb52841dac18250d3b57d499

Observation dd75c079-7ec4-4819-9099-12cd136d5123 · inbound

Light-weight Pronunciation Assessment via Discrete Speech Token Surprisal cites this paper.

Light-weight Pronunciation Assessment via Discrete Speech Token Surprisal XLS-R: Self-supervised Cross-lingual Speech Representation Learning at Scale

Reference 60

Resolution
verified exact
arxiv_id, observed 2026-07-04T03:49:30.231851Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-26T17:37:25.043607Z digest=sha256:8be586d06db7a0ebe6f17f335199c676ea904c94730330446bd92a40fc9e89cf

Observation dedb92c1-9f15-4ce5-8432-63b950b9a2fe · inbound

Beyond Speaker Independence: Evaluating Cross-Lingual Acoustic-to-Articulatory Inversion Across Finnish and Russian cites this paper.

Beyond Speaker Independence: Evaluating Cross-Lingual Acoustic-to-Articulatory Inversion Across Finnish and Russian XLS-R: Self-supervised Cross-lingual Speech Representation Learning at Scale

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-07-04T05:49:37.763673Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-26T15:24:04.594993Z digest=sha256:6d347a9dda76478f86c3733c95aa7eb4814a4535bc60f9dfb3fdcf714ac2a127

Observation e82d9116-99fc-4ed9-b1a5-7b70cbc47ce5 · inbound

Impact Analysis of Speech Representation Learning Models for Acoustic Side-Channel Attack cites this paper.

Impact Analysis of Speech Representation Learning Models for Acoustic Side-Channel Attack XLS-R: Self-supervised Cross-lingual Speech Representation Learning at Scale

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-07-04T06:39:37.331175Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-26T14:19:59.573391Z digest=sha256:9b52a14901fb8c7a2267a9af69207b89e334a97fa56411770a16ba5d03ebbffd

Observation fbff2c09-d503-4660-b16d-78a7dabb35c4 · inbound

Impact Analysis of Speech Representation Learning Models for Acoustic Side-Channel Attack cites this paper.

Impact Analysis of Speech Representation Learning Models for Acoustic Side-Channel Attack XLS-R: Self-supervised Cross-lingual Speech Representation Learning at Scale

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-06-29T19:33:54.406475Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-29T04:44:52.537405Z digest=sha256:4afc21ea861b905695807a45de5dcfd25c02280b6f54c52e448099ecbc3f9bf1

Observation 61adb2a8-dc23-40af-a1f2-68848c557b19 · inbound

From Speech to Text Corpora: Evaluating ASR-Based Data Acquisition for Low-Resource Fongbe and Hausa cites this paper.

From Speech to Text Corpora: Evaluating ASR-Based Data Acquisition for Low-Resource Fongbe and Hausa XLS-R: Self-supervised Cross-lingual Speech Representation Learning at Scale

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-07-04T08:29:42.395446Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-26T11:31:09.747113Z digest=sha256:d62fb9d9dc2edfc0269bf7591f24d72df1801ac87552ba3a64161bbdc80b090c

Observation 5d125a69-5efe-41d1-b65a-f8de9b2f395c · inbound

What Do Deepfake Benchmarks Measure? An Audit Using Frozen Self-Supervised Representations cites this paper.

What Do Deepfake Benchmarks Measure? An Audit Using Frozen Self-Supervised Representations XLS-R: Self-supervised Cross-lingual Speech Representation Learning at Scale

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-07-04T15:49:57.200451Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-26T01:24:46.200846Z digest=sha256:d25dafd4fd71c218fb89e2083c8139117b7fd85e07abd397c371b9450be9262d

Observation 2a86893c-fa6b-413d-9c74-b2607ed4d8f7 · inbound

Syntactic Belief Update as the Driver of Garden Path Processing Difficulty cites this paper.

Syntactic Belief Update as the Driver of Garden Path Processing Difficulty XLS-R: Self-supervised Cross-lingual Speech Representation Learning at Scale

Reference 295

Resolution
metadata mismatch
arxiv_id, observed 2026-06-26T04:38:58.354104Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-26T04:38:01.183423Z digest=sha256:c502c96ac7d11ee763f4a454939638e3caf6e5210a1a823e19f86bc6d2a3eac1

Observation 5e46740a-119f-4c26-ac54-9e04ccec34c6 · inbound

Detecting Audio Deepfakes on the Edge:Lightweight SSL-Based Detection in a Browser Plugin cites this paper.

Detecting Audio Deepfakes on the Edge:Lightweight SSL-Based Detection in a Browser Plugin XLS-R: Self-supervised Cross-lingual Speech Representation Learning at Scale

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-07-01T12:45:44.602950Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-01T01:47:41.684335Z digest=sha256:00b7c0f9efd004b2961fd4e7dad8969f689bd7ec495ae4fb61cbea7e0e2979b5

Observation abbf0d63-cf32-495b-a9cb-570350d25ac4 · inbound

From Objectives to Applications: Aligning Architectural Biases in Audio Self-Supervised Learning cites this paper.

From Objectives to Applications: Aligning Architectural Biases in Audio Self-Supervised Learning XLS-R: Self-supervised Cross-lingual Speech Representation Learning at Scale

Reference 93

Resolution
verified exact
arxiv_id, observed 2026-07-02T05:56:39.846458Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-02T05:52:55.818877Z digest=sha256:19906ab7ac477ae74f90e5efcc14cb7bdd2b596634305d955cf91ba6304a4f80

Observation 5a231095-8343-43cd-9b28-7b84916812e4 · inbound

Speaker-Aware Temporal Aggregation Strategies on Segment Representations for Depression Detection in Dyadic Interaction: A Benchmark Study cites this paper.

Speaker-Aware Temporal Aggregation Strategies on Segment Representations for Depression Detection in Dyadic Interaction: A Benchmark Study XLS-R: Self-supervised Cross-lingual Speech Representation Learning at Scale

Reference 10

Resolution
unresolved
no resolver link, observed 2026-07-12T06:14:42.321767Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T06:14:42.321767Z digest=sha256:37511ec332d2a6f1b4892c1ba5f97dc4f518b6a9e28169d087d0866b57f90f6d

Observation 59a919b8-fb23-4828-adb2-3d3995836e7c · inbound

An Intervention-Based Framework for Shortcut Diagnosis in Spoofing Countermeasures cites this paper.

An Intervention-Based Framework for Shortcut Diagnosis in Spoofing Countermeasures XLS-R: Self-supervised Cross-lingual Speech Representation Learning at Scale

Reference 18

Resolution
unresolved
no resolver link, observed 2026-07-12T04:31:11.771594Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T04:31:11.771594Z digest=sha256:1aaf78e173ad23318a6dfef13904fc6128bdab659ea60f5072908027a486053a

Observation 5265bd2a-3338-4e30-b7ee-b6840c656f53 · inbound

Towards Digital Preservation of Efik: TTS for a Low-Resource African Language cites this paper.

Towards Digital Preservation of Efik: TTS for a Low-Resource African Language XLS-R: Self-supervised Cross-lingual Speech Representation Learning at Scale

Reference 36

Resolution
unresolved
no resolver link, observed 2026-07-11T18:11:36.626189Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T18:11:36.626189Z digest=sha256:a27a916282afae1a3f1102e754a2caf79b0f0a04a877b8980537babe4d96061c

Observation 83eed9c1-9c22-4fb8-b782-3ff5b7d21eb3 · inbound

Evaluating the Effect of Linguistic Relatedness on Cross-Lingual Transfer in Large Multilingual Automatic Speech Recognition cites this paper.

Evaluating the Effect of Linguistic Relatedness on Cross-Lingual Transfer in Large Multilingual Automatic Speech Recognition XLS-R: Self-supervised Cross-lingual Speech Representation Learning at Scale

Reference 2

Resolution
unresolved
no resolver link, observed 2026-07-11T13:07:13.175134Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T13:07:13.175134Z digest=sha256:81a05eafc65fd7829fb044578f127fa7898b0917227b97142343bbdcdcad659d

Observation 087998ef-6c92-4171-8258-2dacc3719058 · inbound

Evaluating the Effect of Linguistic Relatedness on Cross-Lingual Transfer in Large Multilingual Automatic Speech Recognition cites this paper.

Evaluating the Effect of Linguistic Relatedness on Cross-Lingual Transfer in Large Multilingual Automatic Speech Recognition XLS-R: Self-supervised Cross-lingual Speech Representation Learning at Scale

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-02T08:34:55.605107Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:34:55.605107Z digest=sha256:ecee23613a10c9938623ebe5e152e2635698d6aa5edaf24be6b056493c652661

Observation 2d8c6bc4-790e-437b-acb7-3688e45872b3 · inbound

GigaAM Multilingual: Foundation Model for Underrepresented Languages cites this paper.

GigaAM Multilingual: Foundation Model for Underrepresented Languages XLS-R: Self-supervised Cross-lingual Speech Representation Learning at Scale

Reference 17

Resolution
unresolved
no resolver link, observed 2026-07-14T12:15:06.470439Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T12:15:06.470439Z digest=sha256:0a3848ceb97dc5c1dce8188a765a945ce09179a23df1f8fa065586673d57326a

Observation 0b197b28-1724-4dbe-9a52-d2703304a7d9 · inbound

Unified Gradient Projection: Language-Balanced Continual Learning for Multilingual Low-Resource ASR cites this paper.

Unified Gradient Projection: Language-Balanced Continual Learning for Multilingual Low-Resource ASR XLS-R: Self-supervised Cross-lingual Speech Representation Learning at Scale

Reference 8

Resolution
unresolved
no resolver link, observed 2026-07-14T06:34:55.089754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T06:34:55.089754Z digest=sha256:a19ed681d48c7083fbeb0dd45ab69cdcd675902a52ee883448e16c7a5da30c15

Observation 397ebc1a-c572-4365-b947-30a876530d83 · inbound

DONDO: Open w2v-BERT Speech-Recognition Base Models for African Languages cites this paper.

DONDO: Open w2v-BERT Speech-Recognition Base Models for African Languages XLS-R: Self-supervised Cross-lingual Speech Representation Learning at Scale

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-01T07:10:39.837818Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:10:39.837818Z digest=sha256:58123987cbbd5b79f03934f4a9beb933a9fe4d49222fe3aab144d89f9c96754c

Observation af26c9ad-3b23-4deb-be16-701e9507960f · inbound

Leveraging Gradient Reversal Loss and Multitask Learning for Datasets-Aware Audio Deepfake Detection cites this paper.

Leveraging Gradient Reversal Loss and Multitask Learning for Datasets-Aware Audio Deepfake Detection XLS-R: Self-supervised Cross-lingual Speech Representation Learning at Scale

Reference 33

Resolution
unresolved
no resolver link, observed 2026-07-31T23:27:02.078077Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T23:27:02.078077Z digest=sha256:99985e489e08effe1c26e4cb2f078154be19e6b58c7359d1dec92975962ab240

Observation 46196a50-20d5-4b09-b76d-b2dd78425e91 · inbound

Teffic-Audio: Tell Fact from Fiction cites this paper.

Teffic-Audio: Tell Fact from Fiction XLS-R: Self-supervised Cross-lingual Speech Representation Learning at Scale

Reference 15

Resolution
unresolved
no resolver link, observed 2026-07-31T10:28:22.087327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T10:28:22.087327Z digest=sha256:153278f6c99577335a333a9bb4d629928265a4b24e65973ae6325a2bdadbb53d

Observation d97e4775-c32e-4e8e-a2a1-0ac095920b20 · inbound

Multi-Backbone Self-Supervised Ensembles for Audio Deepfake Detection and a Cross-Track Analysis of Generation-Detection Asymmetry cites this paper.

Multi-Backbone Self-Supervised Ensembles for Audio Deepfake Detection and a Cross-Track Analysis of Generation-Detection Asymmetry XLS-R: Self-supervised Cross-lingual Speech Representation Learning at Scale

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-04T20:44:55.686214Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:44:55.686214Z digest=sha256:e2a8f29dfee3157a0b694591faee321c6eb06de5bbec38e427be76c9125f4425

Observation d2c4190e-0f77-431f-9e6a-f25ddb3fb6be · inbound

AffectDF: The Most Comprehensive Benchmark for Speech Deepfake Detection against Emotionally Expressive Attacks cites this paper.

AffectDF: The Most Comprehensive Benchmark for Speech Deepfake Detection against Emotionally Expressive Attacks XLS-R: Self-supervised Cross-lingual Speech Representation Learning at Scale

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-08T11:50:20.959680Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T11:50:20.959680Z digest=sha256:3b55faf04d42e4deee38d45638042c24fb4e0e7d0bfa4b34a04b02d8c72a46e9