Pith. sign in

Paper Citation Record · LEDGER

XLS-R: Self-supervised Cross-lingual Speech Representation Learning at Scale

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 55 inbound Pith citation observations for arXiv:2111.09296.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2111.09296 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 55 of 55 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 55 of 55 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T19:15:52.095800Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 7aaf16de-cdf7-47dd-9d8b-0e43c321bd87 · inbound

Enhancing Expressive Voice Conversion with Discrete Pitch-Conditioned Flow Matching Model cites this paper.

Enhancing Expressive Voice Conversion with Discrete Pitch-Conditioned Flow Matching Model XLS-R: Self-supervised Cross-lingual Speech Representation Learning at Scale

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-08T19:15:52.095800Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T19:15:52.095800Z digest=sha256:fe9e26088b61e9df82d29ffada4ca766f8ba68feae9050f81094f91fc7e97a0c

Observation fb63d79a-1d1c-4c4c-b71f-796f33f5281d · inbound

Recreating Neural Activity During Speech Production with Language and Speech Model Embeddings cites this paper.

Recreating Neural Activity During Speech Production with Language and Speech Model Embeddings XLS-R: Self-supervised Cross-lingual Speech Representation Learning at Scale

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T15:42:19.207531Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:42:19.207531Z digest=sha256:e459c6410569a57a8898873166b873abc976aaa6f9c2327eee8de1a39fe76fd2

Observation 5e4f22e8-cb45-4434-ba3e-6316c3a68fce · inbound

PSRB: A Comprehensive Benchmark for Evaluating Persian ASR Systems cites this paper.

PSRB: A Comprehensive Benchmark for Evaluating Persian ASR Systems XLS-R: Self-supervised Cross-lingual Speech Representation Learning at Scale

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T13:42:54.835853Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:42:54.835853Z digest=sha256:9100b062053267c95117fdad7970e2b77226ef10818dfb3476f2a070f1dd18da

Observation 36098707-1984-449d-af69-2d68b3e84bea · inbound

ZIPA: A family of efficient models for multilingual phone recognition cites this paper.

ZIPA: A family of efficient models for multilingual phone recognition XLS-R: Self-supervised Cross-lingual Speech Representation Learning at Scale

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-07T12:55:39.944247Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:55:39.944247Z digest=sha256:ed921d34854bc74571984bca9058248c6d08acfb2213ebde15de285b72f9c334

Observation 8c6ee2ef-5abc-41d7-a341-6ff54fb01f02 · inbound

Few-Shot Speech Deepfake Detection Adaptation with Gaussian Processes cites this paper.

Few-Shot Speech Deepfake Detection Adaptation with Gaussian Processes XLS-R: Self-supervised Cross-lingual Speech Representation Learning at Scale

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T12:45:29.628714Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:45:29.628714Z digest=sha256:0bd5c17a2f484111921e9ed80ba91dc326c5a8d1fb4711a665908591f21b76ae

Observation ad37c472-53a9-44e6-baed-4e1850d22a49 · inbound

Assortment of Attention Heads: Accelerating Federated PEFT with Head Pruning and Strategic Client Selection cites this paper.

Assortment of Attention Heads: Accelerating Federated PEFT with Head Pruning and Strategic Client Selection XLS-R: Self-supervised Cross-lingual Speech Representation Learning at Scale

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T12:05:51.843072Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:05:51.843072Z digest=sha256:05cfde47ee0182cbb53169d955eb5e290e81e55cd164ef17fe32aa186c575937

Observation 38d25a94-b079-4b34-b82c-8025f1e96b67 · inbound

Rhythm Controllable and Efficient Zero-Shot Voice Conversion via Shortcut Flow Matching cites this paper.

Rhythm Controllable and Efficient Zero-Shot Voice Conversion via Shortcut Flow Matching XLS-R: Self-supervised Cross-lingual Speech Representation Learning at Scale

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T11:57:52.905967Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:57:52.905967Z digest=sha256:880785d26e6597ad9e6a655324362437b895e45978c02c27dc1436b36489881f

Observation df060ada-c158-40dd-a58e-c297ca8c74c8 · inbound

Towards Generalized Source Tracing for Codec-Based Deepfake Speech cites this paper.

Towards Generalized Source Tracing for Codec-Based Deepfake Speech XLS-R: Self-supervised Cross-lingual Speech Representation Learning at Scale

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T05:43:08.358475Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:43:08.358475Z digest=sha256:ea2dcaa3ac2aa1865f3bed2a3ea9db6ba4dc36020b1d7b135ff594b300ec1da4

Observation 97d95895-7355-4b31-83b6-d50e6d76d3c4 · inbound

Joint ASR and Speaker Role Tagging with Serialized Output Training cites this paper.

Joint ASR and Speaker Role Tagging with Serialized Output Training XLS-R: Self-supervised Cross-lingual Speech Representation Learning at Scale

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T04:32:30.227311Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:32:30.227311Z digest=sha256:3c9aa87d0216402a65a203bfd25570ad7811ce3f3e677bae1b9303a8b86f4d11

Observation e892ec6f-7acc-49c8-a6fa-fde74123eb15 · inbound

From Sharpness to Better Generalization for Speech Deepfake Detection cites this paper.

From Sharpness to Better Generalization for Speech Deepfake Detection XLS-R: Self-supervised Cross-lingual Speech Representation Learning at Scale

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T04:08:06.077547Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:08:06.077547Z digest=sha256:c59bd376882efcb82f4b01e01dc16067968ee5147903875a83e624015b11b98e

Observation ff93de16-16f8-419b-9e1a-54ae05522f3d · inbound

Pushing the Performance of Synthetic Speech Detection with Kolmogorov-Arnold Networks and Self-Supervised Learning Models cites this paper.

Pushing the Performance of Synthetic Speech Detection with Kolmogorov-Arnold Networks and Self-Supervised Learning Models XLS-R: Self-supervised Cross-lingual Speech Representation Learning at Scale

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T00:22:26.715238Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:22:26.715238Z digest=sha256:7924a41cd42a1ab405b9a97048bc093549255ea7e4a63f2803cd282da4bbf5fc

Observation 4ba1af66-387e-4ddc-bbca-aa4840297d0d · inbound

Breaking the Transcription Bottleneck: Fine-tuning ASR Models for Extremely Low-Resource Fieldwork Languages cites this paper.

Breaking the Transcription Bottleneck: Fine-tuning ASR Models for Extremely Low-Resource Fieldwork Languages XLS-R: Self-supervised Cross-lingual Speech Representation Learning at Scale

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T23:35:28.171349Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:35:28.171349Z digest=sha256:8cbd71d17590882099574c57a94707763bfaec2bede189c226c76ef6caf8198b

Observation bacce3da-fc5d-4cf5-bbbf-bfeef09b4d93 · inbound

Data Quality Issues in Multilingual Speech Datasets: The Need for Sociolinguistic Awareness and Proactive Language Planning cites this paper.

Data Quality Issues in Multilingual Speech Datasets: The Need for Sociolinguistic Awareness and Proactive Language Planning XLS-R: Self-supervised Cross-lingual Speech Representation Learning at Scale

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T23:34:49.009463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:34:49.009463Z digest=sha256:f016dd3c451898f2e4424060f4a88cb3ae1f27ca6443ec8be5bce4897a3ca455

Observation 045ac136-f4e3-4f8b-8dce-fd38e6707afd · inbound

Thinking Beyond Tokens: From Brain-Inspired Intelligence to Cognitive Foundations for Artificial General Intelligence and its Societal Impact cites this paper.

Thinking Beyond Tokens: From Brain-Inspired Intelligence to Cognitive Foundations for Artificial General Intelligence and its Societal Impact XLS-R: Self-supervised Cross-lingual Speech Representation Learning at Scale

Reference 159

Resolution
unresolved
no resolver link, observed 2026-08-06T21:07:11.272341Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:07:11.272341Z digest=sha256:a1e52c44e896cab7ba045b25c10514066a0652e77ad2d1043e1e81b34a7e7991

Observation dad648f9-9aa1-4934-b17b-0668ee6d121e · inbound

DiceHuBERT: Distilling HuBERT with a Self-Supervised Learning Objective cites this paper.

DiceHuBERT: Distilling HuBERT with a Self-Supervised Learning Objective XLS-R: Self-supervised Cross-lingual Speech Representation Learning at Scale

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T23:02:14.927874Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:02:14.927874Z digest=sha256:94d40be70bc463c1708bcb7dcd32a41f647081926e8cb4892b8927655acea96b

Observation 845ae894-6d3c-49bf-8e4e-53e6d3f9e739 · inbound

Word stress in self-supervised speech models: A cross-linguistic comparison cites this paper.

Word stress in self-supervised speech models: A cross-linguistic comparison XLS-R: Self-supervised Cross-lingual Speech Representation Learning at Scale

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T19:44:26.885657Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:44:26.885657Z digest=sha256:1bbb92cd115ea4cd6835f0ba912b4e214de6d90ee1b2ebe5dac4ee3216fdd9ec

Observation 6910d54b-696a-4824-9f91-cfc860254baf · inbound

RepeaTTS: Towards Feature Discovery through Repeated Fine-Tuning cites this paper.

RepeaTTS: Towards Feature Discovery through Repeated Fine-Tuning XLS-R: Self-supervised Cross-lingual Speech Representation Learning at Scale

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T20:01:35.492150Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:01:35.492150Z digest=sha256:2b2e7e479a622ecd1a13a9d78a86a2a1ba6bc89479c2f483636f6c109008b9d7

Observation 1991a958-5633-48f9-ac24-634cc4e0fab1 · inbound

On Barriers to Archival Audio Processing cites this paper.

On Barriers to Archival Audio Processing XLS-R: Self-supervised Cross-lingual Speech Representation Learning at Scale

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T18:14:08.779769Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:14:08.779769Z digest=sha256:adbc2365af7a12a67210211c218bcdbbe271398d2aa2c86831dd7a0e61c9a64f

Observation f1c21fd1-9842-447b-8a1c-14db27d29faa · inbound

SpeechFake: A Large-Scale Multilingual Speech Deepfake Dataset Incorporating Cutting-Edge Generation Methods cites this paper.

SpeechFake: A Large-Scale Multilingual Speech Deepfake Dataset Incorporating Cutting-Edge Generation Methods XLS-R: Self-supervised Cross-lingual Speech Representation Learning at Scale

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T12:49:16.380954Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T12:49:16.380954Z digest=sha256:037c0413e469f981b69def35870c04ab9fd6f6e168b2c3f3658a2396d0197665

Observation ddfdcf68-6bbb-418c-92ca-b7e5ef60aa85 · inbound

CAM\~OES: A Comprehensive Automatic Speech Recognition Benchmark for European Portuguese cites this paper.

CAM\~OES: A Comprehensive Automatic Speech Recognition Benchmark for European Portuguese XLS-R: Self-supervised Cross-lingual Speech Representation Learning at Scale

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-05T15:39:21.516370Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:39:21.516370Z digest=sha256:144e3e0ad03413fca3c938e274340eb958752ca9c6ad66a6a79dc3962b2b1775

Observation b1e5f998-c1c8-4785-91e5-dd8b66b957c6 · inbound

Generalizable Audio Spoofing Detection using Non-Semantic Representations cites this paper.

Generalizable Audio Spoofing Detection using Non-Semantic Representations XLS-R: Self-supervised Cross-lingual Speech Representation Learning at Scale

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-05T13:55:18.277810Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:55:18.277810Z digest=sha256:0c3f01f8680b5b9af52ae73e5efec11d7528a43f389ba430e031c5aa58c9279c

Observation b423e264-8805-408e-8884-c868e8f5ad10 · inbound

Forensic Similarity for Speech Deepfakes cites this paper.

Forensic Similarity for Speech Deepfakes XLS-R: Self-supervised Cross-lingual Speech Representation Learning at Scale

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-18T10:31:14.902664Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-18T10:28:12.467606Z digest=sha256:89fd750832d276f8700411d522a0f492ec4e192e572c1292671b87f3267f9202

Observation 7ea9d4ba-0fda-47c3-ada1-5f2e3c64aae3 · inbound

SONAR: Spectral-Contrastive Audio Residuals for Generalizable Deepfake Detection cites this paper.

SONAR: Spectral-Contrastive Audio Residuals for Generalizable Deepfake Detection XLS-R: Self-supervised Cross-lingual Speech Representation Learning at Scale

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-03T20:05:29.992627Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:05:29.992627Z digest=sha256:03bc4268262b3848a5eef0dcd0653aeed5ed70dae5ec6e9f7c9e81a012ba238a

Observation 98b8ee19-d213-48f5-9aa8-f181363ba2b3 · inbound

A SUPERB-Style Benchmark of Self-Supervised Speech Models for Audio Deepfake Detection cites this paper.

A SUPERB-Style Benchmark of Self-Supervised Speech Models for Audio Deepfake Detection XLS-R: Self-supervised Cross-lingual Speech Representation Learning at Scale

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-15T17:26:22.179791Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-15T17:25:27.137423Z digest=sha256:70982b27b225c55624e46f8c54b08563038cd1affa2fa2b5e17c1d54c375f603

Observation be6fd803-7145-4683-a19c-101dd423dd21 · inbound

Benchmarking Multilingual Speech Models on Pashto: Zero-Shot ASR, Script Failure, and Cross-Domain Evaluation cites this paper.

Benchmarking Multilingual Speech Models on Pashto: Zero-Shot ASR, Script Failure, and Cross-Domain Evaluation XLS-R: Self-supervised Cross-lingual Speech Representation Learning at Scale

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-10T22:35:48.700833Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T19:44:30.762851Z digest=sha256:fc9c978aaaa2f14c87698f5a94afaa61365d454efee4f139ece0a4eca68468d6

Observation db07feee-6407-4425-8621-09c452858f2c · inbound

Giving Voice to the Constitution: Low-Resource Text-to-Speech for Quechua and Spanish Using a Bilingual Legal Corpus cites this paper.

Giving Voice to the Constitution: Low-Resource Text-to-Speech for Quechua and Spanish Using a Bilingual Legal Corpus XLS-R: Self-supervised Cross-lingual Speech Representation Learning at Scale

Reference 7

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T11:21:00.993257Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T15:01:20.902986Z digest=sha256:a8e4fdc0949b8d51bf59869ca954ee00fa487cf32c3f3b955ae5a9ec34c0bba0

Observation dbb3df5e-639c-4e03-a9f6-4c0a3e50176d · inbound

Similarity Choice and Negative Scaling in Supervised Contrastive Learning for Deepfake Audio Detection cites this paper.

Similarity Choice and Negative Scaling in Supervised Contrastive Learning for Deepfake Audio Detection XLS-R: Self-supervised Cross-lingual Speech Representation Learning at Scale

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-12T08:46:25.862041Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-07T13:50:53.553261Z digest=sha256:42a538a5a88ab762fbabe418cb8019dd26a5d689267bb3a3fe60a0896b0011e8

Observation e98638b2-ad8c-42d9-8f32-79e8ac4438e6 · inbound

Profiling the Voice: Speaker-Specific Phoneme Fingerprinting for Speech Deepfake Detection cites this paper.

Profiling the Voice: Speaker-Specific Phoneme Fingerprinting for Speech Deepfake Detection XLS-R: Self-supervised Cross-lingual Speech Representation Learning at Scale

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-20T01:22:55.817247Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-20T01:20:04.609575Z digest=sha256:3854bd505b34feff76ff55de88a72b330455c1b4c294ac082201c13d9afdebeb

Observation cce6bf73-5148-4a1d-b5a5-29cb6cbcb3d4 · inbound

MixFake: Benchmarking and Enhancing Audio Deepfake Detection in Diverse Real-world Mixed Audio cites this paper.

MixFake: Benchmarking and Enhancing Audio Deepfake Detection in Diverse Real-world Mixed Audio XLS-R: Self-supervised Cross-lingual Speech Representation Learning at Scale

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-25T03:15:17.023264Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-25T03:14:08.880072Z digest=sha256:f5445008d47140a4529c15d2b904fc6584dab98c5c6e48da1a23beeb65354b2e

Observation ddaf9454-c96e-42c0-b5fe-c0ab99677897 · inbound

Escaping the Linearity Trap: Manifold Detours for Black-Box Adversarial Attacks on Singing Audio Deepfake Detection cites this paper.

Escaping the Linearity Trap: Manifold Detours for Black-Box Adversarial Attacks on Singing Audio Deepfake Detection XLS-R: Self-supervised Cross-lingual Speech Representation Learning at Scale

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-06-30T19:05:01.112227Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-30T18:58:31.318883Z digest=sha256:75ff6c5d119545cccbc381704a676cdf6dd713454819b7619f39af9e7cc44e07

Observation 92c547c6-ca7c-4952-967d-fb9b26bfd232 · inbound

Dual-Branch Gated Fusion for Open-Set Audio Deepfake Source Tracing cites this paper.

Dual-Branch Gated Fusion for Open-Set Audio Deepfake Source Tracing XLS-R: Self-supervised Cross-lingual Speech Representation Learning at Scale

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-07-03T03:47:35.572305Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-27T14:45:45.359616Z digest=sha256:553a028071681316c7523162bd494ee38240a5957ba9b33ad0f671769c678fea

Observation f4260727-f010-4480-8756-d5f2a3dee562 · inbound

Pretrained self-supervised speech models can recognize unseen consonants cites this paper.

Pretrained self-supervised speech models can recognize unseen consonants XLS-R: Self-supervised Cross-lingual Speech Representation Learning at Scale

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-07-03T10:07:56.070600Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-27T10:14:47.932613Z digest=sha256:00bba196411521f418e0808924d92d0755eca42cda75863a19de0dbfc77f62cc

Observation cbe71276-4f07-4fdb-8176-4a33c86cc761 · inbound

SpAArSIST: Sparsified AASIST for Efficient and Reliable Anti-Spoofing cites this paper.

SpAArSIST: Sparsified AASIST for Efficient and Reliable Anti-Spoofing XLS-R: Self-supervised Cross-lingual Speech Representation Learning at Scale

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-07-03T12:58:08.668281Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-27T08:35:24.321197Z digest=sha256:3ec13835c29930e7fe15062d02f1a19d58c1b63410c45c78aa5976f61af2ab1c

Observation 18ed6e2d-fcd1-4e0b-9178-354240c9df74 · inbound

Responsible ASR: Overcoming Challenges of Foundational Models in Narrow-Band and Low-Resource Settings cites this paper.

Responsible ASR: Overcoming Challenges of Foundational Models in Narrow-Band and Low-Resource Settings XLS-R: Self-supervised Cross-lingual Speech Representation Learning at Scale

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-07-04T01:59:26.658181Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-26T20:03:28.491546Z digest=sha256:b7830810918d82bbacd3ff291133abeea890209d7c88ebf166e54e8965388d79

Observation dd75c079-7ec4-4819-9099-12cd136d5123 · inbound

Light-weight Pronunciation Assessment via Discrete Speech Token Surprisal cites this paper.

Light-weight Pronunciation Assessment via Discrete Speech Token Surprisal XLS-R: Self-supervised Cross-lingual Speech Representation Learning at Scale

Reference 60

Resolution
verified exact
arxiv_id, observed 2026-07-04T03:49:30.231851Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-26T17:37:25.043607Z digest=sha256:9fb47b53ebdfea98a8b0907d97cf3a6cde799525c30ef0dc5d7f674530c67513

Observation dedb92c1-9f15-4ce5-8432-63b950b9a2fe · inbound

Beyond Speaker Independence: Evaluating Cross-Lingual Acoustic-to-Articulatory Inversion Across Finnish and Russian cites this paper.

Beyond Speaker Independence: Evaluating Cross-Lingual Acoustic-to-Articulatory Inversion Across Finnish and Russian XLS-R: Self-supervised Cross-lingual Speech Representation Learning at Scale

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-07-04T05:49:37.763673Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-26T15:24:04.594993Z digest=sha256:984c5b0babab4548569d53e3fd840519d932fc9b37fda176a27dceb90461d1f5

Observation e82d9116-99fc-4ed9-b1a5-7b70cbc47ce5 · inbound

Impact Analysis of Speech Representation Learning Models for Acoustic Side-Channel Attack cites this paper.

Impact Analysis of Speech Representation Learning Models for Acoustic Side-Channel Attack XLS-R: Self-supervised Cross-lingual Speech Representation Learning at Scale

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-07-04T06:39:37.331175Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-26T14:19:59.573391Z digest=sha256:c65e48b8d912d48f69ad3fd8110de93d2ec4a816ed47d2f3a6c6658b5af2cbc7

Observation fbff2c09-d503-4660-b16d-78a7dabb35c4 · inbound

Impact Analysis of Speech Representation Learning Models for Acoustic Side-Channel Attack cites this paper.

Impact Analysis of Speech Representation Learning Models for Acoustic Side-Channel Attack XLS-R: Self-supervised Cross-lingual Speech Representation Learning at Scale

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-06-29T19:33:54.406475Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-29T04:44:52.537405Z digest=sha256:338f3bb97c500eaf57b2e456b3c13e4b982897b6167b906322f2c6681e4dba44

Observation 61adb2a8-dc23-40af-a1f2-68848c557b19 · inbound

From Speech to Text Corpora: Evaluating ASR-Based Data Acquisition for Low-Resource Fongbe and Hausa cites this paper.

From Speech to Text Corpora: Evaluating ASR-Based Data Acquisition for Low-Resource Fongbe and Hausa XLS-R: Self-supervised Cross-lingual Speech Representation Learning at Scale

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-07-04T08:29:42.395446Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-26T11:31:09.747113Z digest=sha256:4d161984d798059b43515793d9a4115e6bbba742254b4974f7306432feb40f25

Observation 5d125a69-5efe-41d1-b65a-f8de9b2f395c · inbound

What Do Deepfake Benchmarks Measure? An Audit Using Frozen Self-Supervised Representations cites this paper.

What Do Deepfake Benchmarks Measure? An Audit Using Frozen Self-Supervised Representations XLS-R: Self-supervised Cross-lingual Speech Representation Learning at Scale

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-07-04T15:49:57.200451Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-26T01:24:46.200846Z digest=sha256:98708a55e57206f7206f218985f8880b6de514d9e3c8c5181df7f543b625844a

Observation 2a86893c-fa6b-413d-9c74-b2607ed4d8f7 · inbound

Syntactic Belief Update as the Driver of Garden Path Processing Difficulty cites this paper.

Syntactic Belief Update as the Driver of Garden Path Processing Difficulty XLS-R: Self-supervised Cross-lingual Speech Representation Learning at Scale

Reference 295

Resolution
metadata mismatch
arxiv_id, observed 2026-06-26T04:38:58.354104Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-26T04:38:01.183423Z digest=sha256:fa5f3ef7c159f93ab45d33684a858ab421c75b657e16c5b51744fe0f0abca9e1

Observation 5e46740a-119f-4c26-ac54-9e04ccec34c6 · inbound

Detecting Audio Deepfakes on the Edge:Lightweight SSL-Based Detection in a Browser Plugin cites this paper.

Detecting Audio Deepfakes on the Edge:Lightweight SSL-Based Detection in a Browser Plugin XLS-R: Self-supervised Cross-lingual Speech Representation Learning at Scale

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-07-01T12:45:44.602950Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-01T01:47:41.684335Z digest=sha256:403ca92962382656b3a796d3c5a1571f575e8b496d6541f792ace54ad8fc4802

Observation abbf0d63-cf32-495b-a9cb-570350d25ac4 · inbound

From Objectives to Applications: Aligning Architectural Biases in Audio Self-Supervised Learning cites this paper.

From Objectives to Applications: Aligning Architectural Biases in Audio Self-Supervised Learning XLS-R: Self-supervised Cross-lingual Speech Representation Learning at Scale

Reference 93

Resolution
verified exact
arxiv_id, observed 2026-07-02T05:56:39.846458Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-02T05:52:55.818877Z digest=sha256:a807e85dc9f8e7fe539105af05bcbb3ee734d79c57d15f06817803c6f98be8b3

Observation 5a231095-8343-43cd-9b28-7b84916812e4 · inbound

Speaker-Aware Temporal Aggregation Strategies on Segment Representations for Depression Detection in Dyadic Interaction: A Benchmark Study cites this paper.

Speaker-Aware Temporal Aggregation Strategies on Segment Representations for Depression Detection in Dyadic Interaction: A Benchmark Study XLS-R: Self-supervised Cross-lingual Speech Representation Learning at Scale

Reference 10

Resolution
unresolved
no resolver link, observed 2026-07-12T06:14:42.321767Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T06:14:42.321767Z digest=sha256:37511ec332d2a6f1b4892c1ba5f97dc4f518b6a9e28169d087d0866b57f90f6d

Observation 59a919b8-fb23-4828-adb2-3d3995836e7c · inbound

An Intervention-Based Framework for Shortcut Diagnosis in Spoofing Countermeasures cites this paper.

An Intervention-Based Framework for Shortcut Diagnosis in Spoofing Countermeasures XLS-R: Self-supervised Cross-lingual Speech Representation Learning at Scale

Reference 18

Resolution
unresolved
no resolver link, observed 2026-07-12T04:31:11.771594Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T04:31:11.771594Z digest=sha256:1aaf78e173ad23318a6dfef13904fc6128bdab659ea60f5072908027a486053a

Observation 5265bd2a-3338-4e30-b7ee-b6840c656f53 · inbound

Towards Digital Preservation of Efik: TTS for a Low-Resource African Language cites this paper.

Towards Digital Preservation of Efik: TTS for a Low-Resource African Language XLS-R: Self-supervised Cross-lingual Speech Representation Learning at Scale

Reference 36

Resolution
unresolved
no resolver link, observed 2026-07-11T18:11:36.626189Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T18:11:36.626189Z digest=sha256:a27a916282afae1a3f1102e754a2caf79b0f0a04a877b8980537babe4d96061c

Observation 83eed9c1-9c22-4fb8-b782-3ff5b7d21eb3 · inbound

Evaluating the Effect of Linguistic Relatedness on Cross-Lingual Transfer in Large Multilingual Automatic Speech Recognition cites this paper.

Evaluating the Effect of Linguistic Relatedness on Cross-Lingual Transfer in Large Multilingual Automatic Speech Recognition XLS-R: Self-supervised Cross-lingual Speech Representation Learning at Scale

Reference 2

Resolution
unresolved
no resolver link, observed 2026-07-11T13:07:13.175134Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T13:07:13.175134Z digest=sha256:81a05eafc65fd7829fb044578f127fa7898b0917227b97142343bbdcdcad659d

Observation 087998ef-6c92-4171-8258-2dacc3719058 · inbound

Evaluating the Effect of Linguistic Relatedness on Cross-Lingual Transfer in Large Multilingual Automatic Speech Recognition cites this paper.

Evaluating the Effect of Linguistic Relatedness on Cross-Lingual Transfer in Large Multilingual Automatic Speech Recognition XLS-R: Self-supervised Cross-lingual Speech Representation Learning at Scale

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-02T08:34:55.605107Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:34:55.605107Z digest=sha256:ecee23613a10c9938623ebe5e152e2635698d6aa5edaf24be6b056493c652661

Observation 2d8c6bc4-790e-437b-acb7-3688e45872b3 · inbound

GigaAM Multilingual: Foundation Model for Underrepresented Languages cites this paper.

GigaAM Multilingual: Foundation Model for Underrepresented Languages XLS-R: Self-supervised Cross-lingual Speech Representation Learning at Scale

Reference 17

Resolution
unresolved
no resolver link, observed 2026-07-14T12:15:06.470439Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T12:15:06.470439Z digest=sha256:0a3848ceb97dc5c1dce8188a765a945ce09179a23df1f8fa065586673d57326a

Observation 0b197b28-1724-4dbe-9a52-d2703304a7d9 · inbound

Unified Gradient Projection: Language-Balanced Continual Learning for Multilingual Low-Resource ASR cites this paper.

Unified Gradient Projection: Language-Balanced Continual Learning for Multilingual Low-Resource ASR XLS-R: Self-supervised Cross-lingual Speech Representation Learning at Scale

Reference 8

Resolution
unresolved
no resolver link, observed 2026-07-14T06:34:55.089754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T06:34:55.089754Z digest=sha256:a19ed681d48c7083fbeb0dd45ab69cdcd675902a52ee883448e16c7a5da30c15

Observation 397ebc1a-c572-4365-b947-30a876530d83 · inbound

DONDO: Open w2v-BERT Speech-Recognition Base Models for African Languages cites this paper.

DONDO: Open w2v-BERT Speech-Recognition Base Models for African Languages XLS-R: Self-supervised Cross-lingual Speech Representation Learning at Scale

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-01T07:10:39.837818Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:10:39.837818Z digest=sha256:58123987cbbd5b79f03934f4a9beb933a9fe4d49222fe3aab144d89f9c96754c

Observation af26c9ad-3b23-4deb-be16-701e9507960f · inbound

Leveraging Gradient Reversal Loss and Multitask Learning for Datasets-Aware Audio Deepfake Detection cites this paper.

Leveraging Gradient Reversal Loss and Multitask Learning for Datasets-Aware Audio Deepfake Detection XLS-R: Self-supervised Cross-lingual Speech Representation Learning at Scale

Reference 33

Resolution
unresolved
no resolver link, observed 2026-07-31T23:27:02.078077Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T23:27:02.078077Z digest=sha256:99985e489e08effe1c26e4cb2f078154be19e6b58c7359d1dec92975962ab240

Observation 46196a50-20d5-4b09-b76d-b2dd78425e91 · inbound

Teffic-Audio: Tell Fact from Fiction cites this paper.

Teffic-Audio: Tell Fact from Fiction XLS-R: Self-supervised Cross-lingual Speech Representation Learning at Scale

Reference 15

Resolution
unresolved
no resolver link, observed 2026-07-31T10:28:22.087327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T10:28:22.087327Z digest=sha256:153278f6c99577335a333a9bb4d629928265a4b24e65973ae6325a2bdadbb53d

Observation d97e4775-c32e-4e8e-a2a1-0ac095920b20 · inbound

Multi-Backbone Self-Supervised Ensembles for Audio Deepfake Detection and a Cross-Track Analysis of Generation-Detection Asymmetry cites this paper.

Multi-Backbone Self-Supervised Ensembles for Audio Deepfake Detection and a Cross-Track Analysis of Generation-Detection Asymmetry XLS-R: Self-supervised Cross-lingual Speech Representation Learning at Scale

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-04T20:44:55.686214Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:44:55.686214Z digest=sha256:e2a8f29dfee3157a0b694591faee321c6eb06de5bbec38e427be76c9125f4425

Observation d2c4190e-0f77-431f-9e6a-f25ddb3fb6be · inbound

AffectDF: The Most Comprehensive Benchmark for Speech Deepfake Detection against Emotionally Expressive Attacks cites this paper.

AffectDF: The Most Comprehensive Benchmark for Speech Deepfake Detection against Emotionally Expressive Attacks XLS-R: Self-supervised Cross-lingual Speech Representation Learning at Scale

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-08T11:50:20.959680Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T11:50:20.959680Z digest=sha256:2ed663c50af87783a3d1d6aae73dcca4a1b8165dac2f2e08f790a80adaaff568