Pith. sign in

Paper Citation Record · LEDGER

Google USM: Scaling Automatic Speech Recognition Beyond 100 Languages

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 52 inbound Pith citation observations for arXiv:2303.01037.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2303.01037 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 52 of 52 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 52 of 52 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-09T05:54:18.845300Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

112
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 6ac8d2d5-b915-4ac9-bc4a-e8f79f515d34 · inbound

AudioPaLM: A Large Language Model That Can Speak and Listen cites this paper.

AudioPaLM: A Large Language Model That Can Speak and Listen Google USM: Scaling Automatic Speech Recognition Beyond 100 Languages

Reference 42

Resolution
malformed identifier
arxiv_id, observed 2026-05-16T07:07:57.909447Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-16T07:07:57.800866Z digest=sha256:72d85175ba5d09c5ba338dfe5f51a0327c7d8353d25f4ec410dcb2e64aa46d6b

Observation 53ccf5ab-dffe-411b-b761-9485a447d772 · inbound

Gemini: A Family of Highly Capable Multimodal Models cites this paper.

Gemini: A Family of Highly Capable Multimodal Models Google USM: Scaling Automatic Speech Recognition Beyond 100 Languages

Reference 132

Resolution
verified exact
arxiv_id, observed 2026-05-24T05:03:55.455193Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-24T05:00:28.453838Z digest=sha256:e11c5e663064e3d01eb84855ce965022d9943d5168a25a784b469fa1d4db3fa3

Observation 85311310-be97-4427-974d-8a767e1eb694 · inbound

Moshi: a speech-text foundation model for real-time dialogue cites this paper.

Moshi: a speech-text foundation model for real-time dialogue Google USM: Scaling Automatic Speech Recognition Beyond 100 Languages

Reference 115

Resolution
verified exact
arxiv_id, observed 2026-05-12T08:13:22.185653Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-12T08:13:21.962488Z digest=sha256:4f31413047deb721c6e66e03372aa579dae05b679dc08cc94d8e361b4b2eb8b9

Observation 541e5b64-f4bd-4067-85eb-02487c24564f · inbound

Small Language Models (SLMs) Can Still Pack a Punch: A survey (updated 2026) cites this paper.

Small Language Models (SLMs) Can Still Pack a Punch: A survey (updated 2026) Google USM: Scaling Automatic Speech Recognition Beyond 100 Languages

Reference 150

Resolution
verified exact
arxiv_id, observed 2026-05-23T05:52:37.677273Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-23T05:47:48.488826Z digest=sha256:addcf7a6e8a349f45c4d61eabf2fd6731569abbdd82ae9c3b32aed544d3b80be

Observation 9f16420d-e4da-4bed-9af3-3f740a5cdd4b · inbound

Metis: A Foundation Speech Generation Model with Masked Generative Pre-training cites this paper.

Metis: A Foundation Speech Generation Model with Masked Generative Pre-training Google USM: Scaling Automatic Speech Recognition Beyond 100 Languages

Reference 106

Resolution
unresolved
no resolver link, observed 2026-08-09T05:54:18.845300Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T05:54:18.845300Z digest=sha256:817b78c6b2ec81a03be5a777dccb95b267cbae1149403d20b33deb6d57542db5

Observation b00cafe7-f8bd-45a9-a2a9-ceaa16be63be · inbound

OWLS: Scaling Laws for Multilingual Speech Recognition and Translation Models cites this paper.

OWLS: Scaling Laws for Multilingual Speech Recognition and Translation Models Google USM: Scaling Automatic Speech Recognition Beyond 100 Languages

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-07T18:32:15.157420Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T18:32:15.157420Z digest=sha256:1a39d9fc90aba713ea940794d307e7a7e3352af8f6175334529b8d0af3378f17

Observation 16597fcf-abb6-4a3c-bea0-5a3f9a1145e4 · inbound

On The Landscape of Spoken Language Models: A Comprehensive Survey cites this paper.

On The Landscape of Spoken Language Models: A Comprehensive Survey Google USM: Scaling Automatic Speech Recognition Beyond 100 Languages

Reference 52

Resolution
metadata mismatch
arxiv_id, observed 2026-05-22T20:45:08.124653Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-22T20:44:57.476464Z digest=sha256:477c93bd80774d9a6f46f8be634a2b094053b91c2624ec52649454a566f69c17

Observation a431a375-596e-4089-90d1-1ab2db9ec0e3 · inbound

Kimi-Audio Technical Report cites this paper.

Kimi-Audio Technical Report Google USM: Scaling Automatic Speech Recognition Beyond 100 Languages

Reference 87

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T19:21:27.168360Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-11T19:21:26.933349Z digest=sha256:76570da9e513bee56c0f8068028b3ce67b9ac24a1f5ef479dcf971994811d4e7

Observation 48b325f5-22a0-4758-bcab-bcf4935c0bb3 · inbound

Scaling and Enhancing LLM-based AVSR: A Sparse Mixture of Projectors Approach cites this paper.

Scaling and Enhancing LLM-based AVSR: A Sparse Mixture of Projectors Approach Google USM: Scaling Automatic Speech Recognition Beyond 100 Languages

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T15:39:29.177686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:39:29.177686Z digest=sha256:16cee5ff53e0b0e2df0940bde39b448591bcb20fa4ff406e7773111e3a2e4430

Observation ec704d4a-f35c-42d3-87e9-b3517eb435d4 · inbound

Selective Invocation for Multilingual ASR: A Cost-effective Approach Adapting to Speech Recognition Difficulty cites this paper.

Selective Invocation for Multilingual ASR: A Cost-effective Approach Adapting to Speech Recognition Difficulty Google USM: Scaling Automatic Speech Recognition Beyond 100 Languages

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T15:10:49.929741Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:10:49.929741Z digest=sha256:b325a79c74c404c6b5d30e2106cab7d2157de18020ee2400b79d68a529b45739

Observation 1a5185df-5743-4224-b270-77f4de0e01cf · inbound

DuRep: Dual-Mode Speech Representation Learning via ASR-Aware Distillation cites this paper.

DuRep: Dual-Mode Speech Representation Learning via ASR-Aware Distillation Google USM: Scaling Automatic Speech Recognition Beyond 100 Languages

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-07T14:12:47.058460Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:12:47.058460Z digest=sha256:4831120a69a717e1a57c29fa6442928d45e1bb1eb9625e81552ce6bb1f0b91e9

Observation 5aa3de75-d092-4a8b-b414-25c96d26b77d · inbound

Leveraging LLM and Self-Supervised Training Models for Speech Recognition in Chinese Dialects: A Comparative Analysis cites this paper.

Leveraging LLM and Self-Supervised Training Models for Speech Recognition in Chinese Dialects: A Comparative Analysis Google USM: Scaling Automatic Speech Recognition Beyond 100 Languages

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T13:45:15.871141Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:45:15.871141Z digest=sha256:b4c2eb081c3e367bf95635a52845570a5bfd1cdd545b566a9a319f883a67f127

Observation 5adc19bc-4bbb-4c92-a428-b302d83d4cb9 · inbound

VietASR: Achieving Industry-level Vietnamese ASR with 50-hour labeled data and Large-Scale Speech Pretraining cites this paper.

VietASR: Achieving Industry-level Vietnamese ASR with 50-hour labeled data and Large-Scale Speech Pretraining Google USM: Scaling Automatic Speech Recognition Beyond 100 Languages

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T14:40:56.281578Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:40:56.281578Z digest=sha256:534dc9f6b29f531d793c291ca9ab7b328dd562df5afd1cc1113c00ec31079ab3

Observation a3224f44-2c51-44de-a819-e6ea0a11d0d2 · inbound

The Warmup Dilemma: How Learning Rate Strategies Impact Speech-to-Text Model Convergence cites this paper.

The Warmup Dilemma: How Learning Rate Strategies Impact Speech-to-Text Model Convergence Google USM: Scaling Automatic Speech Recognition Beyond 100 Languages

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T12:52:39.598918Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:52:39.598918Z digest=sha256:42c22beda66e55a4df4ac0aec9f38c57cb7a4bdac825c2b0c0356cc5b4312539

Observation a49829e8-8114-4c2a-b598-09edd3f5f677 · inbound

OWSM v4: Improving Open Whisper-Style Speech Models via Data Scaling and Cleaning cites this paper.

OWSM v4: Improving Open Whisper-Style Speech Models via Data Scaling and Cleaning Google USM: Scaling Automatic Speech Recognition Beyond 100 Languages

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T12:12:17.790407Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:12:17.790407Z digest=sha256:c045e3e5e47d6e1cd18ef1341b31a97ab9443eb40cf809774a0f5399efedfce0

Observation f1ef71db-56d8-4cd7-bc49-777c03c8b47a · inbound

NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction cites this paper.

NTPP: Generative Speech Language Modeling for Dual-Channel Spoken Dialogue via Next-Token-Pair Prediction Google USM: Scaling Automatic Speech Recognition Beyond 100 Languages

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-07T11:58:42.528594Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:58:42.528594Z digest=sha256:956df1cd2e460b4c38fc01d87acfbc5cf017557100a6e4d77176f536df9dfecb

Observation 43890719-de93-4fd0-bea7-7fb8e72d6b50 · inbound

GigaAM: Efficient Self-Supervised Learner for Speech Recognition cites this paper.

GigaAM: Efficient Self-Supervised Learner for Speech Recognition Google USM: Scaling Automatic Speech Recognition Beyond 100 Languages

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-07T11:55:10.553872Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:55:10.553872Z digest=sha256:b1f173ae47504d105d4413cbd5213455c54dc5c71118dbda275708cddc955129

Observation 4f187dca-064d-4cd8-ab66-bb2d60499661 · inbound

OWSM-Biasing: Contextualizing Open Whisper-Style Speech Models for Automatic Speech Recognition with Dynamic Vocabulary cites this paper.

OWSM-Biasing: Contextualizing Open Whisper-Style Speech Models for Automatic Speech Recognition with Dynamic Vocabulary Google USM: Scaling Automatic Speech Recognition Beyond 100 Languages

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T04:53:57.801970Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:53:57.801970Z digest=sha256:34902a2be6da8ebc4bcd221b4c7e7f1ea0fe3c9d59bf4cfad2fbd645ebf39882

Observation 15f02759-05b4-497e-9bda-c790d208db1e · inbound

A Self-Refining Framework for Enhancing ASR Using TTS-Synthesized Data cites this paper.

A Self-Refining Framework for Enhancing ASR Using TTS-Synthesized Data Google USM: Scaling Automatic Speech Recognition Beyond 100 Languages

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T05:02:43.054744Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:02:43.054744Z digest=sha256:d64094b9dc2def8bed10836051991690932304bff22bae7fc8696cb45f872c5e

Observation 95397d90-3010-4bef-8bdb-cc5a507f6417 · inbound

Early Attentive Sparsification Accelerates Neural Speech Transcription cites this paper.

Early Attentive Sparsification Accelerates Neural Speech Transcription Google USM: Scaling Automatic Speech Recognition Beyond 100 Languages

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T23:50:40.931247Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:50:40.931247Z digest=sha256:f91feaed7b7f54a6c77fc221de722793185cd0e725184c63fa40db8044f3f276

Observation 518fa8b1-c762-4f4b-b67a-bab509c1639b · inbound

Data Quality Issues in Multilingual Speech Datasets: The Need for Sociolinguistic Awareness and Proactive Language Planning cites this paper.

Data Quality Issues in Multilingual Speech Datasets: The Need for Sociolinguistic Awareness and Proactive Language Planning Google USM: Scaling Automatic Speech Recognition Beyond 100 Languages

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-06T23:34:54.366652Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:34:54.366652Z digest=sha256:5ce5b35737dad2fa1ae1d39f8dbf7973a3fc106fa37144530f74ea080034f93d

Observation 4fb310f8-f465-4fad-a2e9-64901f495397 · inbound

Efficient Multilingual ASR Finetuning via LoRA Language Experts cites this paper.

Efficient Multilingual ASR Finetuning via LoRA Language Experts Google USM: Scaling Automatic Speech Recognition Beyond 100 Languages

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T04:53:12.980349Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:53:12.980349Z digest=sha256:005677ccb8c0c95f2911d6a929067bd9cdc6838bbf85ac7f2d83d789fd3487b7

Observation 08cdef0c-3e3c-484e-ab24-2624ceee4be9 · inbound

A Self-Training Approach for Whisper to Enhance Long Dysarthric Speech Recognition cites this paper.

A Self-Training Approach for Whisper to Enhance Long Dysarthric Speech Recognition Google USM: Scaling Automatic Speech Recognition Beyond 100 Languages

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T22:02:27.761816Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:02:27.761816Z digest=sha256:30661d6f5b69b136c652f4bab39a732d3738e644f8dd2ab4da885f0849561929

Observation 5762d61d-65b1-4767-ac42-aa0f67418fbd · inbound

NIRANTAR: Continual Learning with New Languages and Domains on Real-world Speech Data cites this paper.

NIRANTAR: Continual Learning with New Languages and Domains on Real-world Speech Data Google USM: Scaling Automatic Speech Recognition Beyond 100 Languages

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T21:19:28.630761Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:19:28.630761Z digest=sha256:6564074ca70176635a4a088c64c02025f31422f46842adafef436d20e24a2640

Observation cb95bc1c-5fa2-47f9-a954-7e91f7579075 · inbound

OpenBEATs: A Fully Open-Source General-Purpose Audio Encoder cites this paper.

OpenBEATs: A Fully Open-Source General-Purpose Audio Encoder Google USM: Scaling Automatic Speech Recognition Beyond 100 Languages

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T16:13:18.084243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:13:18.084243Z digest=sha256:2b922cc53c6e81307474ad01abc25582f0445df6f1e962b73043595fea4ac58c

Observation e4edf93b-1b17-4d22-8e53-a22e7a2ede58 · inbound

Identifying Hearing Difficulty Moments in Conversational Audio cites this paper.

Identifying Hearing Difficulty Moments in Conversational Audio Google USM: Scaling Automatic Speech Recognition Beyond 100 Languages

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T10:41:18.090704Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T10:41:18.090704Z digest=sha256:92bdc376671f5870c7d3f6aa6c6fba566cb7814ca244b777c6d782510b77570f

Observation 42b609ea-7c56-4341-b4fb-66f173c059f4 · inbound

LCS-CTC: Leveraging Soft Alignments to Enhance Phonetic Transcription Robustness cites this paper.

LCS-CTC: Leveraging Soft Alignments to Enhance Phonetic Transcription Robustness Google USM: Scaling Automatic Speech Recognition Beyond 100 Languages

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T01:10:15.649692Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T01:10:15.649692Z digest=sha256:6b70c7386b379eeaa3d6af9ead7b6a0cd37d2602d3c5ba9c0fbfc384e9d1206b

Observation 09c581e2-35bf-45f6-987c-1413ba893253 · inbound

Geolocation-Aware Robust Spoken Language Identification cites this paper.

Geolocation-Aware Robust Spoken Language Identification Google USM: Scaling Automatic Speech Recognition Beyond 100 Languages

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-05T17:06:13.310210Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:06:13.310210Z digest=sha256:2cde7aec73645b4d1aee4a9507e4c7347f65d436edbe11567973c5dcb103e6b4

Observation 320834ef-cbe0-4226-9681-7cc23b4bb6c9 · inbound

CAM\~OES: A Comprehensive Automatic Speech Recognition Benchmark for European Portuguese cites this paper.

CAM\~OES: A Comprehensive Automatic Speech Recognition Benchmark for European Portuguese Google USM: Scaling Automatic Speech Recognition Beyond 100 Languages

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-05T15:39:21.275265Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:39:21.275265Z digest=sha256:ec4258cc69b358822a4791654b1a662c4e63ed7d935081f331216dc3dba1ef24

Observation 7c66f95f-54e2-4129-8f2b-fcf6e94d283b · inbound

OLMoASR: Open Models and Data for Training Robust Speech Recognition Models cites this paper.

OLMoASR: Open Models and Data for Training Robust Speech Recognition Models Google USM: Scaling Automatic Speech Recognition Beyond 100 Languages

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-05T14:49:35.963931Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T14:49:35.963931Z digest=sha256:64d23e24e289ea2f80f2928a94d30ecf91521c4f15101df009e5e6ce5f0f1c7a

Observation 1fddbe4f-57f1-4a2c-8618-9ce03aa9efe8 · inbound

Towards Improved Speech Recognition through Optimized Synthetic Data Generation cites this paper.

Towards Improved Speech Recognition through Optimized Synthetic Data Generation Google USM: Scaling Automatic Speech Recognition Beyond 100 Languages

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-05T14:11:07.418384Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:11:07.418384Z digest=sha256:e9b396717f0a565ad3acb0e13302dc74b724ebc6d03d297a3a4d4a5e33290ef4

Observation 5cd6c1fe-12f6-46d5-9058-6cbcc06f8e41 · inbound

UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models cites this paper.

UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models Google USM: Scaling Automatic Speech Recognition Beyond 100 Languages

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-04T11:29:37.769261Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T11:29:37.769261Z digest=sha256:1ccc24f8ab0dcf47c5936a69656c183679646e1563742e0d6050844ec4321365

Observation e296ebd2-d033-4311-9d0f-5337db128193 · inbound

BEST-RQ-Based Self-Supervised Learning for Whisper Domain Adaptation cites this paper.

BEST-RQ-Based Self-Supervised Learning for Whisper Domain Adaptation Google USM: Scaling Automatic Speech Recognition Beyond 100 Languages

Reference 16

Resolution
metadata mismatch
arxiv_id, observed 2026-05-18T03:00:48.629522Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-18T02:57:03.681175Z digest=sha256:36dad6a7a362aa705609f68e8bb8346d832e2cd94893027e8fb2aa5ec9863fc8

Observation d90a938c-6437-4c1b-a526-e633096ec97a · inbound

A Study of Data Selection Strategies for Pre-training Self-Supervised Speech Models cites this paper.

A Study of Data Selection Strategies for Pre-training Self-Supervised Speech Models Google USM: Scaling Automatic Speech Recognition Beyond 100 Languages

Reference 13

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T10:17:43.819953Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-16T10:17:15.726726Z digest=sha256:ee91c6d7a633298cd1aab5f8633f26e83aa2e958cd44961d5b16e8043240a9ea

Observation 28af6dfc-fdeb-403b-a4aa-f7575db52074 · inbound

A Semi-spontaneous Dutch Speech Dataset for Speech Enhancement and Speech Recognition cites this paper.

A Semi-spontaneous Dutch Speech Dataset for Speech Enhancement and Speech Recognition Google USM: Scaling Automatic Speech Recognition Beyond 100 Languages

Reference 39

Resolution
unresolved
no resolver link, observed 2026-07-15T00:03:31.986628Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T00:03:31.986628Z digest=sha256:f7d8732465636846bb3ea0f301aef0c751f7969e5134f217732aa2f9b0667d4b

Observation 36f3c996-20b4-449c-88be-33a3d04dd6cb · inbound

PoM: A Linear-Time Replacement for Attention with the Polynomial Mixer cites this paper.

PoM: A Linear-Time Replacement for Attention with the Polynomial Mixer Google USM: Scaling Automatic Speech Recognition Beyond 100 Languages

Reference 80

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T23:05:49.268917Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T19:20:52.810648Z digest=sha256:1f54ae85f82d8605efd4d9ca3d6463380a407854f3b785fcadba55e427dd2882

Observation aab95778-cb03-4e83-9b20-f35d9465e0ba · inbound

BlasBench: An Open Benchmark for Irish Speech Recognition cites this paper.

BlasBench: An Open Benchmark for Irish Speech Recognition Google USM: Scaling Automatic Speech Recognition Beyond 100 Languages

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-11T09:41:01.511355Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-10T15:53:54.092426Z digest=sha256:548ddb5e5fdf38ba94e5f1a37cc31ab535bf09f49bd235a2348c84e58e0f92dd

Observation 48146589-1f80-4847-9813-463959d8def0 · inbound

UniPASE: A Generative Model for Universal Speech Enhancement with High Fidelity and Low Hallucinations cites this paper.

UniPASE: A Generative Model for Universal Speech Enhancement with High Fidelity and Low Hallucinations Google USM: Scaling Automatic Speech Recognition Beyond 100 Languages

Reference 28

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T10:19:19.994947Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T10:18:16.972414Z digest=sha256:e6d6e8fba62595614e144530e863dd599ab581ea25d79c7be5f9a4812ae5f9d3

Observation b772d669-7687-4bd8-9823-9a58468dbd54 · inbound

UniPASE: A Generative Model for Universal Speech Enhancement with High Fidelity and Low Hallucinations cites this paper.

UniPASE: A Generative Model for Universal Speech Enhancement with High Fidelity and Low Hallucinations Google USM: Scaling Automatic Speech Recognition Beyond 100 Languages

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-02T16:16:36.937293Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T16:16:36.937293Z digest=sha256:f383cdfcee799ef00ba3063292bac68f247115dcf78035ffb81d2e4355f8dc4b

Observation 109d954a-66ba-4253-b6d7-496cf7303ac3 · inbound

Dolphin-CN-Dialect: Where Chinese Dialects Matter cites this paper.

Dolphin-CN-Dialect: Where Chinese Dialects Matter Google USM: Scaling Automatic Speech Recognition Beyond 100 Languages

Reference 19

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T02:16:16.155299Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-12T02:13:42.268229Z digest=sha256:fe5096d72ec268a39d378b12b4b7c5e587d70211deef04a8de83fa616286e19e

Observation a09ac056-8755-42b1-9c2e-d00b44e30b7f · inbound

MegaScale-Omni: A Hyper-Scale, Workload-Resilient System for MultiModal LLM Training in Production cites this paper.

MegaScale-Omni: A Hyper-Scale, Workload-Resilient System for MultiModal LLM Training in Production Google USM: Scaling Automatic Speech Recognition Beyond 100 Languages

Reference 62

Resolution
verified exact
arxiv_id, observed 2026-05-12T02:06:14.936678Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-12T02:04:07.344134Z digest=sha256:f6083640ac826b4a9449f795fb2cf859cc305faaae139ad6532955667db14780

Observation 1abdcf3b-a987-42ad-9141-03dcbf4c7396 · inbound

Mind the Gap: Impact of Synthetic Conversational Data on Multi-Talker ASR and Speaker Diarization cites this paper.

Mind the Gap: Impact of Synthetic Conversational Data on Multi-Talker ASR and Speaker Diarization Google USM: Scaling Automatic Speech Recognition Beyond 100 Languages

Reference 13

Resolution
metadata mismatch
arxiv_id, observed 2026-05-19T14:42:37.457490Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-19T14:39:27.382652Z digest=sha256:a47051178781267a74469b38974b01644a5ae77f6fcd9b9c348e651f813c1337

Observation f0ee6108-8585-47a5-8092-0652071699b3 · inbound

FalAR: A Large-scale Speaker-Annotated European Portuguese Speech Corpus of Parliamentary Sessions cites this paper.

FalAR: A Large-scale Speaker-Annotated European Portuguese Speech Corpus of Parliamentary Sessions Google USM: Scaling Automatic Speech Recognition Beyond 100 Languages

Reference 24

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T18:13:48.543485Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-29T18:08:58.631775Z digest=sha256:bc0cf081b1df275075c6afa020c9b1a09322f87ad54bf7af0217ceb28e7a92f1

Observation 0537089c-ee9c-4be9-b377-e695f6d48450 · inbound

Online Predictive Coding for Dual-Mode Self-Supervised Speech Model cites this paper.

Online Predictive Coding for Dual-Mode Self-Supervised Speech Model Google USM: Scaling Automatic Speech Recognition Beyond 100 Languages

Reference 11

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T07:39:38.778504Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-26T13:15:04.922107Z digest=sha256:2683496f22403d5f8fae134a68b92328b673b4f4578452659995345f55baf707

Observation 4b6e0059-ec66-4566-9414-6175ea0bd66e · inbound

ESPnet3: Infrastructure for Scalable Speech and Audio Research in the Foundation Model Era cites this paper.

ESPnet3: Infrastructure for Scalable Speech and Audio Research in the Foundation Model Era Google USM: Scaling Automatic Speech Recognition Beyond 100 Languages

Reference 12

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T08:19:44.216294Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-26T11:54:08.457573Z digest=sha256:206471867ecb2b59629fc9b01c4dbe9954509aac8a50e4c3239cccfe0e744875

Observation 2be08545-17ae-4b07-9ba5-c0be6153dbc1 · inbound

Data Scale, Not Latency, Shapes Cross-Lingual Encoder Transfer in Streaming ASR cites this paper.

Data Scale, Not Latency, Shapes Cross-Lingual Encoder Transfer in Streaming ASR Google USM: Scaling Automatic Speech Recognition Beyond 100 Languages

Reference 7

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T16:49:57.378232Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-26T00:15:33.685313Z digest=sha256:29cebb05a190cd49fa27fc9de80f4b8693be0d57c4526be4fbcc4d520b02071f

Observation 32d94fb4-7f1f-42e8-a9f6-547dc20c817f · inbound

Comparing Human and Automatic Recognition of Dutch Dysarthric Continuous Speech: A Case Study cites this paper.

Comparing Human and Automatic Recognition of Dutch Dysarthric Continuous Speech: A Case Study Google USM: Scaling Automatic Speech Recognition Beyond 100 Languages

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T06:34:18.648828Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-30T06:33:01.811696Z digest=sha256:a0fa3fc90eb03fd9318774cacfbe07e23abc6d42bd68ff6c3730ad54f02feb75

Observation 1036fbf8-ea9a-4572-91ef-7a2975d37618 · inbound

Enhancing BEST-RQ Pseudo-Label Quality through Online Refinement for Automatic Speech Recognition cites this paper.

Enhancing BEST-RQ Pseudo-Label Quality through Online Refinement for Automatic Speech Recognition Google USM: Scaling Automatic Speech Recognition Beyond 100 Languages

Reference 19

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T06:55:29.107996Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-01T06:51:41.995108Z digest=sha256:1b211cfe9035083bbf84ee4e7fc5d0097538e02cff2e4d337c504518ef17f161

Observation 757e4a3a-12bf-40f3-9509-1fcef7226c86 · inbound

Gemma 4 Technical Report cites this paper.

Gemma 4 Technical Report Google USM: Scaling Automatic Speech Recognition Beyond 100 Languages

Reference 21

Resolution
unresolved
no resolver link, observed 2026-07-12T07:10:06.281832Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T07:10:06.281832Z digest=sha256:0cd7ecfc092b387331587d287f37a365753319465170d4d7ed63e80930a12965

Observation e9f7d5bc-f704-408a-9c4b-c39d313442b4 · inbound

GigaAM Multilingual: Foundation Model for Underrepresented Languages cites this paper.

GigaAM Multilingual: Foundation Model for Underrepresented Languages Google USM: Scaling Automatic Speech Recognition Beyond 100 Languages

Reference 8

Resolution
unresolved
no resolver link, observed 2026-07-14T12:15:06.470439Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T12:15:06.470439Z digest=sha256:0473f8ea019f05b2e52d31cd139de8541422a16733686677bb105177f6eaf675

Observation f3c861d1-95eb-4c31-bbb6-2c589dd7f7b0 · inbound

DONDO: Open w2v-BERT Speech-Recognition Base Models for African Languages cites this paper.

DONDO: Open w2v-BERT Speech-Recognition Base Models for African Languages Google USM: Scaling Automatic Speech Recognition Beyond 100 Languages

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-01T07:10:39.848318Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:10:39.848318Z digest=sha256:2ba03129b6dbceede2f5539a6aac10b443b5d068ad5ca8e5fe04c0beca99d6f6

Observation a2914072-a04c-4beb-880d-1e0681699564 · inbound

MoLGE: Mixture of Language Group Experts for Efficient Scaling of Massively Multilingual Speech Recognition cites this paper.

MoLGE: Mixture of Language Group Experts for Efficient Scaling of Massively Multilingual Speech Recognition Google USM: Scaling Automatic Speech Recognition Beyond 100 Languages

Reference 28

Resolution
unresolved
no resolver link, observed 2026-07-31T23:18:49.234557Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T23:18:49.234557Z digest=sha256:71bdffc26c1967212fab8e938a193abb3bb8599f0fa49e71868a80922b5c4972