Pith. sign in

Paper Citation Record · LEDGER

Whale: Large-Scale multilingual ASR model with w2v-BERT and E-Branchformer with large speech data

As of 8 August 2026, this Paper Citation Record lists 39 of 39 outbound references and 1 inbound Pith citation observation for arXiv:2506.01439.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.01439 v1

Coverage vector

measured 39 of 39 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:52:13.429143Z

measured 40 of 40 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:52:11.502278Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-07T11:52:13.916425Z

Reference resolution

39 of 39 outbound references displayed

  • verified exact3
  • verified fuzzy24
  • unresolved10
  • parse uncertain0
  • malformed identifier2
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation a252876d-78d3-4a80-840d-ecb6af7f9330 · outbound

This paper cites Whale: Large-Scale multilingual ASR model with w2v-BERT and E-Branchformer with large speech data.

Whale: Large-Scale multilingual ASR model with w2v-BERT and E-Branchformer with large speech data Whale: Large-Scale multilingual ASR model with w2v-BERT and E-Branchformer with large speech data

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-08-07T11:52:13.943649Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:52:11.502278Z digest=sha256:a5d1cf65bd499274ebebb4f8c3fcec1f90044b1b91597686d063dab0e17fd3fe

Observation 0f77035d-0147-4b16-baad-e1fd041d2959 · outbound

This paper cites Next, acoustic features are extracted via SSL, w2v-BERT [18].

Whale: Large-Scale multilingual ASR model with w2v-BERT and E-Branchformer with large speech data Next, acoustic features are extracted via SSL, w2v-BERT [18]

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:52:16.978017Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:52:11.547934Z digest=sha256:bea9cc9b12823c36a7e366e8476e43672c4c34166d7a054ef469b185f4034be3

Observation 2f9d5ede-8e03-4a08-92ef-592195b8e6d6 · outbound

This paper cites Training environments The training of the Whale model was conducted on an inter- nal server infrastructure.

Whale: Large-Scale multilingual ASR model with w2v-BERT and E-Branchformer with large speech data Training environments The training of the Whale model was conducted on an inter- nal server infrastructure

Reference 3

Resolution
malformed identifier
raw_fallback, observed 2026-08-07T11:52:16.830710Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:52:11.605647Z digest=sha256:214580bc7aacfd8fac8357a3342e26840ba96c90f0661ae52b2919892fad2798

Observation 3433c08c-7cc0-4712-b862-592fbe1d45be · outbound

This paper cites Our experiments compare Whale against state-of-the-art systems such as Whis- per [14], OWSM [15], and OWSM CTC [17].

Whale: Large-Scale multilingual ASR model with w2v-BERT and E-Branchformer with large speech data Our experiments compare Whale against state-of-the-art systems such as Whis- per [14], OWSM [15], and OWSM CTC [17]

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:52:16.736939Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:52:11.643182Z digest=sha256:e5428c9ec2c9ea4e9d5d3106426c18a7c6baa40dc8ca040d49311bc6f3dd906e

Observation 1417886d-095d-4671-90cd-e6b61244d221 · outbound

This paper cites Through extensive experiments, we demonstrated that Whale achieves highly competitive performance on four benchmarks.

Whale: Large-Scale multilingual ASR model with w2v-BERT and E-Branchformer with large speech data Through extensive experiments, we demonstrated that Whale achieves highly competitive performance on four benchmarks

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:52:16.636351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:52:11.720267Z digest=sha256:60c56495388c3969d336dddd40c2aac91ce01b638e3b089c97066cd396c8c87e

Observation bc8a5bb4-9708-468d-9cd0-81121f30004f · outbound

This paper cites Common V oice: A Massively-Multilingual Speech Corpus,.

Whale: Large-Scale multilingual ASR model with w2v-BERT and E-Branchformer with large speech data Common V oice: A Massively-Multilingual Speech Corpus,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:52:16.513195Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:52:11.799504Z digest=sha256:4c933d70aa7868afd691ec55868b853b202aa3be64cf06f9047a13703a3f2043

Observation 9a9c02f2-0d4f-409a-88b5-5148b6a2e0c6 · outbound

This paper cites MUST-C: a multilingual speech translation corpus,.

Whale: Large-Scale multilingual ASR model with w2v-BERT and E-Branchformer with large speech data MUST-C: a multilingual speech translation corpus,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:52:16.435266Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:52:11.841152Z digest=sha256:63b03619147a349be345dc784521d3573f0080bb7fc271ecfa1e66b0fe49d612

Observation 2f0bcc1e-e966-47a4-af83-e2e4a12150fe · outbound

This paper cites The Multilingual TEDx Corpus for Speech Recognition and Translation,.

Whale: Large-Scale multilingual ASR model with w2v-BERT and E-Branchformer with large speech data The Multilingual TEDx Corpus for Speech Recognition and Translation,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:52:16.351072Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:52:11.886282Z digest=sha256:630a5a71786011106ac7b4f5485b2b83aebab15d3e9b3ff142921516967a771f

Observation f1adea2b-c5bf-46b2-bfae-55fb7b7d90ea · outbound

This paper cites MLS: A Large-Scale Multilingual Dataset for Speech Research,.

Whale: Large-Scale multilingual ASR model with w2v-BERT and E-Branchformer with large speech data MLS: A Large-Scale Multilingual Dataset for Speech Research,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:52:16.249441Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:52:11.943874Z digest=sha256:7a0fb3104f7d21f03d269702988dd6f343fd5c79f71517457ebb4ece282a2614

Observation 98fe1142-8d90-4b5c-9864-114a16c15be5 · outbound

This paper cites YODAS: YouTube-oriented dataset for audio and speech,.

Whale: Large-Scale multilingual ASR model with w2v-BERT and E-Branchformer with large speech data YODAS: YouTube-oriented dataset for audio and speech,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:52:16.131035Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:52:11.987325Z digest=sha256:0cdf8f05931ec920c435d2a3d510031cfe00bbf03dccd5545c16df390a3c0373

Observation 6b362eb3-980c-4b06-9e84-72d49be6ac16 · outbound

This paper cites FLEURS: Few-shot learn- ing evaluation of universal representations of speech,.

Whale: Large-Scale multilingual ASR model with w2v-BERT and E-Branchformer with large speech data FLEURS: Few-shot learn- ing evaluation of universal representations of speech,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:52:16.043916Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:52:12.044617Z digest=sha256:9775555ce81e91bffbc275331390be9a6fd655b3b49cb390b33570930beab557

Observation 73f4e7b9-83cf-4188-a662-2f5856a3235e · outbound

This paper cites JTubeSpeech: corpus of Japanese speech collected from YouTube for speech recognition and speaker verification.

Whale: Large-Scale multilingual ASR model with w2v-BERT and E-Branchformer with large speech data JTubeSpeech: corpus of Japanese speech collected from YouTube for speech recognition and speaker verification

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T11:52:12.107639Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:52:12.107639Z digest=sha256:e4e6f4207094201898afb801c740e352c9be75778ccd5d8c2dff5be7395a7d75

Observation 650a4083-b23e-48c1-b995-12d781bcf2ab · outbound

This paper cites CSTR VCTK Corpus: English Multi-speaker Corpus for CSTR V oice Cloning Toolkit,.

Whale: Large-Scale multilingual ASR model with w2v-BERT and E-Branchformer with large speech data CSTR VCTK Corpus: English Multi-speaker Corpus for CSTR V oice Cloning Toolkit,

Reference 13

Resolution
malformed identifier
doi_truncated, observed 2026-08-07T11:52:13.540449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:52:12.188880Z digest=sha256:ff46f53103806f3bb9c8d15a9137579f602370762f760af031088fd69e934583

Observation 78e257a2-8c00-4ad1-b187-59b56b892abb · outbound

This paper cites Multilingual speech recognition with a single end-to-end model,.

Whale: Large-Scale multilingual ASR model with w2v-BERT and E-Branchformer with large speech data Multilingual speech recognition with a single end-to-end model,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:52:15.911834Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:52:12.220385Z digest=sha256:b9570f6b475c22bd3ad39fbe996f6b4c6f66d694424c40b17ba91691c064f006

Observation fa894e38-556e-40b8-99e1-a2f9d11da06f · outbound

This paper cites Bytes are all you need: End-to-end multilingual speech recognition and synthe- sis with bytes,.

Whale: Large-Scale multilingual ASR model with w2v-BERT and E-Branchformer with large speech data Bytes are all you need: End-to-end multilingual speech recognition and synthe- sis with bytes,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:52:15.841719Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:52:12.267009Z digest=sha256:68de3feec924332daaf3d873dd82ed40addf7202d4d439ef550eef5e1e4b9908

Observation 5bfcfb12-120f-4947-b09e-cd170152f762 · outbound

This paper cites An end-to-end language-tracking speech recognizer for mixed- language speech,.

Whale: Large-Scale multilingual ASR model with w2v-BERT and E-Branchformer with large speech data An end-to-end language-tracking speech recognizer for mixed- language speech,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:52:15.766281Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:52:12.301077Z digest=sha256:0dec677bc370cfcf3a8a2887a98c48d1d9fb2186c0135eb83f1a1affa91ffa98

Observation 311203d9-ad28-40e3-9bf6-0d263da0384b · outbound

This paper cites Scaling speech technology to 1,000+ languages,.

Whale: Large-Scale multilingual ASR model with w2v-BERT and E-Branchformer with large speech data Scaling speech technology to 1,000+ languages,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:52:15.610423Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:52:12.341225Z digest=sha256:77a6b366ccbb7bf631bd52942d2740aaa4cc3c5c4811185bc48d9896f5811219

Observation 2563d107-58fd-442a-a3a4-0dafbc238378 · outbound

This paper cites Less is More: Accu- rate Speech Recognition & Translation without Web-Scale Data,.

Whale: Large-Scale multilingual ASR model with w2v-BERT and E-Branchformer with large speech data Less is More: Accu- rate Speech Recognition & Translation without Web-Scale Data,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:52:15.517326Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:52:12.393041Z digest=sha256:23d64b692450c675268b29876b27b361513695ac8eabfcb639e0972b494a67b2

Observation 3265a550-d7cc-4745-8430-3c2bfaacfb83 · outbound

This paper cites Robust speech recognition via large-scale weak supervision,.

Whale: Large-Scale multilingual ASR model with w2v-BERT and E-Branchformer with large speech data Robust speech recognition via large-scale weak supervision,

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T11:52:12.424757Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:52:12.424757Z digest=sha256:cfde9396d75fd4df8650b4cbdaee1b9afd1ec67bd04358807c2f7a632ef7991c

Observation 605a18c1-61fa-42f9-94ba-d87ffb96b2bc · outbound

This paper cites OWSM v3.1: Better and Faster Open Whisper-Style Speech Models based on E-Branchformer.

Whale: Large-Scale multilingual ASR model with w2v-BERT and E-Branchformer with large speech data OWSM v3.1: Better and Faster Open Whisper-Style Speech Models based on E-Branchformer

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T11:52:12.485510Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:52:12.485510Z digest=sha256:ffed31c28a02176930503853c837e6dd6e795df559bc6bcee0a435fe02326c81

Observation fc862e2b-0f58-424c-aa00-ae35ff0b686e · outbound

This paper cites Reproducing whisper-style train- ing using an open-source toolkit and publicly available data,.

Whale: Large-Scale multilingual ASR model with w2v-BERT and E-Branchformer with large speech data Reproducing whisper-style train- ing using an open-source toolkit and publicly available data,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:52:15.403774Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:52:12.552905Z digest=sha256:bb977ed6aa469452fa63fce6d1f86fe4648bba8d9cb652a9b51faeff1475dc55

Observation 50267203-235d-4053-ad36-6890cc280f96 · outbound

This paper cites OWSM-CTC: An Open Encoder-Only Speech Foundation Model for Speech Recognition, Translation, and Language Identification.

Whale: Large-Scale multilingual ASR model with w2v-BERT and E-Branchformer with large speech data OWSM-CTC: An Open Encoder-Only Speech Foundation Model for Speech Recognition, Translation, and Language Identification

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T11:52:12.584687Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:52:12.584687Z digest=sha256:9d6af5e1e87188e70836e027712cfb900f304186456e6423a412bf7a2f6fe101

Observation f10c69b4-4f34-4a7e-b299-239bd01d81d7 · outbound

This paper cites W2v-BERT: Combin- ing contrastive learning and masked language modeling for self- supervised speech pre-training,.

Whale: Large-Scale multilingual ASR model with w2v-BERT and E-Branchformer with large speech data W2v-BERT: Combin- ing contrastive learning and masked language modeling for self- supervised speech pre-training,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:52:15.277095Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:52:12.616796Z digest=sha256:e741c2865531e3c581dba5893439e60f74021892ee0aa6c5347d76d0704c2311

Observation 392efdd9-04e1-45a4-a835-f1609dda7824 · outbound

This paper cites E-Branchformer: Branchformer with enhanced merging for speech recognition,.

Whale: Large-Scale multilingual ASR model with w2v-BERT and E-Branchformer with large speech data E-Branchformer: Branchformer with enhanced merging for speech recognition,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:52:15.070525Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:52:12.663181Z digest=sha256:bade0f65de50dd096c34c8d2ddbda9269949fc91a0b8c6b8b12ef4fbc8d25e2b

Observation 76402194-0a96-4528-ae06-4065c2867c11 · outbound

This paper cites Joint CTC-attention based end-to-end speech recognition using multi-task learning,.

Whale: Large-Scale multilingual ASR model with w2v-BERT and E-Branchformer with large speech data Joint CTC-attention based end-to-end speech recognition using multi-task learning,

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T11:52:12.696914Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:52:12.696914Z digest=sha256:55f2a6b0795029eaf94b27bfe66c198447d0d9e4a67f189a32ba9b681a69a3c1

Observation fe796955-e6fc-4b6c-9fa6-174890ae8a31 · outbound

This paper cites Joint CTC/attention decoding for end-to-end speech recognition,.

Whale: Large-Scale multilingual ASR model with w2v-BERT and E-Branchformer with large speech data Joint CTC/attention decoding for end-to-end speech recognition,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:52:14.920584Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:52:12.731205Z digest=sha256:2b658935b509509b977412ab638095445d5d42de68ae510eaa6c1c54c959fc21

Observation f7885cbb-39de-4366-a5d4-ac178db87a6d · outbound

This paper cites Curricu- lum learning,.

Whale: Large-Scale multilingual ASR model with w2v-BERT and E-Branchformer with large speech data Curricu- lum learning,

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T11:52:12.768004Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:52:12.768004Z digest=sha256:bd60d119c3e299fa1261838a69bf58da43daa81a498a86707b296da91fd542d5

Observation bfe986cd-b310-4754-9f27-fa52c31e4d7e · outbound

This paper cites Lib- rispeech: an asr corpus based on public domain audio books,.

Whale: Large-Scale multilingual ASR model with w2v-BERT and E-Branchformer with large speech data Lib- rispeech: an asr corpus based on public domain audio books,

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T11:52:12.818208Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:52:12.818208Z digest=sha256:95ac4c397302c756ec08b034043d3db7d359b3dd197acc372bc323181f62b12d

Observation e35690d9-1d5c-4bae-9702-28be96da4503 · outbound

This paper cites Corpus of Spontaneous Japanese: Its design and evaluation,.

Whale: Large-Scale multilingual ASR model with w2v-BERT and E-Branchformer with large speech data Corpus of Spontaneous Japanese: Its design and evaluation,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:52:14.805916Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:52:12.868330Z digest=sha256:b58764e05dbfa8bf681803d8ae360bf89b9b6d37a88470506ff9c478212965c2

Observation 3697c5d3-97ab-4c57-95d5-efef034997a1 · outbound

This paper cites Better Intermediates Im- prove CTC Inference,.

Whale: Large-Scale multilingual ASR model with w2v-BERT and E-Branchformer with large speech data Better Intermediates Im- prove CTC Inference,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:52:14.691919Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:52:12.914300Z digest=sha256:864a6f91e4927b675819948f8a5f1dbc4eb107bcf02868cbd34a87e2a6338e03

Observation 20437515-acb6-4b91-b9d7-5f37dc3b25b5 · outbound

This paper cites BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding.

Whale: Large-Scale multilingual ASR model with w2v-BERT and E-Branchformer with large speech data BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T11:52:12.946413Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:52:12.946413Z digest=sha256:bfcac4ad572ed12bdbf5e7094259ae10c02e41adb0b784c7f2d7dc78a774837f

Observation 99580925-0466-4f65-8f2b-596a39776f07 · outbound

This paper cites mHuBERT-147: A Compact Multilingual HuBERT Model.

Whale: Large-Scale multilingual ASR model with w2v-BERT and E-Branchformer with large speech data mHuBERT-147: A Compact Multilingual HuBERT Model

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T11:52:13.029649Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:52:13.029649Z digest=sha256:d35d19e17aa9fbe2c04100ea0b7bd82b73406f116dc06b100df87efd9c4754af

Observation 1b7f548d-d52f-445d-a734-7b8eab3ee7e0 · outbound

This paper cites WavLM: Large-scale self- supervised pre-training for full stack speech processing,.

Whale: Large-Scale multilingual ASR model with w2v-BERT and E-Branchformer with large speech data WavLM: Large-scale self- supervised pre-training for full stack speech processing,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:52:14.569887Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:52:13.100295Z digest=sha256:910609c9c3c33e52b5c3e0fdcf8fb90ece174c1fad4b5575594edd43763ad9ad

Observation ea4af69f-db0d-4996-b02e-a97bae4e6ee7 · outbound

This paper cites Rapid Language Adaptation for Multilingual E2E Speech Recognition Using Encoder Prompting.

Whale: Large-Scale multilingual ASR model with w2v-BERT and E-Branchformer with large speech data Rapid Language Adaptation for Multilingual E2E Speech Recognition Using Encoder Prompting

Reference 34

Resolution
verified exact
local_arxiv, observed 2026-08-07T11:52:13.768823Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:52:13.181595Z digest=sha256:e7288f1204a8a83ee40b5d7bdcb80141c554cd391f2f279e18dfcdcccc17a0ff

Observation 90e0db55-7011-4431-a972-edaf52a54976 · outbound

This paper cites ESPnet: End-to-End Speech Processing Toolkit,.

Whale: Large-Scale multilingual ASR model with w2v-BERT and E-Branchformer with large speech data ESPnet: End-to-End Speech Processing Toolkit,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:52:14.425893Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:52:13.244378Z digest=sha256:e76737520243c898a9ece1ee28cdc507f364fe5d5a604d23230f87e5d57eb78a

Observation 04f43e6e-ec3e-42fd-b475-676edde6c505 · outbound

This paper cites VoxPopuli: A large-scale multilingual speech corpus for representation learning, semi- supervised learning and interpretation,.

Whale: Large-Scale multilingual ASR model with w2v-BERT and E-Branchformer with large speech data VoxPopuli: A large-scale multilingual speech corpus for representation learning, semi- supervised learning and interpretation,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:52:14.303262Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:52:13.283183Z digest=sha256:1948109f640172bc1c2da719a8d0308d057eac2e9402cc4bef87799cfb3ee759

Observation ab385ad2-7bb1-4b5f-8866-364ceb3501c4 · outbound

This paper cites WenetSpeech: A 10000+ hours multi-domain mandarin corpus for speech recognition,.

Whale: Large-Scale multilingual ASR model with w2v-BERT and E-Branchformer with large speech data WenetSpeech: A 10000+ hours multi-domain mandarin corpus for speech recognition,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:52:14.075308Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:52:13.348188Z digest=sha256:b1719cde7919f22e91bd853feaba847e325ff9d13ac9def6c88cc7987e8c55d5

Observation 7b443f9f-0199-4440-9394-983dbf2e8ece · outbound

This paper cites The People's Speech: A Large-Scale Diverse English Speech Recognition Dataset for Commercial Usage.

Whale: Large-Scale multilingual ASR model with w2v-BERT and E-Branchformer with large speech data The People's Speech: A Large-Scale Diverse English Speech Recognition Dataset for Commercial Usage

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T11:52:13.381090Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:52:13.381090Z digest=sha256:6f86abfd13d4084873ad21ddf4183fefd293b50ad47192b8616515a5b92d1a22

Observation 91d8f7c5-bb41-4732-a8d5-9f66d7a0a411 · outbound

This paper cites The Norwegian Parliamentary Speech Corpus.

Whale: Large-Scale multilingual ASR model with w2v-BERT and E-Branchformer with large speech data The Norwegian Parliamentary Speech Corpus

Reference 39

Resolution
verified exact
local_arxiv, observed 2026-08-07T11:52:13.658600Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:52:13.429143Z digest=sha256:0ac4cf0717857630950491b31d3b7c33a3334b692ac5696e871ca778b59d3de7

Pith citing papers

Observation a252876d-78d3-4a80-840d-ecb6af7f9330 · inbound

Whale: Large-Scale multilingual ASR model with w2v-BERT and E-Branchformer with large speech data cites this paper.

Whale: Large-Scale multilingual ASR model with w2v-BERT and E-Branchformer with large speech data Whale: Large-Scale multilingual ASR model with w2v-BERT and E-Branchformer with large speech data

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-08-07T11:52:13.943649Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T11:52:11.502278Z digest=sha256:a5d1cf65bd499274ebebb4f8c3fcec1f90044b1b91597686d063dab0e17fd3fe