Pith. sign in

Paper Citation Record · LEDGER

OWSM v4: Improving Open Whisper-Style Speech Models via Data Scaling and Cleaning

As of 17 August 2026, this Paper Citation Record lists 47 of 47 outbound references and 1 inbound Pith citation observation for arXiv:2506.00338.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.00338 v1

Coverage vector

measured 47 of 47 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:12:21.174309Z

measured 48 of 48 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:12:17.177212Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-07T12:12:21.729943Z

Reference resolution

47 of 47 outbound references displayed

  • verified exact0
  • verified fuzzy34
  • unresolved10
  • parse uncertain0
  • malformed identifier3
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation da98eed8-8609-4fd1-87f4-9a9d2c0507e2 · outbound

This paper cites OWSM v4: Improving Open Whisper-Style Speech Models via Data Scaling and Cleaning.

OWSM v4: Improving Open Whisper-Style Speech Models via Data Scaling and Cleaning OWSM v4: Improving Open Whisper-Style Speech Models via Data Scaling and Cleaning

Reference 1

Resolution
malformed identifier
local_arxiv, observed 2026-08-07T12:12:21.828937Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T12:12:17.177212Z digest=sha256:6b5e2a8d0e906eb7ce2c8c2a93988a67bdfaf36678baefc1760f0c9382d39179

Observation 8fecf34e-4ab5-4486-8dbe-b28d4b53b12e · outbound

This paper cites YODAS data cleaning The raw YODAS data has not undergone a rigorous cleaning process and may contain annotation errors [24].

OWSM v4: Improving Open Whisper-Style Speech Models via Data Scaling and Cleaning YODAS data cleaning The raw YODAS data has not undergone a rigorous cleaning process and may contain annotation errors [24]

Reference 2

Resolution
malformed identifier
raw_fallback, observed 2026-08-07T12:12:27.493239Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T12:12:17.271724Z digest=sha256:bc4cc78c76422deecc5c3a63b9c2f0372cbd54c43107aa72bb705977e9475e1f

Observation 62ffcb80-8de8-42e8-be83-0757cdfc274f · outbound

This paper cites transcription.

OWSM v4: Improving Open Whisper-Style Speech Models via Data Scaling and Cleaning transcription

Reference 3

Resolution
malformed identifier
raw_fallback, observed 2026-08-07T12:12:21.617665Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T12:12:17.393419Z digest=sha256:6ef6dbe21a1471e1b2f38a541a27186397bb6c1c544d48f4bf89b5a74583c021

Observation e9ce723c-d9be-4b7c-b87b-6e61c7bcaf36 · outbound

This paper cites We reveal that large-scale web-crawled data contains incorrect lan- guage labels and audio-text misalignments.

OWSM v4: Improving Open Whisper-Style Speech Models via Data Scaling and Cleaning We reveal that large-scale web-crawled data contains incorrect lan- guage labels and audio-text misalignments

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:12:27.321548Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T12:12:17.504498Z digest=sha256:71f36c3f7f91206319411560d9b50c74d524f049be47e2b30cc0b0ea09309d18

Observation fc45377e-b74f-4575-826c-4b2c2bab28f7 · outbound

This paper cites an unresolved cited work.

OWSM v4: Improving Open Whisper-Style Speech Models via Data Scaling and Cleaning Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-07T12:12:27.141555Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T12:12:17.586769Z digest=sha256:0ef1a5ce7b3bf3c14ff5296fd80b44822bb269cb0cc3d398640d3dba419b329e

Observation f81a6e80-6bf0-4300-ad75-9e2560fae4d9 · outbound

This paper cites Robust speech recognition via large-scale weak supervision,.

OWSM v4: Improving Open Whisper-Style Speech Models via Data Scaling and Cleaning Robust speech recognition via large-scale weak supervision,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:12:26.954012Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T12:12:17.675610Z digest=sha256:1190502a2f428577e998894f6d384af0de74165e0525adb49521cb57ff563dc9

Observation a49829e8-8114-4c2a-b598-09edd3f5f677 · outbound

This paper cites Google USM: Scaling Automatic Speech Recognition Beyond 100 Languages.

OWSM v4: Improving Open Whisper-Style Speech Models via Data Scaling and Cleaning Google USM: Scaling Automatic Speech Recognition Beyond 100 Languages

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T12:12:17.790407Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:12:17.790407Z digest=sha256:a7b35f83ce62927b2ce1f36db9785ac99127ee493aa1f26fe44a1810668470e7

Observation 1c293df4-4f20-4953-b75d-8f948b50f06e · outbound

This paper cites Scaling speech technology to 1,000+ languages,.

OWSM v4: Improving Open Whisper-Style Speech Models via Data Scaling and Cleaning Scaling speech technology to 1,000+ languages,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:12:26.785010Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T12:12:17.904554Z digest=sha256:91c507a6324e5bc2f55f5429a4b4b3a49817a7764d85420749e6a1429b5ce5b9

Observation cfbe3c66-a058-4b60-9191-a84e9b7324c4 · outbound

This paper cites Less is more: Accurate speech recognition & translation without web- scale data,.

OWSM v4: Improving Open Whisper-Style Speech Models via Data Scaling and Cleaning Less is more: Accurate speech recognition & translation without web- scale data,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:12:26.633111Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T12:12:17.996141Z digest=sha256:01fe2f9af6a4d5945a16744f6264f6c07f9fea09f423cb60888389ec4bed73c5

Observation 538be67f-567f-4d89-9806-f9d4c5892108 · outbound

This paper cites Reproducing Whisper-Style Training Using an Open-Source Toolkit and Pub- licly Available Data,.

OWSM v4: Improving Open Whisper-Style Speech Models via Data Scaling and Cleaning Reproducing Whisper-Style Training Using an Open-Source Toolkit and Pub- licly Available Data,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:12:26.501683Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T12:12:18.091981Z digest=sha256:40c5669ce5f42722b89f22c2c06eec004fecfbd0971faeb4c91403a90901ff91

Observation cef51312-f6bb-4e3d-9d3e-b61f35ee3ab5 · outbound

This paper cites ESPnet: End- to-End Speech Processing Toolkit,.

OWSM v4: Improving Open Whisper-Style Speech Models via Data Scaling and Cleaning ESPnet: End- to-End Speech Processing Toolkit,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:12:26.377366Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T12:12:18.171426Z digest=sha256:05e1d247a8efd7d1253b0ac9dad0ae87fa0e5ecd310f04ed8222e909884e5c4b

Observation 42b5f007-b1a7-4cbb-b9b5-7db73be9266f · outbound

This paper cites Conformer: Convolution-augmented Transformer for Speech Recognition,.

OWSM v4: Improving Open Whisper-Style Speech Models via Data Scaling and Cleaning Conformer: Convolution-augmented Transformer for Speech Recognition,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:12:26.201127Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T12:12:18.265022Z digest=sha256:bca1b85bbb41788c315153b786ace70e66975693854ef6ddf952adabefd6b2f5

Observation efcccd6c-06d7-4d53-b4eb-9218918c66d4 · outbound

This paper cites Branchformer: Parallel MLP-attention architectures to capture local and global context for speech recognition and understanding,.

OWSM v4: Improving Open Whisper-Style Speech Models via Data Scaling and Cleaning Branchformer: Parallel MLP-attention architectures to capture local and global context for speech recognition and understanding,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:12:25.994805Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T12:12:18.358460Z digest=sha256:2f3d2aaf92e46b810eaace663f57c97561d2ab617d3abd67f77cbeb564ba3798

Observation d8abf4c3-f8bf-4c5c-b159-f1501bae89d4 · outbound

This paper cites Zipformer: A faster and better encoder for automatic speech recognition,.

OWSM v4: Improving Open Whisper-Style Speech Models via Data Scaling and Cleaning Zipformer: A faster and better encoder for automatic speech recognition,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:12:25.745680Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T12:12:18.435408Z digest=sha256:0caa49722f58e0fddc2d790029d0c9dfe43f1030802b07956a16d73d3895d9df

Observation cb75094c-5bfc-457f-8ba9-344d1fe37679 · outbound

This paper cites Atten- tion is all you need,.

OWSM v4: Improving Open Whisper-Style Speech Models via Data Scaling and Cleaning Atten- tion is all you need,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:12:25.603809Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T12:12:18.504149Z digest=sha256:b0853e76073b4be9eefced03067441d9921e1d83e6bde2db307bcf715d29b2b1

Observation 9d6c92d4-8967-4486-bcf5-337b9e8732a6 · outbound

This paper cites Squeezeformer: An efficient transformer for automatic speech recognition,.

OWSM v4: Improving Open Whisper-Style Speech Models via Data Scaling and Cleaning Squeezeformer: An efficient transformer for automatic speech recognition,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:12:25.467750Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T12:12:18.599148Z digest=sha256:e3489ac2f5fa4eaf1b9cd94835348ba492b2b9cb93d0e020202fa1451692d91d

Observation fc6f5a37-a3b3-41f0-8d97-7f35e083183d · outbound

This paper cites Fast conformer with linearly scalable attention for efficient speech recognition,.

OWSM v4: Improving Open Whisper-Style Speech Models via Data Scaling and Cleaning Fast conformer with linearly scalable attention for efficient speech recognition,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:12:25.321074Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T12:12:18.726297Z digest=sha256:7b531147f1d8a284e16aedf2af6cd06d4f9a164fa9dffd831c8ff502d6392b2a

Observation 86c09436-cba7-4031-b368-fd2da1233b7f · outbound

This paper cites Sum- maryMixing: A linear-complexity alternative to self-attention for speech recognition and understanding,.

OWSM v4: Improving Open Whisper-Style Speech Models via Data Scaling and Cleaning Sum- maryMixing: A linear-complexity alternative to self-attention for speech recognition and understanding,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:12:25.178642Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T12:12:18.814410Z digest=sha256:d3c439845e28b3e1f8c6a678e52f8527a796651a4893477cd975b1e506a887a0

Observation 21568b9a-d2a0-425e-819e-d272968b80f9 · outbound

This paper cites OWSM v3.1: Bet- ter and faster open whisper-style speech models based on E- Branchformer,.

OWSM v4: Improving Open Whisper-Style Speech Models via Data Scaling and Cleaning OWSM v3.1: Bet- ter and faster open whisper-style speech models based on E- Branchformer,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:12:25.047851Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T12:12:18.852042Z digest=sha256:abfe899a04126bca74557a35512fe488f9757ff46340147d8eb55cb0acb59c65

Observation 837310f8-d203-4448-84fb-80b4cf0d328f · outbound

This paper cites E-Branchformer: Branch- former with enhanced merging for speech recognition,.

OWSM v4: Improving Open Whisper-Style Speech Models via Data Scaling and Cleaning E-Branchformer: Branch- former with enhanced merging for speech recognition,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:12:24.905555Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T12:12:18.855538Z digest=sha256:96747dda06b439f81a2bd6e05785f6aaa5dd02b60c30ed413d7b7e55876e5392

Observation bfcef2ea-205f-4f2c-abe1-7026a12b9920 · outbound

This paper cites A Comparative Study on E-Branchformer vs Conformer in Speech Recognition, Transla- tion, and Understanding Tasks,.

OWSM v4: Improving Open Whisper-Style Speech Models via Data Scaling and Cleaning A Comparative Study on E-Branchformer vs Conformer in Speech Recognition, Transla- tion, and Understanding Tasks,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:12:24.764470Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T12:12:18.888232Z digest=sha256:21b6567846d62b7eedec596e847bd2b21a0f2c9446f05ca9fbcdeddf407d48be

Observation 8ca5f411-4012-44a5-9995-68ca9fbad818 · outbound

This paper cites OWSM-CTC: An open encoder-only speech foundation model for speech recognition, translation, and language identification,.

OWSM v4: Improving Open Whisper-Style Speech Models via Data Scaling and Cleaning OWSM-CTC: An open encoder-only speech foundation model for speech recognition, translation, and language identification,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:12:24.649967Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T12:12:18.949572Z digest=sha256:9c1b3465f458d9f32473f76c2dc86b83507c3f58b32bf5a0a3dc081f439e8d9e

Observation 56d343c9-0199-4c6e-9ab5-09ac9178e1b8 · outbound

This paper cites Connectionist temporal classification: Labelling unsegmented sequence data with recurrent neural networks,.

OWSM v4: Improving Open Whisper-Style Speech Models via Data Scaling and Cleaning Connectionist temporal classification: Labelling unsegmented sequence data with recurrent neural networks,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:12:24.544340Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T12:12:19.036480Z digest=sha256:0b4242a4a4065338a72a79875caf9a5e194d5e4a311d51d89eab00591f694eb7

Observation ea1b3402-bfe7-4ff4-b7a7-7812ed9f1bed · outbound

This paper cites Unsupervised data selection via discrete speech representation for ASR,.

OWSM v4: Improving Open Whisper-Style Speech Models via Data Scaling and Cleaning Unsupervised data selection via discrete speech representation for ASR,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:12:24.424537Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T12:12:19.139139Z digest=sha256:baf1d07cbcbf2e61b40d66aff7d488b65d404c049879589d610de847aa28556b

Observation d971e49d-8a2c-4745-bc46-4fcaf0b7582d · outbound

This paper cites Unsupervised data selec- tion for speech recognition with contrastive loss ratios,.

OWSM v4: Improving Open Whisper-Style Speech Models via Data Scaling and Cleaning Unsupervised data selec- tion for speech recognition with contrastive loss ratios,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:12:24.308824Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T12:12:19.209648Z digest=sha256:3140468e2be3769dbb0e9deccd6f0330447e326cbed4d3f0ae6b079e520e2024

Observation 272d04d1-4869-4c60-b6c8-f692c1539935 · outbound

This paper cites Spgispeech: 5, 000 hours of transcribed financial audio for fully formatted end-to-end speech recognition,.

OWSM v4: Improving Open Whisper-Style Speech Models via Data Scaling and Cleaning Spgispeech: 5, 000 hours of transcribed financial audio for fully formatted end-to-end speech recognition,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:12:24.147242Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T12:12:19.301270Z digest=sha256:a5cae5b804abe58c11af58cd4f7291f311acf47da664ee8af027f78def7b0fb1

Observation 2c835fcb-4178-443b-bed0-c9e0bb01c651 · outbound

This paper cites Gigaspeech: An evolv- ing, multi-domain ASR corpus with 10, 000 hours of transcribed audio,.

OWSM v4: Improving Open Whisper-Style Speech Models via Data Scaling and Cleaning Gigaspeech: An evolv- ing, multi-domain ASR corpus with 10, 000 hours of transcribed audio,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:12:23.980936Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T12:12:19.428013Z digest=sha256:dd33b660dcac888e99af922ed85defbb600666733446f41d0776420310da19f8

Observation 9ec2465d-b686-4696-a457-74c0a22ec0aa · outbound

This paper cites The People's Speech: A Large-Scale Diverse English Speech Recognition Dataset for Commercial Usage.

OWSM v4: Improving Open Whisper-Style Speech Models via Data Scaling and Cleaning The People's Speech: A Large-Scale Diverse English Speech Recognition Dataset for Commercial Usage

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T12:12:19.512688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:12:19.512688Z digest=sha256:0cbbd09e0465f2caffc985610edafcc8cfb3493cbfff517deca828edf7ce0425

Observation 475da79a-c3f1-4fcf-b4df-06d6e81bd146 · outbound

This paper cites YODAS: Youtube-Oriented Dataset for Audio and Speech,.

OWSM v4: Improving Open Whisper-Style Speech Models via Data Scaling and Cleaning YODAS: Youtube-Oriented Dataset for Audio and Speech,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:12:23.795234Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T12:12:19.597084Z digest=sha256:a84e1bed6c88322fdc326adbb07c47d43b78a0f247c3fb6c274ee4189ca59761

Observation 13297370-5350-435d-8613-107046f5c0d7 · outbound

This paper cites On the effects of het- erogeneous data sources on speech-to-text foundation models,.

OWSM v4: Improving Open Whisper-Style Speech Models via Data Scaling and Cleaning On the effects of het- erogeneous data sources on speech-to-text foundation models,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:12:23.611777Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T12:12:19.703596Z digest=sha256:60908502cbdc5bd7377d2da1c18bc783a74afa126706f91733756f8e3d6d3559

Observation 77bf07e5-9d02-4ed0-a257-5038fcbd6f38 · outbound

This paper cites SpeechStew: Simply Mix All Available Speech Recognition Data to Train One Large Neural Network.

OWSM v4: Improving Open Whisper-Style Speech Models via Data Scaling and Cleaning SpeechStew: Simply Mix All Available Speech Recognition Data to Train One Large Neural Network

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T12:12:19.793022Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:12:19.793022Z digest=sha256:d51614e447b6818b5cd5fb20473d89d0095be038181709b49b21d2330abd99fa

Observation 4575edf5-26aa-40ee-8c78-c593b0634b56 · outbound

This paper cites MSR-86K: An Evolving, Multilingual Corpus with 86,300 Hours of Transcribed Audio for Speech Recognition Research,.

OWSM v4: Improving Open Whisper-Style Speech Models via Data Scaling and Cleaning MSR-86K: An Evolving, Multilingual Corpus with 86,300 Hours of Transcribed Audio for Speech Recognition Research,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:12:23.463609Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T12:12:19.884354Z digest=sha256:a553edbaff6e677e324117d854ec25d5e7716166f4f1f845d854b740bbd14da5

Observation c8bfbcdb-e401-4f64-8bbb-f01fba88ba5a · outbound

This paper cites Libriheavy: A 50,000 hours asr corpus with punctuation casing and context,.

OWSM v4: Improving Open Whisper-Style Speech Models via Data Scaling and Cleaning Libriheavy: A 50,000 hours asr corpus with punctuation casing and context,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:12:23.270870Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T12:12:19.960379Z digest=sha256:ac85c49ca6be2cfa2dcb09f7bd3d22a027cd73df93651a624adb74f4620b5b8b

Observation 405fdaac-7213-4d65-b898-1e5404abd152 · outbound

This paper cites GigaSpeech 2: An Evolving, Large-Scale and Multi-domain ASR Corpus for Low-Resource Languages with Automated Crawling, Transcription and Refinement.

OWSM v4: Improving Open Whisper-Style Speech Models via Data Scaling and Cleaning GigaSpeech 2: An Evolving, Large-Scale and Multi-domain ASR Corpus for Low-Resource Languages with Automated Crawling, Transcription and Refinement

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T12:12:20.068090Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:12:20.068090Z digest=sha256:710f06911ba3212366df83388dd58e5315ab5fbab77d74fdd49d179d5f1a1afe

Observation b61a362a-029f-4875-a106-8c302d761e8b · outbound

This paper cites MOSEL: 950,000 Hours of Speech Data for Open-Source Speech Founda- tion Model Training on EU Languages,.

OWSM v4: Improving Open Whisper-Style Speech Models via Data Scaling and Cleaning MOSEL: 950,000 Hours of Speech Data for Open-Source Speech Founda- tion Model Training on EU Languages,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:12:23.100262Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T12:12:20.156850Z digest=sha256:70d04f7653521ed44674e2e7f086186a01dfb97dc594f60d0bdda77403b783d1

Observation 037e1b2f-f68f-4680-9328-3b76ff9c7ab4 · outbound

This paper cites CTC-Segmentation of Large Corpora for German End-to-End Speech Recognition,.

OWSM v4: Improving Open Whisper-Style Speech Models via Data Scaling and Cleaning CTC-Segmentation of Large Corpora for German End-to-End Speech Recognition,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:12:22.956806Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T12:12:20.255300Z digest=sha256:09d72eff5bb92f2a3b9822afb635206581ef1764da58b719f7bb8db9ed55f4b0

Observation ff89a244-e405-41d6-a38f-b4a7e9909921 · outbound

This paper cites Bag of Tricks for Efficient Text Classification.

OWSM v4: Improving Open Whisper-Style Speech Models via Data Scaling and Cleaning Bag of Tricks for Efficient Text Classification

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T12:12:20.365163Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:12:20.365163Z digest=sha256:8675456bce78397c221216407129f64eddbff27c1396b68f6ce3bb7a1be12038

Observation 2207002f-9773-4956-b514-6e8c6148b10c · outbound

This paper cites FastText.zip: Compressing text classification models.

OWSM v4: Improving Open Whisper-Style Speech Models via Data Scaling and Cleaning FastText.zip: Compressing text classification models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T12:12:20.479161Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:12:20.479161Z digest=sha256:790a8a8a8d64f416b2d13b06e2fd78f2cf2855e655bf3966d8845ac1e368f0e6

Observation a1278efc-5e62-4b40-95e9-1fc712b23bd1 · outbound

This paper cites SpeechBrain: A General-Purpose Speech Toolkit.

OWSM v4: Improving Open Whisper-Style Speech Models via Data Scaling and Cleaning SpeechBrain: A General-Purpose Speech Toolkit

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T12:12:20.555853Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:12:20.555853Z digest=sha256:dbb142ca8f82b9de8188ee361141ae54eea7ee0cad21d1363337381b729658a1

Observation 39e7fa83-9333-474a-b8a6-85f1307ad3fb · outbound

This paper cites Common Voice: A Massively-Multilingual Speech Corpus.

OWSM v4: Improving Open Whisper-Style Speech Models via Data Scaling and Cleaning Common Voice: A Massively-Multilingual Speech Corpus

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T12:12:20.661052Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:12:20.661052Z digest=sha256:bbae9e4b6c9a3ef118f45c7be5f4a28239d29ef5ee5c078910d10e5c30704b3f

Observation 2b5be670-959e-4d36-82bc-9f0fba6c716b · outbound

This paper cites Pytorch: An imperative style, high- performance deep learning library,.

OWSM v4: Improving Open Whisper-Style Speech Models via Data Scaling and Cleaning Pytorch: An imperative style, high- performance deep learning library,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:12:22.792853Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T12:12:20.741492Z digest=sha256:5004b491d10aa2470c1a81195cd0c1eb27445820acb7e703b97cec02fc143faf

Observation 8e7d6a8a-75ae-48f8-ab33-a35fe2579cab · outbound

This paper cites FlashAttention-2: Faster Attention with Better Paral- lelism and Work Partitioning,.

OWSM v4: Improving Open Whisper-Style Speech Models via Data Scaling and Cleaning FlashAttention-2: Faster Attention with Better Paral- lelism and Work Partitioning,

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:12:22.648375Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T12:12:20.824229Z digest=sha256:f00e64a3b54ad5891ebc67c18d5e269c99dd4e74bd59b4a65fe0cc414084f817

Observation 59c95f76-2645-43e3-8463-5e7bbf4706eb · outbound

This paper cites Decoupled weight decay regular- ization,.

OWSM v4: Improving Open Whisper-Style Speech Models via Data Scaling and Cleaning Decoupled weight decay regular- ization,

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:12:22.469508Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T12:12:20.908318Z digest=sha256:ef8cb06bc20381a6e0c83963b0070c4aaaaeff4a896aa2036c6921a31c8ab4cc

Observation 9ab61f24-357c-4dfc-915a-db9a72d7dfc7 · outbound

This paper cites CoV oST 2 and Massively Multilingual Speech Translation,.

OWSM v4: Improving Open Whisper-Style Speech Models via Data Scaling and Cleaning CoV oST 2 and Massively Multilingual Speech Translation,

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:12:22.314269Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T12:12:21.008812Z digest=sha256:c7afdf46d251e13fac6476d5db17f777654e2ead4ced7bc43e41d0e3da0a276a

Observation 820967f7-1ead-4b55-b5d4-46b2b45fe252 · outbound

This paper cites FLEURS: Few-Shot Learning Evaluation of Universal Representations of Speech,.

OWSM v4: Improving Open Whisper-Style Speech Models via Data Scaling and Cleaning FLEURS: Few-Shot Learning Evaluation of Universal Representations of Speech,

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:12:22.131499Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T12:12:21.049823Z digest=sha256:398f775e8c78945fbea5e5939f27006a157aa3b97f3e45cc9033b24300d7c74c

Observation d196b4a7-095d-443c-b774-7b1f45f97f0c · outbound

This paper cites MLS: A Large-Scale Multilingual Dataset for Speech Research.

OWSM v4: Improving Open Whisper-Style Speech Models via Data Scaling and Cleaning MLS: A Large-Scale Multilingual Dataset for Speech Research

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T12:12:21.094758Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:12:21.094758Z digest=sha256:d8dab33cdc2c59519be68b6a48a232ff882e0987aa5c5546e576599e56d6bfc4

Observation 48a6f4de-4db2-4861-8697-d222a52175a2 · outbound

This paper cites Srivastav, S.

OWSM v4: Improving Open Whisper-Style Speech Models via Data Scaling and Cleaning Srivastav, S

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:12:21.979286Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T12:12:21.174309Z digest=sha256:18fdd99d04e144a18f2ba353d37e4aead7ecf2f8f844cc9843b8ec3371b5bc36

Pith citing papers

Observation da98eed8-8609-4fd1-87f4-9a9d2c0507e2 · inbound

OWSM v4: Improving Open Whisper-Style Speech Models via Data Scaling and Cleaning cites this paper.

OWSM v4: Improving Open Whisper-Style Speech Models via Data Scaling and Cleaning OWSM v4: Improving Open Whisper-Style Speech Models via Data Scaling and Cleaning

Reference 1

Resolution
malformed identifier
local_arxiv, observed 2026-08-07T12:12:21.828937Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T12:12:17.177212Z digest=sha256:6b5e2a8d0e906eb7ce2c8c2a93988a67bdfaf36678baefc1760f0c9382d39179