{"id":"1ec6e0f9-a7f8-4186-8ee9-24f66d699b4b","arxiv_id":"2606.21237","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":4.0,"correctness_risk":"high","formal_verification":"none","parameter_count":0,"one_line_summary":"OpenWER improves cross-lingual WER robustness via language-specific normalization, compound word detection, and token-based Levenshtein alignment, reporting up to 25% absolute reductions across 52 languages.","lead":"OpenWER is an open-source tool that refines Word Error Rate for automatic speech recognition using language-specific normalization and compound word detection. This could support fairer evaluations of multilingual ASR systems, especially for low-resource languages.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.3","headline":"WER reductions may stem from uncontrolled differences in tokenization/preprocessing vs. claimed normalisation features","rationale":"The reader's weakest assumption directly identifies the attribution gap; the full-text description does not alter this because the claim still rests on an uncontrolled comparison. The concrete test isolates the variable without requiring new data collection.","tokens_in":1607,"tokens_out":275,"duration_ms":13265,"concrete_test":"Recompute WER on the same 52-language test sets using OpenWER and one common library (e.g. jiwer) after forcing identical tokenization, lower-casing, punctuation removal, and compound-word splitting; if the average absolute reduction falls below 5 points, the attribution to the new features is unsupported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The headline result (absolute WER reductions up to 25% across 52 languages) requires that OpenWER's language-specific normalisation and compound-word handling are the causal factors. The comparison is only to 'common libraries'; if those runs used different tokenizers, reference text cleaning, or alignment rules, the delta cannot be attributed to the new components. This assumption is load-bearing because WER is sensitive to exactly these choices, especially cross-lingually, and the abstract provides no explicit statement that all other variables were held fixed.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.3","summary":"The paper introduces OpenWER, an open-source library for computing Word Error Rate (WER) in multilingual ASR. It adds language-specific text normalisation and compound-word detection to improve robustness, implements a token-based Levenshtein alignment that retains complementary token-level metrics and supports metadata embedding, and reports absolute WER reductions of up to 25 % across 52 languages relative to common libraries.","tokens_in":1706,"tokens_out":455,"duration_ms":9871,"significance":"If the reported reductions are shown to arise specifically from the normalisation and compound-word components under controlled conditions, the work would strengthen the reliability of cross-lingual ASR evaluation and support more granular accuracy analyses. The open-source release and preservation of token-level metrics are concrete strengths that could be adopted by the community.","major_comments":[{"comment":"Abstract and §4 (Experiments): the headline claim of up to 25 % absolute WER reduction across 52 languages is load-bearing, yet the manuscript supplies no explicit statement that reference-text cleaning, tokenisation, and alignment rules were held identical between OpenWER and the compared libraries. Because WER is known to be sensitive to exactly these preprocessing choices, the attribution of the delta to the new normalisation features cannot be evaluated from the current description.","section":"Abstract and §4"},{"comment":"§4: no table or figure reports per-language WER values, baseline library names and versions, or any measure of statistical significance or variance across runs. Without these data the cross-lingual claim remains untestable.","section":"§4"}],"minor_comments":[{"comment":"§3.2: the description of the token-based alignment would benefit from a small worked example showing how metadata is embedded and how the resulting token-level scores are aggregated back to WER.","section":"§3.2"},{"comment":"References: several standard multilingual ASR evaluation papers (e.g., on Common Voice or FLEURS) are not cited when discussing cross-lingual challenges.","section":"References"}],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":{"model":"grok-4.3","summary":"We thank the referee for the constructive comments. We address each major point below and agree to revisions that improve clarity and testability of the results.","responses":[{"response":"We agree that an explicit statement is required. All experiments used identical reference transcripts across libraries; the reported deltas arise solely from OpenWER's language-specific normalisation, compound-word detection, and token-based alignment versus the standard rules in the baseline libraries. We will revise §4 to state this control explicitly so that attribution to the new components can be evaluated.","revision_made":"yes","referee_comment":"[Abstract and §4] Abstract and §4 (Experiments): the headline claim of up to 25 % absolute WER reduction across 52 languages is load-bearing, yet the manuscript supplies no explicit statement that reference-text cleaning, tokenisation, and alignment rules were held identical between OpenWER and the compared libraries. Because WER is known to be sensitive to exactly these preprocessing choices, the attribution of the delta to the new normalisation features cannot be evaluated from the current description."},{"response":"We agree that per-language values, library versions, and distributional information would strengthen testability. We will add a supplementary table (referenced from §4) listing per-language WERs for OpenWER and each baseline, together with the exact library names and versions. Because WER computation is deterministic for fixed references and hypotheses, run-to-run variance does not apply; we will instead report the distribution and range of improvements across the 52 languages to support the 'up to 25 %' claim.","revision_made":"yes","referee_comment":"[§4] §4: no table or figure reports per-language WER values, baseline library names and versions, or any measure of statistical significance or variance across runs. Without these data the cross-lingual claim remains untestable."}],"tokens_in":1257,"tokens_out":410,"duration_ms":17750,"standing_objections":[]},"desk_editor":{"model":"grok-4.3","letter":"The main takeaway is that this paper ships a usable open implementation of improved WER normalization plus compound-word handling and token-level alignment, which is genuinely helpful for anyone running multilingual ASR evaluations. The 52-language analysis is also more than most metric papers bother with.\n\nWhat works is the engineering focus. Releasing the code with explicit language rules and keeping token alignment so other metrics stay available is the right move. It directly addresses the English-centric bias in current WER libraries without claiming to replace the metric entirely.\n\nThe soft spot is the headline result. The abstract states absolute WER drops up to 25% versus common libraries, yet gives no indication that tokenization, reference cleaning, or alignment rules were locked down across the comparisons. WER numbers move easily with exactly those choices, especially across languages, so the reductions cannot yet be credited to the normalization and compound detection steps. The stress-test concern lands on the abstract as written.\n\nIf the full paper shows controlled runs with everything else fixed, the claim strengthens. If not, the numbers are suggestive but not yet attributable. No circularity or invented parameters appear.\n\nThis is for ASR groups that evaluate models on many languages and want a drop-in tool rather than a new modeling paper. A reader who needs reproducible cross-lingual scores would find it worth trying. It is coherent on its own terms and deserves referee time to check the experimental controls and code quality.","headline":"OpenWER is a practical open-source WER tool with language-specific normalization, but the 25% reduction claim rests on comparisons that may not isolate the new features from preprocessing differences.","tokens_in":2170,"tokens_out":371,"would_cite":false,"duration_ms":14639,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.3","headline":"OpenWER applies language-specific normalisation and compound word detection to produce more reliable Word Error Rate scores across languages.","keywords":["automatic speech recognition","word error rate","cross-lingual evaluation","multilingual ASR","text normalization","compound word detection","Levenshtein alignment","evaluation metrics"],"falsifier":"Re-run the 52-language evaluation suite with OpenWER while forcing identical tokenization and preprocessing steps as the baseline libraries; if the 25 percent reductions disappear, the central claim does not hold.","tokens_in":2503,"feed_emoji":"📊","tokens_out":645,"duration_ms":12258,"temperature":0.7,"pith_summary":"Automatic speech recognition models now handle many languages at once, yet standard evaluation with Word Error Rate often fails to account for language-specific text features like compound words. The paper presents OpenWER, an open-source tool that normalizes reference and hypothesis texts according to each language's conventions before alignment. This produces lower WER values than common libraries, with absolute reductions reaching 25 percent in tests on 52 languages. The same alignment also supports token-level metrics and metadata attachment for finer-grained accuracy analysis. These changes matter for fair comparisons between models that serve different languages and for more trustworthy results on low-resource languages.","feed_headline":"OpenWER cuts WER by up to 25% across 52 languages","feed_subtitle":"Language-specific normalisation and compound detection produce fairer speech recognition scores for many languages at once.","key_machinery":"OpenWER, the implementation that performs language-specific normalisation plus compound word detection before applying token-based Levenshtein alignment to compute WER.","core_discovery":"OpenWER is an open-source implementation that improves WER robustness through language-specific normalisation and compound word detection. A token-based Levenshtein alignment preserves complementary metrics and allows metadata embedding for granular accuracy scores. Analysis of 52 languages shows absolute WER reductions of up to 25% compared to common libraries.","pith_inferences":["Adoption of OpenWER could gradually shift ASR benchmark reporting toward language-aware normalisation as a default practice.","The token-level alignment approach might be reused for other sequence metrics such as character error rate in future tools.","Model developers could incorporate similar normalisation steps during training to reduce the gap between training and evaluation conditions.","Benchmark organizers might add compound-word handling as a required preprocessing step in shared tasks."],"forward_implications":["Evaluations of multilingual ASR models become more consistent across languages.","Low-resource languages receive more accurate performance estimates than before.","Researchers can attach per-token metadata to produce additional accuracy measures alongside WER.","Cross-lingual model comparisons rest on a more uniform metric foundation.","Standard libraries may need updates to match the normalisation rules introduced here."],"fun_headline_variants":["OpenWER refines WER robustness for 52 languages","Token alignment in OpenWER preserves granular metrics","OpenWER shows up to 25% WER reduction in 52 languages","Language normalisation improves OpenWER cross-lingual scores"],"cache_read_input_tokens":2112,"weakest_assumption_plain":"The reported WER reductions result specifically from the language-specific normalisation and compound word detection rather than from differences in tokenization, data preprocessing, or baseline library configurations.","fun_headline_variants_meta":{"raw":{"variants":["OpenWER refines WER robustness for 52 languages","Token alignment in OpenWER preserves granular metrics","OpenWER shows up to 25% WER reduction in 52 languages","Language normalisation improves OpenWER cross-lingual scores"]},"model":"grok-4.3","cost_usd":0.006245,"raw_usage":{"total_tokens":2888,"prompt_tokens":565,"num_sources_used":0,"completion_tokens":64,"cost_in_usd_ticks":62449500,"prompt_tokens_details":{"text_tokens":565,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":2259,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":565,"tokens_out":64,"duration_ms":19693,"temperature":1.0,"reasoning_tokens":2259,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-06-26T14:37:43.947924+00:00","model_set":{"reader":"grok-4.3"},"falsifier":"Re-run the 52-language evaluation suite with OpenWER while forcing identical tokenization and preprocessing steps as the baseline libraries; if the 25 percent reductions disappear, the central claim does not hold.","supporting_citations":[],"review_version":1}