{"id":"63c4681b-050a-4f0a-9ab7-3ff809ed152c","arxiv_id":"2507.07741","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A systematic review of 127 papers shows code-switching ASR research is concentrated in a few language pairs and fragmented across datasets, metrics, and non-reproducible methods.","lead":"This paper reviews 127 published studies of end-to-end speech recognition for code-switched speech, mapping the languages, datasets, models, and metrics the field uses. It finds heavy concentration on Mandarin-English and a lack of shared benchmarks and reproducible evaluation across the field.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The survey's headline concentration statistics rest on a narrow search query; omitted 'language mixing'/'mixed-language' work could shift the Mandarin-English share.","rationale":"The reader's weakest assumption identified both search coverage and single-annotator coding. I agree that these are the load-bearing points, but I single out the search query as the more fundamental premise because it determines the universe from which all statistics are drawn. Even perfect annotation cannot fix a systematically missing subset of the literature. The proposed test is a direct computational check: re-run the search with expanded terminology and measure how the language-pair distribution moves. A small change would confirm the qualitative conclusions; a large change would weaken the precise percentages while possibly preserving the broader message about concentration and lack of benchmarking. The lack of inter-annotator agreement is an additional transparency concern, but it is secondary and also fixable. Since the existing CONDITIONAL verdict already asks for such auditing, my read does not move the verdict.","tokens_in":26530,"tokens_out":7016,"duration_ms":80475,"concrete_test":"Re-run the Section 2 Semantic Scholar search with the original query and with an expanded query that also includes 'language mixing', 'mixed-language', 'language switching', 'code-alternation', and 'bilingual speech recognition'. After applying the same screening criteria, compare the two paper sets by title/abstract. If the expanded query adds more than 5% new papers and the Mandarin-English share drops by more than 5 percentage points, the concentration claim is not robust to search terminology.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The most load-bearing premise is the completeness of the Semantic Scholar query in Section 2. The query requires (ASR | Speech recognition) AND (code-switch* | codeswitch* | code switch* | code-mix* | codemix*). Papers that study the same phenomenon but use alternative terminology such as 'language mixing', 'mixed-language', 'language switching', or 'code-alternation' would be omitted even if they are squarely about code-switched speech recognition. If the omitted papers are not distributed like the 127 included papers, the headline statistics ('Mandarin-English ~55%', 'top-3 ~76%') change directly. The paper also states that only peer-reviewed venues were kept but does not describe how venue status was verified, and each paper was annotated by a single annotator with no inter-annotator agreement reported. The central claim is a quantitative concentration claim, so a systematically biased sample changes the numerator and denominator. This is not a detected error, but it is an unverified premise that the review does not address.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents a systematic literature review of end-to-end (E2E) automatic speech recognition (ASR) for code-switching (CS). The authors query the Semantic Scholar API with a documented boolean query, retain 127 peer-reviewed papers, and manually annotate them across dimensions such as language pairs, datasets, model architecture, decoding strategy, and evaluation metrics. The central empirical claims are that research is heavily concentrated in a few language pairs (Mandarin–English ~55%, top-3 pairs ~76%), that ~77% of papers use accessible datasets, and that the field lacks standardized benchmarking and evaluation practices. The paper also discusses challenges and opportunities around data scarcity, coverage disparities, evaluation, and reproducibility.","tokens_in":26710,"tokens_out":5731,"duration_ms":61734,"significance":"If the quantitative findings are reliable, this review fills a real gap: it is the first systematic survey focused specifically on E2E ASR for code-switching, and it provides a DOI-level enumeration of all 127 surveyed papers (Table 7), a documented search query, and a transparent annotation scheme. The qualitative picture — that Mandarin–English dominates because of dataset availability and that evaluation is fragmented — is plausible and useful for guiding future research. The paper also makes honest statements about its own coding assumptions (e.g., Section 4.6, the greedy-decoding assumption), which is a strength. However, the headline percentages are load-bearing, and they rest on two methodological premises: search-query recall and single-annotator coding reliability. These premises are not adequately addressed in the manuscript, which limits the strength of the quantitative claims even though the overall qualitative conclusions may well survive.","major_comments":[{"comment":"The search query is limited to variants of 'code-switch*' and 'code-mix*'. Work that studies the same phenomenon under alternative terminology — 'language mixing', 'mixed-language speech', 'language switching', or 'code-alternation' — would be omitted unless those terms co-occur with 'code-switch' in the indexed fields. Because the central quantitative claims in Section 3 (Mandarin–English ~55%, top-3 ~76%, accessible-dataset share ~77%) are computed from the 127-paper sample, the review needs to assess recall, for example by running a broader query and reporting how the distribution shifts, or by explicitly identifying known relevant papers that the query missed. Without such a sensitivity check, the headline statistics are not robust to plausible alternative terminologies.","section":"Section 2 (Data Collection) and Section 3"},{"comment":"Each paper was annotated by one annotator, and no inter-annotator agreement (IAA) is reported. The review's conclusions are expressed as percentages over categorical codes (language pair, dataset accessibility, architecture class, etc.), so the reliability of those codes is load-bearing. The authors should report IAA on a subsample, provide dual coding for the attributes that feed the central statistics, or otherwise justify the consistency of single-annotator coding across five annotators. Without this, the precision of numbers like 55% and 76% is unverified.","section":"Section 2 (Annotation)"},{"comment":"The inclusion criterion 'published in peer-reviewed venues' is asserted but not operationalized, and Table 7's note says the year corresponds to the earliest available version, 'which can be a preprint'. Some entries in the reference list are clearly preprint-form (CoRR), and the manuscript does not explain how venue status was verified or whether preprint-only items were included. Since the 127-paper corpus is the basis for all of the review's statistics, the paper should clarify the verification procedure and state explicitly whether preprints were admitted or excluded.","section":"Section 2 and Table 7"}],"minor_comments":[{"comment":"'ITTG-HingCos' should be 'IITG-HingCos' to match the reference and Table 3.","section":"Section 3.2"},{"comment":"'languauge independent vocabulary' contains a typo ('languauge' should be 'language').","section":"Section 4.3"},{"comment":"'with the exception of of a small set of languages' has a duplicated 'of'.","section":"Section 7"},{"comment":"The phrase 'the combined system is retained using monolingual data' should read 'retrained'.","section":"Section 4.1"},{"comment":"The assumption that unmentioned decoding strategies are greedy is stated transparently, but only seven papers explicitly mention greedy decoding; the review would be stronger if it reported the sensitivity of any decoding-related conclusions to this assumption.","section":"Section 4.6"},{"comment":"A PRISMA-style flow diagram (378 retrieved, exclusions by reason) would improve reproducibility and make the filtering process easier to verify.","section":"Section 2"},{"comment":"The phrase 'or to used TTS synthesis' is grammatically incomplete; it should read 'or to use TTS synthesis'.","section":"Section 5.2"},{"comment":"Several reference entries contain OCR-style spacing errors, e.g., 'V oice' in Grand View Research, 'V ancouver' in Zhang et al. (2018), and 'V enice' in Zhu et al. (2017).","section":"References"},{"comment":"The entry 'Kilkarn' should be 'Killkan' to match the cited dataset name in the reference list.","section":"Table 7"}],"recommendation":"major_revision","confidential_remarks":"The paper includes several self-citations (Mixat, PolyWER) but treats them as data points in the corpus, so I do not see a circularity problem. The main concern is that the quantitative claims, which are the paper's headline, depend on search recall and annotation reliability that are not yet demonstrated. These are fixable within the manuscript's scope, hence major_revision rather than reject. The paper would also benefit from releasing the annotated dataset, or at least a coded appendix, to support verification."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"This is a solid, useful map of end-to-end code-switching ASR. The qualitative findings are correct: research is concentrated in a few language pairs, dataset availability drives the agenda, and evaluation is fragmented. If you work in this area, the paper will save you a week of literature chasing.\n\nWhat's actually new: it is the first E2E-only systematic review, covering 127 papers from 2018 to 2024, and it ships a DOI-level list of every surveyed paper in Table 7. That table alone is worth the price of admission, because it makes the corpus auditable and gives other researchers a concrete starting point. The annotation scheme across languages, datasets, model choices, and metrics is reasonable, and the discussion of challenges around resources and evaluation is balanced, not preachy.\n\nThe soft spots are real but not fatal. The Semantic Scholar query requires 'code-switch*' or 'code-mix*' terms, so work framed as 'language mixing' or 'mixed-language' would be missed. The stress-test note is right that this could shift the exact percentages. My honest reading is that it would not reverse the main concentration finding: Mandarin-English dominance is so pronounced in the included sample, and so consistent with the broader speech literature, that a broader query would trim 55% to maybe 50% rather than change the ranking. Still, the authors should run a sensitivity check with expanded terminology and report whether the top-3 share moves.\n\nBigger issue: each paper was coded by a single annotator with no inter-annotator agreement reported. For a review whose headline claims are quantitative, that is a transparency gap. The DOI list allows some independent spot-checking, but the annotation sheets are not released. These are fixable conditions, not errors. The 'best result' annotations are inherently noisy because papers use different test sets, but the paper handles that by contextualizing Table 4. The self-citations (Mixat, PolyWER) appear as data points, not as load-bearing evidence, so I do not see circularity.\n\nWho is this for? Anyone doing CS ASR, especially newcomers deciding which language pair or dataset to work on. The qualitative takeaways are reliable; the specific percentages should be cited as approximate. It deserves a serious referee, and with the transparency conditions addressed it would be a good contribution.\n\nRecommendation: send to peer review with a request for inter-annotator agreement on a subsample, release of the annotation data or a detailed codebook, and an explicit sensitivity analysis on the search query. The core argument holds up.","headline":"Useful E2E code-switching ASR field map; the 55% Mandarin-English share is directionally right but not a precise figure.","tokens_in":27249,"tokens_out":1745,"would_cite":true,"duration_ms":21786,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A systematic review of 127 peer-reviewed papers finds end-to-end code-switching ASR concentrated in a few language pairs—Mandarin-English alone is about 55% of studies—with no consistent benchmarking across the field.","keywords":["code-switching","end-to-end automatic speech recognition","systematic literature review","language-pair coverage","code-switched speech datasets","ASR evaluation metrics","multilingual speech models","benchmarking"],"falsifier":"Re-run the collection through a second scholarly literature index with the query expanded by terms like 'mixed-language,' 'language mixing,' and 'code-alternation,' and have two annotators independently code a random sample of 30 papers. If the Mandarin-English share moves by more than a few percentage points, or if agreement on core fields (languages covered, dataset accessibility, model type) is poor, the review's headline statistics would need revision, even if the qualitative picture of concentration and fragmentation stands.","tokens_in":26330,"feed_emoji":"🗣️","tokens_out":13598,"duration_ms":126352,"temperature":0.7,"pith_summary":"This paper asks, after a decade of growth, which languages, datasets, models, and metrics actually populate end-to-end speech recognition (ASR) for code-switched speech, and answers by collecting 127 peer-reviewed papers and manually annotating each one across four dimensions. Its central finding is that the field is lopsided and unstandardized: Mandarin-English alone covers about 55% of papers, the top three language pairs about 76%, and most of the 38 identified datasets each cover a single language pair, with 77% of studies relying on accessible data. It also finds no common training or evaluation recipe, so results across studies cannot be directly compared, and reports that none of the best-performing models on popular datasets uses zero-shot evaluation. A sympathetic reader should care because code-switching is a global everyday practice and speech recognition is a mass-market technology: if the review is right, the field's bottleneck is not new architectures but accessible datasets and shared benchmarks for under-studied language pairs.","feed_headline":"55% of code-switching ASR studies cover one language pair","feed_subtitle":"A 127-paper review shows dataset access, not linguistic need, drives the field—and no shared benchmarks exist.","key_machinery":"The machinery is the review protocol itself. The authors queried the Semantic Scholar API for papers matching (ASR | Speech recognition) together with (code-switch* | codeswitch* | code switch* | code-mix* | codemix*), kept only peer-reviewed work describing end-to-end code-switching systems (127 papers), and coded each paper along four annotated dimensions: problem setup and data (languages, datasets, accessibility), model design (monolingual vs multilingual modeling, language identification, text units, architecture, pretrained models, loss, decoding), training and evaluation (augmentation, translation, zero-shot, metrics), and best reported performance. Each paper was annotated by one of five annotators working from agreed dimensions in scheduled sessions. That grid is what turns heterogeneous papers into comparable statistics like the ~55% and ~77% figures, and every conclusion in the paper flows through it.","core_discovery":"On the paper's own terms, the discovery is a measured snapshot of a young field. From 378 candidate papers retrieved through the Semantic Scholar API, filtering to peer-reviewed venues and end-to-end systems—neural ASR trained jointly in a single computational graph—leaves 127 papers published between 2018 and 2024, over half of them since 2022. Manual annotation across problem setup, model design, training and evaluation, and performance yields three headline facts. First, research concentrates in a few pairs: Mandarin-English is covered by ~55% of papers; adding Hindi-English and Arabic-English reaches ~76%, and momentum tracks the availability of accessible datasets such as SEAME (~24% of papers) and the ASRU 2019 challenge (~11%). Second, ~77% of papers use accessible datasets while the rest rely on proprietary or unspecified data, and the Japanese-English cluster (~5%) has no public dedicated dataset at all, leaning on synthetic speech built from a machine-translation corpus. Third, no experimental standard exists: roughly 60% of studies add monolingual data, more than a third use multilingual modeling, ~47% use encoder-decoder architectures, ~33% incorporate language identification, ~45% an external language model, ~30% data augmentation, and best reported results on different datasets share no common recipe. The authors conclude that efforts are 'mostly sporadic, with no consistent benchmarking or clarity on directions for future research,' and note that no state-of-the-art result on the popular datasets comes from zero-shot evaluation, which they read as evidence that dedicated code-switched training data remains essential.","pith_inferences":["The census itself is shaped by its search terms: work framed as 'mixed-language,' 'language mixing,' or dialectal alternation without code-switch/code-mix vocabulary would be invisible to the query, so the true literature is probably larger and more linguistically diverse than 127 papers; the concentration percentages could shift if that framing gap were closed.","A quantitative extension the authors leave implicit: measuring citation continuity—whether later papers on the same dataset cite and improve on earlier baselines—would convert the 'sporadic efforts' claim from an impression into a number, and the paper's own reference table makes this computation straightforward.","The paper's logic supports a testable, supply-side prediction: releasing a public Japanese-English or Indonesian-English corpus of roughly SEAME's scale would plausibly generate a visible publication cluster within two years; future bibliographies could confirm or refute this.","Read with the fairness discussion, the findings imply a concrete allocation rule rather than a research direction: directing benchmark and dataset funding at the roughly thirty language pairs that each appear in a single paper would rebalance the field more than further architectural innovation."],"forward_implications":["Dataset availability, not linguistic need, drives the research agenda: the Mandarin-English cluster grew around SEAME and the ASRU 2019 challenge, so creating one accessible corpus for an under-studied pair is the highest-leverage intervention the data support.","Because there is no consistent benchmarking, positive findings from earlier studies have not been replicated on newer datasets; reported gains should be treated as provisional until a shared evaluation framework appears.","Dedicated code-switched training data remains necessary: no state-of-the-art system in the surveyed tables relies on zero-shot evaluation, so multilingual pretraining alone does not substitute for CS data.","Evaluation itself is unsettled: plain WER penalizes mixed-script output, which is why mixed metrics (MER, TER) and transliteration-aware metrics (toWER, poWER, PolyWER) are appearing; standardizing among them is part of the open problem.","The field is still young and accelerating—over half the surveyed papers appeared since 2022—so the window for establishing standardized benchmarks is open now."],"supporting_citations":[{"why":"Supplies SEAME, the most-used dataset in the survey (~24% of papers); the Mandarin-English concentration claim largely rests on its availability.","marker":"Lyu et al., 2010"},{"why":"Describes the ASRU 2019 code-switching challenge dataset (~11%), the second anchor of the Mandarin-English cluster.","marker":"Shi et al., 2020"},{"why":"The MUCS 2021 shared-task dataset that ~38% of Indic-language papers rely on; supports the Hindi-English and Bengali-English share statistics.","marker":"Diwan et al., 2021"},{"why":"The only prior survey of code-switching in ASR; the paper's scope and sample size (24 vs 127 papers) are defined against it.","marker":"Mustafa et al., 2022"},{"why":"The complementary decades-spanning NLP code-switching survey; establishes that the present review fills the end-to-end-ASR-only gap.","marker":"Winata et al., 2023"},{"why":"Supplies the working definition of end-to-end ASR that determines which of the 378 retrieved papers pass the inclusion filter.","marker":"Prabhavalkar et al., 2024"},{"why":"Whisper is the most commonly used pretrained model (~17% of papers); underpins the pretrained-model and zero-shot evaluation findings.","marker":"Radford et al., 2023"}],"fun_headline_variants":["No shared benchmarks for code-switching ASR, review of 127 finds","Code-switching ASR: sporadic efforts, no consistent benchmarking","Dataset access, not linguistic need, shapes code-switching ASR","Mandarin-English dominates code-switching ASR studies at 55%","Code-switching ASR research follows datasets, not languages"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The headline numbers rest on the assumption that the search query plus the peer-review filter captured essentially the whole universe of end-to-end code-switching ASR work, and that single-annotator coding is accurate enough to support figures like 55%, 76%, and 77%.","fun_headline_variants_meta":{"raw":{"variants":["No shared benchmarks for code-switching ASR, review of 127 finds","Code-switching ASR: sporadic efforts, no consistent benchmarking","Dataset access, not linguistic need, shapes code-switching ASR","Mandarin-English dominates code-switching ASR studies at 55%","Code-switching ASR research follows datasets, not languages"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000505,"raw_usage":{"total_tokens":2476,"prompt_tokens":967,"completion_tokens":1509,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":583,"completion_tokens_details":{"reasoning_tokens":1418}},"tokens_in":583,"tokens_out":1509,"duration_ms":12868,"temperature":1.0,"reasoning_tokens":1418,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T18:33:00.253826+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run the collection through a second scholarly literature index with the query expanded by terms like 'mixed-language,' 'language mixing,' and 'code-alternation,' and have two annotators independently code a random sample of 30 papers. If the Mandarin-English share moves by more than a few percentage points, or if agreement on core fields (languages covered, dataset accessibility, model type) is poor, the review's headline statistics would need revision, even if the qualitative picture of concentration and fragmentation stands.","supporting_citations":[],"review_version":1}