{"id":"11ec3b78-31d0-42e1-990d-20614b9d743c","arxiv_id":"2508.17753","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Four leading speech recognition systems score similarly on standard benchmarks but diverge sharply on accuracy, hallucinations, and bias across eight HRI-relevant datasets spanning noise, accent, age, impairment, and spontaneity.","lead":"This paper tests four speech recognition systems on eight public datasets spanning hard conditions such as noise, accents, child or impaired speech, and spontaneous talk, and finds the systems diverge sharply even when standard benchmarks look similar. Why read it: anyone building voice-controlled robots or assistants needs to know that benchmark scores can hide big differences in how systems mishear or hallucinate.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Divergence claim hinges on uncontrolled comparison; protocol transparency is the key missing evidence.","rationale":"I agree with the reader that the central claim is structurally separate from the evaluation protocol and that the protocol is the weakest point. My concern sharpens this from 'dataset selection after seeing results' to the more specific and testable issue of per-system configuration control. The reader's abstract-only review had no access to the protocol; my attack is therefore about what is missing, not about an observed flaw. Given that the paper's practical recommendation depends on the divergence being intrinsic, and given that a single concrete test (identical decoding settings) could settle it, I recommend a CONDITIONAL verdict: the claim should be accepted only if the authors provide the configuration table and pass the stability check. I did not see any evidence of internal inconsistency, and I give credit for using public datasets (reproducibility in principle), but the comparison fairness remains unverified. This does not require rejecting the paper; it requires transparency plus a robustness check.","tokens_in":814,"tokens_out":1823,"duration_ms":24064,"concrete_test":"Request the authors' full configuration table for each system (model checkpoint/API version, sampling rate, decoding parameters, prompt/context, temperature, any normalized-text or punctuation settings) and their mapping of the eight datasets to the six dimensions. Then rerun all datasets with a single, identical decoding configuration across systems (e.g., greedy decoding with no language model, or a fixed beam/LM weight) and check whether the rank order of WER and hallucination rate remains stable. If any system's ranking flips or the magnitude of the divergence shrinks below significance, the original comparison was not controlled, and the central claim weakens.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central claim—that four SOTA ASR systems with similar standard-benchmark scores diverge on accuracy, hallucination, and bias across six difficulty dimensions—requires the comparison to be apples-to-apples. The abstract does not state which systems, which benchmarks were used to assert 'similar scores,' how each system was configured (model version, sampling rate, beam width, language-model weight, prompt/context, temperature, or dataset-specific preprocessing), or how the eight datasets map to the six dimensions. If systems were tuned per dataset or used different decoding settings, the observed variations reflect test-harness artifacts rather than intrinsic model properties. This is especially load-bearing because the paper's practical recommendation—choose ASR by application-specific failure analysis—is only valid if the divergence persists under fixed, HRI-realistic conditions. The reader's weakest assumption (dataset selection after observing differences, unfair configuration) is exactly the same risk. Without a detailed protocol, the central finding is unverifiable from the abstract.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper evaluates four state-of-the-art ASR systems on eight public datasets spanning six dimensions of speech difficulty (domain-specific, accented, noisy, age-variant, impaired, and spontaneous). It claims that, despite similar performance on standard benchmarks, the systems exhibit significant variations in recognition accuracy, hallucination tendencies, and biases, with implications for human-robot interaction (HRI). The practical message is that HRI developers should select ASR systems based on application-specific failure analysis rather than benchmark rank.","tokens_in":898,"tokens_out":2259,"duration_ms":29667,"significance":"If the central claim holds, the paper makes a valuable contribution to both ASR evaluation and HRI design: it would demonstrate that standard benchmark scores are insufficient to predict HRI-relevant behavior, and it would motivate failure-oriented model selection. The breadth of eight datasets across six difficulty dimensions is a constructive attempt to cover conditions relevant to HRI. However, the result's significance depends entirely on the rigor and fairness of the comparison protocol, which cannot be assessed from the abstract alone.","major_comments":[{"comment":"The central claim of 'significant variations' is unverifiable from the information provided. The abstract does not identify the four ASR systems, the eight datasets, the underlying standard benchmarks used to assert 'similar scores,' or the statistical measures supporting 'significant.' More importantly, the protocol for comparing systems is unspecified: same decoding settings, model versions, prompts, sampling rates, and preprocessing? If systems were tuned per dataset or used different configurations, the observed divergence would be an artifact of the test harness. This is load-bearing because the recommendation to choose ASR by application-specific failure analysis is only valid under a controlled, HRI-realistic comparison.","section":"Abstract"},{"comment":"The terms 'hallucination tendencies' and 'inherent biases' are not operationalized. Hallucination could refer to insertion errors, repeated phrases, or fluent non-sense; bias could refer to demographic performance gaps, lexical bias, or dataset-specific artifacts. Without explicit definitions and evaluation metrics, the reader cannot determine whether the claimed differences are substantively meaningful or the result of the chosen scoring rules. The full text must provide precise definitions and, ideally, examples of hallucinated or biased outputs.","section":"Abstract"},{"comment":"The leap from public ASR datasets to HRI implications is not justified in the abstract. HRI audio conditions include close-talk robotic microphones, motor noise, far-field capture, barge-in, and real-time constraints. If the eight datasets are generic read or broadcast speech, the 'uniquely challenging recognition environment' claim is an extrapolation. The paper should either include HRI-specific test conditions or explicitly argue why the selected datasets approximate those conditions. This is essential for the practical recommendation's validity.","section":"Abstract"}],"minor_comments":[{"comment":"The word 'significant' should be reserved for statistical significance with reported effect sizes and confidence intervals; otherwise, use 'substantial' or report descriptive statistics.","section":"Abstract"},{"comment":"Consider naming the four systems (or at least their families) in the abstract to allow readers to gauge relevance to current HRI practice.","section":"Abstract"},{"comment":"The phrase 'inherent biases' implies intrinsic model properties, but the evaluation can only measure dataset- and task-conditional outcomes. Rephrase to 'observed biases' or 'bias patterns' to avoid overclaiming.","section":"Abstract"},{"comment":"A table mapping each of the eight datasets to the six difficulty dimensions would improve transparency and help readers evaluate coverage.","section":"General"}],"recommendation":"uncertain","confidential_remarks":"This review is based solely on the abstract, as the full text was not available to me. The central claim is plausible and important, but its validity hinges on protocol details that the abstract does not disclose. I recommend that the editor obtain the full manuscript and, if possible, an independent check of the experimental setup (system configurations, dataset preprocessing, scoring rules) before making a final decision. My 'uncertain' recommendation reflects the limitation of the review artifact, not a judgment about the underlying work."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take on arXiv:2508.17753. The headline result—four state-of-the-art ASR systems with similar standard-benchmark scores diverge on accuracy, hallucination, and bias across eight datasets spanning six HRI-relevant dimensions—is exactly the kind of measurement our field needs. If it holds up, it's a concrete argument that HRI people should stop picking ASR by leaderboard and start doing application-specific failure analysis. That's a practical, trust- and safety-relevant claim, and I don't think it's been cleaned up this way before.\n\nWhat's genuinely new: the combination of systems, datasets, and the specific focus on hallucination and bias as first-class evaluation dimensions, rather than just word error rate. The eight-dataset/six-dimension design is a sensible way to operationalize the acoustic chaos of real HRI deployments. Credit for doing an empirical evaluation on public datasets rather than a hand-wavy position paper.\n\nWhere I get nervous: the abstract gives me zero protocol. No system names, no dataset names, no decoding settings, no prompt or temperature, no effect sizes or error bars. The central comparison is only meaningful if all four systems are run under controlled, comparable conditions. If they were tuned per dataset or configured differently, the 'significant variations' are harness artifacts. The abstract also doesn't say which 'standard benchmarks' produced the similar scores, so I can't tell whether those benchmarks are even relevant to the six dimensions. These are open questions, not proven flaws—the full paper may well cover all of this—but as it stands, the load-bearing claim is unverifiable. The stress-test note about dataset selection after the fact is a real worry too, though I'd want to see the paper before accusing them of that.\n\nBottom line: this is a serious paper to engage with, not a desk reject. The topic is important, the approach is empirical, and the potential payoff for HRI practitioners is high. What's missing from the abstract is exactly the information a referee needs to judge fairness. I'd send it to peer review conditional on seeing a detailed protocol. For my own work, I wouldn't cite it until I've checked the full text.","headline":"Useful, safety-relevant ASR comparison for HRI; claims are plausible but unverifiable from the abstract—send to peer review if the full paper shows a fair protocol.","tokens_in":1457,"tokens_out":1896,"would_cite":false,"duration_ms":20090,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Benchmark scores do not predict how ASR systems behave in robot interactions.","keywords":["automatic speech recognition","human-robot interaction","speech foundation models","hallucination","benchmark evaluation","accented speech","noisy speech","spontaneous speech"],"falsifier":"Re-run the same four systems on the same eight datasets using the opposite decoding configurations and prompts; if the performance and hallucination gaps invert or vanish, the comparison was an artifact of settings rather than a property of the models. Alternatively, audit the dataset-selection process: if the datasets were chosen after seeing which ones amplified differences, the claimed variation is overstated.","tokens_in":638,"feed_emoji":"🗣️","tokens_out":1936,"duration_ms":26777,"temperature":0.7,"pith_summary":"This paper argues that four state-of-the-art speech recognition systems can score similarly on standard benchmarks yet diverge sharply when tested on the kinds of audio that matter for human-robot interaction. Across eight public datasets covering domain-specific, accented, noisy, age-variant, impaired, and spontaneous speech, the systems show significant differences in performance, hallucination tendencies, and inherent biases. The authors' point is that a high or similar benchmark number does not tell you whether a model will safely and reliably understand a child, an accented speaker, or a user in a noisy room. For robot designers, this means model choice should follow from failure analysis on the actual interaction conditions, not from leaderboard rank. If the claim holds, standard ASR evaluation practice is insufficient for HRI deployment decisions.","feed_headline":"Speech benchmarks don't predict real robot hearing","feed_subtitle":"Four top ASR systems diverge sharply on accented, noisy, child, and impaired speech despite similar scores.","key_machinery":"The evaluation design itself is the load-bearing mechanism: four ASR systems are compared across eight publicly available datasets, each selected to operationalize one of six specified dimensions of difficulty. The comparison is organized around performance, hallucination tendency, and bias, rather than a single aggregate score, so that differences hidden by benchmark averages become visible as per-condition failure profiles.","core_discovery":"The central claim is that four state-of-the-art ASR systems, despite similar scores on standard benchmarks, exhibit substantial and systematic variation when evaluated on eight datasets that capture six dimensions of real-world speech difficulty: domain-specific content, accent, noise, age variation, speech impairment, and spontaneity. The paper finds that these variations appear not only in raw recognition accuracy but also in hallucination tendencies and inherent biases, which are particularly dangerous in human-robot interaction because recognition errors can break task execution, erode user trust, and create safety hazards. The authors conclude that standard benchmark performance is not","pith_inferences":["An implicit consequence is that two ASR systems with identical aggregate word error rates can produce very different user experiences and safety risks, because errors may cluster on particular speakers, words, or acoustic conditions.","The observed variation probably traces to differences in training data composition among the models, which suggests a testable extension: probing each system with controlled accent and noise perturbations to build a per-system error distribution.","A further extension would be to map hallucination events to interaction outcomes, such as a robot acting on a nonexistent command, turning a static model evaluation into a deployability test.","The paper's logic also applies beyond speech: any foundation model evaluated only on aggregate benchmarks may hide systematic failures on the exact subpopulations that matter in embodied interaction."],"forward_implications":["Robot developers should test ASR systems on the specific user groups and acoustic conditions of their deployment instead of relying on standard benchmark rankings.","Voice interfaces for children, older adults, accented speakers, and users with speech impairments need condition-specific failure analysis before they are used in real interactions.","Hallucination rates, not just error rates, should be a reported characteristic of any ASR model considered for safety-relevant robot tasks.","A single leaderboard score is insufficient for model selection in HRI; per-dataset and per-user-group error profiles are the useful unit of comparison.","Recognition failures in HRI are not just usability problems but potential safety problems, so the way ASR is evaluated must change accordingly."],"supporting_citations":[],"fun_headline_variants":["Benchmark scores hide robot speech failures","Top ASR models flunk real-world robot talk","Robot hearing fails on accents, noise, and kids","Speech AI's hidden failure in robot interactions","When robot ears misfire: benchmark mirage exposed"],"cache_read_input_tokens":2688,"weakest_assumption_plain":"The result depends on the assumption that the eight selected datasets validly cover the six difficulty dimensions and that all four systems were compared under equivalent decoding, prompting, and scoring conditions, so that the observed differences reflect the systems themselves rather than the test setup.","fun_headline_variants_meta":{"raw":{"variants":["Benchmark scores hide robot speech failures","Top ASR models flunk real-world robot talk","Robot hearing fails on accents, noise, and kids","Speech AI's hidden failure in robot interactions","When robot ears misfire: benchmark mirage exposed"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000149,"raw_usage":{"total_tokens":967,"prompt_tokens":617,"completion_tokens":350,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":361,"completion_tokens_details":{"reasoning_tokens":279}},"tokens_in":361,"tokens_out":350,"duration_ms":5230,"temperature":1.0,"reasoning_tokens":279,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-05T16:44:34.403739+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run the same four systems on the same eight datasets using the opposite decoding configurations and prompts; if the performance and hallucination gaps invert or vanish, the comparison was an artifact of settings rather than a property of the models. Alternatively, audit the dataset-selection process: if the datasets were chosen after seeing which ones amplified differences, the claimed variation is overstated.","supporting_citations":[],"review_version":1}