{"id":"56e53668-b717-4f26-b7f2-d22259380735","arxiv_id":"2508.12733","paper_version":2,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":6.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":0,"one_line_summary":"A multilingual safety benchmark of 45k prompts in 12 languages shows LLM safety and helpfulness vary by language and domain.","lead":"This paper introduces LinguaSafe, a safety testing set of 45,000 prompts in 12 languages for large language models. It reports that models behave differently, and often less safely, when tested outside English.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Cross-lingual prompt equivalence is unvalidated; cross-language score gaps may be artifacts of translation/transcreation.","rationale":"The abstract makes a strong empirical claim about cross-language variation in safety scores, but the validity of that claim depends on the translated/transcreated/natively-sourced prompts being equivalent in difficulty and intent across languages. The reader's weakest_assumption correctly identifies this as the key unvalidated assumption. Since the full text is unavailable, we cannot verify whether the authors performed equivalence checks, but the abstract itself provides no evidence. This is exactly the load-bearing concern: if equivalence fails, the headline finding about cross-language variation is confounded. The proposed concrete test—human or model-based severity calibration on parallel prompts—would settle whether the concern lands. The verdict remains UNVERDICTED because the paper's internal methodology is not accessible; our concern reinforces that judgment without moving it.","tokens_in":784,"tokens_out":1623,"duration_ms":18923,"concrete_test":"Extract a matched sample of N=100 prompts per language that are parallel translations of the same seed prompts. Have bilingual annotators rate (a) semantic equivalence to the English source and (b) severity/offensiveness on a common scale. If mean severity differs significantly across languages (e.g., ANOVA p<0.05), then raw safety scores must be normalized before cross-language claims; if not, the concern is resolved.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central empirical claim—that safety scores vary significantly across languages even at similar resource levels—rests on the assumption that the translated, transcreated, and natively-sourced prompts are comparable in content and difficulty across all 12 languages. The abstract reports 45k entries but gives no evidence of equivalence validation: no human review, no back-translation checks, no severity calibration. If, for example, the Hungarian prompts are systematically more explicit or more culturally salient than the Malay prompts, the observed cross-language variation could reflect prompt difficulty rather than differences in model safety alignment. This is not an internal inconsistency but an unverified measurement assumption. Because the full text is unavailable, the decisive validation step cannot be assessed.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces LinguaSafe, a multilingual safety benchmark for LLMs containing 45k entries in 12 languages, built from translated, transcreated, and natively-sourced content. It claims a multidimensional framework covering direct and indirect safety as well as oversensitivity, and reports that evaluation results vary significantly across domains and languages, even among languages with similar resource levels. The abstract also states that the dataset and code are publicly released. The central claims are that the benchmark is comprehensive and linguistically authentic, and that the observed cross-language variation reflects genuine differences in model safety alignment.","tokens_in":934,"tokens_out":1686,"duration_ms":21794,"significance":"If the claims hold, LinguaSafe could be a valuable community resource: a large multilingual safety benchmark covering under-represented languages with fine-grained safety and helpfulness dimensions, released publicly. The emphasis on linguistic authenticity and native sourcing is a strength, and the multidimensional structure could enable more balanced safety alignment research. However, the current manuscript—as represented by the abstract—does not provide the methodological evidence needed to assess whether the cross-language variation is real or an artifact of prompt construction. The significance of the contribution therefore remains conditional on validation that is not visible in the abstract.","major_comments":[{"comment":"The central empirical finding—that safety scores 'vary significantly' across languages even with similar resource levels—rests on the assumption that translated, transcreated, and natively-sourced prompts are semantically, culturally, and difficulty-matched across the 12 languages. The abstract provides no evidence of equivalence validation: no human review, no back-translation checks, no severity calibration, and no comparison of prompt distributions. Without such validation, cross-language score differences could reflect prompt difficulty or cultural salience rather than model safety differences. This is the load-bearing measurement assumption and it is unverified in the abstract.","section":"Abstract (Data Curation)"},{"comment":"The claim that results 'vary significantly' is a statistical assertion, but the abstract reports neither effect sizes nor confidence intervals nor any significance-testing methodology. If the full text contains such analysis, it should be summarized or at least referenced in the abstract; as written, the abstract-level claim is unsupported. The multidimensional framework (direct safety, indirect safety, oversensitivity) is named but not defined, so it is impossible to judge whether the reported metrics measure what is claimed.","section":"Abstract (Evaluation Claims)"},{"comment":"The abstract states that the benchmark 'fills the void' in multilingual safety evaluation and is 'comprehensive,' but no comparison against existing multilingual benchmarks (e.g., existing safety datasets covering some of the same languages) is presented. Comprehensiveness is a relative claim; without a baseline or coverage analysis, this assertion is not yet supported. If the full text includes such comparisons, they should be visible in the abstract to allow assessment of the contribution's novelty.","section":"Abstract (Comprehensiveness)"}],"minor_comments":[{"comment":"The phrase 'from Hungarian to Malay' is used twice and is vague as a language list; framing a language set by two endpoints is misleading and should be replaced by an explicit list or a clearer characterization (e.g., '12 languages including Hungarian, Malay, and others').","section":"Abstract (Wording)"},{"comment":"The phrase 'fills the void' is promotional. The abstract would be more persuasive with neutral language such as 'addresses a gap' or 'provides coverage for under-represented languages.'","section":"Abstract (Style)"},{"comment":"The sentence beginning 'Curated using a combination of translated, transcreated, and natively-sourced data' has a dangling modifier: the dataset is curated, not 'our dataset addresses.' Rephrase for clarity.","section":"Abstract (Grammar)"},{"comment":"The statement 'Our dataset and code are released to the public' is not accompanied by a URL or repository identifier in the abstract. While this may be intentional for anonymized submission, the claim should be verifiable in the final version.","section":"Abstract (Release Statement)"}],"recommendation":"uncertain","confidential_remarks":"The review is based solely on the abstract because the full text was not made available. The central concern—cross-lingual prompt equivalence—cannot be resolved without seeing the full methodology. If the full text provides human validation or back-translation checks, the paper may be sound; if not, the main claim would be at risk. I would need to see the full manuscript to make a definitive recommendation."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nYou asked for my read on LinguaSafe. I'm working from the abstract only, so everything below is provisional. The headline: it's a plausible, potentially useful resource, but the abstract alone doesn't give enough to judge whether the central claim of cross-language variation is real.\n\nWhat's genuinely new: 45k entries across 12 languages, with a mix of translated, transcreated, and natively-sourced prompts, plus a fine-grained framework separating direct, indirect, and oversensitivity evaluations. That combination isn't standard, and the public release is a plus. If the dataset is clean, it would be a solid addition to the multilingual safety toolkit.\n\nWhere it's soft: the abstract offers no validation of the prompts' cross-language equivalence. The stress-test note is right: if Hungarian and Malay prompts differ in explicitness or cultural salience, the reported differences in model scores could be an artifact. No human review, back-translation, or severity calibration is mentioned. Also, there's no comparison with existing multilingual safety benchmarks, so the novelty claim is under-supported. And the central finding of 'significant variation' is stated without numbers or error bars. These are not fatal flaws, but they are the load-bearing parts, and they're unaddressed in the abstract.\n\nThe paper may well do all this in the full text. But as it stands, the abstract reads like a well-scoped engineering contribution, not a validated measurement study.\n\nRecommendation: this deserves peer review, not desk rejection, because the dataset could be valuable and the authors are making it public. The referees should press hard on equivalence validation and on whether the benchmark's construction biases the cross-language comparisons. For my own work, I wouldn't cite it until I've seen the validation, but I'd bring it to the reading group to see what the full paper offers.","headline":"A plausible multilingual safety benchmark that deserves referee scrutiny, but the abstract alone doesn't support the cross-language variation claim without equivalence validation.","tokens_in":1398,"tokens_out":1617,"would_cite":false,"duration_ms":17257,"reading_group":"maybe","serious_thinker":"unclear","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A 45,000-prompt benchmark finds LLM safety varies widely across 12 languages.","keywords":["multilingual safety benchmark","LLM safety","safety alignment","indirect safety","oversensitivity","under-represented languages","cross-lingual evaluation","transcreated prompts"],"falsifier":"Have bilingual raters back-translate the natively-sourced prompts to English and score semantic and cultural equivalence against their translated counterparts; also compare model safety scores on translated versus natively-sourced prompts within the same harm category. If one language's prompts are systematically more explicit or more ambiguous, the reported cross-language differences would reflect prompt artifacts rather than model safety alignment.","tokens_in":717,"feed_emoji":"🌐","tokens_out":2029,"duration_ms":23636,"temperature":0.7,"pith_summary":"This paper introduces LinguaSafe, a multilingual safety benchmark built from 45,000 prompts in 12 languages, mixing translated, transcreated, and natively-sourced content. The authors aim to fill a gap in safety evaluation for under-represented languages and to provide fine-grained measurements of direct and indirect safety, plus oversensitivity. Their central claim is that model safety alignment varies significantly across domains and languages, even among languages with similar resource levels. If true, this means global deployment of LLMs needs language-specific safety auditing rather than assuming one alignment standard suffices.","feed_headline":"Safety alignment varies sharply across languages, 45k-prompt benchmark finds","feed_subtitle":"LinguaSafe tests direct and indirect safety plus oversensitivity in 12 languages, from Hungarian to Malay.","key_machinery":"LinguaSafe itself is the central mechanism: a dataset and evaluation framework that combines three prompt sources—translated, transcreated, and natively-sourced—to achieve linguistic authenticity, and a multidimensional rubric separating direct safety, indirect safety, and oversensitivity. It provides the data and metrics that make cross-language comparison possible; the three-source curation is designed to ensure prompts are culturally and linguistically natural rather than merely literal translations.","core_discovery":"The core discovery is that a multidimensional multilingual benchmark can reveal substantial cross-lingual variation in LLM safety behavior. LinguaSafe organizes 45,000 prompts across 12 languages, scoring models on direct safety (refusing harmful requests), indirect safety (handling subtly unsafe contexts), and oversensitivity (over-refusals of benign requests). The authors report that results vary significantly across domains and languages, including languages with comparable resource levels, which they interpret as evidence that current safety alignment is not balanced across languages.","pith_inferences":["The design of the benchmark implicitly argues that transcreated and natively-sourced prompts capture culturally specific harms that translation alone misses; a testable extension is to compare safety scores on the three prompt types to quantify how much cultural adaptation matters.","If translation artifacts inflate cross-language variation, the reported differences might partly reflect prompt difficulty rather than alignment; the paper does not provide equivalence validation, so this confound remains open.","A practical consequence the authors leave implicit: safety regulators and deployers could use LinguaSafe-style benchmarks to require per-language safety reporting for LLMs serving multilingual populations."],"forward_implications":["If LinguaSafe is valid, multilingual safety evaluation should include indirect safety and oversensitivity, not just direct refusals.","The observed cross-language variation implies that a model that passes safety tests in one language cannot be assumed safe in another, even when the languages have similar resource levels.","The public release of the dataset and code enables researchers and developers to audit their own models across these 12 languages and target specific domains or languages for alignment improvements.","The fine-grained framework allows per-domain and per-language scoring, which could help identify whether safety failures come from cultural misunderstanding, translation artifacts, or genuine alignment gaps."],"supporting_citations":[],"fun_headline_variants":["LLM safety varies by language, 45k-prompt benchmark shows","Multilingual safety benchmark finds uneven alignment across 12 languages","LinguaSafe benchmark reveals safety gaps between languages","Safety alignment inconsistent across languages, new benchmark finds","Benchmark of 45k prompts shows LLM safety differs per language"],"cache_read_input_tokens":2816,"weakest_assumption_plain":"The benchmark's validity depends on the assumption that translated, transcreated, and natively-sourced prompts are equally difficult and culturally meaningful across the 12 languages, yet the paper offers no evidence of equivalence validation such as human review or back-translation checks.","fun_headline_variants_meta":{"raw":{"variants":["LLM safety varies by language, 45k-prompt benchmark shows","Multilingual safety benchmark finds uneven alignment across 12 languages","LinguaSafe benchmark reveals safety gaps between languages","Safety alignment inconsistent across languages, new benchmark finds","Benchmark of 45k prompts shows LLM safety differs per language"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000624,"raw_usage":{"total_tokens":2722,"prompt_tokens":736,"completion_tokens":1986,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":480,"completion_tokens_details":{"reasoning_tokens":1902}},"tokens_in":480,"tokens_out":1986,"duration_ms":15806,"temperature":1.0,"reasoning_tokens":1902,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-05T19:17:23.548111+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Have bilingual raters back-translate the natively-sourced prompts to English and score semantic and cultural equivalence against their translated counterparts; also compare model safety scores on translated versus natively-sourced prompts within the same harm category. If one language's prompts are systematically more explicit or more ambiguous, the reported cross-language differences would reflect prompt artifacts rather than model safety alignment.","supporting_citations":[],"review_version":1}