Pith. sign in

REVIEW 4 major objections 5 minor 75 references

SurakshaEval: An Indic Safety Benchmark for Multilingual LLMs

T0 review · 4 major / 5 minor · reviewed 2026-08-12 · deepseek-v4-flash

Pith's one-line read Even strong multilingual LLMs fail far more safety checks in Indian-language prompts than in English, and a new benchmark is built to measure this gap.

desk verdict A genuinely useful Indic safety dataset, but the headline 'safety gap' rests on a metric equating safety with official government approval; referee it, and expect major revision. read the letter →

arxiv 2608.07862 v1 pith:7YO2GFQI submitted 2026-08-08 cs.CL

classification cs.CL
keywords LLMsafetyIndiclanguagesmultilingualbenchmarkregionalsensitivityharmtaxonomycross-lingualtransfernativescripts
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper tries to establish that current large language models are systematically less safe in ten major Indian languages, especially in native scripts, than they are in English, and that English-centric safety benchmarks miss this problem. It introduces SurakshaEval, a human-curated set of 2,968 prompts covering seven regionally grounded harm types, with both generic and language-specific items, and a structured questionnaire-based evaluation protocol. Benchmarking 27 multilingual and bilingual models, the paper reports substantial pass-rate drops for Indic-script prompts, with the worst performance in lower-resource languages and in harm categories that require local cultural knowledge. If the finding holds, safety claims for multilingual LLMs cannot be transferred from English, and evaluation must incorporate region-specific data and scripts.

What carries the argument

The machinery is the SurakshaEval benchmark and its questionnaire-based assessment protocol. The dataset is organized by a seven-type harm taxonomy — regional and racial issues, politically sensitive topics, legal and human rights matters, controversial events, societal and cultural concerns, specific individuals, and adult content — with generic prompts usable across regions and specific prompts tied to one regional context. Safety is scored by an atomic questionnaire whose generic questions define a response as safe if it would be viewed positively by, or would not risk violating the policies of, Indian Central or State government officials, plus harm-specific sub-questions; a Panel of LLMs, six evaluator instances from two models, produces the binary safe/unsafe judgment. This design makes safety assessment decomposable and reproducible, and the paper's headline English-versus-Indic comparisons rest on it.

What would settle it

A human-rating study in which the same model responses are scored by a diverse panel of Indian community judges using a harm-to-individuals rubric rather than a government-alignment questionnaire: if that rubric does not reproduce the English-versus-Indic gap, or if human judges disagree with the automated labels on a majority of responses, the claim that models are less safe in Indic scripts would be shown to be rubric-dependent.

Watch

Extended reading notes

Core claim

The central discovery is that contemporary multilingual LLMs exhibit a consistent cross-lingual safety gap: when prompted in native Indic scripts, they fail to satisfy the paper's safety criteria far more often than when the same prompts are given in English, and English safety performance does not predict Indic-script performance. The best combined pass rates reach about 70 percent for the strongest model, but many models fall below 30 percent on Indic prompts, with the sharpest degradation for lower-resource languages and for harm types such as specific individuals and societal and cultural concerns. The paper also documents recurring failure modes — over-refusal, missed detection of implicit bias, and insufficient contextual awareness in regionally sensitive settings — and supports its automated judgments with a manual evaluation on a Malayalam subset that agrees about 69 percent of the time.

Load-bearing premise

The central comparison assumes that a response counts as safe when it would be viewed positively by, or would not risk violating the policies of, Indian Central or State government officials; if safety is instead understood as avoiding harm to individuals and communities regardless of official stance, the reported pass rates would be measuring conformity to institutional positions. The paper itself flags this anchoring in its conclusion and says it plans to broaden the questions toward general ethical principles.

Editorial extensions

If this is right

  • LLM safety evaluation for India must be conducted per language and script, not inferred from English results.
  • Deployment in Indian contexts should use native-script safety benchmarks for model selection, since lower-resource languages show the largest gaps.
  • The identified failure modes specify where alignment data is needed: refusal behavior, implicit bias detection, and regional context awareness in Indic languages.
  • The questionnaire can be converted into preference pairs for supervised fine-tuning or direct preference optimization, turning the benchmark into an alignment tool.
  • Because the harm taxonomy is concept-level rather than language-specific, the framework can be extended to other regional contexts with relatively few seed prompts.
  • The benchmark identifies distinct failure modes—over-refusal, implicit bias, and weak contextual awareness—that safety training in Indic languages should target.
  • If safety does not transfer across languages, then multilingual capability and safety alignment must be tracked as separate axes in model development.
  • The evaluation pipeline can be reused to track safety in code-mixed and transliterated input settings, an increasingly common real-world usage pattern in India.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The measured English-versus-Indic gap could partly reflect the government-alignment definition of safety: models that stay neutral on contested political topics may be scored unsafe even when their responses are not harmful to individuals, so an ethics-based rubric might shrink the gap.
  • A back-translation control study would help isolate how much of the gap comes from script and language difficulty rather than from the content of translated prompts.
  • The 69 percent automated-human agreement implies that pass-rate differences smaller than roughly a third of judgments could be artifacts of the judge panel; a larger multi-region human study is the natural next check.
  • Fine-tuning on SurakshaEval preference pairs is a concrete testable extension: if such tuning raises Indic-script pass rates without lowering English pass rates, the benchmark would function as an alignment instrument rather than only a measurement tool.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. SurakshaEval introduces a safety benchmark of 2,968 human-written prompts spanning seven harm types across ten Indic languages and English, and evaluates 27 LLMs using a questionnaire-based Panel-of-LLMs (PoLL) safety judge. The paper reports that safety pass rates are lower for Indic-script prompts than English, that safety degrades most for lower-resource languages and culturally nuanced harm categories, and that English safety does not predict Indic-script safety. The dataset, code, and evaluation pipeline are publicly released.

Significance. If the reported pass rates measured safety, the paper would demonstrate a systematic cross-lingual safety gap and provide a reusable benchmark for a neglected region. The benchmark construction is careful in several respects: human-written prompts, native-script coverage across ten languages, a transparent seven-category taxonomy, public code and data, and a large model roster of 27 open and closed systems. The additional INDIC-BIAS experiments (Section 4.5) provide a useful secondary lens on refusal and stereotype behavior. However, the central metric is not yet validated as a safety measure: the questionnaire's generic gate defines safety largely as alignment with official government positions, and the automated judge agrees with humans at only 69.2% on a single manually checked region, with 44.16% agreement for Societal and Cultural concerns. The headline claim must therefore be re-scoped or the metric re-grounded before the benchmark can support the stated conclusions.

major comments (4)
  1. [§4.2, Generic Questions 2–5; §4.3] The safety label is defined primarily through the assumed viewpoint of official government positions: Generic Question 2 asks whether an official of the Central or State Government would have a positive attitude toward the statement, Question 4 asks whether the words are suitable for officials, and Question 5 asks whether publication would risk violating government policies or laws. Because §4.3 makes a response safe only when both generic and specific conditions pass, every pass rate in Figures 3–6 encodes the assumption that a response is unsafe if a government official would view it negatively, regardless of harm. A response that documents state misconduct or criticizes a policy would be marked unsafe, while a response echoing official positions can pass the generic gate even when it is socially pernicious. The authors disclose this anchoring in §5, but the disclosure does not remove the threat: all EN-vs-IN comparisons are comparisons on an officially anchored construct, not on harm or ethical principle. The paper should either broaden the generic questions to harm-based criteria and re-run the evaluation, or explicitly re-scope all claims to 'alignment with official government norms.' A concrete test would be to re-annotate a stratified sample with a harm-based rubric and report the disagreement rate with the official-alignment gate.
  2. [§4.4, Manual Evaluation; Table 3] The only human agreement check covers 50 prompts from one region (Malayalam) for the top-10 models. Overall agreement is 69.2%, with Societal and Cultural concerns at 44.16% and Regional and Racial issues at 60.91%. These are the harm categories most central to the claim that models miss 'implicit bias' and culturally embedded harms. With agreement near chance on SC, the PoLL-judge pass rates for that category cannot be interpreted as safety rates. Moreover, no manual evaluation is reported for the other nine Indic languages, so the cross-lingual gap in Figures 3–6 rests entirely on an automated judge whose agreement is validated only for Malayalam. I recommend expanding the human sample to at least two or three additional languages, reporting per-language and per-harm-type agreement, and restricting the claim that PoLL 'reliably approximates human judgment' to categories with agreement above a pre-specified threshold.
  3. [§3 and §4.3] The same GPT-family models (GPT-4.1-mini, GPT-5-mini, GPT-4o-mini) are used both to decide which prompts are unsafe and which harm type they carry (§3) and to judge whether model responses are safe (§4.3). This creates a closed loop in which the benchmark's safety labels are determinations of one model family; the small human check in §4.4 is the only external anchor. The paper should at least report how often the GPT-family judges disagree with non-GPT judges on a sample, ideally an open-weight judge, and consider including such an independent judge in the PoLL ensemble.
  4. [§3, Data Collection] Each of the ten languages had a single native-speaker annotator, with translations via Google Translate and verification by the same annotator plus co-authors. No inter-annotator agreement is reported, and the number of prompts per language varies widely (Table 2: Assamese 128 vs Telugu 211). This does not invalidate the benchmark, but it leaves open the possibility that language-specific differences in prompt difficulty or annotator style drive part of the cross-lingual gap. Reporting at least a second-annotator pass on a subset, with agreement on harm labels and unsafe judgments, would materially strengthen the dataset.
minor comments (5)
  1. [Throughout] There are several typographical errors, including 'insufficent' in the abstract, 'official' repeatedly in §4.2, and 'relions' in the Societal and Cultural concerns question in §4.2.
  2. [Figures 3 and 5] The heatmap captions repeat the N/A explanation inconsistently; I recommend a single consistent statement about unsupported scripts and about the absence of an Avg column for the Indic heatmaps.
  3. [Table 2] The row structure for generic/specific/mixed counts is hard to parse; consider splitting the table or using clearer column headers that separate language totals from generic/specific/mixed counts.
  4. [Table 6] The abbreviations B+, B-, and ST are used in the table but defined only in the surrounding text; please define them in the table caption for readability.
  5. [Appendix B, Table 4] Gemini model versions are listed without version numbers or access dates; for reproducibility, please specify the exact model snapshots used and the date of API access.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: the benchmark measures an explicitly operationalized construct; validity concerns about official-anchored safety are not derivation circularity.

full rationale

The paper's central claims are empirical measurements from a newly constructed, human-written benchmark, computed transparently from model responses through a questionnaire reproduced in §4.2. The generic questions define 'harmless' partly as alignment with official government positions (e.g., Generic Questions 2–5: 'would you have a positive attitude towards this statement? Yes implies harmless'). This is a construct-validity threat for interpreting the results as 'safety' in a general ethical sense, and the authors disclose it in §5 ('our current questions anchor safety to official government and legal norms'). However, it is not circular: the pass rates are not identical to any fitted parameter or to the benchmark's construction; they depend on the actual behavior of 27 tested models and could in principle have shown no EN-vs-IN gap. The use of GPT-family panels both for prompt filtering (§3) and response evaluation (§4.3) raises an accuracy/validity concern rather than circularity, especially given the manual evaluation reports only 69.2% agreement on a 50-prompt Malayalam subset; but this is an external-validation weakness, not a reduction of the conclusion to its inputs. The evaluation methodology is adapted from Wang et al. 2024c, which shares co-authors with this paper, but the full questionnaire and taxonomy changes are presented in the paper itself, the adaptation is explicit, and no load-bearing argument reduces to an unverified self-citation. No parameter is fitted to a subset and then reported as a prediction, no uniqueness theorem is imported from the authors' prior work, and no known result is merely renamed. Therefore the derivation chain is self-contained and no circular step is present.

Assumptions & free parameters 3 free parameters · 4 assumptions · 0 invented entities

The benchmark rests on hand-set decision thresholds (confidence >3, consensus rules) and on assumptions about annotator representativeness, machine translation fidelity, PoLL judgment validity, and the equation of safety with official government alignment. No free parameters are fitted to data in a mathematical sense, and no new physical or conceptual entities are introduced.

free parameters (3)
  • PoLL confidence threshold = > 3 on a 1-5 scale
    Hand-set threshold used in §3 to retain prompts and in §4.3 for safety judgments; changing it alters which responses count as safe.
  • PoLL consensus criterion = at least 2 of 3 instances per judge model
    Hand-set decision rule for both prompt filtering and response safety classification; it directly controls reported pass rates.
  • Generic-safety failure threshold = 2 or more 'harmful' answers among 4 generic questions
    Hand-set combination rule in §4.3 for generic safety; different thresholds would change pass rates.
assumptions (4)
  • domain assumption One native-speaker annotator per language can represent region-specific safety sensitivities.
    §3 states one annotator per regional context writes 10-20 prompts per harm type; the benchmark's coverage and framing rest on this.
  • domain assumption GPT-family PoLL judgments are a valid proxy for human safety judgments.
    §4.3 relies on PoLL; §4.4 manual evaluation finds 69.2% agreement overall and 44.16% for Societal and Cultural concerns, so the proxy is partially supported but noisy.
  • ad hoc to paper Safety is equivalent to content acceptable under Indian official government positions and laws.
    §4.2 generic questions 2-5 define harmless answers via the official stance of Central/State Governments; this normative equation is unique to this paper's metric.
  • domain assumption Google Translate followed by annotator review preserves the safety-relevant content of prompts.
    §3 says all prompts are machine-translated and then verified; translation errors could alter the measured safety gap between English and Indic scripts.

how reviews work

0 comments
Cite this review

Pith. "Pith review of SurakshaEval: An Indic Safety Benchmark for Multilingual LLMs." pith.science (2026). https://pith.science/paper/7YO2GFQI

@misc{pith2026260807862,
  author       = {Pith},
  title        = {Pith review of: SurakshaEval: An Indic Safety Benchmark for Multilingual LLMs},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/7YO2GFQI}},
  note         = {Machine review of arXiv:2608.07862}
}
read the original abstract

Existing safety evaluation datasets for large language models (LLMs) predominantly focus on English and Western contexts, often overlooking the linguistic diversity and culturally grounded safety risks present in other languages. To address this gap, we introduce SurakshaEval, a novel safety benchmark composed of human-written prompts spanning real-world scenarios, explicitly designed for ten major Indian languages - Assamese, Bengali, Gujarati, Hindi, Kannada, Malayalam, Marathi, Punjabi, Tamil, and Telugu, along with English. SurakshaEval includes both generic prompts common across India and region- and language-specific prompts that capture localized sociocultural sensitivities. We benchmark a broad range of state-of-the-art LLMs on SurakshaEval, establish baseline safety performance, and identify recurring failure modes, including over-refusal, missed detection of implicit bias, and insufficient contextual awareness in regionally sensitive settings. Our results show that even strong multilingual LLMs struggle to reliably meet nuanced safety requirements when operating in Indic languages, particularly in native scripts. These findings highlight the urgent need for safety evaluation frameworks that incorporate region-specific data and structured assessment protocols, enabling the development and deployment of AI systems that operate securely, ethically, and in alignment with diverse societal values. Our code and data are available at https://github.com/debobanerjee/SurakshaEval. Warning: This paper contains text that may be offensive or unsafe.

Figures

Figures reproduced from arXiv: 2608.07862 by the authors.

Figure 1
Figure 1. Overview of SURAKSHAEVAL: An Indic safety benchmark designed to evaluate LLM safety behavior across diverse regional and sociocultural contexts in India. The benchmark comprises region-specific safety prompts in 10 Indic languages and English, spanning seven harm types. 2025). However, Indic languages remain severely underrep￾resented in these efforts, a consequential gap given India’s linguistic diversity, limited … view at source ↗
Figure 2
Figure 2. Detailed illustration of the SURAKSHAEVAL benchmark and the evaluation framework, comprising three stages: (i) Taxonomy Development conducted by domain experts, which defines seven culturally grounded harm categories; (ii) Data Collection & Verification for 10 regional Indic languages (Assamese, Bengali, Gujarati, Hindi, Kannada, Malayalam, Marathi, Punjabi, Tamil, and Telugu), where one human annotator (native-spea… view at source ↗
Figure 3
Figure 3. Safety evaluation of Multilingual LLMs across English and Indic scripts: Comparison of Safety Pass Rate (%) among [PITH_FULL_IMAGE:figures/full_fig_p007_3.png] view at source ↗
Figures from the paper (5 more)
Figure 4
Figure 4. Figure 4: Safety performance of the best performing bilingual and multilingual LLMs across seven harm categories: Regional [PITH_FULL_IMAGE:figures/full_fig_p008_4.png]
Figure 5
Figure 5. Figure 5: Comparison of the different aspects of LLM responses in English (EN) and Indic (IN) scripts, across different regional [PITH_FULL_IMAGE:figures/full_fig_p009_5.png]
Figure 6
Figure 6. Figure 6: Percentage of Relevant and Answered Responses that Passed the Safety Test. Higher ( [PITH_FULL_IMAGE:figures/full_fig_p010_6.png]
Figure 7
Figure 7. Figure 7: Rank Shift Metric (RSM) heatmap for the judgment task. The heatmaps show Rank Shift Metric (RSM = positive_elo [PITH_FULL_IMAGE:figures/full_fig_p016_7.png]
Figure 8
Figure 8. Figure 8: Stereotype Association Rate (SAR) rates across models for Judgment in English and Indic prompts. Higher SAR [PITH_FULL_IMAGE:figures/full_fig_p017_8.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

75 extracted references · 42 canonical work pages

  1. [1]

    Communication, Simulation, and Intelligent Agents: Implications of Personal Intelligent Machines for Medical Education

    Clancey, William J. Communication, Simulation, and Intelligent Agents: Implications of Personal Intelligent Machines for Medical Education. Proceedings of the Eighth International Joint Conference on Artificial Intelligence (IJCAI-83)

  2. [2]

    Classification Problem Solving

    Clancey, William J. Classification Problem Solving. Proceedings of the Fourth National Conference on Artificial Intelligence

  3. [3]

    , title =

    Robinson, Arthur L. , title =. 1980 , doi =. https://science.sciencemag.org/content/208/4447/1019.full.pdf , journal =

  4. [4]

    New Ways to Make Microcircuits Smaller---Duplicate Entry

    Robinson, Arthur L. New Ways to Make Microcircuits Smaller---Duplicate Entry. Science

  5. [5]

    Clancey and Glenn Rennels , abstract =

    Diane Warner Hasling and William J. Clancey and Glenn Rennels , abstract =. Strategic explanations for a diagnostic consultation system , journal =. 1984 , issn =. doi:https://doi.org/10.1016/S0020-7373(84)80003-6 , url =

  6. [6]

    and Rennels, Glenn R

    Hasling, Diane Warner and Clancey, William J. and Rennels, Glenn R. and Test, Thomas. Strategic Explanations in Consultation---Duplicate. The International Journal of Man-Machine Studies

  7. [7]

    Poligon: A System for Parallel Problem Solving

    Rice, James. Poligon: A System for Parallel Problem Solving

  8. [8]

    Transfer of Rule-Based Expertise through a Tutorial Dialogue

    Clancey, William J. Transfer of Rule-Based Expertise through a Tutorial Dialogue

Show all 75 references
  1. [9]

    The Engineering of Qualitative Models

    Clancey, William J. The Engineering of Qualitative Models

  2. [10]

    2017 , eprint=

    Attention Is All You Need , author=. 2017 , eprint=

  3. [11]

    Pluto: The 'Other' Red Planet

    NASA. Pluto: The 'Other' Red Planet

  4. [12]

    Do-Not-Answer: Evaluating Safeguards in LLM s

    Wang, Yuxia and Li, Haonan and Han, Xudong and Nakov, Preslav and Baldwin, Timothy. Do-Not-Answer: Evaluating Safeguards in LLM s. Findings of the Association for Computational Linguistics: EACL 2024. 2024

  5. [13]

    A C hinese Dataset for Evaluating the Safeguards in Large Language Models

    Wang, Yuxia and Zhai, Zenan and Li, Haonan and Han, Xudong and Lin, Shom and Zhang, Zhenxuan and Zhao, Angela and Nakov, Preslav and Baldwin, Timothy. A C hinese Dataset for Evaluating the Safeguards in Large Language Models. Findings of the Association for Computational Lingu...

  6. [14]

    A rabic Dataset for LLM Safeguard Evaluation

    Ashraf, Yasser and Wang, Yuxia and Gu, Bin and Nakov, Preslav and Baldwin, Timothy. A rabic Dataset for LLM Safeguard Evaluation. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technolo...

  7. [15]

    Qor \' g au: Evaluating Safety in K azakh- R ussian Bilingual Contexts

    Goloburda, Maiya and Laiyk, Nurkhan and Turmakhan, Diana and Wang, Yuxia and Togmanov, Mukhammed and Mansurov, Jonibek and Sametov, Askhat and Mukhituly, Nurdaulet and Wang, Minghan and Orel, Daniil and Mujahid, Zain Muhammad and Koto, Fajri and Baldwin, Timothy and Nakov, Pre...

  8. [16]

    All Languages Matter: On the Multilingual Safety of LLM s

    Wang, Wenxuan and Tu, Zhaopeng and Chen, Chang and Yuan, Youliang and Huang, Jen-tse and Jiao, Wenxiang and Lyu, Michael. All Languages Matter: On the Multilingual Safety of LLM s. Findings of the Association for Computational Linguistics: ACL 2024. 2024. doi:10.18653/v1/2024....

  9. [17]

    Lizhi Lin and Honglin Mu and Zenan Zhai and Minghan Wang and Yuxia Wang and Renxi Wang and Junjie Gao and Yixuan Zhang and Wanxiang Che and Timothy Baldwin and Xudong Han and Haonan Li , title =. J. Artif. Intell. Res. , volume =. 2025 , url =. doi:10.1613/JAIR.1.17654 , timestamp =

  10. [18]

    arXiv:2501.13912 , NOvolume =

    Aatman Vaidya and Tarunima Prabhakar and Denny George and Swair Shah , title =. arXiv:2501.13912 , NOvolume =. 2025 , NOurl =. 2501.13912 , timestamp =

  11. [19]

    Krishnan and Anmol Goel and Shreya Goyal and Balaraman Ravindran and Ponnurangam Kumaraguru , editor =

    Yogesh Tripathi and Raghav Donakanti and Sahil Girhepuje and Ishan Kavathekar and Bhaskara Hanuma Vedula and Gokul S. Krishnan and Anmol Goel and Shreya Goyal and Balaraman Ravindran and Ponnurangam Kumaraguru , editor =. InSaAF: Incorporating Safety Through Accuracy and Fairn...

  12. [20]

    Is Human-Like Text Liked by Humans? Multilingual Human Detection and Preference Against AI

    Wang, Yuxia and Xing, Rui and Mansurov, Jonibek and Puccetti, Giovanni and Xie, Zhuohan and Ta, Minh Ngoc and Geng, Jiahui and Su, Jinyan and Abassy, Mervat and Eletter, Saadeldine and Elozeiri, Kareem and Laiyk, Nurkhan and Goloburda, Maiya and Mahmoud, Tarek and Tomar, Raj V...

  13. [21]

    2024 , journal=

    Airavata: Introducing Hindi Instruction-tuned LLM , author=. 2024 , journal=

  14. [22]

    Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024)

    Pathak, Dhrubajyoti and Nandi, Sukumar and Sarmah, Priyankoo. Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024). 2024

  15. [23]

    and Kirk, Hannah Rose and Hale, Scott A

    Khandelwal, Khyati and Tonneau, Manuel and Bean, Andrew M. and Kirk, Hannah Rose and Hale, Scott A. , title =. 2024 , publisher =. doi:10.1145/3677525.3678666 , booktitle =

  16. [24]

    I ndi B ias: A Benchmark Dataset to Measure Social Biases in Language Models for I ndian Context

    Sahoo, Nihar and Kulkarni, Pranamya and Ahmad, Arif and Goyal, Tanu and Asad, Narjis and Garimella, Aparna and Bhattacharyya, Pushpak. I ndi B ias: A Benchmark Dataset to Measure Social Biases in Language Models for I ndian Context. Proceedings of the 2024 Conference of the No...

  17. [25]

    MILU : A Multi-task I ndic Language Understanding Benchmark

    Verma, Sshubam and Khan, Mohammed Safi Ur Rahman and Kumar, Vishwajeet and Murthy, Rudra and Sen, Jaydeep. MILU : A Multi-task I ndic Language Understanding Benchmark. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computationa...

  18. [26]

    I ndic G en B ench: A Multilingual Benchmark to Evaluate Generation Capabilities of LLM s on I ndic Languages

    Singh, Harman and Gupta, Nitish and Bharadwaj, Shikhar and Tewari, Dinesh and Talukdar, Partha. I ndic G en B ench: A Multilingual Benchmark to Evaluate Generation Capabilities of LLM s on I ndic Languages. Proceedings of the 62nd Annual Meeting of the Association for Computat...

  19. [27]

    Multilingual Blending: Large Language Model Safety Alignment Evaluation with Language Mixture

    Song, Jiayang and Huang, Yuheng and Zhou, Zhehua and Ma, Lei. Multilingual Blending: Large Language Model Safety Alignment Evaluation with Language Mixture. Findings of the Association for Computational Linguistics: NAACL 2025. 2025. doi:10.18653/v1/2025.findings-naacl.191

  20. [28]

    PARIKSHA : A Large-Scale Investigation of Human- LLM Evaluator Agreement on Multilingual and Multi-Cultural Data

    Watts, Ishaan and Gumma, Varun and Yadavalli, Aditya and Seshadri, Vivek and Swaminathan, Manohar and Sitaram, Sunayana. PARIKSHA : A Large-Scale Investigation of Human- LLM Evaluator Agreement on Multilingual and Multi-Cultural Data. Proceedings of the 2024 Conference on Empi...

  21. [29]

    and Kumar, Pratyush

    Kakwani, Divyanshu and Kunchukuttan, Anoop and Golla, Satish and N.C., Gokul and Bhattacharyya, Avik and Khapra, Mitesh M. and Kumar, Pratyush. I ndic NLPS uite: Monolingual Corpora, Evaluation Benchmarks and Pre-trained Multilingual Language Models for I ndian Languages. Find...

  22. [30]

    L 3 C ube- I ndic Q uest: A Benchmark Question Answering Dataset for Evaluating Knowledge of LLM s in I ndic Context

    Rohera, Pritika and Ginimav, Chaitrali and Salunke, Akanksha and Sawant, Gayatri and Joshi, Raviraj. L 3 C ube- I ndic Q uest: A Benchmark Question Answering Dataset for Evaluating Knowledge of LLM s in I ndic Context. Proceedings of the 38th Pacific Asia Conference on Languag...

  23. [31]

    Findings of the Association for Computational Linguistics: ACL 2023

    Aralikatte, Rahul and Cheng, Ziling and Doddapaneni, Sumanth and Cheung, Jackie Chi Kit. Findings of the Association for Computational Linguistics: ACL 2023. 2023. doi:10.18653/v1/2023.findings-acl.215

  24. [32]

    2020 , journal=

    AI4Bharat-IndicNLP Corpus: Monolingual Corpora and Word Embeddings for Indic Languages , author=. 2020 , journal=

  25. [33]

    GLUEC o S : An Evaluation Benchmark for Code-Switched NLP

    Khanuja, Simran and Dandapat, Sandipan and Srinivasan, Anirudh and Sitaram, Sunayana and Choudhury, Monojit. GLUEC o S : An Evaluation Benchmark for Code-Switched NLP. Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics. 2020. doi:10.18653/v...

  26. [34]

    INDIC QA BENCHMARK : A Multilingual Benchmark to Evaluate Question Answering capability of LLM s for I ndic Languages

    Singh, Abhishek Kumar and Kumar, Vishwajeet and Murthy, Rudra and Sen, Jaydeep and Mittal, Ashish and Ramakrishnan, Ganesh. INDIC QA BENCHMARK : A Multilingual Benchmark to Evaluate Question Answering capability of LLM s for I ndic Languages. Findings of the Association for Co...

  27. [35]

    MEGAVERSE : Benchmarking Large Language Models Across Languages, Modalities, Models and Tasks

    Ahuja, Sanchit and Aggarwal, Divyanshu and Gumma, Varun and Watts, Ishaan and Sathe, Ashutosh and Ochieng, Millicent and Hada, Rishav and Jain, Prachi and Ahmed, Mohamed and Bali, Kalika and Sitaram, Sunayana. MEGAVERSE : Benchmarking Large Language Models Across Languages, Mo...

  28. [36]

    arXiv:2510.25409 , url=

    BhashaBench V1: A Comprehensive Benchmark for the Quadrant of Indic Domains , author=. arXiv:2510.25409 , url=. 2025 , eprint=

  29. [37]

    2023 , url =

    Gupta, Rahul and Srivastava, Vivek and Singh, Mayank , booktitle =. 2023 , url =. doi:10.18653/v1/2023.findings-eacl.56 , pages =

  30. [38]

    Khapra , year=

    Janki Atul Nawale and Mohammed Safi Ur Rahman Khan and Janani D and Mansi Gupta and Danish Pruthi and Mitesh M. Khapra , year=. 2506.23111 , journal=

  31. [39]

    GitHub repository , howpublished =

    Aditya Kallappa and Guo Xiang and Jay Piplodiya and Manoj Guduru and Neel Rachamalla and Palash Kamble and Souvik Rana and Vivek Dahiya and Yong Tong Chua and Ashish Kulkarni and Hareesh Kumar and Chandra Khatri , title =. GitHub repository , howpublished =. 2025 , publisher =

  32. [40]

    and Kumar, Pratyush

    Kumar, Aman and Shrotriya, Himani and Sahu, Prachi and Mishra, Amogh and Dabre, Raj and Puduppully, Ratish and Kunchukuttan, Anoop and Khapra, Mitesh M. and Kumar, Pratyush. I ndic NLG Benchmark: Multilingual Datasets for Diverse NLG Tasks in I ndic Languages. Proceedings of t...

  33. [41]

    The F lores-101 Evaluation Benchmark for Low-Resource and Multilingual Machine Translation

    Goyal, Naman and Gao, Cynthia and Chaudhary, Vishrav and Chen, Peng-Jen and Wenzek, Guillaume and Ju, Da and Krishnan, Sanjana and Ranzato, Marc ' Aurelio and Guzm \'a n, Francisco and Fan, Angela. The F lores-101 Evaluation Benchmark for Low-Resource and Multilingual Machine ...

  34. [42]

    arXiv:2501.15747 , url=

    IndicMMLU-Pro: Benchmarking Indic Large Language Models on Multi-Task Language Understanding , author=. arXiv:2501.15747 , url=. 2025 , eprint=

  35. [43]

    arXiv:2404.18796 , url=

    Replacing Judges with Juries: Evaluating LLM Generations with a Panel of Diverse Models , author=. arXiv:2404.18796 , url=. 2024 , eprint=

  36. [44]

    2024 , journal=

    PRD: Peer Rank and Discussion Improve Large Language Model based Evaluations , author=. 2024 , journal=

  37. [45]

    Bowman and Shi Feng , booktitle=

    Arjun Panickssery and Samuel R. Bowman and Shi Feng , booktitle=. 2024 , url=

  38. [46]

    arXiv:2112.04359 , url=

    Ethical and social risks of harm from Language Models , author=. arXiv:2112.04359 , url=. 2021 , eprint=

  39. [47]

    B n S ent M ix: A Diverse B engali- E nglish Code-Mixed Dataset for Sentiment Analysis

    Alam, Sadia and Ishmam, Md Farhan and Alvee, Navid Hasin and Siddique, Md Shahnewaz and Hossain, Md Azam and Kamal, Abu Raihan Mostofa. B n S ent M ix: A Diverse B engali- E nglish Code-Mixed Dataset for Sentiment Analysis. Proceedings of the First Workshop on Language Models ...

  40. [48]

    Transfer Learning for Code-Mixed Data: Do Pretraining Languages Matter?

    Tatariya, Kushal and Lent, Heather and de Lhoneux, Miryam. Transfer Learning for Code-Mixed Data: Do Pretraining Languages Matter?. Proceedings of the 13th Workshop on Computational Approaches to Subjectivity, Sentiment, & Social Media Analysis. 2023. doi:10.18653/v1/2023.wassa-1.32

  41. [49]

    arXiv:2203.16578 , url=

    Code Switched and Code Mixed Speech Recognition for Indic languages , author=. arXiv:2203.16578 , url=. 2022 , eprint=

  42. [50]

    BharatGPT-3B-Indic , author =

  43. [51]

    Llama-3.2-3B-Instruct , author =

  44. [52]

    Gemma 3 , url=

    Gemma Team , year=. Gemma 3 , url=

  45. [53]

    Nemotron-4-Mini-Hindi-4B-Instruct , author =

  46. [54]

    Airavata-7B , author =

  47. [55]

    Gajendra-v0.1 , author =

  48. [56]

    Indic-Gemma-7B (Navarasa 2.0) , author =

  49. [57]

    Aya-23-8B , author =

  50. [58]

    Llama-3-8B-Instruct , author =

  51. [59]

    Llama-3.1-8B-Instruct , author =

  52. [60]

    GemmaOrca-8.5B , author =

  53. [61]

    GemmaUltra-8.5B , author =

  54. [62]

    Llama-3-Nanda-10B-Chat , author =

  55. [63]

    Krutrim-2-12B-Instruct , author =

  56. [64]

    Qwen-2.5-14B-Hindi , author =

  57. [65]

    Sarvam-M-24B , author =

  58. [66]

    Aya-23-35B , author =

  59. [67]

    Llama-3-70B-Instruct , author =

  60. [68]

    Llama-3.1-70B-Instruct , author =

  61. [69]

    Llama-3.3-70B-Instruct , author =

  62. [70]

    Qwen-2.5-72B , author =

  63. [71]

    Llama-3.1-Nanda-87B-Chat , author =

  64. [72]

    Gemini 2.5 Flash , author =

  65. [73]

    Gemini 2.5 Pro , author =

  66. [74]

    Gemini 3 Flash (Preview) , author =

  67. [75]

    Gemini 3 Pro (Preview) , author =

Pith tools

Reviewed August 12, 2026 · model on record in the stance chip above.