REVIEW 4 major objections 6 minor 100 references
IberBench: LLM Evaluation on Iberian Languages
T0 review · 4 major / 6 minor · reviewed 2026-08-16 · deepseek-v4-flash
Pith's one-line read IberBench's evaluation of 23 LLMs on 101 Iberian-language datasets finds that industry-relevant tasks are hardest, with Basque and Galician trailing.
desk verdict IberBench is a genuinely useful benchmark resource whose headline empirical claims are confounded by its own acknowledged task-language imbalance; the resource deserves publication, the findings deserve caution. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is IberBench's measurement pipeline, built on a standardized evaluation harness extended by the authors. Each dataset is normalized to a common format; classification tasks are scored by the model's likelihood of the correct label using Macro-F1, summarization by ROUGE-1, and the sequence-labeling task by chunk-level F1 on a custom tag-wrapping annotation schema. All models are evaluated zero-shot with one prompt per task, and each task is anchored by a uniform random baseline. The analytical core is the aggregation of these per-dataset scores into averages over task categories, relevance classes, and languages, which is what generates the paper's comparative findings.
What would settle it
Recompute the four headline findings after reweighting the datasets so each language and each task category contributes equal weight; if the industry-versus-fundamental gap or the Basque/Galician difficulty shrinks to noise, those findings are artifacts of the benchmark's imbalanced coverage rather than stable model properties.
Extended reading notes
Core claim
The paper's central claim is that a broad, community-extensible benchmark for Iberian languages reveals a consistent capability profile across 23 open LLMs. IberBench assembles 101 datasets, most from evaluation campaigns and the rest from recent LLM benchmarks, spanning 22 task categories and covering Spanish, Portuguese, Catalan, Basque, Galician, English, and five Spanish varieties. Evaluated zero-shot, the models score higher on fundamental tasks such as reading comprehension, question answering, commonsense reasoning, and textual entailment than on industry-relevant tasks such as sentiment, toxicity, stance, author profiling, summarization, and intent classification. The four headline findings are: industry-relevant tasks trail fundamental ones; Galician and Basque are the hardest languages; lexical borrowing chunking, Basque intent classification, and machine-generated-text detection sit near random; and on sentiment, humor, and fake-news detection LLMs beat random but fall below the best published shared-task systems. The paper frames the gap as evidence that existing fundamental-only leaderboards overstate the practical usefulness of LLMs for industrial NLP in these languages.
Load-bearing premise
The headline comparisons assume that averaging scores over 101 datasets is a fair basis for ranking tasks and languages even though Spanish supplies about 60% of the samples and some task categories exist in only one language.
Editorial extensions
If this is right
- Leaderboards that test only fundamental skills will keep overstating how ready models are for industry uses such as moderation, profiling, and summarization in Iberian languages.
- Galician and Basque need deliberate resource building: generic multilingual scaling alone leaves most models at or near random in these languages.
- Tasks near random, including lexical borrowing detection, Basque intent classification, and machine-generated-text detection, define the current reliability frontier and should be the focus of task-specific work.
- Zero-shot evaluation understates models relative to fine-tuned systems; the gap to shared-task results is the available headroom for prompting or adaptation.
- Because model rankings are nearly identical across languages, a model selected on Spanish data will likely rank similarly in Catalan and Portuguese, simplifying deployment choices.
Reading between the lines
- Beyond the paper: the imbalanced language-task matrix means aggregate findings should be re-tested on a matched subset, for instance only tasks that exist in both Spanish and Basque, before guiding resource decisions.
- Beyond the paper: a few-shot variant of the same benchmark would directly test whether the industry-versus-fundamental gap shrinks when models are given examples, which the zero-shot design likely exaggerates.
- Beyond the paper: the near-random machine-text-detection results offer a built-in contamination check; a sudden jump in future submissions could signal that test data leaked into training.
- Beyond the paper: the Portuguese-tuned model's strong Galician transfer suggests deliberately pairing under-resourced languages with close relatives is a cheaper path than building new data from scratch for Galician.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. IberBench is a multi-language benchmark for evaluating LLMs on Iberian languages (Spanish, Portuguese, Catalan, Basque, Galician, English, and several Spanish varieties), integrating 101 datasets from shared tasks (IberLEF, IberEval, TASS, PAN) and existing benchmarks, organized into 22 task categories split into 'fundamental' and 'industry-relevant' groups. The paper describes a complete evaluation infrastructure: dataset normalization, private hosting, a leaderboard UI, an organization committee, and a modified lm-evaluation-harness that supports sequence labeling, incremental evaluation, and on-premise execution. It reports zero-shot evaluations of 23 LLMs (0.1B–14B parameters) using Macro-F1, ROUGE-1, and seqeval F1, and claims (i) LLMs underperform on industry-relevant tasks relative to fundamental ones, (ii) Galician and Basque are harder than other Iberian languages, (iii) several tasks (lexical borrowing, intent, MGT detection) are close to random, and (iv) in other tasks LLMs beat random but lag shared-task systems. The empirical analysis relies heavily on aggregate means over a highly imbalanced (task, language) matrix, which is documented in Table 4 and acknowledged in Section 3.3.
Significance. The paper delivers a genuinely useful resource: a large, curated, standardized collection of Iberian-language evaluation datasets, many previously scattered or difficult to access, with a reproducible evaluation pipeline and public leaderboard infrastructure. The release of prompts, YAML configs, normalization code, and a caching mechanism is a concrete strength, as is the custom annotation scheme that enables sequence-labeling evaluation with LLMs. The empirical insights, if robust, would be valuable for practitioners choosing models for Iberian languages. However, the headline aggregate conclusions (especially the language-difficulty ranking and the fundamental-vs-industry gap) are not yet supported by the analysis as presented, because the aggregation is confounded by the sparse and imbalanced task-language matrix. The benchmark contribution itself is solid; the analysis needs additional controls or, failing that, appropriately hedged claims.
major comments (4)
- [Section 4.2.3, Figures 7–8, Table 4] The claim that "Galician and Basque present greater challenges than other languages" rests on unadjusted averages over a highly imbalanced task-language matrix. Table 4 shows that Basque's 42.6k samples are concentrated in difficult fundamental tasks (EusTrivia, EusProficiency, QNLI, BHTC, FMTODeu) and that Galician's 12.8k samples are almost entirely fundamental tasks from Proxecto Nós plus MGT detection, whereas Spanish's 165.2k samples span many more categories. Since the averages in Figures 7 and 8 are not stratified by task category and are not restricted to datasets that exist in multiple languages, the observed language gap could be driven by which datasets happen to be available in each language rather than by intrinsic language difficulty. The manuscript itself acknowledges this risk in Section 3.3 ("LLMs might appear weaker or stronger in cross-lingual and cross-task comparisons depending on how well they perform on overrepresented or underrepresented combinations"). The authors should add a controlled comparison—e.g., per-category language means, equal-weighting of (task, language) cells, or a matched-dataset analysis—and temper the claim if the gap does not survive that control.
- [Section 4.2.1, Figure 6] The headline finding that LLMs perform worse on industry-relevant than fundamental tasks is subject to the same aggregation issue. Industry datasets are predominantly Spanish (139.6k of 247.1k industry samples) and include notoriously hard categories (MGT, intent, stance, author profiling), while fundamental datasets include easier categories (Commonsense Reasoning, Question Answering) and are more dispersed across languages. A direct comparison of the two relevance groups without task or language controls conflates task difficulty with the relevance label. The authors should either match fundamental/industry pairs within language and task family, or reframe the conclusion as a claim about these specific collections, explicitly citing the per-category medians in Figure 5 rather than the aggregate in Figure 6.
- [Section 4.2.2, paragraph "Shared task participants still lead..."] The comparison with shared-task results is not apples-to-apples: LLMs are evaluated zero-shot on a single prompt, while the "best published models" are fine-tuned on task-specific training data, and the averaging protocol for "best published results" is unspecified (which submissions, which metrics, how many datasets per category). The reported numbers, e.g., 72.67% vs. 84.26% for humor detection, are presented without adjustment for this asymmetry. While the text mentions the zero-shot caveat, the abstract and conclusion repeat the gap without it. The authors should provide a per-dataset comparison using the same metric and test split, and state the fine-tuning/zero-shot asymmetry explicitly where the comparison is summarized.
- [Section 4.2.2, Figure 5 (MGT Detection and Attribution)] Figure 5 shows that the random baseline outperforms every LLM on MGT detection and attribution. The paper reports this without analysis, yet it contradicts the summary claim (iii) that these tasks are "close to random"—they are actually worse than random, which often signals a methodological artifact such as a systematic label bias in likelihood scoring or a prompt-induced majority-class response. The authors should investigate the cause (e.g., report the distribution of predicted labels, check the tokenization of label strings, verify that Macro-F1 is computed identically for the random baseline and the models) and either explain the phenomenon or soften the claim in the abstract and conclusion.
minor comments (6)
- [Section 3.4, Equation (1)] The product notation in Equation (1) appears corrupted in the current rendering ("|y|Y j=0"); it should read ∏_{j=1}^{|y|} pLLM(y_j | s_i, y_<j). Please also define y_<j explicitly as the token prefix before position j.
- [Section 3.3, reference [15]] The name of the evaluation series is written inconsistently: "IberEVAL" in the text appears as "IberEV AL" and the reference [15] uses "IberEval". Please standardize the spelling throughout the manuscript.
- [Table 5] For gemma-2-2b-it, the pre-training language column lists only English, although the model is reported to have broad multilingual coverage. Please clarify the selection criterion (e.g., whether the table lists primary training languages or the languages that meet some coverage threshold) to avoid misleading readers.
- [Section 4.2.3, paragraph on ranking consistency] The statement "in all cases, we find τ≈ρ >0.99 with p<1×10−11" is surprisingly strong given only 23 models and the presence of tied ranks in some language plots. Please report the exact statistics for each pair, state how ties are handled in Kendall's τ, and list the number of models included per language (since not all models have results for all languages).
- [Appendix B, Figures B.10 and B.11] The heatmaps are difficult to read because the numeric annotations are small and the color contrast between adjacent ranks is weak. Consider using a discrete color scale with larger font, or a table with a small multiples layout, so that the rankings are legible in print.
- [Section 3.3, dataset preparation paragraph] The mojibake example ("café" recovered from "café") contains a visible OCR/encoding artifact in the printed text ("caf ˜A©"); please check the encoding of this example in the final PDF.
Circularity Check
No significant circularity: IberBench reports external benchmark measurements, and its aggregation caveats are disclosed rather than disguised.
full rationale
IberBench is a benchmark-curation and empirical-evaluation paper rather than a derivation chain, so there is no fitted parameter or constructed prediction that could reduce to its own inputs. Model scores are produced by the explicit likelihood and decoding procedures in Eqs. (1)-(2) and then compared with the reference labels of 101 existing datasets; the findings (i)-(iv) are descriptive aggregates of those measured scores. The fundamental-versus-industry relevance split is defined a priori in Section 3.3 by economic significance and task provenance, not by observed model performance, so the industry-versus-fundamental gap is an empirical comparison over a pre-registered grouping rather than a definitional artifact. The claim that Galician and Basque are harder is an average over the available (task, language) cells, and the paper itself flags the fragility of such averages: Section 3.3 states that the task-language matrix is sparse and that LLMs 'might appear weaker or stronger in cross-lingual and cross-task comparisons depending on how well they perform on overrepresented or underrepresented combinations,' and the Limitations section says that 'the task types and distributions across languages are imbalanced... All of these aspects can introduce bias to the evaluations and analyses.' That is an honest validity caveat about aggregation, not circularity, because the language-level means are not defined in terms of the conclusion that those languages are harder. The only author-overlapping citation is the IberAuTexTification overview used as a data source; the benchmark evaluates LLMs against the competition's held-out labels, so the citation functions as external, falsifiable evidence rather than as a load-bearing self-reference. No self-citation chain, uniqueness import, ansatz-smuggling, or renaming-of-known-result pattern is present. The appropriate finding is therefore no significant circularity.
Assumptions & free parameters
assumptions (3)
- domain assumption Aggregate mean scores across 101 datasets give a fair basis for comparing task categories and languages.
- domain assumption Zero-shot evaluation with a single prompt per task reflects realistic model performance.
- domain assumption Holding normalized datasets in private repositories prevents test contamination.
Cite this review
Pith. "Pith review of IberBench: LLM Evaluation on Iberian Languages." pith.science (2026). https://pith.science/paper/X5U7JSKF
@misc{pith2026250416921,
author = {Pith},
title = {Pith review of: IberBench: LLM Evaluation on Iberian Languages},
year = {2026},
howpublished = {\url{https://pith.science/paper/X5U7JSKF}},
note = {Machine review of arXiv:2504.16921}
}
read the original abstract
Large Language Models (LLMs) remain difficult to evaluate comprehensively, particularly for languages other than English, where high-quality data is often limited. Existing benchmarks and leaderboards are predominantly English-centric, with only a few addressing other languages. These benchmarks fall short in several key areas: they overlook the diversity of language varieties, prioritize fundamental Natural Language Processing (NLP) capabilities over tasks of industrial relevance, and are static. With these aspects in mind, we present IberBench, a comprehensive and extensible benchmark designed to assess LLM performance on both fundamental and industry-relevant NLP tasks, in languages spoken across the Iberian Peninsula and Ibero-America. IberBench integrates 101 datasets from evaluation campaigns and recent benchmarks, covering 22 task categories such as sentiment and emotion analysis, toxicity detection, and summarization. The benchmark addresses key limitations in current evaluation practices, such as the lack of linguistic diversity and static evaluation setups by enabling continual updates and community-driven model and dataset submissions moderated by a committee of experts. We evaluate 23 LLMs ranging from 100 million to 14 billion parameters and provide empirical insights into their strengths and limitations. Our findings indicate that (i) LLMs perform worse on industry-relevant tasks than in fundamental ones, (ii) performance is on average lower for Galician and Basque, (iii) some tasks show results close to random, and (iv) in other tasks LLMs perform above random but below shared task systems. IberBench offers open-source implementations for the entire evaluation pipeline, including dataset normalization and hosting, incremental evaluation of LLMs, and a publicly accessible leaderboard.
Figures
Figures from the paper (6 more)
Reference graph
Works this paper leans on
-
[1]
T. Eloundou, S. Manning, P. Mishkin, D. Rock, Gpts are gpts: An early look at the labor market impact potential of large language models (2023). arXiv:2303.10130
arXiv 2023
-
[2]
Yenduri, M
G. Yenduri, M. Ramalingam, G. C. Selvi, Y . Supriya, G. Srivastava, P. K. R. Maddikunta, G. D. Raj, R. H. Jhaveri, B. Prabadevi, W. Wang, A. V . Vasilakos, T. R. Gadekallu, Gpt (generative pre-trained transformer)— a comprehensive review on enabling technologies, potential applications, emerging challenges, and future directions, IEEE Access 12 (2024) 54608–54649
2024
-
[3]
A. Grattafiori, A. Dubey, A. Jauhri, A. Pandey, A. Kadian, The llama 3 herd of models (2024). arXiv:2407.21783
arXiv 2024
-
[4]
A. Yang, B. Yang, B. Zhang, B. Hui, B. Zheng, Qwen2.5 technical report (2025). arXiv:2412.15115
arXiv 2025
-
[5]
D. Guo, D. Yang, H. Zhang, J. Song, R. Zhang, Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning (2025). arXiv:2501.12948
arXiv 2025
- [6]
-
[7]
P. Georgiev, V . I. Lei, R. Burnell, L. Bai, A. Gulati, Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context (2024). arXiv:2403.05530
arXiv 2024
-
[8]
Liang, R
P. Liang, R. Bommasani, T. Lee, D. Tsipras, D. Soylu, et al., Holistic Evaluation of Language Models, Transactions on Machine Learning Research (2023)
2023
Show all 100 references
-
[9]
Fourrier, N
C. Fourrier, N. Habib, A. Lozovskaya, K. Szafer, T. Wolf, Open llm leaderboard v2, https://huggingface.co/spaces/ open-llm-leaderboard/open_llm_leaderboard (2024)
2024
-
[10]
Chiang, L
W.-L. Chiang, L. Zheng, Y . Sheng, A. N. Angelopoulos, T. Li, D. Li, B. Zhu, H. Zhang, M. Jordan, J. E. Gonzalez, I. Stoica, Chatbot arena: An open platform for evaluating llms by human preference, in: Proceedings of the International Conference on Machine Learning, 2024, pp. ...
2024
-
[11]
Grandury, J
M. Grandury, J. Aula-Blasco, C. Fourrier, M. Gonz ´alez, G. Mart ´ınez, G. Santamar ´ıa, A. Vaca, La leaderboard: Leaderboard of spanish varieties and official languages, https://huggingface.co/spaces/la-leaderboard/la-leaderboard (2024)
2024
-
[12]
Baucells, J
I. Baucells, J. Aula-Blasco, I. de Dios-Flores, S. Paniagua Su ´arez, N. Perez, A. Salles, S. Sotelo Docio, J. Falc ˜ao, J. J. Saiz, R. Sepul- veda Torres, J. Barnes, P. Gamallo, A. Gonzalez-Agirre, G. Rigau, M. Villegas, IberoBench: A benchmark for LLM evaluation in Iberian l...
2025
-
[13]
Amig ´o, J
E. Amig ´o, J. C. de Albornoz, A. Fern ´andez, J. Gonzalo, M. Lucas, G. Marco, R. Morante, J. Pedrosa, L. Plaza, E. S ´anchez, A. Villa, Proyecto espacio de observaci´on de inteligencia artificial en espa˜nol (odesia), https://leaderboard.odesia.uned.es/leaderboard/ tablesV2 (2025)
2025
-
[14]
Chiruzzo, S
L. Chiruzzo, S. M. Jim ´enez-Zafra, F. Rangel, Overview of iberlef 2024: Natural language processing challenges for spanish and other iberian languages, in: CEUR Workshop Proceedings, V ol. 3756, CEUR-WS, 2024
2024
-
[15]
Rosso, J
P. Rosso, J. Gonzalo, R. Mart ´ınez, S. Montalvo, J. C. de Albornoz (Eds.), Proceedings of the Third Workshop on Evaluation of Human Language Technologies for Iberian Languages (IberEval 2018) co-located with 34th Conference of the Spanish Society for Natural Language Processi...
2018
-
[16]
E. M. C ´amara, Y . Almeida-Cruz, M. C. D´ıaz-Galiano, S. Est´evez-Velarde, M. ´A. G. Cumbreras, M. G. Vega, Y . Guti´errez, A. Montejo-R´aez, A. Montoyo, R. Mu˜noz, A. Piad-Morffis, J. Villena-Rom´an (Eds.), Proceedings of TASS 2018: Workshop on Semantic Analysis at SEPLN, TA...
2018
-
[17]
Stamatatos, M
E. Stamatatos, M. Potthast, F. M. R. Pardo, P. Rosso, B. Stein, Overview of the PAN /CLEF 2015 evaluation lab, in: J. Mothe, J. Savoy, J. Kamps, K. Pinel-Sauvagnat, G. J. F. Jones, E. SanJuan, L. Cappellato, N. Ferro (Eds.), Experimental IR Meets Multilinguality, Multi- modali...
2015
-
[18]
Etxaniz, O
J. Etxaniz, O. Sainz, N. Miguel, I. Aldabe, G. Rigau, E. Agirre, A. Ormazabal, M. Artetxe, A. Soroa, Latxa: An open language model and evaluation suite for Basque, in: Proceedings of the Annual Meeting of the Association for Computational Linguistics (V olume 1: Long Papers), ...
2024
-
[19]
Biderman, H
S. Biderman, H. Schoelkopf, L. Sutawika, L. Gao, J. Tow, B. Abbasi, A. F. Aji, P. S. Ammanamanchi, S. Black, J. Clive, A. DiPofi, J. Etxaniz, B. Fattori, J. Z. Forde, C. Foster, J. Hsu, M. Jaiswal, W. Y . Lee, H. Li, C. Lovering, N. Muennigho ff, E. Pavlick, J. Phang, A. Skowr...
2024 arXiv
-
[20]
Abdin, J
M. Abdin, J. Aneja, H. Behl, S. Bubeck, R. Eldan, S. Gunasekar, M. Harrison, R. J. Hewett, M. Javaheripi, P. Kau ffmann, J. R. Lee, Y . T. Lee, Y . Li, W. Liu, C. C. T. Mendes, A. Nguyen, E. Price, G. de Rosa, O. Saarikivi, A. Salim, S. Shah, X. Wang, R. Ward, Y . Wu, D. Yu, C...
2024 arXiv
-
[21]
Gonzalez-Agirre, M
A. Gonzalez-Agirre, M. P `amies, J. Llop, I. Baucells, S. D. Dalt, D. Tamayo, J. J. Saiz, F. Espu ˜na, J. Prats, J. Aula-Blasco, M. Mina, I. Pikabea, A. Rubio, A. Shvets, A. Sall´es, I. Lacunza, J. Palomar, J. Falc˜ao, L. Tormo, L. Vasquez-Reina, M. Marimon, O. Pareras, V . Ru...
2025 arXiv
-
[22]
P. H. Martins, P. Fernandes, J. Alves, N. M. Guerreiro, R. Rei, D. M. Alves, J. Pombal, A. Farajian, M. Faysse, M. Klimaszewski, P. Colombo, B. Haddow, J. G. de Souza, A. Birch, A. F. Martins, Eurollm: Multilingual language models for europe, Procedia Computer Science 255 (2025) 53–62
2025
-
[23]
J. Xu, D. Ju, M. Li, Y .-L. Boureau, J. Weston, E. Dinan, Bot-adversarial dialogue for safe conversational agents, in: Proceedings of the Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, 2021, pp. 2950–2968
2021
-
[24]
A. Wei, N. Haghtalab, J. Steinhardt, Jailbroken: How does LLM safety training fail?, in: Proceedings of the Conference on Neural Informa- tion Processing Systems, 2023, pp. 80079–80110
2023
-
[25]
G. Bai, J. Liu, X. Bu, Y . He, J. Liu, Z. Zhou, Z. Lin, W. Su, T. Ge, B. Zheng, W. Ouyang, MT-bench-101: A fine-grained benchmark 22 for evaluating large language models in multi-turn dialogues, in: Proceedings of the Annual Meeting of the Association for Computational Linguis...
2024
-
[26]
Li, W.-L
T. Li, W.-L. Chiang, E. Frick, L. Dunlap, T. Wu, B. Zhu, J. E. Gonzalez, I. Stoica, From crowdsourced data to high-quality benchmarks: Arena-hard and benchbuilder pipeline (2024). arXiv:2406.11939
2024 arXiv
-
[27]
M. Wu, A. F. Aji, Style over substance: Evaluation biases for large language models, in: Proceedings of the International Conference on Computational Linguistics, 2025, pp. 297–312
2025
-
[28]
this is a problem, don‘t you agree?
S. Schoch, D. Yang, Y . Ji, “this is a problem, don‘t you agree?” framing and bias in human evaluation for natural language generation, in: Proceedings of the Workshop on Evaluating NLG Evaluation, 2020, pp. 10–16
2020
-
[29]
Baumann, Universal jailbreak backdoors in large language model alignment, in: Neurips Safe Generative AI Workshop 2024, 2024
T. Baumann, Universal jailbreak backdoors in large language model alignment, in: Neurips Safe Generative AI Workshop 2024, 2024
2024
-
[30]
A. Wang, A. Singh, J. Michael, F. Hill, O. Levy, S. Bowman, GLUE: A multi-task benchmark and analysis platform for natural language understanding, in: Proceedings of the EMNLP Workshop BlackboxNLP: Analyzing and Interpreting Neural Networks for NLP, 2018, pp. 353–355
2018
-
[31]
Hendrycks, C
D. Hendrycks, C. Burns, S. Basart, A. Zou, M. Mazeika, D. Song, J. Steinhardt, Measuring massive multitask language understanding, Proceedings of the International Conference on Learning Representations (2021) 1–30
2021
-
[32]
Bavaresco, R
A. Bavaresco, R. Bernardi, L. Bertolazzi, D. Elliott, R. Fern ´andez, A. Gatt, E. Ghaleb, M. Giulianelli, M. Hanna, A. Koller, A. F. T. Martins, P. Mondorf, V . Neplenbroek, S. Pezzelle, B. Plank, D. Schlangen, A. Suglia, A. K. Surikuchi, E. Takmaz, A. Testoni, Llms instead of...
2024 arXiv
-
[33]
Zheng, W.-L
L. Zheng, W.-L. Chiang, Y . Sheng, S. Zhuang, Z. Wu, Y . Zhuang, Z. Lin, Z. Li, D. Li, E. P. Xing, H. Zhang, J. E. Gonzalez, I. Stoica, Judging llm-as-a-judge with mt-bench and chatbot arena, in: Proceedings of the International Conference on Neural Information Processing Syst...
2023
-
[34]
G. H. Chen, S. Chen, Z. Liu, F. Jiang, B. Wang, Humans or LLMs as the judge? a study on judgement bias, in: Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, 2024, pp. 8301–8327
2024
-
[35]
M. Ali, M. Fromm, K. Thellmann, J. Ebert, A. A. Weber, R. Rutmann, C. Jain, M. L ¨ubbering, D. Steinigen, J. Leveling, K. Klug, J. S. Buschhoff, L. Jurkschat, H. Abdelwahab, B. J. Stein, K.-H. Sylla, P. Denisov, N. Brandizzi, Q. Saleem, A. Bhowmick, L. Helmer, C. John, P. O. S...
2024 arXiv
-
[36]
Singh, N
H. Singh, N. Gupta, S. Bharadwaj, D. Tewari, P. Talukdar, IndicGenBench: A multilingual benchmark to evaluate generation capabilities of LLMs on Indic languages, in: Proceedings of the Annual Meeting of the Association for Computational Linguistics (V olume 1: Long Papers), 20...
2024
-
[37]
Susanto, A
Y . Susanto, A. V . Hulagadri, J. R. Montalan, J. G. Ngui, X. B. Yong, W. Leong, H. Rengarajan, P. Limkonchotiwat, Y . Mai, W. C. Tjhi, Sea-helm: Southeast asian holistic evaluation of language models (2025). arXiv:2502.14301
2025 arXiv
-
[38]
Almazrouei, R
E. Almazrouei, R. Cojocaru, M. Baldo, Q. Malartic, H. Alobeidli, D. Mazzotta, G. Penedo, G. Campesan, M. Farooq, M. Alhammadi, J. Lau- nay, B. Noune, AlGhafa evaluation benchmark for Arabic language models, in: Proceedings of the Arabic Natural Language Processing Conference, ...
2023
-
[39]
Duki ´c, J
D. Duki ´c, J. Snajder, Looking right is sometimes right: Investigating the capabilities of decoder-only LLMs for sequence labeling, in: Findings of the Association for Computational Linguistic, 2024, pp. 14168–14181
2024
-
[40]
S. Wang, X. Sun, X. Li, R. Ouyang, F. Wu, T. Zhang, J. Li, G. Wang, Gpt-ner: Named entity recognition via large language models (2023). arXiv:2304.10428
2023 arXiv
-
[41]
Houlsby, A
N. Houlsby, A. Giurgiu, S. Jastrzebski, B. Morrone, Q. de Laroussilhe, A. Gesmundo, M. Attariyan, S. Gelly, Parameter-e fficient transfer learning for nlp, in: Proceedings of the International Conference on Machine Learning, 2019, pp. 2790–2799
2019
-
[42]
Frantar, S
E. Frantar, S. Ashkboos, T. Hoefler, D. Alistarh, OPTQ: Accurate quantization for generative pre-trained transformers, in: The Eleventh International Conference on Learning Representations, 2023
2023
-
[43]
Rangel, P
F. Rangel, P. Rosso, M. Potthast, B. Stein, Overview of the 5th author profiling task at pan 2017: Gender and language variety identification in twitter, Working notes papers of the CLEF 48 (2017)
2017
-
[44]
Bandarkar, D
L. Bandarkar, D. Liang, B. Muller, M. Artetxe, S. N. Shukla, D. Husa, N. Goyal, A. Krishnan, L. Zettlemoyer, M. Khabsa, The belebele benchmark: a parallel reading comprehension dataset in 122 language variants, in: Proceedings of the Annual Meeting of the Association for Compu...
2024
-
[45]
Schae ffer, Pretraining on the test set is all you need, arXiv preprint arXiv:2309.08632 (2023)
R. Schae ffer, Pretraining on the test set is all you need, arXiv preprint arXiv:2309.08632 (2023)
2023 arXiv
-
[46]
Lhoest, A
Q. Lhoest, A. Villanova del Moral, Y . Jernite, A. Thakur, P. von Platen, Datasets: A community library for natural language processing, in: H. Adel, S. Shi (Eds.), Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing: System Demonstrations, A...
2021
-
[47]
Y . Hu, Q. Chen, J. Du, X. Peng, V . K. Keloth, X. Zuo, Y . Zhou, Z. Li, X. Jiang, Z. Lu, K. Roberts, H. Xu, Improving large language models for clinical named entity recognition via prompt engineering, Journal of the American Medical Informatics 31 (9) (2024) 1812–1820
2024
-
[48]
Zubiaga, I
A. Zubiaga, I. S. Vicente, P. Gamallo, J. R. Pichel, I. Alegria, N. Aranberri, A. Ezeiza, V . Fresno, Tweetlid: a benchmark for tweet language identification, Language Resources and Evaluation 50 (2016) 729–766
2016
-
[49]
Garcia-Vega, M
M. Garcia-Vega, M. D ´ıaz-Galiano, M. Garc´ıa-Cumbreras, F. M. P. Del Arco, A. Montejo-Ra´ez, S. Jim´enez-Zafra, E. M. C´amara, C. Aguilar, M. Cabezudo, L. Chiruzzo, et al., Overview of tass 2020: Introducing emotion detection, in: Proceedings of the Iberian Languages Evalua- ...
2020
-
[50]
F. M. Rangel Pardo, F. Celli, P. Rosso, M. Potthast, B. Stein, W. Daelemans, Overview of the author profiling task at pan 2015, in: Working notes papers of the CLEF, 2015, pp. 1–8
2015
-
[51]
Taul ´e, F
M. Taul ´e, F. M. R. Pardo, M. A. Mart´ı, P. Rosso, Overview of the task on multimodal stance detection in tweets on catalan# 1oct referendum, in: Proceedings of the Workshop on Evaluation of Human Language Technologies for Iberian Languages, 2018, pp. 149–166
2018
-
[52]
Chiruzzo, S
L. Chiruzzo, S. Castro, M. Etcheverry, D. Garat, J. J. Prada, A. Ros ´a, Overview of haha at iberlef 2019: Humor analysis based on human annotation, in: Proceedings of the Iberian Languages Evaluation Forum, 2019, pp. 132–144. 23
2019
-
[53]
Ortega-Bueno, F
R. Ortega-Bueno, F. Rangel, D. Hern ´andez Farıas, P. Rosso, M. Montes-y G ´omez, J. E. Medina Pagola, Overview of the task on irony detection in spanish variants, in: Proceedings of the Iberian languages evaluation forum, co-located with conference of the Spanish Society for ...
2019
-
[54]
M. E. Arag ´on, M. ´A. ´Alvarez-Carmona, M. Montes-y G´omez, H. J. Escalante, L. V . Pineda, D. Moctezuma, Overview of mex-a3t at iberlef 2019: Authorship and aggressiveness analysis in mexican spanish tweets., in: Proceedings of the Iberian Languages Evaluation Forum, 2019, p...
2019
-
[55]
Taul ´e, A
M. Taul ´e, A. Ariza, M. Nofre, E. Amig ´o, P. Rosso, Overview of detoxis at iberlef 2021: Detection of toxicity in comments in spanish, Procesamiento del Lenguaje Natural 67 (2021) 209–221
2021
-
[56]
F. M. Plaza-del Arco, S. M. Jim´enez-Zafra, A. Montejo-R´aez, M. D. Molina-Gonz´alez, L. A. Ure˜na-L´opez, M. T. Mart´ın-Valdivia, Overview of the emoevales task on emotion detection for spanish at iberlef 2021, Procesamiento del Lenguaje Natural 67 (2021) 155–161
2021
-
[57]
F. J. Rodr ´ıguez-S´anchez, J. Carrillo-de Albornoz, L. Plaza, J. Gonzalo, P. Rosso, M. Comet, T. Donoso, Overview of exist 2021: sexism identification in social networks, Procesamiento del Lenguaje Natural 67 (2021) 195–207
2021
-
[58]
G ´omez-Adorno, J
H. G ´omez-Adorno, J. P. Posadas-Dur ´an, G. Bel Enguix, C. Porto Capetillo, Overview of fakedes at iberlef 2021: Fake news detection in spanish shared task, Procesamiento del Lenguaje Natural 67 (2021) 223–231
2021
-
[59]
Chiruzzo, S
L. Chiruzzo, S. Castro, S. G ´ongora, A. Ros ´a, J. Meaney, R. Mihalcea, Overview of haha at iberlef 2021: Detecting, rating and analyzing humor in spanish, Procesamiento del Lenguaje Natural 67 (2021) 257–268
2021
-
[60]
Jarqu ´ın-V´asquez, L
H. Jarqu ´ın-V´asquez, L. Villase˜nor Pineda, F. M. Plaza-del Arco, M. Casavantes, H. J. Escalante, M. T. Mart ´ın Valdivia, A. Montejo R´aez, M. Montes y G´omez, Overview of meoffendes at iberlef 2021: Offensive language detection in spanish variants, Procesamiento del Lengua...
2021
-
[61]
M. ´A. ´Alvarez Carmona, R. Aranda, S. Arce-Cardenas, D. Fajardo-Delgado, R. Guerrero-Rodr ´ıguez, A. P. L ´opez-Monroy, J. Mart´ınez- Miranda, H. P ´erez-Espinosa, A. Y . Rodr´ıguez-Gonz´alez, Overview of rest-mex at iberlef 2021: Recommendation system for text mexican touris...
2021
-
[62]
Agerri Gasc ´on, R
R. Agerri Gasc ´on, R. Centeno S´anchez, M. Espinosa, J. Fern´andez de Landa, ´A. Rodrigo Yuste, Vaxxstance@iberlef 2021: Overview of the task on going beyond text in cross-lingual stance detection, Procesamiento del Lenguaje Natural 67 (2021) 173–181
2021
-
[63]
Bel-Enguix, G
G. Bel-Enguix, G. Sierra, H. G ´omez-Adorno, J.-M. Torres-Moreno, J.-G. Ortiz-Barajas, J. V ´asquez, Overview of par-mex at iberlef 2022: Paraphrase detection in spanish shared task, Procesamiento del Lenguaje Natural 69 (2022) 255–263
2022
-
[64]
M. ´A. ´Alvarez Carmona, ´A. D´ıaz-Pacheco, R. Aranda, A. Y . Rodr´ıguez-Gonz´alez, D. Fajardo-Delgado, R. Guerrero-Rodr´ıguez, L. Bustio- Mart´ınez, Overview of rest-mex at iberlef 2022: Recommendation system, sentiment analysis and covid semaphore prediction for mexican tour...
2022
-
[65]
A. M. M ´armol-Romero, A. Moreno-Mu ˜noz, F. M. Plaza-del Arco, M. D. Molina-Gonz ´alez, M. T. Mart ´ın-Valdivia, L. A. Ure ˜na-L´opez, A. Montejo-R´aez, Overview of mentalriskes at iberlef 2023: Early detection of mental disorders risk in spanish, Procesamiento del Lenguaje N...
2023
-
[66]
Labadie Tamayo, B
R. Labadie Tamayo, B. Chulvi, P. Rosso, Everybody hurts, sometimes. overview of hurtful humour at iberlef 2023: Detection of humour spreading prejudice in twitter, Procesamiento del Lenguaje Natural 71 (2023) 383–395
2023
-
[67]
W. S. S.-N. y Pol Pastells y Simona Frenda y Alejandro Ariza-Casabona y Mireia Farr ´us y Paolo Rosso y Mariona Taul ´e, Overview of detests-dis at iberlef 2024: Detection and classification of racial stereotypes in spanish - learning with disagreement, Procesamiento del Lengu...
2024
-
[68]
A. M. S. y Jos ´e ´Angel Gonz ´alez y Francisco Rangel y Paolo Rosso y Marc Franco-Salvador, Overview of iberautextification at iberlef 2024: Detection and attribution of machine-generated text on languages of the iberian peninsula, Procesamiento del Lenguaje Natural 73 (0) (2...
2024
-
[69]
Y . Yang, Y . Zhang, C. Tar, J. Baldridge, PAWS-X: A cross-lingual adversarial dataset for paraphrase identification, in: Proceedings of the Conference on Empirical Methods in Natural Language Processing and the International Joint Conference on Natural Language Processing, 20...
2019
-
[70]
A. I. Vladu, I. de Dios-Flores, C. Magari ˜nos, J. E. Ortega, J. R. Pichel, M. Garcia, P. Gamallo, E. F. Rei, A. Bugar ´ın, M. G. Gonz ´alez, et al., Proxecto n´os: Artificial intelligence at the service of the galician language, in: Proceedings of the Annual Conference of the...
2022
-
[71]
Gonzalez-Agirre, M
A. Gonzalez-Agirre, M. Marimon, C. Rodriguez-Penagos, J. Aula-Blasco, I. Baucells, C. Armentano-Oller, J. Palomar-Giner, B. Kulebi, M. Villegas, Building a data infrastructure for a mid-resource language: The case of Catalan, in: Proceedings of the Joint International Conferen...
2024
-
[72]
Hasan, A
T. Hasan, A. Bhattacharjee, M. S. Islam, K. Mubasshir, Y .-F. Li, Y .-B. Kang, M. S. Rahman, R. Shahriyar, XL-sum: Large-scale multilingual abstractive summarization for 44 languages, in: Findings of the Association for Computational Linguistics, 2021, pp. 4693–4703
2021
-
[73]
URL https://projecteaina.cat/
Barcelona Supercomputing Center, Projecte aina: Recursos ling ¨u´ıstics i tecnol`ogics per al catal`a a l’era digital (2021). URL https://projecteaina.cat/
2021
-
[74]
Urbizu, I
G. Urbizu, I. San Vicente, X. Saralegi, R. Agerri, A. Soroa, Basqueglue: A natural language understanding benchmark for basque, in: Proceedings of the Language Resources and Evaluation Conference, Marseille, France, 2022, pp. 1603–1612
2022
-
[75]
A. V . Serrano, I. L. Montalb´an, D. V . Vel´azquez, M. Moreno, Clindiagnoses (2024). URL https://huggingface.co/datasets/LenguajeNaturalAI/ClinDiagnosES
2024
-
[76]
R ¨ottger, H
P. R ¨ottger, H. Seelawi, D. Nozza, Z. Talat, B. Vidgen, Multilingual HateCheck: Functional tests for multilingual hate speech detection models, in: Proceedings of the Workshop on Online Abuse and Harms, 2022, pp. 154–169
2022
-
[77]
´Alvarez Mellado, L
E. ´Alvarez Mellado, L. Espinosa Anke, J. G. Arroyo, C. Lignos, J. Porta Zamorano, Overview of adobo 2021: Automatic detection of unassimilated borrowings in the spanish press, Procesamiento del Lenguaje Natural 67 (2021) 277–285
2021
-
[78]
Armengol-Estap ´e, C
J. Armengol-Estap ´e, C. P. Carrino, C. Rodriguez-Penagos, O. de Gibert Bonet, C. Armentano-Oller, A. Gonzalez-Agirre, M. Melero, M. Villegas, Are multilingual models the best choice for moderately under-resourced languages? A comprehensive assessment for Catalan, in: Findings...
2021
-
[79]
https://proyectoilenia.es/
Secretar ´ıa de Estado de Digitalizaci´on e Inteligencia Artificial, Proyecto del impulso de las lenguas en la inteligencia artificial (ilenia) (2022). 24 URL "https://proyectoilenia.es/"
2022
-
[80]
N. Bel, M. Punsola, V . Ruiz-Fern ´andez, EsCoLA: Spanish corpus of linguistic acceptability, in: Proceedings of the Joint International Conference on Computational Linguistics, Language Resources and Evaluation, 2024, pp. 6268–6277
2024
-
[81]
N. Bel, M. Punsola, V . Ruiz-Fern ´andez, Catcola, catalan corpus of linguistic acceptability, Procesamiento del Lenguaje Natural 73 (2024) 177–190
2024
-
[82]
Mayor-Rocher, N
M. Mayor-Rocher, N. Melero, E. Merino-G ´omez, M. Gonz ´alez, R. Ferrando, J. Conde, P. Reviriego, Spanish Language Benchmark for Artificial Intelligence Models (TELEIA) (2024). URL https://doi.org/10.5281/zenodo.12571763
2024 doi
-
[83]
Conneau, R
A. Conneau, R. Rinott, G. Lample, A. Williams, S. R. Bowman, H. Schwenk, V . Stoyanov, Xnli: Evaluating cross-lingual sentence repre- sentations, in: Proceedings of the Conference on Empirical Methods in Natural Language Processing, 2018, pp. 2475–2485
2018
-
[84]
Lin, ROUGE: A package for automatic evaluation of summaries, in: Text Summarization Branches Out, Association for Computational Linguistics, 2004, pp
C.-Y . Lin, ROUGE: A package for automatic evaluation of summaries, in: Text Summarization Branches Out, Association for Computational Linguistics, 2004, pp. 74–81
2004
-
[85]
T. He, J. Zhang, T. Wang, S. Kumar, K. Cho, J. Glass, Y . Tsvetkov, On the blind spots of model-based evaluation metrics for text generation, in: Proceedings of the Annual Meeting of the Association for Computational Linguistics (V olume 1: Long Papers), 2023, pp. 12067–12097
2023
-
[86]
Nakayama, seqeval: A python framework for sequence labeling evaluation (2018)
H. Nakayama, seqeval: A python framework for sequence labeling evaluation (2018). URL https://github.com/chakki-works/seqeval
2018
-
[87]
Rafailov, A
R. Rafailov, A. Sharma, E. Mitchell, C. D. Manning, S. Ermon, C. Finn, Direct preference optimization: Your language model is secretly a reward model, in: Advances in Neural Information Processing Systems, V ol. 36, Curran Associates, Inc., 2023, pp. 53728–53741
2023
-
[88]
Schulman, F
J. Schulman, F. Wolski, P. Dhariwal, A. Radford, O. Klimov, Proximal policy optimization algorithms (2017). arXiv:1707.06347
2017 arXiv
-
[89]
Abouelenin, A
A. Abouelenin, A. Ashfaq, A. Atkinson, H. Awadalla, N. Bach, Phi-4-mini technical report: Compact yet powerful multimodal language models via mixture-of-loras (et al., 2025). arXiv:2503.01743
2025 arXiv
-
[90]
A. Q. Jiang, A. Sablayrolles, A. Mensch, C. Bamford, D. S. Chaplot, D. de las Casas, F. Bressand, G. Lengyel, G. Lample, L. Saulnier, L. R. Lavaud, M.-A. Lachaux, P. Stock, T. L. Scao, T. Lavril, T. Wang, T. Lacroix, W. E. Sayed, Mistral 7b (2023).arXiv:2310.06825
2023 arXiv
-
[91]
Radford, J
A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, I. Sutskever, Language models are unsupervised multitask learners, technical report (2019)
2019
-
[92]
Riviere, S
M. Riviere, S. Pathak, P. G. Sessa, C. Hardin, S. Bhupatiraju, Gemma 2: Improving open language models at a practical size (2024). arXiv:2408.00118
2024 arXiv
-
[93]
G. S. G ´omez, G. G. Subies, P. G. Ruiz, M. G. Valero, N. Fuertes, H. M. Zamorano, C. M. Sanz, L. R. Plaza, N. A. Garc ´ıa, D. B. S ´anchez, K. Sushkova, M. G. Nieto, ´Alvaro Barbero Jim´enez, Rigochat 2: an adapted language model to spanish using a bounded dataset and reduced...
2025 arXiv
-
[94]
Petrea, Catallama, https://huggingface.co/catallama/CataLlama-v0.2-Instruct-SFT (2024)
L. Petrea, Catallama, https://huggingface.co/catallama/CataLlama-v0.2-Instruct-SFT (2024)
2024
-
[95]
Pires, H
R. Pires, H. Abonizio, T. S. Almeida, R. Nogueira, Sabi ´a: Portuguese large language models, in: Intelligent Systems, 2023, pp. 226–240
2023
-
[96]
Language, I. S. Group, Aitana, https://huggingface.co/gplsi/Aitana-6.3B (2024)
2024
-
[97]
Q. Dong, L. Li, D. Dai, C. Zheng, J. Ma, R. Li, H. Xia, J. Xu, Z. Wu, B. Chang, X. Sun, L. Li, Z. Sui, A survey on in-context learning, in: Proceedings of the Conference on Empirical Methods in Natural Language Processing, 2024, pp. 1107–1128
2024
-
[98]
Ho ffmann, S
J. Ho ffmann, S. Borgeaud, A. Mensch, E. Buchatskaya, T. Cai, E. Rutherford, D. de Las Casas, L. A. Hendricks, J. Welbl, A. Clark, et al., Training compute-optimal large language models, in: Proceedings of the International Conference on Neural Information Processing Systems, ...
2022
-
[99]
O’Rourke, The galician language in the twenty-first century, A companion to Galician culture 344 (2014) 73
B. O’Rourke, The galician language in the twenty-first century, A companion to Galician culture 344 (2014) 73
2014
-
[100]
Tovar, H
A. Tovar, H. P. Houghton, The Basque Language, University of Pennsylvania Press, 1957. 25 Appendix A. Datasets and Sources Table A.6 lists the main URLs of the workshops, shared tasks and other sources we include in IberBench as industry-relevant tasks. Table A.7 shows the URL...
1957
Reviewed August 16, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.