Pith. sign in

REVIEW 4 major objections 6 minor 100 references

IberBench: LLM Evaluation on Iberian Languages

T0 review · 4 major / 6 minor · reviewed 2026-08-16 · deepseek-v4-flash

Pith's one-line read IberBench's evaluation of 23 LLMs on 101 Iberian-language datasets finds that industry-relevant tasks are hardest, with Basque and Galician trailing.

desk verdict IberBench is a genuinely useful benchmark resource whose headline empirical claims are confounded by its own acknowledged task-language imbalance; the resource deserves publication, the findings deserve caution. read the letter →

arxiv 2504.16921 v1 pith:X5U7JSKF submitted 2025-04-23 cs.CL

classification cs.CL
keywords LLMevaluationIberianlanguagesmultilingualbenchmarkindustry-relevantNLPBasqueGalicianmachine-generatedtextdetectionzero-shot
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Large language models are usually tested on English-centric benchmarks that emphasize reading comprehension and general knowledge, so their practical value for languages like Spanish, Catalan, Basque, and Galician is largely unknown. This paper introduces IberBench, a benchmark that assembles 101 existing datasets across 22 task categories covering six languages and five Spanish varieties, and evaluates 23 open models from 100 million to 14 billion parameters. Its central claim is that current LLMs are measurably weaker at industry-relevant tasks such as sentiment analysis, intent classification, toxicity detection, and machine-generated-text detection than at fundamental language tasks, and that the gap is especially wide for Basque and Galician. The benchmark also reports that on several tasks, including Basque intent classification and machine-text detection, top models barely beat random guessing, and that zero-shot LLMs trail the best fine-tuned shared-task systems on tasks like humor and fake-news detection. If these findings hold, they imply that LLM leaderboards built on fundamental tasks alone overstate readiness for real-world deployment in Iberian languages.

What carries the argument

The load-bearing object is IberBench's measurement pipeline, built on a standardized evaluation harness extended by the authors. Each dataset is normalized to a common format; classification tasks are scored by the model's likelihood of the correct label using Macro-F1, summarization by ROUGE-1, and the sequence-labeling task by chunk-level F1 on a custom tag-wrapping annotation schema. All models are evaluated zero-shot with one prompt per task, and each task is anchored by a uniform random baseline. The analytical core is the aggregation of these per-dataset scores into averages over task categories, relevance classes, and languages, which is what generates the paper's comparative findings.

What would settle it

Recompute the four headline findings after reweighting the datasets so each language and each task category contributes equal weight; if the industry-versus-fundamental gap or the Basque/Galician difficulty shrinks to noise, those findings are artifacts of the benchmark's imbalanced coverage rather than stable model properties.

Watch

Extended reading notes

Core claim

The paper's central claim is that a broad, community-extensible benchmark for Iberian languages reveals a consistent capability profile across 23 open LLMs. IberBench assembles 101 datasets, most from evaluation campaigns and the rest from recent LLM benchmarks, spanning 22 task categories and covering Spanish, Portuguese, Catalan, Basque, Galician, English, and five Spanish varieties. Evaluated zero-shot, the models score higher on fundamental tasks such as reading comprehension, question answering, commonsense reasoning, and textual entailment than on industry-relevant tasks such as sentiment, toxicity, stance, author profiling, summarization, and intent classification. The four headline findings are: industry-relevant tasks trail fundamental ones; Galician and Basque are the hardest languages; lexical borrowing chunking, Basque intent classification, and machine-generated-text detection sit near random; and on sentiment, humor, and fake-news detection LLMs beat random but fall below the best published shared-task systems. The paper frames the gap as evidence that existing fundamental-only leaderboards overstate the practical usefulness of LLMs for industrial NLP in these languages.

Load-bearing premise

The headline comparisons assume that averaging scores over 101 datasets is a fair basis for ranking tasks and languages even though Spanish supplies about 60% of the samples and some task categories exist in only one language.

Editorial extensions

If this is right

  • Leaderboards that test only fundamental skills will keep overstating how ready models are for industry uses such as moderation, profiling, and summarization in Iberian languages.
  • Galician and Basque need deliberate resource building: generic multilingual scaling alone leaves most models at or near random in these languages.
  • Tasks near random, including lexical borrowing detection, Basque intent classification, and machine-generated-text detection, define the current reliability frontier and should be the focus of task-specific work.
  • Zero-shot evaluation understates models relative to fine-tuned systems; the gap to shared-task results is the available headroom for prompting or adaptation.
  • Because model rankings are nearly identical across languages, a model selected on Spanish data will likely rank similarly in Catalan and Portuguese, simplifying deployment choices.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Beyond the paper: the imbalanced language-task matrix means aggregate findings should be re-tested on a matched subset, for instance only tasks that exist in both Spanish and Basque, before guiding resource decisions.
  • Beyond the paper: a few-shot variant of the same benchmark would directly test whether the industry-versus-fundamental gap shrinks when models are given examples, which the zero-shot design likely exaggerates.
  • Beyond the paper: the near-random machine-text-detection results offer a built-in contamination check; a sudden jump in future submissions could signal that test data leaked into training.
  • Beyond the paper: the Portuguese-tuned model's strong Galician transfer suggests deliberately pairing under-resourced languages with close relatives is a cheaper path than building new data from scratch for Galician.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 6 minor

Summary. IberBench is a multi-language benchmark for evaluating LLMs on Iberian languages (Spanish, Portuguese, Catalan, Basque, Galician, English, and several Spanish varieties), integrating 101 datasets from shared tasks (IberLEF, IberEval, TASS, PAN) and existing benchmarks, organized into 22 task categories split into 'fundamental' and 'industry-relevant' groups. The paper describes a complete evaluation infrastructure: dataset normalization, private hosting, a leaderboard UI, an organization committee, and a modified lm-evaluation-harness that supports sequence labeling, incremental evaluation, and on-premise execution. It reports zero-shot evaluations of 23 LLMs (0.1B–14B parameters) using Macro-F1, ROUGE-1, and seqeval F1, and claims (i) LLMs underperform on industry-relevant tasks relative to fundamental ones, (ii) Galician and Basque are harder than other Iberian languages, (iii) several tasks (lexical borrowing, intent, MGT detection) are close to random, and (iv) in other tasks LLMs beat random but lag shared-task systems. The empirical analysis relies heavily on aggregate means over a highly imbalanced (task, language) matrix, which is documented in Table 4 and acknowledged in Section 3.3.

Significance. The paper delivers a genuinely useful resource: a large, curated, standardized collection of Iberian-language evaluation datasets, many previously scattered or difficult to access, with a reproducible evaluation pipeline and public leaderboard infrastructure. The release of prompts, YAML configs, normalization code, and a caching mechanism is a concrete strength, as is the custom annotation scheme that enables sequence-labeling evaluation with LLMs. The empirical insights, if robust, would be valuable for practitioners choosing models for Iberian languages. However, the headline aggregate conclusions (especially the language-difficulty ranking and the fundamental-vs-industry gap) are not yet supported by the analysis as presented, because the aggregation is confounded by the sparse and imbalanced task-language matrix. The benchmark contribution itself is solid; the analysis needs additional controls or, failing that, appropriately hedged claims.

major comments (4)
  1. [Section 4.2.3, Figures 7–8, Table 4] The claim that "Galician and Basque present greater challenges than other languages" rests on unadjusted averages over a highly imbalanced task-language matrix. Table 4 shows that Basque's 42.6k samples are concentrated in difficult fundamental tasks (EusTrivia, EusProficiency, QNLI, BHTC, FMTODeu) and that Galician's 12.8k samples are almost entirely fundamental tasks from Proxecto Nós plus MGT detection, whereas Spanish's 165.2k samples span many more categories. Since the averages in Figures 7 and 8 are not stratified by task category and are not restricted to datasets that exist in multiple languages, the observed language gap could be driven by which datasets happen to be available in each language rather than by intrinsic language difficulty. The manuscript itself acknowledges this risk in Section 3.3 ("LLMs might appear weaker or stronger in cross-lingual and cross-task comparisons depending on how well they perform on overrepresented or underrepresented combinations"). The authors should add a controlled comparison—e.g., per-category language means, equal-weighting of (task, language) cells, or a matched-dataset analysis—and temper the claim if the gap does not survive that control.
  2. [Section 4.2.1, Figure 6] The headline finding that LLMs perform worse on industry-relevant than fundamental tasks is subject to the same aggregation issue. Industry datasets are predominantly Spanish (139.6k of 247.1k industry samples) and include notoriously hard categories (MGT, intent, stance, author profiling), while fundamental datasets include easier categories (Commonsense Reasoning, Question Answering) and are more dispersed across languages. A direct comparison of the two relevance groups without task or language controls conflates task difficulty with the relevance label. The authors should either match fundamental/industry pairs within language and task family, or reframe the conclusion as a claim about these specific collections, explicitly citing the per-category medians in Figure 5 rather than the aggregate in Figure 6.
  3. [Section 4.2.2, paragraph "Shared task participants still lead..."] The comparison with shared-task results is not apples-to-apples: LLMs are evaluated zero-shot on a single prompt, while the "best published models" are fine-tuned on task-specific training data, and the averaging protocol for "best published results" is unspecified (which submissions, which metrics, how many datasets per category). The reported numbers, e.g., 72.67% vs. 84.26% for humor detection, are presented without adjustment for this asymmetry. While the text mentions the zero-shot caveat, the abstract and conclusion repeat the gap without it. The authors should provide a per-dataset comparison using the same metric and test split, and state the fine-tuning/zero-shot asymmetry explicitly where the comparison is summarized.
  4. [Section 4.2.2, Figure 5 (MGT Detection and Attribution)] Figure 5 shows that the random baseline outperforms every LLM on MGT detection and attribution. The paper reports this without analysis, yet it contradicts the summary claim (iii) that these tasks are "close to random"—they are actually worse than random, which often signals a methodological artifact such as a systematic label bias in likelihood scoring or a prompt-induced majority-class response. The authors should investigate the cause (e.g., report the distribution of predicted labels, check the tokenization of label strings, verify that Macro-F1 is computed identically for the random baseline and the models) and either explain the phenomenon or soften the claim in the abstract and conclusion.
minor comments (6)
  1. [Section 3.4, Equation (1)] The product notation in Equation (1) appears corrupted in the current rendering ("|y|Y j=0"); it should read ∏_{j=1}^{|y|} pLLM(y_j | s_i, y_<j). Please also define y_<j explicitly as the token prefix before position j.
  2. [Section 3.3, reference [15]] The name of the evaluation series is written inconsistently: "IberEVAL" in the text appears as "IberEV AL" and the reference [15] uses "IberEval". Please standardize the spelling throughout the manuscript.
  3. [Table 5] For gemma-2-2b-it, the pre-training language column lists only English, although the model is reported to have broad multilingual coverage. Please clarify the selection criterion (e.g., whether the table lists primary training languages or the languages that meet some coverage threshold) to avoid misleading readers.
  4. [Section 4.2.3, paragraph on ranking consistency] The statement "in all cases, we find τ≈ρ >0.99 with p<1×10−11" is surprisingly strong given only 23 models and the presence of tied ranks in some language plots. Please report the exact statistics for each pair, state how ties are handled in Kendall's τ, and list the number of models included per language (since not all models have results for all languages).
  5. [Appendix B, Figures B.10 and B.11] The heatmaps are difficult to read because the numeric annotations are small and the color contrast between adjacent ranks is weak. Consider using a discrete color scale with larger font, or a table with a small multiples layout, so that the rankings are legible in print.
  6. [Section 3.3, dataset preparation paragraph] The mojibake example ("café" recovered from "café") contains a visible OCR/encoding artifact in the printed text ("caf ˜A©"); please check the encoding of this example in the final PDF.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: IberBench reports external benchmark measurements, and its aggregation caveats are disclosed rather than disguised.

full rationale

IberBench is a benchmark-curation and empirical-evaluation paper rather than a derivation chain, so there is no fitted parameter or constructed prediction that could reduce to its own inputs. Model scores are produced by the explicit likelihood and decoding procedures in Eqs. (1)-(2) and then compared with the reference labels of 101 existing datasets; the findings (i)-(iv) are descriptive aggregates of those measured scores. The fundamental-versus-industry relevance split is defined a priori in Section 3.3 by economic significance and task provenance, not by observed model performance, so the industry-versus-fundamental gap is an empirical comparison over a pre-registered grouping rather than a definitional artifact. The claim that Galician and Basque are harder is an average over the available (task, language) cells, and the paper itself flags the fragility of such averages: Section 3.3 states that the task-language matrix is sparse and that LLMs 'might appear weaker or stronger in cross-lingual and cross-task comparisons depending on how well they perform on overrepresented or underrepresented combinations,' and the Limitations section says that 'the task types and distributions across languages are imbalanced... All of these aspects can introduce bias to the evaluations and analyses.' That is an honest validity caveat about aggregation, not circularity, because the language-level means are not defined in terms of the conclusion that those languages are harder. The only author-overlapping citation is the IberAuTexTification overview used as a data source; the benchmark evaluates LLMs against the competition's held-out labels, so the citation functions as external, falsifiable evidence rather than as a load-bearing self-reference. No self-citation chain, uniqueness import, ansatz-smuggling, or renaming-of-known-result pattern is present. The appropriate finding is therefore no significant circularity.

Assumptions & free parameters 0 free parameters · 3 assumptions · 0 invented entities

No free parameters are fitted; the benchmark is empirical. Three domain assumptions carry the interpretability of the results: fair aggregation over imbalanced data, representative zero-shot prompting, and contamination prevention through private data hosting.

assumptions (3)
  • domain assumption Aggregate mean scores across 101 datasets give a fair basis for comparing task categories and languages.
    Used in Section 4.2.1 and Figures 3-9. The paper itself notes the (task, language) matrix is highly imbalanced (Section 3.3, Table 4), which can bias cross-task and cross-language comparisons.
  • domain assumption Zero-shot evaluation with a single prompt per task reflects realistic model performance.
    Stated as the evaluation protocol in Section 4.1; the limitations section acknowledges LLMs are sensitive to prompt phrasing and that zero-shot may underestimate performance.
  • domain assumption Holding normalized datasets in private repositories prevents test contamination.
    Section 3.3. The paper acknowledges this cannot prevent prior public release of source datasets by original authors.

how reviews work

0 comments
Cite this review

Pith. "Pith review of IberBench: LLM Evaluation on Iberian Languages." pith.science (2026). https://pith.science/paper/X5U7JSKF

@misc{pith2026250416921,
  author       = {Pith},
  title        = {Pith review of: IberBench: LLM Evaluation on Iberian Languages},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/X5U7JSKF}},
  note         = {Machine review of arXiv:2504.16921}
}
read the original abstract

Large Language Models (LLMs) remain difficult to evaluate comprehensively, particularly for languages other than English, where high-quality data is often limited. Existing benchmarks and leaderboards are predominantly English-centric, with only a few addressing other languages. These benchmarks fall short in several key areas: they overlook the diversity of language varieties, prioritize fundamental Natural Language Processing (NLP) capabilities over tasks of industrial relevance, and are static. With these aspects in mind, we present IberBench, a comprehensive and extensible benchmark designed to assess LLM performance on both fundamental and industry-relevant NLP tasks, in languages spoken across the Iberian Peninsula and Ibero-America. IberBench integrates 101 datasets from evaluation campaigns and recent benchmarks, covering 22 task categories such as sentiment and emotion analysis, toxicity detection, and summarization. The benchmark addresses key limitations in current evaluation practices, such as the lack of linguistic diversity and static evaluation setups by enabling continual updates and community-driven model and dataset submissions moderated by a committee of experts. We evaluate 23 LLMs ranging from 100 million to 14 billion parameters and provide empirical insights into their strengths and limitations. Our findings indicate that (i) LLMs perform worse on industry-relevant tasks than in fundamental ones, (ii) performance is on average lower for Galician and Basque, (iii) some tasks show results close to random, and (iv) in other tasks LLMs perform above random but below shared task systems. IberBench offers open-source implementations for the entire evaluation pipeline, including dataset normalization and hosting, incremental evaluation of LLMs, and a publicly accessible leaderboard.

Figures

Figures reproduced from arXiv: 2504.16921 by the authors.

Figure 1
Figure 1. Maps showing the Iberian languages considered in IberBench. [PITH_FULL_IMAGE:figures/full_fig_p003_1.png] view at source ↗
Figure 2
Figure 2. IberBench overview. Users can view rankings, plots, and reports; request LLMs for evaluation through the UI; and propose new datasets [PITH_FULL_IMAGE:figures/full_fig_p005_2.png] view at source ↗
Figure 3
Figure 3. Performance per model size (number of parameters), averaged across all the languages and tasks. Legend shows model families, e.g., the [PITH_FULL_IMAGE:figures/full_fig_p014_3.png] view at source ↗
Figures from the paper (6 more)
Figure 4
Figure 4. Figure 4: Averaged performance per model type. The top-3 LLMs are all from the Qwen-2.5 family. Qwen-2.5-7b-Instruct dominates the benchmark with a mean score of 46.8%, followed closely by RigoChat-7b-v2 with 46.7%, and Qwen-2.5-3b-Instruct with 45.9%. These models are among the…
Figure 5
Figure 5. Figure 5: Performance per task category across all models and languages. The random baseline for each category is marked with a black cross. [PITH_FULL_IMAGE:figures/full_fig_p016_5.png]
Figure 6
Figure 6. Figure 6: Performance in fundamental and industry-relevant tasks, averaged across all the languages and models. [PITH_FULL_IMAGE:figures/full_fig_p017_6.png]
Figure 7
Figure 7. Figure 7: Performance per language and Spanish variety averaged across LLMs and tasks. For clarity, we remove the “ambiguous” Spanish variety. [PITH_FULL_IMAGE:figures/full_fig_p018_7.png]
Figure 8
Figure 8. Figure 8: LLM performances in Iberian languages, averaged across tasks. Vertical lines denote the random baseline. [PITH_FULL_IMAGE:figures/full_fig_p018_8.png]
Figure 9
Figure 9. Figure 9: Model performance in Spanish varieties. Vertical lines denote the random baseline. [PITH_FULL_IMAGE:figures/full_fig_p019_9.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

100 extracted references · 60 canonical work pages

  1. [1]

    Eloundou, S

    T. Eloundou, S. Manning, P. Mishkin, D. Rock, Gpts are gpts: An early look at the labor market impact potential of large language models (2023). arXiv:2303.10130

  2. [2]

    Yenduri, M

    G. Yenduri, M. Ramalingam, G. C. Selvi, Y . Supriya, G. Srivastava, P. K. R. Maddikunta, G. D. Raj, R. H. Jhaveri, B. Prabadevi, W. Wang, A. V . Vasilakos, T. R. Gadekallu, Gpt (generative pre-trained transformer)— a comprehensive review on enabling technologies, potential applications, emerging challenges, and future directions, IEEE Access 12 (2024) 54608–54649

  3. [3]

    Grattafiori, A

    A. Grattafiori, A. Dubey, A. Jauhri, A. Pandey, A. Kadian, The llama 3 herd of models (2024). arXiv:2407.21783

  4. [4]

    A. Yang, B. Yang, B. Zhang, B. Hui, B. Zheng, Qwen2.5 technical report (2025). arXiv:2412.15115

  5. [5]

    D. Guo, D. Yang, H. Zhang, J. Song, R. Zhang, Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning (2025). arXiv:2501.12948

  6. [6]

    Achiam, S

    J. Achiam, S. Adler, S. Agarwal, L. Ahmad, I. Akkaya, Gpt-4 technical report (et al., 2024). arXiv:2303.08774

  7. [7]

    Georgiev, V

    P. Georgiev, V . I. Lei, R. Burnell, L. Bai, A. Gulati, Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context (2024). arXiv:2403.05530

  8. [8]

    Liang, R

    P. Liang, R. Bommasani, T. Lee, D. Tsipras, D. Soylu, et al., Holistic Evaluation of Language Models, Transactions on Machine Learning Research (2023)

Show all 100 references
  1. [9]

    Fourrier, N

    C. Fourrier, N. Habib, A. Lozovskaya, K. Szafer, T. Wolf, Open llm leaderboard v2, https://huggingface.co/spaces/ open-llm-leaderboard/open_llm_leaderboard (2024)

  2. [10]

    Chiang, L

    W.-L. Chiang, L. Zheng, Y . Sheng, A. N. Angelopoulos, T. Li, D. Li, B. Zhu, H. Zhang, M. Jordan, J. E. Gonzalez, I. Stoica, Chatbot arena: An open platform for evaluating llms by human preference, in: Proceedings of the International Conference on Machine Learning, 2024, pp. ...

  3. [11]

    Grandury, J

    M. Grandury, J. Aula-Blasco, C. Fourrier, M. Gonz ´alez, G. Mart ´ınez, G. Santamar ´ıa, A. Vaca, La leaderboard: Leaderboard of spanish varieties and official languages, https://huggingface.co/spaces/la-leaderboard/la-leaderboard (2024)

  4. [12]

    Baucells, J

    I. Baucells, J. Aula-Blasco, I. de Dios-Flores, S. Paniagua Su ´arez, N. Perez, A. Salles, S. Sotelo Docio, J. Falc ˜ao, J. J. Saiz, R. Sepul- veda Torres, J. Barnes, P. Gamallo, A. Gonzalez-Agirre, G. Rigau, M. Villegas, IberoBench: A benchmark for LLM evaluation in Iberian l...

  5. [13]

    Amig ´o, J

    E. Amig ´o, J. C. de Albornoz, A. Fern ´andez, J. Gonzalo, M. Lucas, G. Marco, R. Morante, J. Pedrosa, L. Plaza, E. S ´anchez, A. Villa, Proyecto espacio de observaci´on de inteligencia artificial en espa˜nol (odesia), https://leaderboard.odesia.uned.es/leaderboard/ tablesV2 (2025)

  6. [14]

    Chiruzzo, S

    L. Chiruzzo, S. M. Jim ´enez-Zafra, F. Rangel, Overview of iberlef 2024: Natural language processing challenges for spanish and other iberian languages, in: CEUR Workshop Proceedings, V ol. 3756, CEUR-WS, 2024

  7. [15]

    Rosso, J

    P. Rosso, J. Gonzalo, R. Mart ´ınez, S. Montalvo, J. C. de Albornoz (Eds.), Proceedings of the Third Workshop on Evaluation of Human Language Technologies for Iberian Languages (IberEval 2018) co-located with 34th Conference of the Spanish Society for Natural Language Processi...

  8. [16]

    E. M. C ´amara, Y . Almeida-Cruz, M. C. D´ıaz-Galiano, S. Est´evez-Velarde, M. ´A. G. Cumbreras, M. G. Vega, Y . Guti´errez, A. Montejo-R´aez, A. Montoyo, R. Mu˜noz, A. Piad-Morffis, J. Villena-Rom´an (Eds.), Proceedings of TASS 2018: Workshop on Semantic Analysis at SEPLN, TA...

  9. [17]

    Stamatatos, M

    E. Stamatatos, M. Potthast, F. M. R. Pardo, P. Rosso, B. Stein, Overview of the PAN /CLEF 2015 evaluation lab, in: J. Mothe, J. Savoy, J. Kamps, K. Pinel-Sauvagnat, G. J. F. Jones, E. SanJuan, L. Cappellato, N. Ferro (Eds.), Experimental IR Meets Multilinguality, Multi- modali...

  10. [18]

    Etxaniz, O

    J. Etxaniz, O. Sainz, N. Miguel, I. Aldabe, G. Rigau, E. Agirre, A. Ormazabal, M. Artetxe, A. Soroa, Latxa: An open language model and evaluation suite for Basque, in: Proceedings of the Annual Meeting of the Association for Computational Linguistics (V olume 1: Long Papers), ...

  11. [19]

    Biderman, H

    S. Biderman, H. Schoelkopf, L. Sutawika, L. Gao, J. Tow, B. Abbasi, A. F. Aji, P. S. Ammanamanchi, S. Black, J. Clive, A. DiPofi, J. Etxaniz, B. Fattori, J. Z. Forde, C. Foster, J. Hsu, M. Jaiswal, W. Y . Lee, H. Li, C. Lovering, N. Muennigho ff, E. Pavlick, J. Phang, A. Skowr...

  12. [20]

    Abdin, J

    M. Abdin, J. Aneja, H. Behl, S. Bubeck, R. Eldan, S. Gunasekar, M. Harrison, R. J. Hewett, M. Javaheripi, P. Kau ffmann, J. R. Lee, Y . T. Lee, Y . Li, W. Liu, C. C. T. Mendes, A. Nguyen, E. Price, G. de Rosa, O. Saarikivi, A. Salim, S. Shah, X. Wang, R. Ward, Y . Wu, D. Yu, C...

  13. [21]

    Gonzalez-Agirre, M

    A. Gonzalez-Agirre, M. P `amies, J. Llop, I. Baucells, S. D. Dalt, D. Tamayo, J. J. Saiz, F. Espu ˜na, J. Prats, J. Aula-Blasco, M. Mina, I. Pikabea, A. Rubio, A. Shvets, A. Sall´es, I. Lacunza, J. Palomar, J. Falc˜ao, L. Tormo, L. Vasquez-Reina, M. Marimon, O. Pareras, V . Ru...

  14. [22]

    P. H. Martins, P. Fernandes, J. Alves, N. M. Guerreiro, R. Rei, D. M. Alves, J. Pombal, A. Farajian, M. Faysse, M. Klimaszewski, P. Colombo, B. Haddow, J. G. de Souza, A. Birch, A. F. Martins, Eurollm: Multilingual language models for europe, Procedia Computer Science 255 (2025) 53–62

  15. [23]

    J. Xu, D. Ju, M. Li, Y .-L. Boureau, J. Weston, E. Dinan, Bot-adversarial dialogue for safe conversational agents, in: Proceedings of the Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, 2021, pp. 2950–2968

  16. [24]

    A. Wei, N. Haghtalab, J. Steinhardt, Jailbroken: How does LLM safety training fail?, in: Proceedings of the Conference on Neural Informa- tion Processing Systems, 2023, pp. 80079–80110

  17. [25]

    G. Bai, J. Liu, X. Bu, Y . He, J. Liu, Z. Zhou, Z. Lin, W. Su, T. Ge, B. Zheng, W. Ouyang, MT-bench-101: A fine-grained benchmark 22 for evaluating large language models in multi-turn dialogues, in: Proceedings of the Annual Meeting of the Association for Computational Linguis...

  18. [26]

    Li, W.-L

    T. Li, W.-L. Chiang, E. Frick, L. Dunlap, T. Wu, B. Zhu, J. E. Gonzalez, I. Stoica, From crowdsourced data to high-quality benchmarks: Arena-hard and benchbuilder pipeline (2024). arXiv:2406.11939

  19. [27]

    M. Wu, A. F. Aji, Style over substance: Evaluation biases for large language models, in: Proceedings of the International Conference on Computational Linguistics, 2025, pp. 297–312

  20. [28]

    this is a problem, don‘t you agree?

    S. Schoch, D. Yang, Y . Ji, “this is a problem, don‘t you agree?” framing and bias in human evaluation for natural language generation, in: Proceedings of the Workshop on Evaluating NLG Evaluation, 2020, pp. 10–16

  21. [29]

    Baumann, Universal jailbreak backdoors in large language model alignment, in: Neurips Safe Generative AI Workshop 2024, 2024

    T. Baumann, Universal jailbreak backdoors in large language model alignment, in: Neurips Safe Generative AI Workshop 2024, 2024

  22. [30]

    A. Wang, A. Singh, J. Michael, F. Hill, O. Levy, S. Bowman, GLUE: A multi-task benchmark and analysis platform for natural language understanding, in: Proceedings of the EMNLP Workshop BlackboxNLP: Analyzing and Interpreting Neural Networks for NLP, 2018, pp. 353–355

  23. [31]

    Hendrycks, C

    D. Hendrycks, C. Burns, S. Basart, A. Zou, M. Mazeika, D. Song, J. Steinhardt, Measuring massive multitask language understanding, Proceedings of the International Conference on Learning Representations (2021) 1–30

  24. [32]

    Bavaresco, R

    A. Bavaresco, R. Bernardi, L. Bertolazzi, D. Elliott, R. Fern ´andez, A. Gatt, E. Ghaleb, M. Giulianelli, M. Hanna, A. Koller, A. F. T. Martins, P. Mondorf, V . Neplenbroek, S. Pezzelle, B. Plank, D. Schlangen, A. Suglia, A. K. Surikuchi, E. Takmaz, A. Testoni, Llms instead of...

  25. [33]

    Zheng, W.-L

    L. Zheng, W.-L. Chiang, Y . Sheng, S. Zhuang, Z. Wu, Y . Zhuang, Z. Lin, Z. Li, D. Li, E. P. Xing, H. Zhang, J. E. Gonzalez, I. Stoica, Judging llm-as-a-judge with mt-bench and chatbot arena, in: Proceedings of the International Conference on Neural Information Processing Syst...

  26. [34]

    G. H. Chen, S. Chen, Z. Liu, F. Jiang, B. Wang, Humans or LLMs as the judge? a study on judgement bias, in: Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, 2024, pp. 8301–8327

  27. [35]

    M. Ali, M. Fromm, K. Thellmann, J. Ebert, A. A. Weber, R. Rutmann, C. Jain, M. L ¨ubbering, D. Steinigen, J. Leveling, K. Klug, J. S. Buschhoff, L. Jurkschat, H. Abdelwahab, B. J. Stein, K.-H. Sylla, P. Denisov, N. Brandizzi, Q. Saleem, A. Bhowmick, L. Helmer, C. John, P. O. S...

  28. [36]

    Singh, N

    H. Singh, N. Gupta, S. Bharadwaj, D. Tewari, P. Talukdar, IndicGenBench: A multilingual benchmark to evaluate generation capabilities of LLMs on Indic languages, in: Proceedings of the Annual Meeting of the Association for Computational Linguistics (V olume 1: Long Papers), 20...

  29. [37]

    Susanto, A

    Y . Susanto, A. V . Hulagadri, J. R. Montalan, J. G. Ngui, X. B. Yong, W. Leong, H. Rengarajan, P. Limkonchotiwat, Y . Mai, W. C. Tjhi, Sea-helm: Southeast asian holistic evaluation of language models (2025). arXiv:2502.14301

  30. [38]

    Almazrouei, R

    E. Almazrouei, R. Cojocaru, M. Baldo, Q. Malartic, H. Alobeidli, D. Mazzotta, G. Penedo, G. Campesan, M. Farooq, M. Alhammadi, J. Lau- nay, B. Noune, AlGhafa evaluation benchmark for Arabic language models, in: Proceedings of the Arabic Natural Language Processing Conference, ...

  31. [39]

    Duki ´c, J

    D. Duki ´c, J. Snajder, Looking right is sometimes right: Investigating the capabilities of decoder-only LLMs for sequence labeling, in: Findings of the Association for Computational Linguistic, 2024, pp. 14168–14181

  32. [40]

    S. Wang, X. Sun, X. Li, R. Ouyang, F. Wu, T. Zhang, J. Li, G. Wang, Gpt-ner: Named entity recognition via large language models (2023). arXiv:2304.10428

  33. [41]

    Houlsby, A

    N. Houlsby, A. Giurgiu, S. Jastrzebski, B. Morrone, Q. de Laroussilhe, A. Gesmundo, M. Attariyan, S. Gelly, Parameter-e fficient transfer learning for nlp, in: Proceedings of the International Conference on Machine Learning, 2019, pp. 2790–2799

  34. [42]

    Frantar, S

    E. Frantar, S. Ashkboos, T. Hoefler, D. Alistarh, OPTQ: Accurate quantization for generative pre-trained transformers, in: The Eleventh International Conference on Learning Representations, 2023

  35. [43]

    Rangel, P

    F. Rangel, P. Rosso, M. Potthast, B. Stein, Overview of the 5th author profiling task at pan 2017: Gender and language variety identification in twitter, Working notes papers of the CLEF 48 (2017)

  36. [44]

    Bandarkar, D

    L. Bandarkar, D. Liang, B. Muller, M. Artetxe, S. N. Shukla, D. Husa, N. Goyal, A. Krishnan, L. Zettlemoyer, M. Khabsa, The belebele benchmark: a parallel reading comprehension dataset in 122 language variants, in: Proceedings of the Annual Meeting of the Association for Compu...

  37. [45]

    Schae ffer, Pretraining on the test set is all you need, arXiv preprint arXiv:2309.08632 (2023)

    R. Schae ffer, Pretraining on the test set is all you need, arXiv preprint arXiv:2309.08632 (2023)

  38. [46]

    Lhoest, A

    Q. Lhoest, A. Villanova del Moral, Y . Jernite, A. Thakur, P. von Platen, Datasets: A community library for natural language processing, in: H. Adel, S. Shi (Eds.), Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing: System Demonstrations, A...

  39. [47]

    Y . Hu, Q. Chen, J. Du, X. Peng, V . K. Keloth, X. Zuo, Y . Zhou, Z. Li, X. Jiang, Z. Lu, K. Roberts, H. Xu, Improving large language models for clinical named entity recognition via prompt engineering, Journal of the American Medical Informatics 31 (9) (2024) 1812–1820

  40. [48]

    Zubiaga, I

    A. Zubiaga, I. S. Vicente, P. Gamallo, J. R. Pichel, I. Alegria, N. Aranberri, A. Ezeiza, V . Fresno, Tweetlid: a benchmark for tweet language identification, Language Resources and Evaluation 50 (2016) 729–766

  41. [49]

    Garcia-Vega, M

    M. Garcia-Vega, M. D ´ıaz-Galiano, M. Garc´ıa-Cumbreras, F. M. P. Del Arco, A. Montejo-Ra´ez, S. Jim´enez-Zafra, E. M. C´amara, C. Aguilar, M. Cabezudo, L. Chiruzzo, et al., Overview of tass 2020: Introducing emotion detection, in: Proceedings of the Iberian Languages Evalua- ...

  42. [50]

    F. M. Rangel Pardo, F. Celli, P. Rosso, M. Potthast, B. Stein, W. Daelemans, Overview of the author profiling task at pan 2015, in: Working notes papers of the CLEF, 2015, pp. 1–8

  43. [51]

    Taul ´e, F

    M. Taul ´e, F. M. R. Pardo, M. A. Mart´ı, P. Rosso, Overview of the task on multimodal stance detection in tweets on catalan# 1oct referendum, in: Proceedings of the Workshop on Evaluation of Human Language Technologies for Iberian Languages, 2018, pp. 149–166

  44. [52]

    Chiruzzo, S

    L. Chiruzzo, S. Castro, M. Etcheverry, D. Garat, J. J. Prada, A. Ros ´a, Overview of haha at iberlef 2019: Humor analysis based on human annotation, in: Proceedings of the Iberian Languages Evaluation Forum, 2019, pp. 132–144. 23

  45. [53]

    Ortega-Bueno, F

    R. Ortega-Bueno, F. Rangel, D. Hern ´andez Farıas, P. Rosso, M. Montes-y G ´omez, J. E. Medina Pagola, Overview of the task on irony detection in spanish variants, in: Proceedings of the Iberian languages evaluation forum, co-located with conference of the Spanish Society for ...

  46. [54]

    M. E. Arag ´on, M. ´A. ´Alvarez-Carmona, M. Montes-y G´omez, H. J. Escalante, L. V . Pineda, D. Moctezuma, Overview of mex-a3t at iberlef 2019: Authorship and aggressiveness analysis in mexican spanish tweets., in: Proceedings of the Iberian Languages Evaluation Forum, 2019, p...

  47. [55]

    Taul ´e, A

    M. Taul ´e, A. Ariza, M. Nofre, E. Amig ´o, P. Rosso, Overview of detoxis at iberlef 2021: Detection of toxicity in comments in spanish, Procesamiento del Lenguaje Natural 67 (2021) 209–221

  48. [56]

    F. M. Plaza-del Arco, S. M. Jim´enez-Zafra, A. Montejo-R´aez, M. D. Molina-Gonz´alez, L. A. Ure˜na-L´opez, M. T. Mart´ın-Valdivia, Overview of the emoevales task on emotion detection for spanish at iberlef 2021, Procesamiento del Lenguaje Natural 67 (2021) 155–161

  49. [57]

    F. J. Rodr ´ıguez-S´anchez, J. Carrillo-de Albornoz, L. Plaza, J. Gonzalo, P. Rosso, M. Comet, T. Donoso, Overview of exist 2021: sexism identification in social networks, Procesamiento del Lenguaje Natural 67 (2021) 195–207

  50. [58]

    G ´omez-Adorno, J

    H. G ´omez-Adorno, J. P. Posadas-Dur ´an, G. Bel Enguix, C. Porto Capetillo, Overview of fakedes at iberlef 2021: Fake news detection in spanish shared task, Procesamiento del Lenguaje Natural 67 (2021) 223–231

  51. [59]

    Chiruzzo, S

    L. Chiruzzo, S. Castro, S. G ´ongora, A. Ros ´a, J. Meaney, R. Mihalcea, Overview of haha at iberlef 2021: Detecting, rating and analyzing humor in spanish, Procesamiento del Lenguaje Natural 67 (2021) 257–268

  52. [60]

    Jarqu ´ın-V´asquez, L

    H. Jarqu ´ın-V´asquez, L. Villase˜nor Pineda, F. M. Plaza-del Arco, M. Casavantes, H. J. Escalante, M. T. Mart ´ın Valdivia, A. Montejo R´aez, M. Montes y G´omez, Overview of meoffendes at iberlef 2021: Offensive language detection in spanish variants, Procesamiento del Lengua...

  53. [61]

    M. ´A. ´Alvarez Carmona, R. Aranda, S. Arce-Cardenas, D. Fajardo-Delgado, R. Guerrero-Rodr ´ıguez, A. P. L ´opez-Monroy, J. Mart´ınez- Miranda, H. P ´erez-Espinosa, A. Y . Rodr´ıguez-Gonz´alez, Overview of rest-mex at iberlef 2021: Recommendation system for text mexican touris...

  54. [62]

    Agerri Gasc ´on, R

    R. Agerri Gasc ´on, R. Centeno S´anchez, M. Espinosa, J. Fern´andez de Landa, ´A. Rodrigo Yuste, Vaxxstance@iberlef 2021: Overview of the task on going beyond text in cross-lingual stance detection, Procesamiento del Lenguaje Natural 67 (2021) 173–181

  55. [63]

    Bel-Enguix, G

    G. Bel-Enguix, G. Sierra, H. G ´omez-Adorno, J.-M. Torres-Moreno, J.-G. Ortiz-Barajas, J. V ´asquez, Overview of par-mex at iberlef 2022: Paraphrase detection in spanish shared task, Procesamiento del Lenguaje Natural 69 (2022) 255–263

  56. [64]

    M. ´A. ´Alvarez Carmona, ´A. D´ıaz-Pacheco, R. Aranda, A. Y . Rodr´ıguez-Gonz´alez, D. Fajardo-Delgado, R. Guerrero-Rodr´ıguez, L. Bustio- Mart´ınez, Overview of rest-mex at iberlef 2022: Recommendation system, sentiment analysis and covid semaphore prediction for mexican tour...

  57. [65]

    A. M. M ´armol-Romero, A. Moreno-Mu ˜noz, F. M. Plaza-del Arco, M. D. Molina-Gonz ´alez, M. T. Mart ´ın-Valdivia, L. A. Ure ˜na-L´opez, A. Montejo-R´aez, Overview of mentalriskes at iberlef 2023: Early detection of mental disorders risk in spanish, Procesamiento del Lenguaje N...

  58. [66]

    Labadie Tamayo, B

    R. Labadie Tamayo, B. Chulvi, P. Rosso, Everybody hurts, sometimes. overview of hurtful humour at iberlef 2023: Detection of humour spreading prejudice in twitter, Procesamiento del Lenguaje Natural 71 (2023) 383–395

  59. [67]

    W. S. S.-N. y Pol Pastells y Simona Frenda y Alejandro Ariza-Casabona y Mireia Farr ´us y Paolo Rosso y Mariona Taul ´e, Overview of detests-dis at iberlef 2024: Detection and classification of racial stereotypes in spanish - learning with disagreement, Procesamiento del Lengu...

  60. [68]

    A. M. S. y Jos ´e ´Angel Gonz ´alez y Francisco Rangel y Paolo Rosso y Marc Franco-Salvador, Overview of iberautextification at iberlef 2024: Detection and attribution of machine-generated text on languages of the iberian peninsula, Procesamiento del Lenguaje Natural 73 (0) (2...

  61. [69]

    Y . Yang, Y . Zhang, C. Tar, J. Baldridge, PAWS-X: A cross-lingual adversarial dataset for paraphrase identification, in: Proceedings of the Conference on Empirical Methods in Natural Language Processing and the International Joint Conference on Natural Language Processing, 20...

  62. [70]

    A. I. Vladu, I. de Dios-Flores, C. Magari ˜nos, J. E. Ortega, J. R. Pichel, M. Garcia, P. Gamallo, E. F. Rei, A. Bugar ´ın, M. G. Gonz ´alez, et al., Proxecto n´os: Artificial intelligence at the service of the galician language, in: Proceedings of the Annual Conference of the...

  63. [71]

    Gonzalez-Agirre, M

    A. Gonzalez-Agirre, M. Marimon, C. Rodriguez-Penagos, J. Aula-Blasco, I. Baucells, C. Armentano-Oller, J. Palomar-Giner, B. Kulebi, M. Villegas, Building a data infrastructure for a mid-resource language: The case of Catalan, in: Proceedings of the Joint International Conferen...

  64. [72]

    Hasan, A

    T. Hasan, A. Bhattacharjee, M. S. Islam, K. Mubasshir, Y .-F. Li, Y .-B. Kang, M. S. Rahman, R. Shahriyar, XL-sum: Large-scale multilingual abstractive summarization for 44 languages, in: Findings of the Association for Computational Linguistics, 2021, pp. 4693–4703

  65. [73]

    URL https://projecteaina.cat/

    Barcelona Supercomputing Center, Projecte aina: Recursos ling ¨u´ıstics i tecnol`ogics per al catal`a a l’era digital (2021). URL https://projecteaina.cat/

  66. [74]

    Urbizu, I

    G. Urbizu, I. San Vicente, X. Saralegi, R. Agerri, A. Soroa, Basqueglue: A natural language understanding benchmark for basque, in: Proceedings of the Language Resources and Evaluation Conference, Marseille, France, 2022, pp. 1603–1612

  67. [75]

    A. V . Serrano, I. L. Montalb´an, D. V . Vel´azquez, M. Moreno, Clindiagnoses (2024). URL https://huggingface.co/datasets/LenguajeNaturalAI/ClinDiagnosES

  68. [76]

    R ¨ottger, H

    P. R ¨ottger, H. Seelawi, D. Nozza, Z. Talat, B. Vidgen, Multilingual HateCheck: Functional tests for multilingual hate speech detection models, in: Proceedings of the Workshop on Online Abuse and Harms, 2022, pp. 154–169

  69. [77]

    ´Alvarez Mellado, L

    E. ´Alvarez Mellado, L. Espinosa Anke, J. G. Arroyo, C. Lignos, J. Porta Zamorano, Overview of adobo 2021: Automatic detection of unassimilated borrowings in the spanish press, Procesamiento del Lenguaje Natural 67 (2021) 277–285

  70. [78]

    Armengol-Estap ´e, C

    J. Armengol-Estap ´e, C. P. Carrino, C. Rodriguez-Penagos, O. de Gibert Bonet, C. Armentano-Oller, A. Gonzalez-Agirre, M. Melero, M. Villegas, Are multilingual models the best choice for moderately under-resourced languages? A comprehensive assessment for Catalan, in: Findings...

  71. [79]

    https://proyectoilenia.es/

    Secretar ´ıa de Estado de Digitalizaci´on e Inteligencia Artificial, Proyecto del impulso de las lenguas en la inteligencia artificial (ilenia) (2022). 24 URL "https://proyectoilenia.es/"

  72. [80]

    N. Bel, M. Punsola, V . Ruiz-Fern ´andez, EsCoLA: Spanish corpus of linguistic acceptability, in: Proceedings of the Joint International Conference on Computational Linguistics, Language Resources and Evaluation, 2024, pp. 6268–6277

  73. [81]

    N. Bel, M. Punsola, V . Ruiz-Fern ´andez, Catcola, catalan corpus of linguistic acceptability, Procesamiento del Lenguaje Natural 73 (2024) 177–190

  74. [82]

    Mayor-Rocher, N

    M. Mayor-Rocher, N. Melero, E. Merino-G ´omez, M. Gonz ´alez, R. Ferrando, J. Conde, P. Reviriego, Spanish Language Benchmark for Artificial Intelligence Models (TELEIA) (2024). URL https://doi.org/10.5281/zenodo.12571763

  75. [83]

    Conneau, R

    A. Conneau, R. Rinott, G. Lample, A. Williams, S. R. Bowman, H. Schwenk, V . Stoyanov, Xnli: Evaluating cross-lingual sentence repre- sentations, in: Proceedings of the Conference on Empirical Methods in Natural Language Processing, 2018, pp. 2475–2485

  76. [84]

    Lin, ROUGE: A package for automatic evaluation of summaries, in: Text Summarization Branches Out, Association for Computational Linguistics, 2004, pp

    C.-Y . Lin, ROUGE: A package for automatic evaluation of summaries, in: Text Summarization Branches Out, Association for Computational Linguistics, 2004, pp. 74–81

  77. [85]

    T. He, J. Zhang, T. Wang, S. Kumar, K. Cho, J. Glass, Y . Tsvetkov, On the blind spots of model-based evaluation metrics for text generation, in: Proceedings of the Annual Meeting of the Association for Computational Linguistics (V olume 1: Long Papers), 2023, pp. 12067–12097

  78. [86]

    Nakayama, seqeval: A python framework for sequence labeling evaluation (2018)

    H. Nakayama, seqeval: A python framework for sequence labeling evaluation (2018). URL https://github.com/chakki-works/seqeval

  79. [87]

    Rafailov, A

    R. Rafailov, A. Sharma, E. Mitchell, C. D. Manning, S. Ermon, C. Finn, Direct preference optimization: Your language model is secretly a reward model, in: Advances in Neural Information Processing Systems, V ol. 36, Curran Associates, Inc., 2023, pp. 53728–53741

  80. [88]

    Schulman, F

    J. Schulman, F. Wolski, P. Dhariwal, A. Radford, O. Klimov, Proximal policy optimization algorithms (2017). arXiv:1707.06347

  81. [89]

    Abouelenin, A

    A. Abouelenin, A. Ashfaq, A. Atkinson, H. Awadalla, N. Bach, Phi-4-mini technical report: Compact yet powerful multimodal language models via mixture-of-loras (et al., 2025). arXiv:2503.01743

  82. [90]

    A. Q. Jiang, A. Sablayrolles, A. Mensch, C. Bamford, D. S. Chaplot, D. de las Casas, F. Bressand, G. Lengyel, G. Lample, L. Saulnier, L. R. Lavaud, M.-A. Lachaux, P. Stock, T. L. Scao, T. Lavril, T. Wang, T. Lacroix, W. E. Sayed, Mistral 7b (2023).arXiv:2310.06825

  83. [91]

    Radford, J

    A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, I. Sutskever, Language models are unsupervised multitask learners, technical report (2019)

  84. [92]

    Riviere, S

    M. Riviere, S. Pathak, P. G. Sessa, C. Hardin, S. Bhupatiraju, Gemma 2: Improving open language models at a practical size (2024). arXiv:2408.00118

  85. [93]

    G. S. G ´omez, G. G. Subies, P. G. Ruiz, M. G. Valero, N. Fuertes, H. M. Zamorano, C. M. Sanz, L. R. Plaza, N. A. Garc ´ıa, D. B. S ´anchez, K. Sushkova, M. G. Nieto, ´Alvaro Barbero Jim´enez, Rigochat 2: an adapted language model to spanish using a bounded dataset and reduced...

  86. [94]

    Petrea, Catallama, https://huggingface.co/catallama/CataLlama-v0.2-Instruct-SFT (2024)

    L. Petrea, Catallama, https://huggingface.co/catallama/CataLlama-v0.2-Instruct-SFT (2024)

  87. [95]

    Pires, H

    R. Pires, H. Abonizio, T. S. Almeida, R. Nogueira, Sabi ´a: Portuguese large language models, in: Intelligent Systems, 2023, pp. 226–240

  88. [96]

    Language, I. S. Group, Aitana, https://huggingface.co/gplsi/Aitana-6.3B (2024)

  89. [97]

    Q. Dong, L. Li, D. Dai, C. Zheng, J. Ma, R. Li, H. Xia, J. Xu, Z. Wu, B. Chang, X. Sun, L. Li, Z. Sui, A survey on in-context learning, in: Proceedings of the Conference on Empirical Methods in Natural Language Processing, 2024, pp. 1107–1128

  90. [98]

    Ho ffmann, S

    J. Ho ffmann, S. Borgeaud, A. Mensch, E. Buchatskaya, T. Cai, E. Rutherford, D. de Las Casas, L. A. Hendricks, J. Welbl, A. Clark, et al., Training compute-optimal large language models, in: Proceedings of the International Conference on Neural Information Processing Systems, ...

  91. [99]

    O’Rourke, The galician language in the twenty-first century, A companion to Galician culture 344 (2014) 73

    B. O’Rourke, The galician language in the twenty-first century, A companion to Galician culture 344 (2014) 73

  92. [100]

    Tovar, H

    A. Tovar, H. P. Houghton, The Basque Language, University of Pennsylvania Press, 1957. 25 Appendix A. Datasets and Sources Table A.6 lists the main URLs of the workshops, shared tasks and other sources we include in IberBench as industry-relevant tasks. Table A.7 shows the URL...

Pith tools

Reviewed August 16, 2026 · model on record in the stance chip above.