{"id":"2ae48965-f1dc-4eb0-9e6e-004f9076910b","arxiv_id":"2505.13772","paper_version":2,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"A Greek-focused continual pretraining and post-training pipeline yields Llama-Krikri-8B, which surpasses prior open models on Greek benchmarks and introduces three new Greek evaluation datasets.","lead":"Llama-Krikri-8B is an open 8-billion-parameter language model for Greek, built by further training Meta's Llama 3.1-8B on a 110-billion-token Greek and multilingual corpus. It outperforms other open Greek and multilingual models on Greek understanding, generation, and instruction-following benchmarks, and ships with three newly translated Greek evaluation sets.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The Greek benchmark numbers are credible, but the headline chat-score superiority is dominated by judge/translation effects; the +34.8 IFEval gap is the most robust claim.","rationale":"The reader identified the same load-bearing assumption (judge bias in chat benchmarks and translated-benchmark validity) and correctly conditioned the verdict, so I agree with the CONDITIONAL verdict. I disagree on one small point: the reader treated the IFEval Greek numbers as solid; IFEval is an automatically-verifiable instruction-following metric, so the +21.7 IFEval gain is much more robust than the MT-Bench/ArenaHard gains, which rely on LLM-as-a-judge. The most fragile numbers are the ArenaHard Greek results (Table 7), especially the 31.8 vs 4.0 gap versus Llama-3.1-8B-Instruct, since the paper cites the judge-bias literature. The paper also has a small internal inconsistency: the abstract says 'three novel public benchmarks' cover 'instruction-following, multi-turn chats, code/math tasks', but the three new benchmarks are only IFEval, MT-Bench, and ArenaHard; the code/math and agentic claims rest on no new original-Greek code/math benchmark. This is a claim-support gap, but not the main load-bearing concern. My concrete test (changing the judge) would directly test whether the central chat-superiority claim survives. The base-model results (Table 3) are the strongest independent evidence, and the paper's own limitations section admits the translated-benchmark issue, so the conclusion should remain CONDITIONAL until the judge-bias check is done. I would not reject or accept the paper outright; the engineering contribution and the base-model gains are credible, but the headline 'competitive with 3-4x larger models' relies on scores that are exactly the ones most influenced by judge preference.","tokens_in":30876,"tokens_out":3419,"duration_ms":23778,"concrete_test":"Re-evaluate Llama-Krikri-8B-Instruct, Llama-3.1-8B-Instruct, and Aya Expanse 8B on the same Greek MT-Bench/ArenaHard/IFEval prompts with a different judge (e.g., an independent human panel of Greek speakers, or a different strong LLM judge such as Claude-3.5-Sonnet or a Greek-native judge model) and report per-item agreement (e.g., Cohen's kappa) between judges. If the Krikri vs Llama-3.1-8B-Instruct lead shrinks or reverses when the judge is changed, then the chat-score superiority is not robust and the central claim should be narrowed to the base-model Greek accuracy and IFEval gains. Alternatively, run the same ArenaHard Greek comparison with GPT-4o-Mini models of different families as the baseline model, or replace the judge with one that has not seen Krikri-style training data, to quantify judge-bias sensitivity.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central claim that Llama-Krikri-8B-Instruct 'significantly outperforms' open multilingual models on Greek tasks is carried mostly by Table 5 (IFEval EL +21.7 over Llama-3.1-8B-Instruct, MT-Bench EL 7.96, other models 6.25-7.68) and Table 7 (ArenaHard Greek 31.8 vs 4.0 for Llama-3.1-8B-Instruct). The weakest load-bearing link is the reliability of these translated and LLM-judged benchmarks. The 27.8-point ArenaHard gap versus Llama-3.1-8B-Instruct is especially fragile: it is computed with GPT-4o-Mini as baseline and GPT-4o as judge, while the authors themselves (Section 4.2, Arena Hard paragraph) cite evidence that judges favor teacher-distilled models, and Llama-Krikri was post-trained on MAGPIE synthetic data from such teachers. For a non-Greek-aware evaluator, the Greek ArenaHard scores are inflated by judge bias toward model X's output style over that of Llama-3.1-8B-Instruct, which the paper explicitly states it cannot rule out. The MT-Bench Greek numbers (7.96 vs Llama-3.1-8B-Instruct's 6.46) and the IFEval Greek numbers (67.5 vs 45.8) likewise depend on a translated/post-edited Greek test set whose validity is supported only by a stated human post-editing step without an inter-annotator agreement study or comparison to any original-Greek benchmark. In contrast, the base-model Greek improvements (Table 3: Greek average 59.5 vs 48.7, +10.8) and the English-retention evidence (Table 4: 67.0 vs 66.2) are concrete and reproducible, since they use established benchmarks from the ILSP Greek evaluation suite. The single most load-bearing concern is therefore the unverified assumption in the chat-model evaluation that a translated Greek MT-Bench/ArenaHard/IFEval, together with GPT-4o judging, measures Greek conversational quality rather than translation artifacts or judge preference for the Krikri output style.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces Llama-Krikri-8B, a Greek-focused large language model obtained by continual pretraining of Meta's Llama 3.1-8B on a curated 110B-token corpus, with tokenizer and embedding expansion, an annealing phase, and a multi-stage SFT/DPO post-training pipeline. The paper also presents three new Greek chat benchmarks (IFEval Greek, MT-Bench Greek, Arena-Hard Greek) and evaluates base and instruct models on Greek and English tasks, reporting substantial gains over existing open multilingual models and competitiveness with models 3-4 times larger, as well as an Ancient-Modern Greek translation experiment. The central claim is that Llama-Krikri-8B and Llama-Krikri-8B-Instruct significantly outperform comparable open models on Greek tasks while retaining English capabilities.","tokens_in":31414,"tokens_out":5505,"duration_ms":48633,"significance":"If the evaluation concerns are addressed, this is a useful contribution to Greek NLP: the paper releases open models, a detailed training recipe, tokenizer statistics, and three public Greek evaluation benchmarks, and the base-model Greek improvements (Table 3, +10.8 average over Llama-3.1-8B) and English retention (Table 4, +0.8) are concrete and reproducible with established benchmarks. The strengths are the transparency of the corpus and pipeline description, the public release of benchmarks, and the careful attention to tokenizer fertility for Greek. The main weaknesses are that the headline chat-model superiority rests on translated and LLM-judged benchmarks whose validity is not yet established, and no uncertainty quantification is provided for any of the reported scores.","major_comments":[{"comment":"The Greek Arena-Hard win rates (31.8 for Llama-Krikri-8B-Instruct vs 4.0 for Llama-3.1-8B-Instruct) are computed with GPT-4o-Mini as the baseline and GPT-4o as the judge, and the paper itself cites Li et al. (2025) showing that judges are biased toward student models trained on distilled teacher data. Since Llama-Krikri-8B-Instruct was post-trained on MAGPIE and other synthetic data from GPT-4-class teachers, this judge-bias confound directly affects the headline gap; the claim of significant outperformance on this benchmark is not load-bearing without a human-preference check or a second judge with an independent decision rule. I would ask for a robustness analysis on a random subset of prompts (e.g., 50-100) with human ratings, or at minimum a sensitivity analysis using a different judge and an explicit agreement metric.","section":"Section 4.2, Table 7"},{"comment":"The three newly proposed benchmarks are translations and post-edits of English benchmarks, and Section 6 acknowledges that future evaluation should include original Greek datasets to minimize translation effects. The manuscript reports human post-editing but provides no inter-annotator agreement, no translation quality metrics, and no comparison with an original-Greek benchmark; translated prompts can change difficulty because Greek morphology and word-order constraints differ from English. Therefore the size of the Greek IFEval gap (+21.7 over Llama-3.1-8B-Instruct) and the MT-Bench Greek lead are not yet established as true Greek-capability differences. I would ask for translation quality statistics, a back-translation consistency check, or a comparison against a non-translated Greek instruction-following set before treating these benchmarks as decisive.","section":"Section 4.2, IFEval Greek / MT-Bench Greek / Arena-Hard Greek"},{"comment":"All reported scores are single-run point estimates with no error bars, confidence intervals, or significance tests; differences such as the MT-Bench Greek 7.96 vs 7.68 for Aya Expanse 8B, or the Open LLM Leaderboard average 24.18 vs 23.76, may be within run-to-run noise. The word \"significantly\" is used repeatedly (Section 4.2, Section 5) without statistical support; I would ask for repeated evaluations or bootstrap CIs on the main comparisons, or a downgrading of the wording to \"directionally outperforms.\"","section":"Tables 3-7"},{"comment":"The claim of highly accurate handling of Ancient Greek rests on a 100-sentence set with only BLEU scores (54.66 grc->ell, 20.41 ell->grc) and no baseline scores from Llama-3.1-8B-Instruct, Meltemi, or a translation system, and no human evaluation. This is too small and too weakly benchmarked to support the contribution bullet about Ancient Greek; I would ask for baselines on the same set, a larger evaluation sample, and ideally human judgments of adequacy and fluency.","section":"Section 4.2, Ancient-Modern Greek translation"}],"minor_comments":[{"comment":"The caption contains the typo \"extention\" which should read \"extension.\"","section":"Section 3.2, Appendix A.3, Table 8"},{"comment":"The row labeled \"Llama-Krikri-8B\" should be labeled \"Llama-Krikri-8B-Instruct\" to match the other chat models in the table and avoid ambiguity with the base model.","section":"Section 4.2, Table 5"},{"comment":"The description says the benchmark contains 80 multi-turn conversations but does not state whether the full MT-Bench set was used or how many prompts were discarded during post-editing; please specify the exact source and filtering process.","section":"Section 4.2, MT-Bench Greek"},{"comment":"The manuscript mentions using the style-control version of Arena-Hard but does not give the exact revision or style-control parameters; please add this to the appendix for reproducibility.","section":"Section 4.2, Arena-Hard Greek"},{"comment":"The introduction and contributions mention function calling and agentic behavior, but no evaluation of function calling is reported; either add such an evaluation or soften the claim.","section":"Section 1 and Section 4"},{"comment":"The sentence stating that future benchmarks should include original Greek datasets that are not the result of machine translation and post-editing directly qualifies the central comparative claims and should be moved or expanded in the main evaluation section, not only in the limitations.","section":"Section 6, Limitations"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is within scope for an NLP/AI journal, and the released models and benchmark datasets are valuable assets. My main concern is the gap between the hedged limitations (Sections 4.2 and 6) and the unhedged abstract and conclusion claims; the stress-test concern about Arena-Hard judge bias lands, because the paper itself cites the relevant bias result yet still uses the Arena-Hard Greek numbers to support the central claim. The base-model Greek results and the IFEval Greek gap are more credible, so I would not reject the paper; a major revision that adds robustness analyses, uncertainty quantification, and tempered wording would make it acceptable."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: this is a solid, well-documented Greek LLM release, and the base-model results are the part to trust. The headline chat-model wins over Llama-3.1-8B-Instruct are plausible but softer than the abstract suggests, because the new Greek chat benchmarks are translations judged by GPT-4o.\n\nWhat's actually new: a continual-pretrained Llama-3.1-8B with a 21k-token Greek tokenizer extension, a two-stage SFT + DPO post-training pipeline with synthetic data, and three public Greek evaluation benchmarks (IFEval-el, MT-Bench-el, ArenaHard-el). The model weights, tokenizer, and benchmarks are public. The tokenizer fertility numbers are concrete and useful: Greek fertility drops from 2.73 to 1.65 with no English regression. The base-model evaluation on the existing ILSP Greek suite shows a +10.8 average gain over Llama-3.1-8B, and English retention is good (+0.8 on their six English benchmarks). The annealing ablation (Table 11) is a nice piece of evidence that synthetic QA helps. That part is reproducible and I believe it.\n\nThe soft spots are all in the chat evaluation. The three new Greek benchmarks are translations/post-edits of English tests, and the paper provides no inter-annotator agreement or comparison against any original-Greek benchmark. The MT-Bench and ArenaHard scores use GPT-4o as judge. The authors themselves cite Li et al. (2025) and acknowledge that judges favor teacher-distilled models; Krikri-Instruct was trained on MAGPIE-style synthetic data from such teachers. So the 27.8-point ArenaHard gap over Llama-3.1-8B-Instruct is exactly the number most likely to be inflated by judge preference. There are also no error bars or significance tests anywhere. The Ancient Greek translation eval is 100 sentences, which is anecdotal. None of this kills the paper, but it means the 'significantly outperforms' claim in the abstract should be read as applying to the base-model suite, not the chat numbers.\n\nThe citation pattern is fine; the related work is standard and the self-citations to their own Meltemi work are appropriate. The Limitations section is honest about the translated-benchmark issue.\n\nWho this is for: anyone doing continual pretraining for medium-resource European languages, and anyone building Greek NLP resources. It deserves a serious referee. I'd send it out and ask for the chat-benchmark claims to be scaled back or better supported, but I would not desk-reject.","headline":"The base-model Greek gains are the credible story here; the chat-benchmark wins are real but rest on translated tests and a judge the paper itself admits may be biased.","tokens_in":31868,"tokens_out":2040,"would_cite":true,"duration_ms":19163,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A Greek-focused 8B model, built from Llama 3.1 by continual pretraining and synthetic-data post-training, outperforms open multilingual rivals on Greek tasks and holds its own against models three to four times its size.","keywords":["Greek language model","continual pretraining","tokenizer expansion","instruction tuning","direct preference optimization","synthetic data","Greek benchmarks","multilingual LLM"],"falsifier":"Have native Greek speakers blindly rank Llama-Krikri-8B-Instruct, Aya Expanse 32B, and Gemma 2 27B on the same MT-Bench Greek and Arena-Hard Greek prompts, then compare those rankings with GPT-4o's scores; a ranking reversal would show the headline chat-model lead is partly a judge artifact.","tokens_in":30716,"feed_emoji":"🇬🇷","tokens_out":9276,"duration_ms":77799,"temperature":0.7,"pith_summary":"Llama-Krikri-8B is an open Greek-language model built by taking Meta's Llama 3.1-8B and continually pretraining it on roughly 110 billion tokens, about 60 percent of them Greek, after expanding the tokenizer with 20,992 Greek-domain tokens. The paper's central claim is that this adaptation, followed by a multi-stage post-training pipeline using synthetic instruction data and preference optimization, makes an 8-billion-parameter model the strongest open choice for Greek, ahead of both the previous Greek model and general multilingual models of similar size. The instruct version also posts competitive scores against models with three to four times more parameters on Greek benchmarks. To make the case, the paper introduces three new Greek benchmarks built by translating and post-editing IFEval, MT-Bench, and Arena-Hard. A fair reader would care because low- and medium-resource languages rarely get open models with this level of capability, and the recipe could be reused.","feed_headline":"A Greek-trained 8B model beats bigger open rivals on Greek tasks","feed_subtitle":"Continual pretraining, tokenizer expansion, and synthetic instruction data push Llama-Krikri past 27B-32B models on Greek benchmarks.","key_machinery":"The mechanism is a full adaptation pipeline: a domain-expanded tokenizer (149,248 tokens; Greek fertility falls from 2.73 to 1.65 tokens per word), embedding initialization by averaging the old tokenizer's embeddings over each new token, mixed-curriculum continual pretraining on a 110B-token mix with English and code replay segments, an annealing phase on curated and synthetic question-answer data, then two-stage supervised fine-tuning and direct preference optimization with length normalization for the chat version. The evaluation machinery consists of three newly translated and post-edited Greek benchmarks (IFEval Greek, MT-Bench Greek, Arena-Hard Greek) plus six existing Greek tasks and six English tasks; the Greek Arena-Hard and MT-Bench use GPT-4o family models as judge.","core_discovery":"According to the paper, continual pretraining on a Greek-heavy corpus is enough to make Llama 3.1-8B substantially better at Greek without giving back English competence. The base model averages 59.5 percent accuracy on six Greek benchmarks versus 48.7 percent for Llama-3.1-8B and 47.9 percent for Meltemi-7B, and it even edges the base model on English by 0.8 points on average. The instruct variant reaches 67.5 percent on IFEval Greek, 7.96 on MT-Bench Greek, and a 31.8 percent win rate on Arena-Hard Greek, beating Aya Expanse 8B and Llama-3.1-70B-Instruct while matching Gemma 2 27B on the Greek Arena-Hard. The paper reads these results as evidence that data synthesis and careful post-training can close the resource gap for a medium-resource language.","pith_inferences":["Because the three new chat benchmarks are translations of English tests, part of the measured gain may reflect improved English-style instruction following rather than Greek-specific linguistic skill; a fully original Greek benchmark would isolate the language-specific contribution.","The same pipeline of tokenizer expansion, curriculum continual pretraining, synthetic instruction data, and DPO should transfer to other medium-resource European languages that share a strong multilingual base model.","The paper's own caveat about judge bias implies that the Arena-Hard Greek margins over larger models could shrink if a judge not derived from GPT-4-family training were used; a human-preference replication would settle how much of the lead is real."],"forward_implications":["An 8B Greek model reaches 67.5 percent on IFEval Greek, surpassing Llama-3.1-8B-Instruct by 21.7 points and the previous Greek model Meltemi by 34.8 points.","On Arena-Hard Greek, the instruct model's 31.8 percent win rate puts it ahead of Llama-3.1-70B-Instruct at 27.4 percent despite being roughly one-ninth the size.","The base model improves Greek benchmark accuracy by 10.8 points over Llama-3.1-8B while retaining and slightly improving English performance by 0.8 points, so catastrophic forgetting is not observed in this setup.","The expanded tokenizer halves Greek token consumption, directly lowering inference cost for Greek users.","The model extends to polytonic and Ancient Greek, with a reported 54.66 BLEU score for ancient-to-modern Greek translation."],"supporting_citations":[{"why":"Supplies Llama 3.1-8B as the base architecture and the English-capability reference point the model is measured against.","marker":"Grattafiori et al., 2024"},{"why":"Defines Meltemi-7B, the previous open Greek LLM and the main baseline, along with the existing Greek evaluation suite.","marker":"Voukoutis et al., 2024"},{"why":"MAGPIE, the alignment-data synthesis method used to generate Greek instruction and preference data.","marker":"Xu et al., 2024"},{"why":"Direct Preference Optimization, the alignment objective that produces the Instruct model.","marker":"Rafailov et al., 2024"},{"why":"Original IFEval, whose translated and post-edited Greek version is one of the three new benchmarks.","marker":"Zhou et al., 2023"},{"why":"MT-Bench and the LLM-as-judge protocol used to score multi-turn chat quality.","marker":"Zheng et al., 2023"},{"why":"Arena-Hard-Auto v0.1, the adversarial prompt set translated into Greek for the third benchmark.","marker":"Li et al., 2024"},{"why":"Documents judge bias toward student models distilled from the teacher-judge, the caveat that qualifies the Arena-Hard results.","marker":"Li et al., 2025"},{"why":"Belebele, a parallel reading-comprehension benchmark in 122 languages used in the base-model Greek and English evaluations.","marker":"Bandarkar et al., 2024"},{"why":"Defines tokenizer fertility, the metric used to argue the expanded tokenizer halves Greek token cost.","marker":"Csaki et al., 2023"}],"fun_headline_variants":["8B Greek model beats 70B, matches 27B on Greek tasks","Continual pretraining on Greek lifts 8B above 70B rivals","Llama-Krikri: Greek data beats scale—8B edges 70B","Greek-first 8B model beats larger rivals on native benchmarks","Small Greek LLM out scores a 70B model"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the translated Greek benchmarks and the GPT-4o judge measure Greek capability fairly; if the judge favors responses resembling its own training style, or if translation artifacts make the tests easier for models trained on English-distilled data, the reported leads over other open models shrink.","fun_headline_variants_meta":{"raw":{"variants":["8B Greek model beats 70B, matches 27B on Greek tasks","Continual pretraining on Greek lifts 8B above 70B rivals","Llama-Krikri: Greek data beats scale—8B edges 70B","Greek-first 8B model beats larger rivals on native benchmarks","Small Greek LLM out scores a 70B model"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001379,"raw_usage":{"total_tokens":5569,"prompt_tokens":915,"completion_tokens":4654,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":531,"completion_tokens_details":{"reasoning_tokens":4557}},"tokens_in":531,"tokens_out":4654,"duration_ms":32765,"temperature":1.0,"reasoning_tokens":4557,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T20:10:14.000577+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Have native Greek speakers blindly rank Llama-Krikri-8B-Instruct, Aya Expanse 32B, and Gemma 2 27B on the same MT-Bench Greek and Arena-Hard Greek prompts, then compare those rankings with GPT-4o's scores; a ranking reversal would show the headline chat-model lead is partly a judge artifact.","supporting_citations":[],"review_version":1}