{"id":"035e39c2-4e91-444b-9284-b4bb8c04eda8","arxiv_id":"2509.02550","paper_version":1,"verdict":"ACCEPT","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"PalmX 2025 introduces a two-subtask MCQA benchmark for Arabic and Islamic cultural knowledge and shows that task-specific fine-tuning, especially LoRA, improves LLM accuracy.","lead":"The paper presents PalmX 2025, the first shared task benchmarking how well large language models understand Arabic and Islamic culture through multiple-choice questions. It reports that task-specific fine-tuning improves accuracy, with top scores of 72.15% on cultural questions and 84.22% on Islamic knowledge.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Fine-tuning gains are confounded by base-model choice: Tables 4–5 compare fine-tuned models of different sizes and families against a single NileChat-3B zero-shot baseline, so the Section 5.1 claim is not established for most submissions.","rationale":"Good-faith reading: the paper delivers a useful new benchmark and transparently reports system descriptions; I would not reject it. However, the headline finding is a comparative claim, and the comparison is not controlled. The reader's concern about answer-key correctness is important but is partially mitigated by dual linguist review and would affect all systems roughly equally; a systematic key error would change absolute scores, but would not by itself explain the fine-tuning-vs-baseline gap. The baseline mismatch, by contrast, directly determines whether the central claim is true. The fix is modest: add zero-shot scores for each base model used (or at least ALLaM-7B) and re-state the conclusion to say that the winning submissions, which fine-tuned, beat the provided baseline. Under that revision, ACCEPT would be warranted. As written, I would mark it CONDITIONAL.","tokens_in":16599,"tokens_out":3818,"duration_ms":34139,"concrete_test":"Run the organizers' likelihood-based evaluation script on the held-out test sets for zero-shot ALLaM-7B-Instruct, ALLaM-7B-Instruct-preview, Fanar-1-9B-Instruct, and Qwen2.5-7B-Instruct, using the same prompt template as Section 3.2.1. If ALLaM-7B-Instruct scores around or above 82% zero-shot on Subtask 2, then the reported 84.22% fine-tuned result is mostly a base-model advantage; if it scores near 70%, the fine-tuning attribution is supported. This single matched comparison would settle whether the Section 5.1 claim holds.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"Section 5.1 concludes that \"task-specific fine-tuning significantly improves performance over baseline models\" from Tables 4 and 5. The tables do not support this as a general claim. In Subtask 2, the baseline is NileChat-3B zero-shot (75.12%), while the top three systems fine-tune ALLaM-7B variants; no zero-shot ALLaM-7B score is reported. Two fine-tuned Qwen2.5-7B/3B submissions score below the NileChat-3B baseline (74.13% and 70.83%). In Subtask 1, only the first- and second-place systems use the same base model as the baseline; MarsadLab ties the baseline and Hamyaria and Star finish below it. Hence the observed top-system improvements are at least partially explained by choosing stronger base models (e.g., ALLaM, Fanar) rather than by task-specific fine-tuning per se. The claim requires same-base-model comparisons with and without fine-tuning; without those, the central comparative finding is underdetermined.","agreement_with_reader":"disagree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper presents PalmX 2025, which the authors describe as the first shared task for benchmarking LLM cultural competence in Arabic and Islamic domains. The task consists of two MSA multiple-choice subtasks, one on general Arabic culture and one on general Islamic culture, with public training and development splits and a held-out test split. The paper describes the data collection pipeline, which combines Palm-derived items, web-crawled content, and LLM-assisted generation, followed by two-linguist human review; it also presents the likelihood-based evaluation protocol and the results of the nine and six valid submissions to the two subtasks. The central empirical claims are that task-specific fine-tuning substantially improves over zero-shot baselines, that LoRA-style parameter-efficient fine-tuning is the dominant and most effective approach, and that data augmentation helps in the Islamic domain but not in the general cultural domain. The authors report top accuracies of 72.15% for Subtask 1 and 84.22% for Subtask 2, and they release the dataset and evaluation code publicly.","tokens_in":16940,"tokens_out":8742,"duration_ms":79898,"significance":"If the benchmark data are of high quality, this is a useful community resource: it is the first standardized Arabic and Islamic cultural MCQ benchmark that I am aware of, it uses a transparent likelihood-based evaluation harness, it reports results from 15 independently developed systems, and it makes the data and evaluation code publicly available. The human-review design, the private test split, and the explicit accuracy metric are sensible choices, and the participation statistics indicate genuine community interest. The paper's strongest contribution is the resource itself and the documented baseline results. The more general methodological conclusion about task-specific fine-tuning is weaker than the wording suggests, because the baseline is a single small model and most submissions use different base models; the controlled comparisons requested below are needed before that conclusion can be treated as established.","major_comments":[{"comment":"The claim that 'task-specific fine-tuning significantly improves performance over baseline models' is underdetermined by the reported comparisons. The only baseline is NileChat-3B in zero-shot mode, while most submissions use different base models. In Table 5, the top two Islamic-subtask systems fine-tune ALLaM-7B-Instruct and no zero-shot ALLaM-7B score is reported, so the 9.1% improvement over the baseline could be attributable to base-model choice; moreover, two fine-tuned Qwen2.5 submissions (74.13% and 70.83%) fall below the NileChat-3B baseline. In Table 4, only the first- and second-place systems share the baseline base model, while MarsadLabM ties the baseline and Hamyaria and Star finish below it. Please add same-base zero-shot comparisons for ALLaM-7B and Fanar-9B, or explicitly restrict the conclusion to NileChat-3B and the first two Subtask-1 systems.","section":"Section 5.1, Tables 4-5"},{"comment":"The integrity of the evaluation is put in doubt by two entries. First, CultranAI's dataset column in Table 4 lists 'PalmX (test)' among the training sources, which directly contradicts the statement in Section 3.1 that the test set was private and held out; if this is literal, the submission should be disqualified or re-evaluated without the test split. Second, the ISL row describes a 'Retrieval-augmented (Gemini)' approach, and Section 5.4 praises their 'external knowledge retrieval,' which appears to conflict with the rule prohibiting systems with RAG or live internet access; please clarify whether retrieval was used only to construct training data and not at inference time. These points must be resolved because they bear directly on the validity of the leaderboard.","section":"Section 3.1 / Table 4"},{"comment":"No uncertainty quantification or significance testing is provided for any of the accuracy differences. The top four Subtask-1 systems fall within 1.0 percentage point (72.15, 71.65, 71.45, 71.35) on a 2,000-item test, so the ordering may not be stable; for the 1,000-item Islamic test the gap between the top two is also small (84.22 vs. 83.82). The word 'significantly' in the Abstract and Section 5.1 should be supported by exact binomial confidence intervals or a pairwise significance test, or replaced by a weaker formulation.","section":"Sections 4.3 and 5.1"},{"comment":"Because the benchmark's value depends on correct and unambiguous ground-truth answers, the human-review process should be documented quantitatively. The manuscript states that two linguists independently reviewed all questions and that discrepancies were consolidated, but it reports no inter-annotator agreement rate and no post-hoc audit of the final test set; the Limitations section acknowledges that 'small sections' of the dataset were LLM-generated or reformulated and could contain subtle artifacts. Please report the number and percentage of LLM-generated items in each split, the agreement between annotators, and the outcome of discrepancy resolution, so that readers can assess the risk of systematic answer-key errors.","section":"Section 2.1.1 and Limitations"}],"minor_comments":[{"comment":"The statement that parameter-efficient fine-tuning was 'the predominant and most effective approach' is only partially supported: LoRA was predominant, but the overall winner of Subtask 1 used full fine-tuning, so 'most effective' should be softened or justified with a same-base comparison.","section":"Section 5.2"},{"comment":"The sentence 'Test data was shared only after the leaderboard announcement' is confusing, since evaluation requires the organizers to use the test set before announcing the leaderboard; please clarify whether this refers to the public release of the test set after the announcement.","section":"Section 3.1, footnote 9"},{"comment":"The MarsadLabM row is ranked 7th with the same accuracy (67.55%) as the baseline row; an explicit tie-breaking rule would help readers interpret the rankings.","section":"Table 4"}],"recommendation":"major_revision","confidential_remarks":"The paper is a useful shared-task report, but the two integrity ambiguities in Table 4 are serious. If the 'PalmX (test)' entry for CultranAI is not a typo, the paper likely needs rejection or substantial re-evaluation rather than a routine revision; I recommend asking the authors for clarification before making a final decision."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The bottom line: this paper gives the field a genuinely useful resource—the first shared task benchmark for Arabic and Islamic cultural competence, with private test splits, public data, evaluation code, and system papers. If you work on Arabic NLP or cultural alignment, the dataset alone is worth knowing about. But the headline finding, that task-specific fine-tuning substantially beats zero-shot baselines, is only directly supported for one base model in one subtask. The stress-test note is right: the comparison is confounded.\n\nHere's the detail. In Subtask 1, there is a clean same-base comparison: NileChat-3B zero-shot gets 67.55%, and two teams fine-tuning NileChat-3B get 72.15% and 71.65%—about a 4–5 point gain on the same backbone. Beyond that, the table mixes Fanar-9B, Qwen2.5, and Qwen3, with no zero-shot scores for those base models, so their improvements cannot be attributed to fine-tuning per se. In Subtask 2, the top three systems all use ALLaM-7B-Instruct, while the baseline is NileChat-3B; no zero-shot ALLaM number is reported. Given that ALLaM is a 7B Arabic-centric model, part of the 9-point gap likely reflects a stronger base model, not fine-tuning. Two fine-tuned Qwen2.5 submissions actually fall below the NileChat baseline. So the abstract's \"task-specific fine-tuning substantially boosts performance\" is an overgeneralization as written.\n\nWhat the paper does well: the task design is clean—held-out test, likelihood-based scoring with a shared script, reproducible evaluation, clear participation rules. The authors are honest about limitations: country imbalance, MSA-only coverage, LLM-generated sections that could carry artifacts, and automated topic classification with imperfect accuracy. That is exactly the self-knowledge a shared task report should have. The related-work appendix situates PalmX properly against Palm, AraDiCE, ArabCulture, SaudiCulture, and others.\n\nSoft spots beyond the confound: no error bars or significance tests, no random-choice or human-accuracy baselines, and the ground-truth quality for LLM-generated MCQs rests on a human review process that is described but not audited. These are fixable in a revision, not fatal flaws.\n\nI would send this to peer review. It is a valuable resource and a competent report; the main claim just needs qualification: fine-tuning improves over zero-shot when the base model is held fixed, while gains across different base models are confounded. That is a significant but clean revision. Worth citing if you work in this space.","headline":"Useful first benchmark for Arabic and Islamic cultural QA, but the fine-tuning claim is overstated because baseline comparisons mix model sizes and families.","tokens_in":17342,"tokens_out":3025,"would_cite":true,"duration_ms":26644,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"PalmX 2025, the first shared task for benchmarking LLMs on Arabic and Islamic culture, finds that task-specific fine-tuning substantially improves cultural QA, with top systems reaching 72.15% on Arabic culture and 84.22% on Islamic…","keywords":["Arabic culture","Islamic culture","cultural competence","shared task","multiple-choice QA","fine-tuning","parameter-efficient fine-tuning","LLM benchmarking"],"falsifier":"Re-annotate a random sample of, say, 300 test questions per subtask with new professional linguists who have not seen the official key; if agreement with the official answers is materially below perfect, or if model rankings change when only independently confirmed items are scored, the benchmark's central claim is weakened. A cheaper version is to isolate the items the original two reviewers disagreed about before consolidation and compare scores on that subset.","tokens_in":16452,"feed_emoji":"🌙","tokens_out":8543,"duration_ms":71184,"temperature":0.7,"pith_summary":"This paper reports the first shared task designed to measure how well large language models know Arabic and Islamic culture. It contributes PalmX 2025, a benchmark of multiple-choice questions in Modern Standard Arabic covering traditions, food, history, religious practices, and language expressions across 22 Arab countries, with human-reviewed answers. Twenty-six teams registered for the general-culture subtask and nineteen for the Islamic subtask; nine and six valid systems were evaluated. The central finding is that task-specific fine-tuning clearly improves on zero-shot baselines, with parameter-efficient LoRA-style tuning being the most common and most effective approach. The best systems reached 72.15% accuracy on Arabic culture and 84.22% on Islamic knowledge, leaving a large gap that the authors use to motivate continued benchmarking.","feed_headline":"Fine-tuning beats baselines on first Arabic-culture benchmark","feed_subtitle":"Top systems reached 72.15% on Arabic culture and 84.22% on Islamic knowledge, leaving clear headroom.","key_machinery":"The central object is the PalmX 2025 benchmark itself: two sets of multiple-choice questions in Modern Standard Arabic, one on general Arabic culture and one on Islamic culture, each with four options and one human-reviewed correct answer, split into training, dev, and held-out test sets. The mechanism that turns the benchmark into a repeatable measurement is a likelihood-based scoring protocol: for each item, a model receives the question and the four labeled options, and its answer is the label with the highest normalized log-probability as a continuation of the prompt. Participation rules require submitting model weights rather than retrieval systems, cap models at 13 billion parameters, and keep the test set private until evaluation, so the reported accuracies are meant to reflect internalized knowledge plus whatever fine-tuning the teams applied.","core_discovery":"The paper's central claim is that task-specific fine-tuning substantially improves performance over baseline models on both subtasks, as stated in its analysis of results. Against a zero-shot NileChat-3B baseline of 67.55% on general Arabic culture and 75.12% on Islamic culture, the top teams reached 72.15% and 84.22%, respectively, gains of 4.6 and 9.1 points. The same result pattern also supports a methodological sub-claim: parameter-efficient fine-tuning, especially LoRA, is the dominant and most effective strategy, while data augmentation helps in the Islamic domain but did not help the same team on cultural questions. The authors interpret the higher Islamic scores as evidence that Islamic knowledge has more structured, canonical answers than the broader cultural domain.","pith_inferences":["Beyond the paper, a country-stratified accuracy score would test whether the overall numbers hide weak performance on under-represented countries such as Iraq and Algeria, an imbalance the authors acknowledge.","Beyond the paper, the label-only likelihood scorer may favor models with well-calibrated Arabic next-token distributions; a free-text generation variant is a direct robustness check.","Beyond the paper, if the gains replicate in future rounds, the PalmX training split could serve as reusable instruction data for grounding Arabic LLMs culturally, not just as an evaluation set.","Beyond the paper, the 'Islamic knowledge is more structured' explanation predicts that inter-annotator agreement and answer determinism should be measurably higher on Islamic questions; that is testable with the dataset's own review records."],"forward_implications":["Fine-tuning on PalmX's human-reviewed training questions is a reliable route to improving Arabic cultural QA, lifting the top culture score 4.6 points and the top Islamic score 9.1 points over the zero-shot baseline.","Smaller Arabic-centric models can win: the 3-billion-parameter NileChat-3B took first place in the culture subtask through fine-tuning, while larger 7B and 9B models finished close behind, suggesting scale is not the main lever.","Parameter-efficient LoRA fine-tuning is competitive with full fine-tuning, so substantial cultural gains do not require retraining all weights.","Data augmentation is domain-dependent: the AYA team found paraphrase augmentation decisive for Islamic questions but not helpful on the cultural development set.","Because even winning systems answer only 72.15% of cultural and 84.22% of Islamic questions correctly, stable headroom remains for future culturally grounded Arabic LLMs."],"supporting_citations":[{"why":"Supplies the zero-shot NileChat-3B baseline and serves as the base model for the first-place culture system; removing it removes the comparison for the fine-tuning claim.","marker":"Mekki et al., 2025"},{"why":"Provides the Palm instruction dataset from which 4,000 of the cultural questions are generated, forming the backbone of Subtask 1 data.","marker":"Alwajih et al., 2025a"},{"why":"Qwen-30B is the generator used to convert Palm training items into MCQ format.","marker":"Yang et al., 2025"},{"why":"The likelihood-based evaluation approach is adapted from this evaluation harness and defines how accuracy is computed.","marker":"Biderman et al., 2024"},{"why":"ALLaM-7B-Instruct is the base model for the winning Islamic subtask system and several top-two finishes.","marker":"Bari et al., 2024"},{"why":"Fanar-9B is the base model used by multiple top-performing culture subtask systems, including the third- and fourth-place teams.","marker":"Team et al., 2025"},{"why":"System report for the first-place culture submission whose full fine-tuning produced the 72.15% headline result.","marker":"Hossain and Afli, 2025"},{"why":"System report for the first-place Islamic submission whose LoRA fine-tuning plus augmentation produced the 84.22% result.","marker":"Tajrin et al., 2025"}],"fun_headline_variants":["First Arabic-Islamic culture benchmark: fine-tuning wins","PalmX 2025: Fine-tuning beats baselines on culture tasks","LoRA fine-tuning dominates first Arabic culture benchmark","Fine-tuning wins on first Arabic and Islamic culture benchmark","Arabic-Islamic culture benchmark: LoRA fine-tuning leads"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The rankings all rest on one assumption: that the human-reviewed answer key is correct and unambiguous, and the authors themselves flag that LLM-generated or reformulated questions could carry subtle artifacts, so systematic errors in the key would shift every reported accuracy.","fun_headline_variants_meta":{"raw":{"variants":["First Arabic-Islamic culture benchmark: fine-tuning wins","PalmX 2025: Fine-tuning beats baselines on culture tasks","LoRA fine-tuning dominates first Arabic culture benchmark","Fine-tuning wins on first Arabic and Islamic culture benchmark","Arabic-Islamic culture benchmark: LoRA fine-tuning leads"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00097,"raw_usage":{"total_tokens":4140,"prompt_tokens":978,"completion_tokens":3162,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":594,"completion_tokens_details":{"reasoning_tokens":3081}},"tokens_in":594,"tokens_out":3162,"duration_ms":22045,"temperature":1.0,"reasoning_tokens":3081,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T16:35:52.415653+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-annotate a random sample of, say, 300 test questions per subtask with new professional linguists who have not seen the official key; if agreement with the official answers is materially below perfect, or if model rankings change when only independently confirmed items are scored, the benchmark's central claim is weakened. A cheaper version is to isolate the items the original two reviewers disagreed about before consolidation and compare scores on that subset.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"System report for the first-place culture submission whose full fine-tuning produced the 72.15% headline result."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"System report for the first-place Islamic submission whose LoRA fine-tuning plus augmentation produced the 84.22% result."}],"review_version":2}