{"id":"c1acf773-d7c0-4162-bbe2-82cd4a076897","arxiv_id":"2508.03523","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"FilBench introduces a Filipino/Tagalog/Cebuano benchmark across cultural knowledge, classical NLP, reading comprehension, and generation, and shows GPT-4o tops out at 72.23%.","lead":"FilBench is a new benchmark for evaluating large language models on Filipino, Tagalog, and Cebuano. It tests 27 LLMs and finds that even the best, GPT-4o, scores only 72.23%, showing a large gap in Filipino language AI.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 72.23% headline score cannot be evaluated without the task curation and aggregation protocol; the benchmark's capability-gap claim rests on an unverifiable task set and weighting scheme.","rationale":"The paper's contribution is a new benchmark plus an empirical result about 27 LLMs. The load-bearing assumption is that the 72.23% score is a true, comparable measure of GPT-4o's Filipino capability. That requires two separate things to hold: (1) the task/label curation is valid, and (2) the evaluation is standardized across models. The reader's weakest_assumption identified the first condition; my concern additionally foregrounds the second, but both are facets of the same unverifiable methodology. Since only the abstract is available, neither condition can be checked. I deliberately do not claim the benchmark is flawed; rather, the evidence needed to assess the central claim is missing. Therefore the reader's UNVERDICTED verdict is exactly right, and no change is warranted. The concrete reproduction test would give the authors a clear path to resolving the uncertainty: a reproduced 72.23% would validate the score as an evaluation artifact, while annotation and task-weighting details would validate its interpretation as a Filipino-specific capability gap. Until those are released, the correct position is UNVERDICTED at low confidence.","tokens_in":722,"tokens_out":2342,"duration_ms":29431,"concrete_test":"Inspect the full paper or the associated repository and reproduce the headline result: run the released evaluation harness on GPT-4o with the exact prompts, few-shot settings, decoding parameters, and scoring metric, and confirm the reported 72.23%. As a stricter test, recompute GPT-4o's score under different aggregation rules (equal weight per task versus per instance, and per-language breakdown) to see whether the model ordering or the magnitude of the gap changes. In addition, check for annotation guidelines and inter-annotator agreement on the curated tasks; if neither the code nor the annotations are released, the headline score remains unverifiable and the correct verdict stays UNVERDICTED.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that FilBench is challenging because GPT-4o scores only 72.23%, but the abstract provides no evaluation protocol or curation details. For this score to support a capability gap, it must come from a comparable protocol across all 27 LLMs: identical prompts, decoding settings, scoring metrics, and task definitions. The abstract only says tasks were 'carefully curated' to reflect Philippine NLP priorities; it does not report annotation guidelines, inter-annotator agreement, error analysis, or exclusion criteria. If the benchmark over-represents hard categories like Classical NLP or low-resource language variants, or if the scoring is unusually strict or lenient, the 72.23% number is not comparable across models and does not demonstrate a Filipino-specific gap. Because the full text is unavailable, this concern cannot be resolved from the abstract alone. This is not an accusation of flawed curation, only a statement that the central number cannot be interpreted without the methodology behind it.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces FilBench, a Filipino-centric benchmark covering Filipino, Tagalog, and Cebuano across categories such as Cultural Knowledge, Classical NLP, Reading Comprehension, and Generation. The authors evaluate 27 state-of-the-art LLMs and report that GPT-4o achieves the highest score of 72.23%, while the best Southeast Asian language-specific model, SEA-LION v3 70B, reaches only 61.07%. The abstract concludes that FilBench is challenging and demonstrates the need for language-specific LLM benchmarks. The manuscript under review consists solely of this abstract; no full text was provided with the submission.","tokens_in":873,"tokens_out":2997,"duration_ms":37550,"significance":"If the benchmark is well-constructed and the evaluation protocol is sound, this work would make a valuable contribution to multilingual LLM evaluation and to Philippine NLP. The reported performance gap for a underrepresented language family is important, and the observation that region-specific models may underperform on Filipino is a potentially consequential finding. The paper would offer a concrete, reusable resource for tracking progress in Filipino language understanding and generation. However, the significance cannot currently be assessed because the abstract provides no information about task curation, annotation quality, evaluation settings, or scoring methodology. For these reasons, the work is potentially significant but presently unverified.","major_comments":[{"comment":"The central claim that 'FilBench is challenging, with the best model, GPT-4o, achieving only a score of 72.23%' is not interpretable without an explicit evaluation protocol. The abstract does not state whether all 27 LLMs were evaluated with identical prompts, decoding settings, few-shot examples, or scoring metrics. Without this information, the headline score cannot be compared across models, and it is impossible to determine whether the gap reflects Filipino-specific difficulty or artifacts of the evaluation setup.","section":"Abstract"},{"comment":"The claim that tasks were 'carefully curated' is not backed by any description of the curation criteria, data sources, annotation guidelines, inter-annotator agreement, or error-exclusion procedures. As the abstract stands, the benchmark's validity rests entirely on an unverifiable assertion. This is a load-bearing gap because noisy or unrepresentative tasks would undermine the capability rankings and the conclusion about LLM proficiency in Filipino.","section":"Abstract"},{"comment":"The abstract reports aggregated scores (e.g., 72.23% and 61.07%) without defining how these scores are computed. It is unclear whether they are averages over tasks, weighted by task difficulty, or computed with a specific metric (e.g., exact-match, F1, or human preference). Without the aggregation protocol, the numerical comparisons among models and the conclusion that 'several LLMs suffer from reading comprehension and translation capabilities' cannot be verified from the manuscript.","section":"Abstract"}],"minor_comments":[{"comment":"The abstract lists four task categories but does not define them or state how many tasks constitute each category; adding this information would help the reader assess coverage and balance.","section":"Abstract"},{"comment":"The relationship among Filipino, Tagalog, and Cebuano in the benchmark is unspecified. The abstract should clarify whether each task is in a single language, whether all languages appear across tasks, and how language coverage was determined.","section":"Abstract"},{"comment":"The phrase 'the value of curating language-specific LLM benchmarks' is a general claim that would be strengthened by stating what concrete insights or actionable recommendations follow from the FilBench results.","section":"Abstract"}],"recommendation":"uncertain","confidential_remarks":"The submission as provided is an abstract only. The central claims are plausibly useful, but they cannot be evaluated in the absence of the methodology. If the full paper containing task definitions, curation details, and evaluation protocol was intended for review, it is missing from the submission. If the abstract is being considered as a standalone publication, it is insufficient for the claims made. I recommend either obtaining the full manuscript or treating the submission as incomplete."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nQuick read of arXiv:2508.03523, abstract only, since the full text wasn't available. Here's my take.\n\nWhat's new: a Filipino/Tagalog/Cebuano benchmark spanning cultural knowledge, classical NLP, reading comprehension, and generation, evaluated on 27 LLMs. That fills a real gap—most multilingual benchmarks are English-centric or cover high-resource languages. The finding that SEA-specific models underperform, with SEA-LION v3 70B at 61% vs GPT-4o at 72%, is the kind of result that would matter to people building language-specific models. I don't see obvious circularity: the paper compares external LLMs on fixed tasks, so the benchmark is doing its job.\n\nWhat's soft: the abstract reveals no curation details, no annotation process, no evaluation protocol. The headline number—72.23% for GPT-4o—is impossible to interpret without knowing task weights, prompting, decoding settings, scoring metrics, and whether all models got the same protocol. The stress-test note gets this right: the capability-gap claim rests entirely on that number, and we can't audit it from the abstract. That's not an accusation; it's just a statement about what's verifiable. The paper's own description of \"carefully curate\" is a claim, not evidence.\n\nAlso, the novelty needs full-text confirmation. Many language-specific benchmarks exist; the abstract doesn't cite prior Filipino benchmarks, so the incremental advance is plausible but unproven.\n\nOverall: this is a serious-looking benchmark paper that deserves real refereeing. The central idea is sound, the results are interesting if the methodology holds up, and the failure mode—unverifiable curation—is exactly what peer review should catch. I'd send it out with a request for detailed methodology and dataset release. For me, I'd put it on the reading-group list once the full text is out, but I wouldn't cite it until I can see the annotation guidelines and aggregation details.\n\nRecommendation: accept for peer review; require the authors to show task construction, annotation quality, and a consistent evaluation protocol.","headline":"FilBench is a plausible, useful addition to low-resource benchmark suites, but the abstract alone can't support the headline capability-gap numbers; worth sending to referees with the full methodology.","tokens_in":1349,"tokens_out":1513,"would_cite":false,"duration_ms":15464,"reading_group":"maybe","serious_thinker":"unclear","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper introduces FilBench, a Filipino-centric benchmark, and claims that the best tested model, GPT-4o, reaches only 72.23% accuracy, exposing a substantial gap in LLM performance on Filipino, Tagalog, and Cebuano tasks.","keywords":["Filipino NLP","LLM benchmark","multilingual evaluation","Tagalog","Cebuano","reading comprehension","machine translation","cultural knowledge"],"falsifier":"An independent replication in which fluent Filipino speakers re-annotate a random sample of FilBench items and recompute the model scores would settle the claim: if the re-annotation finds high label disagreement or if human performance on the benchmark is no better than GPT-4o's 72.23%, the benchmark's difficulty and its implication of a model capability gap would be called into question.","tokens_in":562,"feed_emoji":"🇵🇭","tokens_out":2212,"duration_ms":28352,"temperature":0.7,"pith_summary":"The paper seeks to measure how well large language models understand and generate Filipino, a language that is largely absent from mainstream LLM evaluation. To do this, it introduces FilBench, a benchmark built from Filipino NLP research priorities and spanning cultural knowledge, classical NLP, reading comprehension, and generation. The central finding is that the best model, GPT-4o, scores only 72.23%, and that models trained specifically for Southeast Asian languages do even worse, with SEA-LION v3 70B at 61.07%. The paper argues that this gap demonstrates the need for language-specific benchmarks to drive progress in Filipino NLP and to include Philippine languages in model development.","feed_headline":"LLMs top out at 72% on new Filipino benchmark","feed_subtitle":"GPT-4o leads but still lags; region-specific models do worse, exposing a clear gap in Filipino, Tagalog, and Cebuano.","key_machinery":"The central object is FilBench itself, a curated set of Filipino-language tasks organized into four families: Cultural Knowledge, Classical NLP, Reading Comprehension, and Generation. It carries the argument by providing a single benchmark on which all 27 models are scored, allowing direct comparison and exposing task-specific weaknesses; the aggregate score is the evidence for the paper's claim that current LLMs are not yet proficient in Filipino.","core_discovery":"FilBench is a challenging benchmark for LLMs in Filipino, Tagalog, and Cebuano. Across 27 state-of-the-art models, the highest score is GPT-4o's 72.23%, and no model comes close to ceiling performance. Models specialized for Southeast Asian languages, such as SEA-LION v3 70B, underperform general-purpose models, reaching only 61.07%. The paper attributes this gap to weaknesses in reading comprehension and translation, and concludes that curated, language-specific evaluation is necessary to reveal and address these gaps.","pith_inferences":["A score of 72.23% on a carefully curated benchmark leaves substantial headroom, which suggests that fine-tuning on Filipino data, not just multilingual pretraining, may push models closer to ceiling performance.","Because the benchmark was designed to reflect Philippine NLP research trends, the reported rankings may also reflect the difficulty of tasks Filipinos actually care about, making them more informative than a generic multilingual test.","A natural next step would be to compare model scores with a human baseline on the same tasks; that would clarify whether the gap represents a real deficiency or merely hard questions.","The relative underperformance of SEA-LION models could be probed by testing whether it stems from training data coverage, tokenization, or evaluation protocol, which this paper does not isolate."],"forward_implications":["Current LLMs, including top general-purpose models, have a meaningful performance gap in Filipino-language tasks, which suggests that improvements are needed before reliable deployment in Filipino-speaking contexts.","The finding that SEA-LION v3 70B underperforms indicates that regionally specialized models do not automatically bring better performance in specific Philippine languages.","The benchmark provides a reusable evaluation suite for tracking progress in Filipino NLP, allowing future models to be measured against the same 72.23% ceiling.","Weaknesses in reading comprehension and translation point to concrete capability areas that model developers should target to improve Filipino language understanding."],"supporting_citations":[],"fun_headline_variants":["New Filipino benchmark caps LLMs at 72%","SEA-specific LLMs perform worse on Filipino tasks","Filipino benchmark reveals LLM language gap","Even GPT-4o struggles with Filipino benchmark","FilBench: LLMs fall short on Tagalog and Cebuano"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The entire conclusion rests on the assumption that FilBench's tasks, labels, and scoring accurately reflect real Filipino language use and NLP priorities; if the curation is flawed or the labels are noisy, the reported model scores and rankings would not reliably show how well LLMs handle Filipino.","fun_headline_variants_meta":{"raw":{"variants":["New Filipino benchmark caps LLMs at 72%","SEA-specific LLMs perform worse on Filipino tasks","Filipino benchmark reveals LLM language gap","Even GPT-4o struggles with Filipino benchmark","FilBench: LLMs fall short on Tagalog and Cebuano"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000206,"raw_usage":{"total_tokens":1358,"prompt_tokens":867,"completion_tokens":491,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":483,"completion_tokens_details":{"reasoning_tokens":413}},"tokens_in":483,"tokens_out":491,"duration_ms":5598,"temperature":1.0,"reasoning_tokens":413,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T04:21:53.087572+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"An independent replication in which fluent Filipino speakers re-annotate a random sample of FilBench items and recompute the model scores would settle the claim: if the re-annotation finds high label disagreement or if human performance on the benchmark is no better than GPT-4o's 72.23%, the benchmark's difficulty and its implication of a model capability gap would be called into question.","supporting_citations":[],"review_version":1}