{"id":"ce21c2fe-b2f0-49e5-b056-4c798858e17e","arxiv_id":"2608.03095","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"The VIVID benchmark shows that current LLMs, including GPT-4o, interpret Vietnamese idioms and proverbs at less than half of the maximum score, exposing a cultural competence gap.","lead":"VIVID is a new benchmark of 1,636 Vietnamese idioms and proverbs, annotated for linguistic and cultural difficulty, that tests whether large language models truly understand figurative Vietnamese. Across eight models, even the best scores below half of the maximum, and Vietnamese-specialized models show no advantage over general multilingual models.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Human-judge agreement study is internally inconsistent: 150+150−50 = 250 distinct annotations cannot fit the stated 200-item sample, leaving the κ=0.792 that anchors the generative scorecard without a clear evidential basis.","rationale":"The reader identified the LLM-as-a-judge reliability assumption as the weakest point. I agree that this is load-bearing: the generative scorecard and the 'less than 50% correctness' headline depend entirely on the GPT-4.1 judge, and the validation evidence is thin. My stress-test sharpens this into a concrete, checkable defect: the annotation counts in Section 4.1.1 are arithmetically inconsistent, and inter-human agreement is asserted but never quantified. If the human-validation sample is actually smaller or differently structured than stated, κ=0.792 cannot be trusted as a calibration of the judge, and all generative scores become uncertain. I do not think this warrants rejection: the discriminative results in Table 5 (e.g., GPT-4o 0.524 vs. Vistral-7B 0.000 on topic classification) show the same qualitative gap without relying on the judge, and the dataset itself is a useful contribution. The appropriate disposition is CONDITIONAL: the authors should publish the human-annotation data, resolve the sample-size contradiction, report inter-human kappa, and ideally add an open-source judge sensitivity check before the generative claims are accepted at face value. My partial agreement with the reader reflects that their stated weakest assumption (judge bias against concise explanations or same-family favoritism) is plausible but secondary; the first-order problem is that the validation study's own sample arithmetic is not coherent, which must be fixed before any bias question can be meaningfully assessed.","tokens_in":18578,"tokens_out":3658,"duration_ms":36358,"concrete_test":"Inspect the released repository or request the human-annotation records for the 200-item subset. Count unique idiom IDs annotated by each person and the intersection; verify whether the union is 200, 250, or something else. Recompute Cohen's kappa between the two annotators on the overlapping items only, then recompute GPT-4.1-versus-human agreement for each of the four prompting strategies on the correctly defined sample. If the overlap is only 50 items or the 200-item set cannot be reconstructed from the records, the reliability claim in Section 4.1.1 is unsupported and Table 1 must be regenerated on a properly sized, adjudicated sample. As a complementary check, rerun the judge on a 100-item subset with an open-source judge (e.g., Qwen-3-14B) to see whether the model rankings and the 'below 50%' headline survive a change of judge family.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"Section 4.1.1 validates the GPT-4.1 judge on 200 randomly sampled idioms, then states: 'Two native Vietnamese speakers independently annotate 150 idiomatic expressions each, with a 50-sample overlap.' If each annotator labeled 150 items and they shared 50, the union is 250 items, not 200. If instead each annotated 150 of the 200 items, the overlap must be 100, not 50. The reported counts cannot both be true. This matters because every generative score in Tables 3 and 4, and the abstract's 'less than 50% correctness' headline, is produced by this judge. The paper also never reports the inter-human kappa, saying only that reliability was 'strong,' so there is no verified human baseline against which κ=0.792 was computed. Section 7 concedes the 200-item subset 'may not fully capture the complete diversity,' but the more immediate problem is that the size and composition of the validation set are internally contradictory. If the human labels were actually collected on a smaller or differently overlapping subset, the agreement estimate could be far noisier than reported, and the judge's scores—and therefore the model rankings—could shift. The discriminative tasks (Table 5) independently support the existence of a large model gap, so the central claim is probably not false; however, the generative evaluation as published does not currently have a defensible reliability certificate.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"VIVID is a benchmark of 1,636 Vietnamese idioms and proverbs annotated with five complexity traits and seven semantic themes. The authors evaluate eight language models on two tasks: a generative explanation task scored by GPT-4.1 as a judge (with four prompting strategies compared against human ratings on a subset), and two discriminative classification tasks (topic and linguistic-characteristic identification). They report that all models, including the best-performing GPT-4o (2.46/5), achieve less than half of the maximum score; that Vietnamese-specialized models do not outperform similarly sized multilingual models; and that few-shot prompting degrades GPT-4o due to stylistic overfitting. The dataset and code are released.","tokens_in":18852,"tokens_out":7171,"duration_ms":62913,"significance":"If the reliability concerns are addressed, VIVID would fill a real gap: no existing benchmark targets Vietnamese figurative language, and the paper's dual-layer annotation (complexity traits plus semantic themes) is a useful resource for the community. The discriminative results (Table 5) independently show a large performance gap and thus provide partial support for the paper's broad conclusion. The systematic comparison of four LLM-as-a-judge prompting strategies (Table 1) is also a useful methodological contribution. However, the generative scorecard, which drives the headline 'below 50%' claim, rests on a judge whose human-agreement study is under-specified and internally inconsistent; the reliability certificate for the central claim therefore needs significant repair.","major_comments":[{"comment":"The reported human-annotation design is internally inconsistent. The text states that 200 idioms are randomly sampled for the reliability experiment, then says 'Two native Vietnamese speakers independently annotate 150 idiomatic expressions each, with a 50-sample overlap.' With 150+150−50 = 250 distinct annotations, this cannot be a 200-item sample; if each annotator covered 150 of the 200 items, the overlap would be 100, not 50. The paper also never reports inter-human agreement, saying only that reliability was 'strong.' Because the κ = 0.792 and the human baseline mean in Table 1 are the sole evidence that the GPT-4.1 judge is aligned with human judgment, the actual sample sizes, the overlap design, and the inter-human kappa must be reported precisely.","section":"§4.1.1"},{"comment":"There is a same-generator bias risk that is not addressed. GPT-4.1 is used both to generate the initial complexity-trait labels in Stage 1 of data validation and to score all generative outputs, and the evaluated models include GPT-4o, a sibling model from the same family. The 200-item human-agreement study does not rule out a systematic preference for GPT-family prose style, which could inflate GPT-4o's scores relative to smaller open-source models. I am not treating this as a circularity claim, but as an empirical correctness risk. The authors should report judge-human agreement separately for each evaluated model (at least for GPT-4o versus one or two open models) on the 200-item subset, and re-score a random subset with an alternative judge (e.g., an open-weight LLM or a second human annotator) to show that the rankings in Tables 3 and 4 are stable.","section":"§4.1 and §3.3"},{"comment":"The headline 'less than 50% correctness' is not supported by the measurement protocol. The judge prompt (Figure 7) asks GPT-4.1 to rate 'overall similarity' between the model explanation and the gold human explanation on a 0–5 scale. A mean of 2.46 is 49.2% of the maximum score, but that does not imply that 49.2% of explanations are correct; the judge is not producing a binary correctness label. The abstract, introduction, and Section 5.1 should say 'less than 50% of the maximum judge-assigned similarity score' or otherwise avoid the term 'correctness' for a score that measures similarity to a reference explanation.","section":"Abstract and §5.1"},{"comment":"The paper needs to separate the full human validation of all 1,636 annotations (Section 3.3) from the 200-item judge-reliability subset (Section 4.1.1). Section 3.3 says two native speakers 'independently reviewed all annotations,' while Section 7 limits the human–LLM agreement analysis to 200 samples and concedes that this subset 'may not fully capture the complete diversity.' These statements are not necessarily contradictory, but the manuscript should state explicitly whether the 200-item reliability subset is a subsample of the fully validated set and whether the same human labels were used in both places. Currently the reader cannot tell whether the judge was validated against the same gold-standard labels that constitute the benchmark, which is essential for interpreting κ = 0.792.","section":"§3.3 and §7"}],"minor_comments":[{"comment":"The parameter count for Llama-4-Scout is inconsistent: Table 2 lists 'Llama-4-Scout (109B)', while Tables 3 and 4 and Section 5.1 refer to 'Llama-4-Scout-17B-16E'. The model should be described consistently with its actual configuration (17B active parameters, 16 experts).","section":"Table 2 vs Tables 3–4"},{"comment":"Several references are duplicated under different citation tags: Le et al. 2022a/2022b (VIMQA), Ho et al. 2019a/2019b, Liu et al. 2022a/2022b, Tang et al. 2024a/2024b, and Wang et al. 2025a/2025b are the same papers in each pair. Also, Section 2.2 cites both ViGLUE and VLUE as '(Tran et al., 2024)' although the reference list contains distinct entries for Tran et al. (ViGLUE) and Do et al. (VLUE).","section":"References"},{"comment":"The 'Human Baseline 3.47' row in Table 1 would be easier to interpret if the paper stated whether this is the mean of the two annotators' scores on the sampled items, how disagreements were resolved in computing this mean, and whether the human scores are the same gold explanations used in the full benchmark.","section":"§4.1.1"},{"comment":"The claim that few-shot prompting degrades GPT-4o is supported by aggregate means (2.46 zero-shot vs. 2.04 few-shot in Table 3) but no variance estimate or significance test is provided. Given that the two illustrative examples are anecdotal, a paired statistical test across the 1,636 items would substantially strengthen this secondary claim.","section":"§5.1.2"},{"comment":"The annotator qualification is described only as 'two native Vietnamese speakers with expertise in Vietnamese linguistics.' Given that the benchmark's validity depends on their judgments, the paper should briefly report their linguistic training or experience, and should describe the adjudication procedure for disagreements beyond 'discussion' (e.g., whether a third expert was consulted).","section":"§3.3"},{"comment":"The manuscript contains numerous ligature and spelling artifacts (e.g., 'diﬀicult', 'oﬀicial', 'suﬀicient') and inconsistent capitalization of model names (e.g., 'Vinallama' vs. 'VinaLLaMA'). These should be fixed in the camera-ready version.","section":"Throughout"}],"recommendation":"major_revision","confidential_remarks":"The paper addresses a relevant gap and the broad conclusion is likely correct, but the generative evaluation's reliability certificate is currently not defensible due to the inconsistent annotation counts in §4.1.1 and the failure to report inter-human agreement. The 'correctness' wording also overstates what a 0–5 similarity score means. These are fixable with additional experiments (per-model judge agreement, an alternative judge, and proper significance tests), so I recommend major revision rather than rejection. The reference list also needs deduplication before publication."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The paper delivers something genuinely useful: VIVID, the first systematic Vietnamese benchmark for figurative language, with 1,636 idiom-proverb pairs from two complementary dictionaries, dual-layer annotation (complexity traits and semantic themes), and a sensible two-task evaluation design. The empirical findings are interesting and probably right: Vietnamese-specialized models do not beat multilingual ones at the 14B scale, and GPT-4o's few-shot degradation via stylistic overfitting is a concrete, well-illustrated observation. The discriminative results in Table 5 independently support the central claim of a large model gap, which matters because the generative scorecard is the weaker link.\n\nThe soft spot is real and is exactly where the stress-test note lands. Section 4.1.1 says the reliability study uses 200 idioms, then reports two native speakers each annotating 150 items with a 50-item overlap. Those numbers cannot both be true: the union would be 250, not 200. The paper never reports the inter-human kappa, only that reliability was 'strong,' so the κ=0.792 against human judgment has no verified human baseline. Since every generative score in Tables 3 and 4 flows through the GPT-4.1 judge, this is not a cosmetic issue. That said, it is fixable: report the actual sample size, give the inter-human agreement number, and add a sensitivity check with an open-weight judge to address the same-family bias (GPT-4.1 annotating and judging, GPT-4o evaluated).\n\nOther concerns are minor. The abstract's 'less than 50% correctness' overstates what a 0–5 similarity score means; 'less than 50% of maximum score' is accurate. There is a model-naming inconsistency (Vistral-7B vs Viet-Mistral-7B) and a dangling reference to earlier Korean benchmarks (KorID vs KULTURE bench) in the related work. The gold explanations come from two dictionaries, which is a reasonable assumption for a benchmark, though the paper should acknowledge that some idioms admit more than one legitimate interpretation.\n\nBottom line: the resource is valuable, the central finding is plausible and supportable by the discriminative tasks, and the flaws are addressable. I would send this to peer review, not desk-reject it. The authors need to clean up the reliability analysis before publication, but this is the kind of benchmark the field can actually build on for Vietnamese and for low-resource figurative-language evaluation more broadly.","headline":"First systematic Vietnamese figurative-language benchmark with real empirical value, but the judge-reliability section has an internal inconsistency that needs fixing before the numbers are fully trusted.","tokens_in":19397,"tokens_out":1079,"would_cite":true,"duration_ms":11680,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"VIVID, a 1,636-item benchmark of Vietnamese idioms and proverbs, shows that current language models, including GPT-4o, explain them at less than half the maximum quality score.","keywords":["Vietnamese idioms","proverbs","figurative language","cultural competence","LLM-as-a-Judge","benchmark","few-shot prompting","GPT-4o"],"falsifier":"Select a random sample of several hundred VIVID explanations, have native Vietnamese speakers score them on the same 0–5 scale, and compare model rankings with GPT-4.1's rankings; if human rankings disagree on which models are better, or if human scores for GPT-4o's explanations rise above half the maximum, the central 50% ceiling claim would fail.","tokens_in":18376,"feed_emoji":"📚","tokens_out":9010,"duration_ms":72706,"temperature":0.7,"pith_summary":"VIVID is the first systematic benchmark for culturally grounded figurative language in Vietnamese: 1,636 idioms and proverbs, each paired with a dictionary-sourced explanation and labeled for five linguistic complexity traits and seven semantic themes. The paper uses it to argue that current language models lack cultural competence: on open-ended explanation generation, the best model tested, GPT-4o, averages 2.46 out of 5, below half of the maximum, while Vietnamese-specialized VinaLLaMA-7B manages only 0.13. A GPT-4.1 judge with aspect-based prompting, validated against human raters on 200 items (Cohen's kappa = 0.792), supplies the scores. The results also show that few-shot prompting is not universally helpful, degrading GPT-4o from 2.46 to 2.04 through stylistic overfitting. If correct, VIVID provides a measurable way to track whether future models are becoming culturally aware rather than merely fluent.","feed_headline":"Scores under half: GPT-4o tops Vietnamese idiom benchmark at 2.46","feed_subtitle":"The new VIVID benchmark shows top models explain Vietnamese idioms at less than half quality, and few-shot prompts can make GPT-4o worse.","key_machinery":"The load-bearing mechanism is the evaluation loop built around VIVID. The benchmark itself is a curated set of 1,636 idiom/proverb–explanation pairs from two Vietnamese dictionaries, annotated with five complexity traits (literal expressions, pragmatic nuances, Sino-Vietnamese terms, uncommon vocabulary, folk knowledge) and seven semantic themes. On the generative side, models explain a phrase in Vietnamese, and GPT-4.1 scores the explanation 0–5 against the gold explanation using aspect-based prompting — explicit criteria for core-meaning accuracy, nuance capture, clarity, and completeness, with the rule that a score of 0 on core meaning forces an overall 0; this judge was validated against human raters on 200 items (kappa = 0.792) and chosen after comparing four prompting strategies. On the discriminative side, 3-shot exact-match classification tasks for topic and for complexity traits measure whether models can categorize the same cultural knowledge. The combination turns cultural competence into a number that can be compared across models and prompting choices.","core_discovery":"At the paper's core is the claim that cultural grounding is a distinct capability that current LLMs have not acquired, and that it can be benchmarked. On each of 1,636 Vietnamese idioms and proverbs, models are asked to write a short Vietnamese explanation; the explanations are scored 0–5 by GPT-4.1 against the dictionary explanation, with the judge asked to consider core-meaning accuracy, nuance, clarity, and completeness. Every model falls short: GPT-4o reaches 2.46, Gemini Flash 2.5 reaches 2.29, the best open-source model Llama-4-Scout reaches 1.24 with few-shot prompting, and Vietnamese-specialized models score at or below 1.09. The paper also identifies four recurring failure modes — literal over-interpretation, lexical gaps on archaic words, cultural disconnection on folk-knowledge idioms, and pragmatic flattening of sarcasm or irony — and shows that GPT-4o's few-shot degradation is driven by imitating the moralizing tone of the examples instead of preserving meaning. VIVID is offered as the first systematic tool for exposing and measuring those failures in Vietnamese.","pith_inferences":["Because the gold explanations come from a single dictionary entry per idiom, VIVID scores similarity to one canonical reading; a model that gives a different but culturally valid interpretation would be marked down, so a human acceptability study across multiple paraphrases could show that the true gap is smaller than 50% for some idioms.","The judge validation on 200 entries leaves open whether GPT-4.1 ranks models the same way humans would on harder categories such as folk knowledge; a larger human sample or an open-source judge could change the ranking.","The same benchmark design could be transplanted to other low-resource languages with rich oral traditions, with the aspect-based judge prompt as a reusable template; the 50% ceiling may be a general property of figurative language rather than Vietnamese-specific.","The exact-match scoring on discriminative tasks likely understates models that understand the trait but format the answer differently (one model scored zero on both tasks); relaxed matching could separate format failures from comprehension failures."],"forward_implications":["If VIVID measures what it claims, then language-specific pretraining is not a shortcut: Vietnamese-specialized GreenMind-14B and multilingual Qwen-3-14B score identically (1.09), so scale and data breadth, not Vietnamese tuning, drive figurative understanding.","The less-than-50% ceiling means state-of-the-art models cannot yet be trusted to explain or translate Vietnamese idioms without human oversight.","Few-shot prompting should be reported alongside zero-shot, because in this setting it materially lowers GPT-4o's score (2.46 to 2.04) by inducing stylistic imitation.","The taxonomy of failure modes — literal over-interpretation, lexical gaps, cultural disconnection, and pragmatic flattening — gives developers concrete targets: a model improves on VIVID only if those four error types decrease.","Researchers can reuse VIVID's 1,636 pairs and the validated judge prompt as a repeatable testbed for culturally aware NLP in Vietnamese."],"supporting_citations":[{"why":"Supplies the primary dictionary explanations used as gold references for most of the 1,636 idiom/proverb pairs in VIVID.","marker":"Lân, 2010"},{"why":"Adds regional and variant idioms so the benchmark covers dialectal and alternative phrasings beyond the standard dictionary.","marker":"Dung et al., 2000"},{"why":"Provides the English FLUTE benchmark this work contrasts with, showing prior figurative-language resources have explanations but no Vietnamese coverage.","marker":"Chakrabarty et al., 2022"},{"why":"Provides the Korean idiom benchmark KorID, establishing that neighboring cultural benchmarks exist while Vietnamese was missing.","marker":"Wang et al., 2024"},{"why":"ProverbEval is the low-resource proverb benchmark the paper extends to Vietnamese, sharing the goal of cultural grounding.","marker":"Azime et al., 2025"},{"why":"Establishes the LLM-as-a-Judge paradigm that VIVID's GPT-4.1 scoring procedure builds on.","marker":"Zheng et al., 2023"},{"why":"Shows LLM judges align with human ratings for figurative idiom translation, and supplies the human-annotation protocol the reliability study follows.","marker":"S. Rezaeimanesh, 2025"},{"why":"Documents why surface metrics like BLEU fail on idioms and that GPT-based evaluation tracks human judgment for East Asian figurative language.","marker":"Tang et al., 2024a"},{"why":"Supports the paper's finding that few-shot prompting can push models toward stylistic imitation and away from figurative interpretation.","marker":"Phelps et al., 2024"},{"why":"Provides the rationale for open-ended explanation tasks instead of multiple choice, avoiding superficial pattern matching.","marker":"Griot et al., 2025"}],"fun_headline_variants":["Vietnamese idiom benchmark: top models score under half","GPT-4o leads but fails Vietnamese figurative language test","Cultural gap: Vietnamese models worse than multilingual on idioms","Few-shot prompts can degrade GPT-4o on idiom understanding","New benchmark exposes AI's literal streak on Vietnamese proverbs"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The whole scorecard depends on GPT-4.1's ratings matching what a native Vietnamese speaker would say about an explanation, but that match was checked on only 200 of the 1,636 entries.","fun_headline_variants_meta":{"raw":{"variants":["Vietnamese idiom benchmark: top models score under half","GPT-4o leads but fails Vietnamese figurative language test","Cultural gap: Vietnamese models worse than multilingual on idioms","Few-shot prompts can degrade GPT-4o on idiom understanding","New benchmark exposes AI's literal streak on Vietnamese proverbs"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001283,"raw_usage":{"total_tokens":5278,"prompt_tokens":1017,"completion_tokens":4261,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":633,"completion_tokens_details":{"reasoning_tokens":4182}},"tokens_in":633,"tokens_out":4261,"duration_ms":26730,"temperature":1.0,"reasoning_tokens":4182,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T14:52:22.696139+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Select a random sample of several hundred VIVID explanations, have native Vietnamese speakers score them on the same 0–5 scale, and compare model rankings with GPT-4.1's rankings; if human rankings disagree on which models are better, or if human scores for GPT-4o's explanations rise above half the maximum, the central 50% ceiling claim would fail.","supporting_citations":[],"review_version":2}