{"id":"e4509704-e5a1-45a5-927b-c2bcf0583f41","arxiv_id":"2506.01305","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"VM14K provides the first Vietnamese medical question-answering benchmark, with 14,000 multiple-choice questions across 34 specialties and four difficulty levels, along with an open pipeline for building similar benchmarks.","lead":"This paper introduces VM14K, a benchmark of 14,000 multiple-choice medical questions in Vietnamese, covering 34 specialties and four difficulty levels. It aims to let researchers test how well AI language models understand Vietnamese medical knowledge, filling a gap for Vietnamese-speaking users of medical AI.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 'expert-verified' label in the central claim is not substantiated by the verification description in Section 3.2.2; the protocol may allow un-reviewed items into the gold set.","rationale":"The reader's weakest assumption identifies the verification process as the primary risk, and I agree. Without a transparent audit trail, the benchmark cannot be considered a gold standard, which is the core of the contribution. The proposed test is a direct measurement of label quality. A conditional accept is appropriate because the issue is fixable by releasing annotation metadata and/or a re-audit, not a fundamental theoretical flaw. I therefore leave the verdict unchanged.","tokens_in":11101,"tokens_out":9500,"duration_ms":98126,"concrete_test":"First locate the actual public release (no URL appears in the text); then take the 10k public set, draw a random 500-item sample, and have two independent Vietnamese physicians, blinded to the gold answers and the paper's difficulty labels, answer the items. Compute exact agreement and Cohen's kappa between each physician and the gold label. If expert–gold agreement falls below 95% (or kappa below 0.9), the 'expert-verified' claim is not supported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim that VM14K is a 14,000-item, expert-verified benchmark requires that every gold answer and label be reliable. Section 3.2.2 describes a three-source verification (source data, LLM answers, medical experts), but omits the counts, qualifications, and adjudication rules for the experts; it also never states whether each of the 14,000 questions was actually human-reviewed. The sentence 'The first questions to be verified by human experts are the easiest ones, where the original source answers and the answers from foundation models differ' suggests that items with source–LLM agreement were not necessarily passed through a human. The Introduction separately calls the annotations 'by medical students' while the Abstract says 'expert-verified,' which further weakens the claim. If the final set contains a material fraction of un-reviewed items, then the benchmark's gold correctness—and therefore every model ranking and the difficulty/topic analysis in Section 4—rests on a mix of source and LLM judgments. This is an internal documentation failure of the paper's central argument, not an external disagreement.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"VM14K is a proposed benchmark for evaluating large language models on Vietnamese medical knowledge. The paper describes a three-stage construction pipeline: crawling and extracting Vietnamese medical exams, textbooks, and public medical records; using GPT-4o and Gemini 2.0 Flash to parse raw text into structured multiple-choice fields (question, options, correct answer, difficulty, medical topic); and a verification pipeline that compares source answers, LLM answers, and human annotation. The final resource contains 14,000 multiple-choice questions across 34 medical specialties and four difficulty levels, split into a 4k sample public set, a 10k full public set, and a 2k private leaderboard set. The paper reports pass@k and ensemble accuracies for several general-purpose and medical-specialized models, together with a cost-performance analysis and breakdowns by difficulty and medical topic.","tokens_in":11261,"tokens_out":4534,"duration_ms":50620,"significance":"If the verification claims are adequately substantiated, VM14K would fill a genuine gap: it is a native Vietnamese medical benchmark rather than a translated English one, it is substantially larger than most existing non-English medical QA resources, and its three-way split with a private test set supports controlled evaluation. The open-source release of the pipeline and the stated intention to support scalability to other languages are also strengths. However, the current manuscript does not provide enough evidence for the 'expert-verified' label that appears in the abstract, and the difficulty and topic labels used in the analyses are generated by LLMs without reported validation. The resource's value as a gold standard therefore remains to be demonstrated.","major_comments":[{"comment":"The verification process is not described in enough detail to support the central claim that all 14,000 questions are expert-verified. The text says that the first questions to be verified by human experts are the easiest ones where the original source answers and the foundation model answers differ, which implies that questions with source-LLM agreement may have bypassed human review. The manuscript does not state how many human annotators participated, whether they were medical students or licensed physicians, whether each question received at least one independent human review, how disagreements were adjudicated, or what the inter-annotator agreement was. These details are load-bearing because every model ranking in Section 4 assumes that the gold answers are correct. Please report the full verification protocol and per-item human review counts, or revise the 'expert-verified' characterization accordingly.","section":"Section 3.2.2"},{"comment":"There is an unresolved inconsistency in who created the labels. The Introduction says the benchmark was 'eventually annotated by medical students,' while the Abstract and Conclusion call it 'expert-verified' and 'expert-annotated.' Medical students under qualified supervision may be a perfectly acceptable annotation workforce, but the qualification and oversight process must be stated explicitly. Please unify the terminology and describe the expertise required of the annotators, because this directly affects the benchmark's credibility as a gold standard.","section":"Sections 1, 5"},{"comment":"The difficulty levels and medical topics are assigned by GPT-4o and Gemini 2.0 Flash, as described in Section 3.2.2 and Table 1, yet Sections 4.1 and 4.2 interpret model performance by these labels (Figure 3, Table 4). Without any validation of these automatic labels on a human-annotated subset, the difficulty and topic analyses should be explicitly framed as analyses of LLM-assigned labels rather than as intrinsic properties of the benchmark. Additionally, Table 1 defines the extracted field correctOption as 'The correct answer according to GPT-4o'; please clarify how this field relates to the final gold answer after the verification step.","section":"Section 3.2.2, Tables 2-4"}],"minor_comments":[{"comment":"The abstract contains a typo and grammatical issue in 'communities,therefore help ensuring the quality'; please revise the sentence.","section":"Abstract"},{"comment":"Model names are inconsistent: 'Geimini Flash' appears in Section 3.2.2 and 'Gemin 2.0 Flash' appears in Table 3; use the same spelling throughout.","section":"Section 3.2.2, Table 3"},{"comment":"Table 2 lists 'Deepseek-R1[?]' with a missing citation placeholder; the reference list has no DeepSeek entry.","section":"Table 2"},{"comment":"The metric description says pass@k with k=3 and k=1, but several models in Table 2 have no pass@3 entry; please clarify which models were evaluated with pass@3 and the reason for the omission.","section":"Section 4"},{"comment":"The cost-performance analysis does not state the pricing snapshot, the API used, or whether reasoning-token overhead for models like DeepSeek-R1 and o3-mini is included; please specify these details for comparability.","section":"Figure 4"},{"comment":"Table 5 lists 34 medical categories but the paper does not report per-category question counts; a distribution table would help readers assess the 'balanced representation' claim.","section":"Appendix A"}],"recommendation":"major_revision","confidential_remarks":"The main blocking issue is verification transparency: the paper's central claim rests on the 14,000 questions being expert-verified, but Section 3.2.2 leaves it ambiguous whether all or only a subset received human review. This is fixable with a supplementary protocol description and, ideally, release of annotation metadata. The 'first Vietnamese medical benchmark' claim also depends on a literature search that is not documented in detail, but I did not treat that as a blocking issue. If the authors provide the missing verification details, the paper could become suitable for publication."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this is the first Vietnamese medical QA benchmark, and that alone makes it worth a serious look. The dataset (14k MCQs, 34 specialties, four difficulty tiers) is new, built from Vietnamese exams and textbooks rather than translated English items, and they release a private test split for leaderboard evaluation. The pipeline is open-source and could transfer to other languages. Credit where due: the paper identifies a real gap and makes a concrete attempt to fill it.\n\nThe soft spots are the usual ones for resource papers, but they matter more here because the central claim is 'expert-verified.' Section 3.2.2 describes a three-source verification (source, LLM, human), but never says how many experts looked at each question, what their qualifications were, how disagreements were adjudicated, or whether every one of the 14,000 questions actually passed through human review. The line about prioritizing easy questions where source and LLM answers differ is worrying: it reads as though items with source–LLM agreement may not have been human-checked at all. If that's the case, the gold standard is partly LLM-confirmed, not expert-verified. Also, the abstract says 'medical experts' while the introduction says 'medical students' — not the same thing, and the paper should be consistent.\n\nThe difficulty and topic labels are assigned by GPT-4o and Gemini, and Section 4 then analyzes model performance by those very labels. That's a mild circularity, not a fatal one, but it should be acknowledged as a limitation. The evaluation tables report accuracy without error bars or significance tests; a few percentage points difference between models may not be real.\n\nThe citation pattern looks fine, and the related-work coverage of non-English medical benchmarks is reasonable. No suspicious references.\n\nBottom line: a useful new resource with a load-bearing documentation gap. The right fix is not to throw the dataset away but to make the verification process transparent and public. I'd send this to review with a request for major revision — the potential value for Vietnamese medical NLP is real.","headline":"A genuinely new Vietnamese medical QA resource, but the 'expert-verified' claim needs numbers, not adjectives.","tokens_in":11823,"tokens_out":2323,"would_cite":true,"duration_ms":24352,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper introduces VM14K, the first Vietnamese medical benchmark: 14,000 expert-verified multiple-choice questions across 34 specialties and four difficulty levels, built by a scalable, open-source pipeline.","keywords":["Vietnamese medical benchmark","large language models","medical question answering","multilingual evaluation","expert annotation","difficulty levels","benchmark construction","data curation pipeline"],"falsifier":"Run an independent, double-blind re-annotation of a random sample of VM14K questions by Vietnamese medical specialists and check agreement with the released answers and difficulty labels; if agreement is at or below the noise level of the original three-source verification, the benchmark cannot serve as ground truth for leaderboard comparisons.","tokens_in":10890,"feed_emoji":"🩺","tokens_out":14044,"duration_ms":124666,"temperature":0.7,"pith_summary":"The paper aims to show that a trustworthy medical benchmark for Vietnamese can be built from locally verified educational sources, rather than by translating English tests. It introduces VM14K, 14,000 expert-verified multiple-choice questions spanning 34 medical specialties and four difficulty levels, and argues this is the first benchmark of its kind for Vietnamese. The authors also demonstrate the benchmark's use by evaluating a range of general and medical-specialized language models, finding that the best performer, DeepSeek R1, still leaves clear room for improvement and that model rankings on VM14K differ from rankings on English medical benchmarks. This matters because, without native-language benchmarks, medical AI for Vietnamese speakers cannot be measured fairly, and the open-source pipeline offers a route to build similar benchmarks for other underserved languages.","feed_headline":"First Vietnamese medical AI benchmark: 14,000 expert-checked questions","feed_subtitle":"Native questions measure clinical AI more fairly than translated English tests; the pipeline is open source.","key_machinery":"The load-bearing mechanism is the data-construction pipeline and the benchmark schema it feeds. A Python crawler pulls raw material from Vietnamese medical textbooks, exams, online quizzes, and clinical records, with a PostgreSQL hash table tracking duplicates; Spark transforms the text to JSON; PDF and Word extraction tools handle documents; then GPT-4o and Gemini 2.0 Flash convert raw question text into a structured field set and assign each item a difficulty level and one or more of 34 medical topics. Deduplication layers hashing with Levenshtein-distance clustering to absorb accent variants and minor character differences. Verification brings together three sources of truth—the original source answer, foundation-model answers, and medical-expert answers—and the annotation queue is ordered by mutual disagreement and difficulty, so the easiest disputed items are checked first. The four-tier difficulty rubric and the 34-specialty taxonomy are the scaffold that makes scores interpretable by breadth and depth, and the whole pipeline is released to be reusable in other languages.","core_discovery":"The central claim is that VM14K is the first Vietnamese medical question benchmark and that it offers a valid, reusable measure of medical knowledge for language models serving Vietnamese speakers. To establish this, the authors collected about 100,000 raw questions from Vietnamese medical textbooks, exams, online quizzes, and clinical records; used GPT-4o and Gemini 2.0 Flash to extract structured questions and assign difficulty and topic labels; removed duplicates with hashing and Levenshtein-distance clustering; and had medical experts verify answers, prioritizing questions where the source answer and the model answer disagreed. From the verified pool they sampled 14,000 questions balanced across 34 specialties and four difficulty levels, released as a 4,000-question sample set, a 10,000-question public set, and a 2,000-question private leaderboard set. Evaluations of general and medical-specialized language models show DeepSeek R1 with the highest accuracy, and the authors read the divergence from English-benchmark rankings as evidence that VM14K measures Vietnamese-specific medical language and knowledge, not just recycled English material.","pith_inferences":["Editorial inference: the ranking inversion against English benchmarks implies that a hospital or product team selecting a model for Vietnamese patients using English medical scores could pick a substantially weaker model; VM14K should be part of procurement decisions.","Editorial inference: publishing inter-annotator agreement and per-question confidence would turn VM14K from a one-shot resource into a reusable measurement instrument; without those numbers, difficulty-label effects could partly reflect annotation noise.","Editorial inference: the 100,000 raw questions that were not expert-verified are a natural training resource; using them for fine-tuning while holding the 14,000 verified questions for evaluation would be a direct test of the pipeline's value beyond benchmarking.","Editorial inference: the private 2,000-question set is the only guard against benchmark contamination; its integrity depends on the authors never releasing it, and the community should treat any public appearance of those questions as a contamination event."],"forward_implications":["Vietnamese-language medical AI can be evaluated directly on native clinical language rather than through translated English tests, avoiding the terminology and cultural mismatches that translation introduces.","Model developers get a concrete target from the public and private splits: the best current model scores about 78% pass@1 on VM14K, so a model approaching 90% or above would be a clear advance for Vietnamese medical applications.","The ensemble results imply that a single pass understates model capability, so future evaluations should consider shuffled-choice voting as a standard protocol.","The difficulty levels let buyers and researchers see where a model breaks: strong recall on Easy items with a sharp drop on Challenging and Hard items identifies a model that cannot perform clinical reasoning in Vietnamese."],"supporting_citations":[{"why":"Supplies evidence that translated or cross-lingual medical evaluation loses clinical nuance, motivating a native-language benchmark.","marker":"[15]"},{"why":"Provides MedQA, the English medical exam benchmark whose model rankings are contrasted with VM14K results.","marker":"[13]"},{"why":"GPT-4o is one of the two LLMs used to extract structured questions and to assign difficulty and topic labels.","marker":"[21]"},{"why":"Gemini 2.0 Flash is the other extraction and labeling model and is also one of the evaluated models.","marker":"[9]"},{"why":"Documents the Chinese CMBExam benchmark and its lack of a standardized difficulty schema, a gap VM14K addresses.","marker":"[19]"},{"why":"Documents the Korean KorMedMCQA benchmark, another non-English medical benchmark without a standardized difficulty structure.","marker":"[17]"},{"why":"MedMCQA is a large multi-subject medical MCQ dataset from Indian exams, used as a structural and evaluative comparison point.","marker":"[23]"},{"why":"Meditron-70B is one of the medical-specialized open-source models evaluated to test whether specialized training transfers to Vietnamese medical QA.","marker":"[7]"},{"why":"UltraMedical models are medical-specialized baselines that show the effect of scale and specialization on Vietnamese medical question answering.","marker":"[36]"}],"fun_headline_variants":["VM14K: Vietnam's first 14k-question medical AI benchmark","First Vietnamese medical benchmark: 14k expert-verified across 34 fields","14k Vietnamese medical questions, four difficulty levels, open-source","Open-source pipeline creates Vietnam's first medical AI benchmark","Vietnam's first medical AI benchmark: 14k native questions, open-source"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The benchmark's validity as a gold standard rests on the expert verification step being consistent and correct, but the paper does not report the number of experts, their agreement rates, or how disagreements with source and model answers were resolved.","fun_headline_variants_meta":{"raw":{"variants":["VM14K: Vietnam's first 14k-question medical AI benchmark","First Vietnamese medical benchmark: 14k expert-verified across 34 fields","14k Vietnamese medical questions, four difficulty levels, open-source","Open-source pipeline creates Vietnam's first medical AI benchmark","Vietnam's first medical AI benchmark: 14k native questions, open-source"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001246,"raw_usage":{"total_tokens":5135,"prompt_tokens":993,"completion_tokens":4142,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":609,"completion_tokens_details":{"reasoning_tokens":4048}},"tokens_in":609,"tokens_out":4142,"duration_ms":33793,"temperature":1.0,"reasoning_tokens":4048,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T11:43:50.379985+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run an independent, double-blind re-annotation of a random sample of VM14K questions by Vietnamese medical specialists and check agreement with the released answers and difficulty labels; if agreement is at or below the noise level of the original three-source verification, the benchmark cannot serve as ground truth for leaderboard comparisons.","supporting_citations":[{"cited_title":"Better to ask in english: Cross-lingual evaluation of large language models for healthcare queries, 2023","cited_arxiv_id":null,"evidence_quote":"Supplies evidence that translated or cross-lingual medical evaluation loses clinical nuance, motivating a native-language benchmark."},{"cited_title":"Introducing gemini 2.0: our new ai model for the agentic era, 2024","cited_arxiv_id":null,"evidence_quote":"Gemini 2.0 Flash is the other extraction and labeling model and is also one of the evaluated models."},{"cited_title":"Benchmarking large language models on cmexam – a comprehensive chinese medical exam dataset, 2023","cited_arxiv_id":null,"evidence_quote":"Documents the Chinese CMBExam benchmark and its lack of a standardized difficulty schema, a gap VM14K addresses."},{"cited_title":"Kormedmcqa: Multi-choice question answering benchmark for korean healthcare professional licensing examinations, 2024","cited_arxiv_id":null,"evidence_quote":"Documents the Korean KorMedMCQA benchmark, another non-English medical benchmark without a standardized difficulty structure."},{"cited_title":"Meditron- 70b: Scaling medical pretraining for large language models, 2023","cited_arxiv_id":null,"evidence_quote":"Meditron-70B is one of the medical-specialized open-source models evaluated to test whether specialized training transfers to Vietnamese medical QA."}],"review_version":1}