{"id":"16667fca-20a8-4b1b-9c47-a83a0b27d473","arxiv_id":"2508.06851","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":6.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":0,"one_line_summary":"MDK12-Bench evaluates multimodal LLMs on 141K K-12 exam questions with a dynamic framework that introduces unfamiliar shifts to reduce contamination.","lead":"This paper introduces MDK12-Bench, a large benchmark of 141,000 real K-12 exam questions across six subjects for testing multimodal AI models. It also proposes a dynamic evaluation method that changes question formats and wording to reduce the chance that models succeed merely by memorizing training data.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The dynamic evaluation's core claim—that visual/textual/question-form shifts are 'unfamiliar'—is asserted, not established; contamination mitigation and the generalization findings stand or fall on this.","rationale":"The abstract is the only substantive material available; the manuscript body appears blank in the provided record. The paper's central claim is that MDK12-Bench provides a large-scale, contamination-resistant evaluation via dynamic unfamiliar shifts. Among the assumptions supporting that claim, the most load-bearing is that the visual/textual/question-form shifts are genuinely unfamiliar to the tested models. The reader's weakest_assumption identifies exactly this, and I agree; I would sharpen it by noting that the shift-generation process itself could introduce a second source of contamination. This is not a disagreement with consensus; it is a correctness risk internal to the method. The proposed check—a leakage audit combining reconstructability, n-gram overlap with public corpora, and a matched comparison against truly novel questions—would settle whether the unfamiliarity assumption holds. If the audit fails, the benchmark's core selling point is invalidated. If it passes, the central generalization findings have a much firmer basis. Since the abstract alone cannot establish this necessary condition, the appropriate disposition is conditional acceptance: release or perform the leakage audit before the contamination-resistance and generalization claims are accepted. This is a slight sharpening of the reader's UNVERDICTED, so I recommend CONDITIONAL rather than UNCHANGED.","tokens_in":732,"tokens_out":4760,"duration_ms":54317,"concrete_test":"Run a leakage audit on a random sample of 500 shifted MDK12 items: for each, query a strong open MLLM (not among the evaluated models) to reconstruct the original source exam question from the shifted text, and compute exact/paraphrastic n-gram overlap between the shifted text and public webpages dated before the benchmark's release. Independently, evaluate the same MLLMs on a matched set of 500 newly authored, never-published K-12 questions with the same knowledge points. If reconstruction succeeds for a substantial fraction of items, or if model accuracy on the shifted set is significantly higher than on the truly novel set, the 'unfamiliar' assumption is falsified and the contamination-mitigation claim collapses; if not, the concern is resolved.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"MDK12-Bench's central contribution is a contamination-resistant evaluation via 'unfamiliar visual, textual, and question form shifts.' This is load-bearing: the benchmark's objectivity and longevity, and all four reported evaluation dimensions, are interpreted through the lens of those shifts being genuinely novel to the tested models. The abstract provides no evidence for that unfamiliarity. If the shifts are shallow paraphrases or template substitutions, or if the original K-12 exams already appear in pretraining corpora, then models can answer shifted items by recognizing the source, and measured performance reflects memorization rather than robust generalization. The transformation process itself is also unexamined: if an LLM generates the shifts, that LLM's training data may contain the same exams or near-duplicates. No leakage tests, no construction details, and no comparison against truly novel holdout questions are described. The provided full text appears empty, so no additional validation is available. In short, the core empirical claim is not internally inconsistent, but it rests on an unverified assumption about data novelty that is directly testable.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces MDK12-Bench, a large-scale multimodal benchmark constructed from real-world K-12 exams across six disciplines, comprising 141K instances and 6,225 knowledge points organized in a six-layer taxonomy. It claims to evaluate MLLMs along four dimensions: difficulty levels, temporal shifts, contextual shifts, and knowledge-driven reasoning, and proposes a dynamic evaluation framework that introduces 'unfamiliar' visual, textual, and question-form shifts to mitigate data contamination. The paper also introduces knowledge-point reference-augmented generation (KP-RAG) and reports key findings about current MLLMs' limitations.","tokens_in":1030,"tokens_out":2867,"duration_ms":32429,"significance":"If the claims are substantiated, MDK12-Bench would be a valuable and timely contribution: it offers a large, real-exam, multidisciplinary benchmark, a dynamic evaluation scheme aimed at contamination resistance, and a mechanism (KP-RAG) for probing knowledge use. The four-dimensional evaluation framework addresses an important gap in static benchmark design. However, the provided manuscript consists only of the abstract; none of the quantitative claims, construction details, or experimental findings can be verified. The significance therefore remains conditional on a full methods and results presentation.","major_comments":[{"comment":"The manuscript as supplied contains only the abstract. There is no description of data collection, quality control, answer key verification, taxonomy construction, model evaluation protocols, or statistical analysis. All central claims—141K instances, 6,225 knowledge points, six-layer taxonomy, five question formats, and the reported findings—are unverifiable. This is a load-bearing omission: without the full methodology, the paper cannot be assessed as a benchmark contribution.","section":"Entire manuscript (provided text is abstract only)"},{"comment":"The central contamination-resistance claim rests on the assertion that the introduced visual, textual, and question-form shifts are 'unfamiliar' to the evaluated models. No evidence is provided: the transformation process is unspecified, the source and date of the underlying exams are not given, and no leakage checks are described. If shifts are template-based or LLM-generated from corpora that already contain these exams, performance on shifted items may reflect memorization, not generalization. The paper must provide concrete leakage analyses, e.g., n-gram overlap with training corpora, behavior on truly novel holdout items, and validation that human raters consider the shifts novel.","section":"Abstract, dynamic evaluation framework"},{"comment":"The 6,225 knowledge points and six-layer taxonomy are the backbone for the 'knowledge-driven reasoning' claims, but the abstract provides no information about how they were created, validated, or assigned to questions. There is no inter-annotator agreement, error audit, or evidence that the taxonomy is exhaustive or consistent across disciplines. Errors in knowledge-point assignment or answer keys would propagate to all conclusions about MLLMs' reasoning abilities. This must be addressed with a detailed validation protocol.","section":"Abstract, knowledge-point taxonomy and answer keys"},{"comment":"The abstract reports 'key findings' without giving the list of evaluated models, prompt formats, inference settings, sampling methods, or statistical comparisons. No confidence intervals, significance tests, or multiple-comparison corrections are mentioned. Since the paper aims to compare MLLMs across several dimensions, the absence of any statistical treatment makes the findings non-assessable. Full evaluation details and raw results are required.","section":"Abstract, evaluation methodology"}],"minor_comments":[{"comment":"The phrase 'multidisciplinary' and 'six disciplines' is repeated; please specify the disciplines explicitly (e.g., mathematics, physics, chemistry, biology, history, geography) to avoid ambiguity.","section":"Abstract"},{"comment":"The term 'unfamiliar visual, textual, and question form shifts' is vague. Please define what constitutes each shift type and provide representative examples in the full paper.","section":"Abstract"},{"comment":"Minor wording: 'while improving benchmark objectivity and longevity by mitigating data contamination' could be read as if contamination is the only threat to objectivity; consider rewording for accuracy.","section":"Abstract"}],"recommendation":"uncertain","confidential_remarks":"The provided full text is empty; this may be an artifact of how the paper was submitted to the reviewing system, but as a referee I can only assess what is in front of me. The abstract makes plausible, high-impact claims, yet none of the technical substance is available. My 'uncertain' recommendation reflects the inability to verify any load-bearing component. If a full manuscript exists, I would need it to evaluate the contamination-resistance argument, the taxonomy validity, and the experimental findings. I also note that the reader's strongest concern—the unverified novelty of the dynamic evaluation shifts—is exactly the point that must be settled empirically before the benchmark's longevity claim can be accepted."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"You asked about MDK12-Bench. Bottom line: this is a resource paper that could be genuinely useful, but what I have in front of me is an abstract and no full text—the submission appears empty. So the only honest verdict is unverified, not incorrect.\n\nWhat the paper does well on paper: the scale is real. 141K instances across six K-12 subjects with 6,225 knowledge points in a six-layer taxonomy is a solid step beyond small static benchmarks. The five question formats and the year/difficulty annotations let you slice results in ways that matter for AI tutoring. The dynamic evaluation idea—introducing unfamiliar visual, textual, and question-form shifts to fight contamination—is a sensible response to a well-known problem, and the KP-RAG experiments ask a useful question about whether retrieved knowledge actually changes behaviour. I would trust those design choices if the implementation checks out.\n\nNow the soft spot, and it is the one the stress test flagged: the whole framework rests on the claim that the shifts are genuinely unfamiliar to the tested models. That is asserted, not established. No leakage tests, no transformation protocol, no comparison against truly novel holdout questions. If the shifts are shallow paraphrases or templates, or if the base K-12 exams are already in pretraining corpora, the four evaluation dimensions become memorization studies rather than generalization studies. This is fixable—you can test for leakage directly—but it is load-bearing and currently unsupported. I also cannot tell you anything about label quality, sampling balance, or whether the six-layer taxonomy was validated by humans, because there are no methods to read.\n\nI agree with the reader's unverdictable take and I think the stress-test concern stands. The weakness is missing evidence, not internal contradiction. Who is this for? People building or evaluating MLLMs for education, and anyone working on contamination-resistant benchmarks. They will get value from the resource if the full paper supports the claims.\n\nRecommendation: send it to peer review only if the full manuscript exists and includes construction and leakage details. If the empty full text is just an indexing artifact, it deserves a serious referee who can pressure-test the unfamiliarity claim. If the paper really is abstract-only, invite resubmission. Either way, the idea is not dead on arrival—it just needs to show its work.","headline":"The benchmark is a good idea and worth engaging, but with the full text empty and the key 'unfamiliar shifts' claim unsupported, the paper is unverified rather than wrong.","tokens_in":1471,"tokens_out":1566,"would_cite":false,"duration_ms":16755,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"MDK12-Bench is a 141K-question, six-discipline K-12 benchmark that evaluates multimodal large language models on difficulty, temporal, contextual, and knowledge-driven reasoning by generating unfamiliar question variants.","keywords":["multimodal large language model benchmark","K-12 examinations","dynamic evaluation","data contamination","knowledge points","knowledge-point retrieval-augmented generation","multidisciplinary evaluation","MLLM robustness"],"falsifier":"Do a near-duplicate search for MDK12-Bench items and their answer keys in the pretraining corpora of the evaluated MLLMs; if a substantial share of shifted items appears verbatim or near-verbatim, or if models score as high on shifted items as on original ones, the central contamination-mitigation claim fails.","tokens_in":707,"feed_emoji":"📚","tokens_out":5658,"duration_ms":57079,"temperature":0.7,"pith_summary":"This paper tries to establish that multimodal large language models (MLLMs) can be evaluated on real K-12 exam material in a way that reflects generalization rather than memorization. To that end, it introduces MDK12-Bench, with 141K questions across six disciplines and 6,225 knowledge points organized in a six-layer taxonomy, annotated by difficulty, year, and question format. The benchmark's dynamic evaluation framework generates unfamiliar visual, textual, and question-form variants, so a model cannot simply recall a seen exam answer. If the design works, MDK12-Bench gives developers and educators a contamination-resistant, fine-grained measurement of MLLM capability on school-level problems, and it can guide work on robustness and AI-assisted education.","feed_headline":"141K K-12 exam questions stress-test multimodal LLMs","feed_subtitle":"New benchmark shifts visuals, wording, and question form to expose real reasoning gaps, not memorized answers.","key_machinery":"The central object is MDK12-Bench itself: a benchmark of 141K questions from real K-12 exams, organized around 6,225 knowledge points in a six-layer taxonomy and annotated with difficulty, year, and question format. The mechanism that carries the argument is the dynamic evaluation framework: it generates unfamiliar variants of each item by shifting visual style, wording, and question form, so that correct answers require handling novelty rather than retrieving memorized content. The knowledge-point reference-augmented generation (KP-RAG) setup is the supporting mechanism: it retrieves the relevant knowledge point as context, letting the authors separate failures caused by missing knowledge f","core_discovery":"MDK12-Bench is a large-scale benchmark assembled from genuine K-12 examinations, covering six disciplines with 141K instances and 6,225 knowledge points in a six-layer taxonomy. Each item carries annotations for difficulty, year, and one of five question formats, enabling evaluations along four dimensions: difficulty, cross-year temporal shift, contextual shift, and knowledge-driven reasoning. The accompanying dynamic evaluation framework creates changed visual presentations, reworded text, and altered question forms for otherwise-same knowledge points, making previously seen answers less useful and mitigating data contamination. The paper reports that current MLLMs show measurable limitatio","pith_inferences":["Editorial extension: the same shift-generation recipe (visual, textual, and format perturbations per item) could be applied to other benchmark suites as a generic anti-contamination layer.","Editorial extension: if performance declines smoothly with the year of the exam, MDK12-Bench could double as a probe for when a model's pretraining knowledge ends, something the paper does not develop.","Editorial extension: the six-layer knowledge taxonomy could feed diagnostic tutoring systems that attribute a student's missed answer to a missing knowledge node; that product is not part of this paper.","Editorial extension: a testable prediction follows: KP-RAG should help most on high-difficulty, knowledge-heavy questions and least on simple visual matching; grouping results by knowledge point and difficulty would confirm or refute this."],"forward_implications":["If the dynamic shifts do block memorization, then MDK12-Bench scores will remain informative even as future models are trained on more public web data.","The per-discipline, per-difficulty, and per-knowledge-point annotations let users localize exactly where a model breaks, from visual parsing to content knowledge.","KP-RAG results imply that adding a short knowledge-point reference to the prompt can change performance, pointing to knowledge retrieval as a practical lever for improving MLLM answers.","The four evaluation dimensions provide a structured way to report model progress beyond one aggregate accuracy score."],"supporting_citations":[],"fun_headline_variants":["Benchmark adds visual and text twists to K-12 exams for AI","New K-12 exam benchmark tests LLMs with shifting questions","Multimodal LLMs struggle on 141K real K-12 exam questions","Dynamic K-12 benchmark reveals AI gaps under changed conditions"],"cache_read_input_tokens":2816,"weakest_assumption_plain":"The load-bearing premise is that the unfamiliar visual, textual, and question-form variants really are unfamiliar to the tested models; if any of those variants or their answer keys already appear in the models' pretraining data, the benchmark's contamination-resistance and generalization claims weaken.","fun_headline_variants_meta":{"raw":{"variants":["Benchmark adds visual and text twists to K-12 exams for AI","New K-12 exam benchmark tests LLMs with shifting questions","Multimodal LLMs struggle on 141K real K-12 exam questions","Dynamic K-12 benchmark reveals AI gaps under changed conditions"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000485,"raw_usage":{"total_tokens":2230,"prompt_tokens":746,"completion_tokens":1484,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":490,"completion_tokens_details":{"reasoning_tokens":1408}},"tokens_in":490,"tokens_out":1484,"duration_ms":9527,"temperature":1.0,"reasoning_tokens":1408,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-05T22:28:04.499978+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Do a near-duplicate search for MDK12-Bench items and their answer keys in the pretraining corpora of the evaluated MLLMs; if a substantial share of shifted items appears verbatim or near-verbatim, or if models score as high on shifted items as on original ones, the central contamination-mitigation claim fails.","supporting_citations":[],"review_version":1}