{"id":"ff049544-15c1-4361-89b8-9248be7d2bc8","arxiv_id":"2508.11894","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":4.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":0,"one_line_summary":"QuarkMed is described as a medical foundation model scoring 70% on the Chinese Medical Licensing Examination, but the abstract is the only usable evidence because the body text is a different paper.","lead":"This report describes QuarkMed, a medical foundation model that reports 70% accuracy on the Chinese Medical Licensing Examination via curated data, retrieval-augmented generation, and reinforcement learning. The supplied full text is an unrelated robotics preprint, so only the abstract could be assessed.","discovery_kind":"unclear","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 70% exam accuracy is unauditable: the submitted full text is a different paper, leaving the evaluation protocol, contamination controls, and verifier design completely unspecified.","rationale":"The reader's verdict is UNVERDICTED with low confidence, based on the abstract-only evidence and the obvious mismatch between the abstract and the supplied full text. My stress-test confirms that the single load-bearing assumption is the validity of the 70% exam measurement. Without the actual technical report, there is no way to verify the most basic safeguards: no training/test contamination, a meaningful reward signal in RL, and a defined benchmark suite. The submission's full text being an unrelated paper makes the gap even more severe, because even the minimal methods section that would normally accompany such a claim is absent. This is not a challenge to the authors' integrity; it is a straightforward observation that the claim cannot be audited from the provided materials. The reader's weakest_assumption identified the same issue, so I agree. The appropriate verdict remains UNVERDICTED, because the abstract alone neither proves nor disproves the claim. My concrete test would resolve the uncertainty by obtaining the full report and examining the evaluation protocol. If the report is unavailable, the claim should remain unverified until it is provided.","tokens_in":2375,"tokens_out":2485,"duration_ms":28024,"concrete_test":"Retrieve the actual QuarkMed technical report (e.g., from the arXiv listing or by author request). If found, check the evaluation section: confirm the Chinese Medical Licensing Examination questions were not in the training set, the answer extraction is deterministic, the scoring follows the official rubric, and the RL verifier is validated against human expert labels. If the report is unavailable, treat the 70% claim as unverified and set correctness_risk to high.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is the 70% accuracy on the Chinese Medical Licensing Examination. For that number to support 'strong generalization across diverse medical benchmarks,' we need to know (i) the exam items were held out from pretraining and RL data; (ii) the RL reward verifier independently checks medical correctness rather than pattern-matching; (iii) the benchmark suite and scoring protocol are defined. None of this is present in the submission. The supplied full text is arXiv 2508.11898, OmniD, a robot manipulation paper; it contains no mention of QuarkMed, medical data, or the licensing exam. Per the reviewing rule, this is a mechanical flag: a central assertion with no supporting methods or evaluation section. The abstract's claims about RAG, verifiable RL, and consumer deployment are likewise unbacked. This is not an internal inconsistency but a missing-evidence gap: the strongest claim is entirely unsupported. The number 70% might be accurate, but there is no way to assess contamination, verifier reliability, or benchmark validity from the available text.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The abstract announces QuarkMed, a medical foundation model built on curated medical data, medical-content Retrieval-Augmented Generation (RAG), and a large-scale verifiable reinforcement learning pipeline, and reports 70% accuracy on the Chinese Medical Licensing Examination, claiming strong generalization across diverse medical benchmarks and deployment to millions of users. However, the full text supplied with the submission is not a technical report for QuarkMed: it is an unrelated robotics preprint, OmniD, on bird's-eye-view representations for robot manipulation. The body contains no mention of QuarkMed, medical data, the licensing examination, RAG, or reinforcement learning. Consequently, the submission provides no methods, no evaluation protocol, no benchmark definitions, and no support for any of the abstract's claims.","tokens_in":2482,"tokens_out":3445,"duration_ms":40277,"significance":"If correct, a medical foundation model reaching 70% on the Chinese Medical Licensing Examination and generalizing across medical benchmarks would be practically significant, especially if the model is deployable at consumer scale. However, the submission as it stands offers no verifiable evidence for these claims. The significance assessment is therefore conditional on future provision of a proper technical report. The current manuscript cannot advance the field because no part of the claimed methodology or evaluation is present in the submitted text.","major_comments":[{"comment":"The central claim—70% accuracy on the Chinese Medical Licensing Examination—is completely unsupported by the body of the manuscript. The body is the OmniD robotics paper, which does not mention QuarkMed, medical licensing, medical data, or any evaluation of a medical model. No evaluation protocol, dataset description, or scoring methodology is given. This is a load-bearing omission: the single accuracy number is the paper's main result, and the submitted text provides no way to audit it.","section":"Abstract"},{"comment":"The generalization claim, 'demonstrating strong generalization across diverse medical benchmarks,' is not supported by any named benchmark, baseline comparison, error bar, or statistical test. Even if the 70% figure were valid for one exam, no evidence is presented that it transfers to other medical tasks. The manuscript must define the benchmark suite and provide per-benchmark results with appropriate uncertainty estimates.","section":"Abstract"},{"comment":"The 'large-scale, verifiable reinforcement learning pipeline' is neither described nor verified. The manuscript gives no details of the verifier, the reward model, the training data, or the procedures used to prevent train/test contamination (e.g., whether exam items were excluded from pretraining and RL data). Without this information, the 70% figure is unauditable. The verifier's medical correctness checking must be described concretely to rule out pattern-matching or reward hacking.","section":"Abstract"},{"comment":"The full text is a different manuscript with a different title, abstract, and subject matter. This is not a missing section or a presentation issue; it means the submission does not contain the claimed technical report. The authors must provide the actual QuarkMed report, including architecture, data curation, training procedure, evaluation protocol, and results. As submitted, the paper cannot be reviewed for scientific soundness.","section":"Full Text (OmniD)"}],"minor_comments":[{"comment":"The title and abstract describe a medical foundation model, while the full text is a robotics paper. This mismatch should be resolved before any resubmission.","section":"Abstract vs. Full Text"},{"comment":"The claim that the model is 'already serving over millions of users at ai.quark.cn' is not a technical result and cannot substitute for evaluation. If intended as a deployment claim, it should be separated from the scientific evaluation.","section":"Abstract"},{"comment":"There are no references to medical benchmarks, prior medical LLMs, or related work on medical licensing examinations. A proper technical report must cite and compare to relevant baselines.","section":"Global"},{"comment":"The term 'QuarkMed' is not defined anywhere in the submitted text; no architecture, parameter count, or model family is given. This information is essential for reproducibility.","section":"Global"}],"recommendation":"reject","confidential_remarks":"In my view, this submission has a fundamental integrity problem: the body text is a different paper. The abstract's claims are entirely disconnected from the supplied full text. Even if the authors later provide the actual QuarkMed technical report, this version should be replaced and treated as a new submission. I would not solicit a revision of the present text."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The thing you should know first: this submission is not a paper about QuarkMed. The abstract describes a medical foundation model, but the full text supplied is OmniD, a robot manipulation preprint (arXiv 2508.11898). There is zero overlap. So the only assessable content is the abstract itself, and that is not enough to evaluate anything beyond a claim.\n\nWhat is the abstract actually saying? QuarkMed uses curated medical data, medical RAG, and a verifiable RL pipeline to reach 70% accuracy on the Chinese Medical Licensing Examination, and it already serves millions of users. If a real technical report backed that up, it would be a meaningful within-subfield result for medical AI. The abstract is readable and the claim is specific. That is about the only credit I can give.\n\nThe soft spots are not soft; they are load-bearing and entirely missing. The 70% number has no evaluation protocol, no contamination controls, no verifier design, no benchmark definitions, and no baselines. One number cannot license \"strong generalization across diverse medical benchmarks.\" The \"verifiable\" RL pipeline is asserted, not shown. There are no prior results cited, so there is no way to anchor novelty. The reader's suspicion about benchmark contamination is reasonable; we cannot rule it out. The submitted body being a different paper makes the whole thing mechanically unauditable. This is a submission-integrity problem, not a methods weakness.\n\nHonestly, there is nothing here to referee. The manuscript is internally incoherent: the title and abstract are about QuarkMed, the body is about OmniD. No serious editor should send this to peer review. If the authors later post the actual QuarkMed technical report, then that document would deserve a careful look, provided it includes the evaluation details and contamination controls. But this submission, as it stands, is a desk reject.\n\nWho could get value from this? Only someone studying how technical reports can go wrong in the arXiv pipeline. Not worth your reading-group time, and not worth citing.","headline":"The submitted full text is an unrelated robotics preprint; the QuarkMed claims are abstract-only and unauditable.","tokens_in":3113,"tokens_out":1511,"would_cite":false,"duration_ms":18612,"reading_group":"no","serious_thinker":"no","would_accept_peer_review":false},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper's central claim is that QuarkMed, a medical foundation model that combines curated data, medical RAG, and verifiable reinforcement learning, reaches 70% accuracy on the Chinese Medical Licensing Examination and generalizes across","keywords":["QuarkMed","medical foundation model","Chinese Medical Licensing Examination","retrieval-augmented generation","verifiable reinforcement learning","medical AI","large language model","generalization"],"falsifier":"Administer QuarkMed a fresh, never-before-used form of the Chinese Medical Licensing Examination, ensuring no overlap with its pretraining, RAG corpus, or RL verification data, and compare the independently measured accuracy with 70%; also audit the RL verifier for leaked answer keys or reward hacking. If the score falls materially short or the verifier cannot distinguish correct medical reasoning from superficially plausible answers, the generalization claim fails.","tokens_in":2173,"feed_emoji":"🩺","tokens_out":4969,"duration_ms":52867,"temperature":0.7,"pith_summary":"This medical-AI report's central claim is that a consumer-scale foundation model, QuarkMed, can pass a high-stakes medical exam: it reports 70% accuracy on the Chinese Medical Licensing Examination and says the same recipe generalizes across diverse medical benchmarks. The model is built from curated medical data processing, medical-content retrieval-augmented generation, and a large-scale 'verifiable' reinforcement learning pipeline. If the claim is right, a personally deployable medical AI can reach near-human exam competence and support consultation, diagnostic-assistance, and medical-search products at scale. For the reader's orientation: the supplied manuscript body is a different preprint on robot manipulation, so the evaluation protocol behind the 70% figure is not present in this text.","feed_headline":"QuarkMed claims 70% on China's medical licensing exam","feed_subtitle":"A medical foundation model bets on curated data, retrieval augmentation, and verifiable RL to pass exam-level questions.","key_machinery":"The central mechanism is the named 'verifiable reinforcement learning pipeline'—reinforcement learning in which the model's medical outputs are checked for correctness by an automated verifier—combined with medical-content Retrieval-Augmented Generation, which retrieves relevant medical passages before generating an answer. The paper claims these two components, on top of curated medical data processing, are what allow the model to reach license-exam-level accuracy while remaining broadly general.","core_discovery":"On its own terms, the paper claims that QuarkMed, a medical foundation model, achieves 70% accuracy on the Chinese Medical Licensing Examination and, on that basis, demonstrates strong generalization across diverse medical benchmarks. The path to this result is described as curated medical data processing, medical-content Retrieval-Augmented Generation, and a large-scale, verifiable reinforcement learning pipeline. The implied discovery is that a combination of grounded retrieval and verifiable RL feedback can push a general-purpose language model to professional-level medical accuracy at a scale suitable for serving millions of users.","pith_inferences":["A fair reader should treat the 70% figure as a pending claim: the abstract names no exam split, no contamination check, and no benchmark list, and the supplied body text is unrelated to the medical model.","A testable consequence of the paper's implied recipe is that removing the medical-content RAG module should measurably lower exam accuracy; an ablation comparing QuarkMed with and without RAG would expose whether retrieval is truly load-bearing.","If the verifiable RL step genuinely verifies medical correctness, the same training pattern could extend to other high-stakes, answer-checkable domains such as legal or financial certification exams, though the paper does not claim this."],"forward_implications":["If the 70% figure is measured on a clean held-out exam, QuarkMed demonstrates a level of medical knowledge that makes it viable as an AI-powered medical consultation and diagnostic-assistance tool.","If the generalization claim holds, the same RAG-plus-verifiable-RL recipe transfers across diverse medical benchmarks rather than overfitting to one exam format.","The model's reported deployment at ai.quark.cn would make this the first consumer-scale medical foundation model carrying a claimed license-exam-level accuracy.","The architectural recipe described in the abstract gives a template for other medical foundation models seeking verifiable accuracy rather than raw fluency."],"supporting_citations":[],"fun_headline_variants":["QuarkMed scores 70% on China's medical licensing exam","Medical foundation model hits 70% on Chinese exam","QuarkMed's 70% exam score: RAG plus verifiable RL","RAG and RL push QuarkMed to 70% on medical exam"],"cache_read_input_tokens":2816,"weakest_assumption_plain":"The load-bearing premise is that the reported 70% accuracy comes from a sound, held-out evaluation with no training-data overlap and with a verification mechanism that genuinely checks medical correctness.","fun_headline_variants_meta":{"raw":{"variants":["QuarkMed scores 70% on China's medical licensing exam","Medical foundation model hits 70% on Chinese exam","QuarkMed's 70% exam score: RAG plus verifiable RL","RAG and RL push QuarkMed to 70% on medical exam"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00028,"raw_usage":{"total_tokens":1430,"prompt_tokens":612,"completion_tokens":818,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":356,"completion_tokens_details":{"reasoning_tokens":741}},"tokens_in":356,"tokens_out":818,"duration_ms":8120,"temperature":1.0,"reasoning_tokens":741,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-05T19:42:08.893463+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Administer QuarkMed a fresh, never-before-used form of the Chinese Medical Licensing Examination, ensuring no overlap with its pretraining, RAG corpus, or RL verification data, and compare the independently measured accuracy with 70%; also audit the RL verifier for leaked answer keys or reward hacking. If the score falls materially short or the verifier cannot distinguish correct medical reasoning from superficially plausible answers, the generalization claim fails.","supporting_citations":[],"review_version":1}