REVIEW 4 major objections 7 minor 31 references
MTCMB: A Multi-Task Benchmark Framework for Evaluating LLMs on Knowledge, Reasoning, and Safety in Traditional Chinese Medicine
T0 review · 4 major / 7 minor · reviewed 2026-08-07 · deepseek-v4-flash
Pith's one-line read MTCMB, a 12-sub-dataset benchmark built with certified TCM practitioners, demonstrates that current LLMs master factual TCM knowledge but fall short in clinical reasoning, prescription planning, and safety compliance.
desk verdict Useful new benchmark resource for TCM evaluation, but the headline capability-gap claim is not yet proven because the scoring metrics and safety judge are not adequately validated. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the MTCMB task suite itself: 12 sub-datasets drawn from national attending and practitioner exams, real clinical cases, classical texts, and official safety guidelines, curated and reviewed with certified TCM practitioners. The five categories measure distinct competences: TCM-ED-A, TCM-ED-B, and TCM-FT for knowledge; TCMeEE, TCM-CHGD, and TCM-LitData for language understanding; TCM-MSDD and TCM-Diagnosis for diagnostic reasoning; TCM-PR and TCM-FRD for prescription generation; and TCM-SE-A plus TCM-SE-B for safety. The evaluation machinery combines exact-match accuracy, averages of BLEU, ROUGE, and BERTScore, one official competition score, and LLM-based grading for open safety items, all mapped to a 0-100 scale. Expert-human ratings correlate significantly with the automatic scores, most strongly on formula recommendation, which the paper uses as evidence that the benchmark's automated metrics track clinically relevant quality.
What would settle it
Run a leakage-controlled version of MTCMB using exam and case items dated after the tested models' training cutoffs, or apply membership-inference and n-gram overlap checks to the released items. If the high TCM-ED-A and TCM-ED-B scores collapse while reasoning scores stay flat, the paper's core knowledge-versus-reasoning finding is an artifact of memorized test items rather than TCM competence.
Extended reading notes
Core claim
The paper's central claim is that MTCMB is a clinically grounded testbed that reveals a systematic split in current LLM ability: strong factual recall and entity extraction, weak higher-order clinical reasoning. On licensing-exam-style knowledge sets such as TCM-ED-A and TCM-ED-B, top general models exceed 85% accuracy, while multi-label syndrome and disease classification (TCM-MSDD) and diagnostic report generation stay around 40% or lower for most models. Prescription tasks show similar limits, and although safety multiple-choice scores look high, qualitative review shows models miss rare contraindications and context-specific restrictions. Medical-specific and reasoning-optimized models give only modest improvements, so the paper concludes that neither domain fine-tuning nor prompting alone closes the reasoning and safety gap.
Load-bearing premise
The benchmark's findings assume the MTCMB test items were not memorized by the evaluated models during pretraining, since the paper reports no contamination or leakage analysis.
Editorial extensions
If this is right
- If MTCMB measures what it claims, factual knowledge scores cannot be used as evidence of clinical readiness in TCM.
- Few-shot and chain-of-thought prompting produce only incremental gains, so closing the reasoning gap will require domain-aligned training data or additional structured components.
- TCM safety evaluation needs its own items: high accuracy on standard safety multiple-choice coexists with missed rare contraindications in qualitative checks.
- The significant expert-automation correlations on TCM-FRD support using MTCMB's automated metrics as a proxy for human grading in formula-recommendation tasks, making larger evaluations feasible.
Reading between the lines
- An implicit consequence is that the five-task template (knowledge, language, diagnosis, prescription, safety) with expert curation could be rebuilt for other under-standardized medical systems such as Ayurveda or Kampo.
- The reported knowledge-versus-reasoning split may partly reflect memorization of test items rather than genuine competence, because the paper reports no contamination analysis; a temporal holdout version would settle this.
- The low diagnosis scores suggest a concrete development target: if a model's syndrome-differentiation accuracy on TCM-MSDD crosses a clinician baseline, that would be a stronger signal of clinical utility than the knowledge QA numbers.
- Because the safety dimension contains only 100 items and one of its two sets is graded by another LLM, it is the least saturated measurement; expanding it would likely reveal even wider gaps.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces MTCMB, a multi-task benchmark for evaluating LLMs on Traditional Chinese Medicine, comprising 12 sub-datasets across knowledge QA, language understanding, diagnosis, prescription recommendation, and safety. The data are sourced from licensing exams, clinical cases, classical texts, and safety guidelines, with claimed expert curation. The authors evaluate 14 general, medical-specialized, and reasoning-oriented LLMs under zero-shot, few-shot, and chain-of-thought prompting, and report that models perform well on factual knowledge but fall short in clinical reasoning, prescription generation, and safety compliance. A small human-correlation study on three datasets is also reported.
Significance. If validated, MTCMB would fill a genuine gap as a publicly released, multi-task, expert-curated benchmark for TCM, and the paper's release of datasets, code, and evaluation tools is a concrete strength. The comparison of 14 models across three prompting settings is useful. However, the headline capability claims are currently under-supported by the evaluation protocol: the absence of contamination checks, expert or chance baselines, and validation of the LLM-based safety judge leaves open the possibility that the reported gaps are partly metric artifacts. The benchmark resource itself is valuable, but the paper's central empirical conclusion needs substantially stronger evidence before it can be accepted.
major comments (4)
- [Sections 4.1 and 5.1] The paper never reports any contamination or leakage analysis for the exam, Q&A, and clinical-case items, even though several sources are public (attending physician exam bank, TCM Q&A Bank, Aliyun Tianchi dataset). Section 5.1 describes the evaluation setup but contains no n-gram overlap check, temporal cutoff analysis, or exclusion of memorized items. The high knowledge-QA scores could therefore reflect memorization of items present in pretraining corpora rather than TCM competence, so the 'knowledge strong' claim is not yet supported.
- [Sections B.3-B.5 and Tables 2-4] Free-form diagnosis (TCM-Diagnosis), prescription (TCM-FRD, TCM-PR), and language-understanding (TCMeEE, TCM-CHGD) scores are computed as averages of BLEU, ROUGE, and/or BERTScore against a single expert reference (Equations 1-5). No expert baseline, expert-expert agreement, or chance level is reported for any task. Because valid alternative diagnoses or equivalent formulas are penalized by construction, the low absolute scores in reasoning and prescription tasks may be metric artifacts rather than capability gaps; an expert performance floor is needed to interpret these numbers.
- [Section 4.1 (Task 5) and Section B.6] TCM-SE-A safety scores are produced by GLM-4-Air-250414 as an LLM judge, with no human validation of the judge on the 50 fill-in-the-blank items. Section 5.3 acknowledges that 'all models, including specialized medical LLMs, frequently missed rare contraindications' despite high TCM-SE-A/B scores, which is internally inconsistent if the safety metric were well calibrated. Either the LLM judge is too lenient, or the qualitative claim needs quantitative support; without human-rated safety outputs or a validated judge, the safety-compliance gap is not established.
- [Section 5.4] The expert-correlation analysis uses only 20 test instances per dataset, aggregates 14 model responses by trimmed mean, and covers only TCM-CHGD, TCM-Diagnosis, and TCM-FRD. Correlation at this scale can support ranking of models, but it does not validate the absolute adequacy of scores or the thresholds used to conclude that models 'fall short' in Section 5.3. Moreover, this validation is not applied to the TCM-SE-A LLM-judge scores, which is the safety metric most in need of external validation.
minor comments (7)
- [Section 5.1] The paper states that default decoding hyperparameters are used for fairness, but the actual temperature, top-p, and maximum token settings are not reported, which limits reproducibility given API-version-dependent variability.
- [Introduction] The introduction says 'we curate eleven sub-datasets', while Table 1 lists twelve; this discrepancy should be corrected.
- [Sections 3.2 and A.4] There are typos: 'follwed' in Section 3.2, 'laten' in Section 3.2, and 'Chain-of-Though' in Appendix A.4.
- [Figure 1] The figure caption says 'four primary dimensions' and excludes knowledge QA, but the caption does not state this exclusion explicitly; please clarify.
- [Table 2] Some rows in Table 2 have numbers running together (e.g., '52.055.064.0' for GPT-4.1 zero-shot on eEE/CHGD/LitData), making the table difficult to read.
- [Model names throughout] There are model-name inconsistencies: 'Gemini 3' in the introduction versus 'Gemini-2.5' in Section 5.1, 'O4-Mini' versus 'O4-mini', 'Deepseek-R1' versus 'DeepSeek-r1' in tables, and 'Gemini Ultra' in Section 5.3 does not appear in the model list.
- [Section B.3] Section B.3 refers to 'TCM-DiagData' while the dataset is called 'TCM-Diagnosis' in Table 1 and Section 4.1; please unify the name.
Circularity Check
No circularity: MTCMB scores are measured against external expert-curated labels and references; the LLM-judge and metric-validity concerns are validity issues, not circular reductions.
full rationale
MTCMB is a benchmark construction and evaluation paper, not a derivation. The twelve sub-datasets are built from external sources: national licensing exam banks, CCL25-Eval competition data, real clinical cases, classical TCM texts, and expert-written safety questions (Table 1, Section 4.1). No parameter in the paper is fitted to the evaluated models' outputs, and no sub-dataset is defined in terms of the headline conclusions. The 'prediction' that LLMs underperform on reasoning, prescription, and safety is an empirical finding obtained by comparing model outputs to fixed gold references via accuracy, BLEU/ROUGE, BERTScore, and competition scores; it is not a quantity that is equal to its own inputs by construction. The one reflexive element is TCM-SE-A, which is scored by GLM-4-Air-250414, an LLM judge (Table 1, Section 4.1, Appendix B.6). That is a measurement-validity concern: the LLM judge is not validated against expert ratings, no chance-level baseline is reported, and the paper's own qualitative analysis (Section 5.3) notes that high SE-A/B scores coexist with missed contraindications. However, this is not circularity under the defined patterns. The judge is not a fitted parameter, is not among the 14 evaluated models, the safety dimension also includes 50 expert-written multiple-choice questions (TCM-SE-B), and the low safety scores are not forced by construction. Similarly, the Section 5.4 correlation evidence (20 instances per dataset) bears on whether lexical metrics are adequate proxies for expert judgment, but moderate correlation does not make the low reasoning scores definitionally forced. There are no self-citations used as load-bearing premises and no imported uniqueness theorems. The central claims are therefore self-contained against external labels, and the circularity score is 0.
Assumptions & free parameters
free parameters (2)
- Composite metric equal weighting =
1/3 BLEU, 1/3 ROUGE, 1/3 BERTScore
- LLM judge model for TCM-SE-A =
GLM-4-Air-250414
assumptions (4)
- domain assumption Benchmark gold labels from TCM experts, exam banks, and classical texts are correct and consistent.
- domain assumption Evaluated LLMs have not memorized MTCMB items during pretraining.
- domain assumption Automated metrics such as BLEU, ROUGE, and BERTScore are valid proxies for TCM clinical quality.
- domain assumption Default decoding hyperparameters of each model API provide a fair comparison.
Cite this review
Pith. "Pith review of MTCMB: A Multi-Task Benchmark Framework for Evaluating LLMs on Knowledge, Reasoning, and Safety in Traditional Chinese Medicine." pith.science (2026). https://pith.science/paper/P5PL3VNE
@misc{pith2026250601252,
author = {Pith},
title = {Pith review of: MTCMB: A Multi-Task Benchmark Framework for Evaluating LLMs on Knowledge, Reasoning, and Safety in Traditional Chinese Medicine},
year = {2026},
howpublished = {\url{https://pith.science/paper/P5PL3VNE}},
note = {Machine review of arXiv:2506.01252}
}
read the original abstract
Traditional Chinese Medicine (TCM) is a holistic medical system with millennia of accumulated clinical experience, playing a vital role in global healthcare-particularly across East Asia. However, the implicit reasoning, diverse textual forms, and lack of standardization in TCM pose major challenges for computational modeling and evaluation. Large Language Models (LLMs) have demonstrated remarkable potential in processing natural language across diverse domains, including general medicine. Yet, their systematic evaluation in the TCM domain remains underdeveloped. Existing benchmarks either focus narrowly on factual question answering or lack domain-specific tasks and clinical realism. To fill this gap, we introduce MTCMB-a Multi-Task Benchmark for Evaluating LLMs on TCM Knowledge, Reasoning, and Safety. Developed in collaboration with certified TCM experts, MTCMB comprises 12 sub-datasets spanning five major categories: knowledge QA, language understanding, diagnostic reasoning, prescription generation, and safety evaluation. The benchmark integrates real-world case records, national licensing exams, and classical texts, providing an authentic and comprehensive testbed for TCM-capable models. Preliminary results indicate that current LLMs perform well on foundational knowledge but fall short in clinical reasoning, prescription planning, and safety compliance. These findings highlight the urgent need for domain-aligned benchmarks like MTCMB to guide the development of more competent and trustworthy medical AI systems. All datasets, code, and evaluation tools are publicly available at: https://github.com/Wayyuanyuan/MTCMB.
Figures
Figures from the paper (25 more)
Reference graph
Works this paper leans on
- [1]
-
[2]
Zhijie Bao, Wei Chen, Shengze Xiao, Kuang Ren, Jiaao Wu, Cheng Zhong, Jiajie Peng, Xuanjing Huang, and Zhongyu Wei. Disc-medllm: Bridging general large language models and real-world medical consultation.arXiv preprint arXiv:2308.14346, 2023
arXiv 2023
-
[3]
Huatuogpt-o1, towards medical complex reasoning with llms.arXiv preprint arXiv:2412.18925, 2024
Junying Chen, Zhenyang Cai, Ke Ji, Xidong Wang, Wanlong Liu, Rongsheng Wang, Jianye Hou, and Benyou Wang. Huatuogpt-o1, towards medical complex reasoning with llms.arXiv preprint arXiv:2412.18925, 2024
arXiv 2024
-
[4]
Behonest: Benchmarking honesty in large language models.arXiv preprint arXiv:2406.13261, 2024
Steffi Chern, Zhulin Hu, Yuqing Yang, Ethan Chern, Yuan Guo, Jiahe Jin, Binjie Wang, and Pengfei Liu. Behonest: Benchmarking honesty in large language models.arXiv preprint arXiv:2406.13261, 2024
arXiv 2024
-
[5]
Measuring massive multitask language understanding.arXiv preprint arXiv:2009.03300, 2020
Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt. Measuring massive multitask language understanding.arXiv preprint arXiv:2009.03300, 2020
arXiv 2009
-
[6]
Bing Hu, Shuang-Shuang Wang, and Qin Du. Traditional chinese medicine for prevention and treatment of hepatocarcinoma: from bench to bedside.World Journal of Hepatology, 7(9):1209, 2015
work page 2015
-
[7]
What disease does this patient have? a large-scale open domain question answering dataset from medical exams.Applied Sciences, 11(14):6421, 2021
Di Jin, Eileen Pan, Nassim Oufattole, Wei-Hung Weng, Hanyi Fang, and Peter Szolovits. What disease does this patient have? a large-scale open domain question answering dataset from medical exams.Applied Sciences, 11(14):6421, 2021
2021
-
[8]
Qiao Jin, Bhuwan Dhingra, Zhengping Liu, William W. Cohen, and Xinghua Lu. Pubmedqa: A dataset for biomedical research question answering. In Kentaro Inui, Jing Jiang, Vincent Ng, and Xiaojun Wan, editors,Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Pr...
work page 2019
Show all 31 references
-
[9]
Deepseek-v3 technical report.arXiv preprint arXiv:2412.19437, 2024
Aixin Liu, Bei Feng, Bing Xue, Bingxuan Wang, Bochao Wu, and others. Deepseek-v3 technical report.arXiv preprint arXiv:2412.19437, 2024
2024 arXiv
-
[10]
Benchmarking large language models on cmexam-a comprehensive chinese medical exam dataset.Advances in Neural Information Processing Systems, 36:52430–52452, 2023
Junling Liu, Peilin Zhou, Yining Hua, Dading Chong, Zhongyu Tian, Andrew Liu, Helin Wang, Chenyu You, Zhenhua Guo, Lei Zhu, et al. Benchmarking large language models on cmexam-a comprehensive chinese medical exam dataset.Advances in Neural Information Processing Systems, 36:52...
2023
-
[11]
Taiyi: a bilingual fine-tuned large language model for diverse biomedical tasks.Journal of the American Medical Informatics Association, 31(9):1865–1874, 2024
Ling Luo, Jinzhong Ning, Yingwen Zhao, Zhijun Wang, Zeyuan Ding, Peng Chen, Weiru Fu, Qinyu Han, Guangtao Xu, Yunzhi Qiu, et al. Taiyi: a bilingual fine-tuned large language model for diverse biomedical tasks.Journal of the American Medical Informatics Association, 31(9):1865–...
2024
-
[12]
Sg-bench: Evaluating llm safety generalization across diverse tasks and prompt types.Advances in Neural Information Processing Systems, 37:123032–123054, 2024
Yutao Mou, Shikun Zhang, and Wei Ye. Sg-bench: Evaluating llm safety generalization across diverse tasks and prompt types.Advances in Neural Information Processing Systems, 37:123032–123054, 2024
2024
-
[13]
Gpt-4.1, 2025
OpenAI. Gpt-4.1, 2025. Accessed: 2025-05-14. 12
2025
-
[14]
Openai o4-mini, 2025
OpenAI. Openai o4-mini, 2025. Accessed: 2025-05-14
2025
-
[15]
Medmcqa: A large-scale multi-subject multi-choice dataset for medical domain question answering
Ankit Pal, Logesh Kumar Umapathi, and Malaikannan Sankarasubbu. Medmcqa: A large-scale multi-subject multi-choice dataset for medical domain question answering. In Gerardo Flores, George H Chen, Tom Pollard, Joyce C Ho, and Tristan Naumann, editors,Proceedings of the Conferenc...
2022
-
[16]
Medhallu: A comprehensive benchmark for detecting medical hallucinations in large language models.arXiv preprint arXiv:2502.14302, 2025
Shrey Pandit, Jiawei Xu, Junyuan Hong, Zhangyang Wang, Tianlong Chen, Kaidi Xu, and Ying Ding. Medhallu: A comprehensive benchmark for detecting medical hallucinations in large language models.arXiv preprint arXiv:2502.14302, 2025
2025 arXiv
-
[17]
Qwen3-235b-a22b, 2025
Qwen Team. Qwen3-235b-a22b, 2025. Accessed: 2025-05-14
2025
-
[18]
Large language models encode clinical knowledge.Nature, 620(7972):172–180, 2023
Karan Singhal, Shekoofeh Azizi, Tao Tu, S Sara Mahdavi, Jason Wei, Hyung Won Chung, Nathan Scales, Ajay Tanwani, Heather Cole-Lewis, Stephen Pfohl, et al. Large language models encode clinical knowledge.Nature, 620(7972):172–180, 2023
2023
-
[19]
Toward expert-level medical question answering with large language models.Nature Medicine, pages 1–8, 2025
Karan Singhal, Tao Tu, Juraj Gottweis, Rory Sayres, Ellery Wulczyn, Mohamed Amin, Le Hou, Kevin Clark, Stephen R Pfohl, Heather Cole-Lewis, et al. Toward expert-level medical question answering with large language models.Nature Medicine, pages 1–8, 2025
2025
-
[20]
Qwen2.5 technical report.arXiv preprint arXiv:2412.15115, 2024
Qwen Team. Qwen2.5 technical report.arXiv preprint arXiv:2412.15115, 2024
2024 arXiv
-
[21]
Large language models in medicine.Nature medicine, 29(8):1930–1940, 2023
Arun James Thirunavukarasu, Darren Shu Jeng Ting, Kabilan Elangovan, Laura Gutierrez, Ting Fang Tan, and Daniel Shu Wei Ting. Large language models in medicine.Nature medicine, 29(8):1930–1940, 2023
1930
-
[22]
Alex Wang, Yada Pruksachatkun, Nikita Nangia, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel R. Bowman. Superglue: A stickier benchmark for general-purpose language understanding systems, 2020
2020
-
[23]
Baichuan-m1: Pushing the medical capability of large language models.arXiv preprint arXiv:2502.12671, 2025
Bingning Wang, Haizhou Zhao, Huozhi Zhou, Liang Song, Mingyu Xu, Wei Cheng, Xiangrong Zeng, Yupeng Zhang, Yuqi Huo, Zecheng Wang, et al. Baichuan-m1: Pushing the medical capability of large language models.arXiv preprint arXiv:2502.12671, 2025
2025 arXiv
-
[24]
Cmb: A comprehensive medical benchmark in chinese.arXiv preprint arXiv:2308.08833, 2023
Xidong Wang, Guiming Hardy Chen, Dingjie Song, Zhiyi Zhang, Zhihong Chen, Qingying Xiao, Feng Jiang, Jianquan Li, Xiang Wan, Benyou Wang, et al. Cmb: A comprehensive medical benchmark in chinese.arXiv preprint arXiv:2308.08833, 2023
2023 arXiv
-
[25]
Tcmeval-sdt: a benchmark dataset for syndrome differentiation thought of traditional chinese medicine.Scientific Data, 12(1):437, 2025
Zhe Wang, Meng Hao, Suyuan Peng, Yuyan Huang, Yiwei Lu, Keyu Yao, Xiaolin Yang, and Yan Zhu. Tcmeval-sdt: a benchmark dataset for syndrome differentiation thought of traditional chinese medicine.Scientific Data, 12(1):437, 2025
2025
-
[26]
Wingpt2-14b-chat, 2025
WiNGPT Team. Wingpt2-14b-chat, 2025. Accessed: 2025-05-14
2025
-
[27]
Catalysing ancient wisdom and modern science for the health and well-being of people and planet, 2024
World Health Organization. Catalysing ancient wisdom and modern science for the health and well-being of people and planet, 2024. Accessed: 2025-05-13
2024
-
[28]
Tcmd: A traditional chinese medicine qa dataset for evaluating large language models.arXiv preprint arXiv:2406.04941, 2024
Ping Yu, Kaitao Song, Fengchen He, Ming Chen, and Jianfeng Lu. Tcmd: A traditional chinese medicine qa dataset for evaluating large language models.arXiv preprint arXiv:2406.04941, 2024
2024 arXiv
-
[29]
Tcmbench: A comprehensive benchmark for evaluating large language models in traditional chinese medicine.arXiv preprint arXiv:2406.01126, 2024
Wenjing Yue, Xiaoling Wang, Wei Zhu, Ming Guan, Huanran Zheng, Pengfei Wang, Changzhi Sun, and Xin Ma. Tcmbench: A comprehensive benchmark for evaluating large language models in traditional chinese medicine.arXiv preprint arXiv:2406.01126, 2024
2024 arXiv
-
[30]
Qibo: A large language model for traditional chinese medicine.arXiv preprint arXiv:2403.16056, 2024
Heyi Zhang, Xin Wang, Zhaopeng Meng, Zhe Chen, Pengwei Zhuang, Yongzhe Jia, Dawei Xu, and Wenbin Guo. Qibo: A large language model for traditional chinese medicine.arXiv preprint arXiv:2403.16056, 2024
2024 arXiv
-
[31]
Safetybench: Evaluating the safety of large language models.arXiv preprint arXiv:2309.07045, 2023
Zhexin Zhang, Leqi Lei, Lindong Wu, Rui Sun, Yongkang Huang, Chong Long, Xiao Liu, Xuanyu Lei, Jie Tang, and Minlie Huang. Safetybench: Evaluating the safety of large language models.arXiv preprint arXiv:2309.07045, 2023. 13 Figure 5: The figure illustrates the design of promp...
2023 arXiv
Reviewed August 7, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.