Pith. sign in

REVIEW 5 cited by

TCMBench: A Comprehensive Benchmark for Evaluating Large Language Models in Traditional Chinese Medicine

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2406.01126 v1 pith:FUWV4HZO submitted 2024-06-03 cs.CL cs.AI

classification cs.CLcs.AI
keywords llmsbenchmarkcomprehensivedomainevaluatingincludinglanguagemodels
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Large language models (LLMs) have performed remarkably well in various natural language processing tasks by benchmarking, including in the Western medical domain. However, the professional evaluation benchmarks for LLMs have yet to be covered in the traditional Chinese medicine(TCM) domain, which has a profound history and vast influence. To address this research gap, we introduce TCM-Bench, an comprehensive benchmark for evaluating LLM performance in TCM. It comprises the TCM-ED dataset, consisting of 5,473 questions sourced from the TCM Licensing Exam (TCMLE), including 1,300 questions with authoritative analysis. It covers the core components of TCMLE, including TCM basis and clinical practice. To evaluate LLMs beyond accuracy of question answering, we propose TCMScore, a metric tailored for evaluating the quality of answers generated by LLMs for TCM related questions. It comprehensively considers the consistency of TCM semantics and knowledge. After conducting comprehensive experimental analyses from diverse perspectives, we can obtain the following findings: (1) The unsatisfactory performance of LLMs on this benchmark underscores their significant room for improvement in TCM. (2) Introducing domain knowledge can enhance LLMs' performance. However, for in-domain models like ZhongJing-TCM, the quality of generated analysis text has decreased, and we hypothesize that their fine-tuning process affects the basic LLM capabilities. (3) Traditional metrics for text generation quality like Rouge and BertScore are susceptible to text length and surface semantic ambiguity, while domain-specific metrics such as TCMScore can further supplement and explain their evaluation results. These findings highlight the capabilities and limitations of LLMs in the TCM and aim to provide a more profound assistance to medical research.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. TreeProbe : A Tibetan Medicine Benchmark for Cultural Bias in LLMs

    cs.CL 2026-08 conditional novelty 7.0 of 10

    TreeProbe, a 4,719-item Tibetan-medicine benchmark, shows LLMs score 40–60% and systematically drift to TCM or biomedical reasoning.

  2. MTCMB: A Multi-Task Benchmark Framework for Evaluating LLMs on Knowledge, Reasoning, and Safety in Traditional Chinese Medicine

    cs.CL 2025-06 conditional novelty 6.0 of 10

    MTCMB is a 12-dataset benchmark for evaluating LLMs on Traditional Chinese Medicine knowledge, reasoning, and safety, with results showing models still fail at clinical reasoning and safe prescriptions.

  3. OphthBench: A Comprehensive Benchmark for Evaluating Large Language Models in Chinese Ophthalmology

    cs.CL 2025-02 conditional novelty 6.0 of 10

    OphthBench is a new 591-question Chinese ophthalmology benchmark on which 39 LLMs score around 70% (after normalization), showing a clear gap between current models and clinical readiness.

  4. PARA: Parameter-Efficient Fine-tuning with Prompt Aware Representation Adjustment

    cs.CL 2025-02 conditional novelty 6.0 of 10

    PARA generates prompt-conditioned scaling vectors for Q, V, and FFN activations, outperforming (IA)^3 and LoRA-style baselines on several benchmarks with similar parameter counts and lower multi-tenant inference latency.

  5. Improving TCM Question Answering through Tree-Organized Self-Reflective Retrieval with LLMs

    cs.CL 2025-02 conditional novelty 4.0 of 10

    A tree-organized, self-reflective retrieval framework over a TCM knowledge base lifts GPT-4 accuracy on a 600-question licensing-exam sample by 19.85 absolute percentage points.

Pith tools