Pith. sign in

REVIEW 9 cited by

Baichuan-M1: Pushing the Medical Capability of Large Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2502.12671 v2 pith:YOMSJOND submitted 2025-02-18 cs.CL

classification cs.CL
keywords medicalbaichuan-m1modelsgenerallanguagelargellmsmodel
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The current generation of large language models (LLMs) is typically designed for broad, general-purpose applications, while domain-specific LLMs, especially in vertical fields like medicine, remain relatively scarce. In particular, the development of highly efficient and practical LLMs for the medical domain is challenging due to the complexity of medical knowledge and the limited availability of high-quality data. To bridge this gap, we introduce Baichuan-M1, a series of large language models specifically optimized for medical applications. Unlike traditional approaches that simply continue pretraining on existing models or apply post-training to a general base model, Baichuan-M1 is trained from scratch with a dedicated focus on enhancing medical capabilities. Our model is trained on 20 trillion tokens and incorporates a range of effective training methods that strike a balance between general capabilities and medical expertise. As a result, Baichuan-M1 not only performs strongly across general domains such as mathematics and coding but also excels in specialized medical fields. We have open-sourced Baichuan-M1-14B, a mini version of our model, which can be accessed through the following links.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 9 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Aligning Clinical Needs and AI Capabilities: A Survey on LLMs for Medical Reasoning

    cs.AI 2026-07 accept novelty 6.0 of 10

    A dual clinical-computational taxonomy for medical LLM reasoning plus a five-level 5k-sample benchmark showing specialists excel at diagnosis and general models at decision support/dialogue.

  2. Explore-Execute Chain: Towards an Efficient Structured Reasoning Paradigm

    cs.LG 2025-09 conditional novelty 6.0 of 10

    E2C decouples LLM reasoning into a stochastic planning phase and a deterministic execution phase, achieving similar or better accuracy with far fewer generated tokens.

  3. MedBLINK: Probing Basic Perception in Multimodal Language Models for Medicine

    cs.AI 2025-08 conditional novelty 6.0 of 10

    Current medical multimodal models, including GPT-4o and Claude 3.5 Sonnet, fail simple perceptual tasks on medical images that human experts solve almost perfectly.

  4. VerIF: Verification Engineering for Reinforcement Learning in Instruction Following

    cs.CL 2025-06 conditional novelty 6.0 of 10

    A hybrid verifier that combines rule-based code checks and a reasoning-LLM judge enables reinforcement learning to improve LLM instruction following on several benchmarks.

  5. MTCMB: A Multi-Task Benchmark Framework for Evaluating LLMs on Knowledge, Reasoning, and Safety in Traditional Chinese Medicine

    cs.CL 2025-06 conditional novelty 6.0 of 10

    MTCMB is a 12-dataset benchmark for evaluating LLMs on Traditional Chinese Medicine knowledge, reasoning, and safety, with results showing models still fail at clinical reasoning and safe prescriptions.

  6. DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models

    cs.CL 2025-05 conditional novelty 6.0 of 10

    DiagnosisArena, a 1,113-case benchmark from top journals, shows state-of-the-art LLMs achieve at most 51% top-1 diagnostic accuracy, far below clinical-level competence.

  7. Baichuan-M2: Scaling Medical Capability with Large Verifier System

    cs.LG 2025-09 conditional novelty 5.0 of 10

    Baichuan-M2, a 32B medical LLM trained with a patient simulator and a clinical rubric generator as RL verifiers, reports state-of-the-art HealthBench scores (60.1 overall, 34.7 hard), ahead of all open-source models.

  8. CX-Mind: A Pioneering Multimodal Large Language Model for Interleaved Reasoning in Chest X-ray via Curriculum-Guided Reinforcement Learning

    cs.LG 2025-07 conditional novelty 5.0 of 10

    CX-Mind combines curriculum reinforcement learning and rule-based process rewards to train a chest X-ray vision-language model that produces interleaved think-answer reasoning and reports state-of-the-art results acro...

  9. Silence is Not Consensus: Disrupting Agreement Bias in Multi-Agent LLMs via Catfish Agent for Clinical Decision Making

    cs.CL 2025-05 conditional novelty 5.0 of 10

    A "catfish" agent that injects structured dissent into multi-agent LLM teams improves clinical question-answering accuracy by reducing premature consensus.

Pith tools