Pith. sign in

REVIEW 11 cited by

Med42-v2: A Suite of Clinical LLMs

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2408.06142 v1 pith:L37KPPPF submitted 2024-08-12 cs.CL cs.AI

classification cs.CLcs.AI
keywords clinicalmodelsmed42-v2llmsgenerichttpshuggingfacellama3
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Med42-v2 introduces a suite of clinical large language models (LLMs) designed to address the limitations of generic models in healthcare settings. These models are built on Llama3 architecture and fine-tuned using specialized clinical data. They underwent multi-stage preference alignment to effectively respond to natural prompts. While generic models are often preference-aligned to avoid answering clinical queries as a precaution, Med42-v2 is specifically trained to overcome this limitation, enabling its use in clinical settings. Med42-v2 models demonstrate superior performance compared to the original Llama3 models in both 8B and 70B parameter configurations and GPT-4 across various medical benchmarks. These LLMs are developed to understand clinical queries, perform reasoning tasks, and provide valuable assistance in clinical environments. The models are now publicly available at \href{https://huggingface.co/m42-health}{https://huggingface.co/m42-health}.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 11 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. MedJudgeRAG: Option-Wise Evidence Judgment with Dynamic Knowledge Graphs for Medical MCQA

    cs.IR 2026-07 conditional novelty 6.0 of 10

    MedJudgeRAG fine-tunes a medical MCQA model to emit per-option evidence verdicts and choose grounded, elimination, or parametric reasoning, improving over vanilla RAG by up to 17 accuracy points.

  2. Auditing Evidence Use in Medical LLM Diagnosis

    cs.AI 2026-07 conditional novelty 6.0 of 10

    Behavioral auditing of five medical LLMs shows most mined evidence interactions are clinically plausible, while adjudicated shortcut-like failures concentrate in negated or absent findings and clinically local evidence.

  3. Ask Before You Diagnose: Safe-Psych, a Sequential Evaluation Benchmark for LLMs in Psychiatry

    cs.CL 2026-05 conditional novelty 6.0 of 10

    A new benchmark of 1,048 psychiatric notes shows LLMs frequently diagnose too early under incomplete evidence, and safety prompting only trades premature diagnoses for excessive abstention.

  4. HIVMedQA: Benchmarking large language models for HIV medical decision support

    cs.CL 2025-07 conditional novelty 6.0 of 10

    The HIVMedQA benchmark finds that Gemini 2.5 Pro leads on most clinical reasoning dimensions, medical fine-tuning does not guarantee gains, and LLM judges are more informative than lexical overlap.

  5. Diagnosing our datasets: How does my language model learn clinical information?

    cs.CL 2025-05 conditional novelty 6.0 of 10

    The frequency of clinical jargon in pretraining corpora predicts how well open-source LLMs interpret that jargon, but hospital notes use abbreviations that appear only rarely online.

  6. C-MIG: Multi-view Information Gain-based Retrieval-Augmented Generation for Clinical Diagnosis Reasoning

    cs.AI 2026-05 unverdicted novelty 5.0 of 10

    C-MIG uses multi-view information gain from retrieved documents and refinements to supervise RAG-RL for clinical diagnosis, claiming top performance on four medical benchmarks.

  7. Quantized Large Language Models in Biomedical Natural Language Processing: Evaluation and Recommendation

    cs.CL 2025-09 conditional novelty 5.0 of 10

    Quantizing LLMs to 4 or 8 bits cuts GPU memory by up to 75% with generally small performance changes across eight biomedical NLP benchmarks.

  8. Evaluating LLMs Across Multi-Cognitive Levels: From Medical Knowledge Mastery to Scenario-Based Problem Solving

    cs.CL 2025-06 conditional novelty 5.0 of 10

    A new three-level medical benchmark shows LLM accuracy falls sharply from factual recall (up to 78%) to full clinical diagnosis (max 19%), with larger models and inference-time scaling helping most at intermediate levels.

  9. Spurious Correlations and Beyond: Understanding and Mitigating Shortcut Learning in SDOH Extraction with Large Language Models

    cs.CL 2025-05 conditional novelty 5.0 of 10

    LLMs infer drug use from alcohol or smoking mentions in clinical notes, producing gender-skewed false positives that prompting only partially corrects.

  10. Medicine on the Edge: Comparative Performance Analysis of On-Device LLMs for Clinical Reasoning

    cs.CL 2025-02 conditional novelty 5.0 of 10

    On-device LLMs reach about half the AMEGA clinical-reasoning score of large cloud models, with Med42 and Aloe highest (about 490/1000) and Phi-3 Mini the best accuracy-per-memory trade-off.

  11. Test-Time-Scaling for Zero-Shot Diagnosis with Visual-Language Reasoning

    cs.CV 2025-06 conditional novelty 4.0 of 10

    Sampling multiple VLM-generated visual descriptions and letting a text-only LLM vote on the diagnosis improves zero-shot medical image classification on three MedMNIST datasets.

Pith tools