REVIEW 11 cited by
Med42-v2: A Suite of Clinical LLMs
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Med42-v2 introduces a suite of clinical large language models (LLMs) designed to address the limitations of generic models in healthcare settings. These models are built on Llama3 architecture and fine-tuned using specialized clinical data. They underwent multi-stage preference alignment to effectively respond to natural prompts. While generic models are often preference-aligned to avoid answering clinical queries as a precaution, Med42-v2 is specifically trained to overcome this limitation, enabling its use in clinical settings. Med42-v2 models demonstrate superior performance compared to the original Llama3 models in both 8B and 70B parameter configurations and GPT-4 across various medical benchmarks. These LLMs are developed to understand clinical queries, perform reasoning tasks, and provide valuable assistance in clinical environments. The models are now publicly available at \href{https://huggingface.co/m42-health}{https://huggingface.co/m42-health}.
Forward citations
Cited by 11 Pith papers
-
MedJudgeRAG: Option-Wise Evidence Judgment with Dynamic Knowledge Graphs for Medical MCQA
MedJudgeRAG fine-tunes a medical MCQA model to emit per-option evidence verdicts and choose grounded, elimination, or parametric reasoning, improving over vanilla RAG by up to 17 accuracy points.
-
Auditing Evidence Use in Medical LLM Diagnosis
Behavioral auditing of five medical LLMs shows most mined evidence interactions are clinically plausible, while adjudicated shortcut-like failures concentrate in negated or absent findings and clinically local evidence.
-
Ask Before You Diagnose: Safe-Psych, a Sequential Evaluation Benchmark for LLMs in Psychiatry
A new benchmark of 1,048 psychiatric notes shows LLMs frequently diagnose too early under incomplete evidence, and safety prompting only trades premature diagnoses for excessive abstention.
-
HIVMedQA: Benchmarking large language models for HIV medical decision support
The HIVMedQA benchmark finds that Gemini 2.5 Pro leads on most clinical reasoning dimensions, medical fine-tuning does not guarantee gains, and LLM judges are more informative than lexical overlap.
-
Diagnosing our datasets: How does my language model learn clinical information?
The frequency of clinical jargon in pretraining corpora predicts how well open-source LLMs interpret that jargon, but hospital notes use abbreviations that appear only rarely online.
-
C-MIG: Multi-view Information Gain-based Retrieval-Augmented Generation for Clinical Diagnosis Reasoning
C-MIG uses multi-view information gain from retrieved documents and refinements to supervise RAG-RL for clinical diagnosis, claiming top performance on four medical benchmarks.
-
Quantized Large Language Models in Biomedical Natural Language Processing: Evaluation and Recommendation
Quantizing LLMs to 4 or 8 bits cuts GPU memory by up to 75% with generally small performance changes across eight biomedical NLP benchmarks.
-
Evaluating LLMs Across Multi-Cognitive Levels: From Medical Knowledge Mastery to Scenario-Based Problem Solving
A new three-level medical benchmark shows LLM accuracy falls sharply from factual recall (up to 78%) to full clinical diagnosis (max 19%), with larger models and inference-time scaling helping most at intermediate levels.
-
Spurious Correlations and Beyond: Understanding and Mitigating Shortcut Learning in SDOH Extraction with Large Language Models
LLMs infer drug use from alcohol or smoking mentions in clinical notes, producing gender-skewed false positives that prompting only partially corrects.
-
Medicine on the Edge: Comparative Performance Analysis of On-Device LLMs for Clinical Reasoning
On-device LLMs reach about half the AMEGA clinical-reasoning score of large cloud models, with Med42 and Aloe highest (about 490/1000) and Phi-3 Mini the best accuracy-per-memory trade-off.
-
Test-Time-Scaling for Zero-Shot Diagnosis with Visual-Language Reasoning
Sampling multiple VLM-generated visual descriptions and letting a text-only LLM vote on the diagnosis improves zero-shot medical image classification on three MedMNIST datasets.
Discussion (0). Continue with ORCID to comment.