Pith. sign in

REVIEW 33 cited by

MedReason: Eliciting Factual Medical Reasoning Steps in LLMs via Knowledge Graphs

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2504.00993 v2 pith:FQM3AFQ6 submitted 2025-04-01 cs.CL cs.AI

MedReason: Eliciting Factual Medical Reasoning Steps in LLMs via Knowledge Graphs

classification cs.CL cs.AI
keywords medicalreasoningdatasetmedreasonclinicalmodelsdatasetsdetailed
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

Medical tasks such as diagnosis and treatment planning require precise and complex reasoning, particularly in life-critical domains. Unlike mathematical reasoning, medical reasoning demands meticulous, verifiable thought processes to ensure reliability and accuracy. However, there is a notable lack of datasets that provide transparent, step-by-step reasoning to validate and enhance the medical reasoning ability of AI models. To bridge this gap, we introduce MedReason, a large-scale high-quality medical reasoning dataset designed to enable faithful and explainable medical problem-solving in large language models (LLMs). We utilize a structured medical knowledge graph (KG) to convert clinical QA pairs into logical chains of reasoning, or ``thinking paths'', which trace connections from question elements to answers via relevant KG entities. Each path is validated for consistency with clinical logic and evidence-based medicine. Our pipeline generates detailed reasoning for various medical questions from 7 medical datasets, resulting in a dataset of 32,682 question-answer pairs, each with detailed, step-by-step explanations. Experiments demonstrate that fine-tuning with our dataset consistently boosts medical problem-solving capabilities, achieving significant gains of up to 7.7% for DeepSeek-Ditill-8B. Our top-performing model, MedReason-8B, outperforms the Huatuo-o1-8B, a state-of-the-art medical reasoning model, by up to 4.2% on the clinical benchmark MedBullets. We also engage medical professionals from diverse specialties to assess our dataset's quality, ensuring MedReason offers accurate and coherent medical reasoning. Our data, models, and code is available at https://github.com/UCSC-VLAA/MedReason.

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 33 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. eCREAM-MedCorpus A Large-Scale Corpus of Clinical Notes for Italian

    cs.CL 2026-06 unverdicted novelty 8.0

    Introduces the largest freely available Italian clinical notes corpus with 4M notes and expert-annotated subset for a new CRF-filling benchmark.

  2. eCREAM-MedCorpus A Large-Scale Corpus of Clinical Notes for Italian

    cs.CL 2026-06 unverdicted novelty 8.0

    EDEN releases the largest freely available Italian clinical notes corpus (4M notes, 6k annotated) and proposes CRF-filling as a structured extraction benchmark with zero-shot baselines from Gemma models.

  3. EpiGraph: Building Generalists for Evidence-Intensive Epilepsy Reasoning in the Wild

    cs.AI 2026-05 conditional novelty 7.0

    EpiGraph creates a heterogeneous epilepsy knowledge graph that boosts LLM performance on clinical reasoning tasks by 30-41% in pharmacogenomics when used with Graph-RAG.

  4. EpiGraph: Building Generalists for Evidence-Intensive Epilepsy Reasoning in the Wild

    cs.AI 2026-05 unverdicted novelty 7.0

    EpiGraph is a new epilepsy knowledge graph with 24,324 entities and 32,009 triplets that improves LLM performance on clinical tasks by up to 41% when used in Graph-RAG.

  5. BioMedArena: An Open-source Toolkit for Building and Evaluating Biomedical Deep Research Agents

    cs.AI 2026-05 unverdicted novelty 7.0

    BioMedArena supplies a standardized open toolkit with 166 biomedical benchmarks, 75 tools, 6 harnesses, and 6 context strategies that improve 12 backbones and surpass prior SOTA by 15.01 points on average across 8 benchmarks.

  6. ClinFusion: A Vision-Centric Multimodal LLM System for Holistic Medical Understanding

    cs.CV 2026-07 conditional novelty 6.5

    A cascaded multi-encoder medical MLLM with native 3D fusion and RoI-grounded report metrics claims SOTA on most 2D/3D medical benchmarks and highest radiologist report rankings.

  7. LoMeVQA: A Comprehensive Benchmark for Longitudinal Medical VQA

    cs.CV 2026-07 conditional novelty 6.0

    A 206K multi-task longitudinal medical VQA benchmark shows current MLLMs fail at temporal reasoning, while fine-tuned MedLong-8B sets a strong baseline.

  8. Aligning Clinical Needs and AI Capabilities: A Survey on LLMs for Medical Reasoning

    cs.AI 2026-07 accept novelty 6.0

    A dual clinical-computational taxonomy for medical LLM reasoning plus a five-level 5k-sample benchmark showing specialists excel at diagnosis and general models at decision support/dialogue.

  9. Online Data Selection for Instruction Tuning via Gaussian Processes

    cs.LG 2026-06 unverdicted novelty 6.0

    GAIA models continuous utility with Gaussian processes across semantic space and applies fixed-share Hedge updates to achieve dynamic regret guarantees while outperforming baselines on three datasets.

  10. eCREAM-MedCorpus A Large-Scale Corpus of Clinical Notes for Italian

    cs.CL 2026-06 conditional novelty 6.0

    eCREAM-MedCorpus releases ~4M anonymized Italian ED clinical notes and a 6k-note 132-item CRF annotation set, with zero-shot Gemma/MedGemma CRF-filling baselines.

  11. Reasoning4Sciences: Bridging Reasoning Language Models to All Scientific Branches

    cs.AI 2026-05 unverdicted novelty 6.0

    Survey of RLM adoption in 28 disciplines reveals maturity disparities via a new assessment framework, with focus on development, evaluation, and public resources.

  12. Reasoning4Sciences: Bridging Reasoning Language Models to All Scientific Branches

    cs.AI 2026-05 unverdicted novelty 6.0

    A survey of RLM use in 28 disciplines reveals uneven adoption and introduces a maturity assessment framework showing larger gaps when limited to public resources.

  13. SDR: Set-Distance Rewards for Radiology Report Generation

    cs.AI 2026-05 unverdicted novelty 6.0

    Set-to-set distances on sentence embeddings provide a permutation-invariant reward signal that improves GRPO training and enables efficient test-time scaling for vision-language models generating chest X-ray reports.

  14. ClinSeekAgent: Automating Multimodal Evidence Seeking for Agentic Clinical Reasoning

    cs.CL 2026-05 unverdicted novelty 6.0

    ClinSeekAgent automates active multimodal evidence seeking for clinical reasoning, improving LLM performance on raw EHR and CXR tasks while enabling distillation into smaller models.

  15. RadThinking: A Dataset for Longitudinal Clinical Reasoning in Radiology

    cs.CV 2026-05 unverdicted novelty 6.0

    RadThinking releases a large longitudinal CT VQA dataset stratified into foundation perception questions, single-rule reasoning questions, and compositional multi-step chains grounded in clinical reporting standards f...

  16. CLR-voyance: Reinforcing Open-Ended Reasoning for Inpatient Clinical Decision Support with Outcome-Aware Rubrics

    cs.CL 2026-05 unverdicted novelty 6.0

    CLR-voyance reformulates inpatient reasoning as POMDP with clinician-validated outcome rubrics, yielding an 8B model that outperforms larger frontier models on the authors' new benchmark.

  17. BioMedArena: An Open-source Toolkit for Building and Evaluating Biomedical Deep Research Agents

    cs.AI 2026-05 conditional novelty 6.0

    BioMedArena releases a standardized toolkit with 147 biomedical benchmarks, 75 tools, and six harnesses that achieve SOTA results on eight tasks with a +15.03 percentage point average lift.

  18. MedSynapse-V: Bridging Visual Perception and Clinical Intuition via Latent Memory Evolution

    cs.CV 2026-04 unverdicted novelty 6.0

    MedSynapse-V proposes meta-query prior memorization, causal counterfactual refinement via RL, and dual-branch memory transition to evolve implicit diagnostic memories in medical VLMs and boost accuracy over chain-of-t...

  19. MedSynapse-V: Bridging Visual Perception and Clinical Intuition via Latent Memory Evolution

    cs.CV 2026-04 unverdicted novelty 6.0

    MedSynapse-V evolves latent diagnostic memories via meta queries, causal counterfactual refinement with RL, and dual-branch memory transition to outperform prior medical VLM methods in diagnostic accuracy.

  20. VeriLLMed: Interactive Visual Debugging of Medical Large Language Models with Knowledge Graphs

    cs.CL 2026-04 conditional novelty 6.0

    VeriLLMed uses biomedical knowledge graphs to turn medical LLM reasoning into comparable paths and automatically flags three recurring error types: relation, branch, and missing errors.

  21. Improving Clinical Diagnosis with Counterfactual Multi-Agent Reasoning

    cs.CL 2026-03 unverdicted novelty 6.0

    A new counterfactual multi-agent framework improves LLM diagnostic accuracy by quantifying confidence shifts from edited clinical findings and guiding specialist discussions.

  22. MASPOB: Bandit-Based Prompt Optimization for Multi-Agent Systems with Graph Neural Networks

    cs.LG 2026-03 conditional novelty 6.0

    MASPOB combines a GNN surrogate, LinUCB-style uncertainty, and coordinate ascent to optimize prompts in fixed-topology multi-agent LLM systems, beating AFlow and MIPRO on average across six benchmarks.

  23. Structured In-context Environment Scaling for Large Language Model Reasoning

    cs.CL 2025-09 conditional novelty 6.0

    SIE framework automatically constructs scalable, verifiable reasoning environments from structured data, improving in-domain performance and enabling generalization to out-of-domain math and logic tasks.

  24. Query-Conditioned Knowledge Alignment for Reliable Cross-System Medical Reasoning

    cs.AI 2026-05 conditional novelty 5.0

    QCEA reformulates entity alignment as a query-conditioned ranking task with semantic encoding, graph learning, and direction-aware transformation to handle context-dependent, asymmetric correspondences in medical know...

  25. ReMedi: Reasoner for Medical Clinical Prediction

    cs.CL 2026-05 unverdicted novelty 5.0

    ReMedi boosts LLM performance on EHR clinical predictions by up to 19.9% F1 through ground-truth-guided rationale regeneration and fine-tuning.

  26. MedSynapse-V: Bridging Visual Perception and Clinical Intuition via Latent Memory Evolution

    cs.CV 2026-04 unverdicted novelty 5.0

    MedSynapse-V proposes a latent diagnostic memory evolution framework using Meta Query, Causal Counterfactual Refinement, and Intrinsic Memory Transition to improve medical VLM diagnostic accuracy over chain-of-thought...

  27. VeriLLMed: Interactive Visual Debugging of Medical Large Language Models with Knowledge Graphs

    cs.CL 2026-04 unverdicted novelty 5.0

    VeriLLMed is an interactive visual debugging tool that maps LLM diagnostic reasoning to knowledge graphs to identify and categorize relation, branch, and missing errors.

  28. DeepER-Med: Advancing Deep Evidence-Based Research in Medicine Through Agentic AI

    cs.AI 2026-04 unverdicted novelty 5.0

    DeepER-Med introduces a three-module agentic AI workflow for evidence-based medical research that outperforms production platforms on a new expert-curated dataset of 100 questions and matches clinical recommendations ...

  29. Medical Reasoning with Large Language Models: A Survey and MR-Bench

    cs.CL 2026-03 accept novelty 5.0

    LLMs show strong exam performance on medical tasks but exhibit a clear gap in accuracy on authentic clinical decision-making as measured by the new MR-Bench benchmark and unified evaluations.

  30. InfiMed-ORBIT: Aligning LLMs on Open-Ended Complex Tasks via Rubric-Based Incremental Training

    cs.CL 2025-10 conditional novelty 5.0

    Rubric-based incremental RL with LLM-generated case-specific checklists lifts Qwen3-4B's HealthBench-Hard score from 7.0 to 27.5 with 2k samples, and improves InfoBench instruction-following from 42.0 to 82.9.

  31. The Path to Self-Evolving Clinical Systems: Scaling Medical Agents from Assistance to Autonomy

    cs.AI 2026-07 conditional novelty 4.5

    Medical agents should be scaled mainly by richer clinical environments and self-evolution loops, not parameter growth alone, under a three-level autonomy taxonomy.

  32. Reasoning4Sciences: Bridging Reasoning Language Models to All Scientific Branches

    cs.AI 2026-05 unverdicted novelty 4.0

    A survey of reasoning language model adoption across 28 ERC scientific disciplines finds large maturity gaps, especially when only public resources are counted.

  33. MedSynapse-V: Bridging Visual Perception and Clinical Intuition via Latent Memory Evolution

    cs.CV 2026-04 unverdicted novelty 4.0

    MedSynapse-V proposes a latent memory evolution framework with meta-query prior retrieval, causal counterfactual refinement via RL, and intrinsic memory transition to improve diagnostic accuracy over chain-of-thought ...