Pith. sign in

REVIEW 9 cited by

FineMedLM-o1: Enhancing Medical Knowledge Reasoning Ability of LLM from Supervised Fine-Tuning to Test-Time Training

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2501.09213 v3 pith:OUNK4VOG submitted 2025-01-16 cs.CL

classification cs.CL
keywords medicalreasoningdatafinemedlm-o1capabilitiesdeepdiagnosisdialogue
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Recent advancements in large language models (LLMs) have shown promise in medical applications such as disease diagnosis and treatment planning. However, most existing medical LLMs struggle with the deep reasoning required for complex medical problems, such as differential diagnosis and medication recommendations. We propose FineMedLM-o1, which leverages high-quality medical synthetic data and long-form reasoning data for Supervised Fine-Tuning (SFT) and Direct Preference Optimization (DPO), enabling advanced dialogue and deep reasoning capabilities. Additionally, we introduce Test-Time Training (TTT) in the medical domain for the first time, facilitating domain adaptation and ensuring reliable, accurate reasoning. Experimental results demonstrate that FineMedLM-o1 achieves a 23% average performance improvement over prior models on key medical benchmarks. Furthermore, the introduction of TTT provides an additional 14% performance boost, highlighting its effectiveness in enhancing medical reasoning capabilities. To support this process, we also propose a novel method for synthesizing medical dialogue. Compared to other open-source datasets, our dataset stands out as superior in both quality and complexity. The project and data will be released on GitHub.

Discussion (0). Sign in to comment.

Forward citations

Cited by 9 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. MetaboLLM: a metabolomics-specialized large language model for biochemical knowledge integration and predictive metabolite graph construction

    cs.LG 2026-08 conditional novelty 6.0 of 10

    A metabolomics-specialized LLM that generates biochemical descriptions, converted into a metabolite graph, improved patient-level metabolomics classification over standard baselines in two datasets.

  2. CLARity: Reasoning Consistency Alone Can Teach Reinforced Experts

    cs.CL 2025-10 conditional novelty 6.0 of 10

    A consistency reward parsed by a 7B LLM, plus a two-stage refine-then-monitor pipeline and reformulated easy questions, improves MCQ-RL accuracy-with-consistency in law and medicine (58.9 vs 51.4 average Acc+).

  3. Kernel-Based Sparse Additive Nonlinear Model Structure Detection through a Linearization Approach

    eess.SY 2025-08 unverdicted novelty 6.0 of 10

    The paper uses an LPV linearization and sparse RKHS estimators to detect the additive structure of continuous-time nonlinear models.

  4. Med-U1: Incentivizing Unified Medical Reasoning in LLMs via Large-scale Reinforcement Learning

    cs.CL 2025-06 conditional novelty 6.0 of 10

    A pure reinforcement learning recipe with mixed rule-based rewards and length control improves Qwen2.5-based models across diverse medical QA formats.

  5. How Far Are We from Optimal Reasoning Efficiency?

    cs.CL 2025-06 conditional novelty 6.0 of 10

    The authors define a reasoning efficiency frontier and a gap metric (REG), then train models with REO-RL to shrink the gap by at least 50% with only small accuracy losses.

  6. DiagnosisArena: Benchmarking Diagnostic Reasoning for Large Language Models

    cs.CL 2025-05 conditional novelty 6.0 of 10

    DiagnosisArena, a 1,113-case benchmark from top journals, shows state-of-the-art LLMs achieve at most 51% top-1 diagnostic accuracy, far below clinical-level competence.

  7. Organ-Agents: Virtual Human Physiology Simulator via LLMs

    cs.LG 2025-08 unverdicted novelty 5.0 of 10

    A nine-agent LLM framework trained on sepsis patient records generates plausible multi-organ trajectories, reproduces critical events, and simulates alternative treatments, validated on held-out and external ICU patients.

  8. One Missing Piece for Open-Source Reasoning Models: A Dataset to Mitigate Cold-Starting Short CoT LLMs in RL

    cs.CL 2025-06 conditional novelty 5.0 of 10

    A pipeline converts short-CoT LLM outputs into o1-style long chain-of-thought rationales using 1K seed reasoning flows, and SFT on the resulting dataset improves downstream RLVR cold-start.

  9. Improving Medical Reasoning with Curriculum-Aware Reinforcement Learning

    cs.AI 2025-05 conditional novelty 4.0 of 10

    A curriculum-based GRPO schedule that first trains on close-ended medical VQA and then on open-ended VQA improves benchmark scores over joint training and vanilla RL, though the open-ended metric is the training objec...

Pith tools