REVIEW 11 cited by
MMedAgent: Learning to Use Medical Tools with Multi-modal Agent
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Multi-Modal Large Language Models (MLLMs), despite being successful, exhibit limited generality and often fall short when compared to specialized models. Recently, LLM-based agents have been developed to address these challenges by selecting appropriate specialized models as tools based on user inputs. However, such advancements have not been extensively explored within the medical domain. To bridge this gap, this paper introduces the first agent explicitly designed for the medical field, named \textbf{M}ulti-modal \textbf{Med}ical \textbf{Agent} (MMedAgent). We curate an instruction-tuning dataset comprising six medical tools solving seven tasks across five modalities, enabling the agent to choose the most suitable tools for a given task. Comprehensive experiments demonstrate that MMedAgent achieves superior performance across a variety of medical tasks compared to state-of-the-art open-source methods and even the closed-source model, GPT-4o. Furthermore, MMedAgent exhibits efficiency in updating and integrating new medical tools. Codes and models are all available.
Forward citations
Cited by 11 Pith papers
-
Evaluating Agentic Harness Systems for Autonomous Computational Pathology
Across 369 agent trajectories on 41 pathology tasks, formal end-to-end workflow completion was rare (10/369), with tool execution, result binding, and reflection far weaker than planning and reporting.
-
Aligning Clinical Needs and AI Capabilities: A Survey on LLMs for Medical Reasoning
A dual clinical-computational taxonomy for medical LLM reasoning plus a five-level 5k-sample benchmark showing specialists excel at diagnosis and general models at decision support/dialogue.
-
Beyond the Final Answer: Evaluating the Reasoning Trajectories of Tool-Augmented Agents
TRACE uses an evidence bank to score tool-augmented LLM agents on efficiency, hallucination, and adaptivity without ground-truth trajectories.
-
Improving Alignment in LVLMs with Debiased Self-Judgment
A contrastive self-judgment score that subtracts a model's image-free confidence from its visual confidence is used to guide decoding, safety moderation, and DPO training, improving hallucination and safety metrics ac...
-
WSI-Agents: A Collaborative Multi-Agent System for Multi-Modal Whole Slide Image Analysis
A route, verify, and summarize agent system uses existing pathology models and a knowledge base to select the best whole-slide image answer.
-
AT-CXR: Uncertainty-Aware Agentic Triage for Chest X-rays
An uncertainty-aware agentic router with abstention improves pulmonary edema triage on a small balanced chest x-ray subset, but the selective-prediction gains rest on under-specified evaluation.
-
AURA: A Multi-Modal Medical Agent for Understanding, Reasoning & Annotation
AURA is an agentic system that orchestrates chest X-ray tools to produce self-evaluated visual and textual explanations via counterfactual image generation.
-
A Comprehensive Survey of Electronic Health Record Modeling: From Deep Learning Approaches to Large Language Models
A survey that taxonomizes EHR modeling research into data-centric, architectural, learning-focused, multimodal, and LLM-based categories, with datasets and metrics.
-
ADAgent: LLM Agent for Alzheimer's Disease Analysis with Collaborative Coordinator
A large language model agent that plans, runs, and aggregates multiple off-the-shelf Alzheimer's imaging models achieves small but consistent accuracy gains over individual models on the ADNI dataset.
-
A Call for Collaborative Intelligence: Why Human-Agent Systems Should Precede AI Autonomy
A position paper arguing that LLM-based human-agent systems, not fully autonomous agents, should be the immediate goal for AI development.
-
MRGAgents: A Multi-Agent Framework for Improved Medical Report Generation with Med-LVLMs
MRGAgents fine-tunes one agent per chest X-ray disease and merges their sentences, reporting higher text metrics, but its evaluation uses oracle disease sentences and is not end-to-end.
Discussion (0). Continue with ORCID to comment.