REVIEW 21 cited by
Large Multimodal Agents: A Survey
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Large language models (LLMs) have achieved superior performance in powering text-based AI agents, endowing them with decision-making and reasoning abilities akin to humans. Concurrently, there is an emerging research trend focused on extending these LLM-powered AI agents into the multimodal domain. This extension enables AI agents to interpret and respond to diverse multimodal user queries, thereby handling more intricate and nuanced tasks. In this paper, we conduct a systematic review of LLM-driven multimodal agents, which we refer to as large multimodal agents ( LMAs for short). First, we introduce the essential components involved in developing LMAs and categorize the current body of research into four distinct types. Subsequently, we review the collaborative frameworks integrating multiple LMAs , enhancing collective efficacy. One of the critical challenges in this field is the diverse evaluation methods used across existing studies, hindering effective comparison among different LMAs . Therefore, we compile these evaluation methodologies and establish a comprehensive framework to bridge the gaps. This framework aims to standardize evaluations, facilitating more meaningful comparisons. Concluding our review, we highlight the extensive applications of LMAs and propose possible future research directions. Our discussion aims to provide valuable insights and guidelines for future research in this rapidly evolving field. An up-to-date resource list is available at https://github.com/jun0wanan/awesome-large-multimodal-agents.
Forward citations
Cited by 21 Pith papers
-
A Comprehensive Study of Implementation Bugs in Multi-modal Agents
First systematic taxonomy of 158 multi-modal agent bugs plus a runtime analyzer that recovers most open issues and surfaces 31 new ones.
-
MentalThink: Shaping Thoughts in Mental SVG World
MLLMs that generate and render SVG sketches as multi-turn intermediate reasoning steps reach 55.1% on VSIBench and 76.0% on MindCube, far above the Qwen2.5-VL-7B backbone.
-
FieldWorkArena: Agentic AI Benchmark for Real Field Work Tasks
A new benchmark dataset and evaluation framework for testing multimodal AI agents on real field work tasks derived from on-site data and worker interviews.
-
OmniPresent: Generating Coherent Presentation Suites from Scientific Papers
A multi-agent HTML pipeline with shared knowledge and cross-artifact verify-and-repair generates coherent poster/slides/video/page suites from papers and beats specialized baselines on OmniPreBench.
-
Auto-Scaling Heterogeneous Neural Processing Units for Energy and Cost-Efficient LLM Serving
NeuScale routes LLM inference requests to the most energy/cost-efficient configuration of heterogeneous NPU chips using roofline allocation and runtime auto-scaling.
-
Human-Agent Collaborative Paper-to-Page Crafting
AutoPage converts academic PDFs into interactive project webpages via a planner-generator-checker pipeline, and outperforms end-to-end LLMs on the new PageBench benchmark.
-
Can VLMs Recall Factual Associations From Visual References?
Vision-language models recall facts much better when the entity is named in text than when the same entity appears only in an image, and hidden-state probes can detect many of these recall failures.
-
Separating Shared and Domain-Specific LoRAs for Multi-Domain Learning
Shared and domain-specific LoRAs are constrained to the column and left null spaces of pretrained weights, but experimental benefits are mixed.
-
MLA-Trust: Benchmarking Trustworthiness of Multimodal LLM Agents in GUI Environments
MLA-Trust introduces 34 tasks and an evaluation toolbox showing that GUI-interacting multimodal agents are substantially less trustworthy than static multimodal chat models.
-
Harnessing Vision Models for Time Series Analysis: A Survey
A survey organizing existing methods that encode time series as images and apply vision models, with a dual-view taxonomy of imaging and modeling approaches.
-
AgentX: Towards Orchestrating Robust Agentic Workflow Patterns with FaaS-hosted MCP Services
AgentX, a stage-planner-executor agentic workflow, matches ReAct and Magentic-One on output quality in three applications while cutting token use on web search, and MCP servers deployed on AWS Lambda run at negligible...
-
SasAgent: Multi-Agent AI System for Small-Angle Scattering Data Analysis
SasAgent connects a large language model to SasView tools through four agents, letting users calculate SLDs, generate synthetic scattering curves, and fit experimental SAS data from text prompts.
-
Efficient Deployment of Vision-Language Models on Mobile Devices: A Case Study on OnePlus 13R
A side-by-side mobile benchmark shows VLM runtimes on a OnePlus 13R leave accelerators idle, push CPUs to thermal limits, and achieve order-of-magnitude power savings only when the GPU handles image and language kernels.
-
H2HTalk: Evaluating Large Language Models as Emotional Companion
H2HTalk is a new 4,650-scenario benchmark that scores LLM emotional companions on dialogue, memory, and itinerary planning, and finds models struggle with implicit needs and long-horizon memory.
-
Graphs Meet AI Agents: Taxonomy, Progress, and Future Opportunities
A survey that groups graph-empowered AI agent research into planning, execution, memory, and multi-agent coordination, plus agents-for-graphs and applications.
-
Do BERT-Like Bidirectional Models Still Perform Better on Text Classification in the Era of LLMs?
Fine-tuned BERT-like models outperform zero-shot and internal-state LLM methods on four of six challenging text classification datasets.
-
A Phased Development Framework Enabling Islanded Operation of Sustainable AI Data Centers With Onsite Grid-Following and Grid-Forming Energy Architectures
An islanded-first, phased construction framework for AI data centers — on-site gas turbines plus grid-forming batteries until grid interconnection matures — is shown via EMT simulation to track 300 MW AI training load swings.
-
SGTA: Scene-Graph Based Multi-Modal Traffic Agent for Video Understanding
A ReAct-style multi-modal agent answers traffic video questions by interleaving Cypher queries over a constructed scene graph with VLM calls on cropped frames.
-
Continual Learning for Generative AI: From LLMs to MLLMs and Beyond
A survey that categorizes continual learning methods for generative models into architecture-based, regularization-based, and replay-based paradigms across four model families.
-
A Call for Collaborative Intelligence: Why Human-Agent Systems Should Precede AI Autonomy
A position paper arguing that LLM-based human-agent systems, not fully autonomous agents, should be the immediate goal for AI development.
-
AI Agents and Agentic AI-Navigating a Plethora of Concepts for Future Manufacturing
A survey that clarifies the overlapping concepts of AI agents, LLM agents, and Agentic AI, and maps their potential roles in smart manufacturing.
Discussion (0). Continue with ORCID to comment.