Pith. sign in

REVIEW 25 cited by

MM-LLMs: Recent Advances in MultiModal Large Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2401.13601 v5 pith:QOTNTPGB submitted 2024-01-24 cs.CL

classification cs.CL
keywords mm-llmsmodelstrainingformulationslanguagelargellmsmultimodal
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

In the past year, MultiModal Large Language Models (MM-LLMs) have undergone substantial advancements, augmenting off-the-shelf LLMs to support MM inputs or outputs via cost-effective training strategies. The resulting models not only preserve the inherent reasoning and decision-making capabilities of LLMs but also empower a diverse range of MM tasks. In this paper, we provide a comprehensive survey aimed at facilitating further research of MM-LLMs. Initially, we outline general design formulations for model architecture and training pipeline. Subsequently, we introduce a taxonomy encompassing 126 MM-LLMs, each characterized by its specific formulations. Furthermore, we review the performance of selected MM-LLMs on mainstream benchmarks and summarize key training recipes to enhance the potency of MM-LLMs. Finally, we explore promising directions for MM-LLMs while concurrently maintaining a real-time tracking website for the latest developments in the field. We hope that this survey contributes to the ongoing advancement of the MM-LLMs domain.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 25 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. VISA: VLM-Guided Instance Semantic Auditing for 3D Occupancy World Models

    cs.CV 2026-06 unverdicted novelty 7.0 of 10

    VISA improves closed-set 3D occupancy mIoU on nuScenes by using VLM instance audits as reliability-weighted semantic supervisors during training of existing world models.

  2. Rectified Schr\"odinger Bridge Matching for Few-Step Visual Navigation

    cs.RO 2026-04 unverdicted novelty 7.0 of 10

    RSBM exploits velocity field invariance across regularization levels to achieve over 94% cosine similarity and 92% success in visual navigation using only 3 integration steps.

  3. Just Noticeable Difference for Large Multimodal Models

    cs.CV 2025-07 conditional novelty 7.0 of 10

    Large multimodal models have measurable just-noticeable-difference thresholds, and most lag far behind humans in seeing small image changes.

  4. CARE-X: Towards Clinically Useful Radiology VLMs with Auxiliary Supervision, Reward-Aligned Learning, and Tool-Augmented Measurement

    cs.CV 2026-08 conditional novelty 6.0 of 10

    A chest X-ray VLM co-trained with classification and grounding heads, tuned with DAPO reinforcement learning, and augmented with deterministic measurement tools outperforms prior radiology VLMs on report generation, V...

  5. Personalize Your Large Vision-language Models With In-context Prompt Tuning

    cs.CV 2026-05 unverdicted novelty 6.0 of 10

    ICPT converts a few reference images of a personalized concept into an adaptive-length visual prompt plus a label embedding, letting a frozen LVLM add and reason about multiple concepts on the fly.

  6. LaRe: Latent Refocusing for Multimodal Reasoning

    cs.CV 2025-11 reject novelty 6.0 of 10

    LaRe performs iterative visual refocusing in latent space and reports accuracy gains with fewer tokens, but its main experiments compare against baselines trained with less data.

  7. E-THER: A Multimodal Dataset for Empathic AI -- Towards Emotional Mismatch Awareness

    cs.HC 2025-09 reject novelty 6.0 of 10

    E-THER is a small annotated therapy-video dataset for verbal-visual incongruence, but the claimed empathy gains are supported mainly by author-built keyword metrics with statistical inconsistencies.

  8. LLaPa: A Vision-Language Model Framework for Counterfactual-Aware Procedural Planning

    cs.CL 2025-07 conditional novelty 6.0 of 10

    A VLM-based planner with task-oriented segmentation reranking and a clause-level condition retriever reports state-of-the-art scores on ActPlan-1K and ALFRED, though the reported ablation numbers are internally inconsistent.

  9. Graft: Integrating the Domain Knowledge via Efficient Parameter Synergy for MLLMs

    cs.CL 2025-06 conditional novelty 6.0 of 10

    Graft merges two domain-specialized multimodal models by combining channel-wise gating, entropy-based global weighting, and an activation compatibility score to improve fusion without retraining.

  10. Fuzzing: Randomness? Reasoning! Efficient Directed Fuzzing via Large Language Models

    cs.SE 2025-06 conditional novelty 6.0 of 10

    LLM-generated reachable seeds and bug-specific mutators speed up directed fuzzing on 14 known CVEs, with several bugs found in under a minute.

  11. CausalVQA: A Physically Grounded Causal Reasoning Benchmark for Video Models

    cs.CV 2025-06 conditional novelty 6.0 of 10

    CausalVQA provides 793 paired real-video causal reasoning questions on which the best multimodal model scores 61.66% versus 84.78% for humans, with the largest gaps on anticipation and hypothetical questions.

  12. ID-Align: RoPE-Conscious Position Remapping for Dynamic High-Resolution Adaptation in Vision-Language Models

    cs.CV 2025-05 conditional novelty 6.0 of 10

    ID-Align improves high-resolution VLM performance by reusing thumbnail position IDs for high-resolution image tokens, yielding small but positive gains on several benchmarks.

  13. TUNA: Comprehensive Fine-grained Temporal Understanding Evaluation on Dense Dynamic Videos

    cs.CV 2025-05 conditional novelty 6.0 of 10

    TUNA introduces a 1,000-video benchmark with dense temporal captions and 1,432 multiple-choice questions, and finds that current video LMMs are weakest at camera motion, action sequences, and multi-subject scenes.

  14. LLMs for LLMs: A Structured Prompting Methodology for Long Legal Documents

    cs.AI 2025-09 reject novelty 5.0 of 10

    On CUAD legal contracts, a prompt-engineered QWEN-2 pipeline with chunking and two answer-selection heuristics reportedly outperforms the fine-tuned DeBERTa-large baseline by about 9%, reaching claimed state-of-the-ar...

  15. GoVector: An I/O-Efficient Caching Strategy for High-Dimensional Vector Nearest Neighbor Search

    cs.DB 2025-08 reject novelty 5.0 of 10

    A GoVector abstract claims a hybrid static/dynamic cache plus disk reordering improves disk-based ANN search, but the manuscript body is an unrelated chart/table benchmark paper.

  16. MapAgent: Trajectory-Constructed Memory-Augmented Planning for Mobile Task Automation

    cs.HC 2025-07 conditional novelty 5.0 of 10

    A memory-augmented LLM planner that stores and retrieves page-level summaries from past trajectories improves success rates on mobile GUI task benchmarks.

  17. EdgeLoRA: An Efficient Multi-Tenant LLM Serving System on Edge Devices

    cs.DC 2025-07 conditional novelty 5.0 of 10

    EdgeLoRA combines automatic adapter routing, LRU caching with a memory pool, and grouped LoRA batching to serve thousands of LoRA adapters on edge devices with up to 4x higher throughput than llama.cpp.

  18. From Street Views to Urban Science: Discovering Road Safety Factors with Multimodal Large Language Models

    cs.LG 2025-06 reject novelty 5.0 of 10

    UrbanX uses LLM-generated visual questions, MLLM answers, and linear regression to predict crash rates, claiming better performance than ResNet/ViT while keeping features interpretable.

  19. NextG-GPT: Leveraging GenAI for Advancing Wireless Networks and Communication Research

    cs.ET 2025-05 conditional novelty 5.0 of 10

    A RAG-enhanced LLM assistant for wireless research testbeds is built and evaluated, with LLaMa3.1-70B scoring best, though the abstract mislabels a faithfulness score as correctness.

  20. Enhancing Large Multimodal Models with Adaptive Sparsity and KV Cache Compression

    cs.AI 2025-07 conditional novelty 4.0 of 10

    A TPE-guided search over layer-wise pruning ratios and KV cache bit-widths compresses LLaVA-1.5 7B/13B with small accuracy loss, outperforming Wanda and SparseGPT on most tested benchmarks.

  21. HiBerNAC: Hierarchical Brain-emulated Robotic Neural Agent Collective for Disentangling Complex Manipulation

    cs.RO 2025-06 reject novelty 4.0 of 10

    HiBerNAC, a multi-agent 'brain-inspired' planner layered on a reactive VLA, is claimed to cut long-horizon task time by 23% and reach 12-31% success where VLA baselines fail, but the supporting data are inconsistent.

  22. Towards Efficient Multi-LLM Inference: Characterization and Analysis of LLM Routing and Hierarchical Techniques

    cs.LG 2025-06 unverdicted novelty 4.0 of 10

    A survey of LLM routing and hierarchical inference techniques that proposes an unvalidated unified evaluation metric called the Inference Efficiency Score.

  23. Towards LLM-Centric Multimodal Fusion: A Survey on Integration Strategies and Techniques

    cs.CL 2025-06 conditional novelty 4.0 of 10

    Multimodal LLMs can be organized by fusion mechanism, fusion level, representation paradigm, and training paradigm, with 125 models classified accordingly.

  24. Chain-of-Thought for Autonomous Driving: A Comprehensive Survey and Future Prospects

    cs.RO 2025-05 conditional novelty 4.0 of 10

    A survey that classifies chain-of-thought methods for autonomous driving into modular, logical, and reflective pipelines, and proposes three evolutionary stages from direct prompting to reinforcement learning.

  25. R-Genie: Reasoning-Guided Generative Image Editing

    cs.CV 2025-05 conditional novelty 4.0 of 10

    R-Genie couples a multimodal LLM with a discrete diffusion model to perform image edits that require commonsense reasoning, and introduces a 1,070-triple benchmark called REditBench.

Pith tools