Pith. sign in

REVIEW 9 cited by

A Survey of Multimodal Large Language Model from A Data-centric Perspective

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2405.16640 v2 pith:3D6QE4XH submitted 2024-05-26 cs.AI cs.CLcs.CVcs.MM

classification cs.AIcs.CLcs.CVcs.MM
keywords mllmsdatalanguagelargemodelsmultimodalsurveydata-centric
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Multimodal large language models (MLLMs) enhance the capabilities of standard large language models by integrating and processing data from multiple modalities, including text, vision, audio, video, and 3D environments. Data plays a pivotal role in the development and refinement of these models. In this survey, we comprehensively review the literature on MLLMs from a data-centric perspective. Specifically, we explore methods for preparing multimodal data during the pretraining and adaptation phases of MLLMs. Additionally, we analyze the evaluation methods for the datasets and review the benchmarks for evaluating MLLMs. Our survey also outlines potential future research directions. This work aims to provide researchers with a detailed understanding of the data-driven aspects of MLLMs, fostering further exploration and innovation in this field.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 9 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Agent Security Needs Redefinition through a Holistic Framework

    cs.CR 2026-07 conditional novelty 6.0 of 10

    Agent security should be redefined around four contextual authorization properties instead of the content of the action performed.

  2. BYO-Eval: Build Your Own Dataset for Fine-Grained Visual Assessment of Multimodal Language Models

    cs.CV 2025-06 conditional novelty 6.0 of 10

    A Blender-based diagnostic toolkit that tests VLMs on fine-grained visual skills by varying one visual attribute at a time, exposing failure modes that coarse benchmarks miss.

  3. Rethinking Machine Unlearning in Image Generation Models

    cs.AI 2025-06 conditional novelty 6.0 of 10

    A new taxonomy and multi-aspect evaluation framework for image generation unlearning, with a curated dataset, shows that ten existing unlearning methods perform poorly on preservation and robustness.

  4. SynthRL: Scaling Visual Reasoning with Verifiable Data Synthesis

    cs.LG 2025-06 conditional novelty 6.0 of 10

    SynthRL synthesizes harder, answer-preserving visual math questions from easy seed questions and reports small but mixed out-of-domain RLVR gains for Qwen2.5-VL-7B.

  5. ID-Align: RoPE-Conscious Position Remapping for Dynamic High-Resolution Adaptation in Vision-Language Models

    cs.CV 2025-05 conditional novelty 6.0 of 10

    ID-Align improves high-resolution VLM performance by reusing thumbnail position IDs for high-resolution image tokens, yielding small but positive gains on several benchmarks.

  6. Failure-Informed Image Self-Augmentation for Multimodal Large Language Model Self-Improvement

    cs.AI 2026-08 conditional novelty 5.0 of 10

    FISA improves MLLM visual question answering by generating and filtering augmented images based on the model's own failure cases.

  7. DecoupleMix: Decoupled Ratio Search and Convex Allocation for Scalable VLM Data Recipes

    cs.CV 2026-07 conditional novelty 5.0 of 10

    Decoupling inter-class ratio search from intra-class convex allocation yields VLM data recipes that beat stacking and transfer from small proxies to larger scales.

  8. DataPrep-Bench: Benchmarking LLMs as Training Data Preparators

    cs.LG 2026-05 conditional novelty 5.0 of 10

    DataPrep-Bench jointly benchmarks data construction and data-quality evaluation for LLMs across six domains with downstream fine-tuning performance as ground truth.

  9. OpenWorldLib: A Unified Codebase and Definition of Advanced World Models

    cs.CV 2026-04 unverdicted novelty 4.0 of 10

    OpenWorldLib offers a standardized codebase and definition for world models that combine perception, interaction, and memory to understand and predict the world.

Pith tools