Pith. sign in

REVIEW 11 cited by

Efficient Multimodal Learning from Data-centric Perspective

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2402.11530 v3 pith:CJZGWNA3 submitted 2024-02-18 cs.CV

classification cs.CV
keywords trainingdatalanguagemllmsmodelsmultimodalbunnyefficient
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Multimodal Large Language Models (MLLMs) have demonstrated notable capabilities in general visual understanding and reasoning tasks. However, their deployment is hindered by substantial computational costs in both training and inference, limiting accessibility to the broader research and user communities. A straightforward solution is to leverage smaller pre-trained vision and language models, which inevitably cause significant performance drops. In this paper, we demonstrate the possibility of training a smaller but better MLLM with high-quality training data. Specifically, we introduce Bunny, a family of lightweight MLLMs with flexible vision and language backbones for efficient multimodal learning from selected training data. Experiments show that our Bunny-4B/8B outperforms the state-of-the-art large MLLMs on multiple benchmarks. We expect that this work can provide the community with a clean and flexible open-source tool for further research and development. The code, models, and data can be found in https://github.com/BAAI-DCAI/Bunny.

Discussion (0). Sign in to comment.

Forward citations

Cited by 11 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 13 citations worldwide. Full citation record

  1. OLMoASR: Open Models and Data for Training Robust Speech Recognition Models

    cs.SD 2025-08 conditional novelty 7.0 of 10

    An open 1M-hour English speech dataset plus Whisper-architecture models trained on it match Whisper's word error rates on short and long-form benchmarks.

  2. RoboInter1.5: A Holistic Intermediate Representation Suite for Embodied World Modeling and Robotic Manipulation

    cs.RO 2026-07 conditional novelty 6.0 of 10

    Dense per-frame intermediate representations (traces, masks, grasp poses, subtasks) improve embodied VQA, VLA action generation, and world-model video prediction in the new 230k-episode RoboInter-Data suite.

  3. Leaderless Collective Motion in Affine Formation Control over the Complex Plane

    cs.RO 2026-04 unverdicted novelty 6.0 of 10

    Modifying Laplacian weights yields leaderless affine collective motion of planar robot swarms, with closed-form eigenvectors/eigenvalues designed via complex-plane analysis.

  4. Modality Gap-Driven Subspace Alignment Training Paradigm For Multimodal Large Language Models

    cs.CV 2026-02 unverdicted novelty 6.0 of 10

    ReAlign corrects the modality gap in unpaired data to let MLLMs learn visual distributions from text alone before instruction tuning, reducing dependence on expensive paired corpora.

  5. LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs

    cs.CV 2025-07 conditional novelty 6.0 of 10

    LLaVA-SP adds six spatial tokens produced by multi-scale cropping or pooling and cross-attention to MLLMs, improving 10/11 benchmarks over LLaVA-1.5 with nearly unchanged latency.

  6. AutoJudger: An Agent-Driven Framework for Efficient Benchmarking of MLLMs

    cs.CL 2025-05 conditional novelty 6.0 of 10

    An agent-driven framework adaptively selects a small subset of benchmark questions for MLLMs, preserving over 90% ranking accuracy with roughly 4-5% of the data.

  7. TimeCausality: Evaluating the Causal Ability in Time Dimension for Vision Language Models

    cs.CV 2025-05 conditional novelty 6.0 of 10

    Introduces a 700-pair benchmark for temporal causal reasoning in VLMs, revealing large open-source vs. closed-source gaps and strong position bias.

  8. FAS-R1: A Unified Multi-Task MLLM for Reasoning Face Anti-Spoofing

    cs.CV 2026-07 conditional novelty 5.0 of 10

    FAS-R1 combines long-CoT supervised fine-tuning with difficulty-aware GRPO and degradation-simulated augmentation to improve multi-task face anti-spoofing and explainable rationales.

  9. Magic-MM-Embedding: Towards Visual-Token-Efficient Universal Multimodal Embedding with MLLMs

    cs.CV 2026-02 conditional novelty 5.0 of 10

    Visual token compression (4x fewer tokens) plus a three-stage generative/contrastive/judge-curated training pipeline yields state-of-the-art MLLM-based retrieval accuracy at lower inference cost.

  10. Towards Mitigating Hallucinations in Large Vision-Language Models by Refining Textual Embeddings

    cs.CV 2025-11 conditional novelty 5.0 of 10

    Injecting an average-pooled visual embedding into every text token improves hallucination-benchmark scores of Video-LLaVA by small single-digit amounts.

  11. Investigating Structural Pruning and Recovery Techniques for Compressing Multimodal Large Language Models: An Empirical Study

    cs.CL 2025-07 conditional novelty 5.0 of 10

    Structural pruning with finetuning plus hidden-state distillation recovers most performance in multimodal LLMs, with 5% of training data sufficient at moderate compression levels.

Pith tools