Pith. sign in

REVIEW 1 cited by

Improving Multimodal Large Language Models Using Continual Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2410.19925 v2 pith:O5GHIO4I submitted 2024-10-25 cs.CL cs.CVcs.LG

classification cs.CLcs.CVcs.LG
keywords continuallearningmultimodallanguagelinguisticperformancewhilecapabilities
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Generative large language models (LLMs) exhibit impressive capabilities, which can be further augmented by integrating a pre-trained vision model into the original LLM to create a multimodal LLM (MLLM). However, this integration often significantly decreases performance on natural language understanding and generation tasks, compared to the original LLM. This study investigates this issue using the LLaVA MLLM, treating the integration as a continual learning problem. We evaluate five continual learning methods to mitigate forgetting and identify a technique that enhances visual understanding while minimizing linguistic performance loss. Our approach reduces linguistic performance degradation by up to 15% over the LLaVA recipe, while maintaining high multimodal accuracy. We also demonstrate the robustness of our method through continual learning on a sequence of vision-language tasks, effectively preserving linguistic skills while acquiring new multimodal capabilities. Project webpage: https://shikhar-srivastava.github.io/cl-for-improving-mllms

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Efficient Few-Shot Continual Learning in Vision-Language Models

    cs.CV 2025-02 conditional novelty 5.0 of 10

    LoRSU selects the most gradient-informative attention heads and MLP parameters of a frozen CLIP encoder to achieve few-shot continual VQA gains with low forgetting and a 25x compute cut.

Pith tools