Pith. sign in

REVIEW 4 cited by

Cobra: Extending Mamba to Multi-Modal Large Language Model for Efficient Inference

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2403.14520 v4 pith:MFG5AZXP submitted 2024-03-21 cs.CV

classification cs.CV
keywords cobraefficientcomplexitylanguagemambamllmmodelachieves
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

In recent years, the application of multimodal large language models (MLLM) in various fields has achieved remarkable success. However, as the foundation model for many downstream tasks, current MLLMs are composed of the well-known Transformer network, which has a less efficient quadratic computation complexity. To improve the efficiency of such basic models, we propose Cobra, a linear computational complexity MLLM. Specifically, Cobra integrates the efficient Mamba language model into the visual modality. Moreover, we explore and study various modal fusion schemes to create an effective multi-modal Mamba. Extensive experiments demonstrate that (1) Cobra achieves extremely competitive performance with current computationally efficient state-of-the-art methods, e.g., LLaVA-Phi, TinyLLaVA, and MobileVLM v2, and has faster speed due to Cobra's linear sequential modeling. (2) Interestingly, the results of closed-set challenging prediction benchmarks show that Cobra performs well in overcoming visual illusions and spatial relationship judgments. (3) Notably, Cobra even achieves comparable performance to LLaVA with about 43% of the number of parameters. We will make all codes of Cobra open-source and hope that the proposed method can facilitate future research on complexity problems in MLLM. Our project page is available at: https://sites.google.com/view/cobravlm.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Boosting Embodied AI Agents through Perception-Generation Disaggregation and Asynchronous Pipeline Execution

    cs.AI 2025-09 conditional novelty 6.0 of 10

    Auras, a perception-generation disaggregation framework with a public context buffer and asynchronous pipeline executor, raises embodied-agent throughput by 2.54x on average without losing accuracy (102.7%).

  2. scMamba: A Scalable Foundation Model for Single-Cell Multi-Omics Integration Beyond Highly Variable Feature Selection

    q-bio.CB 2025-06 conditional novelty 6.0 of 10

    A Mamba-based model integrates paired single-cell RNA and ATAC data using all genomic features, with patch tokenization and contrastive learning, and reports better integration and downstream performance than seven ex...

  3. Finding Change in Satellite Archives from Text: How to Combine Before-and-After Images Efficiently

    cs.CV 2026-07 accept novelty 5.5 of 10

    A training-free subtraction-then-attention cascade matches or beats full fusion recall on LEVIR-CC at 10–15× lower query cost; Mamba is no faster than attention at L=196; TBF cuts parameters 2.3× for a 0.007 BLEU-1 cost.

  4. HAMF: A Hybrid Attention-Mamba Framework for Joint Scene Context Understanding and Future Motion Representation Learning

    cs.CV 2025-05 conditional novelty 5.0 of 10

    HAMF feeds learnable future motion tokens into the scene encoder alongside road and agent tokens, then uses a Mamba decoder to output six diverse trajectories, achieving competitive Argoverse 2 results with 3.0M parameters.

Pith tools