Pith. sign in

REVIEW 5 cited by

Multimodal Large Language Models: A Survey

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2311.13165 v1 pith:CNQ7QIOA submitted 2023-11-22 cs.AI

classification cs.AI
keywords multimodalmodelslanguagedataalgorithmsaspectsdevelopmentlarge
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The exploration of multimodal language models integrates multiple data types, such as images, text, language, audio, and other heterogeneity. While the latest large language models excel in text-based tasks, they often struggle to understand and process other data types. Multimodal models address this limitation by combining various modalities, enabling a more comprehensive understanding of diverse data. This paper begins by defining the concept of multimodal and examining the historical development of multimodal algorithms. Furthermore, we introduce a range of multimodal products, focusing on the efforts of major technology companies. A practical guide is provided, offering insights into the technical aspects of multimodal models. Moreover, we present a compilation of the latest algorithms and commonly used datasets, providing researchers with valuable resources for experimentation and evaluation. Lastly, we explore the applications of multimodal models and discuss the challenges associated with their development. By addressing these aspects, this paper aims to facilitate a deeper understanding of multimodal models and their potential in various domains.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. M3LLM: Model Context Protocol-aided Mixture of Vision Experts For Multimodal LLMs in Networks

    cs.NI 2025-08 reject novelty 6.0 of 10

    M3LLM routes each multimodal query to the semantically most suitable, wirelessly reachable vision expert using protocol-aided retrieval and a decoupled reinforcement learning agent.

  2. Smart Glasses for CVI: Co-Designing Extended Reality Solutions to Support Environmental Perception by People with Cerebral Visual Impairment

    cs.HC 2025-06 conditional novelty 6.0 of 10

    A co-design study with two adults with CVI found that smart glasses with visual overlays can support locating objects, reading, recognizing people, conversations, and stress management.

  3. Vision-Based Assistive Technologies for People with Cerebral Visual Impairment: A Review and Focus Study

    cs.HC 2025-05 accept novelty 6.0 of 10

    A scoping review and focus groups show that vision-based assistive technology has largely ignored cerebral visual impairment, and identify seven challenges and device opportunities for this group.

  4. Broadening Our View: Assistive Technology for Cerebral Visual Impairment

    cs.HC 2025-05 conditional novelty 5.0 of 10

    A review finds that assistive technology research for cerebral visual impairment is scarce, with only 14 papers and one directly CVI-focused assistance study.

  5. Advancing Responsible Innovation in Agentic AI: A study of Ethical Frameworks for Household Automation

    cs.AI 2025-07 conditional novelty 2.0 of 10

    A review that argues ethical principles for household agentic AI must be converted into concrete design patterns for tailored explainability, granular consent, and user override, especially for vulnerable groups.

Pith tools