Pith. sign in

REVIEW 6 cited by

Rec-GPT4V: Multimodal Recommendation with Large Vision-Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2402.08670 v1 pith:XM7JYNT2 submitted 2024-02-13 cs.AI

classification cs.AI
keywords lvlmsimageuserlargemodelsmultimodalvision-languageaddress
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The development of large vision-language models (LVLMs) offers the potential to address challenges faced by traditional multimodal recommendations thanks to their proficient understanding of static images and textual dynamics. However, the application of LVLMs in this field is still limited due to the following complexities: First, LVLMs lack user preference knowledge as they are trained from vast general datasets. Second, LVLMs suffer setbacks in addressing multiple image dynamics in scenarios involving discrete, noisy, and redundant image sequences. To overcome these issues, we propose the novel reasoning scheme named Rec-GPT4V: Visual-Summary Thought (VST) of leveraging large vision-language models for multimodal recommendation. We utilize user history as in-context user preferences to address the first challenge. Next, we prompt LVLMs to generate item image summaries and utilize image comprehension in natural language space combined with item titles to query the user preferences over candidate items. We conduct comprehensive experiments across four datasets with three LVLMs: GPT4-V, LLaVa-7b, and LLaVa-13b. The numerical results indicate the efficacy of VST.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. RecGOAT: Graph Optimal Adaptive Transport for LLM-Enhanced Multimodal Recommendation with Dual Semantic Alignment

    cs.IR 2026-01 reject novelty 6.0 of 10

    RecGOAT aligns LLM and vision item features with collaborative ID embeddings via instance-level contrastive learning and distribution-level optimal transport, reporting state-of-the-art results on three Amazon benchmarks.

  2. Timeripple: Accelerating vDiTs by Understanding the Spatio-Temporal Correlations in Latent Space

    cs.AR 2025-11 conditional novelty 6.0 of 10

    Timeripple cuts vDiT self-attention compute by up to 85% by reusing partial attention scores of spatially and temporally correlated tokens across channels, with VBench quality essentially unchanged.

  3. Multimodal Generative Engine Optimization: Rank Manipulation for Vision-Language Model Rankers

    cs.CL 2026-01 conditional novelty 5.0 of 10

    A coordinated image-plus-text attack moves a target product up VLM shopping rankings by about 2.3 positions on average, more than unimodal attacks or commercial-AI rewriting.

  4. RAG-VisualRec: An Open Resource for Vision- and Text-Enhanced Retrieval-Augmented Generation in Recommendation

    cs.IR 2025-06 conditional novelty 5.0 of 10

    RAG-VisualRec is an open, auditable multimodal benchmark and pipeline for movie recommendation that fuses LLM-generated text with trailer embeddings and reports accuracy and beyond-accuracy metrics.

  5. REGen: Multimodal Retrieval-Embedded Generation for Long-to-Short Video Editing

    cs.CV 2025-05 conditional novelty 5.0 of 10

    REGen generates documentary teasers by fine-tuning an LLM to write a script with <QUOTE> markers, then a trained retriever fills each marker with the most relevant clip from the source video.

  6. ViLLA-MMBench: A Unified Benchmark Suite for LLM-Augmented Multimodal Movie Recommendation

    cs.IR 2025-08 unverdicted novelty 4.0 of 10

    ViLLA-MMBench is an open, YAML-configured benchmark combining audio, visual, and LLM-enriched text embeddings for movie recommendation, reporting cold-start and coverage gains from LLM augmentation.

Pith tools