REVIEW 6 cited by
Rec-GPT4V: Multimodal Recommendation with Large Vision-Language Models
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
The development of large vision-language models (LVLMs) offers the potential to address challenges faced by traditional multimodal recommendations thanks to their proficient understanding of static images and textual dynamics. However, the application of LVLMs in this field is still limited due to the following complexities: First, LVLMs lack user preference knowledge as they are trained from vast general datasets. Second, LVLMs suffer setbacks in addressing multiple image dynamics in scenarios involving discrete, noisy, and redundant image sequences. To overcome these issues, we propose the novel reasoning scheme named Rec-GPT4V: Visual-Summary Thought (VST) of leveraging large vision-language models for multimodal recommendation. We utilize user history as in-context user preferences to address the first challenge. Next, we prompt LVLMs to generate item image summaries and utilize image comprehension in natural language space combined with item titles to query the user preferences over candidate items. We conduct comprehensive experiments across four datasets with three LVLMs: GPT4-V, LLaVa-7b, and LLaVa-13b. The numerical results indicate the efficacy of VST.
Forward citations
Cited by 6 Pith papers
-
RecGOAT: Graph Optimal Adaptive Transport for LLM-Enhanced Multimodal Recommendation with Dual Semantic Alignment
RecGOAT aligns LLM and vision item features with collaborative ID embeddings via instance-level contrastive learning and distribution-level optimal transport, reporting state-of-the-art results on three Amazon benchmarks.
-
Timeripple: Accelerating vDiTs by Understanding the Spatio-Temporal Correlations in Latent Space
Timeripple cuts vDiT self-attention compute by up to 85% by reusing partial attention scores of spatially and temporally correlated tokens across channels, with VBench quality essentially unchanged.
-
Multimodal Generative Engine Optimization: Rank Manipulation for Vision-Language Model Rankers
A coordinated image-plus-text attack moves a target product up VLM shopping rankings by about 2.3 positions on average, more than unimodal attacks or commercial-AI rewriting.
-
RAG-VisualRec: An Open Resource for Vision- and Text-Enhanced Retrieval-Augmented Generation in Recommendation
RAG-VisualRec is an open, auditable multimodal benchmark and pipeline for movie recommendation that fuses LLM-generated text with trailer embeddings and reports accuracy and beyond-accuracy metrics.
-
REGen: Multimodal Retrieval-Embedded Generation for Long-to-Short Video Editing
REGen generates documentary teasers by fine-tuning an LLM to write a script with <QUOTE> markers, then a trained retriever fills each marker with the most relevant clip from the source video.
-
ViLLA-MMBench: A Unified Benchmark Suite for LLM-Augmented Multimodal Movie Recommendation
ViLLA-MMBench is an open, YAML-configured benchmark combining audio, visual, and LLM-enriched text embeddings for movie recommendation, reporting cold-start and coverage gains from LLM augmentation.
Discussion (0). Continue with ORCID to comment.