Pith. sign in

REVIEW 13 cited by

MiniGPT4-Video: Advancing Multimodal LLMs for Video Understanding with Interleaved Visual-Textual Tokens

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2404.03413 v1 pith:XDYU3LP7 submitted 2024-04-04 cs.CV

classification cs.CV
keywords modelminigpt4-videovisualunderstandingbenchmarksmultimodaltextualvideo
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

This paper introduces MiniGPT4-Video, a multimodal Large Language Model (LLM) designed specifically for video understanding. The model is capable of processing both temporal visual and textual data, making it adept at understanding the complexities of videos. Building upon the success of MiniGPT-v2, which excelled in translating visual features into the LLM space for single images and achieved impressive results on various image-text benchmarks, this paper extends the model's capabilities to process a sequence of frames, enabling it to comprehend videos. MiniGPT4-video does not only consider visual content but also incorporates textual conversations, allowing the model to effectively answer queries involving both visual and text components. The proposed model outperforms existing state-of-the-art methods, registering gains of 4.22%, 1.13%, 20.82%, and 13.1% on the MSVD, MSRVTT, TGIF, and TVQA benchmarks respectively. Our models and code have been made publicly available here https://vision-cair.github.io/MiniGPT4-video/

Discussion (0). Sign in to comment.

Forward citations

Cited by 13 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. ObjectStream: Latent Objects as Memory Anchors for Streaming Video Understanding

    cs.CV 2026-07 conditional novelty 6.0 of 10

    Training-free latent-object memory anchors let frozen Video-LLMs retain object histories under a tight token budget and improve streaming and long-video QA.

  2. Sign Language Question Answering: A New Task, Benchmark, and Baseline for Sign Language Understanding

    cs.AI 2026-07 conditional novelty 6.0 of 10

    Sign Language QA benchmarks are introduced from PHOENIX14T and CSL-Daily via template-generated questions, and a question-conditioned baseline outperforms video-language and cascaded baselines.

  3. TimeThink: Reasoning with Time for Video LLMs

    cs.CV 2026-07 accept novelty 6.0 of 10

    TimeThink adds step-wise temporal process rewards (max IoU of referenced intervals) to GRPO for Video-LLMs, improving grounding and reasoning over outcome-only RL baselines.

  4. VUDG: A Dataset for Video Understanding Domain Generalization

    cs.CV 2025-05 conditional novelty 6.0 of 10

    VUDG is a domain-generalization benchmark for video understanding with 11 domains and 36,388 QA pairs, and it shows that current large video-language models lose accuracy across visual domains.

  5. Multimodal Large Language Model-Enabled Video Translation: A Role-Oriented Survey

    cs.CV 2026-04 unverdicted novelty 5.0 of 10

    The paper offers the first focused review of MLLM-based video translation organized by a three-role taxonomy of Semantic Reasoner, Expressive Performer, and Visual Synthesizer, plus open challenges.

  6. LLM-Guided Semantic Relational Reasoning for Multimodal Intent Recognition

    cs.MM 2025-09 conditional novelty 5.0 of 10

    LGSRR uses LLM-generated semantic descriptions and rankings to improve multimodal intent recognition, reporting SOTA results on MIntRec2.0 and IEMOCAP-DA with gains around 0.5-1.3%.

  7. Describe What You See with Multimodal Large Language Models to Enhance Video Recommendations

    cs.IR 2025-08 conditional novelty 5.0 of 10

    Replacing raw video and audio features with MLLM-generated natural-language captions improves hit rate and nDCG for two-tower and SASRec recommenders on MicroLens-100K.

  8. Multi-modal brain encoding models for multi-modal stimuli

    q-bio.NC 2025-05 conditional novelty 5.0 of 10

    On movie-watching fMRI data, multi-modal vision-audio transformers predict brain activity better than unimodal video or speech models, with video dominating cross-modal alignment and video plus audio jointly contribut...

  9. MCAM: Multimodal Causal Analysis Model for Ego-Vehicle-Level Driving Video Understanding

    cs.CV 2025-07 reject novelty 4.0 of 10

    MCAM is a video captioning model combining 3DResNet and VidSwin features with a graph-inspired fusion module, reporting mixed gains on BDD-X and CoVLA but failing to implement the promised causal reasoning.

  10. ViFusion: In-Network Tensor Fusion for Scalable Video Feature Indexing

    cs.MM 2025-06 reject novelty 4.0 of 10

    ViFusion combines dynamic tensor fusion with hierarchical AllReduce to speed up distributed video feature indexing, but the 8-22x throughput claim is an overstatement of bandwidth gains over a self-defined baseline.

  11. LeanPO: Lean Preference Optimization for Likelihood Alignment in Video-LLMs

    cs.CV 2025-06 conditional novelty 4.0 of 10

    LeanPO improves Video-LLM alignment by using a reference-free average-likelihood reward, self-generated winning/losing pairs, and dynamic label smoothing, yielding gains on six video benchmarks.

  12. From Motion to Behavior: Hierarchical Modeling of Humanoid Generative Behavior Control

    cs.RO 2025-05 reject novelty 4.0 of 10

    A new 124K-clip dataset with hierarchical text annotations, plus a pipeline that couples an LLM planner, a text-to-pose VAE, diffusion in-betweening, and physics control to generate long-horizon human behaviors.

  13. Vision Generalist Model: A Survey

    cs.CV 2025-06 conditional novelty 3.0 of 10

    A structured review of vision generalist models, classifying them into encoding-based and sequence-to-sequence frameworks and summarizing datasets, benchmarks, techniques, and open problems.

Pith tools