REVIEW 17 cited by
Direct Preference Optimization of Video Large Multimodal Models from Language Model Reward
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Preference modeling techniques, such as direct preference optimization (DPO), has shown effective in enhancing the generalization abilities of large language model (LLM). However, in tasks involving video instruction-following, providing informative feedback, especially for detecting hallucinations in generated responses, remains a significant challenge. Previous studies have explored using large large multimodal models (LMMs) as reward models to guide preference modeling, but their ability to accurately assess the factuality of generated responses compared to corresponding videos has not been conclusively established. This paper introduces a novel framework that utilizes detailed video captions as a proxy of video content, enabling language models to incorporate this information as supporting evidence for scoring video Question Answering (QA) predictions. Our approach demonstrates robust alignment with OpenAI GPT-4V model's reward mechanism, which directly takes video frames as input. Furthermore, we show that applying this tailored reward through DPO significantly improves the performance of video LMMs on video QA tasks.
Forward citations
Cited by 17 Pith papers
-
Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding
EviSelect uses the target multimodal model's internal attention as a prior to dynamically select frames, sampling rates, and resolutions, achieving about 50% token reduction and a 3.9x speedup with better benchmark accuracy.
-
Twins: Learn to Predict Unified Representations with Focal Loss
Channel-wise concatenation of SigLIP2 and Flux VAE features into one token, trained with a focal-style flow-matching loss, yields a unified representation with 1.59 gFID on ImageNet 256 and VAE-level reconstruction.
-
Conan-embedding-v3: Fusing Modality-Specific Models for Omni-Modal Embedding
The paper presents a fusion framework for omni-modal embeddings that identifies and repairs Projector Drift in audio modalities, reporting 74.9 on MMEB and 55.61 on MAEB.
-
Cambrian-P: Pose-Grounded Video Understanding
Cambrian-P adds per-frame camera pose tokens and a regression head to video MLLMs, delivering 4.5-6.5% gains on spatial benchmarks, generalization to other video QA tasks, and SOTA streaming pose estimation on ScanNet.
-
Omni-RRM: Advancing Omni Reward Modeling via Automatic Rubric-Grounded Preference Synthesis
A single rubric-grounded reward model trained on automatically synthesized, teacher-reconciled preference pairs claims 80.2% on ShareGPT-Video, 66.8% on a synthetic audio benchmark, and 71.8% overall.
-
Mitigating Hallucinations in Multimodal LLMs via Object-aware Preference Optimization
CHAIR-DPO labels DPO preference pairs with CHAIR hallucinations computed from detector outputs, reducing object hallucinations in LLaVA models without proprietary judges.
-
DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs
DisCo assigns each visual token to a unique concept from the caption and aligns its attention across frames, improving video MLLM accuracy and token efficiency.
-
HaploOmni: Unified Single Transformer for Multimodal Video Understanding and Generation
A training recipe that initializes different depth segments of one transformer from pretrained ViT, LLM, and DiT models, then jointly tunes them to do multimodal understanding and generation.
-
Zooming from Context to Cue: Hierarchical Preference Optimization for Multi-Image MLLMs
Context-to-Cue Direct Preference Optimization (CcDPO) reduces multi-image hallucinations in 7B multimodal LLMs by training on perturbed full-sequence captions and region-focused visual prompts, improving average multi...
-
TUNA: Comprehensive Fine-grained Temporal Understanding Evaluation on Dense Dynamic Videos
TUNA introduces a 1,000-video benchmark with dense temporal captions and 1,432 multiple-choice questions, and finds that current video LMMs are weakest at camera motion, action sequences, and multi-subject scenes.
-
Clapper: Compact Learning and Video Representation in VLMs
Clapper achieves 13x visual token compression in video VLMs with maintained or improved QA accuracy using a slow-fast representation and a TimePerceiver module.
-
Investigating and Enhancing the Robustness of Large Multimodal Models Against Temporal Inconsistency
Large multimodal models are shown to rely on prior knowledge and text cues rather than video order under temporal inconsistency, and a benchmark plus preference-optimization method partially correct this.
-
Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs
A CLIP-scored, Gumbel-Max frame sampler with per-frame multi-resolution allocation improves long-video question answering in Video-LLMs under a fixed token budget.
-
VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation
An open-ended short-answer long-video benchmark, built by converting MCQ questions from four existing tests, shows large accuracy drops and different model rankings versus multiple-choice evaluation.
-
SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement
A three-stage coarse-to-fine training recipe for vision backbones produces consistent benchmark gains for lightweight multimodal LLMs.
-
LeanPO: Lean Preference Optimization for Likelihood Alignment in Video-LLMs
LeanPO improves Video-LLM alignment by using a reference-free average-likelihood reward, self-generated winning/losing pairs, and dynamic label smoothing, yielding gains on six video benchmarks.
-
DPO Learning with LLMs-Judge Signal for Computer Use Agents
An LLM-as-Judge pipeline that scores synthetic GUI interaction trajectories and fine-tunes a 2B model with DPO yields a local computer-use agent that beats its base model on 15-step OSWorld tasks.
Discussion (0). Continue with ORCID to comment.