Pith. sign in

REVIEW 17 cited by

Direct Preference Optimization of Video Large Multimodal Models from Language Model Reward

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2404.01258 v2 pith:4O35KQMF submitted 2024-04-01 cs.CV cs.AI

classification cs.CVcs.AI
keywords videolargemodelspreferencerewardlanguagemodeldirect
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Preference modeling techniques, such as direct preference optimization (DPO), has shown effective in enhancing the generalization abilities of large language model (LLM). However, in tasks involving video instruction-following, providing informative feedback, especially for detecting hallucinations in generated responses, remains a significant challenge. Previous studies have explored using large large multimodal models (LMMs) as reward models to guide preference modeling, but their ability to accurately assess the factuality of generated responses compared to corresponding videos has not been conclusively established. This paper introduces a novel framework that utilizes detailed video captions as a proxy of video content, enabling language models to incorporate this information as supporting evidence for scoring video Question Answering (QA) predictions. Our approach demonstrates robust alignment with OpenAI GPT-4V model's reward mechanism, which directly takes video frames as input. Furthermore, we show that applying this tailored reward through DPO significantly improves the performance of video LMMs on video QA tasks.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 17 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Evidence-Driven Dynamic Visual Selector for Efficient Long Video Understanding

    cs.CV 2026-08 conditional novelty 6.0 of 10

    EviSelect uses the target multimodal model's internal attention as a prior to dynamically select frames, sampling rates, and resolutions, achieving about 50% token reduction and a 3.9x speedup with better benchmark accuracy.

  2. Twins: Learn to Predict Unified Representations with Focal Loss

    cs.CV 2026-07 conditional novelty 6.0 of 10

    Channel-wise concatenation of SigLIP2 and Flux VAE features into one token, trained with a focal-style flow-matching loss, yields a unified representation with 1.59 gFID on ImageNet 256 and VAE-level reconstruction.

  3. Conan-embedding-v3: Fusing Modality-Specific Models for Omni-Modal Embedding

    cs.MM 2026-06 unverdicted novelty 6.0 of 10

    The paper presents a fusion framework for omni-modal embeddings that identifies and repairs Projector Drift in audio modalities, reporting 74.9 on MMEB and 55.61 on MAEB.

  4. Cambrian-P: Pose-Grounded Video Understanding

    cs.CV 2026-05 unverdicted novelty 6.0 of 10

    Cambrian-P adds per-frame camera pose tokens and a regression head to video MLLMs, delivering 4.5-6.5% gains on spatial benchmarks, generalization to other video QA tasks, and SOTA streaming pose estimation on ScanNet.

  5. Omni-RRM: Advancing Omni Reward Modeling via Automatic Rubric-Grounded Preference Synthesis

    cs.CL 2026-01 conditional novelty 6.0 of 10

    A single rubric-grounded reward model trained on automatically synthesized, teacher-reconciled preference pairs claims 80.2% on ShareGPT-Video, 66.8% on a synthetic audio benchmark, and 71.8% overall.

  6. Mitigating Hallucinations in Multimodal LLMs via Object-aware Preference Optimization

    cs.CV 2025-08 conditional novelty 6.0 of 10

    CHAIR-DPO labels DPO preference pairs with CHAIR hallucinations computed from detector outputs, reducing object hallucinations in LLaVA models without proprietary judges.

  7. DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs

    cs.CV 2025-07 conditional novelty 6.0 of 10

    DisCo assigns each visual token to a unique concept from the caption and aligns its attention across frames, improving video MLLM accuracy and token efficiency.

  8. HaploOmni: Unified Single Transformer for Multimodal Video Understanding and Generation

    cs.CV 2025-06 conditional novelty 6.0 of 10

    A training recipe that initializes different depth segments of one transformer from pretrained ViT, LLM, and DiT models, then jointly tunes them to do multimodal understanding and generation.

  9. Zooming from Context to Cue: Hierarchical Preference Optimization for Multi-Image MLLMs

    cs.CV 2025-05 conditional novelty 6.0 of 10

    Context-to-Cue Direct Preference Optimization (CcDPO) reduces multi-image hallucinations in 7B multimodal LLMs by training on perturbed full-sequence captions and region-focused visual prompts, improving average multi...

  10. TUNA: Comprehensive Fine-grained Temporal Understanding Evaluation on Dense Dynamic Videos

    cs.CV 2025-05 conditional novelty 6.0 of 10

    TUNA introduces a 1,000-video benchmark with dense temporal captions and 1,432 multiple-choice questions, and finds that current video LMMs are weakest at camera motion, action sequences, and multi-subject scenes.

  11. Clapper: Compact Learning and Video Representation in VLMs

    cs.CV 2025-05 conditional novelty 6.0 of 10

    Clapper achieves 13x visual token compression in video VLMs with maintained or improved QA accuracy using a slow-fast representation and a TimePerceiver module.

  12. Investigating and Enhancing the Robustness of Large Multimodal Models Against Temporal Inconsistency

    cs.CV 2025-05 conditional novelty 6.0 of 10

    Large multimodal models are shown to rely on prior knowledge and text cues rather than video order under temporal inconsistency, and a benchmark plus preference-optimization method partially correct this.

  13. Q-Frame: Query-aware Frame Selection and Multi-Resolution Adaptation for Video-LLMs

    cs.CV 2025-06 conditional novelty 5.0 of 10

    A CLIP-scored, Gumbel-Max frame sampler with per-frame multi-resolution allocation improves long-video question answering in Video-LLMs under a fixed token budget.

  14. VideoEval-Pro: Robust and Realistic Long Video Understanding Evaluation

    cs.CV 2025-05 conditional novelty 5.0 of 10

    An open-ended short-answer long-video benchmark, built by converting MCQ questions from four existing tests, shows large accuracy drops and different model rankings versus multiple-choice evaluation.

  15. SAILViT: Towards Robust and Generalizable Visual Backbones for MLLMs via Gradual Feature Refinement

    cs.CV 2025-07 conditional novelty 4.0 of 10

    A three-stage coarse-to-fine training recipe for vision backbones produces consistent benchmark gains for lightweight multimodal LLMs.

  16. LeanPO: Lean Preference Optimization for Likelihood Alignment in Video-LLMs

    cs.CV 2025-06 conditional novelty 4.0 of 10

    LeanPO improves Video-LLM alignment by using a reference-free average-likelihood reward, self-generated winning/losing pairs, and dynamic label smoothing, yielding gains on six video benchmarks.

  17. DPO Learning with LLMs-Judge Signal for Computer Use Agents

    cs.AI 2025-06 conditional novelty 4.0 of 10

    An LLM-as-Judge pipeline that scores synthetic GUI interaction trajectories and fine-tunes a 2B model with DPO yields a local computer-use agent that beats its base model on 15-step OSWorld tasks.

Pith tools