Pith. sign in

REVIEW 4 cited by

VQA$^2$: Visual Question Answering for Video Quality Assessment

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2411.03795 v4 pith:ZSMJM7PO submitted 2024-11-06 cs.CV cs.AI

classification cs.CVcs.AI
keywords qualityvisualvideotasksansweringassessmentmodelsquestion
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The advent and proliferation of large multi-modal models (LMMs) have introduced new paradigms to computer vision, transforming various tasks into a unified visual question answering framework. Video Quality Assessment (VQA), a classic field in low-level visual perception, focused initially on quantitative video quality scoring. However, driven by advances in LMMs, it is now progressing toward more holistic visual quality understanding tasks. Recent studies in the image domain have demonstrated that Visual Question Answering (VQA) can markedly enhance low-level visual quality evaluation. Nevertheless, related work has not been explored in the video domain, leaving substantial room for improvement. To address this gap, we introduce the VQA2 Instruction Dataset - the first visual question answering instruction dataset that focuses on video quality assessment. This dataset consists of 3 subsets and covers various video types, containing 157,755 instruction question-answer pairs. Then, leveraging this foundation, we present the VQA2 series models. The VQA2 series models interleave visual and motion tokens to enhance the perception of spatial-temporal quality details in videos. We conduct extensive experiments on video quality scoring and understanding tasks, and results demonstrate that the VQA2series models achieve excellent performance in both tasks. Notably, our final model, the VQA2-Assistant, exceeds the renowned GPT-4o in visual quality understanding tasks while maintaining strong competitiveness in quality scoring tasks. Our work provides a foundation and feasible approach for integrating low-level video quality assessment and understanding with LMMs.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. VQ-Insight: Teaching VLMs for AI-Generated Video Quality Understanding via Progressive Visual Reinforcement Learning

    cs.CV 2025-06 conditional novelty 6.0 of 10

    VQ-Insight uses progressive reinforcement learning with temporal shuffle and task rewards to teach a vision-language model to score and compare AI-generated videos, with gains on multiple video quality benchmarks.

  2. Scaling-up Perceptual Video Quality Assessment

    cs.CV 2025-05 conditional novelty 6.0 of 10

    A new pipeline plus datasets (OmniVQA-Chat-400K, OmniVQA-MOS-20K, OmniVQA-FG-Benchmark) yield LMMs with state-of-the-art video quality understanding and rating.

  3. Towards Holistic Visual Quality Assessment of AI-Generated Videos: A LLM-Based Multi-Dimensional Evaluation Model

    cs.CV 2025-06 conditional novelty 5.0 of 10

    AIGVEval combines BLIP, 3D Swin Transformer, and SlowFast features with a LoRA-tuned LLM to predict AI-generated video quality, hitting second place on the NTIRE 2025 Track 2 leaderboard.

  4. Engagement Prediction of Short Videos with Large Multimodal Models

    cs.CV 2025-08 conditional novelty 4.0 of 10

    Large multimodal models, especially one that also processes audio, predict short-video engagement better than the prior feature-based baseline on the SnapUGC test set.

Pith tools