Pith. sign in

REVIEW 18 cited by

VideoScore: Building Automatic Metrics to Simulate Fine-grained Human Feedback for Video Generation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2406.15252 v3 pith:TDS6J7DZ submitted 2024-06-21 cs.CV cs.AI

classification cs.CVcs.AI
keywords videohumanvideoscoremetricsautomaticfeedbackgenerationmodels
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The recent years have witnessed great advances in video generation. However, the development of automatic video metrics is lagging significantly behind. None of the existing metric is able to provide reliable scores over generated videos. The main barrier is the lack of large-scale human-annotated dataset. In this paper, we release VideoFeedback, the first large-scale dataset containing human-provided multi-aspect score over 37.6K synthesized videos from 11 existing video generative models. We train VideoScore (initialized from Mantis) based on VideoFeedback to enable automatic video quality assessment. Experiments show that the Spearman correlation between VideoScore and humans can reach 77.1 on VideoFeedback-test, beating the prior best metrics by about 50 points. Further result on other held-out EvalCrafter, GenAI-Bench, and VBench show that VideoScore has consistently much higher correlation with human judges than other metrics. Due to these results, we believe VideoScore can serve as a great proxy for human raters to (1) rate different video models to track progress (2) simulate fine-grained human feedback in Reinforcement Learning with Human Feedback (RLHF) to improve current video generation models.

Discussion (0). Sign in to comment.

Forward citations

Cited by 18 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Zoom-IQA: Image Quality Assessment with Reliable Region-Aware Reasoning

    cs.CV 2026-01 conditional novelty 7.0 of 10

    Zoom-IQA lets a vision-language model iteratively crop and zoom into image regions before giving a quality score, improving reasoning and restoration guidance over single-pass IQA models.

  2. GeneVA: A Dataset of Human Annotations for Generative Text to Video Artifacts

    cs.CV 2025-09 conditional novelty 7.0 of 10

    GeneVA is the first large-scale benchmark with human-annotated bounding boxes and text descriptions for artifacts in text-to-video generation.

  3. CachedSearch: Training-Free Cached Exploration for Test-Time Search in Video Diffusion

    cs.AI 2026-07 conditional novelty 6.0 of 10

    Using cached drafts to select the winner and regenerating only that winner captures 94.7% of best-of-8 search gain at 63% of the cost.

  4. FantasyTalking2: Timestep-Layer Adaptive Preference Optimization for Audio-Driven Portrait Animation

    cs.CV 2025-08 unverdicted novelty 6.0 of 10

    A three-part system, Talking-Critic, Talking-NSQ, and TLPO, aligns diffusion portrait animation models to human preferences and improves lip-sync, motion naturalness, and visual quality.

  5. Runtime Failure Hunting for Physics Engine Based Software Systems: How Far Can We Go?

    cs.CV 2025-07 conditional novelty 6.0 of 10

    A new 1,000-video benchmark, PhysiXFails, with a 17-category taxonomy of physics failures, shows prompt-tuned large multimodal models outperform video anomaly detectors at detecting and naming physics rule violations ...

  6. "PhyWorldBench": A Comprehensive Evaluation of Physical Realism in Text-to-Video Models

    cs.CV 2025-07 conditional novelty 6.0 of 10

    PhyWorldBench evaluates 12 text-to-video models on 1,050 physics prompts; the best model passes both semantic adherence and physical commonsense checks in only 26.2% of videos.

  7. ADIEE: Automatic Dataset Creation and Scorer for Instruction-Guided Image Editing Evaluation

    cs.CV 2025-07 conditional novelty 6.0 of 10

    An automatically generated training dataset and a fine-tuned LLaVA-NeXT model produce an image editing evaluation scorer that aligns with human preference and serves as a reward model for improving editing models.

  8. MedGen: Unlocking Medical Video Generation by Scaling Granularly-annotated Medical Videos

    cs.CV 2025-07 conditional novelty 6.0 of 10

    MedVideoCap-55K, a 55,803-clip caption-rich medical video dataset, enables MedGen, a LoRA fine-tune of HunyuanVideo that reports top open-source scores and near-commercial quality on medical video benchmarks.

  9. Fake it till You Make it: Reward Modeling as Discriminative Prediction

    cs.CV 2025-06 conditional novelty 6.0 of 10

    GAN-RM trains a CLIP-based discriminator to distinguish a few hundred preference proxy images from model outputs, then uses it for Best-of-N selection, SFT, and DPO.

  10. OpenS2V-Nexus: A Detailed Benchmark and Million-Scale Dataset for Subject-to-Video Generation

    cs.CV 2025-05 conditional novelty 6.0 of 10

    A benchmark and five-million-clip dataset for evaluating and training subject-to-video generation models, with three new metrics for subject consistency, naturalness, and text alignment.

  11. Scaling Image and Video Generation via Test-Time Evolutionary Search

    cs.CV 2025-05 conditional novelty 6.0 of 10

    Evolutionary search over denoising trajectories improves image and video generation quality and diversity as test-time compute increases, without retraining the generative model.

  12. InfLVG: Reinforce Inference-Time Consistent Long Video Generation with GRPO

    cs.CV 2025-05 conditional novelty 6.0 of 10

    InfLVG uses a GRPO-optimized context selection policy to choose top-K relevant video tokens for consistent, prompt-aligned long video generation.

  13. AIGVE-MACS: Unified Multi-Aspect Commenting and Scoring Model for AI-Generated Video Evaluation

    cs.CV 2025-07 conditional novelty 5.0 of 10

    A finetuned vision-language model jointly predicts nine aspect scores and written comments for AI-generated videos, with a new benchmark and claims of state-of-the-art alignment with human judgment.

  14. GigaVideo-1: Advancing Video Generation via Automatic Feedback with 4 GPU-Hours Fine-Tuning

    cs.CV 2025-06 conditional novelty 5.0 of 10

    GigaVideo-1 fine-tunes Wan2.1 on synthetic weakness-targeted prompts with VLM reward reweighting and reports ~4% average VBench-2.0 gains per dimension at 4 GPU-hours each, though joint training gains less.

  15. AlignHuman: Improving Motion and Fidelity via Timestep-Segment Preference Optimization for Audio-Driven Human Animation

    cs.CV 2025-06 conditional novelty 5.0 of 10

    Timestep-segment preference optimization with separate motion and fidelity LoRAs improves audio-driven human animation quality and allows a 3.3x inference speedup.

  16. Towards Holistic Visual Quality Assessment of AI-Generated Videos: A LLM-Based Multi-Dimensional Evaluation Model

    cs.CV 2025-06 conditional novelty 5.0 of 10

    AIGVEval combines BLIP, 3D Swin Transformer, and SlowFast features with a LoRA-tuned LLM to predict AI-generated video quality, hitting second place on the NTIRE 2025 Track 2 leaderboard.

  17. Robust Single-Stage Fully Sparse 3D Object Detection via Detachable Latent Diffusion

    cs.CV 2025-08 unverdicted novelty 4.0 of 10

    The submitted manuscript's abstract and full text are mismatched; the claimed 3D detection method is not present in the body.

  18. Reinforcement Learning: From Algorithms To Foundation Models

    cs.AI 2026-07 conditional novelty 3.0 of 10

    A dissertation uniting the author's published results: non-exploitable Nash-DQN policies and the FightLadder benchmark for games, plus diffusion/consistency-model world models for RL — a compilation rather than new results.

Pith tools