Pith. sign in

REVIEW 5 cited by

Video-STaR: Self-Training Enables Video Instruction Tuning with Any Supervision

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2407.06189 v1 pith:M7A6QVI5 submitted 2024-07-08 cs.CV cs.AI

classification cs.CVcs.AI
keywords videovideo-starinstructionlabelslvlmssupervisiontuningdatasets
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

The performance of Large Vision Language Models (LVLMs) is dependent on the size and quality of their training datasets. Existing video instruction tuning datasets lack diversity as they are derived by prompting large language models with video captions to generate question-answer pairs, and are therefore mostly descriptive. Meanwhile, many labeled video datasets with diverse labels and supervision exist - however, we find that their integration into LVLMs is non-trivial. Herein, we present Video Self-Training with augmented Reasoning (Video-STaR), the first video self-training approach. Video-STaR allows the utilization of any labeled video dataset for video instruction tuning. In Video-STaR, an LVLM cycles between instruction generation and finetuning, which we show (I) improves general video understanding and (II) adapts LVLMs to novel downstream tasks with existing supervision. During generation, an LVLM is prompted to propose an answer. The answers are then filtered only to those that contain the original video labels, and the LVLM is then re-trained on the generated dataset. By only training on generated answers that contain the correct video labels, Video-STaR utilizes these existing video labels as weak supervision for video instruction tuning. Our results demonstrate that Video-STaR-enhanced LVLMs exhibit improved performance in (I) general video QA, where TempCompass performance improved by 10%, and (II) on downstream tasks, where Video-STaR improved Kinetics700-QA accuracy by 20% and action quality assessment on FineDiving by 15%.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. MOSAIC: Adaptive Inter-layer Composition for Efficient Heterogeneous Vision-Language Models

    cs.CV 2026-07 accept novelty 6.0 of 10

    MOSAIC uses multi-objective MIP search over linear/sparse/low-rank operators plus two-stage distillation to convert a homogeneous VLM into a hardware-aware heterogeneous model that matches teacher performance at 2.5× ...

  2. Temporal Preference Optimization for Long-Form Video Understanding

    cs.CV 2025-01 conditional novelty 6.0 of 10

    TPO trains video-LMMs to prefer answers generated from complete, relevant frames over answers from incomplete or irrelevant frames, improving temporal grounding on LongVideoBench, MLVU, and Video-MME.

  3. Apollo: An Exploration of Video Understanding in Large Multimodal Models

    cs.CV 2024-12 conditional novelty 6.0 of 10

    Design decisions for video-LMMs can be made on 2-4B models and datasets and transfer to larger models, yielding efficient Apollo models, though some SOTA claims are contradicted by the paper's own table.

  4. STEP: Enhancing Video-LLMs' Compositional Reasoning by Spatio-Temporal Graph-guided Self-Training

    cs.CV 2024-11 conditional novelty 6.0 of 10

    A graph-guided self-training method lets video-language models generate their own reasoning training data from raw videos, improving multi-step compositional reasoning accuracy.

  5. See, Think, Learn: A Self-Taught Multimodal Reasoner

    cs.CV 2025-12 conditional novelty 4.0 of 10

    A self-training loop that structures model rationales into image-caption, reasoning, and conclusion, and adds negative rationales, improves VLM accuracy on M3CoT over answer-only and STaR baselines.

Pith tools