Pith. sign in

REVIEW 7 cited by

SPOT! Revisiting Video-Language Models for Event Understanding

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2311.12919 v2 pith:Z7DEXKUI submitted 2023-11-21 cs.CV cs.AI

classification cs.CVcs.AI
keywords modelseventunderstandingvideo-languagecaptionseventsvideovideos
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Understanding videos is an important research topic for multimodal learning. Leveraging large-scale datasets of web-crawled video-text pairs as weak supervision has become a pre-training paradigm for learning joint representations and showcased remarkable potential in video understanding tasks. However, videos can be multi-event and multi-grained, while these video-text pairs usually contain only broad-level video captions. This raises a question: with such weak supervision, can video representation in video-language models gain the ability to distinguish even factual discrepancies in textual description and understand fine-grained events? To address this, we introduce SPOT Prober, to benchmark existing video-language models's capacities of distinguishing event-level discrepancies as an indicator of models' event understanding ability. Our approach involves extracting events as tuples (<Subject, Predicate, Object, Attribute, Timestamps>) from videos and generating false event tuples by manipulating tuple components systematically. We reevaluate the existing video-language models with these positive and negative captions and find they fail to distinguish most of the manipulated events. Based on our findings, we propose to plug in these manipulated event captions as hard negative samples and find them effective in enhancing models for event understanding.

Discussion (0). Sign in to comment.

Forward citations

Cited by 7 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Self-Evolving Multi-Agent Systems via Textual Backpropagation

    cs.LG 2025-06 reject novelty 6.0 of 10

    A text-feedback-based framework for automatically optimizing teams of LLM agents outperforms several existing multi-agent systems across coding, math, data analysis, and writing benchmarks.

  2. ReflectRL: Learning from Golden Negative Trajectories via Reflective-to-Direct Reasoning

    cs.AI 2026-08 conditional novelty 5.0 of 10

    ReflectRL repurposes failed expert reasoning traces as reflective scaffolding during RL and distillation training, then transitions the policy to direct reasoning, improving math and science benchmark scores.

  3. Uncertainty-Aware Abstention in Large Language Models with Provable Alignment Guarantees

    cs.CL 2026-07 conditional novelty 4.0 of 10

    CIC selects the largest uncertainty threshold whose Hoeffding or Clopper–Pearson upper bound on acceptance-conditioned error stays ≤ α, guaranteeing finite-sample risk control under exchangeability.

  4. TECP: Token-Entropy Conformal Prediction for LLMs

    cs.CL 2025-08 reject novelty 4.0 of 10

    TECP applies split conformal prediction with token-entropy nonconformity scores to LLM question answering and reports reliable coverage, but its implementation requires the token probabilities it claims to avoid.

  5. FADE: Adversarial Concept Erasure in Flow Models

    cs.CV 2025-07 reject novelty 4.0 of 10

    FADE combines adversarial training with trajectory preservation to erase concepts from diffusion models, reporting state-of-the-art erasure on Stable Diffusion benchmarks, but the evidence is incomplete and the theore...

  6. Conformal Sets in Multiple-Choice Question Answering under Black-Box Settings with Provable Coverage Guarantees

    cs.CL 2025-08 conditional novelty 3.0 of 10

    Repeatedly sampling an LLM and using the entropy of answer frequencies yields conformal prediction sets for multiple-choice questions with empirical miscoverage near the target, and AUROC comparable to logit-based scores.

  7. Conformal P-Value in Multiple-Choice Question Answering Tasks with Provable Risk Control

    cs.CL 2025-08 reject novelty 2.0 of 10

    A p-value reformulation of split conformal prediction for LLM multiple-choice QA achieves nominal miscoverage control on MMLU and MMLU-Pro.

Pith tools