Pith. sign in

REVIEW 4 cited by

ViCaS: A Dataset for Combining Holistic and Pixel-level Video Understanding using Captions with Grounded Segmentation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2412.09754 v3 pith:6SDP3FBD submitted 2024-12-12 cs.CV

classification cs.CV
keywords segmentationvideomodelsunderstandingbenchmarkcaptionsdatasethigh-level
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Recent advances in multimodal large language models (MLLMs) have expanded research in video understanding, primarily focusing on high-level tasks such as video captioning and question-answering. Meanwhile, a smaller body of work addresses dense, pixel-precise segmentation tasks, which typically involve category-guided or referral-based object segmentation. Although both directions are essential for developing models with human-level video comprehension, they have largely evolved separately, with distinct benchmarks and architectures. This paper aims to unify these efforts by introducing ViCaS, a new dataset containing thousands of challenging videos, each annotated with detailed, human-written captions and temporally consistent, pixel-accurate masks for multiple objects with phrase grounding. Our benchmark evaluates models on both holistic/high-level understanding and language-guided, pixel-precise segmentation. We also present carefully validated evaluation measures and propose an effective model architecture that can tackle our benchmark. Project page: https://ali2500.github.io/vicas-project/

Discussion (0). Sign in to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. MSC: A Marine Wildlife Video Dataset with Grounded Segmentation and Clip-Level Captioning

    cs.CV 2025-08 reject novelty 6.0 of 10

    A new marine wildlife video dataset with clip-level captions and grounded segmentation masks, benchmarked on captioning, grounding, and video generation.

  2. VideoMolmo: Spatio-Temporal Grounding Meets Pointing

    cs.CV 2025-06 conditional novelty 6.0 of 10

    A video language model that conditions each frame on earlier frames via a temporal attention module, predicts text-requested object points, and uses SAM2-based bidirectional mask fusion to outperform prior models on v...

  3. Chain-of-Talkers (CoTalk): Fast Human Annotation of Dense Image Captions

    cs.CL 2025-05 conditional novelty 6.0 of 10

    Chain-of-Talkers (CoTalk) has annotators sequentially dictate only the missing visual details, and it reports modest gains in annotation speed and caption density over parallel typed annotation.

  4. Reasoning Segmentation for Images and Videos: A Survey

    cs.CV 2025-05 conditional novelty 3.0 of 10

    The paper organizes the field of reasoning segmentation into image and video tracks, cataloging 26 methods, 12 metrics, and 29 datasets.

Pith tools