Pith. sign in

REVIEW 1 cited by

MMSum: A Dataset for Multimodal Summarization and Thumbnail Generation of Videos

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2306.04216 v2 pith:KJF3KJU7 submitted 2023-06-07 cs.CV cs.MM

classification cs.CVcs.MM
keywords datasetmultimodalsummarizationmmsumtextitcategorizationchallengesdata
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Multimodal summarization with multimodal output (MSMO) has emerged as a promising research direction. Nonetheless, numerous limitations exist within existing public MSMO datasets, including insufficient maintenance, data inaccessibility, limited size, and the absence of proper categorization, which pose significant challenges. To address these challenges and provide a comprehensive dataset for this new direction, we have meticulously curated the \textbf{MMSum} dataset. Our new dataset features (1) Human-validated summaries for both video and textual content, providing superior human instruction and labels for multimodal learning. (2) Comprehensively and meticulously arranged categorization, spanning 17 principal categories and 170 subcategories to encapsulate a diverse array of real-world scenarios. (3) Benchmark tests performed on the proposed dataset to assess various tasks and methods, including \textit{video summarization}, \textit{text summarization}, and \textit{multimodal summarization}. To champion accessibility and collaboration, we will release the \textbf{MMSum} dataset and the data collection tool as fully open-source resources, fostering transparency and accelerating future developments. Our project website can be found at~\url{https://mmsum-dataset.github.io/}

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. MS4UI: A Dataset for Multi-modal Summarization of User Interface Instructional Videos

    cs.CV 2025-06 conditional novelty 7.0 of 10

    MS4UI provides a new benchmark for summarizing UI instructional videos into step-by-step text and key frames, where existing models perform poorly.

Pith tools