Pith. sign in

REVIEW 1 cited by

MSVD-Indonesian: A Benchmark for Multimodal Video-Text Tasks in Indonesian

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2306.11341 v2 pith:GBN4BTUQ submitted 2023-06-20 cs.MM cs.CLcs.CVcs.LGeess.IV

classification cs.MMcs.CLcs.CVcs.LGeess.IV
keywords indonesiandatasetenglishmultimodalretrievaltasksdatasetslearning
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Multimodal learning on video and text has seen significant progress, particularly in tasks like text-to-video retrieval, video-to-text retrieval, and video captioning. However, most existing methods and datasets focus exclusively on English. Despite Indonesian being one of the most widely spoken languages, multimodal research in Indonesian remains under-explored, largely due to the lack of benchmark datasets. To address this gap, we introduce the first public Indonesian video-text dataset by translating the English captions in the MSVD dataset into Indonesian. Using this dataset, we evaluate neural network models which were developed for the English video-text dataset on three tasks, i.e., text-to-video retrieval, video-to-text retrieval, and video captioning. Most existing models rely on feature extractors pretrained on English vision-language datasets, raising concerns about their applicability to Indonesian, given the scarcity of large-scale pretraining resources in the language. We apply a cross-lingual transfer learning approach by leveraging English-pretrained extractors and fine-tuning models on our Indonesian dataset. Experimental results demonstrate that this strategy improves performance across all tasks and metrics. We release our dataset publicly to support future research and hope it will inspire further progress in Indonesian multimodal learning.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Leveraging Auxiliary Information in Text-to-Video Retrieval: A Review

    cs.CV 2025-05 conditional novelty 3.0 of 10

    A structured review of 81 text-to-video retrieval papers that leverage auxiliary information, organized by a taxonomy and compared on standard benchmarks.

Pith tools