Pith. sign in

REVIEW 2 cited by

VIST-GPT: Ushering in the Era of Visual Storytelling with LLMs?

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2504.19267 v3 pith:J7YAFU5V submitted 2025-04-27 cs.CL cs.AIcs.CVcs.LG

classification cs.CLcs.AIcs.CVcs.LG
keywords visualstorytellingmetricsevaluationmodelsmultimodalnarrativesnovel
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Visual storytelling is an interdisciplinary field combining computer vision and natural language processing to generate cohesive narratives from sequences of images. This paper presents a novel approach that leverages recent advancements in multimodal models, specifically adapting transformer-based architectures and large multimodal models, for the visual storytelling task. Leveraging the large-scale Visual Storytelling (VIST) dataset, our VIST-GPT model produces visually grounded, contextually appropriate narratives. We address the limitations of traditional evaluation metrics, such as BLEU, METEOR, ROUGE, and CIDEr, which are not suitable for this task. Instead, we utilize RoViST and GROOVIST, novel reference-free metrics designed to assess visual storytelling, focusing on visual grounding, coherence, and non-redundancy. These metrics provide a more nuanced evaluation of narrative quality, aligning closely with human judgment.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. LEMUR 2: Unlocking Neural Network Diversity for AI

    cs.LG 2026-07 conditional novelty 5.5 of 10

    LEMUR 2 releases a multi-generator, multi-task neural-architecture corpus with real-device latency metadata intended as fuel for LLM-driven AutoML.

  2. AugmentGest: Can Random Data Cropping Augmentation Boost Gesture Recognition Performance?

    cs.CV 2025-06 conditional novelty 3.0 of 10

    Random cropping, rotation, zoom, and brightness/contrast augmentation on 2D skeleton gesture images improves accuracy on SHREC'17, DHG14/28, and JHMDB by up to 4.5% across three models.

Pith tools