Pith. sign in

REVIEW 1 cited by

DiffuVST: Narrating Fictional Scenes with Global-History-Guided Denoising Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2312.07066 v1 pith:ZUNYL24N submitted 2023-12-12 cs.CL cs.CV

classification cs.CLcs.CV
keywords diffuvstinferencemodelsscenesvisualautoregressivedenoisingfictional
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Recent advances in image and video creation, especially AI-based image synthesis, have led to the production of numerous visual scenes that exhibit a high level of abstractness and diversity. Consequently, Visual Storytelling (VST), a task that involves generating meaningful and coherent narratives from a collection of images, has become even more challenging and is increasingly desired beyond real-world imagery. While existing VST techniques, which typically use autoregressive decoders, have made significant progress, they suffer from low inference speed and are not well-suited for synthetic scenes. To this end, we propose a novel diffusion-based system DiffuVST, which models the generation of a series of visual descriptions as a single conditional denoising process. The stochastic and non-autoregressive nature of DiffuVST at inference time allows it to generate highly diverse narratives more efficiently. In addition, DiffuVST features a unique design with bi-directional text history guidance and multimodal adapter modules, which effectively improve inter-sentence coherence and image-to-text fidelity. Extensive experiments on the story generation task covering four fictional visual-story datasets demonstrate the superiority of DiffuVST over traditional autoregressive models in terms of both text quality and inference speed.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Structured Relevance Assessment for Robust Retrieval-Augmented Language Models

    cs.AI 2025-07 reject novelty 4.0 of 10

    A retrieval-augmented framework that scores document relevance, balances internal and external knowledge, and abstains when uncertain, claims to cut hallucinations, but the reported evidence is thin and inconsistent.

Pith tools