Pith. sign in

REVIEW 1 cited by

Prefix-diffusion: A Lightweight Diffusion Model for Diverse Image Captioning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2309.04965 v2 pith:ORXZ6GFM submitted 2023-09-10 cs.CV cs.AIcs.CL

classification cs.CVcs.AIcs.CL
keywords imagediffusioncaptioningmodelcaptionsprefix-diffusiondesigndiverse
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

While impressive performance has been achieved in image captioning, the limited diversity of the generated captions and the large parameter scale remain major barriers to the real-word application of these systems. In this work, we propose a lightweight image captioning network in combination with continuous diffusion, called Prefix-diffusion. To achieve diversity, we design an efficient method that injects prefix image embeddings into the denoising process of the diffusion model. In order to reduce trainable parameters, we employ a pre-trained model to extract image features and further design an extra mapping network. Prefix-diffusion is able to generate diverse captions with relatively less parameters, while maintaining the fluency and relevance of the captions benefiting from the generative capabilities of the diffusion model. Our work paves the way for scaling up diffusion models for image captioning, and achieves promising performance compared with recent approaches.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. DIR: Retrieval-Augmented Image Captioning with Comprehensive Understanding

    cs.CV 2024-12 conditional novelty 6.0 of 10

    DIR improves out-of-domain image captioning by guiding image features with a frozen diffusion model and retrieving text decomposed into objects, actions, and environments.

Pith tools