Pith. sign in

REVIEW 1 cited by

Iconographic Image Captioning for Artworks

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2102.03942 v1 pith:DET5S3VD submitted 2021-02-07 cs.CV

classification cs.CV
keywords imagecaptioningcaptionsmodeldatasetimagesresultsannotations
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Image captioning implies automatically generating textual descriptions of images based only on the visual input. Although this has been an extensively addressed research topic in recent years, not many contributions have been made in the domain of art historical data. In this particular context, the task of image captioning is confronted with various challenges such as the lack of large-scale datasets of image-text pairs, the complexity of meaning associated with describing artworks and the need for expert-level annotations. This work aims to address some of those challenges by utilizing a novel large-scale dataset of artwork images annotated with concepts from the Iconclass classification system designed for art and iconography. The annotations are processed into clean textual description to create a dataset suitable for training a deep neural network model on the image captioning task. Motivated by the state-of-the-art results achieved in generating captions for natural images, a transformer-based vision-language pre-trained model is fine-tuned using the artwork image dataset. Quantitative evaluation of the results is performed using standard image captioning metrics. The quality of the generated captions and the model's capacity to generalize to new data is explored by employing the model on a new collection of paintings and performing an analysis of the relation between commonly generated captions and the artistic genre. The overall results suggest that the model can generate meaningful captions that exhibit a stronger relevance to the art historical context, particularly in comparison to captions obtained from models trained only on natural image datasets.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Automating Iconclass: LLMs and RAG for Large-Scale Classification of Religious Woodcuts

    cs.IR 2025-10 conditional novelty 5.0 of 10

    Full-page LLM descriptions matched to Iconclass via vector search and RAG classify 590 early-modern religious woodcuts at 87-92% precision, roughly tripling the F1 of the image-only baseline (0.30 to 0.82).

Pith tools