Pith. sign in

REVIEW 1 cited by

AI Safety in Practice: Enhancing Adversarial Robustness in Multimodal Image Captioning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2407.21174 v1 pith:6OE5M4VA submitted 2024-07-30 cs.CV cs.AIeess.AS

classification cs.CVcs.AIeess.AS
keywords adversarialmultimodalrobustnesstrainingattackscaptioningimagemodels
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Multimodal machine learning models that combine visual and textual data are increasingly being deployed in critical applications, raising significant safety and security concerns due to their vulnerability to adversarial attacks. This paper presents an effective strategy to enhance the robustness of multimodal image captioning models against such attacks. By leveraging the Fast Gradient Sign Method (FGSM) to generate adversarial examples and incorporating adversarial training techniques, we demonstrate improved model robustness on two benchmark datasets: Flickr8k and COCO. Our findings indicate that selectively training only the text decoder of the multimodal architecture shows performance comparable to full adversarial training while offering increased computational efficiency. This targeted approach suggests a balance between robustness and training costs, facilitating the ethical deployment of multimodal AI systems across various domains.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Exploring Visual Embedding Spaces Induced by Vision Transformers for Online Auto Parts Marketplaces

    cs.CV 2025-02 conditional novelty 3.0 of 10

    Visual embeddings from a pretrained ViT produce poorly separated clusters of auto parts images (silhouette 0.015), far below the 0.38 reported for a multimodal model on similar data.

Pith tools