Pith. sign in

REVIEW 2 cited by

Data-free Multi-label Image Recognition via LLM-powered Prompt Tuning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2403.01209 v1 pith:3V2W3LE3 submitted 2024-03-02 cs.CV

classification cs.CV
keywords multi-labelrecognitionframeworkpromptpromptsclassificationclipdata-free
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

This paper proposes a novel framework for multi-label image recognition without any training data, called data-free framework, which uses knowledge of pre-trained Large Language Model (LLM) to learn prompts to adapt pretrained Vision-Language Model (VLM) like CLIP to multilabel classification. Through asking LLM by well-designed questions, we acquire comprehensive knowledge about characteristics and contexts of objects, which provides valuable text descriptions for learning prompts. Then we propose a hierarchical prompt learning method by taking the multi-label dependency into consideration, wherein a subset of category-specific prompt tokens are shared when the corresponding objects exhibit similar attributes or are more likely to co-occur. Benefiting from the remarkable alignment between visual and linguistic semantics of CLIP, the hierarchical prompts learned from text descriptions are applied to perform classification of images during inference. Our framework presents a new way to explore the synergies between multiple pre-trained models for novel category recognition. Extensive experiments on three public datasets (MS-COCO, VOC2007, and NUS-WIDE) demonstrate that our method achieves better results than the state-of-the-art methods, especially outperforming the zero-shot multi-label recognition methods by 4.7% in mAP on MS-COCO.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. OpenLKA: an open dataset of lane keeping assist from market autonomous vehicles

    cs.RO 2025-01 conditional novelty 6.0 of 10

    OpenLKA is an open dataset showing that commercial lane keeping assist systems deviate significantly on sharp curves and in low-contrast, adverse conditions.

  2. Decoding Neighborhood Environments with Large Language Models

    cs.AI 2025-05 conditional novelty 4.0 of 10

    Majority voting across three commercial LLMs achieves roughly 88 percent accuracy in detecting six neighborhood-environment indicators from Google Street View images, below a trained YOLOv11 detector's 99 percent mAP50.

Pith tools