A retrieval pipeline that augments CLIP queries with LLM-written entity visual descriptions, selected by a retriever-trained rewriter, improves image-text retrieval over CLIP baselines.
Title resolution pending
1 Pith paper cite this work. Polarity classification is still indexing.
1
Pith paper citing it
fields
cs.CV 1years
2025 1verdicts
CONDITIONAL 1representative citing papers
citing papers explorer
-
EvdCLIP: Improving Vision-Language Retrieval with Entity Visual Descriptions from Large Language Models
A retrieval pipeline that augments CLIP queries with LLM-written entity visual descriptions, selected by a retriever-trained rewriter, improves image-text retrieval over CLIP baselines.