Pith. sign in

REVIEW 5 cited by

Vision-by-Language for Training-Free Compositional Image Retrieval

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2310.09291 v2 pith:XOYM2JFX submitted 2023-10-13 cs.CV

classification cs.CV
keywords imageretrievaltargetzs-circirevlcompositionallanguagemodels
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Given an image and a target modification (e.g an image of the Eiffel tower and the text "without people and at night-time"), Compositional Image Retrieval (CIR) aims to retrieve the relevant target image in a database. While supervised approaches rely on annotating triplets that is costly (i.e. query image, textual modification, and target image), recent research sidesteps this need by using large-scale vision-language models (VLMs), performing Zero-Shot CIR (ZS-CIR). However, state-of-the-art approaches in ZS-CIR still require training task-specific, customized models over large amounts of image-text pairs. In this work, we propose to tackle CIR in a training-free manner via our Compositional Image Retrieval through Vision-by-Language (CIReVL), a simple, yet human-understandable and scalable pipeline that effectively recombines large-scale VLMs with large language models (LLMs). By captioning the reference image using a pre-trained generative VLM and asking a LLM to recompose the caption based on the textual target modification for subsequent retrieval via e.g. CLIP, we achieve modular language reasoning. In four ZS-CIR benchmarks, we find competitive, in-part state-of-the-art performance - improving over supervised methods. Moreover, the modularity of CIReVL offers simple scalability without re-training, allowing us to both investigate scaling laws and bottlenecks for ZS-CIR while easily scaling up to in parts more than double of previously reported results. Finally, we show that CIReVL makes CIR human-understandable by composing image and text in a modular fashion in the language domain, thereby making it intervenable, allowing to post-hoc re-align failure cases. Code will be released upon acceptance.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. CR-Refiner: An Object-Centric Optimal Transport Reranker for Edit-Conditioned 3D Scene Retrieval

    cs.CV 2026-07 conditional novelty 6.0 of 10

    An unbalanced optimal-transport reranker with structural priors and an LLM verifier improves hard-subset 3D scene retrieval, evaluated on the new synthetic 3D-CER benchmark.

  2. Beyond Simple Edits: Composed Video Retrieval with Dense Modifications

    cs.CV 2025-08 conditional novelty 6.0 of 10

    A new benchmark with much longer, denser modification texts, plus a single-encoder fusion model, raises composed video retrieval Recall@1 by 3.4 points on its own test set.

  3. FACap: A Large-scale Fashion Dataset for Fine-grained Composed Image Retrieval

    cs.LG 2025-07 conditional novelty 6.0 of 10

    FACap contributes 227,680 fashion CIR triplets with VLM/LLM-generated modification texts, and FashionBLIP-2 trained on it reaches 44.63 average Recall on FashionIQ without downstream fine-tuning and 65.97 with fine-tuning.

  4. DetailFusion: A Dual-branch Framework with Detail Enhancement for Composed Image Retrieval

    cs.CV 2025-05 conditional novelty 6.0 of 10

    A dual-branch CIR framework that pre-trains a detail-focused branch on InstructPix2Pix editing data, fuses global and detail features with an adaptive compositor, and reports state-of-the-art on CIRR and FashionIQ.

  5. MCoT-RE: Multi-Faceted Chain-of-Thought and Re-Ranking for Training-Free Zero-Shot Composed Image Retrieval

    cs.CV 2025-07 conditional novelty 4.0 of 10

    A training-free composed image retrieval method uses multi-faceted chain-of-thought prompting to produce modification-focused and integration-focused captions, then filters and re-ranks CLIP candidates with a weighted fusion.

Pith tools