Pith. sign in

REVIEW 2 cited by

Exploiting the Textual Potential from Vision-Language Pre-training for Text-based Person Search

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2303.04497 v1 pith:HAY2VL3P submitted 2023-03-08 cs.CV

classification cs.CV
keywords textualmodalitypre-trainingcorpusfine-grainedmodelmodelsperformance
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Text-based Person Search (TPS), is targeted on retrieving pedestrians to match text descriptions instead of query images. Recent Vision-Language Pre-training (VLP) models can bring transferable knowledge to downstream TPS tasks, resulting in more efficient performance gains. However, existing TPS methods improved by VLP only utilize pre-trained visual encoders, neglecting the corresponding textual representation and breaking the significant modality alignment learned from large-scale pre-training. In this paper, we explore the full utilization of textual potential from VLP in TPS tasks. We build on the proposed VLP-TPS baseline model, which is the first TPS model with both pre-trained modalities. We propose the Multi-Integrity Description Constraints (MIDC) to enhance the robustness of the textual modality by incorporating different components of fine-grained corpus during training. Inspired by the prompt approach for zero-shot classification with VLP models, we propose the Dynamic Attribute Prompt (DAP) to provide a unified corpus of fine-grained attributes as language hints for the image modality. Extensive experiments show that our proposed TPS framework achieves state-of-the-art performance, exceeding the previous best method by a margin.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Blurring Modal Boundaries: A Unified Survey from Single- to Multi-Modal Person Re-ldentification

    cs.CV 2026-07 conditional novelty 5.0 of 10

    A unified taxonomy and survey of single- to multi-modal person ReID, plus a Transformer-based VI-ReID baseline that is solid but not state-of-the-art.

  2. Enhancing Visual Representation for Text-based Person Searching

    cs.CV 2024-12 conditional novelty 5.0 of 10

    VFE-TPS adds text-guided masked image modeling and identity-supervised feature calibration to CLIP and reports state-of-the-art Rank-1 accuracy on three text-based person search benchmarks, though not above a cited Ra...

Pith tools