REVIEW 5 cited by
GraFT: Gradual Fusion Transformer for Multimodal Re-Identification
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Object Re-Identification (ReID) is pivotal in computer vision, witnessing an escalating demand for adept multimodal representation learning. Current models, although promising, reveal scalability limitations with increasing modalities as they rely heavily on late fusion, which postpones the integration of specific modality insights. Addressing this, we introduce the \textbf{Gradual Fusion Transformer (GraFT)} for multimodal ReID. At its core, GraFT employs learnable fusion tokens that guide self-attention across encoders, adeptly capturing both modality-specific and object-specific features. Further bolstering its efficacy, we introduce a novel training paradigm combined with an augmented triplet loss, optimizing the ReID feature embedding space. We demonstrate these enhancements through extensive ablation studies and show that GraFT consistently surpasses established multimodal ReID benchmarks. Additionally, aiming for deployment versatility, we've integrated neural network pruning into GraFT, offering a balance between model size and performance.
Forward citations
Cited by 5 Pith papers
-
Multi-Modal Object Re-Identification with Prompt-S6 and Semantic-Aware Knowledge Guidance
Prompt-S6 plus semantic token pruning and progressive tri-modal fusion improves multi-spectral object ReID accuracy and efficiency on four benchmarks.
-
ICPL-ReID: Identity-Conditional Prompt Learning for Multi-Spectral Object Re-Identification
An identity-conditioned, online prompt learning framework with low-rank adapters sets new state-of-the-art results on five multi-spectral person and vehicle re-identification benchmarks.
-
MambaPro: Multi-Modal Object Re-Identification with Mamba Aggregation and Synergistic Prompt
MambaPro combines CLIP with a parallel adapter, synergistic residual prompts, and Mamba aggregation to achieve state-of-the-art mAP on RGBNT201, RGBNT100, and MSVR310.
-
DeMo: Decoupled Feature-Based Mixture of Experts for Multi-Modal Object Re-Identification
DeMo improves multi-modal object re-identification by decoupling RGB, NIR, and TIR features into seven attention-derived streams and weighting them with an attention-triggered mixture of experts.
-
Multi-Modal Object Re-Identification with Dual Semantic Guidance and Global-Local Mutual Modulation
A dual-semantic (text + soft mask) global-local mutual modulation framework reports SOTA mAP/Rank-1 on RGBNT201, RGBNT100, and MSVR310.
Discussion (0). Continue with ORCID to comment.