Pith. sign in

REVIEW 1 cited by

Few-Shot Learning Meets Transformer: Unified Query-Support Transformers for Few-Shot Classification

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2208.12398 v1 pith:4V4TEXFP submitted 2022-08-26 cs.CV

classification cs.CV
keywords learningfew-shottransformerclassificationimageproposedquerysupport
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Few-shot classification which aims to recognize unseen classes using very limited samples has attracted more and more attention. Usually, it is formulated as a metric learning problem. The core issue of few-shot classification is how to learn (1) consistent representations for images in both support and query sets and (2) effective metric learning for images between support and query sets. In this paper, we show that the two challenges can be well modeled simultaneously via a unified Query-Support TransFormer (QSFormer) model. To be specific,the proposed QSFormer involves global query-support sample Transformer (sampleFormer) branch and local patch Transformer (patchFormer) learning branch. sampleFormer aims to capture the dependence of samples in support and query sets for image representation. It adopts the Encoder, Decoder and Cross-Attention to respectively model the Support, Query (image) representation and Metric learning for few-shot classification task. Also, as a complementary to global learning branch, we adopt a local patch Transformer to extract structural representation for each image sample by capturing the long-range dependence of local image patches. In addition, a novel Cross-scale Interactive Feature Extractor (CIFE) is proposed to extract and fuse multi-scale CNN features as an effective backbone module for the proposed few-shot learning method. All modules are integrated into a unified framework and trained in an end-to-end manner. Extensive experiments on four popular datasets demonstrate the effectiveness and superiority of the proposed QSFormer.

Discussion (0). Sign in to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. ViT-ProtoNet for Few-Shot Image Classification: A Multi-Benchmark Evaluation

    cs.CV 2025-07 reject novelty 2.0 of 10

    ViT-ProtoNet, a Prototypical Network with a ViT-Small encoder, is reported to reach 95-97% 5-shot accuracy on three benchmarks and 81.88% on FC100, but the evaluation lacks critical baselines.

Pith tools