Pith. sign in

REVIEW 4 cited by

Learning to Merge Tokens in Vision Transformers

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2202.12015 v1 pith:KUTWRTSS submitted 2022-02-24 cs.CV cs.LG

classification cs.CVcs.LG
keywords computationalpatchmergerperformancetokenstransformersvisionwhileachieves
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Transformers are widely applied to solve natural language understanding and computer vision tasks. While scaling up these architectures leads to improved performance, it often comes at the expense of much higher computational costs. In order for large-scale models to remain practical in real-world systems, there is a need for reducing their computational overhead. In this work, we present the PatchMerger, a simple module that reduces the number of patches or tokens the network has to process by merging them between two consecutive intermediate layers. We show that the PatchMerger achieves a significant speedup across various model sizes while matching the original performance both upstream and downstream after fine-tuning.

Discussion (0). Sign in to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Segment This Thing: Foveated Tokenization for Efficient Point-Prompted Segmentation

    cs.CV 2025-06 conditional novelty 7.0 of 10

    A point-prompted segmentation model gains efficiency by foveated tokenization, cutting tokens from 4096 to 172 while staying competitive on mIoU benchmarks.

  2. Training-free Token Reduction for Vision Mamba

    cs.CV 2025-07 conditional novelty 6.0 of 10

    MTR uses Mamba's timescale parameter Δ as a token importance score to merge unimportant tokens, giving training-free inference speedups with small accuracy loss.

  3. Detecting Regional Spurious Correlations in Vision Transformers via Token Discarding

    cs.CV 2025-09 conditional novelty 5.0 of 10

    A token-discarding method for vision transformers measures whether predictions rely on features outside the object's bounding box, identifying spurious correlations and problematic ImageNet classes.

  4. Token Compression Meets Compact Vision Transformers: A Survey and Comparative Evaluation for Edge AI

    cs.CV 2025-07 reject novelty 4.0 of 10

    A survey and one-model benchmark concluding that token compression methods hurt compact Vision Transformers when used off the shelf, a conclusion supported only by an unverified experiment on AutoFormer-S.

Pith tools