Pith. sign in

REVIEW 8 cited by

TokenLearner: What Can 8 Learned Tokens Do for Images and Videos?

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2106.11297 v4 pith:GXENAZTU submitted 2021-06-21 cs.CV cs.LG

classification cs.CVcs.LG
keywords tokensvisualresultsvideoapproachattentionimageimages
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

In this paper, we introduce a novel visual representation learning which relies on a handful of adaptively learned tokens, and which is applicable to both image and video understanding tasks. Instead of relying on hand-designed splitting strategies to obtain visual tokens and processing a large number of densely sampled patches for attention, our approach learns to mine important tokens in visual data. This results in efficiently and effectively finding a few important visual tokens and enables modeling of pairwise attention between such tokens, over a longer temporal horizon for videos, or the spatial content in images. Our experiments demonstrate strong performance on several challenging benchmarks for both image and video recognition tasks. Importantly, due to our tokens being adaptive, we accomplish competitive results at significantly reduced compute amount. We obtain comparable results to the state-of-the-arts on ImageNet while being computationally more efficient. We also confirm the effectiveness of the approach on multiple video datasets, including Kinetics-400, Kinetics-600, Charades, and AViD. The code is available at: https://github.com/google-research/scenic/tree/main/scenic/projects/token_learner

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 8 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. ATS-ToDMA: Adaptive Token Selection and Token-Domain Multiple Access for Cross-Modal Semantic Communications

    cs.IT 2026-07 conditional novelty 6.0 of 10

    ATS-ToDMA jointly selects semantic tokens, schedules them under a similarity-based SSINR interference model, and allocates power, yielding higher simulated semantic throughput and accuracy than OMA and Semantic NOMA.

  2. VolumeDP: Modeling Volumetric Representation for Manipulation Policy Learning

    cs.RO 2026-03 conditional novelty 6.0 of 10

    An RGB-only diffusion policy that builds an explicit volumetric representation, distills it into spatial tokens, and conditions a multi-token decoder reaches 88.8% average success on LIBERO.

  3. Think Proprioceptively: State-Grounded Visual Token Selection for VLA Policies

    cs.RO 2026-02 conditional novelty 6.0 of 10

    Using tokenized proprioception plus instruction to select ~15% of visual patches matches or beats full-token VLA baselines and cuts latency by ~58%.

  4. A Comprehensive Study on Visual Token Redundancy for Discrete Diffusion-based Multimodal Large Language Models

    cs.CV 2025-11 conditional novelty 6.0 of 10

    Visual token pruning causes severe loss in discrete diffusion MLLMs; only from-scratch models on long-answer tasks recover via late denoising, so redundancy is recoverability, not dispensability.

  5. Cross-Modal Dual-Causal Learning for Long-Term Action Recognition

    cs.CV 2025-07 reject novelty 6.0 of 10

    CMDCL debiases text embeddings by back-door adjustment and deconfounds video features by front-door adjustment, achieving state-of-the-art long-term action recognition on three benchmarks.

  6. Brain-Aligned Multi-Stream Video Transformers with Sparse Self-Selection

    cs.CV 2026-07 conditional novelty 5.0 of 10

    A split-and-fuse video transformer with sparse winner-takes-all token selection reports 82.55% top-1 on Kinetics-400, Pareto-efficient inference, and peak EEG RSA of 0.18 (about 78% of the noise ceiling).

  7. Rethinking Visual Autoregressive Sampling with Information-Grounding Guidance

    cs.CV 2025-09 conditional novelty 5.0 of 10

    IGG, an attention-based reweighting of classifier-free guidance, concentrates guidance on important tokens and modestly improves FID/IS in scale-wise autoregressive image generation.

  8. Lightweight Backbone Networks Only Require Adaptive Lightweight Self-Attention Mechanisms

    cs.CV 2025-08 conditional novelty 5.0 of 10

    LOLViT, a GhostNet-based lightweight backbone using adaptive window attention, reports CNN-like CPU speed with MobileViT-level accuracy.

Pith tools