Pith. sign in

REVIEW 2 cited by

Emerging Property of Masked Token for Effective Pre-training

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2404.08330 v1 pith:ZFIPTBN6 submitted 2024-04-12 cs.CV

classification cs.CV
keywords maskedtokenspre-trainingefficiencytokenapproachinherentmodeling
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Driven by the success of Masked Language Modeling (MLM), the realm of self-supervised learning for computer vision has been invigorated by the central role of Masked Image Modeling (MIM) in driving recent breakthroughs. Notwithstanding the achievements of MIM across various downstream tasks, its overall efficiency is occasionally hampered by the lengthy duration of the pre-training phase. This paper presents a perspective that the optimization of masked tokens as a means of addressing the prevailing issue. Initially, we delve into an exploration of the inherent properties that a masked token ought to possess. Within the properties, we principally dedicated to articulating and emphasizing the `data singularity' attribute inherent in masked tokens. Through a comprehensive analysis of the heterogeneity between masked tokens and visible tokens within pre-trained models, we propose a novel approach termed masked token optimization (MTO), specifically designed to improve model efficiency through weight recalibration and the enhancement of the key property of masked tokens. The proposed method serves as an adaptable solution that seamlessly integrates into any MIM approach that leverages masked tokens. As a result, MTO achieves a considerable improvement in pre-training efficiency, resulting in an approximately 50% reduction in pre-training epochs required to attain converged performance of the recent approaches.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. TADFormer : Task-Adaptive Dynamic Transformer for Efficient Multi-Task Learning

    cs.CV 2025-01 conditional novelty 6.0 of 10

    A prompt-and-dynamic-filter PEFT design for multi-task dense prediction beats MTLoRA on PASCAL-Context with fewer trainable parameters.

  2. Improving Generative Pre-Training: An In-depth Study of Masked Image Modeling and Denoising Models

    cs.CV 2024-12 conditional novelty 6.0 of 10

    Noise-based and masking-based visual pre-training work better together when corruption is inside the encoder, noise is added at lower-layer features, and masked and noised tokens are explicitly separated.

Pith tools