Pith. sign in

REVIEW 1 cited by

Efficient Self-supervised Learning with Contextualized Target Representations for Vision, Speech and Language

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2212.07525 v2 pith:27FVDDJ6 submitted 2022-12-14 cs.LG cs.CLcs.SDeess.AS

classification cs.LGcs.CLcs.SDeess.AS
keywords data2vecaccuracylearningrepresentationsself-supervisedtimecontextualizedfast
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Current self-supervised learning algorithms are often modality-specific and require large amounts of computational resources. To address these issues, we increase the training efficiency of data2vec, a learning objective that generalizes across several modalities. We do not encode masked tokens, use a fast convolutional decoder and amortize the effort to build teacher representations. data2vec 2.0 benefits from the rich contextualized target representations introduced in data2vec which enable a fast self-supervised learner. Experiments on ImageNet-1K image classification show that data2vec 2.0 matches the accuracy of Masked Autoencoders in 16.4x lower pre-training time, on Librispeech speech recognition it performs as well as wav2vec 2.0 in 10.6x less time, and on GLUE natural language understanding it matches a retrained RoBERTa model in half the time. Trading some speed for accuracy results in ImageNet-1K top-1 accuracy of 86.8\% with a ViT-L model trained for 150 epochs.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. HyPASE: Hyperbolic Geometry for Parameter-Efficient Speech Emotion Fine-Tuning Framework for Large Audio-Language Models

    cs.SD 2026-08 conditional novelty 5.0 of 10

    A hyperbolic parameter-efficient fine-tuning framework for audio-language models improves speech emotion recognition over Euclidean LoRA and Adapter baselines on MELD and IEMOCAP.

Pith tools