Pith. sign in

REVIEW 1 cited by

Blockwise Self-Supervised Learning at Scale

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2302.01647 v2 pith:7NUQYB4L submitted 2023-02-03 cs.CV cs.AIcs.LG

classification cs.CVcs.AIcs.LG
keywords blockwiselearningaccuracybackpropagationself-supervisedend-to-endexplorenetworks
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Current state-of-the-art deep networks are all powered by backpropagation. In this paper, we explore alternatives to full backpropagation in the form of blockwise learning rules, leveraging the latest developments in self-supervised learning. We show that a blockwise pretraining procedure consisting of training independently the 4 main blocks of layers of a ResNet-50 with Barlow Twins' loss function at each block performs almost as well as end-to-end backpropagation on ImageNet: a linear probe trained on top of our blockwise pretrained model obtains a top-1 classification accuracy of 70.48%, only 1.1% below the accuracy of an end-to-end pretrained network (71.57% accuracy). We perform extensive experiments to understand the impact of different components within our method and explore a variety of adaptations of self-supervised learning to the blockwise paradigm, building an exhaustive understanding of the critical avenues for scaling local learning rules to large networks, with implications ranging from hardware design to neuroscience.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. DeInfoReg: A Decoupled Learning Framework for Better Training Throughput

    cs.LG 2025-06 conditional novelty 6.0 of 10

    DeInfoReg trains deep networks with per-module local losses so gradients flow only within each module, improving accuracy and enabling pipeline parallelism, with speedups of up to 1.47x over single-GPU backpropagation.

Pith tools