Pith. sign in

REVIEW 3 cited by

ViTally Consistent: Scaling Biological Representation Learning for Cell Microscopy

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2411.02572 v2 pith:CSV2ZIJV submitted 2024-11-04 cs.LG cs.AIcs.CV

classification cs.LGcs.AIcs.CV
keywords microscopybiologicalcellimagesperturbationsbestbiologicallyblocks
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Large-scale cell microscopy screens are used in drug discovery and molecular biology research to study the effects of millions of chemical and genetic perturbations on cells. To use these images in downstream analysis, we need models that can map each image into a feature space that represents diverse biological phenotypes consistently, in the sense that perturbations with similar biological effects have similar representations. In this work, we present the largest foundation model for cell microscopy data to date, a new 1.9 billion-parameter ViT-G/8 MAE trained on over 8 billion microscopy image crops. Compared to a previous published ViT-L/8 MAE, our new model achieves a 60% improvement in linear separability of genetic perturbations and obtains the best overall performance on whole-genome biological relationship recall and replicate consistency benchmarks. Beyond scaling, we developed two key methods that improve performance: (1) training on a curated and diverse dataset; and, (2) using biologically motivated linear probing tasks to search across each transformer block for the best candidate representation of whole-genome screens. We find that many self-supervised vision transformers, pretrained on either natural or microscopy images, yield significantly more biologically meaningful representations of microscopy images in their intermediate blocks than in their typically used final blocks. More broadly, our approach and results provide insights toward a general strategy for successfully building foundation models for large-scale biological data.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. A Foundation Model for Spatial Proteomics

    cs.CV 2025-06 conditional novelty 6.0 of 10

    KRONOS, a self-supervised foundation model for spatial proteomics, outperforms existing vision models on cell phenotyping, retrieval, region classification, and label-efficient tasks across 11 cohorts.

  2. C3R: Channel Conditioned Cell Representations for unified evaluation in microscopy imaging

    cs.CV 2025-05 conditional novelty 6.0 of 10

    C3R uses a context-concept channel split plus masked context distillation to enable zero-shot cross-dataset cell representation learning.

  3. Fast Vision Mamba: Pooling Spatial Dimensions for Accelerated Processing

    cs.CV 2025-02 conditional novelty 5.0 of 10

    FastVim reduces Vision Mamba's SSM parallel scan steps from log(h^2) to log(h) by alternately mean-pooling tokens across rows or columns, delivering up to a 72.5% inference speedup at 2048x2048 with roughly unchanged ...

Pith tools