REVIEW 5 cited by
ST-Align: A Multimodal Foundation Model for Image-Gene Alignment in Spatial Transcriptomics
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Spatial transcriptomics (ST) provides high-resolution pathological images and whole-transcriptomic expression profiles at individual spots across whole-slide scales. This setting makes it an ideal data source to develop multimodal foundation models. Although recent studies attempted to fine-tune visual encoders with trainable gene encoders based on spot-level, the absence of a wider slide perspective and spatial intrinsic relationships limits their ability to capture ST-specific insights effectively. Here, we introduce ST-Align, the first foundation model designed for ST that deeply aligns image-gene pairs by incorporating spatial context, effectively bridging pathological imaging with genomic features. We design a novel pretraining framework with a three-target alignment strategy for ST-Align, enabling (1) multi-scale alignment across image-gene pairs, capturing both spot- and niche-level contexts for a comprehensive perspective, and (2) cross-level alignment of multimodal insights, connecting localized cellular characteristics and broader tissue architecture. Additionally, ST-Align employs specialized encoders tailored to distinct ST contexts, followed by an Attention-Based Fusion Network (ABFN) for enhanced multimodal fusion, effectively merging domain-shared knowledge with ST-specific insights from both pathological and genomic data. We pre-trained ST-Align on 1.3 million spot-niche pairs and evaluated its performance through two downstream tasks across six datasets, demonstrating superior zero-shot and few-shot capabilities. ST-Align highlights the potential for reducing the cost of ST and providing valuable insights into the distinction of critical compositions within human tissue.
Forward citations
Cited by 5 Pith papers
-
Gene Ontology-Guided Hierarchical Spatial Gene Expression Prediction from Histopathology Images
MSGR's Gene Ontology-guided hierarchical decoder improves spatial gene expression prediction from histology images, with the biological structure adding a +0.027 gain over an equivalent random hierarchy.
-
HierarchicalDAEW: Domain-Aware Edge-Weighted Graph Convolution with Evidential Uncertainty for Multi-Section Spatial Gene Expression Prediction from H&E Histology
HierarchicalDAEW predicts spatial gene expression with expression-derived domain edge typing and calibrated uncertainty, beating 13 baselines on breast Visium sections.
-
SToFM: a Multi-scale Foundation Model for Spatial Transcriptomics
SToFM pretrains an SE(2) Transformer on multi-scale sub-slices of spatial transcriptomics data, using virtual cells to capture tissue structure, and reports state-of-the-art results on several ST benchmarks.
-
SPATIA: Multimodal Generation and Prediction of Spatial Cell Phenotypes
A hierarchical multimodal model fusing morphology, expression, and spatial context that generates target-state cell morphologies from optimal-transport weak pairs, trained on a 25.9M-cell atlas and benchmarked against...
-
Spatial Transcriptomics as Images for Large-Scale Pretraining
Cropping ST slides into fixed multi-channel gene patches preserves local spatial context, multiplies training samples, and beats spot- and slice-based pretraining on domain detection.
Discussion (0). Continue with ORCID to comment.