Pith. sign in

REVIEW 11 cited by

HEST-1k: A Dataset for Spatial Transcriptomics and Histology Image Analysis

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2406.16192 v2 pith:RFGRRETS submitted 2024-06-23 cs.CV

classification cs.CV
keywords hest-1kspatialcancercohortshesthest-benchmarkhest-libraryintroduce
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Spatial transcriptomics enables interrogating the molecular composition of tissue with ever-increasing resolution and sensitivity. However, costs, rapidly evolving technology, and lack of standards have constrained computational methods in ST to narrow tasks and small cohorts. In addition, the underlying tissue morphology, as reflected by H&E-stained whole slide images (WSIs), encodes rich information often overlooked in ST studies. Here, we introduce HEST-1k, a collection of 1,229 spatial transcriptomic profiles, each linked to a WSI and extensive metadata. HEST-1k was assembled from 153 public and internal cohorts encompassing 26 organs, two species (Homo Sapiens and Mus Musculus), and 367 cancer samples from 25 cancer types. HEST-1k processing enabled the identification of 2.1 million expression--morphology pairs and over 76 million nuclei. To support its development, we additionally introduce the HEST-Library, a Python package designed to perform a range of actions with HEST samples. We test HEST-1k and Library on three use cases: (1) benchmarking foundation models for pathology (HEST-Benchmark), (2) biomarker exploration, and (3) multimodal representation learning. HEST-1k, HEST-Library, and HEST-Benchmark can be freely accessed at https://github.com/mahmoodlab/hest.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 11 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 22 citations worldwide. Full citation record

  1. Diffusion Generative Modeling for Spatially Resolved Gene Expression Inference from Histology Images

    cs.CV 2025-01 conditional novelty 7.0 of 10

    Stem uses a conditional diffusion model to infer spatially resolved gene expression from H&E histology images, outperforming regression baselines on several datasets.

  2. Sparser2Sparse: Single-shot Sparser-to-Sparse Learning for Spatial Transcriptomics Imputation with Natural Image Co-learning

    cs.CV 2025-07 conditional novelty 6.0 of 10

    S2S-ST reconstructs dense spatial gene expression from 25% sampled spots using self-supervision plus natural-image co-training, and reports better MAE and SSIM than TESLA, BayesSpace, and DIST on eight Xenium samples.

  3. SPATIA: Multimodal Generation and Prediction of Spatial Cell Phenotypes

    q-bio.QM 2025-07 conditional novelty 6.0 of 10

    A hierarchical multimodal model fusing morphology, expression, and spatial context that generates target-state cell morphologies from optimal-transport weak pairs, trained on a 25.9M-cell atlas and benchmarked against...

  4. MeDi: Metadata-Guided Diffusion Models for Mitigating Biases in Tumor Classification

    eess.IV 2025-06 conditional novelty 6.0 of 10

    Conditioning a histopathology diffusion model on tissue source site metadata improves synthetic image fidelity and enhances downstream tumor classification under subpopulation shift.

  5. Scalable Generation of Spatial Transcriptomics from Histology Images via Whole-Slide Flow Matching

    cs.CV 2025-05 conditional novelty 6.0 of 10

    STFlow uses whole-slide flow matching with local spatial attention to jointly predict gene expression across all spots in a histology image, outperforming prior spot-wise and slide-wise baselines on two benchmarks.

  6. From Pixels to Gigapixels: Bridging Local Inductive Bias and Long-Range Dependencies with Pixel-Mamba

    cs.CV 2024-12 conditional novelty 6.0 of 10

    Pixel-Mamba, an end-to-end Mamba-based architecture with progressive token expansion, reports tumor staging and survival scores on three TCGA datasets that match or exceed several pathology foundation models without p...

  7. MERGE: Multi-faceted Hierarchical Graph-based GNN for Gene Expression Prediction from Whole Slide Histopathology Images

    cs.CV 2024-12 conditional novelty 6.0 of 10

    A hierarchical graph built from both spatial and image-feature clustering improves joint gene expression prediction from whole slide histology images over 1-hop GNNs and prior transformer/CNN baselines.

  8. ST-Align: A Multimodal Foundation Model for Image-Gene Alignment in Spatial Transcriptomics

    cs.CV 2024-11 reject novelty 6.0 of 10

    ST-Align aligns pathology images and gene expression at spot and neighborhood scales, improving spatial domain identification and gene prediction in spatial transcriptomics.

  9. Robustifying pathology foundation models via fine-tuning

    cs.CV 2026-07 reject novelty 5.0 of 10

    A uniform fine-tuning step improves acquisition robustness and downstream performance across ten pathology foundation models, but the paper never discloses the fine-tuning recipe.

  10. Accelerating Data Processing and Benchmarking of AI Models for Pathology

    cs.CV 2025-02 conditional novelty 5.0 of 10

    The authors introduce Trident, a WSI processing package, Patho-Bench, a benchmarking library, and 42 curated pathology tasks to standardize foundation model evaluation.

  11. Distilling foundation models for robust and efficient models in digital pathology

    cs.CV 2025-01 conditional novelty 5.0 of 10

    A distilled 86M-parameter pathology model reaches near state-of-the-art performance on EVA and HEST benchmarks and shows strong robustness to scanner and staining variation.

Pith tools