REVIEW 10 cited by
Lung and Colon Cancer Histopathological Image Dataset (LC25000)
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
The field of Machine Learning, a subset of Artificial Intelligence, has led to remarkable advancements in many areas, including medicine. Machine Learning algorithms require large datasets to train computer models successfully. Although there are medical image datasets available, more image datasets are needed from a variety of medical entities, especially cancer pathology. Even more scarce are ML-ready image datasets. To address this need, we created an image dataset (LC25000) with 25,000 color images in 5 classes. Each class contains 5,000 images of the following histologic entities: colon adenocarcinoma, benign colonic tissue, lung adenocarcinoma, lung squamous cell carcinoma, and benign lung tissue. All images are de-identified, HIPAA compliant, validated, and freely available for download to AI researchers.
Forward citations
Cited by 10 Pith papers
-
A Clinically Validated Foundation Model for Comprehensive Lung Pathology Interpretation
PulmoFoundation achieves 92.3% average AUC on 32 lung pathology tasks in prospective validation and raises pathologist accuracy from 83.8% to 91.7% in a crossover RCT.
-
From Multi-Resolution Cells to Gigapixel Whole Slide Images Foundation Model for Computational Pathology
MRPT, a multi-resolution hierarchical transformer pre-trained on 36K whole-slide images, is reported to outperform prior pathology foundation models on 34 classification, captioning, and VQA datasets.
-
ALICE: Learning a General-Purpose Pathology Foundation Model from Vision, Vision-Language, and Slide-Level Experts
Multi-stage agglomerative distillation consolidates eight vision, vision-language, and slide-level pathology teachers into one backbone that ranks first on average across 96 downstream tasks.
-
MedPMC: A Systematic Framework for Scaling High-Fidelity Medical Multimodal Data for Foundation Models
An automated curation pipeline extracts 11M high-fidelity medical image-text pairs from PMC, yielding CLIP and MLLM vision encoders that outperform baselines on 26 benchmarks and a clinical dermatology retrieval task.
-
A Unified Low-level Foundation Model for Enhancing Pathology Image Quality
A prompt-guided diffusion model pretrained on 190 million pathology patches outperforms task-specific models across most restoration and virtual staining benchmarks.
-
Unifying Biomedical Vision-Language Expertise: Towards a Generalist Foundation Model via Multi-CLIP Knowledge Distillation
A student CLIP model distilled from nine medical CLIP teachers outperforms its teachers across most of 58 biomedical benchmarks.
-
Histopathology Multi-modal Embedding for Pathology Composed Retrieval
HOMIE adapts a general multimodal LLM to retrieve pathology cases from queries that interleave images, text, and video, and outperforms existing models on a newly introduced composed retrieval benchmark.
-
Classification based deep learning models for lung cancer and disease using medical images
ResNet+, a ResNet-D and CBAM hybrid, achieves up to 99.25% accuracy on lung cancer CT classification and 98.14% on histopathology, though the reported metrics contain inconsistencies.
-
Caching Techniques for Reducing the Communication Cost of Federated Learning in IoT Environments
A server-side cache that filters and reuses client updates lowers federated learning communication by up to 20 percent while keeping accuracy roughly unchanged, according to the authors' experiments.
-
Multi-Scale Deep Learning for Colon Histopathology: A Hybrid Graph-Transformer Approach
A hybrid CNN-transformer-graph network is reported to reach 96% accuracy on LC25000, but the paper lacks architectural detail, code, and a described data split.
Discussion (0). Continue with ORCID to comment.