Pith. sign in

REVIEW 10 cited by

Lung and Colon Cancer Histopathological Image Dataset (LC25000)

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1912.12142 v1 pith:7FOQNN7R submitted 2019-12-16 eess.IV cs.CVq-bio.QM

classification eess.IVcs.CVq-bio.QM
keywords imagedatasetslungimagesadenocarcinomaavailablebenigncancer
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The field of Machine Learning, a subset of Artificial Intelligence, has led to remarkable advancements in many areas, including medicine. Machine Learning algorithms require large datasets to train computer models successfully. Although there are medical image datasets available, more image datasets are needed from a variety of medical entities, especially cancer pathology. Even more scarce are ML-ready image datasets. To address this need, we created an image dataset (LC25000) with 25,000 color images in 5 classes. Each class contains 5,000 images of the following histologic entities: colon adenocarcinoma, benign colonic tissue, lung adenocarcinoma, lung squamous cell carcinoma, and benign lung tissue. All images are de-identified, HIPAA compliant, validated, and freely available for download to AI researchers.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 10 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. A Clinically Validated Foundation Model for Comprehensive Lung Pathology Interpretation

    eess.IV 2026-05 unverdicted novelty 7.0 of 10

    PulmoFoundation achieves 92.3% average AUC on 32 lung pathology tasks in prospective validation and raises pathologist accuracy from 83.8% to 91.7% in a crossover RCT.

  2. From Multi-Resolution Cells to Gigapixel Whole Slide Images Foundation Model for Computational Pathology

    cs.CV 2026-08 conditional novelty 6.0 of 10

    MRPT, a multi-resolution hierarchical transformer pre-trained on 36K whole-slide images, is reported to outperform prior pathology foundation models on 34 classification, captioning, and VQA datasets.

  3. ALICE: Learning a General-Purpose Pathology Foundation Model from Vision, Vision-Language, and Slide-Level Experts

    cs.CV 2026-07 accept novelty 6.0 of 10

    Multi-stage agglomerative distillation consolidates eight vision, vision-language, and slide-level pathology teachers into one backbone that ranks first on average across 96 downstream tasks.

  4. MedPMC: A Systematic Framework for Scaling High-Fidelity Medical Multimodal Data for Foundation Models

    cs.CV 2026-07 accept novelty 6.0 of 10

    An automated curation pipeline extracts 11M high-fidelity medical image-text pairs from PMC, yielding CLIP and MLLM vision encoders that outperform baselines on 26 benchmarks and a clinical dermatology retrieval task.

  5. A Unified Low-level Foundation Model for Enhancing Pathology Image Quality

    cs.CV 2025-09 conditional novelty 6.0 of 10

    A prompt-guided diffusion model pretrained on 190 million pathology patches outperforms task-specific models across most restoration and virtual staining benchmarks.

  6. Unifying Biomedical Vision-Language Expertise: Towards a Generalist Foundation Model via Multi-CLIP Knowledge Distillation

    cs.CV 2025-06 conditional novelty 6.0 of 10

    A student CLIP model distilled from nine medical CLIP teachers outperforms its teachers across most of 58 biomedical benchmarks.

  7. Histopathology Multi-modal Embedding for Pathology Composed Retrieval

    cs.CV 2025-02 conditional novelty 5.0 of 10

    HOMIE adapts a general multimodal LLM to retrieve pathology cases from queries that interleave images, text, and video, and outperforms existing models on a newly introduced composed retrieval benchmark.

  8. Classification based deep learning models for lung cancer and disease using medical images

    eess.IV 2025-07 reject novelty 4.0 of 10

    ResNet+, a ResNet-D and CBAM hybrid, achieves up to 99.25% accuracy on lung cancer CT classification and 98.14% on histopathology, though the reported metrics contain inconsistencies.

  9. Caching Techniques for Reducing the Communication Cost of Federated Learning in IoT Environments

    cs.DC 2025-07 reject novelty 3.0 of 10

    A server-side cache that filters and reuses client updates lowers federated learning communication by up to 20 percent while keeping accuracy roughly unchanged, according to the authors' experiments.

  10. Multi-Scale Deep Learning for Colon Histopathology: A Hybrid Graph-Transformer Approach

    cs.CV 2025-09 reject novelty 2.0 of 10

    A hybrid CNN-transformer-graph network is reported to reach 96% accuracy on LC25000, but the paper lacks architectural detail, code, and a described data split.

Pith tools