Pith. sign in

REVIEW 7 cited by

RadGPT: Constructing 3D Image-Text Tumor Datasets

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2501.04678 v2 pith:D23JPFQE submitted 2025-01-08 eess.IV cs.CV

classification eess.IVcs.CV
keywords reportstumordatasetsabdomenatlasgenerationmasksradgptreport
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Cancers identified in CT scans are usually accompanied by detailed radiology reports, but publicly available CT datasets often lack these essential reports. This absence limits their usefulness for developing accurate report generation AI. To address this gap, we present AbdomenAtlas 3.0, the first public, high-quality abdominal CT dataset with detailed, expert-reviewed radiology reports. All reports are paired with per-voxel masks and they describe liver, kidney and pancreatic tumors. AbdomenAtlas 3.0 has 9,262 triplets of CT, mask and report--3,955 with tumors. These CT scans come from 17 public datasets. Besides creating the reports for these datasets, we expanded their number of tumor masks by 4.2x, identifying 3,011 new tumor cases. Notably, the reports in AbdomenAtlas 3.0 are more standardized, and generated faster than traditional human-made reports. They provide details like tumor size, location, attenuation and surgical resectability. These reports were created by 12 board-certified radiologists using our proposed RadGPT, a novel framework that converted radiologist-revised tumor segmentation masks into structured and narrative reports. Besides being a dataset creation tool, RadGPT can also become a fully-automatic, segmentation-assisted report generation method. We benchmarked this method and 5 state-of-the-art report generation vision-language models. Our results show that segmentation strongly improves tumor detection in AI-made reports.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 7 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Learning Segmentation from Radiology Reports

    eess.IV 2025-07 conditional novelty 7.0 of 10

    R-Super converts tumor count, size, and location information from radiology reports into voxel-wise losses that improve CT tumor segmentation beyond training with masks alone.

  2. Are Vision Language Models Ready for Clinical Diagnosis? A 3D Medical Benchmark for Tumor-centric Visual Question Answering

    cs.CV 2025-05 conditional novelty 7.0 of 10

    DeepTumorVQA, a 9,262-volume 3D medical VQA benchmark, shows that current vision-language models handle measurement but remain far from clinical-grade lesion recognition and reasoning.

  3. PanTS: The Pancreatic Tumor Segmentation Dataset

    eess.IV 2025-07 conditional novelty 6.0 of 10

    PanTS is a new large CT dataset with expert-drawn pancreatic tumor and anatomy labels, and models trained on it beat prior public benchmarks.

  4. ShapeKit

    eess.IV 2025-06 reject novelty 5.0 of 10

    ShapeKit, a rule-based post-processing toolkit, reports Dice score improvements of up to 8.8 percentage points on two CT datasets without retraining the segmentation model.

  5. ${\mu}^2$Tokenizer: Differentiable Multi-Scale Multi-Modal Tokenizer for Radiology Report Generation

    cs.LG 2025-06 reject novelty 4.0 of 10

    A tokenizer that combines multi-scale CT image features with text questions, plus DPO training on a clinical metric, is claimed to improve automated radiology report generation.

  6. Region-Aware Multimodal Large Language Model via SlowFast Tokenization and Pseudo-Mask Guidance for 3D CT Report Generation

    eess.IV 2025-06 conditional novelty 4.0 of 10

    MedRegion-CT integrates region-representative tokens, mask-driven segmentation tokens, and patient-specific attribute prompts into a multimodal LLM, reporting state-of-the-art scores on RadGenome-Chest CT report generation.

  7. Domain Specific Benchmarks for Evaluating Multimodal Large Language Models

    cs.LG 2025-06 conditional novelty 3.0 of 10

    A review paper that organizes domain-specific MLLM benchmarks into an eight-discipline taxonomy, with summary tables and performance highlights.

Pith tools