REVIEW 3 major objections 6 minor 20 references
The RSNA Lumbar Degenerative Imaging Spine Classification (LumbarDISC) Dataset
T0 review · 3 major / 6 minor · reviewed 2026-08-07 · deepseek-v4-flash
Pith's one-line read A 2,697-patient MRI dataset now anchors lumbar spine stenosis AI benchmarking.
desk verdict A genuinely useful large lumbar MRI dataset with careful curation, but the single-annotation training labels are the weak link and need an honest caveat or agreement stats. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the annotation schema itself, enforced through a web-based radiology workflow and quality-control scripts. Annotators had to pass a 10-case practice test at 60% agreement before being assigned to one grading task, and an organizer-placed localizer at the L5/S1 disc standardized spinal-level numbering across all studies. Test-set labels were upgraded from single-reader grades to consensus grades by collecting one to three additional annotations until two readers agreed. A stratified split balanced site, age, sex, and severity, and because high-grade disease is rare in nature, the split over-sampled moderate and severe disease at under-represented levels. That machinery is what makes the dataset usable for training without re-curation.
What would settle it
Re-annotate a random sample of the training studies with several expert neuroradiologists, following the same instruction manual, and compute agreement with the released single-reader labels; if agreement at the three-point scale falls well below the 60% threshold used in practice tests, the training ground truth is too noisy for reliable benchmarking.
Extended reading notes
Core claim
The paper's central discovery is a resource: an expert-annotated lumbar spine MRI dataset large enough and diverse enough to train and evaluate deep learning models that grade degenerative narrowing at every intervertebral level. Each study includes a sagittal T2-like series, a sagittal T1 series, and an axial T2 series, and each annotation carries a localizer coordinate marking the center of the graded region. The original four-point severity scale was collapsed to a three-point scale of normal/mild, moderate, and severe for the competition and dataset. Training-study annotations were made once by a single volunteer radiologist, while test-study grades were finalized only after two annotators agreed. The collection deliberately over-samples moderate and severe disease at the usually sparse L1/L2, L2/L3, and L5/S1 levels so models can be evaluated across the full severity range.
Load-bearing premise
The central assumption is that a single volunteer radiologist's grade, made after passing a 60% agreement threshold on ten practice cases, is reliable enough to serve as ground truth for training and benchmarking, since no inter-rater agreement statistics on the actual annotations are reported.
Editorial extensions
If this is right
- Researchers can train deep learning models on thousands of studies to grade spinal canal, neural foraminal, and subarticular stenosis at all five lumbar disc levels.
- A public test set with consensus labels provides a common benchmark, letting different model architectures be compared on the same images.
- The deliberate over-sampling of high-grade disease at L1/L2, L2/L3, and L5/S1 gives models exposure to rare severe cases that natural incidence would under-represent.
- The loose MRI protocol criteria, including multiple T2-like sagittal sequences, should make models trained here more robust to real-world imaging variation.
- The same images can be reused for future research on intervertebral disc and endplate degeneration, which the paper notes were outside the competition scope.
Reading between the lines
- The paper does not report inter-rater agreement for the single-reader training labels; a reader could re-annotate a training subset and produce a useful reliability figure without collecting new data.
- Because test labels are consensus grades, leaderboard results reflect performance against a stronger reference standard than training labels, and benchmark comparisons should state this gap explicitly.
- The localizer coordinates attached to every grade make the dataset usable for detection and localization tasks, not just classification, though the paper does not frame it that way.
- Combining this collection with the closest existing European comparator dataset could let researchers study how imaging protocol and geography affect grading behavior, a question the paper raises implicitly.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper describes the RSNA LumbarDISC dataset: 2,697 adult outpatient lumbar spine MRI studies (8,593 series) from 8 institutions across 6 countries and 5 continents, curated for the RSNA 2024 Lumbar Spine Degenerative Classification challenge. Each study includes sagittal T2-like, sagittal T1, and axial T2 series; stenosis severity (collapsed to normal/mild, moderate, severe from a 4-point scale) is labeled at L1/L2 through L5/S1 for the spinal canal, right and left neural foramina, and right and left subarticular recess, with localizer coordinates for each grade. The paper details inclusion/exclusion criteria, site-specific data extraction, the annotation workflow (volunteer radiologists, a 10-case practice test with a 60% agreement threshold, single-pass training annotations, and multi-annotator consensus for test sets), dataset structure, and a comparison with the Genodisc dataset, claiming this is the largest publicly available dataset of its kind.
Significance. If the resource is as described, this is a substantial contribution to medical imaging machine learning: it is likely the largest public lumbar degeneration MRI dataset with per-level multi-location stenosis grades, offering geographic and protocol diversity, DICOM images, CSV label files, and coordinate localizers. The curation pipeline is described transparently, and the data is accessible via Kaggle and MIRA. The central existence claim is credible, but the dataset's value as benchmark ground truth depends on annotation reliability, which is currently under-documented.
major comments (3)
- [Appendix A, Annotation; Dataset Description and Usage] The training set's single-pass annotations are the load-bearing ground-truth component, but no inter-rater reliability evidence is reported. The manuscript states that each annotator passed a 10-case practice test at '60% agreement with committee consensus' and that 'Data for the training set was annotated only once'; yet it does not report the distribution of practice-test scores, per-annotator or per-class pass rates, or any kappa/AC1 statistic on practice or production data. Given the collapsed 3-point scale's heavy class imbalance (the text reports 85.4% normal/mild for spinal canal grades), a 60% overall-agreement threshold can be passed by an annotator who assigns the majority class to nearly every case. The paper should either provide reliability statistics (e.g., a confusion matrix on the practice set, a subsample of training cases re-annotated by a second reader, or an explicit statement of the intended use and limitation of single-reader labels) or moderate the 'expertly annotated' claim for the training partition.
- [Table 1, Site 7 row] The 'Any location' column for severe narrowing reports 55 cases (58.5%), while the 'Neural Foramen' severe column reports 56 cases (59.5%) for the same site. Since 'any location' is the union of the three locations, it cannot be smaller than a component; this is arithmetically impossible and indicates a data or transcription error. The authors should correct this and run an independent audit of all tables, including Table 2, for internal consistency.
- [Appendix A, Annotation paragraph] The practice-test passing criterion is ambiguous: the text describes grading on a 4-point scale, then states the 4-point scale was 'contracted to three' after annotation, but does not specify whether the 60% passing score was computed on the original 4-point grades or on the collapsed 3-point grades. This distinction matters because the collapsed scale makes the threshold much easier to pass under class imbalance. Please specify the scoring scale and report the actual agreement values (e.g., overall and per-class) achieved by annotators on the 10 practice cases.
minor comments (6)
- [References, item 9] Reference 9 (Battié et al., 'Disc degeneration-related clinical phenotypes') does not appear to be the primary description of the Genodisc dataset, yet the text cites it for the Genodisc count of 2,287 studies; please cite the correct dataset paper or provide a direct source for that count.
- [Appendix A, Annotation] The statement 'Second year neuroradiology fellows that had completed an ACGME-accredited 1 year neuroradiology fellowship also qualified as annotators' is confusing; clarify whether second-year fellows, fully trained neuroradiologists, or both qualified.
- [Table 2 caption and content] Table 2's caption says it reports 'Patient demographics and prevalence of moderate and severe disease across the training, public test, and private test sets,' but the displayed content contains only disease counts by level; please reconcile the caption with the actual table contents.
- [Dataset Description and Usage] The sentence 'For the test datasets, an additional 1 to 3 annotations were acquired until 2 annotators agreed' is ambiguous about the total number of annotations per test case; please specify how many cases required 2, 3, or 4 annotations.
- [Appendix A, Annotation] Please report the number of annotators assigned to each task and summary statistics of annotation workload (e.g., studies per annotator), since the 'crowdsourcing' model is otherwise unquantified.
- [Appendix B] Ethics approvals are mentioned for some sites but not all; please add a consolidated statement on IRB/ethics approval and data-sharing compliance for each contributing institution.
Circularity Check
No significant circularity: the paper is a dataset description with no derived predictions or fitted parameters, and its central existence claim is externally checkable.
full rationale
This manuscript is a dataset descriptor rather than a derivation. Its central claim, that LumbarDISC is the largest publicly available adult MRI lumbar spine degeneration dataset, is supported by direct counts (2,697 patients, 8,593 series) and by an external comparison with the Genodisc dataset (2,287 studies). No equation is derived, no fitted parameter is renamed as a prediction, and no uniqueness theorem or self-citation chain is invoked to force a conclusion. The annotation procedure (Appendix A) does describe a 60% agreement pass threshold on a 10-case practice set and single annotation for training studies, and the absence of reported inter-rater statistics for the training set may be a quality limitation; however, that is an evidentiary or reliability concern, not circularity, because the dataset's ground-truth labels are presented as empirical annotations rather than as outputs of a model or of the paper's own claims. The internal arithmetic inconsistency in Table 1 for Site 7 (56 severe neural foraminal cases versus 55 severe at any location) is likewise a data-consistency issue, not a circular-reasoning issue. The paper is self-contained against external benchmarks in the sense that the dataset is downloadable and auditable through Kaggle and MIRA, so the resource claim is independently falsifiable. No circular step meeting the required standard can be quoted or exhibited.
Assumptions & free parameters
assumptions (3)
- domain assumption The L5/S1 intervertebral disc is identified as the most caudal fully formed disc space, and the localizer method correctly anchors grading levels.
- domain assumption The expert annotators' grades, after collapsing normal/mild, are a valid 3-point ground truth for stenosis severity.
- domain assumption Inclusion/exclusion criteria correctly filter out non-degenerative pathology and postoperative hardware.
Cite this review
Pith. "Pith review of The RSNA Lumbar Degenerative Imaging Spine Classification (LumbarDISC) Dataset." pith.science (2026). https://pith.science/paper/4ZEOQA2M
@misc{pith2026250609162,
author = {Pith},
title = {Pith review of: The RSNA Lumbar Degenerative Imaging Spine Classification (LumbarDISC) Dataset},
year = {2026},
howpublished = {\url{https://pith.science/paper/4ZEOQA2M}},
note = {Machine review of arXiv:2506.09162}
}
read the original abstract
The Radiological Society of North America (RSNA) Lumbar Degenerative Imaging Spine Classification (LumbarDISC) dataset is the largest publicly available dataset of adult MRI lumbar spine examinations annotated for degenerative changes. The dataset includes 2,697 patients with a total of 8,593 image series from 8 institutions across 6 countries and 5 continents. The dataset is available for free for non-commercial use via Kaggle and RSNA Medical Imaging Resource of AI (MIRA). The dataset was created for the RSNA 2024 Lumbar Spine Degenerative Classification competition where competitors developed deep learning models to grade degenerative changes in the lumbar spine. The degree of spinal canal, subarticular recess, and neural foraminal stenosis was graded at each intervertebral disc level in the lumbar spine. The images were annotated by expert volunteer neuroradiologists and musculoskeletal radiologists from the RSNA, American Society of Neuroradiology, and the American Society of Spine Radiology. This dataset aims to facilitate research and development in machine learning and lumbar spine imaging to lead to improved patient care and clinical efficiency.
Figures
Reference graph
Works this paper leans on
-
[1]
The RSNA Lumbar Degenerative Imaging Spine Classification (LumbarDISC) Dataset Authors: Tyler J. Richards, Adam E. Flanders, Errol Colak, Luciano M. Prevedello, Robyn L. Ball, Felipe Kitamura, John Mongan, Maryam Vazirabad, Hui-Ming Lin, Anne Kendell, Thanat Kanthawang, Salita Angkurawaranon, Emre Altinmakas, Hakan Dogan, Paulo Eduardo de Aguiar Kuriki, A...
work page 2024
-
[2]
Fatoye F, Gebrye T, Mbada CE, et al. Clinical and economic burden of low back pain in low- and middle-income countries: a systematic review. BMJ Open 2023;13:e064119
work page 2023
-
[3]
The prevalence of low back pain: a systematic review of the literature from 1966 to
Walker BF. The prevalence of low back pain: a systematic review of the literature from 1966 to
work page 1966
-
[4]
Reliability of Readings of Magnetic Resonance Imaging Features of Lumbar Spinal Stenosis
Lurie JD, Tosteson AN, Tosteson TD, et al. Reliability of Readings of Magnetic Resonance Imaging Features of Lumbar Spinal Stenosis. Spine 2008;33:1605–10
work page 2008
-
[5]
Spinal Stenosis Grading in Magnetic Resonance Imaging Using Deep Convolutional Neural Networks
Won D, Lee H-J, Lee S-J, et al. Spinal Stenosis Grading in Magnetic Resonance Imaging Using Deep Convolutional Neural Networks. Spine 2020;45:804–12
work page 2020
-
[6]
Herzog R, Elgort DR, Flanders AE, et al. Variability in diagnostic error rates of 10 MRI centers performing lumbar spine MRI examinations on the same patient within a 3-week period. Spine J 2017;17:554–61
work page 2017
-
[7]
search their database to identify additional cases of high grade disease at these levels. For the test datasets, an additional 1 to 3 annotations were acquired until 2 annotators agreed on a degree of severity which established the consensus grade. See Figure 3 for severity distributions of high grade stenosis by training and test sets and Appendix A for ...
work page 2024
-
[8]
Causes of failure of surgery on the lumbar spine
Burton CV, Kirkaldy-Willis W, Yong-Hing K, et al. Causes of failure of surgery on the lumbar spine. Clin Orthop Relat Res 1976-2007 1981;157:191–9
work page 1976
Show all 20 references
-
[9]
Could automated machine-learned MRI grading aid epidemiological studies of lumbar spinal stenosis? Validation within the Wakayama spine study
Ishimoto Y, Jamaludin A, Cooper C, et al. Could automated machine-learned MRI grading aid epidemiological studies of lumbar spinal stenosis? Validation within the Wakayama spine study. BMC Musculoskelet Disord 2020;21:158
2020
-
[10]
Deep Learning Model for Automated Detection and Classification of Central Canal, Lateral Recess, and Neural Foraminal Stenosis at Lumbar Spine MRI
Hallinan JTPD, Zhu L, Yang K, et al. Deep Learning Model for Automated Detection and Classification of Central Canal, Lateral Recess, and Neural Foraminal Stenosis at Lumbar Spine MRI. Radiology 2021;300:130–8
2021
-
[12]
Disc degeneration-related clinical phenotypes
Battié MC, Lazáry A, Fairbank J, et al. Disc degeneration-related clinical phenotypes. Eur Spine J Off Publ Eur Spine Soc Eur Spinal Deform Soc Eur Sect Cerv Spine Res Soc 2014;23 Suppl 3:S305-314. Sex At Least Moderate Narrowing Severe Narrowing Site Male Female Age (y) Total...
2014
-
[13]
Flowchart summary of the data acquisition, curation, and annotation process. 0 500 1000 1500 2000 2500 NFSASCNFSASCNFSASCNFSASCNFSASCL1/L2L2/L3L3/L4L4/L5L5/S1 Training Dataset: Incidence of High Grade Spondylosis by Level SevereModerate 050100150200250300350 NFSASCNFSASCNFSASC...
2000
-
[14]
pre-labels
Distribution of moderate and severe disease across the training, public test, and private test sets. The y-axis represents the number of moderate and severe grades by level on either the right or left for neural foramina and subarticular recesses. NF = Neural Foramen; SR = Sub...
2024
-
[15]
The search was restricted to outpatient status, patients 18 years of age or older, and the date range was between January 1, 2022 and August 15,
Chiang Mai University, Thailand The Faculty of Medicine at Chiang Mai University searched their PACS backup archive (Synapse Radiology PACS version 5.7.000; FUJIFILM Medical systems) using RIS (Envision.Net) for MRIs of the lumbar spine without contrast. The search was restric...
2022
-
[17]
The process of evaluation was performed by two experienced radiologists as well as the site's primary investigator
Rigorous analysis of MRI reports was undertaken to confirm that the inclusion criteria was met and cases were excluded according to the exclusion criteria. The process of evaluation was performed by two experienced radiologists as well as the site's primary investigator. Demog...
2023
-
[19]
stenosis
Each study was manually evaluated to ensure that it met the inclusion criteria. For each case, the worst level of spinal canal stenosis, right and left neural foraminal narrowing and right and left subarticular narrowing were recorded. In all cases, data including patient age,...
2019
-
[20]
Afterward, the selected cases were anonymized using RSNA CTP software before being shared
The cases were manually reviewed based on the inclusion criteria. Afterward, the selected cases were anonymized using RSNA CTP software before being shared. University of California San Francisco, USA The primary site investigator searched their institutional RIS (Radiant; Epi...
2023
-
[300]
The studies were downloaded from the PACS in DICOM format and underwent de-identification prior to uploading into the RSNA database
The level of worst spinal canal stenosis, right and left neural foraminal narrowing and right and left subarticular recess narrowing were recorded, in addition to patient age, gender and ethnicity where available. The studies were downloaded from the PACS in DICOM format and u...
2022
-
[1998]
J Spinal Disord 2000;13:205–17
2000
-
[2023]
For each case, the worst level of spinal canal stenosis, right and left neural foramina narrowing and right and left subarticular narrowing were recorded
Each study was manually reviewed by two radiologists to confirm that the inclusion criteria was met and cases were excluded according to the exclusion criteria. For each case, the worst level of spinal canal stenosis, right and left neural foramina narrowing and right and left...
2020
Reviewed August 7, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.