Pith. sign in

REVIEW 3 major objections 6 minor 20 references

The RSNA Lumbar Degenerative Imaging Spine Classification (LumbarDISC) Dataset

T0 review · 3 major / 6 minor · reviewed 2026-08-07 · deepseek-v4-flash

Pith's one-line read A 2,697-patient MRI dataset now anchors lumbar spine stenosis AI benchmarking.

desk verdict A genuinely useful large lumbar MRI dataset with careful curation, but the single-annotation training labels are the weak link and need an honest caveat or agreement stats. read the letter →

arxiv 2506.09162 v1 pith:4ZEOQA2M submitted 2025-06-10 eess.IV cs.CV

classification eess.IVcs.CV
keywords lumbarspineMRIspinalstenosisgradingdegenerativediseasemedicalimagingdatasetmachinelearningbenchmarkradiologyannotationDISC
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper establishes the LumbarDISC dataset: the largest publicly available collection of adult MRI lumbar spine examinations annotated for degenerative narrowing. It contains 2,697 patients and 8,593 image series from eight institutions across six countries and five continents, with stenosis grades at five disc levels for the spinal canal, right and left neural foramina, and right and left subarticular recesses. The dataset was built to support machine-learning model development and benchmarking for lumbar spine stenosis grading, a task where previous public datasets have been too small or too narrow in disease diversity. If the paper's claims hold, researchers gain a free, geographically diverse resource that previous efforts lacked.

What carries the argument

The load-bearing object is the annotation schema itself, enforced through a web-based radiology workflow and quality-control scripts. Annotators had to pass a 10-case practice test at 60% agreement before being assigned to one grading task, and an organizer-placed localizer at the L5/S1 disc standardized spinal-level numbering across all studies. Test-set labels were upgraded from single-reader grades to consensus grades by collecting one to three additional annotations until two readers agreed. A stratified split balanced site, age, sex, and severity, and because high-grade disease is rare in nature, the split over-sampled moderate and severe disease at under-represented levels. That machinery is what makes the dataset usable for training without re-curation.

What would settle it

Re-annotate a random sample of the training studies with several expert neuroradiologists, following the same instruction manual, and compute agreement with the released single-reader labels; if agreement at the three-point scale falls well below the 60% threshold used in practice tests, the training ground truth is too noisy for reliable benchmarking.

Watch

Extended reading notes

Core claim

The paper's central discovery is a resource: an expert-annotated lumbar spine MRI dataset large enough and diverse enough to train and evaluate deep learning models that grade degenerative narrowing at every intervertebral level. Each study includes a sagittal T2-like series, a sagittal T1 series, and an axial T2 series, and each annotation carries a localizer coordinate marking the center of the graded region. The original four-point severity scale was collapsed to a three-point scale of normal/mild, moderate, and severe for the competition and dataset. Training-study annotations were made once by a single volunteer radiologist, while test-study grades were finalized only after two annotators agreed. The collection deliberately over-samples moderate and severe disease at the usually sparse L1/L2, L2/L3, and L5/S1 levels so models can be evaluated across the full severity range.

Load-bearing premise

The central assumption is that a single volunteer radiologist's grade, made after passing a 60% agreement threshold on ten practice cases, is reliable enough to serve as ground truth for training and benchmarking, since no inter-rater agreement statistics on the actual annotations are reported.

Editorial extensions

If this is right

  • Researchers can train deep learning models on thousands of studies to grade spinal canal, neural foraminal, and subarticular stenosis at all five lumbar disc levels.
  • A public test set with consensus labels provides a common benchmark, letting different model architectures be compared on the same images.
  • The deliberate over-sampling of high-grade disease at L1/L2, L2/L3, and L5/S1 gives models exposure to rare severe cases that natural incidence would under-represent.
  • The loose MRI protocol criteria, including multiple T2-like sagittal sequences, should make models trained here more robust to real-world imaging variation.
  • The same images can be reused for future research on intervertebral disc and endplate degeneration, which the paper notes were outside the competition scope.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The paper does not report inter-rater agreement for the single-reader training labels; a reader could re-annotate a training subset and produce a useful reliability figure without collecting new data.
  • Because test labels are consensus grades, leaderboard results reflect performance against a stronger reference standard than training labels, and benchmark comparisons should state this gap explicitly.
  • The localizer coordinates attached to every grade make the dataset usable for detection and localization tasks, not just classification, though the paper does not frame it that way.
  • Combining this collection with the closest existing European comparator dataset could let researchers study how imaging protocol and geography affect grading behavior, a question the paper raises implicitly.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 6 minor

Summary. The paper describes the RSNA LumbarDISC dataset: 2,697 adult outpatient lumbar spine MRI studies (8,593 series) from 8 institutions across 6 countries and 5 continents, curated for the RSNA 2024 Lumbar Spine Degenerative Classification challenge. Each study includes sagittal T2-like, sagittal T1, and axial T2 series; stenosis severity (collapsed to normal/mild, moderate, severe from a 4-point scale) is labeled at L1/L2 through L5/S1 for the spinal canal, right and left neural foramina, and right and left subarticular recess, with localizer coordinates for each grade. The paper details inclusion/exclusion criteria, site-specific data extraction, the annotation workflow (volunteer radiologists, a 10-case practice test with a 60% agreement threshold, single-pass training annotations, and multi-annotator consensus for test sets), dataset structure, and a comparison with the Genodisc dataset, claiming this is the largest publicly available dataset of its kind.

Significance. If the resource is as described, this is a substantial contribution to medical imaging machine learning: it is likely the largest public lumbar degeneration MRI dataset with per-level multi-location stenosis grades, offering geographic and protocol diversity, DICOM images, CSV label files, and coordinate localizers. The curation pipeline is described transparently, and the data is accessible via Kaggle and MIRA. The central existence claim is credible, but the dataset's value as benchmark ground truth depends on annotation reliability, which is currently under-documented.

major comments (3)
  1. [Appendix A, Annotation; Dataset Description and Usage] The training set's single-pass annotations are the load-bearing ground-truth component, but no inter-rater reliability evidence is reported. The manuscript states that each annotator passed a 10-case practice test at '60% agreement with committee consensus' and that 'Data for the training set was annotated only once'; yet it does not report the distribution of practice-test scores, per-annotator or per-class pass rates, or any kappa/AC1 statistic on practice or production data. Given the collapsed 3-point scale's heavy class imbalance (the text reports 85.4% normal/mild for spinal canal grades), a 60% overall-agreement threshold can be passed by an annotator who assigns the majority class to nearly every case. The paper should either provide reliability statistics (e.g., a confusion matrix on the practice set, a subsample of training cases re-annotated by a second reader, or an explicit statement of the intended use and limitation of single-reader labels) or moderate the 'expertly annotated' claim for the training partition.
  2. [Table 1, Site 7 row] The 'Any location' column for severe narrowing reports 55 cases (58.5%), while the 'Neural Foramen' severe column reports 56 cases (59.5%) for the same site. Since 'any location' is the union of the three locations, it cannot be smaller than a component; this is arithmetically impossible and indicates a data or transcription error. The authors should correct this and run an independent audit of all tables, including Table 2, for internal consistency.
  3. [Appendix A, Annotation paragraph] The practice-test passing criterion is ambiguous: the text describes grading on a 4-point scale, then states the 4-point scale was 'contracted to three' after annotation, but does not specify whether the 60% passing score was computed on the original 4-point grades or on the collapsed 3-point grades. This distinction matters because the collapsed scale makes the threshold much easier to pass under class imbalance. Please specify the scoring scale and report the actual agreement values (e.g., overall and per-class) achieved by annotators on the 10 practice cases.
minor comments (6)
  1. [References, item 9] Reference 9 (Battié et al., 'Disc degeneration-related clinical phenotypes') does not appear to be the primary description of the Genodisc dataset, yet the text cites it for the Genodisc count of 2,287 studies; please cite the correct dataset paper or provide a direct source for that count.
  2. [Appendix A, Annotation] The statement 'Second year neuroradiology fellows that had completed an ACGME-accredited 1 year neuroradiology fellowship also qualified as annotators' is confusing; clarify whether second-year fellows, fully trained neuroradiologists, or both qualified.
  3. [Table 2 caption and content] Table 2's caption says it reports 'Patient demographics and prevalence of moderate and severe disease across the training, public test, and private test sets,' but the displayed content contains only disease counts by level; please reconcile the caption with the actual table contents.
  4. [Dataset Description and Usage] The sentence 'For the test datasets, an additional 1 to 3 annotations were acquired until 2 annotators agreed' is ambiguous about the total number of annotations per test case; please specify how many cases required 2, 3, or 4 annotations.
  5. [Appendix A, Annotation] Please report the number of annotators assigned to each task and summary statistics of annotation workload (e.g., studies per annotator), since the 'crowdsourcing' model is otherwise unquantified.
  6. [Appendix B] Ethics approvals are mentioned for some sites but not all; please add a consolidated statement on IRB/ethics approval and data-sharing compliance for each contributing institution.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the paper is a dataset description with no derived predictions or fitted parameters, and its central existence claim is externally checkable.

full rationale

This manuscript is a dataset descriptor rather than a derivation. Its central claim, that LumbarDISC is the largest publicly available adult MRI lumbar spine degeneration dataset, is supported by direct counts (2,697 patients, 8,593 series) and by an external comparison with the Genodisc dataset (2,287 studies). No equation is derived, no fitted parameter is renamed as a prediction, and no uniqueness theorem or self-citation chain is invoked to force a conclusion. The annotation procedure (Appendix A) does describe a 60% agreement pass threshold on a 10-case practice set and single annotation for training studies, and the absence of reported inter-rater statistics for the training set may be a quality limitation; however, that is an evidentiary or reliability concern, not circularity, because the dataset's ground-truth labels are presented as empirical annotations rather than as outputs of a model or of the paper's own claims. The internal arithmetic inconsistency in Table 1 for Site 7 (56 severe neural foraminal cases versus 55 severe at any location) is likewise a data-consistency issue, not a circular-reasoning issue. The paper is self-contained against external benchmarks in the sense that the dataset is downloadable and auditable through Kaggle and MIRA, so the resource claim is independently falsifiable. No circular step meeting the required standard can be quoted or exhibited.

Assumptions & free parameters 0 free parameters · 3 assumptions · 0 invented entities

The central claim is a description of a new dataset. The main axioms are the reliability of expert annotations and the correctness of the L5/S1 localization convention. No free parameters or invented theoretical entities are introduced.

assumptions (3)
  • domain assumption The L5/S1 intervertebral disc is identified as the most caudal fully formed disc space, and the localizer method correctly anchors grading levels.
    Used to standardize labeling across all studies; if incorrect, all level-specific grades shift (Appendix A).
  • domain assumption The expert annotators' grades, after collapsing normal/mild, are a valid 3-point ground truth for stenosis severity.
    The dataset's utility assumes annotation reliability; the paper does not report inter-rater agreement for the training set (Appendix A).
  • domain assumption Inclusion/exclusion criteria correctly filter out non-degenerative pathology and postoperative hardware.
    The criteria rely on manual review and report screening at each site; errors could contaminate labels (Appendix B).

how reviews work

0 comments
Cite this review

Pith. "Pith review of The RSNA Lumbar Degenerative Imaging Spine Classification (LumbarDISC) Dataset." pith.science (2026). https://pith.science/paper/4ZEOQA2M

@misc{pith2026250609162,
  author       = {Pith},
  title        = {Pith review of: The RSNA Lumbar Degenerative Imaging Spine Classification (LumbarDISC) Dataset},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/4ZEOQA2M}},
  note         = {Machine review of arXiv:2506.09162}
}
read the original abstract

The Radiological Society of North America (RSNA) Lumbar Degenerative Imaging Spine Classification (LumbarDISC) dataset is the largest publicly available dataset of adult MRI lumbar spine examinations annotated for degenerative changes. The dataset includes 2,697 patients with a total of 8,593 image series from 8 institutions across 6 countries and 5 continents. The dataset is available for free for non-commercial use via Kaggle and RSNA Medical Imaging Resource of AI (MIRA). The dataset was created for the RSNA 2024 Lumbar Spine Degenerative Classification competition where competitors developed deep learning models to grade degenerative changes in the lumbar spine. The degree of spinal canal, subarticular recess, and neural foraminal stenosis was graded at each intervertebral disc level in the lumbar spine. The images were annotated by expert volunteer neuroradiologists and musculoskeletal radiologists from the RSNA, American Society of Neuroradiology, and the American Society of Spine Radiology. This dataset aims to facilitate research and development in machine learning and lumbar spine imaging to lead to improved patient care and clinical efficiency.

Figures

Figures reproduced from arXiv: 2506.09162 by the authors.

Figure 1
Figure 1. Images A (sagittal T2) and B (sagittal STIR) demonstrate the location of [PITH_FULL_IMAGE:figures/full_fig_p011_1.png] view at source ↗
Figure 2
Figure 2. Flowchart summary of the data acquisition, curation, and annotation [PITH_FULL_IMAGE:figures/full_fig_p012_2.png] view at source ↗
Figure 3
Figure 3. Distribution of moderate and severe disease across the training, public [PITH_FULL_IMAGE:figures/full_fig_p015_3.png] view at source ↗

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

20 extracted references · 20 canonical work pages

  1. [1]

    Richards, Adam E

    The RSNA Lumbar Degenerative Imaging Spine Classification (LumbarDISC) Dataset Authors: Tyler J. Richards, Adam E. Flanders, Errol Colak, Luciano M. Prevedello, Robyn L. Ball, Felipe Kitamura, John Mongan, Maryam Vazirabad, Hui-Ming Lin, Anne Kendell, Thanat Kanthawang, Salita Angkurawaranon, Emre Altinmakas, Hakan Dogan, Paulo Eduardo de Aguiar Kuriki, A...

  2. [2]

    Clinical and economic burden of low back pain in low- and middle-income countries: a systematic review

    Fatoye F, Gebrye T, Mbada CE, et al. Clinical and economic burden of low back pain in low- and middle-income countries: a systematic review. BMJ Open 2023;13:e064119

  3. [3]

    The prevalence of low back pain: a systematic review of the literature from 1966 to

    Walker BF. The prevalence of low back pain: a systematic review of the literature from 1966 to

  4. [4]

    Reliability of Readings of Magnetic Resonance Imaging Features of Lumbar Spinal Stenosis

    Lurie JD, Tosteson AN, Tosteson TD, et al. Reliability of Readings of Magnetic Resonance Imaging Features of Lumbar Spinal Stenosis. Spine 2008;33:1605–10

  5. [5]

    Spinal Stenosis Grading in Magnetic Resonance Imaging Using Deep Convolutional Neural Networks

    Won D, Lee H-J, Lee S-J, et al. Spinal Stenosis Grading in Magnetic Resonance Imaging Using Deep Convolutional Neural Networks. Spine 2020;45:804–12

  6. [6]

    Variability in diagnostic error rates of 10 MRI centers performing lumbar spine MRI examinations on the same patient within a 3-week period

    Herzog R, Elgort DR, Flanders AE, et al. Variability in diagnostic error rates of 10 MRI centers performing lumbar spine MRI examinations on the same patient within a 3-week period. Spine J 2017;17:554–61

  7. [7]

    For the test datasets, an additional 1 to 3 annotations were acquired until 2 annotators agreed on a degree of severity which established the consensus grade

    search their database to identify additional cases of high grade disease at these levels. For the test datasets, an additional 1 to 3 annotations were acquired until 2 annotators agreed on a degree of severity which established the consensus grade. See Figure 3 for severity distributions of high grade stenosis by training and test sets and Appendix A for ...

  8. [8]

    Causes of failure of surgery on the lumbar spine

    Burton CV, Kirkaldy-Willis W, Yong-Hing K, et al. Causes of failure of surgery on the lumbar spine. Clin Orthop Relat Res 1976-2007 1981;157:191–9

Show all 20 references
  1. [9]

    Could automated machine-learned MRI grading aid epidemiological studies of lumbar spinal stenosis? Validation within the Wakayama spine study

    Ishimoto Y, Jamaludin A, Cooper C, et al. Could automated machine-learned MRI grading aid epidemiological studies of lumbar spinal stenosis? Validation within the Wakayama spine study. BMC Musculoskelet Disord 2020;21:158

  2. [10]

    Deep Learning Model for Automated Detection and Classification of Central Canal, Lateral Recess, and Neural Foraminal Stenosis at Lumbar Spine MRI

    Hallinan JTPD, Zhu L, Yang K, et al. Deep Learning Model for Automated Detection and Classification of Central Canal, Lateral Recess, and Neural Foraminal Stenosis at Lumbar Spine MRI. Radiology 2021;300:130–8

  3. [12]

    Disc degeneration-related clinical phenotypes

    Battié MC, Lazáry A, Fairbank J, et al. Disc degeneration-related clinical phenotypes. Eur Spine J Off Publ Eur Spine Soc Eur Spinal Deform Soc Eur Sect Cerv Spine Res Soc 2014;23 Suppl 3:S305-314. Sex At Least Moderate Narrowing Severe Narrowing Site Male Female Age (y) Total...

  4. [13]

    Flowchart summary of the data acquisition, curation, and annotation process. 0 500 1000 1500 2000 2500 NFSASCNFSASCNFSASCNFSASCNFSASCL1/L2L2/L3L3/L4L4/L5L5/S1 Training Dataset: Incidence of High Grade Spondylosis by Level SevereModerate 050100150200250300350 NFSASCNFSASCNFSASC...

  5. [14]

    pre-labels

    Distribution of moderate and severe disease across the training, public test, and private test sets. The y-axis represents the number of moderate and severe grades by level on either the right or left for neural foramina and subarticular recesses. NF = Neural Foramen; SR = Sub...

  6. [15]

    The search was restricted to outpatient status, patients 18 years of age or older, and the date range was between January 1, 2022 and August 15,

    Chiang Mai University, Thailand The Faculty of Medicine at Chiang Mai University searched their PACS backup archive (Synapse Radiology PACS version 5.7.000; FUJIFILM Medical systems) using RIS (Envision.Net) for MRIs of the lumbar spine without contrast. The search was restric...

  7. [17]

    The process of evaluation was performed by two experienced radiologists as well as the site's primary investigator

    Rigorous analysis of MRI reports was undertaken to confirm that the inclusion criteria was met and cases were excluded according to the exclusion criteria. The process of evaluation was performed by two experienced radiologists as well as the site's primary investigator. Demog...

  8. [19]

    stenosis

    Each study was manually evaluated to ensure that it met the inclusion criteria. For each case, the worst level of spinal canal stenosis, right and left neural foraminal narrowing and right and left subarticular narrowing were recorded. In all cases, data including patient age,...

  9. [20]

    Afterward, the selected cases were anonymized using RSNA CTP software before being shared

    The cases were manually reviewed based on the inclusion criteria. Afterward, the selected cases were anonymized using RSNA CTP software before being shared. University of California San Francisco, USA The primary site investigator searched their institutional RIS (Radiant; Epi...

  10. [300]

    The studies were downloaded from the PACS in DICOM format and underwent de-identification prior to uploading into the RSNA database

    The level of worst spinal canal stenosis, right and left neural foraminal narrowing and right and left subarticular recess narrowing were recorded, in addition to patient age, gender and ethnicity where available. The studies were downloaded from the PACS in DICOM format and u...

  11. [1998]

    J Spinal Disord 2000;13:205–17

  12. [2023]

    For each case, the worst level of spinal canal stenosis, right and left neural foramina narrowing and right and left subarticular narrowing were recorded

    Each study was manually reviewed by two radiologists to confirm that the inclusion criteria was met and cases were excluded according to the exclusion criteria. For each case, the worst level of spinal canal stenosis, right and left neural foramina narrowing and right and left...

Pith tools

Reviewed August 7, 2026 · model on record in the stance chip above.