{"id":"87c6fd35-9956-43e0-ad56-f44fdc675377","arxiv_id":"2608.00100","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"SPARC-Rad is a 300-item multimodal benchmark and evaluation pipeline for spatial and anatomical reasoning in radiology vision-language models.","lead":"SPARC-Rad is a new manually curated benchmark of 300 radiology image-question pairs for testing whether vision-language models can reason about anatomy and spatial relationships in CT, MRI, and X-ray images. The paper describes the benchmark and an evaluation pipeline, but releases no dataset link, no code, and no model baselines.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"No evidence that SPARC-Rad questions require image use: a text-only control is missing, so the benchmark may measure general anatomical knowledge rather than image-grounded spatial reasoning.","rationale":"The reader's weakest assumption focused on annotation reliability and expert adjudication, which is a valid concern but secondary. My concern is more load-bearing: even perfect ground-truth labels would not rescue the benchmark if the questions can be answered without looking at the image. The paper presents no control to rule out text-only priors, so the central claim that SPARC-Rad evaluates image-grounded spatial reasoning rests on an untested assumption. However, this does not require changing the verdict: the paper is already CONDITIONAL on validation and data release. The proposed text-only control is a concrete, feasible validation step that would settle the issue. I therefore agree with the reader's overall conditional assessment (partial agreement on the specific weakest point) and see no reason to move the verdict.","tokens_in":12028,"tokens_out":1880,"duration_ms":25460,"concrete_test":"Obtain the full set of 300 SPARC-Rad question texts (without any images) and run a strong text-only LLM (e.g., GPT-4o or Claude with image input disabled) on them, using the same standardized prompt minus the image. Compare its accuracy against a state-of-the-art VLM evaluated with the actual images. If the text-only model achieves accuracy comparable to the VLM (e.g., within 20 percentage points), SPARC-Rad questions are not image-dependent and the benchmark does not measure visual spatial reasoning. Additionally, report human expert accuracy on the image-grounded questions to calibrate whether the questions are answerable from images at all. This single check directly tests the central construct-validity assumption.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that SPARC-Rad tests image-grounded spatial and anatomical reasoning (Section 1, Section 3.4). For this claim to hold, the 300 questions must be answerable only by inspecting the image, not from common anatomical priors or memorized radiology knowledge. The paper provides no control condition to establish this. It instructs models to 'answer using only the provided image' (Section 4.1) and states that questions were 'intended to require image-grounded reasoning' (Section 3.4), but intent is not evidence. No baseline model results, no text-only LLM comparison, and no human performance are reported. If a text-only LLM can answer a large fraction of questions correctly by using anatomical priors (e.g., 'the heart is left-sided', 'the liver is in the right upper quadrant'), then a model could score well on SPARC-Rad without any genuine visual-spatial ability. This is the most load-bearing threat to construct validity because it would invalidate every downstream claim from the benchmark, regardless of annotation quality. The absence of such a control is not merely a reproducibility gap; it leaves the benchmark's core measurement claim unverified. The paper's limitations (Section 7) acknowledge contamination risk and missing expert adjudication but do not address this text-only-prior confound, which is even more fundamental.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript proposes SPARC-Rad, a manually curated multimodal benchmark dataset and evaluation pipeline for spatial and anatomical reasoning in radiology vision-language models. It describes 300 image-question pairs derived from healthy-control TCIA studies across CT, MRI, and radiography, spanning abdomen, chest, breast, neuro, and musculoskeletal categories. Radiology trainees manually designed questions targeting anatomical identification, localization, laterality, regional recognition, inter-structure spatial relationships, and counting/device recognition. The evaluation pipeline includes standardized prompting, structured output collection, normalization, LLM-as-judge grading, human review, binary correctness scoring, and subgroup analysis. The paper's central claim is that SPARC-Rad isolates image-grounded spatial reasoning and provides a reusable framework for evaluating radiology VLMs.","tokens_in":12301,"tokens_out":4014,"duration_ms":46570,"significance":"If the benchmark were released and its validity demonstrated, SPARC-Rad would fill a genuine gap in radiology VQA benchmarking: most existing datasets do not target spatial and anatomical reasoning as a distinct construct. The manual question design, clinically meaningful reasoning taxonomy, modality/anatomical diversity, and statistical recommendation of McNemar testing are positive features. However, the manuscript currently contains no dataset link, no code, no baseline model results, and no text-only or human-performance controls. The paper itself acknowledges missing inter-rater agreement, board-certified review, and contamination risk (Section 7). These are not mere presentation issues: they leave the core measurement claim—that SPARC-Rad scores reflect image-grounded spatial reasoning—unverified.","major_comments":[{"comment":"The paper refers to a 'provided' dataset (Section 3.2, 'The final dataset that we’ve provided...') and to the benchmark as a usable artifact, but no URL, repository, supplement, or data-availability statement is given. Section 6 reports only distribution statistics. For a benchmark dataset paper, the dataset is the central output; without access, readers cannot verify the counts, inspect the items, run the pipeline, or independently assess the design. This must be fixed by releasing the data and code (or clearly stating access conditions) before the paper can support its claims.","section":"3.2, 3.5, 6"},{"comment":"Ground-truth reliability is not established. Sections 3.4 and 3.5 state that radiology trainees authored questions and reference answers and that quality review occurred, but no inter-rater agreement, board-certified adjudication, or item-level validation is reported. Section 7 explicitly acknowledges this: 'additional board-certified radiologist review, multi-reader adjudication, and formal assessment of inter-rater agreement would further strengthen ground-truth reliability.' Since all downstream scores are computed against these reference answers, label noise directly affects every claim the benchmark can make. This is a load-bearing validity threat, not a minor caveat.","section":"3.4, 7"},{"comment":"No evidence is provided that the questions require image inspection. The prompt instructs models to 'answer using only the provided image' (Section 4.1), and Section 3.4 says questions were 'intended to require image-grounded reasoning,' but intent is not evidence. The paper reports no text-only LLM baseline, no image-ablated control, and no human performance on the questions. Many items—e.g., laterality of the heart or liver position—could be answered from standard anatomical priors. If a text-only model obtains high accuracy, then SPARC-Rad measures priors, not visual-spatial reasoning. This confound is more fundamental than annotation noise and is not addressed in Section 7.","section":"4.1, 3.4"},{"comment":"The LLM-as-judge grading method is not validated. Section 4.3 describes using a separate judging model and human review, but the paper gives no agreement statistics between automated and human grading, no error analysis for the judge, and no protocol for resolving disagreements. Since the primary scoring outcome is binary correctness derived from this judge, grader unreliability would propagate to all benchmark scores. The paper's own Section 7 notes 'grader dependence' and calls for reporting disagreement rates, but no such data appear. This is a central component of the pipeline and must be empirically characterized.","section":"4.3, 6"},{"comment":"The paper reports no baseline model evaluations. Section 6, titled 'Results,' contains only dataset distribution statistics; it does not apply the proposed pipeline to any VLM. The statistical framework in Section 5 (accuracy, confidence intervals, McNemar tests) is described but never demonstrated. For a benchmark and evaluation pipeline paper, at least one example evaluation—with a current VLM, reporting overall and subgroup accuracy, and including a text-only control—is necessary to show that the pipeline is operational and that the questions behave as intended. Without this, the paper remains a design proposal, not a validated benchmark.","section":"6, 5"}],"minor_comments":[{"comment":"The phrase 'human review is incorporated' is vague: no criteria are given for which responses trigger human review, how many reviewers are used, or how disagreements are resolved. This belongs in a concrete protocol.","section":"4.3"},{"comment":"The use of multiple subgroup comparisons with chi-square and McNemar tests should address multiple-testing corrections or explicitly justify why they are not needed.","section":"5"},{"comment":"Reference [11] lacks a year and publication venue; reference [8] (3D-RAD) is cited but not clearly related to the design choice. Please check formatting consistency.","section":"References"},{"comment":"The specific TCIA collections used are not named. Listing them would improve reproducibility without requiring the release of derived images.","section":"3.1"},{"comment":"There is a formatting error in affiliation '1.3' and the affiliation for author 6 appears incomplete. These should be corrected.","section":"Author list"}],"recommendation":"major_revision","confidential_remarks":"This manuscript is a benchmark description without the benchmark artifact or any demonstration of its use. The reader's concern about the missing text-only control is confirmed and, in my view, is the most serious issue: it challenges the construct validity of every downstream score the benchmark could produce. The paper's own Section 7 limitation statements corroborate the need for additional validation. I recommend requiring (1) release of data/code, (2) baseline VLM results with a text-only control, (3) inter-rater reliability and judge-agreement statistics, before reconsideration."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"You should know this paper before deciding anything: it proposes a radiology VQA benchmark aimed squarely at spatial and anatomical reasoning, which is a genuine gap in the existing resources. The taxonomy of reasoning types (laterality, localization, inter-structure relationships, etc.) is clear, and the decision to use healthy-control imaging to avoid pathology confounds is reasonable. The evaluation pipeline is also thoughtfully designed, with accepted-synonym mapping, human review, and stratified analysis. I believe the authors are serious and honest; they list many limitations themselves, including the lack of board-certified review and contamination risk.\n\nBut here is the thing: the paper currently contains no dataset link, no code, and no baseline model evaluation. The central deliverable, the 300 pairs, is invisible. You cannot verify a single claim about the questions or answers. That alone would make me hesitant to call this a benchmark paper yet.\n\nThe more fundamental soft spot, which the paper does not mention, is construct validity. The authors state that questions were 'intended to require image-grounded reasoning,' but intent is not evidence. Without a text-only control, we have no idea how many questions can be answered from anatomical priors (heart on the left, liver on the right). A model could score well just from general medical knowledge. That is a load-bearing issue because it undermines the benchmark's stated purpose. It is not just a reproducibility gap; it is a measurement validity gap.\n\nI also agree with the reader that the small size (300) and healthy-only images limit generalization, but those are secondary and appropriately acknowledged. The LLM-as-judge concern is real but secondary too, and the authors do flag it.\n\nWho is this for? People building radiology VQA benchmarks or evaluating VLMs for spatial reasoning would find the taxonomy and reporting framework useful. But they cannot use the resource yet.\n\nIf I were the editor, I would not desk-reject the idea, but I would not send this version to referees either. The authors need to release the data and code, add a text-only baseline control, and ideally some expert validation. If they do that, a revised version could be a useful peer-reviewed resource. As it stands, this is a well-written proposal, not a verifiable benchmark.","headline":"A sensible idea for a radiology spatial-reasoning benchmark, but the paper ships no data, no code, no baselines, and no text-only control, so right now it is a proposal, not a benchmark.","tokens_in":12821,"tokens_out":2576,"would_cite":false,"duration_ms":34455,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":false},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A manually curated set of 300 radiologic image-question pairs is claimed to isolate spatial and anatomical reasoning in vision-language models, with a reusable evaluation pipeline for subgroup and error analysis.","keywords":["radiology","vision-language models","benchmark dataset","spatial reasoning","anatomical reasoning","visual question answering","medical imaging","evaluation pipeline"],"falsifier":"Give the same 300 questions to a strong text-only language model with the images removed or replaced by blank noise and compare its accuracy with the vision-language models; if its accuracy approaches theirs, the benchmark is measuring textual priors rather than spatial reasoning. A second check: have a panel of board-certified radiologists re-annotate all reference answers; low agreement with the trainee ground truth would show that scores reflect annotator variability.","tokens_in":1255,"feed_emoji":"🩻","tokens_out":1259,"duration_ms":49361,"temperature":0.7,"pith_summary":"The paper sets out to measure a capability that existing radiology benchmarks do not isolate: whether vision-language models can reason about anatomy as a spatial system, identifying structures, determining laterality, localizing devices, and describing spatial relationships. To do this it introduces SPARC-Rad, a manually curated set of 300 image-question pairs drawn from healthy-control CT, MRI, and radiography studies across five anatomical regions. The central claim is that these spatially grounded questions, together with an evaluation pipeline that includes prompt standardization, answer normalization, LLM-as-judge grading, human review, binary correctness scoring, and subgroup analysis by modality, anatomy, body region, and reasoning type, provide a reusable way to profile model behavior rather than reducing it to a single leaderboard number. If the claim is right, the benchmark can reveal clinically important spatial failures, such as laterality reversal, mislocalization, and region confusion, that broad visual-question-answering benchmarks miss.","feed_headline":"300 questions test whether radiology AI sees anatomy spatially","feed_subtitle":"SPARC-Rad isolates laterality, localization, and structure relationships that broad medical VQA benchmarks miss.","key_machinery":"The central mechanism is the pairing of each radiologic image with a manually crafted spatially grounded question and a reference answer, plus an evaluation pipeline that converts free-text model responses into comparable binary correctness labels. Load-bearing choices in the pipeline include: a standardized prompt instructing the model to answer only from the image; answer normalization and accepted-synonym mapping so that clinically equivalent phrasing (for example, an abbreviation versus a full device name) does not create false errors; LLM-as-judge grading against the reference answer; human review for responses that are partially correct but spatially incomplete (especially for laterali","core_discovery":"The paper claims that a compact, manually designed benchmark can target spatial perception and anatomical reasoning as a distinct construct. SPARC-Rad contains 300 image-question pairs: 114 radiographs, 98 CT, and 88 MRI; covering abdomen, chest, breast, neuro, and musculoskeletal categories. Radiology trainees designed the questions to require image-grounded reasoning for anatomical identification, localization, laterality, regional recognition, device identification and counting, and inter-structure spatial relationships. The contribution is not a set of model results but the benchmark itself and its evaluation methodology, which the authors argue supports quantitative model comparison, su","pith_inferences":["If SPARC-Rad's questions are truly image-grounded, an editorially useful validation would be to give the same 300 questions to a strong text-only model with images replaced by blank noise; a narrow accuracy gap between that model and the vision-language models would indicate that linguistic priors, not spatial reasoning, are driving performance.","A natural extension not reported in the paper is item-level psychometric analysis: calibrating question difficulty and computing per-question discrimination would separate ambiguous annotation from genuine model weakness, which the binary scoring alone cannot do.","The reasoning-type taxonomy could be extended to volumetric and longitudinal spatial reasoning, such as cross-slice relationships and multi-temporal device tracking, which are the logical next pressure-tests given the healthy-control baseline design."],"forward_implications":["Models with similar overall accuracy may be distinguished by consistency across modality and anatomy; dispersion measures such as macro-average accuracy, per-category standard deviation, and minimum subgroup accuracy can expose unbalanced failures.","Laterality, device-localization, and inter-structure relationship questions require human adjudication because a response can contain relevant terminology while still being spatially incorrect.","SPARC-Rad performance should be interpreted as a measure of foundational anatomical reasoning on healthy anatomy, not as diagnostic competence in pathological, postoperative, or safety-critical settings.","Reproducible evaluation requires reporting the exact prompt, model name and version, inference date, decoding parameters, and raw outputs for every model run.","Because the source imaging is drawn from public collections, controlled-access test sets or hidden evaluation servers are needed to reduce the risk of pretraining contamination."],"fun_headline_variants":["SPARC-Rad: a spatial IQ test for radiology AI","New benchmark probes radiology AI's spatial anatomy skills","300 questions challenge radiology AI on spatial anatomy","SPARC-Rad: measuring if AI sees anatomy in 3D","Radiology AI spatial reasoning gets its own benchmark"],"cache_read_input_tokens":14208,"weakest_assumption_plain":"The benchmark's validity depends on the assumption that the manually designed questions genuinely require image-grounded spatial reasoning and that the trainee-written reference answers are unambiguous and correct enough that a wrong score reflects model reasoning rather than annotation noise; the paper reports no inter-rater agreement and no board-certified adjudication.","fun_headline_variants_meta":{"raw":{"variants":["SPARC-Rad: a spatial IQ test for radiology AI","New benchmark probes radiology AI's spatial anatomy skills","300 questions challenge radiology AI on spatial anatomy","SPARC-Rad: measuring if AI sees anatomy in 3D","Radiology AI spatial reasoning gets its own benchmark"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000426,"raw_usage":{"total_tokens":2012,"prompt_tokens":733,"completion_tokens":1279,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":477,"completion_tokens_details":{"reasoning_tokens":1211}},"tokens_in":477,"tokens_out":1279,"duration_ms":12692,"temperature":1.0,"reasoning_tokens":1211,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-04T01:16:08.035468+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Give the same 300 questions to a strong text-only language model with the images removed or replaced by blank noise and compare its accuracy with the vision-language models; if its accuracy approaches theirs, the benchmark is measuring textual priors rather than spatial reasoning. A second check: have a panel of board-certified radiologists re-annotate all reference answers; low agreement with the trainee ground truth would show that scores reflect annotator variability.","supporting_citations":[],"review_version":1}