{"id":"81ca83a4-493c-40e1-bd1d-34a498cc86c1","arxiv_id":"2607.29253","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"A new public dataset of 1,764 expert-annotated CBCT slices plus a 27-metric image-quality benchmark is presented.","lead":"This paper releases a public dataset of 1,764 cone-beam CT image slices from two body phantoms, each scored by three experts for image quality. It provides a benchmark of 27 quality-assessment measures, giving researchers a shared testbed for automated CBCT image-quality evaluation.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Reference assumption contradicted: the Normal reconstruction used as 'high-quality' reference is reported to be the noisiest; this anchors all expert ratings and full-reference IQA, risking biased benchmark.","rationale":"After reading the full text, I consider the most load-bearing concern to be the unverified and internally contradicted assumption that the default acquisition (90 kV, 8.0 ms, Normal) is the highest-quality reference. This premise is used both in the annotation protocol (expert raters grade test images relative to this reference) and in all full-reference IQA calculations. The paper's own observation that Normal reconstruction has more pronounced noise than Smooth/VSmooth directly undermines the premise. This is not merely a disagreement with consensus; it is an internal inconsistency that could bias the entire benchmark. The reader's weakest_assumption already identifies this, and I agree. A secondary issue is the claim that 'all IQA measures agreed' in Table 7, contradicted by VIF-python's rankings, but that affects only the exploratory IQA-based rating, which the authors themselves call exploratory. The reference assumption, by contrast, affects the core dataset annotations and the headline IQA benchmark. Therefore, the verdict should remain CONDITIONAL: the paper can be accepted once the reference choice is justified or the analyses are repeated with an alternative reference. No change from the reader's verdict.","tokens_in":14222,"tokens_out":7613,"duration_ms":68470,"concrete_test":"Run a blinded forced-choice experiment asking the same three expert readers to select the best-quality image from the three reconstruction types (Normal, Smooth, VSmooth) of the same slice/volume. If Normal is not chosen in a majority of the comparisons, the reference assumption fails. Additionally, recompute the main SRCC tables (Tables 3–4) using Smooth (or a denoised average) as the reference; if the ranking of IQA measures changes substantially, the benchmark results are reference-dependent.","verdict_should_be":"UNCHANGED","load_bearing_attack":"In Section 'Image Acquisition Protocols,' the default acquisition (kV=90, PW=8.0 ms, RT Normal) is declared to 'represent the device-standard, high-quality images' and is used as the reference. This reference is the anchor for both the expert relative-quality ratings (Speedy IQA) and every full-reference IQA measure. However, Section 'Performance of IQA measures in presence of background noise' states that 'noise in the reconstruction type Normal was more pronounced visually compared to the other two types (Smooth and Very Smooth).' This is a direct internal contradiction: the reference is noisier than some images labeled 'degraded.' Consequently, the expert labels and full-reference IQA scores are all computed relative to a potentially non-optimal baseline. The paper's later finding that correlations are driven by reconstruction-type variation could be an artifact of referencing a single noisy Normal volume rather than reflecting true quality differences. Without establishing that the reference is actually highest quality, the dataset's validity as a benchmark for IQA is not secured.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces CBCT-IQ, a publicly available cone-beam CT dataset for image quality assessment. The authors acquired CBCT volumes of thorax and pelvis phantoms on a Siemens ARTIS pheno system, systematically varying tube voltage (60/90/125 kV), pulse width (3.2/5.0/6.4/8.0 ms), and reconstruction type (Normal/Smooth/Very Smooth), resulting in 36 acquisition configurations. From these, they selected 1,764 2D slices (71 per chest configuration, 27 per pelvis configuration) and had three expert raters assign four-level quality scores for both overall image quality and a predefined ROI, with the default acquisition (90 kV, 8.0 ms, Normal) used as the reference. The paper further benchmarks 26 full-reference and no-reference IQA measures (27 implementations) against the expert ratings and proposes an exploratory IQA-measure-based ranking intended to capture subtle quality differences not perceived by human raters. The dataset and annotations are deposited on Zenodo.","tokens_in":14426,"tokens_out":5228,"duration_ms":50866,"significance":"If the annotations and benchmark are trustworthy, CBCT-IQ would fill a clear gap: a standardized, openly available CBCT dataset with expert quality labels, enabling development and comparison of IQA measures in a modality where such resources are scarce. The systematic variation of acquisition and reconstruction parameters, the use of three clinical experts, the provision of OIQ and ROI scores, and the inclusion of computed IQA values and software versions are concrete strengths that support reproducibility and benchmarking. The paper also provides a useful baseline comparison of 26 IQA measures on CBCT data. However, the validity of the expert labels and of the full-reference IQA benchmark depends critically on the choice of the reference image; the manuscript contains an internal contradiction about the noise level of that reference, which raises serious concerns about the benchmark's interpretation.","major_comments":[{"comment":"The paper declares the default acquisition (kV=90, PW=8.0 ms, RT Normal) as the 'device-standard, high-quality' reference (Section 'Image Acquisition Protocols') and uses it as the anchor for all expert relative-quality ratings and for every full-reference IQA measure. Yet Section 'Performance of IQA measures in presence of background noise' states that 'noise in the reconstruction type Normal was more pronounced visually compared to the other two types (Smooth and Very Smooth).' These statements directly contradict each other: Smooth and Very Smooth images, which are labelled as 'degraded', are less noisy than the reference. Since raters were asked to grade each test image relative to this reference, a noisier reference biases the quality labels: a smoother image may be scored lower not because it is diagnostically worse, but because it differs from the reference. Likewise, full-referen","section":"Image Acquisition Protocols / Performance of IQA measures in presence of background noise"},{"comment":"The text claims that 'all IQA measures agreed and identified PW6.4-KV90 combination as the best quality' and that 'Results on the ROI also confirmed this agreement.' Table 7 itself contradicts this. In the VIF-python row, the values are: PW8-KV60=1.258, PW8-KV125=0.696, PW5-KV90=0.934, PW6.4-KV90=0.937, PW3.2-KV90=0.921. For VIF, higher values indicate better quality, so VIF-python selects PW8-KV60 as the best, not PW6.4-KV90. The same issue appears in Table 8 (ROI) where VIF-python gives PW8-KV60=1.475 versus PW6.4-KV90=0.918. Thus the 'strong agreement' claim is not supported by the reported data. The IQA-based rating must be derived from per-measure rankings with a clearly defined aggregation rule, and the text should be corrected to avoid overclaiming consensus.","section":"IQA measure–based rating, Table 7"},{"comment":"The benchmark results report Spearman rank correlation coefficients as point estimates without confidence intervals or significance tests. Many of the SRCC values in these tables are close to each other (e.g., the top performers in Table 3 differ by a few hundredths), and with 1,764 samples but heavily dependent ratings, these differences may not be meaningful. The paper makes practical recommendations about which IQA measures are 'best-performing' (e.g., in the noise-exclusion analysis), so the absence of uncertainty quantification is a substantive omission. The authors should provide bootstrap confidence intervals for the SRCC values and, where appropriate, tests for differences between dependent correlations, so that readers can judge whether the reported ordering of IQA measures is statistically reliable. Without this, the benchmark's conclusions are under-supported.","section":"Tables 3–6: Benchmark statistics"}],"minor_comments":[{"comment":"The text says '36 volumes consisting of 378 slices were generated,' but the dataset contains 1,764 slices (18 chest configurations × 71 slices + 18 pelvis configurations × 27 slices). Please clarify what the 378 refers to or correct the sentence; as written it is internally inconsistent.","section":"Image Acquisition Protocols"},{"comment":"The abstract says '26 full reference- and no reference-based IQA measures,' while the Methods section mentions '26 commonly used as well as task-based IQA measures (27 IQA measure implementations).' Please standardize the count throughout (e.g., '26 measures, 27 implementations') to avoid confusion.","section":"Abstract and Methods"},{"comment":"In the Data Records section, the reference path is described as containing 'kernel EE' (e.g., './DATASETNAME/SLICE/EE/090/8.0/NORMAL.nii.gz'). The term 'kernel' has not been introduced; the paper discusses reconstruction types (Normal, Smooth, VSmooth). Please clarify what 'EE' denotes and whether it is a reconstruction kernel name.","section":"Data Records"},{"comment":"Several typos appear: 'Nomal' in Table 1 and Figure 2/3 captions; 'Norml' in Figure 2; 'boy anatomy' should likely be 'bony anatomy' in the Phantoms section; 'Slides' appears instead of 'slices' in the abstract/Data Records. Please proofread the manuscript.","section":"Figures and Table 1"},{"comment":"The annotation software section says the main task category is 'Image-guided therapy' and subcategory 'Region of Interest (ROI),' but the raters assessed both OIQ and ROI. Clarify how the OIQ task was configured in Speedy IQA, since the described configuration appears to name only the ROI subcategory.","section":"Image annotation and reader study"}],"recommendation":"major_revision","confidential_remarks":"The referee report focuses on the reference-image contradiction, which is the most serious issue. If the authors can convincingly justify the Normal reconstruction as the clinically correct reference despite its higher visual noise, or re-anchor the annotations and IQA calculations, the dataset remains a valuable contribution. The VIF-python disagreement with the 'all agreed' claim is a clear factual inconsistency that must be corrected. The lack of confidence intervals is fixable but important for a benchmarking paper. I see no grounds for rejection if these points are addressed, but the revision needs to be substantial."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThe dataset is the real deliverable here. CBCT-IQ gives the community the first public CBCT image set with expert quality annotations — 1,764 slices from two phantoms, three raters, four-level scores for overall and ROI quality. If the data check out on Zenodo, this fills a clear gap. The annotation protocol is sensible, inter-rater OIQ agreement is high (SRCC above 0.8), and the systematic kV/pulse-width/reconstruction-type variation is well documented. This is worth having.\n\nThe IQA benchmark is secondary and has real soft spots. The biggest one is the reference image. The default protocol (kV90, PW8, RT Normal) is declared the 'device-standard, high-quality' reference and anchors every full-reference measure and every expert relative rating. But later the authors say noise in the Normal reconstruction was more pronounced than in Smooth or Very Smooth. If the reference is noisier than some 'degraded' images, the expert ratings and all full-reference correlations are relative to a baseline of questionable optimality. That does not destroy the dataset, but it undermines the benchmark conclusions as stated, and the authors do not address the tension.\n\nAlso, the claim that 'all IQA measures agreed' on the subtle ranking is contradicted by the VIF-python row in Table 7, which would put PW8-KV60 at the top. The SRCC values have no confidence intervals or significance tests, and there are internal inconsistencies (e.g., '36 volumes consisting of 378 slices' vs. 1,764 slices total). The authors provide software versions but not the analysis code, which limits reproducibility of the benchmark.\n\nTo their credit, they call the IQA-based rating exploratory and are clear that it does not necessarily correspond to diagnostic quality. And the main result — the correlation drop when reconstruction type is held fixed — is an honest and useful finding about what these measures are capturing.\n\nFor peer review: yes, send it. The dataset deserves referee time and is publishable after revision. But the revision needs to address the reference choice explicitly (e.g., justify it or test sensitivity to another reference), fix the overclaims, and add uncertainty quantification for the benchmark.\n\nRecommendation: engage with the dataset, but treat the benchmark table as provisional.","headline":"A genuinely useful first CBCT IQA dataset, but the benchmark conclusions are anchored to a reference image that the paper itself suggests is noisy.","tokens_in":15016,"tokens_out":3512,"would_cite":true,"duration_ms":34229,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper publishes the first open-access cone-beam CT dataset in which 1,764 image slices carry expert quality ratings, providing a shared benchmark for testing image-quality measures against human perception.","keywords":["cone-beam CT","image quality assessment","expert annotations","benchmark dataset","full-reference IQA","no-reference IQA","CBCT","phantom study"],"falsifier":"A concrete check: take a subset of slices and have the same three experts re-rate the degraded images paired against a different reference (for instance, 125 kV or the Smooth reconstruction). If the relative quality rankings of the other variants change substantially, or if full-reference IQA measures no longer correlate with expert scores, then the default-protocol anchor is not a stable ground truth and the dataset's benchmark conclusions weaken.","tokens_in":14093,"feed_emoji":"🩻","tokens_out":5562,"duration_ms":48053,"temperature":0.7,"pith_summary":"The paper aims to give the CBCT research community a much-needed public benchmark: a systematically varied set of CBCT images with expert clinical annotations of image quality. The authors scanned thorax and pelvis phantoms on a C-arm CBCT system, varying tube voltage, pulse width, and reconstruction type, and had three experts rate overall and region-of-interest quality on a four-level scale relative to a device-standard reference image. They computed 26 full-reference and no-reference IQA measures and compared them to the expert scores, finding that the strongest correlations are driven by reconstruction-type differences rather than by tube voltage or pulse width changes. They also introduce an exploratory IQA-based ranking that agrees across measures on subtle quality differences that human raters could not distinguish. If accepted, the dataset gives researchers a shared testbed for developing and validating quantitative image-quality measures, with a baseline for comparison.","feed_headline":"1,764 annotated CBCT slices become a public image-quality benchmark","feed_subtitle":"Three experts graded each slice; 26 IQA measures are benchmarked so new methods can be compared on equal footing.","key_machinery":"The central object is the paired reference/degraded image set: every degraded slice has a matching reference slice from the same phantom and anatomical position, acquired at the device-standard settings (90 kV, 8.0 ms, Normal reconstruction). This pairing enables both the expert relative-quality annotations (each test image is graded while displayed next to its reference) and all full-reference IQA computations. The paper's quantitative machinery is Spearman rank correlation between the three experts' z-scored ratings and each of 27 IQA measure implementations, applied to the whole image, a predefined ROI, and a background-noise patch; the ROI and noise-patch analyses are what reveal which m","core_discovery":"The central discovery is the dataset: 1,764 annotated CBCT slices from 36 volumes with systematic variations in kV, pulse width, and reconstruction type, each graded by three experts for overall and ROI quality on a 1–4 scale, anchored to a device-standard reference acquisition. Benchmarking 27 IQA implementations shows LPIPS, DISTS, FSIM, VSI, VIF, DSS, and Haar-based measures correlate most strongly with expert scores, but several respond mainly to background noise rather than structure. The high overall correlations are driven almost entirely by reconstruction-type variation; within a fixed reconstruction type, expert ratings barely change and IQA correlations drop sharply. To fill this g","pith_inferences":["Extension: if the reference acquisition (90 kV, 8.0 ms, Normal) is not actually the clinically best image for a given task, then both the expert relative ratings and all full-reference IQA correlations shift; re-running the expert study with a different reference (e.g., 125 kV or Smooth reconstruction) would test this.","Extension: the dataset's single scanner and phantom geometry means its IQA rankings may not transfer to other CBCT systems with different detectors or reconstruction algorithms; cross-scanner validation is needed before treating these results as a universal CBCT IQA benchmark.","Extension: the consistent IQA consensus on subtle variants suggests a route toward automated acquisition-protocol optimization, where a set of measures could serve as a proxy objective in closed-loop dose optimization.","Extension: the noise-patch analysis shows some measures respond primarily to background noise, so the dataset could be used to develop structure-aware IQA measures that explicitly penalize noise-only responses—a testable direction the authors hint at but do not pursue."],"forward_implications":["Any new CBCT IQA method can be tested on the dataset and its Spearman correlation against expert OIQ and ROI scores compared directly with the 27 measures benchmarked here.","The finding that LPIPS, DISTS, and RPIAXIS track background noise rather than structure warns researchers against using those measures alone for CBCT protocol evaluation.","Since expert ratings detect reconstruction-type changes but not kV/pulse-width changes, the dataset separates 'visible' quality differences from 'measureable but invisible' ones, providing two complementary test regimes.","The IQA-based consensus rating offers a way to rank acquisition settings for subtle quality differences when human perception is insensitive, potentially informing dose-optimization studies.","Because reference and degraded slices come from the same phantom and slice position, the dataset can also be used for training and validating deep learning models on image-quality prediction."],"fun_headline_variants":["First open CBCT IQA dataset: 1,764 expert-graded slices","1,764 expert-scored CBCT slices launch open IQA benchmark","Public CBCT dataset with expert quality grades to benchmark IQA","CBCT-IQ benchmark: 1,764 slices, 3 experts, 26 IQA metrics","New public benchmark: 1,764 CBCT slices with expert annotations"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The load-bearing premise is that the single reference acquisition (90 kV, 8.0 ms, Normal reconstruction) is the highest-quality image for every phantom and slice, so that all expert relative-quality ratings and all full-reference IQA measures inherit that one anchor.","fun_headline_variants_meta":{"raw":{"variants":["First open CBCT IQA dataset: 1,764 expert-graded slices","1,764 expert-scored CBCT slices launch open IQA benchmark","Public CBCT dataset with expert quality grades to benchmark IQA","CBCT-IQ benchmark: 1,764 slices, 3 experts, 26 IQA metrics","New public benchmark: 1,764 CBCT slices with expert annotations"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001404,"raw_usage":{"total_tokens":5518,"prompt_tokens":755,"completion_tokens":4763,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":499,"completion_tokens_details":{"reasoning_tokens":4660}},"tokens_in":499,"tokens_out":4763,"duration_ms":29383,"temperature":1.0,"reasoning_tokens":4660,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-03T10:43:02.916158+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A concrete check: take a subset of slices and have the same three experts re-rate the degraded images paired against a different reference (for instance, 125 kV or the Smooth reconstruction). If the relative quality rankings of the other variants change substantially, or if full-reference IQA measures no longer correlate with expert scores, then the default-protocol anchor is not a stable ground truth and the dataset's benchmark conclusions weaken.","supporting_citations":[],"review_version":1}