{"id":"a3c4a583-0f12-46cc-b754-77001c992dd4","arxiv_id":"2604.16034","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":5.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":0,"one_line_summary":"Integrated Gradients and DeepLIFT ranked highest among 13 XAI methods for faithfulness, complexity, and plausibility in head and neck cancer outcome prediction on the HECKTOR multi-center dataset.","lead":"The paper evaluated and ranked 13 explainable AI methods for predicting outcomes in head and neck cancer patients using PET/CT scans, testing them on 24 metrics for faithfulness, robustness, complexity, and plausibility. This provides guidance on selecting reliable XAI techniques to support clinical trust and personalized treatment decisions.","discovery_kind":"new_application","skeptic_critique":{"model":"grok-4.3","headline":"Proxy metrics for plausibility and faithfulness lack direct clinical anchoring for outcome prediction","rationale":"The reader's weakest assumption directly identifies the metrics-to-clinical-utility gap. The abstract provides no evidence that the 24 metrics were validated against expert judgment or downstream treatment impact, making this the least secure link in the central claim. No other internal inconsistency (e.g., dataset handling or metric definitions) is visible from the supplied text.","tokens_in":1720,"tokens_out":333,"duration_ms":29128,"concrete_test":"On the HECKTOR test set, obtain independent radiologist rankings (1-5 Likert) of explanation usefulness for 30 cases; recompute the aggregate ranking using only the subset of metrics that correlate >0.6 with the clinician scores. If IG or DL drop out of the top three, the original 24-metric ranking does not generalize to clinical relevance.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The headline result (IG and DL consistently top-ranked across 24 metrics) depends on the assumption that the chosen faithfulness, robustness, complexity, and plausibility proxies are sufficient to rank methods for real HNC treatment decisions. Faithfulness metrics (e.g., insertion/deletion) test pixel-level sensitivity but do not verify whether highlighted regions correspond to biologically relevant PET/CT features such as hypoxic subvolumes or nodal involvement that actually drive prognosis. Plausibility is typically measured via overlap with segmentation masks rather than prospective clinician ratings of explanatory utility. If these proxies diverge from clinical utility, the ranking does not support the claim that IG/DL are preferable for deployment.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.3","summary":"The manuscript claims to be the first comprehensive benchmarking of 13 XAI methods using 24 metrics spanning faithfulness, robustness, complexity, and plausibility for head and neck cancer outcome prediction from PET/CT images on the multi-center HECKTOR dataset. It reports substantial performance variations across methods and that Integrated Gradients and DeepLIFT consistently rank highest in faithfulness, complexity, and plausibility.","tokens_in":1846,"tokens_out":434,"duration_ms":44177,"significance":"If the rankings prove robust, the work supplies empirical guidance for XAI selection in medical imaging and demonstrates the value of multi-aspect evaluation over ad-hoc choices. This could support more interpretable AI models for HNC prognosis, though its impact depends on whether the proxy metrics align with clinical decision-making needs.","major_comments":[{"comment":"The central ranking result depends on the 24 metrics serving as valid proxies for clinical utility in outcome prediction. Faithfulness metrics such as insertion/deletion evaluate pixel-level sensitivity but do not test whether highlighted regions correspond to biologically relevant features (e.g., hypoxic subvolumes or nodal involvement) that drive HNC prognosis. Plausibility is assessed via overlap with segmentation masks rather than clinician ratings of explanatory value for treatment decisions. This disconnect is load-bearing for interpreting the IG/DL rankings as preferable for real-world deployment.","section":null},{"comment":"The abstract and results sections state rankings and 'large variations' without specifying model architectures for the base predictor, exact implementations of the 24 metrics, statistical significance tests, error bars, or preprocessing details. These omissions prevent verification that post-hoc choices did not influence the reported superiority of IG and DL.","section":null}],"minor_comments":[{"comment":"Abstract: 'consistently obtained high rankings' contains a tense inconsistency; rephrase to 'consistently obtain high rankings' or similar for grammatical accuracy.","section":null},{"comment":"The manuscript would benefit from an expanded limitations paragraph explicitly addressing the gap between proxy metrics and prospective clinical validation.","section":null}],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":{"model":"grok-4.3","summary":"We thank the referee for the detailed and constructive review. We address each major comment point-by-point below. Where the comments identify gaps in detail or discussion, we have revised the manuscript accordingly.","responses":[{"response":"We agree that the 24 metrics are established quantitative proxies rather than direct measures of biological relevance or clinical decision utility. Our benchmarking follows standard XAI evaluation protocols from the literature to enable objective, reproducible comparisons across methods. In the revised manuscript we have added a dedicated Limitations subsection in the Discussion that explicitly acknowledges this gap, notes that segmentation-overlap plausibility is a common but imperfect proxy, and states that future work should incorporate clinician ratings and biological validation (e.g., hypoxic subvolume correlation). The core empirical rankings remain unchanged because they are correctly reported as metric-specific results.","revision_made":"yes","referee_comment":"The central ranking result depends on the 24 metrics serving as valid proxies for clinical utility in outcome prediction. Faithfulness metrics such as insertion/deletion evaluate pixel-level sensitivity but do not test whether highlighted regions correspond to biologically relevant features (e.g., hypoxic subvolumes or nodal involvement) that drive HNC prognosis. Plausibility is assessed via overlap with segmentation masks rather than clinician ratings of explanatory value for treatment decisions. This disconnect is load-bearing for interpreting the IG/DL rankings as preferable for real-world deployment."},{"response":"We acknowledge that the original submission omitted several implementation details required for full reproducibility. In the revised manuscript we have: (1) expanded the Methods section with the precise base predictor architecture (3D ResNet-50 with specific hyperparameters), (2) provided references and pseudocode for each of the 24 metrics, (3) added statistical significance testing (paired Wilcoxon tests with p-values and effect sizes) between top-ranked methods, (4) included error bars on all ranking plots, and (5) detailed the full preprocessing pipeline (resampling, normalization, augmentation). These additions allow independent verification and address the concern that post-hoc choices may have influenced the IG/DL rankings.","revision_made":"yes","referee_comment":"The abstract and results sections state rankings and 'large variations' without specifying model architectures for the base predictor, exact implementations of the 24 metrics, statistical significance tests, error bars, or preprocessing details. These omissions prevent verification that post-hoc choices did not influence the reported superiority of IG and DL."}],"tokens_in":1307,"tokens_out":519,"duration_ms":27804,"standing_objections":[]},"desk_editor":{"model":"grok-4.3","letter":"The paper's core contribution is an empirical ranking of explanation techniques for a prognostic model on the public HECKTOR multi-center dataset. They test 13 methods against 24 metrics spanning faithfulness, robustness, complexity, and plausibility, and report that Integrated Gradients and DeepLIFT come out ahead on most of the key dimensions. This is new for the HNC setting, where earlier work tended to pick XAI tools without systematic comparison, and the scale of the evaluation plus the public data make the results usable as a reference point for similar tasks.","headline":"This is a straightforward benchmarking study that ranks 13 XAI methods on HNC outcome prediction from PET/CT and finds IG and DL strongest overall, but the proxy metrics leave a gap to actual clinical utility.","tokens_in":2396,"tokens_out":198,"would_cite":false,"duration_ms":21535,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.3","headline":"A systematic ranking of 13 XAI methods across 24 metrics identifies Integrated Gradients and DeepLIFT as top performers for explaining head and neck cancer outcome predictions from PET/CT images.","keywords":["explainable AI","XAI evaluation","head and neck cancer","outcome prediction","PET/CT","Integrated Gradients","DeepLIFT","HECTOR dataset"],"falsifier":"A replication on an independent multi-center dataset where Integrated Gradients and DeepLIFT no longer rank at the top for faithfulness and plausibility, or where the relative ordering of all 13 methods changes substantially.","tokens_in":2615,"feed_emoji":"📊","tokens_out":509,"duration_ms":59138,"temperature":0.7,"pith_summary":"The authors seek to replace ad-hoc selection of explainable AI techniques in medical imaging with a full comparison. They evaluate thirteen XAI methods on deep learning models that predict head and neck cancer outcomes from PET and CT scans in the multi-center HECKTOR dataset. Each method is scored on twenty-four metrics that measure how faithfully the explanation reflects the model's logic, how stable it remains under small changes, how simple the output is, and how plausible it appears. The results show wide differences in scores across methods, with Integrated Gradients and DeepLIFT placing consistently high on faithfulness, complexity, and plausibility. This matters for clinical use because trustworthy explanations could help doctors understand AI recommendations when choosing personalized treatments.","feed_headline":"XAI ranking places Integrated Gradients and DeepLIFT at top for cancer explanations","feed_subtitle":"Evaluation of 13 methods on 24 metrics across multi-center PET/CT data shows wide differences and consistent leaders in key aspects.","key_machinery":"The ranking framework of 24 metrics grouped into faithfulness, robustness, complexity, and plausibility, applied to explanations from 13 XAI methods on PET/CT-based models for HNC prognosis.","core_discovery":"The paper establishes that a comprehensive evaluation of thirteen XAI methods using twenty-four metrics on the HECKTOR multi-center dataset reveals large performance variations, with Integrated Gradients and DeepLIFT achieving high rankings for faithfulness, complexity, and plausibility when interpreting AI models for head and neck cancer outcome prediction.","pith_inferences":[],"forward_implications":[],"fun_headline_variants":["XAI evaluation ranks Integrated Gradients and DeepLIFT highest for HNC outcomes","13 XAI techniques evaluated across 24 metrics on HECKTOR data IG DeepLIFT lead","Comprehensive XAI ranking reveals Integrated Gradients DeepLIFT as top for cancer AI","IG and DeepLIFT top XAI rankings in evaluation for head and neck cancer prediction"],"cache_read_input_tokens":2112,"weakest_assumption_plain":"That the 24 chosen metrics together capture the qualities that make an explanation useful and trustworthy for real clinical decisions in head and neck cancer.","fun_headline_variants_meta":{"raw":{"variants":["XAI evaluation ranks Integrated Gradients and DeepLIFT highest for HNC outcomes","13 XAI techniques evaluated across 24 metrics on HECKTOR data IG DeepLIFT lead","Comprehensive XAI ranking reveals Integrated Gradients DeepLIFT as top for cancer AI","IG and DeepLIFT top XAI rankings in evaluation for head and neck cancer prediction"]},"model":"grok-4.3","cost_usd":0.005736,"raw_usage":{"total_tokens":2699,"prompt_tokens":594,"num_sources_used":0,"completion_tokens":90,"cost_in_usd_ticks":57362000,"prompt_tokens_details":{"text_tokens":594,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":2015,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":594,"tokens_out":90,"duration_ms":19791,"temperature":1.0,"reasoning_tokens":2015,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-05-10T09:07:58.432368+00:00","model_set":{"reader":"grok-4.3"},"falsifier":"A replication on an independent multi-center dataset where Integrated Gradients and DeepLIFT no longer rank at the top for faithfulness and plausibility, or where the relative ordering of all 13 methods changes substantially.","supporting_citations":[],"review_version":1}