REVIEW 1 major objections 50 references
Appropriate reliance on set-valued AI advice is measured by new rates for classification and quantity-plus-quality metrics for regression.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · grok-4.3
2026-06-28 01:49 UTC pith:ZH47X45Z
load-bearing objection This paper gives the first explicit metrics for appropriate reliance on set-valued AI advice but does not show that the four quantities are sufficient without missing dimensions. the 1 major comments →
A Framework for Measuring Appropriate Reliance on Set-Valued AI Advice
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
The central claim is that appropriate reliance on set-valued AI advice is jointly characterized by the correct reliance rate on AI and the correct reliance rate on self in classification tasks, together with the quantity of AI reliance and the quality of AI reliance in regression tasks, and that these metrics capture important nuances in human-AI collaboration overlooked by existing point-prediction measures.
What carries the argument
The four metrics (correct reliance rate on AI, correct reliance rate on self, quantity of AI reliance, quality of AI reliance) that evaluate whether and how set-valued advice is used correctly in the judge-advisor sequence.
Load-bearing premise
The newly defined metrics are sufficient by themselves to characterize appropriate reliance and to capture the nuances that prior measures miss.
What would settle it
An experiment in which participants receive set-valued AI advice, make decisions, and the four metrics fail to separate cases where reliance behavior is intuitively appropriate from cases where it is not.
If this is right
- Human-AI teams using interval advice in classification can now be scored on whether reliance choices are correct.
- Regression decisions can be broken down into whether the AI set was consulted at all and whether consultation improved accuracy over the initial estimate.
- Different formats of set-valued advice can be compared directly for how well they support appropriate reliance.
- Studies of human-AI collaboration can move beyond point-prediction baselines to account for uncertainty communication.
Where Pith is reading between the lines
- Designers of AI systems might tune set outputs specifically to raise the correct reliance rates or the quality metric.
- Training for decision makers could use these metrics as feedback to improve how people weigh uncertain AI advice.
- The same measurement approach could be tested on multi-step decisions that go beyond the single judge-advisor round described.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper develops the first formal framework for measuring appropriate reliance on set-valued AI advice (discrete sets or continuous intervals) within the sequential judge-advisor paradigm. It spans classification tasks by defining two metrics—correct reliance rate on AI and correct reliance rate on self—that jointly characterize appropriate reliance, and regression tasks by defining quantity of AI reliance and quality of AI reliance. The framework is presented as capturing important nuances in human-AI collaboration overlooked by existing point-prediction measures.
Significance. If the metrics are shown to be jointly sufficient for characterizing appropriate reliance without unmodeled dimensions, the framework would fill a clear gap in human-AI collaboration research by extending beyond point predictions to uncertainty-aware advice. The purely definitional approach with no free parameters or fitted quantities is a methodological strength, as is the explicit coverage of both classification and regression.
major comments (1)
- [Abstract] Abstract: The central claim that the four defined metrics 'jointly characterize appropriate reliance' and 'capture important nuances... that existing measures overlook' is load-bearing but unsupported. The framework provides no argument or test establishing that these quantities (correct reliance rates for classification; quantity/quality for regression) exhaust the space of relevant human decision-making dimensions, such as confidence calibration, set-size sensitivity, or asymmetric error costs. Without such justification, the sufficiency assumption remains unverified.
Simulated Author's Rebuttal
We thank the referee for their constructive review and for identifying an important point about the strength of the claims in the abstract. We respond to the major comment below.
read point-by-point responses
-
Referee: [Abstract] Abstract: The central claim that the four defined metrics 'jointly characterize appropriate reliance' and 'capture important nuances... that existing measures overlook' is load-bearing but unsupported. The framework provides no argument or test establishing that these quantities (correct reliance rates for classification; quantity/quality for regression) exhaust the space of relevant human decision-making dimensions, such as confidence calibration, set-size sensitivity, or asymmetric error costs. Without such justification, the sufficiency assumption remains unverified.
Authors: We agree that the abstract's phrasing is too strong and that the manuscript provides no formal argument or empirical test establishing that the proposed metrics are exhaustive of all relevant dimensions of human decision making. The framework is a definitional contribution that derives the metrics directly from the structure of set-valued advice in the judge-advisor paradigm: for classification, the two rates capture whether a decision maker correctly follows (or does not follow) the AI when the ground truth is inside or outside the set; for regression, quantity and quality capture the magnitude and benefit of adjustment relative to the initial estimate. These dimensions are not addressed by existing point-prediction measures, which is the sense in which the framework captures nuances they overlook. However, we do not claim the metrics are jointly sufficient for every possible aspect of reliance behavior. We will revise the abstract to replace 'jointly characterize appropriate reliance' with 'provide metrics that characterize appropriate reliance along the dimensions of correctness and benefit of reliance decisions' and will add a dedicated limitations paragraph in the discussion section that explicitly notes additional factors (e.g., calibration, set-size sensitivity, asymmetric costs) that may require separate modeling in future work. revision: yes
Circularity Check
No circularity: purely definitional framework with no reductions to fitted inputs or self-citations
full rationale
The paper introduces a new framework by defining four metrics (correct reliance rate on AI/self for classification; quantity/quality of AI reliance for regression) that are stated to jointly characterize appropriate reliance. These are presented as definitional choices within the sequential judge-advisor paradigm, with no equations, fitted parameters, or predictions that reduce by construction to the inputs. No self-citations are invoked as load-bearing for uniqueness or ansatzes, and the central claim rests on the explicit definitions rather than any derivation that collapses into prior results. The framework is therefore self-contained against external benchmarks.
Axiom & Free-Parameter Ledger
axioms (1)
- domain assumption The sequential judge-advisor paradigm is the appropriate setting for evaluating reliance on set-valued AI advice.
read the original abstract
Appropriate reliance on AI advice has become a central research theme in human-AI collaboration. Existing frameworks have focused exclusively on point predictions as AI advice. However, set-valued AI advice (e.g., discrete sets or continuous intervals) is increasingly being used to communicate uncertainty and improve human decision making. In this paper, we develop the first formal framework for measuring appropriate reliance on set-valued AI advice within the sequential judge-advisor paradigm, spanning both classification and regression tasks. For classification, we first introduce the dimensions that are necessary for evaluating set-valued AI advice. We then define two metrics: correct reliance rate on AI and correct reliance rate on self, which jointly characterize appropriate reliance in this setting. For regression, we introduce quantity of AI reliance and quality of AI reliance, which respectively measure whether a decision maker utilized the AI advice and whether their reliance helped them get closer to the ground truth relative to their initial estimate. Through the application of our framework, we demonstrate how these metrics capture important nuances in human-AI collaboration that existing measures overlook.
Figures
Reference graph
Works this paper leans on
-
[1]
Proceedings of the 41st International Conference on Machine Learning , pages=
Conformal prediction sets improve human decision making , author=. Proceedings of the 41st International Conference on Machine Learning , pages=
-
[2]
Evaluating the utility of conformal prediction sets for
Zhang, Dongping and Chatzimparmpas, Angelos and Kamali, Negar and Hullman, Jessica , booktitle=. Evaluating the utility of conformal prediction sets for
-
[3]
Conformal prediction for
Folgado, Duarte and Famiglini, Lorenzo and Campagner, Andrea and Dores, H. Conformal prediction for. International Conference on Artificial Intelligence in Medicine , pages=. 2025 , organization=
2025
-
[4]
Towards human-
De Toni, Giovanni and Okati, Nastaran and Thejaswi, Suhas and Straitouri, Eleni and Gomez-Rodriguez, Manuel , journal=. Towards human-
-
[5]
International Conference on Machine Learning , pages=
Improving expert predictions with conformal prediction , author=. International Conference on Machine Learning , pages=. 2023 , organization=
2023
-
[6]
arXiv preprint arXiv:2410.01888 , year=
Conformal prediction sets can cause disparate impact , author=. arXiv preprint arXiv:2410.01888 , year=
-
[7]
Conformal prediction and human decision making, 2025
Conformal prediction and human decision making , author=. arXiv preprint arXiv:2503.11709 , year=
-
[8]
Proceedings of the 2021 AAAI/ACM Conference on AI, Ethics, and Society , pages=
Uncertainty as a form of transparency: Measuring, communicating, and using uncertainty , author=. Proceedings of the 2021 AAAI/ACM Conference on AI, Ethics, and Society , pages=
2021
-
[9]
Does the whole exceed its parts?
Bansal, Gagan and Wu, Tongshuang and Zhou, Joyce and Fok, Raymond and Nushi, Besmira and Kamar, Ece and Ribeiro, Marco Tulio and Weld, Daniel , booktitle=. Does the whole exceed its parts?
-
[10]
To trust or to think: Cognitive forcing functions can reduce overreliance on
Bu. To trust or to think: Cognitive forcing functions can reduce overreliance on. Proceedings of the ACM on Human-Computer Interaction , volume=. 2021 , publisher=
2021
-
[11]
To engage or not to engage with
Lebovitz, Sarah and Lifshitz-Assaf, Hila and Levina, Natalia , journal=. To engage or not to engage with. 2022 , publisher=
2022
-
[12]
Journal of Artificial Intelligence Research , volume=
Schoeffer, Jakob and Jakubik, Johannes and V. Journal of Artificial Intelligence Research , volume=
-
[13]
Explanations, fairness, and appropriate reliance in human-
Schoeffer, Jakob and De-Arteaga, Maria and Kuehl, Niklas , booktitle=. Explanations, fairness, and appropriate reliance in human-
-
[14]
Too Sure for Our Own Good: A User Study on
Fregosi, Caterina and Vicente, Lucia and Campagner, Andrea and Cabitza, Federico , booktitle=. Too Sure for Our Own Good: A User Study on
-
[15]
``Are you really sure?''
Ma, Shuai and Wang, Xinru and Lei, Ying and Shi, Chuhan and Yin, Ming and Ma, Xiaojuan , booktitle=. ``Are you really sure?''
-
[16]
Understanding uncertainty: How lay decision-makers perceive and interpret uncertainty in human-
Prabhudesai, Snehal and Yang, Leyao and Asthana, Sumit and Huan, Xun and Liao, Q Vera and Banovic, Nikola , booktitle=. Understanding uncertainty: How lay decision-makers perceive and interpret uncertainty in human-
-
[17]
Journal of Forecasting , volume=
Visualizing uncertainty in time series forecasts: The impact of uncertainty visualization on users' confidence, algorithmic advice utilization, and forecasting performance , author=. Journal of Forecasting , volume=. 2025 , publisher=
2025
-
[18]
Balancing the unknown: Exploring human reliance on
Holstein, Joshua and B. Balancing the unknown: Exploring human reliance on. ACM Transactions on Computer-Human Interaction , volume=. 2025 , publisher=
2025
-
[19]
Designing for appropriate reliance: The roles of
Cao, Shiye and Liu, Anqi and Huang, Chien-Ming , journal=. Designing for appropriate reliance: The roles of. 2024 , publisher=
2024
-
[20]
A decision theoretic framework for measuring
Guo, Ziyang and Wu, Yifan and Hartline, Jason D and Hullman, Jessica , booktitle=. A decision theoretic framework for measuring
-
[21]
A survey of
Eckhardt, Sven and K. A survey of. ACM Computing Surveys , volume=. 2025 , publisher=
2025
-
[22]
Proceedings of the AAAI conference on artificial intelligence , volume=
Fair conformal predictors for applications in medical imaging , author=. Proceedings of the AAAI conference on artificial intelligence , volume=
-
[23]
Appropriate reliance on
Schemmer, Max and Kuehl, Niklas and Benz, Carina and Bartos, Andrea and Satzger, Gerhard , booktitle=. Appropriate reliance on
-
[24]
How do we measure over-reliance?
Mikhaylova, Daria and Turchi, Tommaso and Malizia, Alessio , booktitle=. How do we measure over-reliance?
-
[25]
International
Bengio, Yoshua and Clare, Stephen and Prunkl, Carina and Andriushchenko, Maksym and Bucknall, Ben and Murray, Malcolm and Bommasani, Rishi and Casper, Stephen and Davidson, Tom and Douglas, Raymond and others , journal=. International
-
[26]
Do people appropriately rely on
Raees, Muhammad and Khan, Vassilis-Javed and Lykourentzou, Ioanna and Papangelis, Konstantinos , booktitle=. Do people appropriately rely on
-
[27]
Cabitza, Federico and Campagner, Andrea and Angius, Riccardo and Natali, Chiara and Reverberi, Carlo , booktitle=
-
[28]
Human Factors , volume=
Trust in automation: Designing for appropriate reliance , author=. Human Factors , volume=. 2004 , publisher=
2004
-
[29]
Organizational Behavior and Human Decision Processes , volume=
Advice taking and decision-making: An integrative literature review, and implications for the organizational sciences , author=. Organizational Behavior and Human Decision Processes , volume=. 2006 , publisher=
2006
-
[30]
Foundations and Trends in Machine Learning , volume=
Conformal prediction: A gentle introduction , author=. Foundations and Trends in Machine Learning , volume=. 2023 , publisher=
2023
-
[31]
Automation and human performance , pages=
Human decision makers and automated decision aids: Made for each other? , author=. Automation and human performance , pages=. 2018 , publisher=
2018
-
[32]
Journal of Experimental Psychology: General , volume=
Algorithm aversion: People erroneously avoid algorithms after seeing them err , author=. Journal of Experimental Psychology: General , volume=. 2015 , publisher=
2015
-
[33]
Current Psychology , volume=
A meta-analysis of the weight of advice in decision-making , author=. Current Psychology , volume=. 2023 , publisher=
2023
-
[34]
Proceedings of the Conference on Fairness, Accountability, and Transparency , pages=
Bias in bios: A case study of semantic representation bias in a high-stakes setting , author=. Proceedings of the Conference on Fairness, Accountability, and Transparency , pages=
-
[35]
A meta-analysis of the utility of explainable artificial intelligence in human-
Schemmer, Max and Hemmer, Patrick and Nitsche, Maximilian and K. A meta-analysis of the utility of explainable artificial intelligence in human-. Proceedings of the 2022 AAAI/ACM Conference on AI, Ethics, and Society , pages=
2022
-
[36]
Journal of Machine Learning Research , volume=
A tutorial on conformal prediction , author=. Journal of Machine Learning Research , volume=
-
[37]
Advances in Neural Information Processing Systems , volume=
Conformalized quantile regression , author=. Advances in Neural Information Processing Systems , volume=
-
[38]
Automation bias in the
Laux, Johann and Ruschemeier, Hannah , journal=. Automation bias in the. 2025 , publisher=
2025
-
[39]
Proceedings of the 2024 ACM Conference on Fairness, Accountability, and Transparency , pages=
On the quest for effectiveness in human oversight: Interdisciplinary perspectives , author=. Proceedings of the 2024 ACM Conference on Fairness, Accountability, and Transparency , pages=
2024
-
[40]
In search of verifiability: Explanations rarely enable complementary performance in
Fok, Raymond and Weld, Daniel S , journal=. In search of verifiability: Explanations rarely enable complementary performance in. 2024 , publisher=
2024
-
[41]
To err is
He, Gaole and Bharos, Abri and Gadiraju, Ujwal , booktitle=. To err is
-
[42]
Knowing about knowing: An illusion of human competence can hinder appropriate reliance on
He, Gaole and Kuiper, Lucie and Gadiraju, Ujwal , booktitle=. Knowing about knowing: An illusion of human competence can hinder appropriate reliance on
-
[43]
Dealing with uncertainty: Understanding the impact of prognostic versus diagnostic tasks on trust and reliance in human-
Salimzadeh, Sara and He, Gaole and Gadiraju, Ujwal , booktitle=. Dealing with uncertainty: Understanding the impact of prognostic versus diagnostic tasks on trust and reliance in human-
-
[44]
2022 , publisher=
Tejeda, Heliodoro and Kumar, Aakriti and Smyth, Padhraic and Steyvers, Mark , journal=. 2022 , publisher=
2022
-
[45]
2025 , publisher=
Gerlich, Michael , journal=. 2025 , publisher=
2025
-
[46]
Explanations can reduce overreliance on
Vasconcelos, Helena and J. Explanations can reduce overreliance on. Proceedings of the ACM on Human-Computer Interaction , volume=. 2023 , publisher=
2023
-
[47]
You can only verify when you know the answer: Feature-based explanations reduce overreliance on
Zhang, Zelun Tony and Buchner, Felicitas and Liu, Yuanting and Butz, Andreas , booktitle=. You can only verify when you know the answer: Feature-based explanations reduce overreliance on
-
[48]
Effective human oversight of
Langer, Markus and Baum, Kevin and Schlicker, Nadine , journal=. Effective human oversight of. 2024 , publisher=
2024
-
[49]
Computer Law & Security Review , volume=
The flaws of policies requiring human oversight of government algorithms , author=. Computer Law & Security Review , volume=. 2022 , publisher=
2022
-
[50]
Understanding trust and reliance development in
Kahr, Patricia K and Rooks, Gerrit and Willemsen, Martijn C and Snijders, Chris CP , journal=. Understanding trust and reliance development in. 2024 , publisher=
2024
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.