Pith. sign in

REVIEW 3 major objections 5 minor 95 references

Perceived System Predictability: Scale Development and Application

T0 review · 3 major / 5 minor · reviewed 2026-07-11 · grok-4.5

Pith's one-line read Users’ felt ability to anticipate a system’s behavior is a distinct construct from how correctly they actually forecast it, and a validated 6-item scale measures that perception.

desk verdict Solid, usable 6-item scale for perceived predictability with clean psychometrics and a real dissociation finding; external validity to opaque models is the open question, not a hidden flaw. read the letter →

arxiv 2607.05674 v1 pith:2Q4OIQQN submitted 2026-07-06 cs.HC

classification cs.HC
keywords ScaleDevelopmentValidationQuestionnaireExplainableAIHuman-AIInteractionPerceivedSystemPredictabilityTrustMentalModels
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

HCI has long measured whether people can correctly predict a system’s next outputs, but has lacked a clear concept and validated instrument for how predictable the system feels. This paper defines perceived system predictability (PSP) from uncertainty theory—epistemic (enough past observations), aleatory (consistent rather than random behavior), and effective (overall felt ability to anticipate)—and delivers a short 6-item questionnaire that works as a single score or three subscales. In two controlled studies, PSP and objective prediction correctness illuminate different parts of users’ mental models: explanations change how predictable the system feels without changing forecast accuracy, while added noise hurts accuracy without lowering PSP. PSP also relates to trust and situational information-processing awareness without collapsing into either. A sympathetic reader should care because designers who only track accuracy, trust, or usability can miss when users feel (or fail to feel) able to anticipate system behavior—the condition the authors treat as a prerequisite for calibrated reliance and accountable human–AI interaction.

What carries the argument

Perceived system predictability (PSP) and its 6-item scale. PSP is the degree to which a user feels able to predict how a system behaves, decomposed into epistemic, aleatory, and effective facets. Two items per facet yield an overall mean score and optional subscale scores; that instrument, validated via known-groups and nomological analyses, carries the claim that PSP is measurable and distinct from correctness and trust.

What would settle it

Rerun the sentiment study with a large language model and post-hoc explanations of deliberately low faithfulness: if explanation format still moves PSP while the correctness and trust patterns reverse or disappear, the claim that PSP generalizes as a model-agnostic construct is undercut.

Watch

Extended reading notes

Core claim

Perceived system predictability is a user-centered construct that cannot be reduced to objective prediction correctness or to existing subjective measures such as trust. A 6-item scale grounded in epistemic, aleatory, and effective predictability shows strong psychometric properties and supports both unidimensional and three-factor use. In a sentiment-classifier study, PSP itself predicts how correctly users forecast system outputs; explanation format shifts PSP without shifting correctness; and increased stochasticity degrades correctness without lowering PSP. PSP therefore captures mental-model aspects that objective and adjacent subjective measures leave unaddressed.

Load-bearing premise

The results rest on tightly controlled transparent classifiers—a fictional shape mapper and a rule-based sentiment model—rather than opaque modern systems where explanations may not match the true decision process.

Editorial extensions

If this is right

  • HCI evaluations of interactive AI should report PSP alongside objective prediction tasks, trust, and usability rather than treating any one as a proxy for the others.
  • Explanation designers can raise or lower how predictable a system feels (for example heatmap versus bar chart) without necessarily improving users’ actual forecasts of its outputs.
  • Raising system stochasticity can impair forecast accuracy while leaving felt predictability unchanged, so miscalibration may go undetected by self-report alone.
  • The scale enables both brief unidimensional screening and finer three-facet diagnosis when epistemic and aleatory levels are expected to differ.
  • Transparent, trustworthy system design can treat users’ ability to anticipate behavior as a first-class, measurable design target.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If visualization choices can raise PSP without improving forecast skill, some explanation interfaces may create a false sense of predictability that encourages over-reliance in high-stakes settings.
  • Applying the scale to large language models will need explanation faithfulness as an explicit experimental factor, because post-hoc explanations could themselves inflate or deflate PSP.
  • Product teams could treat a lightweight prediction probe plus the PSP scale as paired early-warning signals after model or UX changes.
  • Bias-mitigating chart designs (such as cumulative bars) may be a practical way to keep felt predictability closer to actual predictive skill.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. The paper introduces perceived system predictability (PSP) as a user-centered construct grounded in uncertainty theory, distinguishing epistemic, aleatory, and effective facets. It contributes a theoretical framing relative to trust and understanding, a 6-item scale developed from a 60-item pool via expert review and cognitive interviews, and psychometric validation in a fictional shape-classifier study (N=200) supporting both unidimensional and three-factor hierarchical structures with high reliability and known-groups differentiation. A second sentiment-classifier study (N=200) varies explanation modality and stochasticity and relates PSP to prediction correctness, trust (FOST), SIPA, and need for cognition. The central applied claim is that PSP and objective prediction correctness capture distinct aspects of users’ mental models and can diverge: PSP predicts correctness, explanations shift PSP but not correctness, and increased stochasticity degrades correctness without lowering PSP.

Significance. If the scale is psychometrically sound and the reported dissociations hold under broader conditions, this is a useful contribution to HCI and human-centered XAI: the field has lacked a dedicated, validated instrument for perceived predictability, and the paper shows that neither objective prediction tasks nor adjacent self-report constructs (trust, SIPA) are adequate substitutes. Strengths include a conventional multi-stage scale-development pipeline, CFA with both one- and three-factor models, known-groups tests, concurrent validity against SIPA, two adequately powered between-subjects studies, and GAM analyses that allow non-linear relations. The work also provides a public project page and a clear scoring example. These are genuine assets for cumulative research on calibrated reliance and transparent interactive systems.

major comments (3)
  1. [Abstract; §4.1.1; §4.4; §5; §6] Abstract and §§4–5 state a general dissociation (explanations shift PSP not correctness; noise degrades correctness without lowering PSP). That pattern is demonstrated only with a fully transparent SentiWordNet rule-based classifier whose decision rule coincides with the communicated explanation (§4.1.1) and, earlier, a fictional shape task. The authors themselves attribute noise-invariance of PSP to the illusion of explanatory depth when participants can fall back on their own sentiment judgments (§4.4). In opaque systems with post-hoc explanations, both the explanation-induced PSP shift and the noise-invariance could shrink or reverse. §6 acknowledges the gap but does not scope the abstract/conclusion claims accordingly. The central applied claim needs explicit boundary conditions (transparent vs opaque models; faithful vs post-hoc explanations) or additional evidence; otherwise the ge
  2. [§3.3.6; Table 4; Fig. 7] CFA supports both models, but Pearson correlations among the three facets are extremely high (0.901, 0.889, 0.876; §3.3.6), and parsimony indices favor the one-factor model. Retaining two items per facet is reasonable for reliability estimation, yet the manuscript does not give concrete decision rules for when subscale scores add interpretive value beyond the total mean. Given near-redundancy, claims that the hierarchical structure is practically usable need either (a) clearer guidance and example use-cases where facets diverge, or (b) stronger evidence that subscales discriminate differently under theoretically targeted manipulations (e.g., pure epistemic vs pure aleatory designs).
  3. [§3.3.8; §4.5; Table 10] Concurrent validity with SIPA is very strong (r=0.856, p<.001; §3.3.8 and Table 10), only modestly below internal PSP subscale correlations. The paper argues PSP is more focused and better predicts objective correctness than SIPA, which is an important differentiator, but the nomological network still leaves open how much unique variance PSP captures once SIPA’s predictability subscale is partialled out. A brief hierarchical or residual analysis (PSP total vs SIPA-predictability items predicting correctness/trust) would strengthen the claim that a dedicated instrument is necessary rather than a refined SIPA subscale.
minor comments (5)
  1. [Table 1; §3.3.2; Appendix A.3] Table 1 and the scoring example (Appendix A.3) are clear; consider stating in the main text whether reverse-coded items were considered and rejected, and whether item order is fixed in deployment (noted as fixed in §4 but randomized in validation).
  2. [§4.3; §4.4; Tables 6–9] In §4.3–4.4, multiple GAMs and post-hoc Wald contrasts are reported. A short note on family-wise error control (or why none is applied) would help readers interpret the p-values for explanation-format and noise contrasts.
  3. [Fig. 13] Figure 13 normalizes PSP for visual comparison with correctness; state the normalization explicitly in the caption so the y-axis is not misread as raw Likert means.
  4. [Fig. 1; Table 2; §1.1] Minor wording: “repsponses” appears in Fig. 1 caption; “product [sic]” is already flagged in Table 2; a few long sentences in §1.1 could be split for readability.
  5. [§3.3.3; §4.1.4] Recruitment is restricted to US/AU/UK MTurk workers. A sentence on language and cultural scope of the instrument would help future users decide on translation/adaptation needs.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity: empirical scale development and experimental dissociation, not a derivation that reduces to its own inputs.

full rationale

The paper is instrument development and nomological validation in HCI, not a first-principles derivation. PSP is defined as a user-centered construct (degree to which a user feels able to predict system behavior), decomposed into epistemic/aleatory/effective facets grounded in uncertainty theory and target-population interviews; the 6-item scale is obtained by item generation, expert review, cognitive interviews, and CFA/reliability/known-groups tests on an independent N=200 shape-classifier sample. Concurrent validity with SIPA (r=0.856) is reported as expected theoretical overlap yet weaker than internal subscale correlations, and is not used to force the construct. The application study (sentiment classifier, N=200) treats explanation format and noise as experimental factors and reports empirical dissociations via GAMs (explanations shift PSP but not correctness; noise degrades correctness without lowering PSP; PSP predicts correctness while objective correctness does not predict PSP). These are observed associations, not fitted parameters renamed as predictions, self-definitional identities, uniqueness theorems imported from the authors, or ansatzes smuggled via self-citation. Self-citations to the authors’ prior explanation work are background only and not load-bearing for the scale or the dissociation claims. The paper is therefore self-contained against its own benchmarks; external-validity limits (transparent classifier) are a separate concern, not circularity.

Assumptions & free parameters 0 free parameters · 4 assumptions · 1 invented entities

As an empirical scale-development and experimental paper, the load-bearing commitments are standard psychometric assumptions plus the transfer of epistemic/aleatory uncertainty language from decision theory to user perception. No free parameters are fitted to force a theoretical prediction; the three facets are interpretive constructs supported by interviews and CFA rather than derived quantities.

assumptions (4)
  • domain assumption Epistemic and aleatory uncertainty (Fox & Ülkümen) can be mapped onto users' perceived ability to predict system behavior, with an additional non-additive 'effective' facet.
    Introduced in §3.1.2–3.1.3 and used to structure the item pool and subscales.
  • standard math Standard classical test theory and CFA assumptions (items reflect latent factors, local independence, etc.) hold for the collected Likert responses.
    Implicit throughout §3.3 reliability and CFA analyses.
  • domain assumption MTurk participants from US/AU/UK provide usable data for construct validation of a general interactive-system perception scale.
    Sampling frame for all studies (N=20+25+200+200).
  • ad hoc to paper A transparent rule-based sentiment classifier is a sufficient first test bed for a model-agnostic perception construct.
    Explicit design choice in §4.1.1 justified by faithfulness needs; limits external validity.
invented entities (1)
  • Perceived System Predictability (PSP) construct with three facets (epistemic, aleatory, effective)
    purpose: Provide a dedicated user-centered construct and measurement target distinct from trust, SIPA, and objective prediction accuracy.
    Defined in abstract and §3.1; operationalized by the 6-item scale. Independent evidence is the psychometric and experimental results inside the paper itself; no external physiological or behavioral marker is offered.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Perceived System Predictability: Scale Development and Application." pith.science (2026). https://pith.science/paper/2Q4OIQQN

@misc{pith2026260705674,
  author       = {Pith},
  title        = {Pith review of: Perceived System Predictability: Scale Development and Application},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/2Q4OIQQN}},
  note         = {Machine review of arXiv:2607.05674}
}
abstract

How predictable users perceive an interactive system to be shapes how they interpret, trust, and rely on it, yet HCI lacks both a precise conceptualization and a validated instrument for this perception. We address this gap by introducing perceived system predictability (PSP) as a user-centered construct grounded in uncertainty theory, distinguishing epistemic, aleatory, and effective predictability. We contribute (i) a theoretical framework that situates PSP relative to adjacent constructs such as trust and understanding, (ii) a 6-item PSP scale, derived from a 60-item pool through expert review and cognitive interviews, and validated in a shape-classifier study ($N=200$) that supports both a unidimensional and a three-factor hierarchical structure, and (iii) a sentiment-classifier study ($N=200$) that varies explanations and stochasticity, and relates PSP to the correctness of users' predictions of system behavior, trust, subjective information processing awareness, and need for cognition. We find that PSP and prediction correctness capture distinct aspects of users' mental models and that both can diverge: PSP itself predicts correctness, explanations shift PSP but not correctness, and increased stochasticity degrades correctness without lowering PSP. PSP thus goes beyond existing objective and subjective measures and offers a principled foundation for designing transparent and trustworthy interactive systems.

Figures

Figures reproduced from arXiv: 2607.05674 by the authors.

Figure 1
Figure 1. Objective and perceived predictability capture distinct aspects of a user’s mental model of a system. The metaphor of two [PITH_FULL_IMAGE:figures/full_fig_p003_1.png] view at source ↗
Figure 2
Figure 2. PSP is a subjective self-report measure. We argue that perceived predictability should be assessed alongside objective [PITH_FULL_IMAGE:figures/full_fig_p004_2.png] view at source ↗
Figure 3
Figure 3. We decompose perceived predictability into three (partially overlapping) facets. In contrast to uncertainty in statistics, in [PITH_FULL_IMAGE:figures/full_fig_p008_3.png] view at source ↗
Figures from the paper (16 more)
Figure 4
Figure 4. Figure 4: Overview of our scale development process. For each stage, we report the number of participants in parentheses and the [PITH_FULL_IMAGE:figures/full_fig_p009_4.png]
Figure 5
Figure 5. Figure 5: One of the five fictional classification system prediction scenarios we show to participants. The class prediction for the blue [PITH_FULL_IMAGE:figures/full_fig_p010_5.png]
Figure 6
Figure 6. Figure 6: The three symbols we ask users to predict the system’s output for. Given the scenario shown in Figure [PITH_FULL_IMAGE:figures/full_fig_p010_6.png]
Figure 7
Figure 7. Figure 7: The two models we compare in confirmatory factor analysis (CFA) along with standardized coefficients. The rectangular boxes [PITH_FULL_IMAGE:figures/full_fig_p012_7.png]
Figure 8
Figure 8. Figure 8: The three explanations forms underlying the six explanation modalities used in our experiment. [PITH_FULL_IMAGE:figures/full_fig_p014_8.png]
Figure 9
Figure 9. Figure 9: Two of the three additional interactive explanation modalities. [PITH_FULL_IMAGE:figures/full_fig_p015_9.png]
Figure 10
Figure 10. Figure 10: Subset of the system predictions shown to users in the heatmap conditions. [PITH_FULL_IMAGE:figures/full_fig_p016_10.png]
Figure 11
Figure 11. Figure 11: Partial effect plots for factors with significant effects on perceived predictability in our GAM analysis. The plots show [PITH_FULL_IMAGE:figures/full_fig_p017_11.png]
Figure 12
Figure 12. Figure 12: Partial effect plots for factors with significant effects on objective prediction correctness in our analysis. The plots show [PITH_FULL_IMAGE:figures/full_fig_p020_12.png]
Figure 13
Figure 13. Figure 13: Boxplot showing the distributions of prediction correctness and normalized PSP scores for different levels of system [PITH_FULL_IMAGE:figures/full_fig_p021_13.png]
Figure 14
Figure 14. Figure 14: Enhanced bar chart explanation visualization using cumulative bars as proposed by Kang et al [PITH_FULL_IMAGE:figures/full_fig_p022_14.png]
Figure 15
Figure 15. Figure 15: A scenario with mixed uncertainty, but slightly more aleatory uncertainty than the scenario shown in Figure [PITH_FULL_IMAGE:figures/full_fig_p033_15.png]
Figure 16
Figure 16. Figure 16: A scenario with mixed uncertainty, but twice the number of examples of the scenario shown in Figure [PITH_FULL_IMAGE:figures/full_fig_p033_16.png]
Figure 17
Figure 17. Figure 17: A scenario with strong aleatory uncertainty (i.e., low predictability). We refer to this scenario as [PITH_FULL_IMAGE:figures/full_fig_p034_17.png]
Figure 18
Figure 18. Figure 18: A scenario with a high degree of epistemic and aleatory certainty. We refer to this scenario as [PITH_FULL_IMAGE:figures/full_fig_p034_18.png]
Figure 19
Figure 19. Figure 19: System predictions shown to users. Figure showing examples in the heatmap conditions, sentences were equal across [PITH_FULL_IMAGE:figures/full_fig_p036_19.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

95 extracted references · 19 canonical work pages

  1. [1]

    Alarcon, August A

    Gene M. Alarcon, August A. Capiola, Michael A. Lee, Sasha M. Willis, Izz Aldin Hamdan, Sarah A. Jessup, and Krista N. Harris. 2024. Development and Validation of the System Trustworthiness Scale.Hum. Factors66, 7 (2024), 1893–1913. https://doi.org/10.1177/00187208231189000

  2. [2]

    Stefano Baccianella, Andrea Esuli, and Fabrizio Sebastiani. 2010. SentiWordNet 3.0: An Enhanced Lexical Resource for Sentiment Analysis and Opinion Mining. InProceedings of the Seventh International Conference on Language Resources and Evaluation (LREC’10). European Language Resources Association (ELRA), Valletta, Malta. http://www.lrec-conf.org/proceedin...

  3. [3]

    Gagan Bansal, Tongshuang Wu, Joyce Zhou, Raymond Fok, Besmira Nushi, Ece Kamar, Marco Túlio Ribeiro, and Daniel S. Weld. 2021. Does the Whole Exceed its Parts? The Effect of AI Explanations on Complementary Team Performance. InCHI ’21: CHI Conference on Human Factors in Computing Systems, Virtual Event / Yokohama, Japan, May 8-13, 2021, Yoshifumi Kitamura...

  4. [4]

    Javier A Bargas-Avila and Florian Brühlmann. 2016. Measuring user rated language quality: development and validation of the user interface Language Quality Survey (LQS).International Journal of Human-Computer Studies86 (2016), 1–10

  5. [5]

    Paul C Beatty and Gordon B Willis. 2007. Research synthesis: The practice of cognitive interviewing.Public opinion quarterly71, 2 (2007), 287–311

  6. [6]

    Or Biran and Kathleen R. McKeown. 2017. Human-Centric Justification of Machine Learning Predictions. InProceedings of the Twenty-Sixth International Joint Conference on Artificial Intelligence, IJCAI 2017, Melbourne, Australia, August 19-25, 2017, Carles Sierra (Ed.). ijcai.org, 1461–1467. https://doi.org/10.24963/ijcai.2017/202 26 Schuff, Adel, and Vu

  7. [7]

    Boateng, Torsten B

    Godfred O. Boateng, Torsten B. Neilands, Edward A. Frongillo, Hugo R. Melgar-Quiñonez, and Sera L. Young. 2018. Best Practices for Developing and Validating Scales for Health, Social, and Behavioral Research: A Primer.Frontiers in Public Health6 (2018), 149. https://doi.org/10.3389/fpubh. 2018.00149

  8. [8]

    Bowling, Jason L

    Nathan A. Bowling, Jason L. Huang, Cheyna K. Brower, and Caleb B. Bragg. 2021. The Quick and the Careless: The Construct Validity of Page Time as a Measure of Insufficient Effort Responding to Surveys.Organizational Research Methods26 (2021), 323 – 352

Show all 95 references
  1. [9]

    John Brooke. 1996. SUS: a “quick and dirty’usability.Usability evaluation in industry(1996), 189

  2. [10]

    Gajos, and Elena L

    Zana Buçinca, Phoebe Lin, Krzysztof Z. Gajos, and Elena L. Glassman. 2020. Proxy tasks and subjective measures can be misleading in evaluating explainable AI systems. InIUI ’20: 25th International Conference on Intelligent User Interfaces, Cagliari, Italy, March 17-20, 2020, F...

  3. [11]

    Zana Buçinca, Maja Barbara Malaya, and Krzysztof Z. Gajos. 2021. To Trust or to Think: Cognitive Forcing Functions Can Reduce Overreliance on AI in AI-assisted Decision-making.Proc. ACM Hum. Comput. Interact.5, CSCW1 (2021), 188:1–188:21. https://doi.org/10.1145/3449287

  4. [12]

    Adrian Bussone, Simone Stumpf, and Dympna O’Sullivan. 2015. The Role of Explanations on Trust and Reliance in Clinical Decision Support Systems. In2015 International Conference on Healthcare Informatics, ICHI 2015, Dallas, TX, USA, October 21-23, 2015, Prabhakaran Balakrishnan...

  5. [13]

    Chris Callison-Burch, Miles Osborne, and Philipp Koehn. 2006. Re-evaluating the Role of Bleu in Machine Translation Research. In11th Conference of the European Chapter of the Association for Computational Linguistics. Association for Computational Linguistics, Trento, Italy, 2...

  6. [14]

    Carpinella, Alisa B

    Colleen M. Carpinella, Alisa B. Wyman, Michael A. Perez, and Steven J. Stroessner. 2017. The Robotic Social Attributes Scale (RoSAS): Development and Validation. InProceedings of the 2017 ACM/IEEE International Conference on Human-Robot Interaction. ACM, Vienna Austria, 254–26...

  7. [15]

    Wm Casper, Bryan D Edwards, J Craig Wallace, Ronald S Landis, Dustin A Fife, et al. 2020. Selecting response anchors with equal intervals for summated rating scales.Journal of Applied Psychology105, 4 (2020), 390

  8. [16]

    Eric Chu, Deb Roy, and Jacob Andreas. 2020. Are Visual Explanations Useful? A Case Study in Model-in-the-Loop Prediction.CoRRabs/2007.12248 (2020). arXiv:2007.12248 https://arxiv.org/abs/2007.12248

  9. [17]

    Julien Colin, Thomas Fel, Rémi Cadène, and Thomas Serre. 2022. What I Cannot Predict, I Do Not Understand: A Human-Centered Evaluation Framework for Explainability Methods. InAdvances in Neural Information Processing Systems 35: Annual Conference on Neural Information Processi...

  10. [18]

    Henriette S. M. Cramer, Vanessa Evers, Satyan Ramlal, Maarten van Someren, Lloyd Rutledge, Natalia Stash, Lora Aroyo, and Bob J. Wielinga. 2008. The effects of transparency on trust in and acceptance of a content-based art recommender.User Model. User Adapt. Interact.18, 5 (20...

  11. [19]

    Lee Joseph Cronbach. 1951. Coefficient alpha and the internal structure of tests.Psychometrika16 (1951), 297–334

  12. [20]

    Mary Czerwinski, Eric Horvitz, and Edward Cutrell. 2001. Subjective duration assessment: An implicit probe for software usability. InProceedings of IHM-HCI 2001 conference, Vol. 2. 167–170

  13. [21]

    Gabriel Lins de Holanda Coelho, Paul H. P. Hanel, and Lukas J Wolf. 2018. The Very Efficient Assessment of Need for Cognition: Developing a Six-Item Version*.Assessment27 (2018), 1870 – 1885

  14. [22]

    John P. Deegan. 1978. On the Occurrence of Standardized Regression Coefficients Greater Than One.Educational and Psychological Measurement38 (1978), 873 – 888

  15. [23]

    McCarty, Kelly Miller, Kristina Callaghan, and Greg Kestin

    Louis Deslauriers, Logan S. McCarty, Kelly Miller, Kristina Callaghan, and Greg Kestin. 2019. Measuring actual learning versus feeling of learning in response to being actively engaged in the classroom.Proceedings of the National Academy of Sciences116, 39 (2019), 19251–19257....

  16. [24]

    2021.Scale development: Theory and applications

    Robert F DeVellis and Carolyn T Thorpe. 2021.Scale development: Theory and applications. Sage publications

  17. [25]

    Birk Diedenhofen and Jochen Musch. 2015. cocor: A Comprehensive Solution for the Statistical Comparison of Correlations.PLoS ONE10 (2015)

  18. [26]

    Karen Dunn and Gareth McCray. 2020. The Place of the Bifactor Model in Confirmatory Factor Analysis Investigations Into Construct Dimensionality in Language Testing.Frontiers in Psychology11 (2020)

  19. [27]

    Olive Jean Dunn and Virginia A. Clark. 1969. Correlation Coefficients Measured on the Same Individuals.J. Amer. Statist. Assoc.64 (1969), 366–377

  20. [28]

    Rob Eisinga, Manfred te Grotenhuis, and Ben Pelzer. 2013. The reliability of a two-item scale: Pearson, Cronbach, or Spearman-Brown?International Journal of Public Health58 (2013), 637–642

  21. [29]

    Hannes Eisler. 1976. Experiments on subjective duration 1868-1975: A collection of power function exponents.Psychological Bulletin83, 6 (1976), 1154

  22. [30]

    Mica R Endsley. 1988. Situation awareness global assessment technique (SAGAT). InProceedings of the IEEE 1988 national aerospace and electronics conference. IEEE, 789–795

  23. [31]

    Boyd-Graber

    Shi Feng and Jordan L. Boyd-Graber. 2019. What can AI do for me?: evaluating machine learning interpretations in cooperative play. InProceedings of the 24th International Conference on Intelligent User Interfaces, IUI 2019, Marina del Ray, CA, USA, March 17-20, 2019, Wai-Tat F...

  24. [32]

    Kraig Finstad. 2010. The usability metric for user experience.Interacting with Computers22, 5 (2010), 323–327. Perceived System Predictability: Scale Development and Application 27

  25. [33]

    Kraig Finstad. 2013. Response to commentaries on ’The Usability Metric for User Experience’.Interact. Comput.25, 4 (2013), 327–330. https: //doi.org/10.1093/iwc/iwt005

  26. [34]

    Kenny, and Mark T

    Courtney Ford, Eoin M. Kenny, and Mark T. Keane. 2020. Play MNIST For Me! User Studies on the Effects of Post-Hoc, Example-Based Explanations & Error Rates on Debugging a Deep Learning, Black-Box Classifier.CoRRabs/2009.06349 (2020). arXiv:2009.06349 https://arxiv.org/abs/2009.06349

  27. [35]

    Distinguishing Two Dimensions of Uncertainty,

    Craig R Fox and Gülden Ülkümen. 2011. Distinguishing two dimensions of uncertainty.Fox, Craig R. and Gülden Ülkümen (2011), “Distinguishing Two Dimensions of Uncertainty, ” in Essays in Judgment and Decision Making, Brun, W., Kirkebøen, G. and Montgomery, H., eds. Oslo: Univer...

  28. [36]

    Krems, Viktoria Zott, and Andreas Keinath

    Thomas Franke, Maria Trantow, Madlen Günther, Josef F. Krems, Viktoria Zott, and Andreas Keinath. 2015. Advancing electric vehicle range displays for enhanced user experience: the relevance of trust and adaptability. InProceedings of the 7th International Conference on Automot...

  29. [37]

    2022.Psychometrics: an introduction

    R Michael Furr. 2022.Psychometrics: an introduction. SAGE publications

  30. [38]

    Ana Valeria Gonzalez, Anna Rogers, and Anders Søgaard. 2021. On the Interaction of Belief Bias and Explanations. InFindings of the Association for Computational Linguistics: ACL/IJCNLP 2021, Online Event, August 1-6, 2021 (Findings of ACL, Vol. ACL/IJCNLP 2021), Chengqing Zong...

  31. [39]

    Yash Goyal, Ziyan Wu, Jan Ernst, Dhruv Batra, Devi Parikh, and Stefan Lee. 2019. Counterfactual Visual Explanations. InProceedings of the 36th International Conference on Machine Learning, ICML 2019, 9-15 June 2019, Long Beach, California, USA (Proceedings of Machine Learning ...

  32. [40]

    Ben Green and Yiling Chen. 2019. The Principles and Limits of Algorithm-in-the-Loop Decision Making.Proc. ACM Hum. Comput. Interact.3, CSCW (2019), 50:1–50:24. https://doi.org/10.1145/3359152

  33. [41]

    Jonathan Grudin and Allan MacLean. 1985. Adapting A Psychophysical Method To Measure Performance And Preference Tradeoffs In Human- Computer Interaction. InHuman-Computer Interaction - INTERACT ’84(human-computer interaction - interact ’84 ed.). Elsevier Science Publishers B.V...

  34. [42]

    Felix Haag. 2025. The Effect of Explainable AI-based Decision Support on Human Task Performance: A Meta-Analysis

  35. [43]

    Sandra G Hart and Lowell E Staveland. 1988. Development of NASA-TLX (Task Load Index): Results of empirical and theoretical research. In Advances in psychology. Vol. 52. Elsevier, 139–183

  36. [44]

    Peter Hase and Mohit Bansal. 2020. Evaluating Explainable AI: Which Algorithmic Explanations Help Users Predict Model Behavior?. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics. Association for Computational Linguistics, Online, 5540–...

  37. [45]

    Herlocker, Joseph A

    Jonathan L. Herlocker, Joseph A. Konstan, and John Riedl. 2000. Explaining collaborative filtering recommendations. InCSCW 2000, Proceeding on the ACM 2000 Conference on Computer Supported Cooperative Work, Philadelphia, PA, USA, December 2-6, 2000, Wendy A. Kellogg and Steve ...

  38. [46]

    Kasper Hornbæk. 2006. Current practice in measuring usability: Challenges to usability studies and research.International Journal of Human- Computer Studies64, 2 (2006), 79–102. https://doi.org/10.1016/j.ijhcs.2005.06.002

  39. [47]

    Li-tze Hu and Peter M. Bentler. 1999. Cutoff criteria for fit indexes in covariance structure analysis : Conventional criteria versus new alternatives. Structural Equation Modeling6 (1999), 1–55

  40. [48]

    Alon Jacovi, Hendrik Schuff, Heike Adel, Ngoc Thang Vu, and Yoav Goldberg. 2023. Neighboring Words Affect Human Interpretation of Saliency Explanations. InFindings of the Association for Computational Linguistics: ACL 2023, Anna Rogers, Jordan Boyd-Graber, and Naoaki Okazaki (...

  41. [49]

    Karl G Jöreskog. 1999. How large can a standardized coefficient be. (1999)

  42. [50]

    Jorgensen, Sunthud Pornprasertmanit, Alexander M

    Terrence D. Jorgensen, Sunthud Pornprasertmanit, Alexander M. Schoemann, and Yves Rosseel. 2022. semTools: Useful tools for structural equation modeling. https://CRAN.R-project.org/package=semTools R package version 0.5-6

  43. [51]

    Hyun Seung Kang, Jeayeong Ji, Yeji Yun, and Kwang Hee Han. 2021. Estimating Bar Graph Averages: Overcoming Within-the-Bar Bias.i-Perception 12 (2021)

  44. [52]

    Anjali Khurana, Parsa Alamzadeh, and Parmit K. Chilana. 2021. ChatrEx: Designing Explainable Chatbot Interfaces for Enhancing Usefulness, Transparency, and Trust. InIEEE Symposium on Visual Languages and Human-Centric Computing, VL/HCC 2021, St Louis, MO, USA, October 10-13, 2...

  45. [53]

    Jenia Kim, Henry Maathuis, and Danielle Sent. 2024. Human-centered evaluation of explainable AI applications: a systematic review.Frontiers Artif. Intell.7 (2024). https://doi.org/10.3389/FRAI.2024.1456486

  46. [54]

    Matthias Kirchler, Martin Graf, Marius Kloft, and Christoph Lippert. 2021. Explainability Requires Interactivity.CoRRabs/2109.07869 (2021). arXiv:2109.07869 https://arxiv.org/abs/2109.07869

  47. [55]

    Moritz Körber. 2018. Theoretical considerations and development of a questionnaire to measure trust in automation. InCongress of the International Ergonomics Association. Springer, 13–30

  48. [56]

    Why is ’Chicago’ deceptive?

    Vivian Lai, Han Liu, and Chenhao Tan. 2020. "Why is ’Chicago’ deceptive?" Towards Building Model-Driven Tutorials for Humans. InCHI ’20: CHI Conference on Human Factors in Computing Systems, Honolulu, HI, USA, April 25-30, 2020, Regina Bernhaupt, Florian ’Floyd’ Mueller, David...

  49. [57]

    Vivian Lai and Chenhao Tan. 2019. On Human Predictions with Explanations and Predictions of Machine Learning Models: A Case Study on Deception Detection. InProceedings of the Conference on Fairness, Accountability, and Transparency, FAT* 2019, Atlanta, GA, USA, January 29-31, ...

  50. [58]

    James R. Lewis. 2018. Measuring Perceived Usability: The CSUQ, SUS, and UMUX.International Journal of Human–Computer Interaction34 (2018), 1148 – 1156

  51. [59]

    Qingyu Liang and Jaime Banks. 2025. Perceived shared understanding between humans and artificial intelligence: Development and validation of a self-report scale.Technology, Mind, and Behavior(2025)

  52. [60]

    Q Vera Liao, Yunfeng Zhang, Ronny Luss, Finale Doshi-Velez, and Amit Dhurandhar. 2022. Connecting Algorithmic Research and Usage Contexts: A Perspective of Contextualized Evaluation for Explainable AI. InProceedings of the AAAI Conference on Human Computation and Crowdsourcing...

  53. [61]

    Chia-Wei Liu, Ryan Lowe, Iulian Serban, Mike Noseworthy, Laurent Charlin, and Joelle Pineau. 2016. How NOT To Evaluate Your Dialogue System: An Empirical Study of Unsupervised Evaluation Metrics for Dialogue Response Generation. InProceedings of the 2016 Conference on Empirica...

  54. [62]

    Lundberg, Bala G

    Scott M. Lundberg, Bala G. Nair, Monica S. Vavilala, Mayumi Horibe, Michael J. Eisses, Trevor Adams, David Liston, Daniel King-Wai Low, Shu-Fang Newman, Jerry H. Kim, and Su-In Lee. 2018. Explainable machine-learning predictions for the prevention of hypoxaemia during surgery....

  55. [63]

    2023.sjPlot: Data Visualization for Statistics in Social Science

    Daniel Lüdecke. 2023.sjPlot: Data Visualization for Statistics in Social Science. https://CRAN.R-project.org/package=sjPlot R package version 2.8.14

  56. [64]

    A MacLean, PJ Barnard, and MD Wilson. 1985. Evaluating the human interface of a data entry system: user choice and performance measures yield different tradeoff functions.People and computers: Designing the interface5, 7 (1985), 45–61

  57. [65]

    Andreas Madsen, Siva Reddy, and Sarath Chandar. 2023. Post-hoc Interpretability for Neural NLP: A Survey.ACM Comput. Surv.55, 8 (2023), 155:1–155:42. https://doi.org/10.1145/3546577

  58. [66]

    Mcdonald

    Roderick P. Mcdonald. 1999. Test Theory: A Unified Treatment

  59. [67]

    N Menold and K Bogner. 2016. Design of rating scales in questionnaires.GESIS survey guidelines4 (2016)

  60. [68]

    Newman and Brian J

    George E. Newman and Brian J. Scholl. 2012. Bar graphs depicting averages are perceptually misinterpreted: The within-the-bar bias.Psychonomic Bulletin & Review19 (2012), 601–607

  61. [69]

    Jakob Nielsen and Jonathan Levy. 1994. Measuring Usability: Preference vs. Performance.Commun. ACM37, 4 (1994), 66–75. https://doi.org/10.114 5/175276.175282

  62. [70]

    Mahsan Nourani, Samia Kabir, Sina Mohseni, and Eric D. Ragan. 2019. The Effects of Meaningful and Meaningless Explanations on Trust and Perceived System Accuracy in Intelligent Systems.Proceedings of the AAAI Conference on Human Computation and Crowdsourcing7, 1 (Oct. 2019), 9...

  63. [71]

    Jekaterina Novikova, Ondřej Dušek, Amanda Cercas Curry, and Verena Rieser. 2017. Why We Need New Evaluation Metrics for NLG. InProceedings of the 2017 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, Copenhagen, Denmark...

  64. [72]

    Goldstein, Jake M

    Forough Poursabzi-Sangdeh, Daniel G. Goldstein, Jake M. Hofman, Jennifer Wortman Vaughan, and Hanna M. Wallach. 2021. Manipulating and Measuring Model Interpretability. InCHI ’21: CHI Conference on Human Factors in Computing Systems, Virtual Event / Yokohama, Japan, May 8-13, ...

  65. [73]

    Tenko Raykov. 2001. Estimation of congeneric scale reliability using covariance structure analysis with nonlinear constraints.The British journal of mathematical and statistical psychology54 Pt 2 (2001), 315–23

  66. [74]

    Ehud Reiter. 2018. A Structured Review of the Validity of BLEU.Computational Linguistics44, 3 (Sept. 2018), 393–401. https://doi.org/10.1162/coli _a_00322

  67. [75]

    2022.psych: Procedures for Psychological, Psychometric, and Personality Research

    William Revelle. 2022.psych: Procedures for Psychological, Psychometric, and Personality Research. Northwestern University, Evanston, Illinois. https://CRAN.R-project.org/package=psych R package version 2.2.9

  68. [76]

    Mireia Ribera and Àgata Lapedriza. 2019. Can we do better explanations? A proposal of user-centered explainable AI. InJoint Proceedings of the ACM IUI 2019 Workshops co-located with the 24th ACM Conference on Intelligent User Interfaces (ACM IUI 2019), Los Angeles, USA, March ...

  69. [77]

    Delphine Ribes, Nicolas Henchoz, Hélène Portier, Lara Défayes, Thanh-Trung Phan, Daniel Gatica-Perez, and Andreas Sonderegger. 2021. Trust Indicators and Explainable AI: A Study on User Perceptions. InHuman-Computer Interaction - INTERACT 2021 - 18th IFIP TC 13 International C...

  70. [78]

    Leon Rozenblit and Frank C. Keil. 2002. The misunderstood limits of folk science: an illusion of explanatory depth.Cognitive science26 5 (2002), 521–562. Perceived System Predictability: Scale Development and Application 29

  71. [79]

    Heleen Rutjes, Martijn Willemsen, and Wijnand IJsselsteijn. 2019. Considerations on explainable AI and users’ mental models. InWhere is the Human? Bridging the Gap Between AI and HCI. Association for Computing Machinery, Inc, United States. CHI 2019 Workshop : Where is the Hum...

  72. [80]

    Viktor Schlegel, Erick Mendez Guzman, and Riza Batista-Navarro. 2022. Towards Human-Centred Explainability Benchmarks For Text Classification. InWorkshop Proceedings of the 16th International AAAI Conference on Web and Social Media, ICWSM 2022 Workshops, Atlanta, Georgia, USA ...

  73. [81]

    Tim P P Schrills, Susanne Kargl, Mona Bickel, and Thomas Franke. 2022. Perceive, Understand & Predict - Empirical Indication for Facets in Subjective Information Processing Awareness. https://doi.org/10.31234/osf.io/3n95u

  74. [82]

    Hendrik Schuff, Heike Adel, Peng Qi, and Ngoc Thang Vu. 2022. Challenges in Explanation Quality Evaluation.CoRRabs/2210.07126 (2022). https://doi.org/10.48550/ARXIV.2210.07126 arXiv:2210.07126

  75. [83]

    Hendrik Schuff, Heike Adel, and Ngoc Thang Vu. 2020. F1 is Not Enough! Models and Evaluation Towards User-Centered Explainable Question Answering. InProceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP). Association for Computational Li...

  76. [84]

    Hendrik Schuff, Alon Jacovi, Heike Adel, Yoav Goldberg, and Ngoc Thang Vu. 2022. Human Interpretation of Saliency-Based Explanation Over Text. In2022 ACM Conference on Fairness, Accountability, and Transparency(Seoul, Republic of Korea)(FAccT ’22). Association for Computing Ma...

  77. [85]

    Tenenbaum, David N

    Eric Schulz, Joshua B. Tenenbaum, David N. Reshef, Maarten Speekenbrink, and Samuel Gershman. 2015. Assessing the Perceived Predictability of Functions. InProceedings of the 37th Annual Meeting of the Cognitive Science Society, CogSci 2015, Pasadena, California, USA, July 22-2...

  78. [86]

    Gombolay

    Andrew Silva, Mariah Schrum, Erin Hedlund-Botti, Nakul Gopalan, and Matthew C. Gombolay. 2023. Explainable Artificial Intelligence: Evaluating the Objective and Subjective Impacts of xAI on Human-Agent Interaction.Int. J. Hum. Comput. Interact.39, 7 (2023), 1390–1404. https: /...

  79. [87]

    Elior Sulem, Omri Abend, and Ari Rappoport. 2018. BLEU is Not Suitable for the Evaluation of Text Simplification. InProceedings of the 2018 Conference on Empirical Methods in Natural Language Processing. Association for Computational Linguistics, Brussels, Belgium, 738–744. ht...

  80. [88]

    Rachel L Thomas and David Uminsky. 2022. Reliance on metrics is a fundamental challenge for AI.Patterns3, 5 (2022), 100476. Publisher: Elsevier

  81. [89]

    Noam Tractinsky and Joachim Meyer. 2001. Task structure and the apparent duration of hierarchical search.International Journal of Human-Computer Studies55, 5 (2001), 845–860. https://doi.org/10.1006/ijhc.2001.0506

  82. [90]

    Pei Wang and Nuno Vasconcelos. 2020. SCOUT: Self-Aware Discriminant Counterfactual Explanations. In2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2020, Seattle, W A, USA, June 13-19, 2020. Computer Vision Foundation / IEEE, 8978–8987. https://doi.org...

  83. [91]

    Xinru Wang and Ming Yin. 2021. Are Explanations Helpful? A Comparative Study of the Effects of Explanations in AI-Assisted Decision-Making. In IUI ’21: 26th International Conference on Intelligent User Interfaces, College Station, TX, USA, April 13-17, 2021, Tracy Hammond, Kat...

  84. [92]

    2004.Cognitive interviewing: A tool for improving questionnaire design

    Gordon B Willis. 2004.Cognitive interviewing: A tool for improving questionnaire design. sage publications

  85. [93]

    Yei-Yu Yeh and Christopher D. Wickens. 1988. Dissociation of Performance and Subjective Measures of Workload.Human Factors30, 1 (1988), 111–120. https://doi.org/10.1177/001872088803000110 arXiv:https://doi.org/10.1177/001872088803000110

  86. [94]

    Vera Liao, and Rachel K

    Yunfeng Zhang, Q. Vera Liao, and Rachel K. E. Bellamy. 2020. Effect of confidence and explanation on accuracy and trust calibration in AI-assisted decision making. InFAT* ’20: Conference on Fairness, Accountability, and Transparency, Barcelona, Spain, January 27-30, 2020, Mire...

  87. [95]

    My knowledge about the system behavior is complete

    Guang Yong Zou. 2007. Toward using confidence intervals to compare correlations.Psychological methods12 4 (2007), 399–413. 30 Schuff, Adel, and Vu A ITEM GENERATION A.1 Initial Item Pool Our initial item pool contains 60 items. Concretely, these items are: •"My knowledge about...

Pith tools

Reviewed July 11, 2026 · model on record in the stance chip above.