{"id":"cff7d0ff-8b1b-4927-9549-5eedb67a89bc","arxiv_id":"2607.05674","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.5,"correctness_risk":"low","formal_verification":"none","parameter_count":0,"one_line_summary":"A validated 6-item PSP scale shows perceived system predictability is distinct from objective prediction accuracy and is shifted by explanations but not by added stochasticity.","lead":"Researchers developed and validated a 6-item questionnaire that measures how predictable people feel an interactive system is, distinguishing three facets of that feeling. The scale shows that users' sense of predictability can diverge from how accurately they actually predict the system, and that explanations and randomness affect the two differently.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.5","headline":"The reported dissociation between PSP and objective correctness may be inflated by the transparent, low-stakes classifier and the illusion-of-depth mechanism the authors themselves invoke.","rationale":"The reader correctly identifies the transparent-classifier premise as the weakest load-bearing assumption. My concern is the same premise, sharpened: the dissociation itself (not merely the scale) may be an artifact of that transparency and of the illusion-of-depth mechanism the authors invoke. Because the paper already flags the limitation and the psychometric work is solid for the tested regimes, the appropriate verdict remains CONDITIONAL rather than REJECT. The concrete LLM replication would settle whether the dissociation is general or setting-bound.","tokens_in":30808,"tokens_out":499,"duration_ms":6277,"concrete_test":"Replicate the §4 design with an opaque LLM sentiment classifier, holding the same three explanation modalities (none / saliency / bar) and the same three noise levels, but generating explanations with a fixed post-hoc method (e.g., integrated gradients). If the explanation-format effect on PSP and the noise-level effect on correctness both remain significant while the noise-level effect on PSP remains non-significant, the dissociation claim is robust; if either dissociation collapses, the strongest claim is setting-specific.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central claim is that PSP and prediction correctness capture distinct aspects of users’ mental models and can diverge (explanations shift PSP but not correctness; noise degrades correctness without lowering PSP). That claim is demonstrated only on a fully transparent, rule-based SentiWordNet classifier whose decision rule is identical to the communicated explanation (§4.1.1) and on a fictional shape task. The authors themselves attribute the noise-invariance of PSP to the “illusion of explanatory depth” that arises when participants can fall back on their own sentiment judgments (§4.4). In a setting where the model is opaque and explanations are post-hoc (and therefore potentially unfaithful), both the explanation-induced PSP shift and the noise-invariance of PSP could shrink or reverse. The paper therefore treats as general a dissociation that may be an artifact of the very transparency that made the validation clean. The Limitations section acknowledges the external-validity gap but does not quantify how much of the reported dissociation depends on it.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.5","summary":"The paper introduces perceived system predictability (PSP) as a user-centered construct grounded in uncertainty theory, distinguishing epistemic, aleatory, and effective facets. It contributes a theoretical framing relative to trust and understanding, a 6-item scale developed from a 60-item pool via expert review and cognitive interviews, and psychometric validation in a fictional shape-classifier study (N=200) supporting both unidimensional and three-factor hierarchical structures with high reliability and known-groups differentiation. A second sentiment-classifier study (N=200) varies explanation modality and stochasticity and relates PSP to prediction correctness, trust (FOST), SIPA, and need for cognition. The central applied claim is that PSP and objective prediction correctness capture distinct aspects of users’ mental models and can diverge: PSP predicts correctness, explanations shift PSP but not correctness, and increased stochasticity degrades correctness without lowering PSP.","tokens_in":31074,"tokens_out":1378,"duration_ms":18928,"significance":"If the scale is psychometrically sound and the reported dissociations hold under broader conditions, this is a useful contribution to HCI and human-centered XAI: the field has lacked a dedicated, validated instrument for perceived predictability, and the paper shows that neither objective prediction tasks nor adjacent self-report constructs (trust, SIPA) are adequate substitutes. Strengths include a conventional multi-stage scale-development pipeline, CFA with both one- and three-factor models, known-groups tests, concurrent validity against SIPA, two adequately powered between-subjects studies, and GAM analyses that allow non-linear relations. The work also provides a public project page and a clear scoring example. These are genuine assets for cumulative research on calibrated reliance and transparent interactive systems.","major_comments":[{"comment":"Abstract and §§4–5 state a general dissociation (explanations shift PSP not correctness; noise degrades correctness without lowering PSP). That pattern is demonstrated only with a fully transparent SentiWordNet rule-based classifier whose decision rule coincides with the communicated explanation (§4.1.1) and, earlier, a fictional shape task. The authors themselves attribute noise-invariance of PSP to the illusion of explanatory depth when participants can fall back on their own sentiment judgments (§4.4). In opaque systems with post-hoc explanations, both the explanation-induced PSP shift and the noise-invariance could shrink or reverse. §6 acknowledges the gap but does not scope the abstract/conclusion claims accordingly. The central applied claim needs explicit boundary conditions (transparent vs opaque models; faithful vs post-hoc explanations) or additional evidence; otherwise the ge","section":"Abstract; §4.1.1; §4.4; §5; §6"},{"comment":"CFA supports both models, but Pearson correlations among the three facets are extremely high (0.901, 0.889, 0.876; §3.3.6), and parsimony indices favor the one-factor model. Retaining two items per facet is reasonable for reliability estimation, yet the manuscript does not give concrete decision rules for when subscale scores add interpretive value beyond the total mean. Given near-redundancy, claims that the hierarchical structure is practically usable need either (a) clearer guidance and example use-cases where facets diverge, or (b) stronger evidence that subscales discriminate differently under theoretically targeted manipulations (e.g., pure epistemic vs pure aleatory designs).","section":"§3.3.6; Table 4; Fig. 7"},{"comment":"Concurrent validity with SIPA is very strong (r=0.856, p<.001; §3.3.8 and Table 10), only modestly below internal PSP subscale correlations. The paper argues PSP is more focused and better predicts objective correctness than SIPA, which is an important differentiator, but the nomological network still leaves open how much unique variance PSP captures once SIPA’s predictability subscale is partialled out. A brief hierarchical or residual analysis (PSP total vs SIPA-predictability items predicting correctness/trust) would strengthen the claim that a dedicated instrument is necessary rather than a refined SIPA subscale.","section":"§3.3.8; §4.5; Table 10"}],"minor_comments":[{"comment":"Table 1 and the scoring example (Appendix A.3) are clear; consider stating in the main text whether reverse-coded items were considered and rejected, and whether item order is fixed in deployment (noted as fixed in §4 but randomized in validation).","section":"Table 1; §3.3.2; Appendix A.3"},{"comment":"In §4.3–4.4, multiple GAMs and post-hoc Wald contrasts are reported. A short note on family-wise error control (or why none is applied) would help readers interpret the p-values for explanation-format and noise contrasts.","section":"§4.3; §4.4; Tables 6–9"},{"comment":"Figure 13 normalizes PSP for visual comparison with correctness; state the normalization explicitly in the caption so the y-axis is not misread as raw Likert means.","section":"Fig. 13"},{"comment":"Minor wording: “repsponses” appears in Fig. 1 caption; “product [sic]” is already flagged in Table 2; a few long sentences in §1.1 could be split for readability.","section":"Fig. 1; Table 2; §1.1"},{"comment":"Recruitment is restricted to US/AU/UK MTurk workers. A sentence on language and cultural scope of the instrument would help future users decide on translation/adaptation needs.","section":"§3.3.3; §4.1.4"}],"recommendation":"major_revision","confidential_remarks":"Solid instrument paper with a real gap in the literature. The main risk is overgeneralization of the dissociation findings from a deliberately transparent test bed; if the authors scope claims tightly and add the residual/unique-variance analysis vs SIPA, this is close to a strong accept for an HCI methods/empirical venue. Fit is good for cs.HC; less so if the venue expects LLM-scale systems as the primary evaluation setting."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"Punchline: this is a competent instrument paper that gives HCI a short, validated measure of perceived system predictability, plus evidence that the subjective score and objective prediction correctness can move independently. That dissociation is the part worth remembering.\n\nWhat is actually new is not the word “predictability” — SIPA already had a two-item facet, and people have used ad-hoc sliders — but a dedicated six-item scale with an explicit epistemic/aleatory/effective framing, a full development pipeline (60-item pool, experts, cognitive interviews, CFA, known-groups), and two N=200 studies that relate PSP to trust, SIPA, NFC, and prediction correctness. The psychometrics look fine: high reliability, both one-factor and three-factor models fit, concurrent validity with SIPA is strong but not identity. The GAMs are a good choice; they surface non-linear links instead of forcing Pearson-only stories. The finding that explanations (especially saliency vs bars) shift PSP without moving correctness, while noise moves correctness without moving PSP, is the substantive payoff.\n\nSoft spots, in proportion. The three facets are highly correlated, so unidimensional scoring is what most people will use; the hierarchical story is more theoretical than empirically forced. Both validation settings are deliberately transparent (fictional shapes; SentiWordNet with faithful explanations). The authors say so, and they even invoke illusion of explanatory depth to explain why noise did not dent PSP. The stress-test concern is fair: the clean dissociation may partly depend on that transparency, and transfer to post-hoc explanations on LLMs is unproven. That is a real external-validity limit, not a reason to dismiss the instrument for the settings they tested. MTurk and low-stakes tasks are ordinary for this genre; they do not sink the work.\n\nMath and citation pattern look ordinary and honest — no circular “fit as prediction,” no invented prior art. Who it is for: anyone measuring mental models, explanation effects, or reliance calibration in interactive systems. I would bring it to a methods/XAI reading group, cite the scale when I need a predictability measure, and send it to peer review rather than desk-reject. Engage with it; treat the LLM transfer as the next experiment, not as already settled.","headline":"Solid, usable 6-item scale for perceived predictability with clean psychometrics and a real dissociation finding; external validity to opaque models is the open question, not a hidden flaw.","tokens_in":31637,"tokens_out":563,"would_cite":true,"duration_ms":10106,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"Users’ felt ability to anticipate a system’s behavior is a distinct construct from how correctly they actually forecast it, and a validated 6-item scale measures that perception.","keywords":["Scale Development","Validation","Questionnaire","Explainable AI","Human-AI Interaction","Perceived System Predictability","Trust","Mental Models"],"falsifier":"Rerun the sentiment study with a large language model and post-hoc explanations of deliberately low faithfulness: if explanation format still moves PSP while the correctness and trust patterns reverse or disappear, the claim that PSP generalizes as a model-agnostic construct is undercut.","tokens_in":31716,"feed_emoji":"🧭","tokens_out":991,"duration_ms":20733,"temperature":0.7,"pith_summary":"HCI has long measured whether people can correctly predict a system’s next outputs, but has lacked a clear concept and validated instrument for how predictable the system feels. This paper defines perceived system predictability (PSP) from uncertainty theory—epistemic (enough past observations), aleatory (consistent rather than random behavior), and effective (overall felt ability to anticipate)—and delivers a short 6-item questionnaire that works as a single score or three subscales. In two controlled studies, PSP and objective prediction correctness illuminate different parts of users’ mental models: explanations change how predictable the system feels without changing forecast accuracy, while added noise hurts accuracy without lowering PSP. PSP also relates to trust and situational information-processing awareness without collapsing into either. A sympathetic reader should care because designers who only track accuracy, trust, or usability can miss when users feel (or fail to feel) able to anticipate system behavior—the condition the authors treat as a prerequisite for calibrated reliance and accountable human–AI interaction.","feed_headline":"Felt predictability can diverge from actual forecast skill","feed_subtitle":"A 6-item scale shows explanations change how predictable a system feels without improving users’ forecasts.","key_machinery":"Perceived system predictability (PSP) and its 6-item scale. PSP is the degree to which a user feels able to predict how a system behaves, decomposed into epistemic, aleatory, and effective facets. Two items per facet yield an overall mean score and optional subscale scores; that instrument, validated via known-groups and nomological analyses, carries the claim that PSP is measurable and distinct from correctness and trust.","core_discovery":"Perceived system predictability is a user-centered construct that cannot be reduced to objective prediction correctness or to existing subjective measures such as trust. A 6-item scale grounded in epistemic, aleatory, and effective predictability shows strong psychometric properties and supports both unidimensional and three-factor use. In a sentiment-classifier study, PSP itself predicts how correctly users forecast system outputs; explanation format shifts PSP without shifting correctness; and increased stochasticity degrades correctness without lowering PSP. PSP therefore captures mental-model aspects that objective and adjacent subjective measures leave unaddressed.","pith_inferences":["If visualization choices can raise PSP without improving forecast skill, some explanation interfaces may create a false sense of predictability that encourages over-reliance in high-stakes settings.","Applying the scale to large language models will need explanation faithfulness as an explicit experimental factor, because post-hoc explanations could themselves inflate or deflate PSP.","Product teams could treat a lightweight prediction probe plus the PSP scale as paired early-warning signals after model or UX changes.","Bias-mitigating chart designs (such as cumulative bars) may be a practical way to keep felt predictability closer to actual predictive skill."],"forward_implications":["HCI evaluations of interactive AI should report PSP alongside objective prediction tasks, trust, and usability rather than treating any one as a proxy for the others.","Explanation designers can raise or lower how predictable a system feels (for example heatmap versus bar chart) without necessarily improving users’ actual forecasts of its outputs.","Raising system stochasticity can impair forecast accuracy while leaving felt predictability unchanged, so miscalibration may go undetected by self-report alone.","The scale enables both brief unidimensional screening and finer three-facet diagnosis when epistemic and aleatory levels are expected to differ.","Transparent, trustworthy system design can treat users’ ability to anticipate behavior as a first-class, measurable design target."],"fun_headline_variants":["Felt predictability diverges from users' forecast skill","Explanations shift felt predictability without better forecasts","Systems can feel predictable even when forecasts fail","Stochasticity cuts forecast skill while felt predictability holds","Prediction correctness and felt predictability often diverge"],"cache_read_input_tokens":16512,"weakest_assumption_plain":"The results rest on tightly controlled transparent classifiers—a fictional shape mapper and a rule-based sentiment model—rather than opaque modern systems where explanations may not match the true decision process.","fun_headline_variants_meta":{"raw":{"variants":["Felt predictability diverges from users' forecast skill","Explanations shift felt predictability without better forecasts","Systems can feel predictable even when forecasts fail","Stochasticity cuts forecast skill while felt predictability holds","Prediction correctness and felt predictability often diverge"]},"model":"grok-4.5","effort":"low","cost_usd":0.008144,"raw_usage":{"total_tokens":1930,"prompt_tokens":811,"num_sources_used":0,"completion_tokens":70,"cost_in_usd_ticks":81440000,"prompt_tokens_details":{"text_tokens":811,"audio_tokens":0,"image_tokens":0,"cached_tokens":128},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":1049,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":811,"tokens_out":70,"duration_ms":7997,"temperature":1.0,"reasoning_tokens":1049,"cache_read_input_tokens":128,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-11T03:58:13.690308+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"Rerun the sentiment study with a large language model and post-hoc explanations of deliberately low faithfulness: if explanation format still moves PSP while the correctness and trust patterns reverse or disappear, the claim that PSP generalizes as a model-agnostic construct is undercut.","supporting_citations":[],"review_version":1}