{"id":"d64aacb8-3ca0-4b91-aba6-280c918552c1","arxiv_id":"2501.12001","paper_version":2,"verdict":"REJECT","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"high","formal_verification":"none","parameter_count":4,"one_line_summary":"A progress bar with subtask markers added to a conversational AI chat measurably increased task-specific self-efficacy in a 22-person user study.","lead":"This paper describes a visual progress bar that appears alongside a conversational AI chat, marking each subtask of a goal-oriented task as completed. A 22-person user study reports that users with the progress bar showed larger self-efficacy gains than users with a plain chat interface.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Section 4.3 reports only within-group pre-post t-tests and Cohen's d; no direct between-group test on change scores is provided, so the abstract's claim of 'significant improvements compared to' the control group is not statistically established.","rationale":"The reader's weakest_assumption identifies the unvalidated Progress Feedback Agent as the load-bearing risk. I agree this is a serious construct-validity concern, but the most load-bearing issue is the missing between-group statistical test. The paper's headline claim is explicitly comparative: CPG users showed 'significant improvements in self-efficacy measures compared to' control users. Section 4.3's analyses—within-group pre-post t-tests and Cohen's d values—cannot support a comparative inference. The absence of a test on change scores or a group × time interaction means the central claim is not merely underpowered but untested. This is directly verifiable from the paper's own reported statistics: the authors state no between-group p-values for improvement and Table 4 shows heterogeneous effect sizes, including two items favoring the control group. The unvalidated evaluator is important for explaining the mechanism and for potential bias, but even a perfectly accurate evaluator would not make the current analysis support the comparative claim. Therefore, I concur with the reader's overall REJECT verdict, and my identified concern (statistical omission) differs from the reader's stated weakest assumption (evaluator validity). I selected 'disagree' to reflect that difference in the specific load-bearing concern, while noting the reader did identify the statistical issue in their strongest_claim. A revision adding the missing between-group test could potentially change the verdict to conditional acceptance, but as it stands the paper's primary claim is unsupported.","tokens_in":11898,"tokens_out":3795,"duration_ms":39981,"concrete_test":"Perform an independent-samples t-test (or a non-parametric equivalent such as Mann-Whitney U) on the change scores (post-task minus pre-task) for each of the six self-efficacy items, comparing the experimental and control groups. Alternatively, run a mixed ANOVA with group as a between-subjects factor and time (pre/post) as a within-subjects factor, and report the group × time interaction. If no significant interaction or change-score difference survives multiple-comparison correction, the abstract's comparative claim fails. This requires the raw per-participant data or the change scores, which the paper does not currently provide.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central empirical claim, stated in the abstract and repeated in Section 5.1 ('improved significantly compared to the control group'), requires a direct comparison of improvement between the two groups. However, Section 4.3 reports: (a) pre-task between-group t-tests showing no baseline differences, (b) separate within-group pre-post t-tests for the control and experimental groups, and (c) Cohen's d values in Table 4. No independent-samples test on change scores (post − pre) and no group × time interaction test are reported. Merely showing that each group improved, and that Cohen's d is larger for some items in the experimental group, is not a statistical comparison. Table 4 even shows Q1 and Q5 with larger effect sizes in the control group and Q6 nearly equal, so the pattern is mixed. With only 22 participants total, the claim of significant comparative improvement is unsupported by the analysis as presented. This is the most load-bearing concern because it directly concerns the headline result; the unvalidated Progress Feedback Agent, while relevant to mechanism, would still leave the statistical gap unresolved even if the evaluator were perfect.","agreement_with_reader":"disagree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces the Conversation Progress Guide (CPG), a UI that displays subtask-completion markers on a progress bar during text-based conversational AI interactions. The authors claim that CPG provides mastery experiences and thereby improves self-efficacy. They report a user study with 22 participants (control vs. experimental) performing an RSA encryption task, measuring self-efficacy before and after the task, cognitive load, satisfaction, task time, and interaction count. The paper concludes that CPG significantly improved self-efficacy compared to the default interface, without increasing cognitive load or harming satisfaction or task efficiency.","tokens_in":12043,"tokens_out":5505,"duration_ms":50800,"significance":"The topic is relevant and timely: self-efficacy is known to affect learning and persistence, and conversational AI failures may undermine it. The CPG concept is novel in this application domain, the implementation is functional, and the study includes a control group. The paper also candidly discusses several limitations, including issues with the progress evaluator and the self-efficacy scale. However, the central comparative claim is not supported by the reported statistical analysis, and the measurement design has a circularity problem. As presented, the evidence does not establish that CPG improves self-efficacy beyond the default interface. If properly validated and reanalyzed, the concept could be useful, but the current study is insufficient to support the headline claim.","major_comments":[{"comment":"The central claim that CPG led to 'significant improvements in self-efficacy measures compared to those using a conventional conversational AI' is not supported by the analysis. Section 4.3 reports (a) pre-task between-group t-tests showing no baseline differences, (b) within-group pre-post t-tests for each group, and (c) Cohen's d values in Table 4. No independent-samples test on change scores (post − pre) and no group × time interaction test is reported. Table 4 shows that Cohen's d is larger in the control group for Q1 and Q5 and similar for Q6, so the descriptive pattern is mixed. The conclusion of a significant between-group difference therefore rests on a comparison that was never performed.","section":"Abstract; Section 4.3; Section 5.1"},{"comment":"The Progress Feedback Agent, implemented with GPT-4, is the sole mechanism for lighting subtask markers, but its accuracy is never validated. No ground truth annotations, inter-rater agreement, or error rates are reported for the evaluation rules in Figure 4 and Figure 9. The paper itself documents failures: P7 saw the completion modal without understanding what triggered it, and P14 abandoned the experiment after the system failed to recognize task completion. If the agent frequently misjudges subtask completion, the purported mastery experiences are spurious. The reliability of the intervention needs to be established before its effect on self-efficacy can be interpreted.","section":"Section 3.2; Section 4.2; Table 5"},{"comment":"The self-efficacy survey items largely restate the subtasks displayed by the progress bar (e.g., Q2 asks about obtaining the RSA private key, which corresponds to the 'Private Key Generation' marker). Participants in the experimental group who see a marker light up may simply answer the matching survey item with higher confidence, so the observed gain may reflect direct feedback from the interface rather than a genuine change in self-efficacy. This item-wise overlap between the intervention display and the outcome measure is a measurement confound that threatens construct validity. The study needs a self-efficacy scale that does not directly mirror the displayed subtasks.","section":"Section 3.3; Section 4.1; Table 2"},{"comment":"The paper interprets null results as evidence that the CPG has no effect ('confirming the system does not increase cognitive load'). With 22 participants total (approximately 11 per group), the statistical power to detect differences in cognitive load, satisfaction, task time, and interaction count is very low. The absence of significant p-values should not be equated with evidence of equivalence. Exact p-values and confidence intervals should be reported, and the wording should reflect that these are null findings rather than demonstrations of no impact.","section":"Section 4.3, Cognitive Load and Task Time"}],"minor_comments":[{"comment":"There are typos: 'an web application' should be 'a web application', and 'For the Progress Feedback Agent, employed the GPT-4 model' is missing a subject ('we employed').","section":"Section 4.2"},{"comment":"The within-group comparisons are t-tests, but the type (paired vs. independent) is not explicitly stated; please state that paired t-tests were used for pre-post comparisons.","section":"Section 4.3"},{"comment":"Exact p-values are not reported for the non-significant tests (cognitive load, satisfaction, task time, interaction count); only 'above 0.05' is given. Reporting the exact values would aid interpretation.","section":"Section 4.3"},{"comment":"The keyword 'Conservation Interface' appears to be a typo; it likely should read 'Conversation Interface'.","section":"Keywords"},{"comment":"The version of Cohen's d used (e.g., pooled SD vs. SD of change scores) is not defined; please specify the formula.","section":"Section 4.3, Table 4"},{"comment":"Please report the number of participants in each group; the text only gives the total (22).","section":"Section 4.1"}],"recommendation":"reject","confidential_remarks":"The paper has a clear and useful idea, but the empirical evidence is centrally flawed. The missing between-group test is fixable by reanalysis, but the circular self-efficacy measure and the unvalidated progress evaluator would require a new study design with a different survey instrument and an accuracy assessment of the evaluator. These are not local revisions. I therefore recommend rejection, though I would encourage the authors to rework the study and resubmit."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"You asked about Jeong et al.'s CPG paper. The short version: the idea is simple and worth knowing about, but the central statistical claim does not hold up as reported. The authors show within-group pre-post gains in both conditions and compare Cohen's d values by inspection; they never run a between-group test on change scores, so \"significantly greater gains in the CPG group\" is not actually tested.\n\nWhat's genuinely new: applying visual progress feedback to generative conversational AI, with a working prototype (two GPT agents, progress bar, subtask markers) and a modest user study. The paper is clearly written, the design is well-motivated by Bandura, and the limitations section is unusually candid—it acknowledges the linear-subtask scope, the P7 case where the completion modal fired without comprehension, and the need for better measurement. That honesty counts for something.\n\nThe biggest soft spot is the analysis. With 22 total participants, comparing separate within-group t-tests and effect sizes across groups cannot support a comparative claim. A direct test on change scores or a group-by-time interaction is needed. The stress-test note is correct. Also, the Progress Feedback Agent is unvalidated; P14 abandoned the experiment because the system did not recognize completion, and P7 got a spurious modal. Since the intervention is entirely about those markers, evaluator accuracy is load-bearing. The self-efficacy survey items map one-to-one to the displayed subtasks, so there is a non-trivial circularity risk: the progress bar may cue the survey answers rather than build genuine efficacy. And no data or code are shared, so the reported effect sizes cannot be checked.\n\nWho is this for? Researchers working on human-AI interaction, self-efficacy, or UI patterns for conversational agents. They would get a clean description of one design approach and a useful cautionary example of underpowered HCI evaluation. The paper deserves a serious referee, but with the clear expectation of major revision: rerun the analysis with proper between-group tests, report confidence intervals, validate or at least hedge the evaluator, and ideally release materials. I would not desk-reject it—the idea is viable—but the current evidence does not support the abstract's claim.","headline":"Plausible progress-UI idea with a clean prototype, but the headline comparative claim rests on a between-group test that was never run.","tokens_in":12644,"tokens_out":1861,"would_cite":false,"duration_ms":20096,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper claims that a Conversation Progress Guide—a progress bar with subtask markers in a conversational AI interface—significantly improves users' self-efficacy compared to a conventional chat interface.","keywords":["self-efficacy","progress bar","conversational AI","human-AI interaction","mastery experience","user interface design","user study"],"falsifier":"A direct between-group statistical test on pre-to-post self-efficacy change scores (e.g., independent-samples t-test or ANCOVA with baseline covariate) would settle the comparative claim; if the difference is not significant at the conventional threshold, the claim that the CPG significantly improves self-efficacy relative to a conventional interface is not supported by the data. Additionally, measuring the evaluator's judgments against human raters on the same conversation logs would test the assumption that the markers reflect real progress.","tokens_in":11598,"feed_emoji":"📊","tokens_out":8858,"duration_ms":81049,"temperature":0.7,"pith_summary":"The paper introduces the Conversation Progress Guide (CPG), a visual interface for text-based conversational AI that shows users a progress bar with markers that light up as subtasks of their goal are completed. The central claim is that this interface raises users' self-efficacy—their belief in their ability to succeed at the task—by giving them frequent 'mastery experiences,' the most powerful source of self-efficacy in Bandura's theory. The authors argue that because conversational AI often produces failures and dissatisfaction, users' confidence can erode, and that providing visible partial successes can offset that erosion. A user study with 22 participants compared a CPG-enhanced chat to a conventional chat on a math/encryption task, and the paper reports significantly larger self-efficacy gains in the CPG group, with no significant differences in cognitive load, task time, or satisfaction. If the claim holds, it implies that a lightweight UI addition—not a change to the underlying AI—could make conversational assistants more empowering.","feed_headline":"Visual progress guide lifts self-efficacy in AI conversations","feed_subtitle":"A bar that lights up as subtasks finish made users feel more capable, with no added cognitive load.","key_machinery":"The central object is the Conversation Progress Guide (CPG), a UI layer consisting of a progress bar and subtask markers, driven by a separate Progress Feedback Agent. The agent re-evaluates the conversation history after each exchange and decides, based on hand-crafted evaluation rules, whether a predefined subtask (e.g., 'Multiplication of Primes') has been completed; when it returns a positive judgment, the interface activates the corresponding marker in the predefined order. This design translates the user's actual conversation into an accumulating visual record of partial achievements, which is the concrete carrier of the claimed self-efficacy effect.","core_discovery":"The paper's central claim is that a visual progress guide for text-based conversational AI—a progress bar whose markers light up as the user completes subtasks—produces significantly greater gains in task-specific self-efficacy than a conventional chat interface, without adding cognitive load or reducing satisfaction. The mechanism is the reinforcement of 'mastery experiences': each time a subtask is judged complete, a marker appears, giving the user visible evidence of success. This contrasts with ordinary chat, where failures and ambiguous responses can accumulate with no visible record of partial progress. In the authors' user study, 22 participants performed an RSA encryption task with either a CPG-enhanced GPT-3.5-based chat or the same chat without the progress display; both groups improved on a six-item self-efficacy survey, and the authors report that the CPG group improved significantly more, with larger effect sizes on most items. They conclude that the interface succeeds in strengthening self-efficacy by leveraging partial successes, and that the effect is achieved without harming task performance, efficiency, or user experience.","pith_inferences":["A plausible but untested extension is that the CPG's effect is largest for users who begin with low self-efficacy, since mastery experiences are thought to matter most when confidence is fragile; a stratified analysis by baseline score would reveal this.","The hand-crafted evaluation rules could be replaced by a learning-based subtask detector, which would let the interface generalize beyond the linear, pre-defined tasks used here without human rule authoring.","The present analysis compares within-group changes; a direct between-group test on change scores, such as an ANCOVA with baseline as covariate, would be a natural and more definitive statistical check of the comparative claim."],"forward_implications":["Layering the CPG onto an existing conversational AI requires no change to the underlying model, so any chat service with linearly decomposable tasks could adopt the same visual feedback.","Users who see their subtask completions accumulate may gain confidence in their ability to use conversational AI for multi-turn problem solving, which could increase their willingness to engage with such tools.","Because the interface adds no measurable cognitive load and does not alter task time or satisfaction, it offers a low-cost way to improve user experience in goal-oriented conversations.","The approach is deliberately limited to tasks with well-defined sub-steps; open-ended or purely social conversations would not receive meaningful progress markers."],"supporting_citations":[{"why":"Supplies the self-efficacy theory and the four sources, especially mastery experiences, that motivate the CPG design.","marker":"[4]"},{"why":"Provides the guidelines for constructing task-specific self-efficacy scales used to develop the survey instrument.","marker":"[5]"},{"why":"The NASA-TLX instrument used to measure cognitive load and show the CPG adds no burden.","marker":"[15]"},{"why":"Evidence that effort-based visual feedback strategies build self-efficacy in online learning, a direct precedent for the progress markers.","marker":"[16]"},{"why":"Documents the user dissatisfaction categories in conversational AI that motivate the need to strengthen self-efficacy.","marker":"[18]"},{"why":"Classic work on percent-done progress indicators that motivates the choice of a progress bar UI.","marker":"[24]"}],"fun_headline_variants":["Progress bar lifts self-efficacy in AI conversations","Visual progress guide boosts AI chat confidence","Subtask progress markers enhance AI user self-belief","Mastery feedback via progress bar improves AI self-efficacy","AI chat progress bar raises self-efficacy, no extra load"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The entire claim rests on the automated evaluator that decides when a subtask is finished; if that evaluator frequently misjudges completion, the progress markers would deliver spurious mastery experiences and the self-efficacy effect would not be genuine.","fun_headline_variants_meta":{"raw":{"variants":["Progress bar lifts self-efficacy in AI conversations","Visual progress guide boosts AI chat confidence","Subtask progress markers enhance AI user self-belief","Mastery feedback via progress bar improves AI self-efficacy","AI chat progress bar raises self-efficacy, no extra load"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000493,"raw_usage":{"total_tokens":2382,"prompt_tokens":865,"completion_tokens":1517,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":481,"completion_tokens_details":{"reasoning_tokens":1445}},"tokens_in":481,"tokens_out":1517,"duration_ms":13200,"temperature":1.0,"reasoning_tokens":1445,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T17:35:58.359891+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A direct between-group statistical test on pre-to-post self-efficacy change scores (e.g., independent-samples t-test or ANCOVA with baseline covariate) would settle the comparative claim; if the difference is not significant at the conventional threshold, the claim that the CPG significantly improves self-efficacy relative to a conventional interface is not supported by the data. Additionally, measuring the evaluator's judgments against human raters on the same conversation logs would test the assumption that the markers reflect real progress.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the self-efficacy theory and the four sources, especially mastery experiences, that motivate the CPG design."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the guidelines for constructing task-specific self-efficacy scales used to develop the survey instrument."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Evidence that effort-based visual feedback strategies build self-efficacy in online learning, a direct precedent for the progress markers."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Documents the user dissatisfaction categories in conversational AI that motivate the need to strengthen self-efficacy."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Classic work on percent-done progress indicators that motivates the choice of a progress bar UI."}],"review_version":1}