{"id":"6bce2bb7-492e-4482-a9f8-ff96d8bbf96d","arxiv_id":"2411.15597","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"An 81-participant experiment found that both conventional and scaffolding chatbots improved comprehension of dashboard visualisations, and higher GenAI literacy predicted larger gains, particularly with the conventional chatbot.","lead":"In a study of 81 medical and nursing students, learners who could chat with a conventional or scaffolding chatbot while reading a learning analytics dashboard improved their comprehension scores. Higher generative AI literacy was associated with larger gains, especially with the conventional chatbot, while scaffolding showed a weaker literacy dependence.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"No-chatbot control and counterbalancing are absent, so the within-subject comprehension gains attributed to the chatbots may be practice effects.","rationale":"The reader's weakest_assumption identifies the same core issue: the fixed-order, within-subject design without a no-chatbot control or counterbalanced question sets leaves the comprehension gains open to practice/familiarity explanations. This is indeed the most load-bearing concern because RQ1 is the empirical foundation for the paper's central claim, and the literacy results (RQ2) presuppose that the chatbots produced the improvements. The concern is not about statistical error or internal inconsistency in the reported numbers, but about causal attribution. Since the reader already flagged this and assigned a CONDITIONAL verdict, my stress-test does not change the verdict. I considered other issues, such as the mislabelling of the Wilcoxon test as 'Mann-Whitney U' in Section 4.2 and the ENA comparison at p = 0.04 with a stated Bonferroni correction, but these are secondary or likely typographical, and they do not affect the primary causal claim. The proposed control-arm experiment would directly resolve whether the observed gains are specific to the chatbots.","tokens_in":17553,"tokens_out":4000,"duration_ms":39250,"concrete_test":"Conduct a follow-up experiment with a third, no-chatbot control arm: participants complete the same baseline task and a second task in the same fixed order but without any chatbot, interacting only with the dashboard. If the no-chatbot control group shows a mean Improvement_score comparable to the conventional or scaffolding chatbot groups (e.g., within a pre-specified non-inferiority margin of 0.5 points on the 0-6 scale), the claim that the chatbots causatively improved comprehension is not supported. Optionally, counterbalance the two question sets across phases to additionally rule out set-difficulty effects.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 3.3.3 fixes the order so that every participant first completes a baseline analytical writing task and six evaluation questions, then the assigned chatbot condition with a second set of questions. There is no no-chatbot control arm, and the two question sets are not counterbalanced across phases. The same visualisation types appear in both phases, and the evaluation questions are similar in format, differing only in the underlying data and insights. Consequently, the Baseline_score to Intervention_score improvement reported in Section 4.2 could be driven by practice with the visualisations, familiarisation with the question format, or differences in difficulty between the two sets. The authors note (Section 3.3) that participants had no prior knowledge of the learning context, but this addresses prior content knowledge, not task practice. Because the central claim that 'both conventional and scaffolding GenAI chatbots significantly improved learners' comprehension' rests entirely on this within-subject contrast, the lack of control and counterbalancing is the most load-bearing threat to the paper's main conclusion.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper reports a between-subject experiment with 81 medical/nursing students who completed a baseline analytical writing task and six comprehension questions, were then assigned to either a conventional (reactive) or scaffolding (proactive) GenAI chatbot in a learning analytics dashboard, and completed a second, parallel writing task and question set with chatbot support. GenAI literacy was measured with the GLAT instrument, and learner-chatbot interactions were analysed with epistemic network analysis. The paper reports that both chatbots significantly improved comprehension (RQ1), that higher GenAI literacy was associated with larger comprehension gains, especially for the conventional chatbot (RQ2), and that learners with different literacy levels exhibited different interaction patterns (RQ3). The authors conclude that scaffolding chatbots reduce reliance on learners' GenAI literacy and offer a more equitable design.","tokens_in":17684,"tokens_out":4183,"duration_ms":41295,"significance":"If the causal claims are correct, the study provides useful empirical evidence for the value of GenAI chatbots in learning analytics dashboards and for the importance of GenAI literacy as a learner characteristic. The strengths include an a priori power analysis, random assignment to the two chatbot conditions, a mixed-methods design combining quantitative comprehension gains with epistemic network analysis, and an open repository with materials and the coding scheme. The main limitation is that the headline comprehension gains rest entirely on a within-subject baseline-to-intervention contrast without a no-chatbot control condition or counterbalancing of question sets; this directly threatens the causal interpretation. The moderation claim about scaffolding reducing reliance on GenAI literacy is also not supported by a formal statistical test of the interaction between condition and literacy. These issues are fixable by re-analysis and by softening or re-scoping the causal claims, rather than by adding new data, so the work has a defensible core that merits revision.","major_comments":[{"comment":"The central RQ1 claim that the chatbots caused comprehension improvement is not supported by the design as described. Every participant first completed the baseline task and then the intervention task in a fixed order, with the same visualisation types and similar question formats, and there was no no-chatbot control condition. The observed Baseline_score to Intervention_score gains could therefore be driven by practice with the visualisations, familiarisation with the question format, or differences in difficulty between the two question sets. The authors' statement in Section 3.3 that participants had no prior knowledge of the learning context addresses prior content knowledge, not task practice. To support the causal wording in Section 5, the authors should add a no-chatbot control arm and/or counterbalance the two question sets across phases, or explicitly reframe the result as an observed improvement that is associated with chatbot use rather than caused by it.","section":"§3.3.3 and §4.2"},{"comment":"The conclusion that scaffolding chatbots reduce reliance on GenAI literacy is not formally tested. The authors compare separate OLS coefficients for GenAI_literacy in the conventional group (β = 0.19) and the scaffolding group (β = 0.08), but no test of the condition-by-literacy interaction or of the equality of the coefficients is reported. A more direct test would be a single regression on the Intervention_score with Baseline_score, GenAI_literacy, condition, and the GenAI_literacy × condition interaction. In addition, the current gain-score regression uses Improvement_score = Intervention_score − Baseline_score as the dependent variable while including Baseline_score as a regressor; this creates a mechanical negative association between baseline and improvement that is not fully addressed by the ceiling-effect interpretation. The moderation claim should be supported by an explicit interaction test, or the authors should temper the claim.","section":"§3.5.3 and §4.3"},{"comment":"The ENA significance tests appear to be inconsistent with the stated Bonferroni correction. Section 3.5.4 says Bonferroni correction was applied with an initial alpha of 0.05, but Section 4.4 reports a p-value of 0.04 for the conventional group comparison along the X-axis and calls it statistically significant with alpha 0.05. With multiple comparisons across axes and conditions, the corrected threshold would be substantially smaller than 0.05, so p = 0.04 would not be significant. The authors should report corrected p-values, state the exact number of comparisons, or explicitly justify why no correction is needed for the specific RQ3 contrasts.","section":"§3.5.4 and §4.4"},{"comment":"The within-subject comparisons in Section 4.2 are labelled as 'Mann-Whitney U test' but report a W statistic and compare paired baseline and intervention scores. These should be Wilcoxon signed-rank tests, as stated in Section 3.5.2. If the Mann-Whitney U test was actually misapplied to paired data, the analyses need to be redone; if it is simply a terminological error, the text and table labels should be corrected so that the reported effect sizes and p-values are unambiguously associated with the correct test.","section":"§4.2"}],"minor_comments":[{"comment":"The correction-for-guessing formula is mentioned for both the comprehension scores and the GenAI literacy score, but the paper does not report whether the descriptive statistics in Section 4.2 are corrected or raw scores. Since the correction can change the scale and possibly produce negative values, please clarify the formula and state whether medians, IQRs, and regression inputs use corrected or raw scores.","section":"§3.5"},{"comment":"There is a typo in the scaffolding condition paragraph: 'Chabot.Information' should be 'Chatbot.Information'.","section":"§4.4"},{"comment":"The paper says participants were randomly assigned to one of two intervention groups, but no details are given about the randomisation method or whether allocation was concealed. A sentence describing the allocation procedure would strengthen the report.","section":"§3.3.3"},{"comment":"The median-split categorisation of low and high GenAI literacy is sample-dependent and yields a small high-literacy subgroup in the conventional condition (n = 11). Please acknowledge this as a limitation and, if feasible, report sensitivity of the ENA results to alternative cut points.","section":"§3.5.4"},{"comment":"The phrase 'left–below in Section 4.4' is unclear; please refer to panels of the figure explicitly, e.g., 'left panel, described in Section 4.4'.","section":"Figure 3"}],"recommendation":"major_revision","confidential_remarks":"The manuscript's central claims rest on a within-subject contrast without a no-chatbot control or counterbalancing; this is fixable by re-analysis and by reframing the claims, but it is currently a load-bearing threat. The GLAT instrument is developed by the same research group and is cited as an arXiv preprint; this is worth watching in review, although comprehension is measured independently of the literacy test. The paper otherwise fits the scope of a learning-analytics or HCI venue and has useful materials and mixed-methods analysis."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: this is a solid empirical study of a real question, with one design flaw that keeps me from fully believing the headline causal claim. The new thing is the comparison of conventional vs scaffolding GenAI chatbots in a learning analytics dashboard, tied to a measured GenAI literacy instrument, with both comprehension scores and interaction traces. That is genuinely useful; I don't know another published comparison like it. The paper is also careful in places: a priori power analysis, random assignment, guessing correction, inter-rater reliability above 0.7, and the ENA coding makes sense.\n\nWhere it falls short: the core RQ1 claim that both chatbots improved comprehension is a within-subject baseline-to-intervention contrast with no no-chatbot control and no counterbalancing of question sets. Seeing the same visualisation types twice and answering similar questions is enough to produce part of the gain by practice. The authors explicitly say participants had no prior knowledge of the context, but that doesn't address task practice. This is the load-bearing concern; everything else is secondary. The effect sizes are large (r around 0.9), which makes it less likely that practice alone explains everything, but the design still can't rule it out.\n\nThe secondary issues are real but smaller. The ENA significant difference of p = 0.04 in Section 4.4 appears inconsistent with the stated Bonferroni correction; if they corrected for multiple comparisons, that p-value wouldn't survive. And although they conclude scaffolding reduces dependence on GenAI literacy, they never formally test whether the regression coefficients differ between the two conditions; the difference is only visible in two separate models. That's a claim the data supports only weakly.\n\nOn the self-reliance point: yes, GLAT and VizChat come from the same group, but the outcome measure is a separate comprehension test, so this is a matter of external validation rather than circular reasoning. Not a flaw in the study.\n\nBottom line: I'd send this to review, but I'd insist the no-control/counterbalancing problem be addressed head-on—either with a control arm, a discussion of why it's infeasible, or a more modest claim. For now, treat the comprehension gain as plausible but not established.","headline":"Useful empirical comparison of chatbot designs in a learning analytics dashboard, but the central causal claim is weakened by the lack of a no-chatbot control and non-counterbalanced question order.","tokens_in":18260,"tokens_out":2553,"would_cite":true,"duration_ms":24127,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper claims that adding either a conventional or a scaffolding generative-AI chatbot to a learning analytics dashboard significantly improves learners' comprehension of complex visualisations, and that a learner's generative-AI…","keywords":["learning analytics dashboard","generative AI literacy","generative AI chatbots","data visualisation","scaffolding","human-computer interaction","epistemic network analysis","multimodal learning analytics"],"falsifier":"Run the same two-phase design with a third condition in which participants get no chatbot at all (or only the written context again) while keeping identical visualisations and counterbalancing the question sets; if the no-chatbot condition shows the same baseline-to-intervention gain, then the improvement attributed to the chatbots is not theirs alone.","tokens_in":17350,"feed_emoji":"🤖","tokens_out":6804,"duration_ms":57291,"temperature":0.7,"pith_summary":"Learning analytics dashboards present complex data that learners often struggle to interpret. This paper asks whether adding a generative-AI chatbot—either one that simply answers questions or one that proactively guides users with scaffolding questions—makes the insights in those dashboards more comprehensible, and whether a learner's generative-AI literacy decides how much they benefit. In a study of 81 medical and nursing students, both chatbot types raised comprehension scores from baseline to intervention, with large effects, and higher generative-AI literacy predicted larger improvements. The effect of literacy was more than twice as strong with the conventional chatbot as with the scaffolding chatbot, suggesting that proactive guidance can partly compensate for low literacy. If correct, the result gives designers of educational dashboards a concrete reason to build scaffolding into AI assistants and to measure AI literacy when deploying them.","feed_headline":"Chatbots help learners read complex dashboards, study finds","feed_subtitle":"Scaffolding chatbots narrow the gap for low AI-literacy learners.","key_machinery":"The central mechanism is the contrast between the two chatbot designs built on the VizChat prototype: a conventional chatbot that reactively answers user queries and a scaffolding chatbot that proactively poses expert-designed guiding questions and gives step-by-step feedback on visualisation elements. Comprehension is measured as an improvement score, the difference between a six-question multiple-choice test taken before and after chatbot use, with questions based on Bloom's taxonomy levels 1 and 2 and a correction-for-guessing formula. Generative-AI literacy is measured with the 20-item Generative AI Literacy Assessment Test (GLAT). The statistical core is an OLS regression of improvement on baseline score and literacy, with a tested interaction term, run separately for each condition, and the cognitive analysis uses epistemic network analysis of coded utterances. The competing designs do the argumentative work: the literacy coefficient's difference between conditions is what supports the claim that scaffolding lowers the literacy requirement.","core_discovery":"On its own terms, the paper reports three empirical findings. First, within-subject comprehension scores on six multiple-choice questions improved significantly after interacting with either chatbot (conventional group median from 3 to 4, scaffolding group from 3 to 5), with large rank-biserial effect sizes (0.87 and 0.94). Second, regression analyses of improvement scores found a positive association with generative-AI literacy in both conditions, but the coefficient for literacy was $0.19$ points per literacy point in the conventional group versus $0.08$ in the scaffolding group, and only the conventional model needed a baseline-by-literacy interaction term. Third, epistemic network analysis of user-chatbot utterances showed that high-literacy learners engaged in more information integration and reflection with the conventional chatbot, while low-literacy learners relied on clarification and commands; with the scaffolding chatbot, low-literacy learners showed more exploratory, iterative interaction patterns. The authors interpret these results as evidence that chatbot-assisted dashboards support comprehension and that scaffolding can reduce the literacy gap.","pith_inferences":["An implication the authors leave untested is that a no-chatbot control group might show a similar improvement due to task practice alone; the fixed baseline-then-intervention order without counterbalancing makes this the key confound.","If the mechanism is that scaffolding lowers prompt-crafting demands, the same conventional-versus-scaffolding contrast should replicate in other data-rich domains such as financial or health literacy dashboards.","The measured literacy score could be used as an adaptive routing signal: assign low-literacy learners to scaffolding chatbots automatically, and reserve conventional chatbots for high-literacy users.","The epistemic network analysis differences also suggest a hybrid design: start all learners with scaffolding prompts, then fade them as the user's interaction patterns show integration and reflection."],"forward_implications":["If replication holds, adding either reactive or proactive chatbots to existing dashboards can be expected to raise comprehension of complex visualisations in similar populations, with no significant difference in mean improvement between the two designs.","The literacy effect implies that when a conventional, unguided chatbot is deployed, learners with higher generative-AI literacy will benefit considerably more, so interface designers should not assume equal benefit without measuring literacy.","Scaffolding chatbots' smaller literacy coefficient suggests proactive guidance can serve as an equity mechanism, letting lower-literacy learners approach the gains of higher-literacy peers.","The epistemic network analysis patterns imply that low-literacy learners' interactions focus on clarification and commands, so logging such patterns could help a dashboard detect users who need more supportive prompting."],"supporting_citations":[{"why":"Provides the 20-item Generative AI Literacy Assessment Test used to measure the independent variable.","marker":"[26]"},{"why":"The VizChat prototype from which both conventional and scaffolding chatbots were developed.","marker":"[71]"},{"why":"Defines scaffolding as decomposing complex content and posing guiding questions, the basis for the scaffolding design.","marker":"[21]"},{"why":"Supplies empirical evidence that computer-based scaffolding improves problem-based learning, motivating the scaffolding condition.","marker":"[32]"},{"why":"Introduces explanatory dashboards with narratives and visual markers, the dashboard context the chatbots augment.","marker":"[15]"},{"why":"Supplies the correction-for-guessing formula applied to both comprehension and literacy scores.","marker":"[60]"},{"why":"Frames generative AI in learning analytics and motivates the promise of conversational dashboards.","marker":"[68]"}],"fun_headline_variants":["Scaffolding chatbots close AI literacy gap in dashboards","GenAI literacy shapes how students use dashboard chatbots","Chatbots boost dashboard learning, but literacy matters most","Proactive chatbots level playing field for low AI-literacy users","Chatbot dashboards help, yet GenAI literacy drives gains"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The result rests on the assumption that the comprehension gains measured after chatbot interaction are caused by the chatbot itself; because every participant did the baseline before the intervention with no counterbalancing and no no-chatbot control, practice effects, growing familiarity with the visualisations, or easier question sets could explain the improvement instead.","fun_headline_variants_meta":{"raw":{"variants":["Scaffolding chatbots close AI literacy gap in dashboards","GenAI literacy shapes how students use dashboard chatbots","Chatbots boost dashboard learning, but literacy matters most","Proactive chatbots level playing field for low AI-literacy users","Chatbot dashboards help, yet GenAI literacy drives gains"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000691,"raw_usage":{"total_tokens":3153,"prompt_tokens":993,"completion_tokens":2160,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":609,"completion_tokens_details":{"reasoning_tokens":2080}},"tokens_in":609,"tokens_out":2160,"duration_ms":14742,"temperature":1.0,"reasoning_tokens":2080,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T14:06:47.291033+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same two-phase design with a third condition in which participants get no chatbot at all (or only the written context again) while keeping identical visualisations and counterbalancing the question sets; if the no-chatbot condition shows the same baseline-to-intervention gain, then the improvement attributed to the chatbots is not theirs alone.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"The VizChat prototype from which both conventional and scaffolding chatbots were developed."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Defines scaffolding as decomposing complex content and posing guiding questions, the basis for the scaffolding design."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies empirical evidence that computer-based scaffolding improves problem-based learning, motivating the scaffolding condition."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Introduces explanatory dashboards with narratives and visual markers, the dashboard context the chatbots augment."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the correction-for-guessing formula applied to both comprehension and literacy scores."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Frames generative AI in learning analytics and motivates the promise of conversational dashboards."}],"review_version":1}