{"id":"12de8abf-83d7-46f1-9d97-c2ce958cedf9","arxiv_id":"2506.06225","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A pilot study of a GenAI literacy chatbot for older adults found positive qualitative feedback and a non-significant trend in self-reported AI literacy.","lead":"The paper describes Litti, a chatbot that teaches generative AI literacy to older adults, and reports results from a 12-person pilot study at a senior living center. The quantitative gains were not statistically significant, but interviews showed older adults had varied AI familiarity and a strong desire to learn, pointing to design needs for AI literacy education.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 0.55-point gain is not interpretable as AI literacy because Litti's tasks were built from the same MAILS sub-scales used in the survey; the reported SD also makes the 'not statistically significant' claim internally questionable.","rationale":"The reader's weakest_assumption correctly identifies the instrument-validity and teaching-to-the-test risk as the load-bearing point. I agree that the quantitative conclusion rests entirely on the adapted MAILS measure, which is not validated for this population and is deliberately aligned with the intervention's tasks. My stress-test adds a separate statistical reporting inconsistency: the reported mean and SD would actually yield a significant paired t-test, so the abstract's 'not statistically significant' claim is internally questionable. These concerns do not change the overall CONDITIONAL verdict, because the qualitative findings—positive engagement, useful design feedback, and rich quotes—are plausible and independently reported. The paper should be revised to report the actual significance test and to separate trained-item gains from transfer-item gains before the quantitative trend is used to support any recommendation.","tokens_in":10854,"tokens_out":5938,"duration_ms":60815,"concrete_test":"Reconstruct from Figure 2/Appendix the item-level pre/post responses, separate the three MAILS sub-scales rehearsed in Litti's tasks from untrained MAILS sub-scales (e.g., Create AI, AI Problem Solving), and compute gains on trained vs. untrained items plus a paired t-test or Wilcoxon on the overall gain. If gains concentrate in trained items or if the overall p-value is <0.05, the 'positive trend' claim must be reframed as a practice effect and the 'not statistically significant' wording corrected.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 3 states that Litti's tasks were 'designed around' the MAILS sub-scales 'Know & AI,' 'Detect AI,' and 'AI Ethics'; Section 4.3 states that the pre/post questionnaire was 'developed adopting' the same MAILS framework. The outcome measure is therefore not independent of the intervention content. With no control condition, no psychometric validation of the adapted instrument for adults 75+, and no behavioral transfer measure, the reported average increase of 0.55 points (SD=0.72) in Section 5.2 is equally consistent with practice effects, item recall, and social desirability as with genuine AI literacy learning. The internal statistical reporting adds a second problem: with n=12, a paired t-test on these values gives t≈2.65 (p≈0.023), contradicting the abstract's claim that results were 'not statistically significant'; Section 5.2 only mentions non-significant correlations, not a test of the mean gain. Thus the central quantitative trend—the only quantitative support for the contribution—is uninterpretable as reported.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper reports a mixed-methods study of Litti, a GenAI literacy chatbot, with 12 older adults (mean age 81.1) in a single one-hour session. Participants completed pre/post surveys adapted from the MAILS framework, interacted with the chatbot through five to six tasks, and joined focus group interviews. The paper claims a positive qualitative learning experience and a non-significant positive trend in AI literacy (an average increase of 0.55 points on a 5-point scale, SD=0.72), while also reporting that the intervention did not significantly change trust or safety perceptions. The contribution is framed as an exploratory case study with design recommendations for AI literacy education for older adults.","tokens_in":11027,"tokens_out":5258,"duration_ms":51561,"significance":"The qualitative portion of the study is a useful exploratory contribution: it documents varied prior familiarity with GenAI among adults over 75, highlights trust and safety as persistent concerns, and offers concrete, participant-grounded design suggestions such as larger fonts, multimodal input, and familiar task scenarios. The detailed description of the Litti intervention (built on Claude and FlowXO) supports replication, and the explicit limitations paragraph is a strength. However, the quantitative evidence is not load-bearing as reported. The abstract's 'not statistically significant' statement is internally contradicted by the reported mean and standard deviation, the outcome measure is aligned with the intervention content (a teaching-to-the-test risk), and the adapted MAILS instrument is not validated for this population. The paper's significance therefore rests primarily on the qualitative findings and on the suggestion of a viable chatbot format for future, better-controlled studies, rather than on the claimed quantitative trend.","major_comments":[{"comment":"The printed quantitative results are internally inconsistent. With n=12 and a reported average increase of 0.55 points (SD=0.72) for the pre/post difference, a paired t-test yields t ≈ 2.65 with df=11 and a two-tailed p ≈ 0.023, which is significant at the conventional .05 level. The abstract and Section 5.2 state that the results were 'not statistically significant' and frame the change as a 'positive trend,' but they neither report this test nor explain which test was actually used. The authors must reanalyze and re-report these numbers, or provide a clear justification for a different test that yields non-significance; as it stands, the conclusion and the data contradict each other.","section":"Abstract and Section 5.2"},{"comment":"The outcome measure is not independent of the intervention content. Section 3 states that Litti's tasks were 'designed around' the MAILS sub-scales 'Know & AI,' 'Detect AI,' and 'AI Ethics,' while Section 4.3 states that the pre/post questionnaire was 'developed adopting' the same MAILS framework. The measured gains are therefore plausibly attributable to practice effects, item familiarity, or teaching-to-the-test rather than to generalizable AI literacy learning. The paper should explicitly discuss this alignment as a threat to construct validity; the current text does not acknowledge this circularity. A transfer measure, a control condition, or a delayed post-test would partly mitigate the concern, but none is present.","section":"Section 3 and Section 4.3"},{"comment":"The claim that the 0.55-point overall increase is 'meaningful' is unsupported by any inferential statistics. The section reports only means and standard deviations for the overall score and the three sub-scales, with no test statistics, confidence intervals, or effect sizes; the only explicitly reported statistical results are the non-significant correlations. To support any quantitative claim, the authors should report the paired t-test (or an appropriate nonparametric alternative) result, the associated effect size with a confidence interval, and an interpretation of that magnitude relative to the instrument's scale and measurement properties.","section":"Section 5.2"},{"comment":"The sample and instrument limitations are more severe than the Discussion acknowledges. With n=12, a single one-hour session, and 10 of 12 participants holding graduate degrees, the quantitative results have extremely limited generalizability. In addition, the MAILS framework was originally developed and validated on other populations, and no psychometric evidence is provided for its use with adults aged 75 and over; without measurement invariance, test-retest reliability, or known floor/ceiling behavior, the pre/post change scores may partly reflect measurement artifact. These issues should be moved from 'future work' into the primary limitations of the quantitative analysis.","section":"Section 4.1 and Section 4.3"}],"minor_comments":[{"comment":"The number of tasks is inconsistent: Section 3 says 'four tasks' but then describes five (general search, personal assistance, emotional companion, AI safety, AI ethics), while Section 4.2 enumerates six steps. The authors should align these counts.","section":"Section 3 and Section 4.2"},{"comment":"The sentence beginning 'As enAI tools become...' contains a typo; it should read 'As GenAI tools...'.","section":"Section 6"},{"comment":"The chatbot is referred to as 'Littie' in one participant quote; this should be 'Litti'.","section":"Section 5.1.3"},{"comment":"The citation 'Shandilya and Fan [?]' has a placeholder question mark instead of a reference number; it should be numbered consistently with the bibliography.","section":"Section 6"},{"comment":"Reference '[48]' is cited in the sentence about age-related cognitive and visual impairments, but the reference list only goes up to [24]; the intended citation is likely one of the existing AI-safety references, such as [16] or [24].","section":"Section 2.2"},{"comment":"Figure 2's caption is identical to Figure 1's caption ('AI literacy increase by individual'), which could confuse readers; it should be made more descriptive or referenced explicitly in the text.","section":"Appendix C"}],"recommendation":"major_revision","confidential_remarks":"The internal statistical contradiction—a mean/SD combination that yields p≈0.023 while the abstract claims non-significance—is the most pressing issue and must be fixed before this can be considered for publication. The circularity between MAILS-based intervention tasks and MAILS-based outcome measurement is a deeper validity concern that also needs explicit acknowledgment. Given the exploratory scope and the venue, I do not see this as a reject, but the quantitative claims must be reanalyzed and recalibrated. The manuscript also appears rushed, with inconsistent task counts, a placeholder citation, and a stray reference number."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick read: this extended abstract does something nobody else has done yet—deploy a GenAI literacy chatbot with adults 75+ and record what they actually said. The qualitative material is the paper's real contribution. The participants' own words about voice-clone scams, skepticism, and the desire to learn are vivid and useful for designers. The authors also deserve credit for honestly reporting that the intervention did not move trust or safety perceptions, and for spelling out limitations many authors would have buried.\n\nWhat is new: a chatbot named Litti built on Claude/FlowXO, with tasks tailored to older adults (search, personal assistance, emotional support, safety/ethics), a pre/post MAILS questionnaire, and focus groups. The qualitative analysis surfaces a wide range of prior familiarity—from self-described \"total newbie\" to a participant who already teaches a course with AI—and that range is itself a finding.\n\nThe soft spots are real. The pre/post measure was the MAILS sub-scales, and the tasks were designed around the same MAILS sub-scales. That makes the 0.55-point gain just as consistent with practice or item-recall as with learning. There is no control group, no psychometric validation for adults 75+, and no behavioral transfer measure. The statistics are also internally inconsistent: with n=12, mean gain 0.55, SD 0.72, a paired t-test gives t≈2.65, p≈0.023—significant at the conventional level—yet the abstract and Section 5.2 call it non-significant. Either the test was different (e.g., Wilcoxon) or the reported numbers are wrong; either way, the section as written can't be trusted. There is also a placeholder citation in the Discussion and no data or transcripts shared.\n\nTo be clear about proportion: none of this kills the qualitative value. The focus group material is a legitimate exploratory result, and the design insights (font size, scrolling, latency, voice-only accessibility) are actionable. But the quantitative trend should be reported as a pilot observation, with the circularity and missing significance test acknowledged, not as \"a meaningful improvement.\"\n\nWho should read it: HCI folks working on AI literacy, older adults, or chatbot-based instruction. It's a useful pointer for future work, not a result to build on statistically.\n\nFor review: I'd send it to a serious referee—the topic is timely and the qualitative core is salvageable—but the authors need to fix the statistical reporting and reframe the quantitative claims before publication. If those can't be fixed, strip the quantitative section and present it as a qualitative study. Engage with it, but only for the qualitative findings; don't cite the numbers.","headline":"Useful qualitative pilot on GenAI literacy chatbots for older adults, but the quantitative trend is circular and internally contradictory.","tokens_in":11566,"tokens_out":4663,"would_cite":true,"duration_ms":43530,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Litti, a chatbot built to teach generative AI to older adults, produces a positive learning experience and a positive but statistically non-significant shift in self-reported AI literacy; trust and safety perceptions stay flat.","keywords":["generative AI literacy","older adults","educational chatbot","AI safety and ethics","trust in AI","mixed methods study","assisted living"],"falsifier":"Run a larger pre-registered study with a no-intervention control group and a non-interactive video condition using the same adapted questionnaire; if the chatbot condition does not outperform the video on post-test gain, or if the no-intervention group shows the same gain, the claim that the chatbot drives literacy improvement is falsified. A simpler check is to administer the questionnaire twice a week apart with no intervention; a large practice effect would undermine the measured trend.","tokens_in":10630,"feed_emoji":"🤖","tokens_out":8067,"duration_ms":76698,"temperature":0.7,"pith_summary":"This paper asks whether a purpose-built generative AI chatbot can teach adults aged 75 and older what generative AI is, how to use it, and how to recognize its risks. Twelve residents of a senior assisted-living center spent about an hour with Litti, a chatbot whose guided tasks cover general search, personal assistance, health questions, emotional support, safety, and ethics. Pre- and post-surveys built from an existing AI-literacy questionnaire showed a positive but not statistically significant trend: average self-reported literacy rose 0.55 points on a 5-point scale. Interviews showed participants enjoyed the experience and wanted to learn more, yet trust and safety concerns remained largely unchanged. If the pattern holds, chat-based instruction is a workable entry point for GenAI literacy, but it is not, by itself, a trust-building or safety intervention.","feed_headline":"One-hour chatbot session nudges AI literacy, not trust, in adults 75+","feed_subtitle":"Self-reported AI literacy rose 0.55 points on a 5-point scale, but safety fears did not move.","key_machinery":"The machinery is Litti, a chatbot whose large-language-model responses are shaped by a prompt that fixes its persona as an empathetic AI literacy instructor and sequences six tasks: explaining generative AI, general search, personal assistance, health information seeking, emotional support, and safety and ethics. Around this sits an adapted pre/post questionnaire drawn from a published multi-subscale AI-literacy instrument; the authors prioritized the sub-scales for knowing about AI, detecting AI, and AI ethics, and built the chatbot's tasks from those same sub-scales. The prompt and the aligned questionnaire are what carry the argument: every quantitative claim is a comparison of the same instrument before and after one hour with the bot.","core_discovery":"The paper's central claim is that an interactive chatbot is a viable format for introducing generative AI to older adults, and that a single guided session can move self-reported AI literacy in the positive direction. On the paper's own evidence, the claim is supported only as a trend: 12 participants aged 75 to 90, with a mean age of 81.1, showed an average increase of 0.55 points (SD 0.72) on a 5-point AI-literacy scale after interacting with Litti, with sub-scale increases of 0.72 for knowledge, 0.47 for ethics, and 0.37 for AI detection, none reaching statistical significance. Qualitative focus-group data show broad enthusiasm for the learning experience and a strong desire for more education, alongside persistent skepticism about privacy, scams, and the trustworthiness of AI outputs. The paper therefore establishes, provisionally, that this demographic will engage with chatbot-facilitated GenAI education; it does not establish that such education measurably improves literacy, trust, or safety.","pith_inferences":["The observed effect size, mean gain 0.55 points with SD 0.72, is large enough that a replication with a few dozen participants might reach statistical significance; a pre-registered power calculation based on these numbers would settle whether the trend is worth pursuing.","The inverse relation between pre-test score and gain, if it survives replication, suggests chatbot literacy education functions mainly as a leveling tool for novices; the paper only reports this pattern descriptively.","A natural extension is to test Litti against a non-interactive version with identical content; if gains match, the chatbot's interactivity is not the active ingredient and simpler static tutorials would suffice.","Trust may be harder to move than knowledge because the bot's own polished, human-like output can sharpen concerns about deception; one participant's reaction in the focus group illustrates this, and a future measure of trust after safety-specific modules could test it."],"forward_implications":["If the positive trend is real, a single one-hour chat session is enough to nudge self-reported AI literacy upward in this age group, with the largest gains appearing in participants who started with the lowest scores.","Chatbot-based instruction can be run in an assisted-living setting with laptops and about one hour of supervision, making it a practical low-cost format to deploy more widely.","Because trust and safety perceptions did not improve, AI literacy curricula for older adults need separate, content-rich safety modules and explicit trust-building activities rather than relying on hands-on use alone.","A formal between-subjects trial comparing Litti against a website or video covering the same material would be the natural next test of whether interactivity, rather than content exposure, drives the effect."],"supporting_citations":[{"why":"supplies the AI-literacy questionnaire whose sub-scales structure both Litti's tasks and the pre/post survey.","marker":"[3]"},{"why":"gives the definition of AI literacy that frames what the intervention is trying to teach.","marker":"[13]"},{"why":"documents existing AI literacy levels among older adults and marks the gap this study addresses.","marker":"[11]"},{"why":"identifies perceived usefulness, technical literacy, and apprehension as factors shaping older adults' technology adoption, motivating the chatbot's design.","marker":"[21]"},{"why":"establishes the scam and security threat landscape that motivates the AI safety component of the curriculum.","marker":"[16]"},{"why":"contributes evidence that awareness and education can mitigate AI-related vulnerabilities, the premise for expecting literacy gains to help.","marker":"[24]"},{"why":"shows prior digital-literacy measurement work focused on assessment rather than intervention, framing the study's contribution.","marker":"[18]"}],"fun_headline_variants":["Chatbot nudges AI literacy in elders, but trust lags","Older adults learn AI via chatbot, trust unchanged","AI chatbot boosts literacy trend, not trust, in seniors","Seniors' AI literacy trend up after chatbot, trust flat"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the adapted AI-literacy questionnaire measures real change in adults aged 75 and older, so that the 0.55-point gain reflects learning rather than practice on questions similar to the tasks.","fun_headline_variants_meta":{"raw":{"variants":["Chatbot nudges AI literacy in elders, but trust lags","Older adults learn AI via chatbot, trust unchanged","AI chatbot boosts literacy trend, not trust, in seniors","Seniors' AI literacy trend up after chatbot, trust flat"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000152,"raw_usage":{"total_tokens":1195,"prompt_tokens":929,"completion_tokens":266,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":545,"completion_tokens_details":{"reasoning_tokens":197}},"tokens_in":545,"tokens_out":266,"duration_ms":3085,"temperature":1.0,"reasoning_tokens":197,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T05:57:44.471428+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run a larger pre-registered study with a no-intervention control group and a non-interactive video condition using the same adapted questionnaire; if the chatbot condition does not outperform the video on post-test gain, or if the no-intervention group shows the same gain, the claim that the chatbot drives literacy improvement is falsified. A simpler check is to administer the questionnaire twice a week apart with no intervention; a large practice effect would undermine the measured trend.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"documents existing AI literacy levels among older adults and marks the gap this study addresses."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"identifies perceived usefulness, technical literacy, and apprehension as factors shaping older adults' technology adoption, motivating the chatbot's design."},{"cited_title":"Kim, Minsu Kim, Jaeuk Oh, Sang Hui Chu, and JiYeon Choi","cited_arxiv_id":null,"evidence_quote":"shows prior digital-literacy measurement work focused on assessment rather than intervention, framing the study's contribution."}],"review_version":1}