{"id":"e9543342-4cca-408f-8904-493abd75d710","arxiv_id":"2412.14200","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"ActiveAI provides a large K-12 AI literacy dataset and preliminary evidence of learning gains and gender differences in assessments.","lead":"This paper presents ActiveAI, an online learning platform that delivered AI literacy modules and assessments to over 1,000 K-12 students across 12 schools. The authors report preliminary learning gains and gender differences and plan to release the de-identified dataset for broader research.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The learning-gain claim depends on an unverified equivalence of pre- and post-tests: the paper says they are isomorphic but gives no reliability, equating, or practice-effect control, so the 4-of-7 significant gains may be artifacts.","rationale":"The paper's primary contribution is a dataset, and on that front the reporting is reasonably transparent: the authors state the number of users, the number of completers, the modules, and the statistical tests used. I credit that transparency and the reference to prior design work. The load-bearing assumption is the isomorphism of the pre- and post-tests. The authors assert equivalence in Section 2.1 but provide no psychometric evidence for it. If the post-test is easier or merely more familiar, the Wilcoxon gains are not learning. This is especially important because there is no control group and because learners may improve merely from seeing the format twice. The reader's weakest_assumption identifies exactly this point, so I agree. A second-order issue is the lack of effect sizes and multiple-comparison control in the '4 of 7' claim; this is a reporting gap, not necessarily a fatal flaw. The gender finding is suggestive but also underreported. None of this changes the conditional verdict: the dataset claim can stand if the data are actually released; the learning-gain claim needs either additional statistical evidence or softer wording. Since the reader already issued CONDITIONAL and my concern aligns with the weakest assumption, I recommend UNCHANGED.","tokens_in":2878,"tokens_out":4055,"duration_ms":38896,"concrete_test":"Using the released de-identified item-level data, fit a mixed-effects model of pre/post correctness with fixed effects for time, item ID or item difficulty, and form order, and random intercepts for student and school. If including item-difficulty differences between the two forms removes the significant time effect, then the 'isomorphic' tests are not equated and the learning-gain claim is unsupported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 2.1 states that the pre- and post-tests are 'isomorphic assessments targeting the same learning objectives,' and Section 3 reports significant Wilcoxon gains on 4 of 7 objectives. The central empirical claim—that the platform produces measurable AI-literacy learning—requires that pre-to-post score changes reflect growth in the construct rather than differences in form difficulty, item wording, or test-taking familiarity. The paper provides no internal-consistency reliability, no equivalent-forms correlation, no item-level difficulty or equating analysis, and no control or delayed-post condition to estimate practice effects. Because learners see the same learning objectives twice and no evidence rules out a form-order effect, the reported gains could be inflated by exposure to the test itself. The report of '4 of 7' significant objectives also lacks multiple-comparison correction or effect sizes, so the number of true gains is uncertain. The secondary gender result (module 4, f=6.97 pre and f=6.80 post) is likewise reported without means, SDs, or effect sizes, and the pre-test difference makes the post-test comparison hard to interpret causally. These gaps do not invalidate the dataset contribution, but they do mean the preliminary effectiveness results are not yet established.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents ActiveAI, an online learning platform for K-12 AI literacy education, and reports on a dataset collected from over 1,000 users across 12 schools, with 426 learners completing all components across four modules. The authors report preliminary findings: significant learning gains on 4 of 7 learning objectives based on Wilcoxon tests, gender differences in Module 4 assessment scores, and an association between learning gains and cognitive engagement levels under the ICAP framework. The paper's main contributions are the dataset itself, which the authors plan to make openly available, the platform's standardized data logging, and the AI literacy learning activities.","tokens_in":3153,"tokens_out":2539,"duration_ms":22347,"significance":"If the dataset is released as described, it would be a valuable and scarce resource for AI literacy education research, enabling secondary analyses in a rapidly growing but data-poor field. The platform's logging standards and planned alignment with existing educational data repositories (DataShop, LearnSphere) are concrete strengths, as is the use of a backward-design curriculum framework aligned with AI4K12 standards. The paper also makes an empirical contribution by attempting to link learning gains to engagement via the ICAP framework. However, the statistical evidence presented for the preliminary effectiveness claims is incomplete, and the dataset availability is only promised rather than demonstrated with a link or repository identifier.","major_comments":[{"comment":"The reported Wilcoxon tests on 7 learning objectives lack effect sizes, confidence intervals, and multiple-comparison correction. With 7 tests, the claim of 'significant learning gains in 4 of them' is not yet robust; please report adjusted p-values (e.g., FDR or Bonferroni) and an effect-size measure such as the rank-biserial correlation for each objective.","section":"Sec. 3 (learning gains)"},{"comment":"The paper asserts that pre- and post-tests are 'isomorphic assessments targeting the same learning objectives' but provides no reliability evidence (e.g., internal consistency per form), no equivalent-forms correlation, and no item-level difficulty or equating analysis. Without such evidence, the reported learning gains could reflect differences in form difficulty, item wording, or practice effects from repeated exposure. Please provide these analyses or explicitly temper the learning-gain claim as preliminary and not fully controlled.","section":"Sec. 2.1 and Sec. 3 (pre/post equivalence)"},{"comment":"For the Module 4 gender comparison, the paper reports only F-statistics and p-values. This is insufficient: please report means, standard deviations, and sample sizes per gender, along with an effect size (e.g., eta-squared). Because the pre-test already shows a significant gender difference, also consider an ANCOVA with pre-test score as a covariate to assess whether the post-test difference reflects differential learning rather than prior differences.","section":"Sec. 3 (gender analysis)"},{"comment":"The analysis is based on 426 of over 1,000 users, but the paper does not compare completers with non-completers on demographics, pre-test scores, or other available variables. An attrition analysis is needed to assess potential bias. Additionally, the statement that 'learning gains correlated with cognitive engagement levels (ICAP framework)' is made without reporting any correlation coefficient or test statistic; either provide the quantitative result (e.g., Spearman's rho between ICAP level and gain) or remove the claim.","section":"Sec. 3 (attrition and ICAP claim)"}],"minor_comments":[{"comment":"The paper cites 'Tseng et al., 2024' twice with different co-author lists (one as a 2024 SIGCSE paper and one as an EC-TEL paper). Please ensure these are distinct references and are formatted consistently in the bibliography.","section":"References"},{"comment":"Figure 1 is referenced in the text but has no caption or detailed explanation of its panes; add a caption describing the system design, data flow, and example interfaces so readers can interpret the figure without guessing.","section":"Figure 1"},{"comment":"The paper promises open access to the de-identified dataset but does not provide a repository URL, dataset name, or expected release timeline. Even in a short paper, a data availability statement is essential for a dataset contribution; please add one in the final version.","section":"Sec. 4 (data availability)"},{"comment":"The text contains numerous spacing errors (e.g., 'engagementinK-12AILiteracyeducationhassurged' and 'assessmentdatafromover1,000users'). A careful proofreading pass is needed to restore spaces and correct formatting.","section":"Throughout"}],"recommendation":"major_revision","confidential_remarks":"This is a short companion paper whose primary value is the promised dataset. The statistical gaps in the preliminary results are fixable but require additional analyses, which justifies a major revision rather than a simple minor one. I would recommend considering the paper's framing as a system/dataset description rather than a definitive effectiveness study; the learning-gain and gender claims should be presented with appropriate caution until the requested robustness checks are provided. The authors should also be encouraged to include a sample of the actual data or an access mechanism in the revision, as that is the core contribution."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The kernel of this paper is genuinely useful: a K-12 AI literacy platform deployed at 12 schools with over 1,000 users, logging standardized data for DataShop/LearnSphere, aimed at a domain where public datasets are scarce. If the de-identified data actually ships, that is a real contribution for secondary analysis. I also credit the authors for building on their prior ActiveAI work rather than starting from scratch, and for situating the results against the ICAP framework. The paper is honest about being preliminary.\n\nThe soft spots are in the empirical claims. The central learning-gain result rests on pre- and post-tests being “isomorphic,” but the paper gives no reliability, equating, or practice-effect evidence. That is load-bearing: if the two forms differ in difficulty or if test-taking familiarity inflates the second score, the 4-of-7 significant Wilcoxon gains could be artifacts. The statistical reporting is also thin—no effect sizes, no confidence intervals, no multiple-comparison correction, no attrition analysis. Only 426 of 1,000+ learners completed all components, and we don't know how those 426 differ from the rest. The gender difference in module 4 is reported with an ANOVA pre-test difference already present, so the post-test difference cannot be read as differential learning without controlling for baseline. These are not fatal for a dataset paper, but they are exactly the places a skeptical reviewer will push.\n\nI disagree with the stress-test note only in tone: it says the learning-gain claim “may be artifacts,” and I'd soften that to “are not yet established.” The data may well show real learning, but the measurement evidence is missing. The authors should provide equivalent-forms reliability, an equating analysis, or at least an item-by-item difficulty comparison, plus effect sizes and a completers-versus-non-completers check.\n\nWho should read this? Learning analytics researchers in K-12 AI literacy and anyone looking for an open dataset in that space. For a companion/workshop track, it deserves a serious referee, not a desk reject. For a full archival paper, it needs the missing statistical work and a released dataset URL. My recommendation: accept it as a work-in-progress dataset resource, with the clear expectation that effectiveness claims be downgraded or properly supported.","headline":"A valuable dataset contribution wrapped in a thin preliminary analysis; the learning-gain claims need stronger measurement evidence before they can be taken at face value.","tokens_in":3630,"tokens_out":1570,"would_cite":true,"duration_ms":15861,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper claims that a K-12 AI literacy platform can gather data from over 1,000 students and that preliminary analysis shows significant learning gains in four of seven objectives, with an open dataset for secondary research.","keywords":["AI literacy","K-12 education","learning analytics","online learning platform","pre-post assessment","ICAP framework","gender differences","educational dataset"],"falsifier":"A control design would settle it: give students the pre-test, wait through the same class period without the learning module, administer the post-test, and check whether scores rise as much as the reported gains did. If they do, the gains are an artifact of testing; the paper does not report such a control or reliability/equating statistics, so the claim is checked by a reader doing that comparison or examining item-level difficulty in the open dataset.","tokens_in":2753,"feed_emoji":"🤖","tokens_out":7592,"duration_ms":62017,"temperature":0.7,"pith_summary":"The paper's claim is that a web-based platform can deliver K-12 AI literacy instruction at a scale that has been missing, and that the data it collects can underpin replicable research. It reports deployment in 12 secondary schools with over 1,000 users, of whom 426 completed all surveys, pre-tests, modules, and post-tests. Preliminary analyses show significant pre-to-post learning gains in four of seven objectives, with gains concentrated in interactive activities, and a gender difference on one module. The authors intend the de-identified dataset to be openly available for secondary analysis, which is what makes the claim matter to the learning-analytics community if it holds.","feed_headline":"1,000 students log AI literacy gains in 4 of 7 objectives","feed_subtitle":"Open dataset lets researchers test why interactive activities outperform passive reading in K-12 AI lessons.","key_machinery":"The machinery is the instructional pipeline: students take a demographic survey, then an isomorphic pre-test, then a learning module in which they interact with an AI agent in simulated real-life scenarios (for example, spotting hallucinations in news summaries), then an isomorphic post-test. All interactions are logged. Backward design—starting from learning objectives aligned with the AI4K12 Big Ideas and building assessments and activities from them—ties each module's content to its measured outcome, while the ICAP framework supplies the interpretive link that converts observed score gains into claims about cognitive engagement. The logs are the part that enables scale: standardized records can be aggregated by student, class, activity, and objective.","core_discovery":"The central discovery, as the authors state it, is that a learning platform built on backward-designed AI literacy modules can collect standardized longitudinal learning data at K-12 scale. Using isomorphic pre- and post-tests targeting the same learning objectives, the authors found statistically significant Wilcoxon gains on four of seven objectives and observed that these gains tracked the ICAP engagement level of the activities: interactive scenarios with an AI agent produced significant gains, while passive reading produced smaller ones. In the hallucination-identification module, non-male students scored higher than male students on both the pre-test (F=6.97, p<0.01) and the post-test (F=6.80, p=0.01). The dataset, including demographic surveys, activity interactions, and outcomes, is positioned as the paper's main contribution: a new AI literacy corpus for secondary analysis.","pith_inferences":["Beyond the paper: because no test-reliability or equating evidence is reported, the learning gains may partly reflect practice effects; a secondary analyst could test this using item-level data from the open dataset.","Beyond the paper: the gender gap on the hallucination module may reflect differential prior exposure to generative AI rather than module quality; survey covariates could disentangle these.","Beyond the paper: if the ICAP-aligned pattern is causal, then scaling AI literacy should prioritize scenario-based interactive activities over passive reading, but the current evidence cannot rule out time-on-task or motivation confounds.","Beyond the paper: the planned comparison of standalone IT courses in Asian schools versus integrated STEM-club instruction elsewhere could reveal whether instructional context moderates gains, once the longer-duration data arrive."],"forward_implications":["Four of seven objectives show significant Wilcoxon pre-to-post gains, so the platform's interactive activities can be credited with measurable learning on those objectives.","Because significant gains cluster in interactive, AI-agent activities and not in passive reading, the results support the ICAP prediction that cognitive engagement drives AI literacy learning.","The module 4 gender difference—non-male students outperforming male students on both tests—identifies a design target for interventions that make AI literacy assessment fairer.","With 1,000 users and 426 complete records, the open de-identified dataset becomes one of the first large resources for modeling AI literacy prior knowledge, interaction patterns, and outcomes.","Standardized logging compatible with common educational data repositories lets outside researchers run secondary analyses without new data collection."],"supporting_citations":[{"why":"Systematic review documenting the shortage of large-scale AI literacy datasets and assessment efforts, which frames the paper's contribution.","marker":"Almatrafi et al., 2024"},{"why":"Defines the ICAP cognitive-engagement framework used to explain why interactive activities produced significant learning gains.","marker":"Chi & Wylie, 2014"},{"why":"Articulates the AI4K12 Big Ideas that the modules' learning objectives and assessments are aligned to.","marker":"Touretzky et al., 2023"},{"why":"Reports the earlier ActiveAI tutoring system that this platform extends and whose empirical examples underpin the instructional design.","marker":"Tseng et al., 2024"},{"why":"Supplies the backward-design method used to create objectives, assessments, and activities.","marker":"Wiggins & McTighe, 1998"}],"fun_headline_variants":["AI literacy dataset from 1,000 students shows gains in 4 of 7 goals","1,000 K-12 students: gender gap seen in AI hallucination detection","Backward-designed AI modules log significant gains in 4 of 7 objectives","Open dataset from 12 schools enables AI literacy secondary analysis","AI literacy: 1,000 users, 12 schools, and a new public dataset"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The learning-gain result stands on the assumption that the pre-test and post-test are equivalent in difficulty and measure the same objectives, so that a rise in scores means learning rather than practice, familiarity, or a harder first test.","fun_headline_variants_meta":{"raw":{"variants":["AI literacy dataset from 1,000 students shows gains in 4 of 7 goals","1,000 K-12 students: gender gap seen in AI hallucination detection","Backward-designed AI modules log significant gains in 4 of 7 objectives","Open dataset from 12 schools enables AI literacy secondary analysis","AI literacy: 1,000 users, 12 schools, and a new public dataset"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000243,"raw_usage":{"total_tokens":1466,"prompt_tokens":817,"completion_tokens":649,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":433,"completion_tokens_details":{"reasoning_tokens":545}},"tokens_in":433,"tokens_out":649,"duration_ms":5268,"temperature":1.0,"reasoning_tokens":545,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T15:14:46.167836+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A control design would settle it: give students the pre-test, wait through the same class period without the learning module, administer the post-test, and check whether scores rise as much as the reported gains did. If they do, the gains are an artifact of testing; the paper does not report such a control or reliability/equating statistics, so the claim is checked by a reader doing that comparison or examining item-level difficulty in the open dataset.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Systematic review documenting the shortage of large-scale AI literacy datasets and assessment efforts, which frames the paper's contribution."}],"review_version":1}