{"id":"59986354-f3aa-426a-a8db-829938a1e2de","arxiv_id":"2607.17704","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"CT-SEI is a 27-item scale that measures computational-thinking self-efficacy in a Python context and, in preliminary data, splits into create-solution and evaluate-solution factors.","lead":"Researchers built and tested a 27-item questionnaire for measuring university students' self-efficacy in computational thinking during a Python programming task. A preliminary factor analysis suggests the items group into two broad categories—creating a solution and evaluating a solution—but the validation sample was small and the model was adjusted after looking at the data.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Post-hoc CFA on n=70 may not support the two-category CT-SEI claim; fit indices are uninformative at this sample size.","rationale":"The central claim is not that CT self-efficacy exists (that's theoretically grounded), but that this specific 27-item instrument has a replicable two-category structure and can be used to measure it. The CFA in Section 3.2.3 is the only quantitative support for the two-category structure. Yet the model was revised after inspecting the same data: Model 1 did not converge, Model 2 was a post-hoc combination, and Model 3 was another post-hoc modification. With only 70 cases, the model's excellent fit indices are weak evidence: the chi-square test is underpowered, and the RMSEA is insensitive in small samples. The reader identified the same weakest assumption; I agree. The proposed test — fitting Model 3 to the 200-case PCA subsample — is a feasible cross-validation that directly checks whether the structure generalizes to a larger sample that was not used to specify the model. A second part of the test compares a one-factor model; because Cronbach's alpha is .976, a general-factor model might explain the data just as well, which would undermine the 'two categories' summary. The paper is honest about its limitations, so it remains a reasonable preliminary instrument paper; thus the verdict stays CONDITIONAL rather than being upgraded or rejected.","tokens_in":10739,"tokens_out":8686,"duration_ms":94769,"concrete_test":"Using the existing data, fit Model 3 (as specified in Table 4) to the 200-case PCA subsample used for item reduction, which was not used to calibrate Model 3. Compare the CFA fit statistics (chi-square, RMSEA, CFI, TLI) to those reported in Table 3; also fit a one-factor model on the same 200 cases and a four-factor correlated model. If Model 3's CFI drops below 0.95 or RMSEA rises above 0.08 on the larger 200-case sample, or if the one-factor model fits within ΔCFI<0.01 of Model 3, the two-category structure is not stable and the claim should be downgraded to exploratory. If Model 3's fit is comparable and is not worse than a one-factor model, the concern is mitigated.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's headline claim — that the 27 items 'can be divided into two categories: (1) creating the solution and (2) evaluating the solution' — rests on Model 3 in Table 3, but the model was not specified a priori. Section 3.2.3 reports that Model 1 was inadmissible, Model 2 was constructed by collapsing algorithmic thinking and decomposition, and Model 3 was added 'based on an inspection of the items included in the final item set' after seeing the first two results. The CFA is therefore exploratory in all but name, and with n=70 for 27 items the chi-square test has low power to detect misspecification; p=0.146 and RMSEA=0.050 cannot distinguish a well-specified model from an overfit one. The abstract's two-category phrasing further overstates the analysis: Model 3 fits four first-order factors (AB, AT+DC, GE, EV) and then groups three of them under 'creating'; no model in Table 3 actually constrains the items to exactly two factors. A unidimensional model might account for the data nearly as well given Cronbach's alpha of .976.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents the development and preliminary validation of the Computational Thinking Self-Efficacy Instrument (CT-SEI), a 27-item scale measuring self-efficacy for abstraction, algorithmic thinking, decomposition, evaluation, and generalization in the context of a Python programming assignment. Ninety-one candidate items were reduced by expert review to 54, then by principal component analysis on 200 respondents to 27. A confirmatory factor analysis on the remaining 70 respondents is used to claim the items fall into two categories: creating the solution and evaluating the solution. The paper also reports gender differences favoring men on assignment-level self-efficacy and some items.","tokens_in":11021,"tokens_out":6774,"duration_ms":67902,"significance":"If the factor structure were supported, the CT-SEI would be a useful, theory-aligned instrument for computing education research. The paper has genuine strengths: item wording follows Bandura's self-efficacy principles, content validity is addressed via independent expert review, the assignment context is clearly specified, and the final item set is reported in Table 4. However, the central confirmatory claim is not supported by the reported analyses: the two-category structure was selected post hoc and estimated on a very small sample. The contribution at this stage is a carefully constructed item pool and a set of hypotheses about its structure, not a validated two-factor scale.","major_comments":[{"comment":"The headline two-category result is not a confirmatory finding. Model 1 was inadmissible, Model 2 was formed by collapsing algorithmic thinking and decomposition, and Model 3 was added 'based on an inspection of the items included in the final item set' after seeing the earlier results. All models are fit to the same n=70 sample. At this sample size, χ² has low power, and a model selected after inspecting the data can reproduce plausible-looking RMSEA/CFI/TLI values even if it is overfit. The abstract's 'showed the items can be divided into two categories' overstates what the analysis establishes. Please re-estimate the structure on independent data, or recast the analysis as exploratory and temper the conclusion accordingly. The existing limitation statement in §3.4 acknowledges the sample size but does not address the post-hoc selection problem.","section":"§3.2.3 and Table 3"},{"comment":"The model described as a 'two-factor' model is not a two-factor CFA at the item level. In Model 3, the 'Create solution' factor contains three sub-factors (Abstraction, Generalization, and Algorithmic Thinking & Decomposition), while 'Evaluate solution' contains only the Evaluation items. This is a higher-order structure, not a direct division of 27 items into two factors. The abstract and Section 3.2.3 use 'two categories' loosely. In addition, the text says the first factor consists of 'abstraction, generalization, and a combination of algorithmic thinking and abstraction', but Table 4 labels the combined factor 'Algorithmic Thinking & Decomposition'; this discrepancy should be fixed. Please specify the exact model (first-order factors, second-order loadings, constraints) that produced the fit indices in Table 3. A second-order factor with only one first-order factor (Evaluation) is no","section":"§3.2.3 and Table 4"},{"comment":"The item-reduction PCA was run separately for each of the five CT skills on its own item subset, rather than on the full 54-item set. This procedure cannot reveal cross-loadings or assess whether the five theoretically defined components are empirically distinct. It therefore partly entrenches the five-factor theory instead of testing it, and it makes the later CFA results harder to interpret. Given the strong interfactor correlations implied by the inadmissible Model 1 and the eventual collapsing of AT and DC, an exploratory factor analysis on all 54 items would be a more appropriate item-selection step. Please report the full PCA results (eigenvalues, loadings, cross-loadings) or justify the per-skill approach more strongly.","section":"§3.2.2"},{"comment":"The combined alpha of .976 for 27 items is very high, which raises the possibility that a single general factor accounts for most of the variance. To support the two-category interpretation, the paper should report factor correlations from Model 3 and compare Model 3 against a unidimensional model and/or a two-factor model without the sub-factors. Without discriminant validity evidence, the high internal consistency is ambiguous and does not by itself validate the 'creating vs. evaluating' distinction.","section":"§3.2.2, Cronbach's alpha"}],"minor_comments":[{"comment":"The phrase 'principle component analysis' should be 'principal component analysis' (abstract, §3.2, §3.2.2).","section":"Throughout"},{"comment":"'Kaiser maximization' is presumably 'Kaiser normalization'; please correct.","section":"§3.2.2"},{"comment":"The typo 'combination of algorithmic thinking and abstraction' should read 'algorithmic thinking and decomposition' to match Table 4.","section":"§3.2.3"},{"comment":"The full 91-item pool is said to be 'available from the first author upon request'; for reproducibility, the candidate items and expert review outcomes should be included in a supplement or online repository.","section":"Footnote 1 and Table 4"},{"comment":"The gender comparison makes many Mann-Whitney tests without a multiple-comparison correction. Report effect sizes, and consider a conservative correction for the 27/54 item-level tests.","section":"§3.3"}],"recommendation":"major_revision","confidential_remarks":"The paper is a reasonable scale-development report, but the central confirmatory language is not justified by the evidence. The post-hoc model search on 70 participants is the main obstacle. I would be willing to review a revision that either collects/uses independent validation data or clearly repositions the contribution as an exploratory scale construction paper with hypotheses to be tested."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nQuick take: this is a genuine attempt at a context-specific CT self-efficacy scale, and the item development is the strongest part. But the headline two-category factor structure is not supported by the analysis as reported; it's a post-hoc model fit on 70 responses, and the abstract oversells it.\n\nThe good: The authors ground the instrument in Bandura's self-efficacy principles, use the Selby/Dagiene five-skill framework, generate 91 items from existing scales plus new ones, and get independent expert review. The final 27 items in Table 4 are sensible \"I can\" statements tied to a concrete Python guessing-game assignment. That is a real improvement over generic CT scales. The gender split analysis is a nice exploratory addition, and the paper is explicit about its limitations. The writing is clear; the methods are described in enough detail that a reader can see exactly what was done.\n\nThe soft spots: The CFA is exploratory in all but name. Model 1 doesn't converge, Model 2 is a theoretical regroup, and Model 3 is created after inspecting the items. Running three models on the same 70 people and then reporting the best-fitting one is not confirmatory. At n=70, RMSEA=0.050 and p=0.146 can't distinguish a good model from an overfit one; the confidence interval would be wide. The stress-test point is fair: no model in Table 3 actually tests a two-factor solution; Model 3 has four first-order factors under two higher-order factors, so saying \"items can be divided into two categories\" is misleading. And with Cronbach's alpha at .976, a unidimensional model might fit just as well, which the paper never tests. The data is also not public (items \"available upon request\"), so the analyses can't be independently reproduced.\n\nThat said, I don't think these flaws are fatal to the instrument's potential. The item set is theory-driven and expert-vetted; the PCA reduction was reasonable for a preliminary study; and the authors explicitly call for larger-sample revalidation and test-retest reliability. The problem is the claim in the abstract, which goes beyond the evidence. A careful reviewer would send this back for major revision, not reject it outright.\n\nVerdict: worth sending to peer review. The paper would benefit from a venue that does scale development and is willing to push the authors to reframe Model 3 as exploratory, add a one-factor competitor, and ideally collect a fresh sample before claiming the factor structure. I'd also ask them to post the full item pool and data. For a reading group on instrument validation, this is a great case study.\n\nRegards.","headline":"Useful instrument development with a genuinely good item pool, but the two-factor CFA claim rests on post-hoc model fitting with n=70 and is oversold in the abstract.","tokens_in":11482,"tokens_out":2590,"would_cite":true,"duration_ms":26982,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Computational thinking self-efficacy can be measured with a 27-item, context-bound survey whose responses split into two factors: creating a solution and evaluating it.","keywords":["Computational Thinking","self-efficacy","instrument development","CT-SEI","confirmatory factor analysis","principal component analysis","higher education","gender differences"],"falsifier":"Administer the final 27-item instrument to a new, larger sample of higher-education students in a preregistered confirmatory factor analysis. If the create/evaluate two-factor model shows RMSEA above 0.06 or CFI/TLI below 0.95, or if the two factors correlate near 1, the claimed structure fails to reproduce. A simpler check: low test-retest correlations on the same students would indicate the instrument does not measure a stable self-efficacy belief.","tokens_in":10643,"feed_emoji":"🧩","tokens_out":4587,"duration_ms":46948,"temperature":0.7,"pith_summary":"The paper sets out to measure a psychological belief: whether higher-education students think they can perform computational thinking (CT) skills such as abstraction, algorithmic thinking, decomposition, evaluation, and generalization. Because self-efficacy depends on context, the survey anchors every item in a concrete Python programming assignment. Starting from 91 candidate statements, expert review and principal component analysis reduced the set to 27, and confirmatory factor analysis on 70 additional responses suggests the items form two high-level factors: creating the solution and evaluating it. If this factor structure holds, educators and researchers gain a practical instrument for profiling student confidence and tailoring CT instruction.","feed_headline":"27-item scale splits CT self-efficacy into two factors","feed_subtitle":"A new survey measures whether students believe they can build and check solutions, helping educators tailor computing instruction.","key_machinery":"The instrument's three-part structure: a context-setting assignment, one overall task self-efficacy question, and 27 'I can' statements covering CT sub-skills on a 0-10 scale. The statistical machinery is principal component analysis with oblimin rotation for item reduction, followed by confirmatory factor analysis on a separate 70-response sample. The two-factor CFA model (create versus evaluate) is the load-bearing result that translates the theoretical skill list into a measurable two-dimensional belief structure.","core_discovery":"The central claim is that CT self-efficacy in higher education can be measured as a context-dependent belief, and that when measured with 'I can' statements tied to a Python assignment, the 27 retained items do not reproduce the five theoretical CT skills. Instead, the best-fitting model groups the items into two categories: creating the solution (abstraction, algorithmic thinking, decomposition, and generalization items) and evaluating the solution (evaluation items), with reported fit statistics of RMSEA = 0.050, CFI = 0.972, TLI = 0.969. The paper also reports that women in the sample rated their CT self-efficacy lower than men on most items, a pattern consistent with broader findings in","pith_inferences":["If the two-factor create/evaluate structure replicates, it would suggest that students' self-efficacy does not track the five-skill taxonomy that guided item construction; the taxonomy may describe competence better than perceived competence.","The gender differences here could reflect measurement non-invariance across groups or particularities of the sample; a formal invariance test would show whether the scale measures the same construct for men and women.","Because the CFA sample is small and models were selected after inspecting the data, a preregistered replication on a new sample, including test-retest reliability, is the natural next test of whether the two-factor structure is stable.","The instrument's link to actual behavior is untested; correlating CT-SEI scores with performance on the programming task or course grades would show whether self-reported confidence predicts the outcomes that matter."],"forward_implications":["Educators can use the 27 items to profile a student's CT self-efficacy separately for creating and evaluating solutions, rather than relying on a single overall score.","The scale can serve as an outcome measure in intervention studies, including the authors' planned comparison of worked examples versus practice problems.","Because the instrument is context-dependent, the same items can be adapted to other CT tasks by changing the assignment text, though validity in new contexts still needs testing.","The reported gender differences suggest that CT instruction may need to address self-efficacy directly, especially for women in computing courses.","The CFA results imply that separate self-efficacy beliefs for algorithmic thinking, decomposition, abstraction, and generalization may not be distinguishable in this population."],"fun_headline_variants":["CT self-efficacy splits into creating and evaluating solution factors","27-item scale reveals two-factor CT self-efficacy structure","Study: CT self-efficacy has two dimensions, not five skills","New CT self-efficacy survey: build and check factors emerge"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The load-bearing premise is that a two-factor structure fitted with only 70 respondents and with models chosen after inspecting the data describes a stable structure in the broader student population rather than a small-sample artifact.","fun_headline_variants_meta":{"raw":{"variants":["CT self-efficacy splits into creating and evaluating solution factors","27-item scale reveals two-factor CT self-efficacy structure","Study: CT self-efficacy has two dimensions, not five skills","New CT self-efficacy survey: build and check factors emerge"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000158,"raw_usage":{"total_tokens":1072,"prompt_tokens":764,"completion_tokens":308,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":508,"completion_tokens_details":{"reasoning_tokens":254}},"tokens_in":508,"tokens_out":308,"duration_ms":4079,"temperature":1.0,"reasoning_tokens":254,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-01T17:12:36.219339+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Administer the final 27-item instrument to a new, larger sample of higher-education students in a preregistered confirmatory factor analysis. If the create/evaluate two-factor model shows RMSEA above 0.06 or CFI/TLI below 0.95, or if the two factors correlate near 1, the claimed structure fails to reproduce. A simpler check: low test-retest correlations on the same students would indicate the instrument does not measure a stable self-efficacy belief.","supporting_citations":[],"review_version":1}