{"id":"2655f2b0-6b66-43e2-b6ef-86b01b9a41aa","arxiv_id":"2507.17760","paper_version":1,"verdict":"REJECT","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"high","formal_verification":"none","parameter_count":0,"one_line_summary":"Problem-solving before instruction improved one near-transfer strategy score, while the claimed far-transfer benefit is not supported by the between-group tests reported.","lead":"This study compared learning diagnostic strategies in an online pharmacy simulation when instruction came before or after problem-solving. The problem-solving-first group outperformed the instruction-first group on a near-transfer interpersonal score, but the far-transfer advantage named in the abstract is not present in the reported statistics.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Headline claim is unsupported by the paper's own statistics: no far-transfer between-condition contrast is significant, and the sole significant transfer effect is a near-transfer interpersonal subscore (p=.0438) that is fragile and overgeneralized.","rationale":"The paper is a small two-arm RCT whose central claim is a causal ordering effect on transfer. I read the statistical results as the least secure support for that claim. The most direct test of sequence, a between-condition contrast on far transfer, is never reported as significant; the abstract and discussion claim a far-transfer advantage that the Results section does not establish. The reader's weakest_assumption, that personalization is confounded with sequence, is a real design problem and would invalidate causal interpretation even if the statistics were strong. My primary concern is upstream: the inferential statistics do not reach the claim. I therefore keep the REJECT verdict: the paper should not be accepted with these claims as stated. A focused reanalysis of the between-condition contrasts with multiplicity control would settle whether any transfer effect survives. This is a good-faith critique of the evidence-claim link, not a dispute with the productive-failure literature; the paper's own reported numbers are the basis.","tokens_in":5881,"tokens_out":4929,"duration_ms":51566,"concrete_test":"Re-run the full between-condition analysis on the raw PharmaSim logs: for each strategy (checklist, interpersonal, data interpretation) and each transfer client (B, C1, C2), fit a mixed model with condition as a fixed effect and pretest as a covariate, then apply Benjamini-Hochberg correction to the family of between-condition contrasts. If the corrected Client B interpersonal contrast (reported p = .0438) no longer remains below .05, there is no statistically significant PS-I transfer advantage, and the paper would need to report a null between-condition effect with a possible exploratory subscore difference.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The condition that would have to hold for the central claim, that PS-I leads to significantly higher transfer performance, is a significant between-condition advantage on transfer measures. Section 3 reports exactly one such comparison: the near-transfer Client B interpersonal-relationship subscore, where PS-I (63.6) exceeded I-PS (48.6) with p = .0438. No far-transfer contrast between conditions is reported as significant. The Discussion's far-transfer conclusion is instead inferred from a within-group pattern: the I-PS group scored lower on Client C2 (32.9) than on Client B (49.7, p = .0190) and Client C1 (53.4, p = .0023), while PS-I performance was stable. A within-group decline across scenarios in one arm cannot establish that PS-I outperforms I-PS on far transfer; the between-condition test is the direct evidence, and it is absent. The one significant near-transfer subscore is also one of many post-hoc comparisons across strategies and clients; with no correction for multiplicity, p = .0438 is weak. The Abstract and Introduction additionally claim improvements in \"transfer tasks\" and \"far-transfer performance\" that the paper's own tests do not provide. The Discussion's reference to a \"no-instruction\" baseline is also unsupported because the design has no such control group. Thus the load-bearing statistical evidence for the headline claim is missing or at most a single uncorrected near-transfer subscore.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript reports a between-groups experiment (N = 80 pharmacy apprentices) in the PharmaSim scenario-based learning environment, comparing instruction-before-problem-solving (I-PS) with problem-solving-before-instruction (PS-I). Diagnostic strategy performance is scored on checklist (LINDAFF), interpersonal relationship, and data-interpretation measures across a learning phase, a near-transfer client, and a far-transfer two-client scenario. The headline claim is that PS-I yields significantly higher transfer performance, particularly far transfer, and that the I-PS group did not outperform a no-instruction baseline.","tokens_in":6125,"tokens_out":4152,"duration_ms":42536,"significance":"If the central claim were supported, the study would be a useful contribution to the productive-failure and scenario-based learning literatures, and it would inform practical decisions about instructional sequencing with personalized feedback. The design has real strengths: random assignment, a realistic interactive simulation, process-logged outcomes, a pretest equivalence check, and multilevel models with student random effects. However, the manuscript's own reported statistics do not support the advertised claim. The only significant between-condition contrast is a single uncorrected near-transfer interpersonal-relationship subscore, and the sequencing manipulation is confounded with example personalization. As reported, the evidence is better characterized as exploratory rather than as a demonstration that PS-I improves transfer.","major_comments":[{"comment":"The central claim is contradicted by the reported statistics. The mixed linear models found no effect of experimental group condition and no interactions between scenarios and experimental group across all strategies, and the post-hoc comparisons found no significant differences between conditions except for the near-transfer Client B interpersonal-relationship score (p = .0438). No far-transfer between-condition contrast is significant. The within-group I-PS decline on Client C2 relative to Clients B and C1 cannot establish that PS-I outperforms I-PS on far transfer; that would require a significant group-by-scenario interaction or a direct between-group contrast on the far-transfer measures. Thus the Abstract's assertion of 'significantly higher performance in transfer tasks' and the Introduction's 'PS-I significantly improves far-transfer performance' are not supported by the paper's own tests.","section":"Section 3, MLM results and post-hoc comparisons"},{"comment":"The two conditions differ not only in instructional sequence but also in the content of the illustrative examples. PS-I participants received instruction illustrated with examples drawn from their own prior interaction with Client A, while I-PS participants received instruction with examples based on a hypothetical case because they had not yet interacted with Client A. Any observed advantage for PS-I, including the significant near-transfer interpersonal-relationship difference, is therefore confounded with personalization of examples. A clean test of sequencing would require holding example content constant across conditions or crossing personalization with order. Without such a design, the paper cannot attribute the result to 'problem-solving before instruction.'","section":"Section 2.1, Procedure"},{"comment":"The statement that 'the I-PS group did not outperform a \"no-instruction\" baseline' is unsupported, because the study has no no-instruction control group. For the same reason, the Abstract's claim that both instruction types are beneficial is not directly tested; with only two conditions, the design can compare I-PS with PS-I but cannot establish benefit relative to no instruction. This unsupported baseline comparison should be removed or explicitly marked as not tested by the data.","section":"Section 4, Discussion and Conclusions"},{"comment":"The single significant between-condition result (p = .0438) is drawn from many comparisons across three strategy scores, several scenarios/clients, and within-group contrasts, and no multiplicity correction is reported. An uncorrected p-value near .05 is weak evidence in isolation, and it is not sufficient to carry the paper's strong transfer conclusion. The authors should report adjusted p-values, confidence intervals, or effect sizes for all contrasts, or explicitly frame the finding as hypothesis-generating.","section":"Section 3, post-hoc comparisons"}],"minor_comments":[{"comment":"The sentence reporting 'no significant differences between conditions for all clients (p > 0.5)' appears to contain a typo; the threshold should likely be p > 0.05, and the exact p-values should be stated.","section":"Section 3"},{"comment":"The interpersonal relationships strategy score is described as a scale of 0 to 3 but then as a proportion of fulfilled categories; please clarify whether raw scores or percentages are reported and keep the normalization procedure consistent across all strategy scores.","section":"Section 2.2, Measurement and Analysis"},{"comment":"The term 'post-test' is used both for the test taken immediately after the learning-phase diagnostic conversation and for the data-interpretation score averaged across Clients C1 and C2; these are different measures and should be given distinct names to avoid confusion.","section":"Section 2.2"},{"comment":"The phrases 'transfer tasks' and 'far-transfer performance' are used interchangeably, while the Results distinguish near and far transfer; the claims in the Abstract and Introduction should be aligned with the specific transfer measures actually analyzed.","section":"Abstract and Introduction"}],"recommendation":"reject","confidential_remarks":"The manuscript is within scope for cs.CY and the authors report a reasonable amount of design and measurement detail. The problem is not a trivial wording issue: the headline causal claim about sequencing is contradicted by the paper's own inferential statistics and is further confounded by the personalization difference between conditions. A revision that merely softens the conclusion would still leave the central claim unsupported without a new experiment; I therefore recommend rejection rather than major revision."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: the paper's headline claim, that PS-I leads to significantly higher transfer performance, is not supported by its own statistics. The only significant between-condition difference is a near-transfer interpersonal-relationship subscore (p = .0438), and no far-transfer between-group contrast reaches significance. The Discussion's far-transfer conclusion rests on a within-group decline in the I-PS arm, which cannot show that PS-I beats I-PS.\n\nThat said, this is a real empirical study with a sensible research question. The authors ran a between-groups experiment with 80 pharmacy apprentices in a simulated client environment, checked pretest equivalence, used mixed linear models, and report effect sizes in the figures. Extending the problem-solving-before-instruction paradigm to diagnostic strategy learning in scenario-based environments is a legitimate and underexplored context. The near-transfer interpersonal finding is worth following up, even though it is fragile.\n\nThe soft spots are load-bearing. First, the abstract and introduction claim 'significantly higher performance in transfer tasks' and 'significantly improves far-transfer performance,' but the paper does not report a single significant far-transfer between-condition contrast. The within-group comparison (I-PS students scored lower on Client C2 than on B and C1) shows instability in one arm, not superiority of the other; the direct between-condition test is missing. Second, the one significant between-group effect is a single subscore among many post-hoc comparisons; with no multiplicity correction, p = .0438 is weak. Third, the design confounds instructional sequence with the source of the illustrative examples: PS-I received examples drawn from their own prior interaction, I-PS received examples from a hypothetical case. The authors acknowledge this but then still attribute the difference to sequencing. Fourth, the Discussion's claim that I-PS did not beat a 'no-instruction' baseline is unsupported because no no-instruction control exists.\n\nThe near-transfer subscore is a real observation, but it cannot carry the abstract or the far-transfer conclusion. The paper needs major revision: either reframe the claims around what the data actually show, or add the missing between-condition tests and address the confound. I would send it to peer review rather than desk-reject, because the question and the data are worth a serious look, but as written it should not be accepted.","headline":"The headline claim of far-transfer superiority for PS-I is unsupported by the paper's own statistics; only a fragile near-transfer subscore survives, and the design confounds sequence with example personalization.","tokens_in":6666,"tokens_out":2332,"would_cite":false,"duration_ms":22843,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper claims that problem-solving before instruction, with instruction illustrated by the student's own interaction, leads to better transfer of diagnostic reasoning than instruction-first teaching.","keywords":["diagnostic reasoning","instructional sequencing","problem-solving before instruction","personalized feedback","scenario-based learning","transfer of learning","pharmacy education","strategy learning"],"falsifier":"Run a replication with three conditions: I-PS, PS-I, and a variant where the same personalized example content is used in both timings; if PS-I no longer beats I-PS on far transfer, the sequencing claim is falsified.","tokens_in":5679,"feed_emoji":"🧠","tokens_out":6966,"duration_ms":67103,"temperature":0.7,"pith_summary":"This paper asks whether pharmacy apprentices learn diagnostic reasoning better when they try a realistic case before receiving explicit strategy instruction, or when they receive the instruction first. In an online scenario-based environment called PharmaSim, 80 apprentices were randomly assigned to one of the two sequences, and both groups eventually received the same instruction and personalized feedback. The central finding is that the problem-solving-first group performed significantly better on transfer tasks, and its performance stayed stable in the complex far-transfer scenario while the instruction-first group declined. If the result is right, scenario-based training in professional fields should let novices struggle with a case first and then give them instruction tied to their own actions.","feed_headline":"Problem-first training beats instruction-first for diagnostic skills","feed_subtitle":"Apprentices who solved a case before instruction held up on complex transfer tests; instruction-first students declined.","key_machinery":"The mechanism is the sequence itself: in PS-I, the student first conducts a diagnostic conversation with a simulated client, and the instruction that follows is illustrated with examples taken from that student's own just-completed interaction; in I-PS, the same instruction is given first, illustrated with a hypothetical case, and personalized feedback arrives only after the task. The outcome machinery is a multidimensional scoring of three diagnostic strategies across scenarios of increasing complexity: adherence to the LINDAFF checklist (a seven-category symptom inquiry), questions about interpersonal relationships (e.g., asking about the mother or baby), and data interpretation (listing possible causes, assigning likelihoods, and justifying them).","core_discovery":"The paper's central claim is that providing explicit diagnostic strategy instruction before problem-solving (I-PS) is less effective than letting students solve a case first and then instructing them with examples drawn from their own interaction (PS-I). The advantage shows up specifically in transfer: PS-I students maintained their diagnostic-reasoning scores in the far-transfer scenario, whereas I-PS students' scores dropped, and PS-I also outperformed I-PS on the interpersonal-relationship strategy in the near-transfer scenario. The authors interpret the pattern as evidence that the initial problem-solving attempt helps learners build a provisional understanding that the later instruction and personalized feedback can refine.","pith_inferences":["Not tested by the paper: the difficulty of the first problem-solving attempt may moderate the effect, since too-easy cases would not generate the knowledge gaps that later instruction fills and too-hard cases could overload novices.","Not tested by the paper: holding the example content constant across both timings and varying only the order would isolate sequence from personalization, and if PS-I's transfer advantage disappears, personalization rather than sequence is the driver.","Not tested by the paper: a direct cognitive-load measure, such as time-on-task or help-seeking behavior, could decide whether I-PS students' decline is overload rather than failed transfer.","Not tested by the paper: learners with stronger prior domain knowledge may need less struggle before instruction, a prediction the current design does not address."],"forward_implications":["Scenario-based learning environments aimed at transfer should place a first problem-solving attempt before explicit strategy instruction, rather than front-loading the instruction.","Personalized feedback tied to the student's own prior interaction is a plausible active ingredient; instruction alone, without that feedback loop, may not produce measurable learning gains.","The productive-failure pattern previously shown in math and physics extends to diagnostic strategy learning in vocational healthcare training.","Comparisons between instructional sequences should measure far-transfer performance, because near-transfer alone may be too easy to expose sequencing differences.","The I-PS group's performance decline in the most complex client scenario points to cognitive load rather than lack of knowledge as a possible boundary condition."],"supporting_citations":[{"why":"Supplies the productive-failure evidence that problem-solving before instruction can improve learning.","marker":"[13]"},{"why":"The closest prior comparison of instruction-first versus simulation-first for a causality strategy, which this study extends to diagnostic reasoning.","marker":"[23]"},{"why":"Provides the 'inventing to prepare for future learning' account for why initial struggle deepens transfer.","marker":"[25]"},{"why":"Frames the open question of when and how problem-solving before instruction supports learning.","marker":"[15]"},{"why":"Defines the near-transfer/far-transfer taxonomy used to distinguish the study's transfer scenarios.","marker":"[2]"},{"why":"Supports the productive-failure mechanism with evidence that instruction after failure is effective.","marker":"[28]"}],"fun_headline_variants":["Solve first, learn later: better diagnostic transfer","Problem-first instruction boosts diagnostic transfer","Case-before-lesson improves diagnostic reasoning transfer","PS-I beats I-PS on diagnostic transfer tasks","Solving before instruction improves diagnostic transfer"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The causal claim that sequencing causes the transfer difference assumes the two conditions differ only in the order of instruction and problem-solving, but PS-I students received instruction examples drawn from their own interaction while I-PS students received hypothetical examples, so personalization is tied to the sequence.","fun_headline_variants_meta":{"raw":{"variants":["Solve first, learn later: better diagnostic transfer","Problem-first instruction boosts diagnostic transfer","Case-before-lesson improves diagnostic reasoning transfer","PS-I beats I-PS on diagnostic transfer tasks","Solving before instruction improves diagnostic transfer"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00059,"raw_usage":{"total_tokens":2693,"prompt_tokens":792,"completion_tokens":1901,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":408,"completion_tokens_details":{"reasoning_tokens":1835}},"tokens_in":408,"tokens_out":1901,"duration_ms":13608,"temperature":1.0,"reasoning_tokens":1835,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T23:06:41.169720+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run a replication with three conditions: I-PS, PS-I, and a variant where the same personalized example content is used in both timings; if PS-I no longer beats I-PS on far transfer, the sequencing claim is falsified.","supporting_citations":[{"cited_title":"Instructional Science40(4), 651–672 (2012)","cited_arxiv_id":null,"evidence_quote":"Supplies the productive-failure evidence that problem-solving before instruction can improve learning."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"The closest prior comparison of instruction-first versus simulation-first for a causality strategy, which this study extends to diagnostic reasoning."},{"cited_title":"Cog- nition and Instruction22(2), 129–184 (2004)","cited_arxiv_id":null,"evidence_quote":"Provides the 'inventing to prepare for future learning' account for why initial struggle deepens transfer."},{"cited_title":"Educational Psychology Review29(4), 693– 715 (2017)","cited_arxiv_id":null,"evidence_quote":"Frames the open question of when and how problem-solving before instruction supports learning."},{"cited_title":"Psychological Bulletin128(4), 612–637 (2002)","cited_arxiv_id":null,"evidence_quote":"Defines the near-transfer/far-transfer taxonomy used to distinguish the study's transfer scenarios."},{"cited_title":"Computers & Education: Artificial Intelligence2, 100017 (2021)","cited_arxiv_id":null,"evidence_quote":"Supports the productive-failure mechanism with evidence that instruction after failure is effective."}],"review_version":1}