REVIEW 4 major objections 4 minor 30 references
How Instructional Sequence and Personalized Support Impact Diagnostic Strategy Learning
T0 review · 4 major / 4 minor · reviewed 2026-08-15 · deepseek-v4-flash
Pith's one-line read This paper claims that problem-solving before instruction, with instruction illustrated by the student's own interaction, leads to better transfer of diagnostic reasoning than instruction-first teaching.
desk verdict The headline claim of far-transfer superiority for PS-I is unsupported by the paper's own statistics; only a fragile near-transfer subscore survives, and the design confounds sequence with example personalization. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The mechanism is the sequence itself: in PS-I, the student first conducts a diagnostic conversation with a simulated client, and the instruction that follows is illustrated with examples taken from that student's own just-completed interaction; in I-PS, the same instruction is given first, illustrated with a hypothetical case, and personalized feedback arrives only after the task. The outcome machinery is a multidimensional scoring of three diagnostic strategies across scenarios of increasing complexity: adherence to the LINDAFF checklist (a seven-category symptom inquiry), questions about interpersonal relationships (e.g., asking about the mother or baby), and data interpretation (listing possible causes, assigning likelihoods, and justifying them).
What would settle it
Run a replication with three conditions: I-PS, PS-I, and a variant where the same personalized example content is used in both timings; if PS-I no longer beats I-PS on far transfer, the sequencing claim is falsified.
Extended reading notes
Core claim
The paper's central claim is that providing explicit diagnostic strategy instruction before problem-solving (I-PS) is less effective than letting students solve a case first and then instructing them with examples drawn from their own interaction (PS-I). The advantage shows up specifically in transfer: PS-I students maintained their diagnostic-reasoning scores in the far-transfer scenario, whereas I-PS students' scores dropped, and PS-I also outperformed I-PS on the interpersonal-relationship strategy in the near-transfer scenario. The authors interpret the pattern as evidence that the initial problem-solving attempt helps learners build a provisional understanding that the later instruction and personalized feedback can refine.
Load-bearing premise
The causal claim that sequencing causes the transfer difference assumes the two conditions differ only in the order of instruction and problem-solving, but PS-I students received instruction examples drawn from their own interaction while I-PS students received hypothetical examples, so personalization is tied to the sequence.
Editorial extensions
If this is right
- Scenario-based learning environments aimed at transfer should place a first problem-solving attempt before explicit strategy instruction, rather than front-loading the instruction.
- Personalized feedback tied to the student's own prior interaction is a plausible active ingredient; instruction alone, without that feedback loop, may not produce measurable learning gains.
- The productive-failure pattern previously shown in math and physics extends to diagnostic strategy learning in vocational healthcare training.
- Comparisons between instructional sequences should measure far-transfer performance, because near-transfer alone may be too easy to expose sequencing differences.
- The I-PS group's performance decline in the most complex client scenario points to cognitive load rather than lack of knowledge as a possible boundary condition.
Reading between the lines
- Not tested by the paper: the difficulty of the first problem-solving attempt may moderate the effect, since too-easy cases would not generate the knowledge gaps that later instruction fills and too-hard cases could overload novices.
- Not tested by the paper: holding the example content constant across both timings and varying only the order would isolate sequence from personalization, and if PS-I's transfer advantage disappears, personalization rather than sequence is the driver.
- Not tested by the paper: a direct cognitive-load measure, such as time-on-task or help-seeking behavior, could decide whether I-PS students' decline is overload rather than failed transfer.
- Not tested by the paper: learners with stronger prior domain knowledge may need less struggle before instruction, a prediction the current design does not address.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The manuscript reports a between-groups experiment (N = 80 pharmacy apprentices) in the PharmaSim scenario-based learning environment, comparing instruction-before-problem-solving (I-PS) with problem-solving-before-instruction (PS-I). Diagnostic strategy performance is scored on checklist (LINDAFF), interpersonal relationship, and data-interpretation measures across a learning phase, a near-transfer client, and a far-transfer two-client scenario. The headline claim is that PS-I yields significantly higher transfer performance, particularly far transfer, and that the I-PS group did not outperform a no-instruction baseline.
Significance. If the central claim were supported, the study would be a useful contribution to the productive-failure and scenario-based learning literatures, and it would inform practical decisions about instructional sequencing with personalized feedback. The design has real strengths: random assignment, a realistic interactive simulation, process-logged outcomes, a pretest equivalence check, and multilevel models with student random effects. However, the manuscript's own reported statistics do not support the advertised claim. The only significant between-condition contrast is a single uncorrected near-transfer interpersonal-relationship subscore, and the sequencing manipulation is confounded with example personalization. As reported, the evidence is better characterized as exploratory rather than as a demonstration that PS-I improves transfer.
major comments (4)
- [Section 3, MLM results and post-hoc comparisons] The central claim is contradicted by the reported statistics. The mixed linear models found no effect of experimental group condition and no interactions between scenarios and experimental group across all strategies, and the post-hoc comparisons found no significant differences between conditions except for the near-transfer Client B interpersonal-relationship score (p = .0438). No far-transfer between-condition contrast is significant. The within-group I-PS decline on Client C2 relative to Clients B and C1 cannot establish that PS-I outperforms I-PS on far transfer; that would require a significant group-by-scenario interaction or a direct between-group contrast on the far-transfer measures. Thus the Abstract's assertion of 'significantly higher performance in transfer tasks' and the Introduction's 'PS-I significantly improves far-transfer performance' are not supported by the paper's own tests.
- [Section 2.1, Procedure] The two conditions differ not only in instructional sequence but also in the content of the illustrative examples. PS-I participants received instruction illustrated with examples drawn from their own prior interaction with Client A, while I-PS participants received instruction with examples based on a hypothetical case because they had not yet interacted with Client A. Any observed advantage for PS-I, including the significant near-transfer interpersonal-relationship difference, is therefore confounded with personalization of examples. A clean test of sequencing would require holding example content constant across conditions or crossing personalization with order. Without such a design, the paper cannot attribute the result to 'problem-solving before instruction.'
- [Section 4, Discussion and Conclusions] The statement that 'the I-PS group did not outperform a "no-instruction" baseline' is unsupported, because the study has no no-instruction control group. For the same reason, the Abstract's claim that both instruction types are beneficial is not directly tested; with only two conditions, the design can compare I-PS with PS-I but cannot establish benefit relative to no instruction. This unsupported baseline comparison should be removed or explicitly marked as not tested by the data.
- [Section 3, post-hoc comparisons] The single significant between-condition result (p = .0438) is drawn from many comparisons across three strategy scores, several scenarios/clients, and within-group contrasts, and no multiplicity correction is reported. An uncorrected p-value near .05 is weak evidence in isolation, and it is not sufficient to carry the paper's strong transfer conclusion. The authors should report adjusted p-values, confidence intervals, or effect sizes for all contrasts, or explicitly frame the finding as hypothesis-generating.
minor comments (4)
- [Section 3] The sentence reporting 'no significant differences between conditions for all clients (p > 0.5)' appears to contain a typo; the threshold should likely be p > 0.05, and the exact p-values should be stated.
- [Section 2.2, Measurement and Analysis] The interpersonal relationships strategy score is described as a scale of 0 to 3 but then as a proportion of fulfilled categories; please clarify whether raw scores or percentages are reported and keep the normalization procedure consistent across all strategy scores.
- [Section 2.2] The term 'post-test' is used both for the test taken immediately after the learning-phase diagnostic conversation and for the data-interpretation score averaged across Clients C1 and C2; these are different measures and should be given distinct names to avoid confusion.
- [Abstract and Introduction] The phrases 'transfer tasks' and 'far-transfer performance' are used interchangeably, while the Results distinguish near and far transfer; the claims in the Abstract and Introduction should be aligned with the specific transfer measures actually analyzed.
Circularity Check
No significant circularity: the study is an empirical group comparison with no derivation-to-input loop.
full rationale
This paper reports a between-groups experiment comparing instructional sequences (PS-I vs. I-PS) in an online scenario-based learning environment. There is no mathematical derivation, no fitted parameter later relabeled as a prediction, and no claim that a first-principles model reproduces an outcome by construction. The dependent measures are rubric-based scores of diagnostic strategies; although the rubric uses the same strategy categories that the instruction teaches, this is a measurement alignment, not a circular derivation: the scores are independently scored from student interactions, and the central comparison is an empirical difference between randomized conditions. No load-bearing conclusion rests on a self-citation: references to prior productive-failure and transfer literature are contextual and supportive, not used as a uniqueness theorem or as a substitute for the reported data. The Discussion's assertions about far transfer and a 'no-instruction' baseline are questionable statistically and by design, but those are correctness or evidentiary concerns, not circularity. Therefore, no circular step can be substantiated from the quoted text, and the appropriate finding is no significant circularity.
Assumptions & free parameters
assumptions (3)
- domain assumption Diagnostic quality scores derived from the LINDAFF checklist, interpersonal relationships, and data interpretation are valid measures of the target learning outcomes.
- domain assumption Random assignment and pretest equivalence ensure that differences between conditions can be attributed to the intervention rather than baseline differences.
- domain assumption The personalized feedback system correctly identifies the correct diagnostic process for each scenario.
Cite this review
Pith. "Pith review of How Instructional Sequence and Personalized Support Impact Diagnostic Strategy Learning." pith.science (2026). https://pith.science/paper/IBKISBUA
@misc{pith2026250717760,
author = {Pith},
title = {Pith review of: How Instructional Sequence and Personalized Support Impact Diagnostic Strategy Learning},
year = {2026},
howpublished = {\url{https://pith.science/paper/IBKISBUA}},
note = {Machine review of arXiv:2507.17760}
}
read the original abstract
Supporting students in developing effective diagnostic reasoning is a key challenge in various educational domains. Novices often struggle with cognitive biases such as premature closure and over-reliance on heuristics. Scenario-based learning (SBL) can address these challenges by offering realistic case experiences and iterative practice, but the optimal sequencing of instruction and problem-solving activities remains unclear. This study examines how personalized support can be incorporated into different instructional sequences and whether providing explicit diagnostic strategy instruction before (I-PS) or after problem-solving (PS-I) improves learning and its transfer. We employ a between-groups design in an online SBL environment called PharmaSim, which simulates real-world client interactions for pharmacy technician apprentices. Results indicate that while both instruction types are beneficial, PS-I leads to significantly higher performance in transfer tasks.
Figures
Reference graph
Works this paper leans on
-
[1]
International Journal of Artificial Intelligence in Education34, 974–1007 (2024)
Abdelshiheed, M., Barnes, T., Chi, M.: How and when: The impact of metacog- nitive knowledge instruction and motivation on transfer across intelligent tutoring systems. International Journal of Artificial Intelligence in Education34, 974–1007 (2024)
work page 2024
-
[2]
Psychological Bulletin128(4), 612–637 (2002)
Barnett, S.M., Ceci, S.J.: When and where do we apply what we learn?: A taxon- omy for far transfer. Psychological Bulletin128(4), 612–637 (2002)
work page 2002
-
[3]
The New England Journal of Medicine355(21), 2217–2225 (2006)
Bowen, J.L.: Educational strategies to promote clinical diagnostic reasoning. The New England Journal of Medicine355(21), 2217–2225 (2006)
work page 2006
-
[4]
Cognitive Science5(2), 121–152 (1981)
Chi, M.T., Feltovich, P.J., Glaser, R.: Categorization and representation of physics problems by experts and novices. Cognitive Science5(2), 121–152 (1981)
work page 1981
-
[5]
Academic Medicine78(8), 775–780 (2003)
Croskerry, P.: The importance of cognitive errors in diagnosis and strategies to minimize them. Academic Medicine78(8), 775–780 (2003)
work page 2003
-
[6]
In: International Society of the Learning Sciences (2022)
DeCaro, M.S., McClellan, D.K., Powe, A., Franco, D., Chastain, R.J., Hieb, J.L., Fuselier, L.: Exploring an online simulation before lecture improves undergraduate chemistry learning. In: International Society of the Learning Sciences (2022)
work page 2022
-
[7]
Harvard University Press, Cambridge, MA (1978)
Elstein, A.S., Shulman, L.S., Sprafka, S.A.: Medical Problem Solving: An Analysis of Clinical Reasoning. Harvard University Press, Cambridge, MA (1978)
work page 1978
-
[8]
Diagnosis 11(4), 374–379 (2024)
Eriksen, T., Gögenur, I.: Interprofessional clinical reasoning education. Diagnosis 11(4), 374–379 (2024)
work page 2024
Show all 30 references
-
[9]
Medical Education38(1), 7–14 (2004)
Eva, K.W.: What every teacher needs to know about clinical reasoning. Medical Education38(1), 7–14 (2004)
2004
-
[10]
StatPearls Publishing (2020)
Gillis, A.: Family Dynamics. StatPearls Publishing (2020)
2020
-
[11]
BMJ Quality & Safety21(7), 535–557 (2012)
Graber, M.L., Kissam, S., Payne, V.L., Meyer, A.N.D., Sorensen, A., Lenfestey, N., Tant, E., Henriksen, K., Labresh, K.: Cognitive interventions to reduce diagnostic error: A narrative review. BMJ Quality & Safety21(7), 535–557 (2012)
2012
-
[12]
BMJ Quality & Safety28(6), 495–500 (2019)
Graber, M.L., Trowbridge, R.: Clinical reasoning checklists: A tool to reduce diag- nostic error. BMJ Quality & Safety28(6), 495–500 (2019)
2019
-
[13]
Instructional Science40(4), 651–672 (2012)
Kapur, M.: Productive failure in learning the concept of variance. Instructional Science40(4), 651–672 (2012)
2012
-
[14]
BMJ Case Reports (2011)
Kumar, B., Kanna, B., Kumar, S.: The pitfalls of premature closure: Clinical decision-making in a case of aortic dissection. BMJ Case Reports (2011)
2011
-
[15]
Educational Psychology Review29(4), 693– 715 (2017)
Loibl, K., Roll, I., Rummel, N.: Towards a theory of when and how problem-solving before instruction supports learning. Educational Psychology Review29(4), 693– 715 (2017)
2017
-
[16]
Instructional Science 42(2), 305–326 (2014)
Loibl, K., Rummel, N.: The impact of guidance during problem-solving prior to instruction on students’ inventions and learning outcomes. Instructional Science 42(2), 305–326 (2014)
2014
-
[17]
Instructional Science48(2), 135–175 (2020)
Loibl, K., Tillema, M., Rummel, N., van Gog, T.: The effect of contrasting cases during problem solving prior to and after instruction. Instructional Science48(2), 135–175 (2020)
2020
-
[18]
Springer, 2nd edn
McDaniel, S.H., Campbell, T.L., Hepworth, J.: Family-Oriented Primary Care: A Manual for Medical Providers. Springer, 2nd edn. (2005)
2005
-
[19]
Med- ical Education39(4), 418–427 (2005)
Norman, G.: Research in clinical reasoning: Past history and current trends. Med- ical Education39(4), 418–427 (2005)
2005
-
[20]
PharmaWiki: LINDAAFF.https://www.pharmawiki.ch/wiki/index.php?wiki= LINDAAFF, accessed: 12 February 2025 8 F. B. Güreş et al
2025
-
[21]
Roll, I., Butler, D., Yee, N., Welsh, A., Perez, S., Briseno, A., Perkins, K., Bonn, D.: Understanding the Impact of Guiding Inquiry: The Relationship Between Di- rective Support, Student Attributes, and Transfer of Knowledge, Attitudes, and Behaviours in Inquiry Learning. Ins...
2018
-
[22]
Clark, Richard E
Ruth C. Clark, Richard E. Mayer: Scenario-based e-Learning: Evidence-Based Guidelines for Online Workforce (2012)
2012
-
[23]
Saba, J.,Kapur,M., Roll,I.: TheDevelopment ofMultivariable CausalityStrategy: Instruction or Simulation First? In: Lecture Notes in Computer Science (including subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformat- ics). vol. 13916 LNAI, pp. 41–53...
2023
-
[24]
Educational Psychologist24(2), 113–142 (1989)
Salomon, G., Perkins, D.N.: Rocky roads to transfer: Rethinking the mechanisms of a neglected phenomenon. Educational Psychologist24(2), 113–142 (1989)
1989
-
[25]
Cog- nition and Instruction22(2), 129–184 (2004)
Schwartz, D.L., Martin, T.: Inventing to prepare for future learning: The hidden efficiency of encouraging original student production in statistics instruction. Cog- nition and Instruction22(2), 129–184 (2004)
2004
-
[26]
International Journal of Artificial Intelligence in Education34, 825–861 (2024)
Shabrina, P., Mostafavi, B., Abdelshiheed, M., Chi, M., Barnes, T.: Investigating the impact of backward strategy learning in a logic tutor: Aiding subgoal learning towards improved problem solving. International Journal of Artificial Intelligence in Education34, 825–861 (2024)
2024
-
[27]
Learn- ing and Instruction75(10 2021)
Sinha, T., Kapur, M.: Robust effects of the efficacy of explicit failure-driven scaf- folding in problem-solving prior to instruction: A replication and extension. Learn- ing and Instruction75(10 2021)
2021
-
[28]
Computers & Education: Artificial Intelligence2, 100017 (2021)
Sinha, T., Kapur, M.: When problem solving followed by instruction works: Ev- idence for productive failure. Computers & Education: Artificial Intelligence2, 100017 (2021)
2021
-
[29]
Science185(4157), 1124–1131 (1974)
Tversky, A., Kahneman, D.: Judgment under uncertainty: Heuristics and biases. Science185(4157), 1124–1131 (1974)
1974
-
[30]
International Journal of Artificial Intelligence in Education32, 931–970 (2022)
Zhang, N., Biswas, G., Hutchins, N.: Measuring and analyzing students’ strategic learning behaviors in open-ended learning environments. International Journal of Artificial Intelligence in Education32, 931–970 (2022)
2022
Reviewed August 15, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.