REVIEW 2 major objections 4 minor 25 references
Designing Effective LLM-Assisted Interfaces for Curriculum Development
T0 review · 2 major / 4 minor · reviewed 2026-08-07 · deepseek-v4-flash
Pith's one-line read Clickable, predefined commands outperform free-form chat for LLM-assisted curriculum design.
desk verdict Useful internal comparison between two DM-based LLM interfaces; headline control comparison is confounded by undocumented prompt configuration. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The carrying mechanism is the application of Direct Manipulation to LLM interaction: (DM1) continuous representation of the object of interest, realized as an interactive course-outline table that stays visible; (DM2) physical actions such as clicking, checking, dragging, and dropping instead of typing syntax; (DM3) rapid, incremental, reversible operations with undo/redo and loading effects confined to changed sections; and (DM4) recognition of available commands through menus and buttons rather than recall of prompt syntax. UI Predefined operationalizes this with curated command groups, while UI Open adds a dynamic command chat box with drag-to-localize and save/reuse. Behind the interface, a system prompt defines persona, task, delimiters, and output format, and the UI engine translates user actions into prompts that always include the current course outline.
What would settle it
A replication that captures the raw prompts, system prompt, and model configuration for all three arms would settle the claim: if a standard chat interface given the same system prompt and output format closes the SUS and NASA RTLX gaps, the reported advantage is due to prompt engineering rather than interface design. Alternatively, an in-the-wild study with educators revising real courses under deadline pressure could falsify the lab result if the usability gains disappear.
Extended reading notes
Core claim
The central claim is that interface design is the decisive factor in making LLM-assisted curriculum development usable by educators. The paper claims UI Predefined, an interface that replaces free-form prompts with four groups of expert-curated clickable commands acting on an always-visible interactive course outline, outperformed the ChatGPT-like control on both measured dimensions: usability (SUS 86.75 vs 69.00, p < 0.009) and workload (NASA RTLX 2.25 vs 3.30, p < 0.03). It also claims UI Open, which keeps the same outline view but adds a chat box for dynamic, reusable, and drag-and-drop commands, ranked between the two, with greater flexibility but a steeper learning curve and no statistically significant advantage over ChatGPT. The authors read the results as evidence that applying Direct Manipulation to LLM interaction, with a human check in the loop, is an effective design strategy for educator-LLM collaboration.
Load-bearing premise
The comparison assumes the ChatGPT-like control differed from the two new interfaces only in interface design, with the same underlying system prompt, model settings, and output formatting, but the paper does not document that those were matched.
Editorial extensions
If this is right
- Educators can produce course outlines with less typing and lower cognitive load when LLM commands are predefined and applied directly to the document, rather than expressed as chat prompts.
- The non-significant gap between UI Open and the control suggests that flexibility alone does not buy usability; structure and guidance carry the benefit.
- A hybrid interface that combines predefined commands with the ability to save custom ones, as the paper proposes, would target the flexibility cost seen in UI Open.
- Institutional adoption of LLM tools for curriculum work may hinge on providing expert-curated command sets and visible output formatting, not on prompt-training educators.
Reading between the lines
- A testable extension is to run the same three conditions with the ChatGPT-like control matched to the proposed UIs on system prompt, model version, and output formatting; this would separate interface effects from prompt-engineering effects, which the current report does not fully distinguish.
- The same design pattern likely transfers to other structured authoring tasks, such as assessments, syllabi, and reports, where the object of interest is a visible document with a stable schema, though the paper only tests curriculum outlines.
- The 20-participant sample with 1-21 years of teaching experience suggests the usability gap may be smaller for educators already fluent in prompt engineering; screening by prompt skill could reveal an interaction.
- A field deployment with real course revision deadlines, rather than a lab task, would test whether the workload reduction survives time pressure and context switching.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes two user interfaces, UI Predefined and UI Open, grounded in direct manipulation principles, to reduce the prompt-engineering burden faced by educators when using LLMs for curriculum development. A within-subjects user study with 20 participants compares these UIs against an open-webui-based ChatGPT replica using the System Usability Scale (SUS) and the NASA Raw Task Load Index (RTLX). The results show that UI Predefined achieves the highest SUS score (86.75) and lowest RTLX score (2.25), followed by UI Open (70.75, 3.00) and the ChatGPT replica (69.00, 3.30). Statistical tests using the Wilcoxon signed-rank test are reported for the two winning comparisons, and the authors conclude that UI Predefined significantly outperforms both ChatGPT and UI Open in usability and workload.
Significance. If the control condition is shown to match the treatment arms on model, system prompt, and context provisioning, the study would provide a useful empirical demonstration that GUI-style direct manipulation can improve perceived usability and reduce workload for LLM interaction in a practical educational task. The paper uses standard, externally validated instruments (SUS, NASA RTLX), randomizes the order of interface use, and makes raw data available, which are strengths that support reproducibility. The contribution is moderate but relevant to HCI research on LLM interfaces and to the learning-technology community, as it addresses an actual pain point (complex prompt engineering) for a specific user group (educators).
major comments (2)
- [3.3-3.4, 5.1] The causal attribution in the abstract and Section 5.1 is undermined by an undocumented confound in the control condition. Section 3.3 describes a tailored system prompt for the proposed UIs that defines the persona, delimits prompt components, specifies the course-outline curation task, fixes an output format, and always includes the current course outline and user commands. Section 3.4 describes the control as an open-webui-based ChatGPT replica, but the manuscript never states whether the control used the same model (gpt-4o-2024-08-06), the same system prompt, or the same automatic inclusion of the course outline and user commands. If the control lacked these prompt and context provisions, the reported SUS advantage of 17.75 points and RTLX advantage of 1.05 points (Tables 1 and 2) could arise from prompt engineering or context injection rather than from the direct-manipulation interface itself. The authors should document the control configuration in detail, or run an additional control condition that holds model, system prompt, and context constant while varying only the interface, and then re-analyze the data.
- [Tables 1-2, 5.1] The statistical evidence for the central workload claim is weaker than reported. Tables 1 and 2 present p-values only as inequalities and do not report confidence intervals, effect sizes, or any correction for multiple comparisons. For the NASA RTLX comparison between UI Predefined and ChatGPT, the reported p<0.032 would not survive a Bonferroni correction for the three pairwise comparisons (corrected threshold 0.017), and the comparison between UI Predefined and UI Open (p<0.022) would also not survive. The abstract and Section 5.1 claim significantly reduced task load for UI Predefined, but as reported this claim is not robust. The authors should report exact p-values, effect sizes (e.g., matched rank-biserial correlation), and either a multiple-comparison correction or a pre-specified analysis plan.
minor comments (4)
- [Table 2] The comparison between UI Open and ChatGPT has no reported p-value, yet Section 5.1 describes it as a non-significant improvement; please either provide the test statistic or explicitly state that the test was not performed.
- [Figure 6] The caption says 'The dotted lines are mean,' but the figure legend could clarify which dotted line corresponds to which interface to avoid ambiguity.
- [Reference [15]] The reference lists 'EC-TEL 2020' but the publication year appears to be 2024; please correct the conference or year.
- [Section 6] There is a typo in 'UI Predefined' written as 'UIPredefined' in the sentence 'Our findings revealed that the UIPredefined significantly outperformed ChatGPT.'
Circularity Check
No significant circularity: the paper's claims are empirical user-study results benchmarked against external instruments.
full rationale
This paper is an empirical HCI user study, not a derivation chain. The central claims—that UI Predefined achieves higher SUS scores and lower NASA RTLX workload than a ChatGPT-like control—are supported by participant questionnaire data analyzed with the Wilcoxon signed-rank test. SUS and NASA RTLX are external, established instruments, and the control condition is an open-webui-based ChatGPT replica. There are no fitted parameters, no model equations, and no formal predictions that reduce by construction to their inputs. The direct-manipulation principles are used as design heuristics for building the interfaces, not as a source of the measured outcome. The main weakness identified by the skeptic is that the control condition's model version, system prompt, and context injection are not documented, so the comparison may be confounded by prompt scaffolding. That is a validity or reproducibility concern, not circular reasoning: the SUS and RTLX differences are not definitionally equal to the interface design choices. Citations to the authors' prior work appear in related-work and motivation contexts and are not load-bearing for the experimental comparison. No circular step can be exhibited with specific quoted reductions, so the appropriate score is 0.
Assumptions & free parameters
assumptions (5)
- domain assumption SUS and NASA RTLX are valid measures of usability and workload.
- domain assumption The open-webui interface is a faithful control representing ChatGPT.
- domain assumption Participants' subjective questionnaire responses reflect actual usability and workload differences.
- domain assumption The course outline creation task is representative of curriculum development.
- standard math Wilcoxon signed-rank test is appropriate for paired ordinal data.
Cite this review
Pith. "Pith review of Designing Effective LLM-Assisted Interfaces for Curriculum Development." pith.science (2026). https://pith.science/paper/VP7NS5XU
@misc{pith2026250611767,
author = {Pith},
title = {Pith review of: Designing Effective LLM-Assisted Interfaces for Curriculum Development},
year = {2026},
howpublished = {\url{https://pith.science/paper/VP7NS5XU}},
note = {Machine review of arXiv:2506.11767}
}
read the original abstract
Large Language Models (LLMs) have the potential to transform the way a dynamic curriculum can be delivered. However, educators face significant challenges in interacting with these models, particularly due to complex prompt engineering and usability issues, which increase workload. Additionally, inaccuracies in LLM outputs can raise issues around output quality and ethical concerns in educational content delivery. Addressing these issues requires careful oversight, best achieved through cooperation between human and AI approaches. This paper introduces two novel User Interface (UI) designs, UI Predefined and UI Open, both grounded in Direct Manipulation (DM) principles to address these challenges. By reducing the reliance on intricate prompt engineering, these UIs improve usability, streamline interaction, and lower workload, providing a more effective pathway for educators to engage with LLMs. In a controlled user study with 20 participants, the proposed UIs were evaluated against the standard ChatGPT interface in terms of usability and cognitive load. Results showed that UI Predefined significantly outperformed both ChatGPT and UI Open, demonstrating superior usability and reduced task load, while UI Open offered more flexibility at the cost of a steeper learning curve. These findings underscore the importance of user-centered design in adopting AI-driven tools and lay the foundation for more intuitive and efficient educator-LLM interactions in online learning environments.
Figures
Figures from the paper (3 more)
Reference graph
Works this paper leans on
-
[1]
IEEE Transactions on Learning Technolo- gies14(1), 81–92 (2021)
Abdi, S., Khosravi, H., Sadiq, S., Demartini, G.: Evaluating the quality of learning resources: A learnersourcing approach. IEEE Transactions on Learning Technolo- gies14(1), 81–92 (2021)
work page 2021
-
[2]
In: 2024 International Joint Conference on Neural Networks (IJCNN)
Bahrami, M., Sonoda, R., Srinivasan, R.: Llm diagnostic toolkit: Evaluating llms for ethical issues. In: 2024 International Joint Conference on Neural Networks (IJCNN). pp. 1–8. IEEE (2024)
work page 2024
-
[3]
Bangor, A., Kortum, P.T., Miller, J.T.: An empirical evaluation of the system us- ability scale. Intl. Journal of Human–Computer Interaction24(6), 574–594 (2008)
work page 2008
-
[4]
arXiv preprint arXiv:2310.14735 (2023)
Chen, B., Zhang, Z., Langrené, N., Zhu, S.: Unleashing the potential of prompt engineering in large language models: a comprehensive review. arXiv preprint arXiv:2310.14735 (2023)
arXiv 2023
-
[5]
In: Proceedings of the 2022 CHI Conference on Human Factors in Computing Systems
Dang, H., Mecke, L., Buschek, D.: Ganslider: How users control generative models for images using multiple sliders with and without feedforward information. In: Proceedings of the 2022 CHI Conference on Human Factors in Computing Systems. pp. 1–15 (2022)
work page 2022
-
[6]
arXiv preprint arXiv:2306.10509 (2023)
Denny, P., Khosravi, H., Hellas, A., Leinonen, J., Sarsa, S.: Can we trust ai- generated educational content? comparative analysis of human and ai-generated learning resources. arXiv preprint arXiv:2306.10509 (2023)
arXiv 2023
-
[7]
Canvil: Designerly Adaptation for LLM-Powered User Experiences
Feng, K., Liao, Q.V., Xiao, Z., Vaughan, J.W., Zhang, A.X., McDonald, D.W.: Canvil: Designerly adaptation for llm-powered user experiences. arXiv preprint arXiv:2401.09051 (2024)
work page Pith review arXiv 2024
-
[8]
Gros, B., García-Peñalvo, F.J.: Future trends in the design strategies and techno- logical affordances of e-learning. In: Learning, design, and technology: An interna- tional compendium of theory, research, practice, and policy, pp. 345–367. Springer (2023)
work page 2023
Show all 25 references
-
[9]
Human–computer interaction1(4), 311–338 (1985)
Hutchins, E.L., Hollan, J.D., Norman, D.A.: Direct manipulation interfaces. Human–computer interaction1(4), 311–338 (1985)
1985
-
[10]
In: 2021 IEEE Inter- national Conference on Engineering, Technology & Education (TALE)
Kawamata, T., Matsuda, Y., Sekiya, T., Yamaguchi, K.: Analysis of computer sci- ence textbooks by topic modeling and dynamic time warping. In: 2021 IEEE Inter- national Conference on Engineering, Technology & Education (TALE). p. 865–870. IEEE, Wuhan, Hubei Province, China (De...
2021
-
[11]
ACM Computing Surveys55(13s), 1–39 (2023)
Kosch, T., Karolus, J., Zagermann, J., Reiterer, H., Schmidt, A., Woźniak, P.W.: A survey on measuring cognitive workload in human-computer interaction. ACM Computing Surveys55(13s), 1–39 (2023)
2023
-
[12]
IEEE Transactions on Visualization and Computer Graph- ics (2024)
Liu, Y., Wen, Z., Weng, L., Woodman, O., Yang, Y., Chen, W.: Sprout: an interac- tive authoring tool for generating programming tutorials with the visualization of large language models. IEEE Transactions on Visualization and Computer Graph- ics (2024)
2024
-
[13]
IEEE Transactions on Biomedical Engineering68(2), 685–694 (Feb 2021)
Mariani, A., Pellegrini, E., De Momi, E.: Skill-oriented and performance-driven adaptive curricula for training in robot-assisted surgery using simulators: A feasi- bility study. IEEE Transactions on Biomedical Engineering68(2), 685–694 (Feb 2021)
2021
-
[14]
In: Proceedings of the CHI Conference on Human Factors in Computing Systems
Masson, D., Malacria, S., Casiez, G., Vogel, D.: Directgpt: A direct manipula- tion interface to interact with large language models. In: Proceedings of the CHI Conference on Human Factors in Computing Systems. pp. 1–16 (2024)
2024
-
[15]
Faraji et al
Moein, M., Molavi, M., Faraji, A., Tavakoli, M., Kismihók, G.: Beyond search engines: Can large language models improve curriculum development? In: 19th 14 A. Faraji et al. European Conference on Technology Enhanced Learning, EC-TEL 2020, Krems, Austria, September 16–20, 2024....
2024
-
[16]
Molavi, M., Tavakoli, M., Kismihók, G.: Extracting topics from open educational resources. In: Addressing Global Challenges and Quality Education: 15th Euro- pean Conference on Technology Enhanced Learning, EC-TEL 2020, Heidelberg, Germany, September 14–18, 2020, Proceedings 1...
2020
-
[17]
OpenAI: Prompt engineering - openai api (2024),https://platform.openai.com/ docs/guides/prompt-engineering, [Online; accessed Sep. 2024]
2024
-
[18]
International Journal of Human– Computer Interaction39(3), 391–437 (2023)
Ozmen Garibay, O., Winslow, B., Andolina, S., Antona, M., Bodenschatz, A., Coursaris, C., Falco, G., Fiore, S.M., Garibay, I., Grieman, K., et al.: Six human- centered artificial intelligence grand challenges. International Journal of Human– Computer Interaction39(3), 391–437 (2023)
2023
-
[19]
Computers and Education: Artificial Intelligence7, 100256 (2024)
Padovano,A.,Cardamone,M.:Towardshuman-aicollaborationinthecompetency- based curriculumdevelopmentprocess:The caseof industrialengineering andman- agement education. Computers and Education: Artificial Intelligence7, 100256 (2024)
2024
-
[20]
Computer16(08), 57–69 (1983)
Shneiderman, B.: Direct manipulation: A step beyond programming languages. Computer16(08), 57–69 (1983)
1983
-
[21]
Mol, S., Kismihók, G.: Hybrid human- ai curriculum development for personalised informal learning environments
Tavakoli, M., Faraji, A., Molavi, M., T. Mol, S., Kismihók, G.: Hybrid human- ai curriculum development for personalised informal learning environments. In: LAK22: 12th International Learning Analytics and Knowledge Conference. pp. 563–569 (2022)
2022
-
[22]
Advanced Engineering Informatics52, 101508 (2022)
Tavakoli, M., Faraji, A., Vrolijk, J., Molavi, M.d., Mol, S.T., Kismihók, G.: An ai- based open recommender system for personalized labor market driven education. Advanced Engineering Informatics52, 101508 (2022)
2022
-
[23]
In: Proceedings of the 29th International Conference on Intelligent User Interfaces
Wang, B., Li, Y., Lv, Z., Xia, H., Xu, Y., Sodhi, R.: Lave: Llm-powered agent assistance and language augmentation for video editing. In: Proceedings of the 29th International Conference on Intelligent User Interfaces. pp. 699–714 (2024)
2024
-
[24]
Encyclopedia of Biostatistics8(2005)
Woolson, R.F.: Wilcoxon signed-rank test. Encyclopedia of Biostatistics8(2005)
2005
-
[25]
International Journal of Human–Computer Interaction39(3), 494–518 (2023)
Xu, W., Dainoff, M.J., Ge, L., Gao, Z.: Transitioning to human interaction with ai systems: New challenges and opportunities for hci professionals to enable human- centered ai. International Journal of Human–Computer Interaction39(3), 494–518 (2023)
2023
Reviewed August 7, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.