REVIEW 3 major objections 6 minor 22 references
Insights into User Interface Innovations from a Design Thinking Workshop at deRSE25
T0 review · 3 major / 6 minor · reviewed 2026-08-05 · deepseek-v4-flash
Pith's one-line read A design-thinking workshop produced user-generated UI concepts—branching panels, context weighting, hover interactions, and per-message checkboxes—that now guide the authors' whiteboard-based LLM interface and point to flexible context mana
desk verdict Raw workshop data worth having; the 'advanced our UI' claim is unsupported without a pre-workshop baseline. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is the authors' whiteboard-based LLM interface: prompts and outputs are discrete movable blocks on an infinite canvas, with an always-visible context window that users can fill, reorder, or deactivate, so the model only sees what the user chooses. The discovery instrument is the workshop itself—a short design-thinking protocol with use-case collection, like/dislike cards, a shared-reflection question, and a 15-minute flipchart prototyping task—followed by a two-stage inductive thematic grouping that stays close to participants' verbatim wording. The context window carries the design argument: it turns context from an invisible chronological buffer into a user-managed resou
What would settle it
Run the same design-thinking protocol with a demographically broad group that includes non-programmers and general consumers, and check whether flexible context management and conversation branching still emerge as top needs; alternatively, build a whiteboard-style prototype with and without context checkboxes and weighting, and measure whether users actually prefer or perform better with them than with a linear chat interface.
Extended reading notes
Core claim
The central discovery is that ordinary workshop participants, given the prompt 'How to overcome the limitations of chat based LLM interfaces?', independently generated seven interface concepts that converge on the authors' design direction. The mechanisms include dragging useful answers into new side panels, hover-triggered sub-conversations, an 'auto-forget' command, pinning or highlighting messages to weight their influence, assigning explicit importance scales, output-format dropdowns, and per-message checkboxes that include or exclude items from the active context. The paper integrates these into the whiteboard UI as branch visualization, a user-facing weighting function, and hover-based
Load-bearing premise
The workshop participants were self-selected research-software conference attendees, so the argument that these UI priorities are the right direction for LLM users generally rests on that sample standing in for all users.
Editorial extensions
If this is right
- LLM interfaces should let users include or exclude individual messages from the active context instead of feeding the whole chronological history into the model.
- Branching should become a first-class interaction: parallel conversation threads shown as side panels or a path diagram, not just editing a previous message.
- User-assigned weights, via highlights or explicit scales, should influence which parts of the conversation shape the next output.
- Because combining too many innovations can hurt usability, multiple prototypes with different feature sets should be built and tested with users.
- If model capabilities plateau, interface usability and user experience will become the main differentiator between LLM products.
Reading between the lines
- The participant pool skews toward research-software engineers, so the same 60-minute protocol should be rerun with non-programmers and general consumers before treating these themes as universal; the authors acknowledge the audience emphasis, but the generalizability test is an editorial extension.
- The weighting idea could be implemented without changing the model by reordering, duplicating, or summarizing weighted segments inside the prompt—an operational path the paper does not spell out.
- The absence of UI criticism could alternatively mean the chat interface is good enough for many tasks; distinguishing habituation from satisfaction would require longer-term usage studies, not just a one-hour workshop.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper reports on a 60-minute design thinking workshop held at the deRSE25 conference, in which seven groups of 4–6 attendees provided LLM use cases, positive and negative experiences, and sketches of novel LLM user interfaces. The authors transcribe the workshop outputs verbatim, group them manually into thematic categories, and describe seven participant-generated UI concepts. They then relate these concepts to their own previously developed whiteboard-based LLM interface, selecting three ideas—branching visualization, a weighting function, and hover-based interaction—as particularly valuable for future development. The paper situates this work within the human-centered design (HCD) process and argues that LLM interface innovation should receive more attention.
Significance. If taken as an exploratory, qualitative design case study, this paper is useful: it provides a transparent account of a participatory design workshop, includes verbatim participant data, and offers concrete UI ideas (branching, weighting, hover interactions) that plausibly extend current linear chat interfaces. The authors are explicit about several limitations, including the self-selected conference sample and the manual, interpretive nature of the analysis. The contribution is moderate rather than transformative: it is a single, small-sample workshop report, not a controlled evaluation or a validated design framework. Its main value lies in demonstrating a method for involving users in LLM UI ideation and in articulating specific design directions that could inform future prototypes.
major comments (3)
- [§5.2 and Abstract] The claim that workshop ideas 'advanced our own whiteboard-based UI approach' is not verifiable from the manuscript. Section 5.1 describes the pre-existing whiteboard UI and lists its features (selective context management, dynamic reordering, visibility, blueprint design), but it does not state whether branching visualization, weighting functions, or hover-based interaction were part of the design before the workshop. Without a pre-workshop baseline, the reader cannot assess whether these features were genuinely added because of the workshop or were already planned. Please either provide a dated or versioned pre-workshop feature list, or revise the abstract and §5.2 to say the workshop 'informed', 'validated', or 'inspired' the ongoing design, explicitly distinguishing new ideas from pre-existing ones.
- [§4.1 and §5.2] The thematic grouping of responses and, more importantly, the selection of 'particularly valuable' UI ideas are performed manually by the authors, with no inter-rater reliability, independent coding, or member checking. Since the authors organized the workshop and have a stake in validating their own UI concept, the mapping from participant sketches to the chosen design directions in §5.2 is interpretive. This is a central link in the paper's argument, not a peripheral detail. I recommend adding a reflexivity statement, and ideally having a second coder independently map the seven sketches to the features in §5.2; at minimum, the authors should temper the strength of the causal/inspirational claim and acknowledge the selection as their own design judgment.
- [§3 and §6] The paper generalizes from a single self-selected sample of deRSE25 attendees (seven groups of 4–6, no demographic information) to broader 'future LLM interface development'. While §6 acknowledges the conference-audience emphasis on software development, the broader implications are still framed in general terms. The evidence supports hypothesis generation for research-software users, not a general claim about LLM users at large. Please add explicit boundary conditions, e.g., 'these insights are exploratory and specific to a technical, science-oriented user group', and avoid unqualified statements about general LLM interface needs.
minor comments (6)
- [Table 4 and §4.2] The table header reads 'Lack of functional capabilities (15) / usability (1)' while the text says 'functional capabilities and usability (15 mentions)'. Clarify whether the total is 15 or 16 and make the text and table consistent.
- [UI Solution 5 (Figure 7)] The prose states 'similarities with the previous branching and highlighting visualisation in figure 5'; this likely should refer to Figure 6 (UI Solution 4), not Figure 5. Please correct the cross-reference.
- [UI Solution 7 (Figure 9)] Typo: 'clarification ot terms' should be 'clarification of terms'.
- [§2.3] Typo: 'choosen' should be 'chosen'.
- [§3] The sentence 'we ask the participants' should be 'we asked the participants' for tense consistency.
- [§3] Please provide a stable URL or persistent identifier for the workshop materials on the HiFis conference portal; as written, the citation is not actionable.
Circularity Check
No significant circularity; the workshop insights are reported qualitative data, not derivations fitted to the conclusion.
full rationale
This paper does not present a derivation chain: it reports a design-thinking workshop, transcribes participant responses, and describes thematically grouped results (Section 4). No parameters are fitted, no quantity is predicted from a fitted input, and no formal result is imported from prior work by the same authors. The nearest self-referential element is that the authors already had a whiteboard UI concept (Section 5.1) before the workshop and later selected participant ideas they found 'particularly valuable' (Section 5.2); this is a potential confirmation/interpretation bias and an internal-validity limitation for the causal claim that the workshop 'advanced' their UI, but it is not circular in the sense that the conclusion is equivalent to the input by construction. The cited external methods, such as Hsieh and Shannon (2005) and ISO 9241-210, are used only as methodological framing. The lack of a pre-workshop feature baseline and the self-selected sample are limitations, not circular steps. Therefore no circularity is identified.
Assumptions & free parameters
assumptions (3)
- domain assumption The manual two-stage thematic grouping by the authors, without inter-rater reliability testing, is a valid representation of participant responses.
- domain assumption A 60-minute design-thinking workshop with voluntary participation elicits user needs relevant to LLM interface design.
- domain assumption The deRSE25 attendee sample is sufficiently representative of LLM users to inform UI design.
Cite this review
Pith. "Pith review of Insights into User Interface Innovations from a Design Thinking Workshop at deRSE25." pith.science (2026). https://pith.science/paper/AMW7SUTL
@misc{pith2026250818784,
author = {Pith},
title = {Pith review of: Insights into User Interface Innovations from a Design Thinking Workshop at deRSE25},
year = {2026},
howpublished = {\url{https://pith.science/paper/AMW7SUTL}},
note = {Machine review of arXiv:2508.18784}
}
read the original abstract
Large Language Models have become widely adopted tools due to their versatile capabilities, yet their user interfaces remain limited, often following rigid, linear interaction paradigms. In this paper, we present insights from a design thinking workshop held at the deRSE25 conference aiming at collaboratively developing innovative user interface concepts for LLMs. During the workshop, participants identified common use cases, evaluated the strengths and shortcomings of current LLM interfaces, and created visualizations of new interaction concepts emphasizing flexible context management, dynamic conversation branching, and enhanced mechanisms for user control. We describe how these participant-generated ideas advanced our own whiteboard-based UI approach. The ongoing development of this interface is guided by the human-centered design process - an iterative, user-focused methodology that emphasizes continuous refinement through user feedback. Broader implications for future LLM interface development are discussed, advocating for increased attention to UI innovation grounded in user-centered design principles.
Figures
Figures from the paper (4 more)
Reference graph
Works this paper leans on
-
[1]
[BGMS21] E. M. Bender, T. Gebru, A. McMillan-Major, S. Shmitchell. On the dangers of stochastic parrots: Can language models be too big? In Proceedings of the 2021 ACM conference on fairness, accountability, and transparency. Pp. 610–623
work page 2021
-
[8]
doi:10.36227/techrxiv.23589741.v1 http://dx.doi.org/10.36227/techrxiv.23589741.v1 [HW05] E. Hollnagel, D. D. Woods. Joint Cognitive Systems: Foundations of Cognitive Systems Engineering. CRC Press,
-
[9]
https://arxiv.org/abs/2307.10169 [KR18] T. Kudo, J. Richardson. Sentencepiece: A simple and language independent subword tokenizer and detokenizer for neural text processing. arXiv preprint arXiv:1808.06226,
-
[10]
doi:https://doi.org/10.1016/j.lindif.2023.102274 https://www.sciencedirect.com/science/article/pii/S1041608023000195 [MDW+23] P. Ma, R. Ding, S. Wang, S. Han, D. Zhang. Demonstration of InsightPilot: An LLM-Empowered Automated Data Exploration System
arXiv 2023
-
[11]
Demonstration of InsightPilot: An LLM-Empowered Automated Data Exploration System
https://arxiv.org/abs/2304.00477 [MMBK24] L. Metzger, L. Miller, M. Baumann, J. Kraus. Empowering Calibrated (Dis-)Trust in Conversational Agents: A User Study on the Persuasive Power of Limitation Disclaimers vs. Authoritative Style. In Proceedings of the 2024 CHI Conference on Human Factors in Computing Systems . CHI ’24. Association for Computing Machi...
work page Pith review arXiv 2024
-
[12]
doi:10.1145/3613904.3642122 https://doi.org/10.1145/3613904.3642122 [MMCV24] D. Masson, S. Malacria, G. Casiez, D. V ogel. DirectGPT: A Direct Manipulation Interface to Interact with Large Language Models. In Proceedings of the 2024 CHI Conference on Human Factors in Computing Systems . CHI ’24. Association for Computing Machinery, New York, NY , USA,
arXiv 2024
- [13]
-
[15]
https://arxiv.org/abs/2407.20046 [MP13] F. Macpherson, D. Platchias. Hallucination: Philosophy and psychology . MIT Press,
Show all 22 references
-
[16]
https://arxiv.org/abs/2307.06435 [Nor86] D. A. Norman. Cognitive engineering. User Centered System Design , pp. 31–61,
-
[18]
doi:10.1080/00140139.2024.2434604 https://doi.org/10.1080/00140139.2024.2434604 [RNS+18] A
PMID: 39610202. doi:10.1080/00140139.2024.2434604 https://doi.org/10.1080/00140139.2024.2434604 [RNS+18] A. Radford, K. Narasimhan, T. Salimans, I. Sutskever et al. Improving language understanding by generative pre-training
2024
-
[22]
doi:10.1145/3708359.3712125 https://doi.org/10.1145/3708359.3712125 [ZZL+25] W. X. Zhao, K. Zhou, J. Li, T. Tang, X. Wang, Y . Hou, Y . Min, B. Zhang, J. Zhang, Z. Dong, Y . Du, C. Yang, Y . Chen, Z. Chen, J. Jiang, R. Ren, Y . Li, X. Tang, Z. Liu, P. Liu, J.-Y . Nie, J.-R. We...
-
[23]
https://arxiv.org/abs/2303.18223 28 / 28
-
[1986]
27 / 28 Design Thinking Workshop at deRSE25 [fNB10] D. D. I. f ¨ur Normung (Berlin).ISO 9241-210:2019 Ergonomie der Mensch-System- Interaktion: Teil 210: Prozess zur Gestaltung gebrauchstauglicher interaktiver Systeme (ISO 9241-210:2019) : Ausgabe: 2019-04-01 ; Deutsche Fassun...
2019
-
[1987]
[SDG+24] M. Shen, S. Das, K. Greenewald, P. Sattigeri, G. Wornell, S. Ghosh. Ther- mometer: Towards universal calibration for large language models. arXiv preprint arXiv:2403.08819,
-
[1996]
[GBB+20] L. Gao, S. Biderman, S. Black, L. Golding, T. Hoppe, C. Foster, J. Phang, H. He, A. Thite, N. Nabeshima, S. Presser, C. Leahy. The Pile: An 800GB Dataset of Diverse Text for Language Modeling.arXiv preprint arXiv:2101.00027,
-
[2005]
doi:10.1177/1049732305276687 https://doi.org/10.1177/1049732305276687 [HtQ+23] M
PMID: 16204405. doi:10.1177/1049732305276687 https://doi.org/10.1177/1049732305276687 [HtQ+23] M. U. Hadi, q. a. tashi, R. Qureshi, A. Shah, a. muneer, M. Irfan, A. Zafar, M. B. Shaikh, N. Akhtar, J. Wu, S. Mirjalili. A Survey on Large Language Models: Ap- plications, Challeng...
-
[2017]
http://arxiv.org/abs/1706.03762 [WWPY25] X. Wang, X. Wang, S. Park, Y . Yao. Mental Models of Generative AI Chatbot Ecosystems. In Proceedings of the 30th International Conference on Intelligent User Interfaces . IUI ’25, p. 1016–1031. Association for Computing Machinery, New ...
-
[2020]
Gadiraju, S
[GKD+23] V . Gadiraju, S. Kane, S. Dev, A. Taylor, D. Wang, R. Denton, R. Brewer. ”I wouldn’t say offensive but...”: Disability-Centered Perspectives on Large Language 25 / 28 Design Thinking Workshop at deRSE25 Models. In Proceedings of the 2023 ACM Conference on Fairness, Ac...
2023
-
[2021]
Gallotta, G
doi:https://doi.org/10.1016/j.dss.2021.113515 https://www.sciencedirect.com/science/article/pii/S0167923621000257 [GTZ+24] R. Gallotta, G. Todd, M. Zammit, S. Earle, A. Liapis, J. Togelius, G. N. Yannakakis. Large Language Models and Games: A Survey and Roadmap. IEEE Transacti...
2021
-
[2023]
doi:10.1145/3593013.3593989 https://doi.org/10.1145/3593013.3593989 [GSG21] G. M. Grimes, R. M. Schuetzler, J. S. Giboney. Mental models and expectation violations in conversational AI interactions.Decision Support Systems144:113515,
-
[2024]
doi:10.1109/TG.2024.3461510 [HHS24] M. T. Hicks, J. Humphries, J. Slater. ChatGPT is bullshit. Ethics and Information Technology 26(2):1–10,
2024
-
[2025]
Mart ´ınez, L
https://arxiv.org/abs/2402.06196 [MMR24] P. Mart ´ınez, L. Moreno, A. Ramos. Exploring Large Language Models to generate Easy to Read content
Reviewed August 5, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.