REVIEW 3 major objections 4 minor 55 references
Generating Proto-Personas through Prompt Engineering: A Case Study on Efficiency, Effectiveness and Empathy
T0 review · 3 major / 4 minor · reviewed 2026-08-06 · deepseek-v4-flash
Pith's one-line read A prompt-engineering approach to proto-persona generation, evaluated with 19 practitioners in a real Lean Inception, reduced creation time to about six minutes and was well accepted, but affective and behavioral empathy were only partially achieved.
desk verdict Useful, honest case study with a real speed claim that doesn't quite hold up—worth refereeing, but the 'days to under six minutes' framing needs to be fixed. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
Extended reading notes
Core claim
The paper's central claim, stated in the Synthesis for RQ1, is: 'Our prompt-engineering approach cut proto-persona creation from days to under six minutes, automating repetitive steps and accelerating output.' If true, it demonstrates that a structured prompt workflow can make early-stage product discovery faster and more collaborative, with high user acceptance and strong cognitive empathy, though weaker affective and behavioral empathy.
Load-bearing premise
The approach's effectiveness depends on the assumption that the two input artifacts, Product Vision and the 'Is/Is Not/Does/Does Not' matrix, are sufficient and correctly transcribed to generate contextually accurate proto-personas. The paper checks transcription with a single reviewer (Section 4.4.2) but does not validate the quality or completeness of these artifacts. If these inputs are weak or incomplete, the generated personas will inherit those flaws, and the claimed 'Conformity of the Proto-personas with the Project Context' would not generalize.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes a prompt-engineering-based approach for generating proto-personas during Lean Inception, refining a workflow previously proposed by the authors and evaluating it in a real software project with 19 participants. The empirical design combines timing data, TAM-based questionnaires, empathy questionnaires, and semi-structured interviews, with thematic synthesis used to answer four research questions on efficiency, effectiveness, acceptance, and empathy. The authors report that the approach is efficient, generally well accepted, and yields personas with strong cognitive empathy but weaker affective and behavioral empathy, and they provide an availability link to the study artifacts.
Significance. If the findings are taken as an exploratory qualitative case study, the paper makes a useful contribution: it embeds the prompt-engineering intervention in a real Lean Inception, uses multiple data sources, reports saturation, includes both technical and non-technical participants, and makes its instruments and prompts publicly available. The explicit attention to empathy as a multidimensional construct and the use of an Ishikawa diagram to link causes and effects are strengths. The main limitation is that the headline efficiency claim is broader than what the measurement actually supports, so the paper requires a careful reframing of that claim before its conclusions can be accepted.
major comments (3)
- [§5.0.1 (RQ1 Synthesis), §4.4.5, §4.5] The synthesis statement 'Our prompt-engineering approach cut proto-persona creation from days to under six minutes' is not supported by the measured data. The 5.94-minute average (SD 2.38) reported in §5.0.1 corresponds only to the individual execution of the pre-scripted prompts, as described in §4.4.5. The untimed activities include the LI convergence phase, in which the 19 participants collaboratively merged their outputs into final proto-personas, the preparation and transcription of the Product Vision and Is/Is Not matrix artifacts (§4.4.2), and any human review or refinement. In addition, the comparison in §4.5 is made against an external four-hour figure from [9], not against a measured manual baseline under comparable conditions. The paper should either redefine the claim as 'prompt execution takes under six minutes, while the full proto-persona process includes additional convergence time' or measure and report the total end-to-end time. This is load-bearing because the abstract and RQ1 conclusion currently present the unqualified 'days to under six minutes' as the main efficiency result.
- [§4.3, §5.0.2 (RQ2)] The effectiveness claim in the conclusion ('The approach proved efficient and effective') is stronger than what the evidence supports. The metric defined in §4.3 for RQ2 is the 'Effect of the generated proto-personas to the quality of MVP definition during Lean Inception,' but the analysis in §5.0.2 is based entirely on participant self-reports and researcher validation of the Ishikawa diagram. There is no objective measure of MVP quality, no independent evaluation of the generated personas, and no comparison against personas produced by a manual process. This is acceptable if the claim is explicitly framed as perceived effectiveness, but the current wording conflates perception with outcome. Please either soften the conclusion or add a complementary evaluation that does not rely solely on participants' impressions.
- [§4.4.2 and §5.0.2 (Conformity theme)] The study does not validate the quality or completeness of the two input artifacts, Product Vision and the Is/Is Not/Does/Does Not matrix, beyond a single reviewer checking that the transcription was faithful. The participants' positive assessments of contextual conformity, summarized in the theme 'Conformity of the Proto-personas with the Project Context,' therefore partly depend on the quality of these inputs rather than on the prompt-engineering approach itself. If the inputs are incomplete or biased, the generated personas inherit those flaws. The paper should state this dependency as a limitation and, ideally, include a brief content check of the input artifacts against the project's documentation or stakeholder knowledge.
minor comments (4)
- [§5.0.3] There are several typos and wording issues, including 'colaborate' instead of 'collaborate', 'dilema' instead of 'dilemma', and '5-Likert-Scale' instead of '5-point Likert scale.'
- [§5.0.1 (RQ1)] The text attributes a statement to 'Pereira et al. [19]', but reference [19] is Jackson et al.; please correct the citation or the name used in the text.
- [§4.4.5] The description of the execution phase says both that 'the execution was conducted by two researchers' and that 'each participant executed the approach separately'; please clarify who actually ran the prompts and who recorded the time, as this affects the interpretation of the 5.94-minute figure.
- [§5.0.4] For the empathy Likert results, the paper reports percentages (e.g., 32% neutral, 15% disagreement) without giving the exact counts. Since the sample has only 19 participants, reporting the raw counts alongside percentages would improve transparency and allow readers to judge the precision of the claims.
Circularity Check
No circular derivation: the empirical findings rest on independent participant feedback and measured execution times, with only a minor non-load-bearing self-citation.
full rationale
This paper is an empirical case study rather than a formal derivation, so the circularity patterns involving equations or fitted inputs do not apply. The prompt-engineering approach is described as 'a refined version of the method proposed by Leão et al. [24]' (Section 3), which is a self-citation, but the evidence for the reported results is independent of that prior work: RQ1 uses participant-reported times plus a measured average execution time of 5.94 minutes (Section 5.0.1), RQ2 and RQ3 use thematic synthesis of interviews and TAM Likert responses, and RQ4 uses empathy questionnaire responses. No parameter is fitted to a target result and then presented as a prediction; the efficiency comparison to a 'four hour session [9]' is an external benchmark, not a variable defined by the paper's own outputs. The fact that the convergence phase is untimed is a measurement validity concern, not a circularity. The only self-referential element is the provenance of the approach from the authors' prior work, which does not force any of the empirical findings; therefore, no circular step can be exhibited with a specific reduction.
Assumptions & free parameters
assumptions (3)
- domain assumption The two Lean Inception artifacts (Product Vision and the Is/Is Not/Does/Does Not matrix) are sufficient and accurate inputs for generating contextually appropriate proto-personas.
- domain assumption Self-reported perceptions via TAM and empathy Likert scales measure the intended constructs of acceptance and empathy.
- domain assumption The execution time of the approach (average 5.94 minutes) is a meaningful measure of efficiency even though no manual baseline was measured in the same context.
Cite this review
Pith. "Pith review of Generating Proto-Personas through Prompt Engineering: A Case Study on Efficiency, Effectiveness and Empathy." pith.science (2026). https://pith.science/paper/GLAYGQ2P
@misc{pith2026250708594,
author = {Pith},
title = {Pith review of: Generating Proto-Personas through Prompt Engineering: A Case Study on Efficiency, Effectiveness and Empathy},
year = {2026},
howpublished = {\url{https://pith.science/paper/GLAYGQ2P}},
note = {Machine review of arXiv:2507.08594}
}
read the original abstract
Proto-personas are commonly used during early-stage Product Discovery, such as Lean Inception, to guide product definition and stakeholder alignment. However, the manual creation of proto-personas is often time-consuming, cognitively demanding, and prone to bias. In this paper, we propose and empirically investigate a prompt engineering-based approach to generate proto-personas with the support of Generative AI (GenAI). Our goal is to evaluate the approach in terms of efficiency, effectiveness, user acceptance, and the empathy elicited by the generated personas. We conducted a case study with 19 participants embedded in a real Lean Inception, employing a qualitative and quantitative methods design. The results reveal the approach's efficiency by reducing time and effort and improving the quality and reusability of personas in later discovery phases, such as Minimum Viable Product (MVP) scoping and feature refinement. While acceptance was generally high, especially regarding perceived usefulness and ease of use, participants noted limitations related to generalization and domain specificity. Furthermore, although cognitive empathy was strongly supported, affective and behavioral empathy varied significantly across participants. These results contribute novel empirical evidence on how GenAI can be effectively integrated into software Product Discovery practices, while also identifying key challenges to be addressed in future iterations of such hybrid design processes.
Figures
Reference graph
Works this paper leans on
-
[9]
Paulo Caroli. 2017. Lean inception. São Paulo, BR: Caroli. org (2017)
work page 2017
-
[1]
Iftekhar Ahmed, Aldeida Aleti, Haipeng Cai, Alexander Chatzigeorgiou, Pinjia He, Xing Hu, Mauro Pezzè, Denys Poshyvanyk, and Xin Xia. 2025. Artificial Intelligence for Software Engineering: The Journey so far and the Road ahead. ACM Trans. Softw. Eng. Methodol. (April 2025). https://doi.org/10.1145/3719006 Just Accepted
-
[2]
Chetan Arora, John Grundy, and Mohamed Abdelrazek. 2023. Advancing re- quirements engineering through generative ai: Assessing the role of llms. arXiv preprint arXiv:2310.13976 (2023)
work page Pith review arXiv 2023
-
[3]
Leonardo Banh, Florian Holldack, and Gero Strobel. 2025. Copiloting the Future: How Generative AI Transforms Software Engineering.Information and Software Technology (2025), 107751
work page 2025
-
[4]
Tuomas Bazzan, Benjamin Olojo, Przemysław Majda, Thomas Kelly, Mu- rat Yilmaz, Gerard Marks, and Paul Clarke. 2024. Analysing the Role of Generative AI in Software Engineering – Results from an MLR. In Systems, Software and Services Process Improvement (EuroSPI 2024) (Communications in Computer and Information Science, Vol. 1927). Springer Nature, 163–180...
-
[5]
Lenz Belzner, Thomas Gabor, and Martin Wirsing. 2023. Large language model assisted software engineering: prospects, challenges, and a case study. In International Conference on Bridging the Gap between AI and Reality. Springer, 355–374
2023
-
[6]
Jan Bosch. 2019. From Efficiency to Effectiveness: Delivering Busi- ness Value Through Software. In Software Business. ICSOB 2019 (Lecture Notes in Business Information Processing, Vol. 370), Sami Hyrynsalmi, Mari Suoranta, Anh Nguyen-Duc, Pasi Tyrväinen, and Pekka Abrahamsson (Eds.). Springer, Cham, 3–10. https://doi.org/10.1007/978-3-030-33742-1_1
-
[7]
Anders Bruun, Niels van Berkel, Dimitrios Raptis, and Effie Lai-Chong Law
Show all 55 references
-
[8]
Juan González Calleros, Soraia Prietch, and Josefina Guerrero García. 2024. Using AI Tools for Generating Proto-Personas: An Exploration in the De- sign of Strategies for Promoting Ethical Awareness on Responsible Comput- ing. Avances en Interacción Humano-Computadora 2024 (20...
2024
-
[10]
Daniela Soares Cruzes and Tore Dybå. 2011. Recommended Steps for The- matic Synthesis in Software Engineering. In Proceedings of the International Symposium on Empirical Software Engineering and Measurement. 275–284. https://doi.org/10.1109/ESEM.2011.36
2011 doi
-
[11]
Fred D. Davis. 1989. Perceived Usefulness, Perceived Ease of Use, and User Acceptance of Information Technology. MIS Quarterly 13, 3 (1989), 319–340. https://doi.org/10.2307/249008
1989 doi
-
[12]
Christof Ebert and Panos Louridas. 2023. Generative AI for software practitioners. IEEE Software 40, 4 (2023), 30–38
2023
-
[13]
Angela Fan, Beliz Gokkaya, Mark Harman, Mitya Lyubarskiy, Shubho Sengupta, Shin Yoo, and Jie M Zhang. 2023. Large language models for software engineering: Survey and open problems. arXiv preprint arXiv:2310.03533 (2023)
2023 arXiv
-
[14]
Bruna Moraes Ferreira, Simone D. J. Barbosa, and Tayana Conte. 2016. PA- THY: Using Empathy with Personas to Design Applications that Meet the Users’ Needs. In Human-Computer Interaction. Theory, Design, Development and Practice (Lecture Notes in Computer Science, Vol. 9731). ...
2016 doi
-
[15]
J. Gothelf. 2012. Using proto-personas for executive alignment. https://uxmag. com/articles/using-proto-personas-for-executive-alignment/
2012
-
[16]
O’Reilly Media, Inc
Jeff Gothelf. 2013. Lean UX: Applying lean principles to improve user experience. " O’Reilly Media, Inc. "
2013
-
[17]
Hashini Gunatilake, John Grundy, Rashina Hoda, and Ingo Mueller. 2024. En- ablers and Barriers of Empathy in Software Developer and User Interactions: A Mixed Methods Case Study. ACM Transactions on Software Engineering and Methodology 33, 4 (2024), 109. https://doi.org/10.114...
2024 doi
-
[18]
Hashini Gunatilake, John Grundy, Ingo Mueller, and Rashina Hoda. 2023. Empa- thy Models and Software Engineering — A Preliminary Analysis and Taxonomy. Journal of Systems and Software 203 (2023), 111747. https://doi.org/10.1016/j. jss.2023.111747
2023
-
[19]
Victoria Jackson, Guilherme Vaz Pereira, Rafael Prikladnicki, Andre van der Hoek, Luciane Fortes, Carolina Araújo, Andre Coelho, Ligia Chelli, and Diego Ramos. 2025. Exploring GenAI in Software Development: Insights from a Case Study in a Large Brazilian Company. (2025)
2025
-
[20]
Soon-Gyo Jung, Jisun An, Haewoon Kwak, Moeed Ahmad, Lene Nielsen, and Bernard J Jansen. 2017. Persona generation from aggregated social media data. In Proceedings of the 2017 CHI conference extended abstracts on human factors in computing systems. 1748–1755
2017
-
[21]
Sora Kang, Andreea-Elena Potinteu, and Nadia Said. 2025. ExplainitAI: When do we trust artificial intelligence? The influence of content and explainability in a cross-cultural comparison. In Proceedings of the CHI Conference on Human Factors in Computing Systems (CHI ’25). Ass...
2025 arXiv
-
[22]
Grundy, Tanjila Kanij, Jennifer McIntosh, and Humphrey O
Devi Karolita, John C. Grundy, Tanjila Kanij, Jennifer McIntosh, and Humphrey O. Obie. 2024. Lessons Learned from Persona Usage in Requirements Engineering Practice. 1–10
2024
-
[23]
Devi Karolita, Jennifer McIntosh, Tanjila Kanij, John Grundy, and Humphrey O Obie. 2023. Use of personas in Requirements Engineering: A systematic mapping study. Information and Software Technology (2023), 107264
2023
-
[24]
Raul Leão, Fernando Ayach, Vitor Lameirão, and Awdren Fontão. 2024. A Prompt Engineering-based Process to Build Proto-personas during Lean Inception. In Proceedings of the XXXVIII Brazilian Symposium on Software Engineering. SBC, Porto Alegre, RS, Brasil, 588–594. https://doi....
2024
-
[25]
Jianing Liu, Jia Shi, Jun Xie, Xinyun Zhang, Zichuan Zhang, John Grundy, and Tanjila Kanij. 2022. A curated personas and design guidelines tool for better supporting diverse end-users. In 2022 IEEE 46th Annual Computers, Software, and Applications Conference (COMPSAC). IEEE, 1606–1613
2022
-
[26]
Pengfei Liu, Weizhe Yuan, Jinlan Fu, Zhengbao Jiang, Hiroaki Hayashi, and Graham Neubig. 2023. Pre-train, prompt, and predict: A systematic survey of prompting methods in natural language processing.Comput. Surveys 55, 9 (2023), 1–35
2023
-
[27]
Patricia Losana, John W Castro, Xavier Ferre, Elena Villalba-Mora, and Silvia T Acuña. 2021. A systematic mapping study on integration proposals of the per- sonas technique in agile methodologies. Sensors 21, 18 (2021), 6298
2021
-
[28]
Teixeira, Ariel E
Maylon Macedo, Gabriel V. Teixeira, Ariel E. S. Campos, and Luciana Zaina. 2024. Supporting User-Centered Requirements Elicitation from Lean Personas: A UX Data Visualization-Based Approach. In In Proceedings of the 26th International Conference on Enterprise Information Syste...
2024 doi
-
[29]
Nuno Marques, Rodrigo Rocha Silva, and Jorge Bernardino. 2024. Using Chat- GPT in Software Requirements Engineering: A Comprehensive Review. Future Internet 16, 6 (2024), 180
2024
-
[30]
Jennifer McIntosh and Humphrey O. Obie. 2024. Recommendations for Effective Persona Usage in Requirements Engineering. In 2024 IEEE 32nd International Requirements Engineering Conference (RE). 216–225. https://doi.org/10.1109/ RE60410.2024.00033
2024
-
[31]
Jürgen Münch, Stefan Trieflinger, and Bernd Heisler. 2020. Product Discov- ery - Building the Right Things: Insights from a Grey Literature Review. In 2020 IEEE International Conference on Engineering, Technologyand Innovation (ICE/ITMC). 1–8. https://doi.org/10.1109/ICE/ITMC4...
2020
-
[32]
Anh Nguyen-Duc, Beatriz Cabrero Daniel, Adam Przybylek, Chetan Arora, Dron Khanna, Tomas Herda, Usman Rafiq, Jorge Melegati, Eduardo Guerra, Kai-Kristian Kemell, Mika Saari, Zheying Zhang, Huy Le, Tho Quan, and Pekka Abrahamsson
-
[33]
Landauer
Jakob Nielsen and Thomas K. Landauer. 1993. A Mathematical Model of the Finding of Usability Problems. In Proceedings of the INTERCHI ’93 Conference on Human Factors in Computing Systems. ACM, 206–213. https://doi.org/10. 1145/169059.169166
1993
-
[34]
OpenAI. 2024. Prompt Engineering Guide. https://platform.openai.com/docs/ guides/prompt-engineering. Accessed: 2025-07-03
2024
-
[35]
Eduardo Gouveia Pinheiro, Larissa Albano Lopes, Tayana Uchôa Conte, and Lu- ciana Aparecida Martinez Zaina. 2018. The contribution of non-technical stake- holders on the specification of UX requirements: an experimental study using the proto-persona technique. In Proceedings o...
2018
-
[36]
Wilson Pádua. 2010. Measuring Complexity, Effectiveness and Efficiency in Software Course Projects. In Proceedings of the 32nd ACM/IEEE International Conference on Software Engineering - Volume 1 (ICSE ’10). ACM, 473–482. https://doi.org/10.1145/1806799.1806878
2010
-
[37]
Paul Ralph, Nauman bin Ali, Sebastian Baltes, Domenico Bianculli, Jessica Diaz, Yvonne Dittrich, Neil Ernst, Michael Felderer, Robert Feldt, Antonio Filieri, Breno Bernard Nicolau de França, Carlo Alberto Furia, Greg Gay, Nicolas Gold, Daniel Graziotin, Pinjia He, Rashina Hoda...
2020
-
[38]
John Roberts, Max Baker, and Jane Andrew. 2024. Artificial Intelligence and Qualitative Research: The Promise and Perils of Large Language Model (LLM) ’Assistance’. Critical Perspectives on Accounting 99 (2024), 102722. https://doi. org/10.1016/j.cpa.2024.102722 SBES 2025, Oct...
2024
-
[39]
Per Runeson and Martin Höst. 2009. Guidelines for conducting and reporting case study research in software engineering. Empirical software engineering 14 (2009), 131–164
2009
-
[40]
Daniel Russo. 2024. Navigating the complexity of generative ai adoption in software engineering. ACM Transactions on Software Engineering and Methodology 33, 5 (2024), 1–50
2024
-
[41]
Joni Salminen, Kathleen Wenyun Guan, Soon-Gyo Jung, and Bernard Jansen
-
[42]
Benjamin Saunders, Julius Sim, Tom Kingstone, Shula Baker, Jackie Waterfield, Bernadette Bartlam, Heather Burroughs, and Clare Jinks. 2018. Saturation in qual- itative research: exploring its conceptualization and operationalization. Quality & quantity 52 (2018), 1893–1907
2018
-
[43]
Jaakko Sauvola, Sasu Tarkoma, Mika Klemettinen, Jukka Riekki, and David Doermann. 2024. Future of software development with generative AI.Automated Software Engineering 31, 1 (2024), 26
2024
-
[44]
Gabriel Viana Teixeira and Luciana A. M. Zaina. 2022. Using Lean Personas to the Description of UX-related Requirements: A Study with Software Startup Professionals. In Proceedings of the 24th International Conference on Enterprise Information Systems (ICEIS), Volume 2. SciTePress
2022
-
[45]
Stefan Trieflinger, Dominic Lang, Selina Spies, and Jürgen Münch. 2023. The discovery effort worthiness index: How much product discovery should you do and how can this be integrated into delivery? Information and software technology 157 (2023), 107167
2023
-
[46]
Yi Wang, Chetan Arora, Xiao Liu, Thuong Hoang, Vasudha Malhotra, Ben Cheng, and John Grundy. 2025. Who uses personas in requirements engineering: The practitioners’ perspective. Information and Software Technology 178 (2025), 107609. https://doi.org/10.1016/j.infsof.2024.107609
2025
-
[47]
Jules White, Quchen Fu, Sam Hays, Michael Sandborn, Carlos Olea, Henry Gilbert, Ashraf Elnashar, Jesse Spencer-Smith, and Douglas C Schmidt. 2023. A prompt pattern catalog to enhance prompt engineering with chatgpt. arXiv preprint arXiv:2302.11382 (2023)
2023 arXiv
-
[48]
Lixiang Yan, Vanessa Echeverria, Gloria Milena Fernandez-Nieto, Yueqiao Jin, Zachari Swiecki, Linxuan Zhao, Dragan Gašević, and Roberto Martinez- Maldonado. 2024. Human-AI Collaboration in Thematic Analysis Using Chat- GPT: A User Study and Design Recommendations. In Extended ...
2024
-
[49]
Xishuo Zhang, Lin Liu, Yi Wang, Xiao Liu, Hailong Wang, Chetan Arora, Haichao Liu, Weijia Wang, and Thuong Hoang. 2024. Auto-Generated Personas: Enhanc- ing User-centered Design Practices among University Students. In Extended Abstracts of the CHI Conference on Human Factors i...
2024
-
[50]
Zhang, L
X. Zhang, L. Liu, Y. Wang, X. Liu, H. Wang, A. Ren, and C. Arora. 2023. Per- sonaGen: A tool for generating personas from user feedback. In 2023 IEEE 31st International Requirements Engineering Conference (RE). 353–354. https: //doi.org/10.1109/RE57278.2023.00048
2023
-
[51]
Xishuo Zhang, Lin Liu, Yi Wang, Xiao Liu, Hailong Wang, Anqi Ren, and Chetan Arora. 2023. PersonaGen: A Tool for Generating Personas from User Feedback. In 2023 IEEE 31st International Requirements Engineering Conference (RE). IEEE, 353–354
2023
-
[52]
Yinkun Zhu, Qiwen Liu, and Li Zhao. 2025. Exploring the impact of generative artificial intelligence on students’ learning outcomes: a meta-analysis. Education and Information Technologies (2025). https://doi.org/10.1007/s10639-025-13420- z
2025 doi
-
[2022]
In Proceedings of the 2022 CHI Conference on Human Factors in Computing Systems
Use cases for design personas: A systematic review and new frontiers. In Proceedings of the 2022 CHI Conference on Human Factors in Computing Systems. 1–21
2022
- [2023]
-
[2025]
In Proceedings of the CHI Conference on Human Factors in Computing Systems (CHI ’25)
Coordination Mechanisms in AI Development: Practitioner Experiences on Integrating UX Activities. In Proceedings of the CHI Conference on Human Factors in Computing Systems (CHI ’25). Association for Computing Machinery, New York, NY, USA, 1–14. https://doi.org/10.1145/3706598...
Reviewed August 6, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.