Pith. sign in

REVIEW 2 major objections 5 minor 82 references

A Case Study Investigating the Role of Generative AI in Quality Evaluations of Epics in Agile Software Development

T0 review · 2 major / 5 minor · reviewed 2026-08-15 · deepseek-v4-flash

Pith's one-line read LLM-based evaluations of agile epics are a viable new use of AI, an industry study of 17 product managers argues.

desk verdict Solid qualitative case study whose 'viable today' claim overreaches the concept-test evidence; rubric and practitioner insights are the substance. read the letter →

arxiv 2505.07664 v1 pith:IOUCQFT7 submitted 2025-05-12 cs.SE cs.AIcs.HC

classification cs.SEcs.AIcs.HC
keywords agilesoftwaredevelopmentepicsqualityevaluationlargelanguagemodelsgenerativeAILLM-as-a-judgeproductmanagersrequirementsengineering
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper argues that large language models can act as quality evaluators for agile epics, the high-level requirement documents product managers use to align stakeholders, and that product managers would welcome such evaluations in practice. Through a user study with 17 product managers, it tests a rubric of eight quality criteria alongside a prototype tool, Epic Evaluator, that rates epic elements and suggests improvements. High levels of satisfaction with the evaluations lead the authors to claim that agile epics are a new, viable application of AI evaluation. The paper also identifies barriers, including lack of domain knowledge, rigidity of rubrics, and need for workflow integration, that must be addressed before such tools are adopted.

What carries the argument

The central objects are an eight-element rubric defining quality criteria for agile epic elements, including title, problem statement, product outcome and instrumentation, user stories, requirements, assumptions, non-functional requirements, and out-of-scope, and a prototype tool called Epic Evaluator that uses prompt-based evaluation by an LLM to detect which elements are present, rate each on a High/Medium/Low scale, and generate explanations and recommendations. The rubric operationalizes quality so that an LLM can judge it, the prompts carry the evaluation logic, and the concept-testing interviews with product managers supply the evidence of viability.

What would settle it

Give a group of product managers a working Epic Evaluator integrated into their agile management tool for three months while a matched control group continues without it; if the intervention group shows no measurable improvement in epic quality scores, no reduction in downstream churn or delays, or no increase in actual usage, the viability claim is unsupported.

Watch

Extended reading notes

Core claim

The central claim is that LLM evaluations of epics are viable and can provide value today. Product managers in the study largely agreed with the tool's scores and recommendations (15 of 17), expressed enthusiasm about using such evaluations to augment peer feedback, accelerate epic improvement, and support managers in triaging large numbers of epics, and rated their trust in the LLM evaluation at an average of 3.5 on a 5-point scale. The authors conclude that carefully designed LLM-based tooling can help product managers craft higher-quality epics and thereby reduce downstream churn, communication breakdowns, delays, and cost overruns.

Load-bearing premise

The claim rests on participants' self-reported satisfaction with static, screenshot-based evaluations, assuming that this predicts real-world adoption and improved epic quality in daily workflows.

Editorial extensions

If this is right

  • LLM-based epic evaluation could be injected at four common milestones: after the first draft, before sharing with stakeholders, after stakeholder iteration, and before handoff to development.
  • Product managers want actionable recommendations over scores, suggesting that evaluator and authoring tools will tend to merge.
  • Managers see bulk evaluation of team epics as a way to spot epics needing attention, and as a training aid for novice product managers.
  • To be adopted, such tools must integrate tightly into existing agile management tools so users do not need to copy and paste epic text.
  • Adding domain and stakeholder knowledge, for example through retrieval-augmented generation, is the key next step to improve the specificity and actionability of evaluations.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If viability holds under real adoption, organizations could standardize epic quality across teams without mandating rigid templates, using a core rubric plus team-customizable extensions, a path the paper sketches but does not test.
  • The finding that managers would not use evaluations for performance assessment, while non-managers fear score-based reporting, suggests that governance of AI evaluation must separate quality improvement from personnel evaluation, a question the paper does not address.
  • A testable extension is a randomized field trial in which product managers receive an LLM evaluator integrated into their agile tool or no evaluator, measuring revision cycles and downstream churn rather than self-reported satisfaction.
  • The average trust rating of 3.5 on a 5-point scale and concerns about missing context imply that LLM epic evaluators will likely settle into a human-in-the-loop advisory role rather than an automated gatekeeper role.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

2 major / 5 minor

Summary. This industry case study investigates whether large language models (LLMs) can evaluate the quality of agile epics. The authors developed an eight-element rubric for epic quality, built a prototype tool called Epic Evaluator that uses LLMs to rate elements and provide recommendations, and conducted semi-structured interviews with 17 product managers at a global company. During the interviews they concept-tested the rubric and static screenshots of LLM evaluations of the participants' own epics. The paper reports that participants found the rubric useful, expressed enthusiasm for the tool, identified several concrete uses (e.g., augmenting peer feedback, bulk evaluation for managers, training novices), and raised barriers including lack of domain knowledge, limited actionability, and integration friction. The central claim, stated in the abstract and in Section 5, is that LLM evaluations of epics are 'viable and can provide value today.' The paper also proposes design implications such as stage-aware evaluation, customizable rubrics, and tight integration with existing agile tools.

Significance. If taken as a preliminary qualitative study, this paper makes a useful contribution to the emerging literature on LLM-as-a-Judge for domain-specific text. Its strengths include a transparently described rubric development process with subject matter expert involvement, inclusion of the actual prompts in the appendix, and rich verbatim quotes from practitioners that give insight into real-world epic creation practices. The qualitative findings on perceived value, adoption barriers, and the need for domain knowledge are credible and potentially transferable to other agile organizations. However, the paper's central claim of current 'viability' goes beyond what the evidence can support: the study measures self-reported satisfaction and intended use from a concept test with static screenshots, not actual tool adoption, evaluation accuracy, or downstream improvements in epic quality. The contributions are therefore best framed as perceived value and design implications, not as demonstrated viability.

major comments (2)
  1. [Abstract; Section 5; Section 4.3.2; Section 6] The central claim that LLM evaluations of epics are 'viable and can provide value today' is not supported by the evidence presented. The study used concept testing with static screenshots; participants never interacted directly with the Epic Evaluator (Section 3.3 and acknowledged in Section 6). There is no comparison against expert human evaluations, no baseline condition, and no measurement of whether acting on the LLM recommendations improves epic quality. Section 4.3.2 reports that 15/17 participants 'agreed with aspects' of the score or recommendations, but agreement with aspects of a concept test is not evidence of evaluation accuracy or of downstream value, especially given that average self-reported trust was only 3.5 on a 1-5 scale and that Section 4.3.3 lists substantial barriers including lack of domain knowledge, limited actionability, and workflow integration concerns. The strongest defensible conclusion is that product managers perceive potential value in such a tool and desire it; the claim of current viability should be softened to a claim of perceived potential or preliminary feasibility.
  2. [Section 1 Contributions; Section 3.1; Section 4.2.2] The paper states in the contributions that it introduces and validates a rubric, but the rubric validation consists solely of think-aloud feedback from 17 product managers. No inter-rater reliability is reported, the rubric was not independently applied by multiple raters to a set of epics, and no criterion validity against expert quality judgments or epic outcomes is provided. Furthermore, the participants identified missing elements (e.g., acceptance criteria), rigidity concerns, and worries about being penalized for missing elements that their teams do not use (Section 4.2.2). The term 'validate' therefore overstates the evidence; 'elicit practitioner feedback on a rubric' or 'preliminary evaluation of the rubric' would be more accurate.
minor comments (5)
  1. [Appendix A.3] There is a typo in the evaluation prompt: 'Please rovide a detailed EXPLANATION' should read 'Please provide a detailed EXPLANATION.'
  2. [Table 3 (Appendix A.1)] The example for the Non-Functional Requirements row is incomplete; it ends with 'The system should respond to password reset requests within 1 second for 95' and needs the concluding percentage or unit.
  3. [Section 3.2] The description of model selection and prompt iteration is qualitative. Please specify the temperature setting, the number of evaluation runs used to assess stability, and how 'stable' was operationalized, since LLM outputs are sensitive to sampling parameters.
  4. [Section 4.3.2] The statement that 15/17 participants 'agreed with aspects of the score or recommendations' is vague; please clarify what counted as agreement (e.g., agreement with the rating, with the explanation, with the recommendation, or any combination) and how this was coded from the interview data.
  5. [References [60]] Reference [60] is cited as 'SAFe. User Stories' but the URL points to businessmap.io; please verify that the citation and URL correspond to the intended source.

Circularity Check

0 steps flagged · score 0.0 of 10

No circular derivation: the viability claim rests on human participant feedback, not on the LLM's own outputs or fitted parameters.

full rationale

This paper is a qualitative industry case study rather than a formal derivation chain. The authors developed a rubric from company templates, subject-matter experts, and LLM input; built an LLM tool around that rubric; and gathered interview and concept-test data from 17 product managers. The central claim that LLM-based epic evaluations are viable is supported by participant satisfaction (Section 4.3.2: '15/17 participants agreed with aspects of the score or recommendations') and by participants' stated intended uses and perceived barriers. There is no fitted parameter later renamed as a prediction, no equation in which an output equals an input by construction, and no uniqueness theorem imported from the authors' prior work. The self-citations (EvalAssist, EvaluLLM, and related LLM-as-a-judge papers) appear in related-work and design-inspiration contexts and are not load-bearing: the authors explicitly state they are 'not leveraging any existing LLM-as-a-Judge tools,' and the cited prior work is not used to justify the viability conclusion. The paper itself flags a genuine limitation in Section 6 — participants viewed static screenshots and did not interact directly with the tool — which is a threat to the strength of the empirical claim, but it is a validity limitation, not circularity. The conclusion could be too strong given the evidence, but it is not presupposed by definition or by self-referential evidence.

Assumptions & free parameters 0 free parameters · 4 assumptions · 0 invented entities

The paper contains no numeric free parameters or fitted constants. It introduces a rubric and a prototype tool, but these are artifacts, not scientific postulates. The key assumptions are domain-level: that LLM text evaluation is reliable, that self-report is a valid proxy, and that the company-specific template defines appropriate quality standards.

assumptions (4)
  • domain assumption Agile epics are a key artifact, and poorly defined epics lead to churn, delivery delays, and cost overruns.
    Stated in the introduction, citing prior literature; the study does not measure this causal link directly.
  • domain assumption LLMs can reliably evaluate the quality of natural-language text against a rubric.
    Background cites LLM-as-a-judge literature; the study does not compare LLM ratings against expert human ratings of epic quality.
  • domain assumption Participant self-report of satisfaction and perceived usefulness is a valid measure of tool viability.
    The viability conclusion is based on interview responses and descriptive statistics, with no objective outcome metric such as improved epic quality over time.
  • domain assumption The company's agile template and two subject matter experts define valid quality criteria for epics.
    Rubric development in Section 3.1 relies on internal template and SME iteration; external validity beyond this organization is acknowledged as limited.

how reviews work

0 comments
Cite this review

Pith. "Pith review of A Case Study Investigating the Role of Generative AI in Quality Evaluations of Epics in Agile Software Development." pith.science (2026). https://pith.science/paper/IOUCQFT7

@misc{pith2026250507664,
  author       = {Pith},
  title        = {Pith review of: A Case Study Investigating the Role of Generative AI in Quality Evaluations of Epics in Agile Software Development},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/IOUCQFT7}},
  note         = {Machine review of arXiv:2505.07664}
}
read the original abstract

The broad availability of generative AI offers new opportunities to support various work domains, including agile software development. Agile epics are a key artifact for product managers to communicate requirements to stakeholders. However, in practice, they are often poorly defined, leading to churn, delivery delays, and cost overruns. In this industry case study, we investigate opportunities for large language models (LLMs) to evaluate agile epic quality in a global company. Results from a user study with 17 product managers indicate how LLM evaluations could be integrated into their work practices, including perceived values and usage in improving their epics. High levels of satisfaction indicate that agile epics are a new, viable application of AI evaluations. However, our findings also outline challenges, limitations, and adoption barriers that can inform both practitioners and researchers on the integration of such evaluations into future agile work practices.

Figures

Figures reproduced from arXiv: 2505.07664 by the authors.

Figure 1
Figure 1. The experimental Epic Evaluator tool. The full text of an agile epic is pasted into the input box at the top. When [PITH_FULL_IMAGE:figures/full_fig_p005_1.png] view at source ↗

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

82 extracted references · 45 canonical work pages

  1. [1]

    M. O. Ahmad, J. Markkula, M. Oivo, and P. Kuvaja. 2018. Kanban in software development: A systematic literature review. Journal of Systems and Software 137 (2018), 96–113. doi:10.1016/j.jss.2017.11.049

  2. [2]

    Albertus M. M. Antaputra and Lim Tek Yong. 2024. A Systematic Literature Review on Agile Requirements Engineering and Artificial Intelligence. In 2024 22nd International Conference on ICT and Knowledge Engineering (ICT&KE) . 1–6. doi:10.1109/ICTKE62841.2024.10787176

  3. [3]

    2024.Advancing Require- ments Engineering Through Generative AI: Assessing the Role of LLMs

    Chetan Arora, John Grundy, and Mohamed Abdelrazek. 2024.Advancing Require- ments Engineering Through Generative AI: Assessing the Role of LLMs . Springer Nature Switzerland, Cham, 129–148. doi:10.1007/978-3-031-55642-5_6

  4. [4]

    Johnson, Martin Santil- lan Cooper, Elizabeth M

    Zahra Ashktorab, Michael Desmond, Qian Pan, James M. Johnson, Martin Santil- lan Cooper, Elizabeth M. Daly, Rahul Nair, Tejaswini Pedapati, Swapnaja Achintal- war, and Werner Geyer. 2024. Aligning Human and LLM Judgments: Insights from EvalAssist on Task-Specific Evaluations and AI-assisted Assessment Strategy Preferences. arXiv:2410.00873 [cs.HC] https:/...

  5. [5]

    Atlassian. 2025. Epics in Agile Project Management. https://www.atlassian.com/ agile/project-management/epics Accessed: 2025-02-02

  6. [6]

    Anas BAHI, Jihane GHARIB, and Youssef GAHI. 2024. Integrating Generative AI for Advancing Agile Software Development and Mitigating Project Management Challenges. International Journal of Advanced Computer Science and Applications 15, 3 (2024). doi:10.14569/IJACSA.2024.0150306

  7. [7]

    Martin, Steve Mellor, Ken Schwaber, Jeff Sutherland, and Dave Thomas

    Kent Beck, Mike Beedle, Arie van Bennekum, Alistair Cockburn, Ward Cun- ningham, Martin Fowler, James Grenning, Jim Highsmith, Andrew Hunt, Ron Jeffries, Jon Kern, Brian Marick, Robert C. Martin, Steve Mellor, Ken Schwaber, Jeff Sutherland, and Dave Thomas. 2001. Manifesto for Agile Software Development. https://agilemanifesto.org/. Accessed: 2025-01-24

  8. [8]

    Woubshet Behutiye, Pilar Rodríguez, and Markku Oivo. 2022. Quality requirement documentation guidelines for agile software development. IEEE Access 10 (2022), 70154–70173

Show all 82 references
  1. [9]

    Woubshet Behutiye, Pertti Seppänen, Pilar Rodríguez, and Markku Oivo. 2020. Documentation of quality requirements in agile software development. In Pro- ceedings of the 24th International Conference on Evaluation and Assessment in Software Engineering. 250–259

  2. [10]

    Niels Bik, Garm Lucassen, and Sjaak Brinkkemper. 2017. A reference method for user story requirements in agile systems development. In 2017 IEEE 25th international requirements engineering conference workshops (REW) . IEEE, 292– 298

  3. [11]

    Boehm and R

    B. Boehm and R. Turner. 2004. Balancing Agility and Discipline: A Guide for the Perplexed. Addison-Wesley

  4. [12]

    Michelle Brachman, Amina El-Ashry, Casey Dugan, and Werner Geyer. 2024. How Knowledge Workers Use and Want to Use LLMs in an Enterprise Context. In Extended Abstracts of the CHI Conference on Human Factors in Computing Systems (Honolulu, HI, USA) (CHI EA ’24). Association for ...

  5. [13]

    Jill Burstein. 2003. The E-rater ® Scoring Engine: Automated Essay Scoring with Natural Language Processing. In Automated Essay Scoring: A Cross-disciplinary Perspective, Mark D. Shermis and Jill Burstein (Eds.). Lawrence Erlbaum Asso- ciates, Inc., 113–121

  6. [14]

    Birgitta Böckeler, Adhavan KP, Sayantan Mukhopadhyay, Radhika Sivadass, and Jaiganesh B. 2024. Using AI for requirements analysis: A case study. https://www.thoughtworks.com/en-us/insights/blog/generative-ai/using- ai-requirements-analysis-case-study. Accessed: 2025-01-10

  7. [15]

    Alexia Cambon, Brent Hecht, Ben Edelman, Donald Ngwe, Sonia Jaffe, Amy Heger, Mihaela Vorvoreanu, Sida Peng, Jake Hofman, Alex Farach, et al . 2023. Early LLM-based Tools for Enterprise Information Workers Likely Provide Meaningful Boosts to Productivity. Microsoft Research. M...

  8. [16]

    Andres Campero, Michelle Vaccaro, Jaeyoon Song, Haoran Wen, Abdullah Almaa- touq, and Thomas W. Malone. 2022. A Test for Evaluating Performance in Human- Computer Systems. arXiv:2206.12390 [cs.HC] https://arxiv.org/abs/2206.12390

  9. [17]

    Peter W Cardon, Kristen Getchell, Stephen Carradini, Carolin Fleischmann, and James Stapp. 2023. Generative AI in the workplace: Employee perspectives of ChatGPT benefits and organizational policies. (2023)

  10. [18]

    Haowei Cheng, Jati H Husen, Sien Reeve Peralta, Bowen Jiang, Nobukazu Yosh- ioka, Naoyasu Ubayashi, and Hironori Washizaki. 2024. Generative AI for Requirements Engineering: A Systematic Literature Review. arXiv preprint arXiv:2409.06741 (2024)

  11. [19]

    Cheng-Han Chiang, Wei-Chih Chen, Chun-Yi Kuan, Chienchou Yang, and Hung- yi Lee. 2024. Large language model as an assignment evaluator: Insights, feedback, and challenges in a 1000+ student course. arXiv preprint arXiv:2407.05216 (2024)

  12. [20]

    Cheng-Han Chiang and Hung-yi Lee. 2023. Can large language models be an alternative to human evaluations? arXiv preprint arXiv:2305.01937 (2023)

  13. [21]

    Victoria Clarke and Virginia Braun and. 2017. Thematic analysis. The Journal of Positive Psychology 12, 3 (2017), 297–298. doi:10.1080/17439760.2016.1262613 arXiv:https://doi.org/10.1080/17439760.2016.1262613

  14. [22]

    Fabiano Dalpiaz and Sjaak Brinkkemper. 2021. Agile Requirements Engineering: From User Stories to Software Architectures. In 2021 IEEE 29th International Requirements Engineering Conference (RE) . 504–505. doi:10.1109/RE51729.2021. 00076

  15. [23]

    Hoa Khanh Dam, Truyen Tran, John Grundy, Aditya Ghose, and Yasutaka Kamei

  16. [24]

    Michael Desmond, Zahra Ashktorab, Qian Pan, Casey Dugan, and James M Johnson. 2024. EvaluLLM: LLM assisted evaluation of generative outputs. In Companion Proceedings of the 29th International Conference on Intelligent User Interfaces. 30–32

  17. [25]

    Dingsøyr, S

    T. Dingsøyr, S. Nerur, V. Balijepally, and N. B. Moe. 2012. A decade of agile methodologies: Towards explaining agile software development. Journal of Systems and Software 85, 6 (2012), 1213–1221. doi:10.1016/j.jss.2012.02.033

  18. [26]

    Yann Dubois, Chen Xuechen Li, Rohan Taori, Tianyi Zhang, Ishaan Gulrajani, Jimmy Ba, Carlos Guestrin, Percy S Liang, and Tatsunori B Hashimoto. 2024. Al- pacafarm: A simulation framework for methods that learn from human feedback. Advances in Neural Information Processing Syst...

  19. [27]

    António MS Ferreira, Alberto Rodrigues da Silva, and Ana CR Paiva. 2022. To- wards the Art of Writing Agile Requirements with User Stories, Acceptance CHIWORK ’25, June 23–25, 2025, Amsterdam, Netherlands Geyer et al. Criteria, and Related Constructs.. In ENASE. 477–484

  20. [28]

    International Organization for Standardization. 2018. ISO/IEC/IEEE Interna- tional Standard - Systems and software engineering – Life cycle processes – Requirements engineering. 104 pages

  21. [29]

    Jie Gao, Simret Araya Gebreegziabher, Kenny Tsu Wei Choo, Toby Jia-Jun Li, Simon Tangi Perrault, and Thomas W. Malone. 2024. A Taxonomy for Human- LLM Interaction Modes: An Initial Exploration. In Extended Abstracts of the CHI Conference on Human Factors in Computing Systems . 1–11

  22. [30]

    Geeks For Geeks. 2025. Epics. https://www.geeksforgeeks.org/what-are-epics- user-stories-in-agile/ Accessed: 2025-02-05

  23. [31]

    Gradio Team. 2025. Gradio: Build Machine Learning Apps Quickly and Easily. https://www.gradio.app/ Accessed: 2025-01-24

  24. [32]

    Hussain, E

    A. Hussain, E. O. Mkpojiogu, and F. M. Kamal. 2016. The role of requirements in the success or failure of software projects. International Review of Management and Marketing 6, 7 (2016), 306–311

  25. [33]

    Bill Iuso. 1975. Concept testing: an appropriate approach. Journal of Marketing research 12, 2 (1975), 228–231

  26. [34]

    M. I. Kamata and T. Tamai. 2007. How does requirements quality relate to project success or failure?. In15th IEEE International Requirements Engineering Conference (RE 2007). IEEE, 69–78. doi:10.1109/RE.2007.45

  27. [35]

    Rashidah Kasauli, Eric Knauss, Jennifer Horkoff, Grischa Liebel, and Fran- cisco Gomes de Oliveira Neto. 2021. Requirements engineering challenges and practices in large-scale agile system development.Journal of Systems and Software 172 (2021), 110851

  28. [36]

    Rashidah Kasauli, Grischa Liebel, Eric Knauss, Swathi Gopakumar, and Benjamin Kanagwa. 2017. Requirements Engineering Challenges in Large-Scale Agile System Development. In 2017 IEEE 25th International Requirements Engineering Conference (RE). 352–361. doi:10.1109/RE.2017.60

  29. [37]

    Seungone Kim, Jamin Shin, Yejin Cho, Joel Jang, Shayne Longpre, Hwaran Lee, Sangdoo Yun, Seongjin Shin, Sungdong Kim, James Thorne, et al . 2023. Prometheus: Inducing fine-grained evaluation capability in language models. In The Twelfth International Conference on Learning Rep...

  30. [38]

    Tae Soo Kim, Yoonjoo Lee, Jamin Shin, Young-Ho Kim, and Juho Kim. 2024. Evallm: Interactive evaluation of large language model prompts on user-defined criteria. In Proceedings of the CHI Conference on Human Factors in Computing Systems. 1–21

  31. [39]

    Abhishek Kumar, Sonia Haiduc, Partha Pratim Das, and Partha Pratim Chakrabarti. 2024. LLMs as Evaluators: A Novel Approach to Evaluate Bug Report Summarization. arXiv preprint arXiv:2409.00630 (2024)

  32. [40]

    Patrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni, Vladimir Karpukhin, Naman Goyal, Heinrich Küttler, Mike Lewis, Wen-tau Yih, Tim Rocktäschel, et al. 2020. Retrieval-augmented generation for knowledge-intensive nlp tasks. Advances in Neural Information Processing S...

  33. [41]

    Sebastian Lubos, Alexander Felfernig, Thi Ngoc Trang Tran, Damian Garber, Merfat El Mansi, Seda Polat Erdeniz, and Viet-Man Le. 2024. Leveraging llms for the quality assurance of software requirements. In 2024 IEEE 32nd International Requirements Engineering Conference (RE) . ...

  34. [42]

    Simone Luchini, Nadine T Maliakkal, Paul V DiStefano, John D Patterson, Roger Beaty, and Roni Reiter-Palmon. 2023. Automatic Scoring of Creative Problem- Solving with Large Language Models: A Comparison of Originality and Quality Ratings. (2023)

  35. [43]

    Shengjie Ma, Chong Chen, Qi Chu, and Jiaxin Mao. 2024. Leveraging Large Language Models for Relevance Judgments in Legal Case Retrieval.arXiv preprint arXiv:2403.18405 (2024)

  36. [44]

    Markel, Steven G

    Julia M. Markel, Steven G. Opferman, James A. Landay, and Chris Piech. 2023. GPTeach: Interactive TA Training with GPT-based Students. In Proceedings of the Tenth ACM Conference on Learning @ Scale (Copenhagen, Denmark) (L@S ’23). Association for Computing Machinery, New York,...

  37. [45]

    Nuno Marques, Rodrigo Rocha Silva, and Jorge Bernardino. 2024. Using ChatGPT in Software Requirements Engineering: A Comprehensive Review.Future Internet 16, 6 (2024), 180

  38. [46]

    Mio Mattila. 2023. Assisting product line thinking and information sharing using a real-time visual tracking model for agile epics. (2023)

  39. [47]

    Farias, Manoel Mendonça, Hen- rique Frota Soares, Marcos Kalinowski, and Rodrigo Oliveira Spínola

    Thiago Souto Mendes, Mário André de F. Farias, Manoel Mendonça, Hen- rique Frota Soares, Marcos Kalinowski, and Rodrigo Oliveira Spínola. 2016. Im- pacts of agile requirements documentation debt on software projects: a retrospec- tive study. InProceedings of the 31st annual AC...

  40. [48]

    Shahraz Nasir, Eduardo Guerra, Luciana Zaina, and Jorge Melegati. 2023. An Exploratory Study About Non-functional Requirements Documentation Practices in Agile Teams. In Proceedings of the 38th ACM/SIGAPP Symposium on Applied Computing. 1009–1017

  41. [49]

    Sofia Nilsson and Emilia Palm. 2024. Optimization of Portfolio Management How to prioritize work within different departments and portfolios. (2024)

  42. [50]

    Johannes J Norheim, Eric Rebentisch, Dekai Xiao, Lorenz Draeger, Alain Kerbrat, and Olivier L de Weck. 2024. Challenges in applying large language models to requirements engineering tasks. Design Science 10 (2024), e16

  43. [51]

    Elastic N.V. 2025. Elasticsearch - The Distributed Search and Analytics Engine. https://www.elastic.co/elasticsearch/. Accessed: 2025-02-07

  44. [52]

    Johanna Olander and Johanna Qvist. 2023. Agile Planning Activities and Team Characteristics for On-time Delivery in Software Development Teams: A case study at Ericsson

  45. [53]

    Oswal, Harshil T

    Jay U. Oswal, Harshil T. Kanakia, and Devvrat Suktel. 2024. Transforming Soft- ware Requirements into User Stories with GPT-3.5 -: An AI-Powered Approach. In 2024 2nd International Conference on Intelligent Data Communication Technologies and Internet of Things (IDCIoT). 913–9...

  46. [54]

    Qian Pan, Zahra Ashktorab, Michael Desmond, Martin Santillan Cooper, James Johnson, Rahul Nair, Elizabeth Daly, and Werner Geyer. 2024. Human-Centered Design Recommendations for LLM-as-a-judge. arXiv preprint arXiv:2407.03479 (2024)

  47. [55]

    Krishna Ronanki, Christian Berger, and Jennifer Horkoff. 2023. Investigating ChatGPT’s potential to assist in requirements elicitation processes. In 2023 49th Euromicro Conference on Software Engineering and Advanced Applications (SEAA) . IEEE, 354–361

  48. [56]

    Paul Röttger, Valentin Hofmann, Valentina Pyatkin, Musashi Hinck, Hannah Kirk, Hinrich Schuetze, and Dirk Hovy. 2024. Political Compass or Spinning Arrow? Towards More Meaningful Evaluations for Values and Opinions in Large Language Models. In Proceedings of the 62nd Annual Me...

  49. [57]

    W. W. Royce. 1970. Managing the development of large software systems. In Proceedings of IEEE WESCON . 1–9

  50. [58]

    Cheol Ryu, Seolhwa Lee, Subeen Pang, Chanyeol Choi, Hojun Choi, Myeonggee Min, and Jy-Yong Sohn. 2023. Retrieval-based Evaluation for LLMs: A Case Study in Korean Legal QA. In Proceedings of the Natural Legal Language Processing Workshop 2023. 132–137

  51. [59]

    SAFe. 2025. NFRs. https://scaledagileframework.com/nonfunctional- requirements Accessed: 2025-02-05

  52. [60]

    SAFe. 2025. User Stories. https://businessmap.io/agile/project-management/user- stories Accessed: 2025-02-05

  53. [61]

    Malik Abdul Sami, Muhammad Waseem, Zheying Zhang, Zeeshan Rasheed, Kari Systä, and Pekka Abrahamsson. 2024. AI based Multiagent Approach for Requirements Elicitation and Analysis. arXiv preprint arXiv:2409.00038 (2024)

  54. [62]

    D Schon. 1983. The reflective practitioner. How professionals think in action. Temple Smith, London

  55. [63]

    Schwaber and J

    K. Schwaber and J. Sutherland. 2020. The Scrum Guide. Scrum Alliance. https: //scrumguides.org/

  56. [64]

    Scrum Alliance. 2025. Understanding Epics. https://resources.scrumalliance.org/ Article/epic-agile Accessed: 2025-02-05

  57. [65]

    Orit Shaer, Angelora Cooper, Osnat Mokryn, Andrew L Kun, and Hagit Ben Shoshan. 2024. AI-Augmented Brainwriting: Investigating the use of LLMs in group ideation. In Proceedings of the 2024 CHI Conference on Human Factors in Computing Systems (Honolulu, HI, USA)(CHI ’24). Assoc...

  58. [66]

    Vaishali Siddeshwar, Sanaa Alwidian, and Masoud Makrehchi. 2024. A Systematic Review of AI-Enabled Frameworks in Requirements Elicitation. IEEE Access 12 (2024), 154310–154336. doi:10.1109/ACCESS.2024.3475293

  59. [67]

    Henrique F Soares, Nicolli SR Alves, Thiago S Mendes, Manoel Mendonça, and Rodrigo O Spinola. 2015. Investigating the link between user stories and doc- umentation debt on software projects. In 2015 12th International Conference on Information Technology-New Generations. IEEE, 385–390

  60. [68]

    Yishen Song, Qianta Zhu, Huaibo Wang, and Qinhua Zheng. 2024. Automated Essay Scoring and Revising Based on Open-Source Large Language Models. IEEE Transactions on Learning Technologies (2024)

  61. [69]

    Annalisa Szymanski, Noah Ziems, Heather A Eicher-Miller, Toby Jia-Jun Li, Meng Jiang, and Ronald A Metoyer. 2024. Limitations of the LLM-as-a-Judge Approach for Evaluating LLM Outputs in Expert Knowledge Tasks.arXiv preprint arXiv:2410.20266 (2024)

  62. [70]

    Tech Target. 2025. Agile Software Development. https://www.techtarget.com/ searchsoftwarequality/definition/agile-software-development Accessed: 2025- 02-05

  63. [71]

    Muhammad Aminu Umar and Kevin Lano. 2023. Automated Requirements Engi- neering in Agile Development: A Practitioners Survey. In 2023 3rd International Conference on Electrical, Computer, Communications and Mechatronics Engineering (ICECCME). 1–7. doi:10.1109/ICECCME57830.2023.10253030

  64. [72]

    Nalin Wadhwa, Jui Pradhan, Atharv Sonwane, Surya Prakash Sahu, Nagarajan Natarajan, Aditya Kanade, Suresh Parthasarathy, and Sriram Rajamani. 2023. Frustrated with code quality issues? llms can help!arXiv preprint arXiv:2309.12938 (2023)

  65. [73]

    Kelly B Wagman, Matthew T Dearing, and Marshini Chetty. 2025. Generative AI Uses and Risks for Knowledge Workers in a Science Organization. arXiv preprint arXiv:2501.16577 (2025)

  66. [74]

    Daly, Qian Pan, Martín Santillán Cooper, James M

    Nico Wagner, Michael Desmond, Rahul Nair, Zahra Ashktorab, Elizabeth M. Daly, Qian Pan, Martín Santillán Cooper, James M. Johnson, and Werner A Case Study Investigating the Role of Generative AI in Quality Evaluations of Epics in Agile Software Development CHIWORK ’25, June 23...

  67. [75]

    Chen Wang, Jin Zhao, and Jiaqi Gong. 2024. A Survey on Large Language Models from Concept to Implementation. arXiv:2403.18969 [cs.CL] https://arxiv.org/ abs/2403.18969

  68. [76]

    Z Wang, J Li, G Li, and Z Jin. 2023. ChatCoder: Chat-based Refine Requirement Improves LLMs’ Code Generation. arXiv 2023. arXiv preprint arXiv:2311.00272 (2023)

  69. [77]

    Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Ed Chi, Quoc Le, and Denny Zhou. 2022. Chain of Thought Prompting Elicits Reasoning in Large Language Models. arXiv preprint arXiv:2201.11903 (2022). https: //arxiv.org/abs/2201.11903

  70. [78]

    Young Urban Project. 2025. Epic in Agile. https://www.youngurbanproject.com/ what-is-an-epic-in-agile/ Accessed: 2025-02-05

  71. [79]

    Wayne Xin Zhao, Kun Zhou, Junyi Li, Tianyi Tang, Xiaolei Wang, Yupeng Hou, Yingqian Min, Beichen Zhang, Junjie Zhang, Zican Dong, Yifan Du, Chen Yang, Yushuo Chen, Zhipeng Chen, Jinhao Jiang, Ruiyang Ren, Yifan Li, Xinyu Tang, Zikang Liu, Peiyu Liu, Jian-Yun Nie, and Ji-Rong W...

  72. [80]

    Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric Xing, et al. 2023. Judging llm-as-a-judge with mt-bench and chatbot arena. Advances in Neural Information Processing Systems 36 (2023), 46595–46623

  73. [81]

    Enable Password Reset for Registered Users to Regain Account Access

    Zilliz. 2025. Milvus: Open-source Vector Database for Scalable AI Applications. https://milvus.io/. Accessed: 2025-02-07. CHIWORK ’25, June 23–25, 2025, Amsterdam, Netherlands Geyer et al. A APPENDIX A.1 Rubric Table 3: The rubric developed with subject matter experts. The rub...

  74. [2019]

    In2019 IEEE/ACM 41st international conference on software engineering: new ideas and emerging results (ICSE-NIER)

    Towards effective AI-powered agile project management. In2019 IEEE/ACM 41st international conference on software engineering: new ideas and emerging results (ICSE-NIER). IEEE, 41–44

Pith tools

Reviewed August 15, 2026 · model on record in the stance chip above.