{"id":"8110e439-6e77-416e-a806-41a76d9bd16d","arxiv_id":"2505.07664","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"An LLM-based evaluator for agile epics was built from a new rubric and tested with 17 product managers, who found it useful but limited by lack of domain knowledge and rigid scoring.","lead":"This paper tests whether large language models can grade the quality of 'epics', big blocks of work that product managers write for agile software development. Interviews with 17 product managers at one company suggest the idea is welcome, but the tool needs domain knowledge and flexible, stage-aware scoring to be useful.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The claim that LLM epic evaluations are 'viable and can provide value today' rests entirely on self-reported satisfaction with static screenshots, with no behavioral or outcome evidence linking the tool to improved epic quality.","rationale":"The reader's weakest_assumption identifies self-reported satisfaction with static screenshots as insufficient evidence for real-world viability, and my analysis converges on the same load-bearing point. The paper's qualitative findings are valuable and the authors are transparent about the concept-test design; the rubric development and the diversity of reported practices are genuine contributions. However, the abstract and Section 5 assert a stronger conclusion than the evidence supports. There is no baseline (e.g., human expert evaluation) against which to judge the LLM's ratings, and no metric of whether the evaluations improve downstream outcomes such as churn or delays. The participants' positive reactions show interest and perceived usefulness, but perceived usefulness alone does not establish that the tool is viable today. The concern is not that the study is invalid as a qualitative exploration; it is that the headline claim overstates the current evidence. The proposed deployment study would directly test whether the tool is used voluntarily and whether it improves epic quality as judged by blinded experts, which would either substantiate or refute the viability claim. Since the reader already reached a CONDITIONAL verdict with essentially the same caveat, my stress-test does not change the verdict.","tokens_in":25323,"tokens_out":3036,"duration_ms":32637,"concrete_test":"Run a small deployment study: give 8–10 product managers access to the actual Epic Evaluator for their real epics over 4–6 weeks. Track voluntary usage logs; have two expert PMs, blind to condition, rate the quality of each participant's epics before and after using the tool, using the same rubric; and compare the LLM's ratings on a sample of epics to expert ratings on the same rubric (e.g., Cohen's kappa). If voluntary usage is negligible, expert-rated quality does not improve, or LLM-expert agreement is no better than chance, the 'viable and can provide value today' claim is not supported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim (abstract; Section 5) that LLM evaluations of epics are 'viable and can provide value today' is supported only by participants' stated satisfaction and intended use in a concept test (Section 4.3.2). No participant interacted with the Epic Evaluator; they were shown static screenshots of evaluations of 2–4 elements of their own epic and asked to critique them (Sections 3.3 and 6). The paper reports 15/17 participants 'agreed with aspects of the score or recommendations,' but agreement is not a measure of evaluation accuracy or of downstream value: there is no comparison with expert human evaluation, no baseline condition, and no measurement of whether acting on the LLM's recommendations improves epic quality. The authors themselves list the lack of direct interaction as a limitation (Section 6). Consequently, the leap from 'participants found the concept appealing' to 'viable application providing value today' is underdetermined. The enthusiasm could reflect novelty, courtesy bias, or the perceived usefulness of any structured feedback rather than the specific LLM-based evaluation. For the central claim to hold, one would need evidence that the tool's evaluations are accurate enough to be trusted and that using them changes behavior or epic quality in real workflows. Absent that, the strongest defensible claim is that product managers see potential value in such a tool, not that it is currently viable.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This industry case study investigates whether large language models (LLMs) can evaluate the quality of agile epics. The authors developed an eight-element rubric for epic quality, built a prototype tool called Epic Evaluator that uses LLMs to rate elements and provide recommendations, and conducted semi-structured interviews with 17 product managers at a global company. During the interviews they concept-tested the rubric and static screenshots of LLM evaluations of the participants' own epics. The paper reports that participants found the rubric useful, expressed enthusiasm for the tool, identified several concrete uses (e.g., augmenting peer feedback, bulk evaluation for managers, training novices), and raised barriers including lack of domain knowledge, limited actionability, and integration friction. The central claim, stated in the abstract and in Section 5, is that LLM evaluations of epics are 'viable and can provide value today.' The paper also proposes design implications such as stage-aware evaluation, customizable rubrics, and tight integration with existing agile tools.","tokens_in":25555,"tokens_out":3256,"duration_ms":32438,"significance":"If taken as a preliminary qualitative study, this paper makes a useful contribution to the emerging literature on LLM-as-a-Judge for domain-specific text. Its strengths include a transparently described rubric development process with subject matter expert involvement, inclusion of the actual prompts in the appendix, and rich verbatim quotes from practitioners that give insight into real-world epic creation practices. The qualitative findings on perceived value, adoption barriers, and the need for domain knowledge are credible and potentially transferable to other agile organizations. However, the paper's central claim of current 'viability' goes beyond what the evidence can support: the study measures self-reported satisfaction and intended use from a concept test with static screenshots, not actual tool adoption, evaluation accuracy, or downstream improvements in epic quality. The contributions are therefore best framed as perceived value and design implications, not as demonstrated viability.","major_comments":[{"comment":"The central claim that LLM evaluations of epics are 'viable and can provide value today' is not supported by the evidence presented. The study used concept testing with static screenshots; participants never interacted directly with the Epic Evaluator (Section 3.3 and acknowledged in Section 6). There is no comparison against expert human evaluations, no baseline condition, and no measurement of whether acting on the LLM recommendations improves epic quality. Section 4.3.2 reports that 15/17 participants 'agreed with aspects' of the score or recommendations, but agreement with aspects of a concept test is not evidence of evaluation accuracy or of downstream value, especially given that average self-reported trust was only 3.5 on a 1-5 scale and that Section 4.3.3 lists substantial barriers including lack of domain knowledge, limited actionability, and workflow integration concerns. The strongest defensible conclusion is that product managers perceive potential value in such a tool and desire it; the claim of current viability should be softened to a claim of perceived potential or preliminary feasibility.","section":"Abstract; Section 5; Section 4.3.2; Section 6"},{"comment":"The paper states in the contributions that it introduces and validates a rubric, but the rubric validation consists solely of think-aloud feedback from 17 product managers. No inter-rater reliability is reported, the rubric was not independently applied by multiple raters to a set of epics, and no criterion validity against expert quality judgments or epic outcomes is provided. Furthermore, the participants identified missing elements (e.g., acceptance criteria), rigidity concerns, and worries about being penalized for missing elements that their teams do not use (Section 4.2.2). The term 'validate' therefore overstates the evidence; 'elicit practitioner feedback on a rubric' or 'preliminary evaluation of the rubric' would be more accurate.","section":"Section 1 Contributions; Section 3.1; Section 4.2.2"}],"minor_comments":[{"comment":"There is a typo in the evaluation prompt: 'Please rovide a detailed EXPLANATION' should read 'Please provide a detailed EXPLANATION.'","section":"Appendix A.3"},{"comment":"The example for the Non-Functional Requirements row is incomplete; it ends with 'The system should respond to password reset requests within 1 second for 95' and needs the concluding percentage or unit.","section":"Table 3 (Appendix A.1)"},{"comment":"The description of model selection and prompt iteration is qualitative. Please specify the temperature setting, the number of evaluation runs used to assess stability, and how 'stable' was operationalized, since LLM outputs are sensitive to sampling parameters.","section":"Section 3.2"},{"comment":"The statement that 15/17 participants 'agreed with aspects of the score or recommendations' is vague; please clarify what counted as agreement (e.g., agreement with the rating, with the explanation, with the recommendation, or any combination) and how this was coded from the interview data.","section":"Section 4.3.2"},{"comment":"Reference [60] is cited as 'SAFe. User Stories' but the URL points to businessmap.io; please verify that the citation and URL correspond to the intended source.","section":"References [60]"}],"recommendation":"major_revision","confidential_remarks":"This is a well-written industry case study with useful qualitative depth, and the limitations section is honest about the concept-testing design. The main issue is the gap between the evidence (perceived value from static screenshots) and the claim (viability and value today). If the authors revise the abstract, contributions, and Section 5 to frame the contribution as perceived value and design implications rather than demonstrated viability, the paper would be acceptable for a CHIWORK audience. The self-citation pattern is not a concern; the EvalAssist and EvaluLLM references are directly relevant to LLM-as-a-Judge tooling."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Read the IBM epic evaluator case study. Bottom line: the rubric and the qualitative findings on what product managers want from LLM-based epic feedback are the real contribution. The 'viable and can provide value today' claim, as stated in the abstract and Section 5, is thinner than the data support. The stress-test note is on target. Fifteen of seventeen participants said they agreed with aspects of the score or recommendations, but that is 'aspects,' and they never used the tool—they looked at static screenshots of evaluations of 2–4 elements of their own epic. That supports 'product managers see potential value and want this kind of support,' not 'this is a viable application today.' The authors acknowledge the no-direct-interaction limitation in Section 6, but the discussion still reaches for the stronger claim. Similarly, the contributions say the rubric was 'validated,' when what was measured is satisfaction and perceived usefulness; no inter-rater reliability, no comparison to expert human evaluation.\n\nWhat's good: the rubric is concrete and detailed, grounded in the company's epic template, with SME review and LLM input; the prompts are transparent; the thematic analysis is systematic and the quotes are rich enough to let a reader judge. The insights about stage-aware evaluation, flexibility to accommodate diverse practices, domain knowledge gaps, and integration into existing tools are genuinely useful design guidance for LLM-as-a-judge systems in real work. The paper is also honest about limitations: only IBM-approved open models, prompt sensitivity, single-company scope, no direct interaction.\n\nMy fix: soften the central claim to 'high perceived value and desire for integration' or 'promising concept-test results,' and reserve 'viability' for a study with direct tool use and ideally a human-rater baseline. The RAG idea for domain knowledge is reasonable future work, and the four-stage injection pattern is a nice design insight.\n\nThis paper is for practitioners and HCI researchers working on LLM-assisted knowledge work. It deserves a serious referee; send it out, with the expectation that claims get aligned to the evidence.","headline":"Solid qualitative case study whose 'viable today' claim overreaches the concept-test evidence; rubric and practitioner insights are the substance.","tokens_in":26126,"tokens_out":2148,"would_cite":true,"duration_ms":21068,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"LLM-based evaluations of agile epics are a viable new use of AI, an industry study of 17 product managers argues.","keywords":["agile software development","agile epics","quality evaluation","large language models","generative AI","LLM-as-a-judge","product managers","requirements engineering"],"falsifier":"Give a group of product managers a working Epic Evaluator integrated into their agile management tool for three months while a matched control group continues without it; if the intervention group shows no measurable improvement in epic quality scores, no reduction in downstream churn or delays, or no increase in actual usage, the viability claim is unsupported.","tokens_in":25116,"feed_emoji":"🤖","tokens_out":7644,"duration_ms":58943,"temperature":0.7,"pith_summary":"This paper argues that large language models can act as quality evaluators for agile epics, the high-level requirement documents product managers use to align stakeholders, and that product managers would welcome such evaluations in practice. Through a user study with 17 product managers, it tests a rubric of eight quality criteria alongside a prototype tool, Epic Evaluator, that rates epic elements and suggests improvements. High levels of satisfaction with the evaluations lead the authors to claim that agile epics are a new, viable application of AI evaluation. The paper also identifies barriers, including lack of domain knowledge, rigidity of rubrics, and need for workflow integration, that must be addressed before such tools are adopted.","feed_headline":"Product managers approve AI epic-quality checks","feed_subtitle":"In an industry study, 17 product managers found LLM-based epic evaluations useful and wanted them in their workflow.","key_machinery":"The central objects are an eight-element rubric defining quality criteria for agile epic elements, including title, problem statement, product outcome and instrumentation, user stories, requirements, assumptions, non-functional requirements, and out-of-scope, and a prototype tool called Epic Evaluator that uses prompt-based evaluation by an LLM to detect which elements are present, rate each on a High/Medium/Low scale, and generate explanations and recommendations. The rubric operationalizes quality so that an LLM can judge it, the prompts carry the evaluation logic, and the concept-testing interviews with product managers supply the evidence of viability.","core_discovery":"The central claim is that LLM evaluations of epics are viable and can provide value today. Product managers in the study largely agreed with the tool's scores and recommendations (15 of 17), expressed enthusiasm about using such evaluations to augment peer feedback, accelerate epic improvement, and support managers in triaging large numbers of epics, and rated their trust in the LLM evaluation at an average of 3.5 on a 5-point scale. The authors conclude that carefully designed LLM-based tooling can help product managers craft higher-quality epics and thereby reduce downstream churn, communication breakdowns, delays, and cost overruns.","pith_inferences":["If viability holds under real adoption, organizations could standardize epic quality across teams without mandating rigid templates, using a core rubric plus team-customizable extensions, a path the paper sketches but does not test.","The finding that managers would not use evaluations for performance assessment, while non-managers fear score-based reporting, suggests that governance of AI evaluation must separate quality improvement from personnel evaluation, a question the paper does not address.","A testable extension is a randomized field trial in which product managers receive an LLM evaluator integrated into their agile tool or no evaluator, measuring revision cycles and downstream churn rather than self-reported satisfaction.","The average trust rating of 3.5 on a 5-point scale and concerns about missing context imply that LLM epic evaluators will likely settle into a human-in-the-loop advisory role rather than an automated gatekeeper role."],"forward_implications":["LLM-based epic evaluation could be injected at four common milestones: after the first draft, before sharing with stakeholders, after stakeholder iteration, and before handoff to development.","Product managers want actionable recommendations over scores, suggesting that evaluator and authoring tools will tend to merge.","Managers see bulk evaluation of team epics as a way to spot epics needing attention, and as a training aid for novice product managers.","To be adopted, such tools must integrate tightly into existing agile management tools so users do not need to copy and paste epic text.","Adding domain and stakeholder knowledge, for example through retrieval-augmented generation, is the key next step to improve the specificity and actionability of evaluations."],"supporting_citations":[{"why":"It demonstrates that LLMs can identify quality issues in software requirements and explain them, serving as the direct precedent for the Epic Evaluator design.","marker":"[41]"},{"why":"It outlines the potential for LLMs to support requirements analysis and evaluation, including of epics, motivating the study.","marker":"[3]"},{"why":"It is an industry case study showing that LLM-assisted writing of epics and stories improves quality, providing anecdotal precedent.","marker":"[14]"},{"why":"It establishes LLM-as-a-judge as a method for evaluating text, the technical foundation of the tool.","marker":"[80]"},{"why":"It shows that LLMs can perform decently as evaluators of natural language text, supporting the viability premise.","marker":"[26]"},{"why":"It is a systematic review of agile requirements engineering and AI that frames the opportunity for generative AI in this space.","marker":"[2]"},{"why":"It is a systematic review linking poorly defined requirements to churn, delays, and cost overruns, the problem the paper targets.","marker":"[18]"}],"fun_headline_variants":["LLM epic quality checks gain PM approval","17 PMs find LLM epic checks useful","Generative AI aids epic quality evaluation","AI epic reviews: product managers see value","LLM epic evaluators impress product managers"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The claim rests on participants' self-reported satisfaction with static, screenshot-based evaluations, assuming that this predicts real-world adoption and improved epic quality in daily workflows.","fun_headline_variants_meta":{"raw":{"variants":["LLM epic quality checks gain PM approval","17 PMs find LLM epic checks useful","Generative AI aids epic quality evaluation","AI epic reviews: product managers see value","LLM epic evaluators impress product managers"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000304,"raw_usage":{"total_tokens":1682,"prompt_tokens":816,"completion_tokens":866,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":432,"completion_tokens_details":{"reasoning_tokens":800}},"tokens_in":432,"tokens_out":866,"duration_ms":8606,"temperature":1.0,"reasoning_tokens":800,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T22:09:56.899801+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Give a group of product managers a working Epic Evaluator integrated into their agile management tool for three months while a matched control group continues without it; if the intervention group shows no measurable improvement in epic quality scores, no reduction in downstream churn or delays, or no increase in actual usage, the viability claim is unsupported.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"It demonstrates that LLMs can identify quality issues in software requirements and explain them, serving as the direct precedent for the Epic Evaluator design."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"It is an industry case study showing that LLM-assisted writing of epics and stories improves quality, providing anecdotal precedent."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"It shows that LLMs can perform decently as evaluators of natural language text, supporting the viability premise."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"It is a systematic review of agile requirements engineering and AI that frames the opportunity for generative AI in this space."}],"review_version":1}