{"id":"6fb8d3aa-5040-4755-8875-21a06c9bdbe4","arxiv_id":"2506.13303","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A five-year case study shows that use case adoption at a large company diverges from textbook recommendations and that most quality criteria, except solution-orientation, have little measurable effect on design time.","lead":"An industrial case study of 1,188 business requirements found that a Swedish telecom software company's real-world use of use case descriptions deviates from textbook quality guidelines. Only a few quality factors, especially solution-oriented white-box steps, showed measurable impact on the time spent in the subsequent design phase.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Unit-of-analysis mismatch in RQ2 undermines the paper's headline solution-orientation result.","rationale":"The reader's weakest assumption—unobserved confounders in the causal DAGs—is genuine but generic to any observational case study and is partially acknowledged in the paper's threats-to-validity section. The unit-of-analysis issue is more specific and more directly threatens the one inferential result the abstract singles out: the claimed impact of solution-orientation on design time. It is also checkable from the replication package, which the paper itself provides. I therefore partially agree with the reader: the causal conclusions are not ready as stated, but the more load-bearing problem is not the DAG per se; it is the mismatch between UC-level predictors and a requirement-level outcome. The descriptive RQ1 contribution and the replication package are real strengths and are not affected by this concern. The existing CONDITIONAL verdict remains appropriate, with the additional condition that the authors clarify and re-estimate the RQ2 model at the correct unit of analysis.","tokens_in":16191,"tokens_out":7154,"duration_ms":82098,"concrete_test":"Download the replication package (Zenodo DOI 10.5281/zenodo.15672950) and inspect the data-preparation script for RQ2. Determine whether each row is a use case or an aggregated business requirement; then re-run the solution-orientation model with a hierarchical specification that adds a requirement-level random intercept (if rows are use cases) or with an explicit sum/mode aggregation (if rows are already requirements). Report the posterior mean and 95% credible interval for white-box steps on design time. If the interval crosses zero or the point estimate shrinks substantially, the abstract overstates the solution-orientation result. As a secondary check, add total steps or total interactions as a covariate to test whether the white-box-step effect is distinct from UC length.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central inferential result highlighted in the abstract—that solution-orientation has an actual impact on design time—rests on an unstated and potentially incorrect unit of analysis. The outcome, design time, is measured once per business requirement (Section III-B5), while the key predictors (white-box steps, UC location, explicit actors) are attributes of individual template-style UCs (Tables III–IV). The paper describes aggregation of UC-level attributes to the requirement level only for descriptive visualization (Section III-C1); it never states how the RQ2 regression data were constructed. If the RQ2 models use one row per use case, requirements with multiple template UCs contribute duplicated design-time outcomes, creating non-independence and overconfident credible intervals. If they use one row per requirement, the aggregation method (sum, mean, or mode) is not specified, and information is lost when a requirement mixes UCs in different locations. The paper's own threat-to-validity statement that 'we did not consider more complex hierarchical models' (Section V-C3) makes this omission acute, because a requirement-level random intercept is the standard remedy for exactly this nesting. Without clarifying or re-estimating the model, the abstract's causal claim about solution-orientation cannot be properly assessed.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper reports an industrial case study at a large, globally distributed telecom BSS company. The authors analyze 1,188 business requirements and 1,192 use-case-like descriptions authored between 2020 and 2024, manually annotate the 273 template-style use cases against seven quality heuristics from Phalp et al., and use Bayesian statistical causal inference to estimate (RQ2) how use-case quality attributes affect design-stage duration and (RQ3) how requirements-engineering process factors affect selected quality attributes. The descriptive results show that adoption grew only after active management promotion, that template-style use cases form a minority (22.9%), that many use cases are solution-oriented and located in 'Solution Proposal' sections, and that actual practice deviates from the textbook 7-C guidelines. The inferential results are presented as mostly null, with solution-orientation (white-box steps and UC location) and, less clearly, explicit actors associated with longer design time; RQ3 finds mainly that location and complexity predict white-box steps. The abstract concludes that only a few phenomena such as solution-orientation show an actual impact in practice.","tokens_in":16339,"tokens_out":5669,"duration_ms":63904,"significance":"The paper's main strength is its rare longitudinal, industrially grounded dataset. The manual review of all 1,188 requirements with a documented extraction guideline, the use of established quality criteria from prior literature, the explicit causal DAGs, and the replication package with mocked data are concrete assets that make the descriptive adoption findings credible and reusable. If the inferential claims can be made sound, the paper would make a useful contribution to requirements-quality research by providing field evidence on which textbook UC-quality recommendations matter in practice. The finding that only solution-orientation, not most other 7-C heuristics, predicts downstream design activity is a falsifiable, practically relevant result. At present, however, the unit-of-analysis problem and the overstatement of significance mean that the inferential part is not yet a reliable basis for the abstract's causal wording.","major_comments":[{"comment":"The RQ2 analysis contains a unit-of-analysis mismatch that is load-bearing for the abstract's claim about solution-orientation. Design time is measured once per business requirement (Section III-B5), while predictors such as white-box steps, UC location, and explicit actors are attributes of individual template-style UCs (Tables II-IV). Section III-C1 describes aggregation of UC-level attributes to the requirement level only for descriptive visualization; the manuscript never states how the RQ2 regression data were constructed. If the models use one row per UC, requirements with multiple template UCs contribute duplicated design-time outcomes, violating independence and overstating precision. If they use one row per requirement, the aggregation method (sum, mean, or mode) is unspecified and information about mixed UC locations is lost. Section V-C3's admission that no hierarchical models were considered makes this omission acute. Please report the exact data construction and re-estimate the models with a requirement-level random intercept or an equivalent clustered analysis.","section":"III-C1/III-C2, V-C3"},{"comment":"The paper overstates the statistical support for solution-orientation. In Section IV-B2, Figure 17 shows credible intervals that overlap across locations, and Figure 18 shows wide intervals for both white-box steps and explicit actors; Section IV-B3 explicitly labels the consecutive-steps effect non-significant. Yet Section V-A describes solution-orientation as a significant factor, and Section V-B reports that the managers perceived the effects as significant. The phrase 'significant' should be reserved for a stated posterior criterion, and the relevant effect estimates with credible intervals should be reported in the text rather than only in figures. As written, the abstract's causal claim about solution-orientation rests on mean directions rather than on inferential evidence.","section":"IV-B2, V-A"},{"comment":"The causal-sufficiency assumption of the DAGs is not quantified. The manuscript acknowledges in Section V-C1 that unobserved confounders such as author skill, stakeholder pressure, and requirement importance may affect both UC quality and design time, but it does not assess how strong such confounding would need to be to explain away the reported effects. Because the headline result is precisely a causal claim about white-box steps and UC location increasing design time, a sensitivity analysis (e.g., E-value or bias-formula bounds) or an explicit argument about the plausible direction and magnitude of unobserved confounding would materially strengthen the conclusion.","section":"III-C2a, V-C1"},{"comment":"The selection of the four outcome variables for RQ3 is not principled. Section IV-C states that the authors investigate the four most significant UC quality factors, but the RQ2 results in Section IV-B1 report no meaningful effect for number of use cases, and Section IV-B2 gives no significance statement for explicit actors with all intervals overlapping. The criterion for 'most significant' should be defined (e.g., posterior probability of a positive or negative effect, width of credible interval, or minimum effect size) and applied consistently, otherwise the set of RQ3 outcomes appears selected without a stated rule.","section":"IV-C"}],"minor_comments":[{"comment":"There is a typo in the sentence 'the format in which a UC descirption may be specified'; 'descirption' should be 'description'.","section":"II-A"},{"comment":"The text contains an unresolved placeholder: 'This hints at the negative effect of solution-orientation which we will discuss in ??.' The cross-reference needs to be completed.","section":"IV-B2"},{"comment":"The manual extraction procedure is described as involving independent extraction followed by discussion, but no inter-rater agreement statistic is reported. Since the descriptive and inferential results depend on these manual judgments, reporting a measure such as Cohen's kappa for at least a subsample would increase confidence in the extraction.","section":"III-B2"},{"comment":"The y-axis of Figure 16 is labeled 'Requirements' but the plot shows percentages; the label should be 'Share of requirements' to avoid confusion.","section":"IV-A2, Figure 16"},{"comment":"The figures report credible intervals but none of the marginal-effect plots includes the numerical posterior mean and interval values in the caption or text; adding these values would make the reported overlaps and distinctions easier to assess.","section":"III-C2c, Figures 17-20"}],"recommendation":"major_revision","confidential_remarks":"The paper fits the scope of an empirical software engineering journal and has a solid descriptive contribution. The causal-inference part needs the requested re-analysis before publication; the unit-of-analysis issue is fixable within the manuscript's scope. I see no reason to doubt the authors' integrity, and the replication package is a plus."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Two things to know. The descriptive half of this paper is the real contribution: five years of Jira data, all 1,188 requirements manually screened, 273 template-style use cases scored against the 7 Cs, and adoption trends plotted over time. That is rare industrial evidence, and the finding that UC usage only stabilised after active management push, and that many 'use cases' are just titles, is worth publishing on its own. The authors also did the right things around transparency: explicit causal DAGs, uninformative priors, posterior predictive checks, a replication package, and validation discussions with two senior managers.\n\nThe soft spots are all in the inferential part. Most causal effect estimates overlap zero or each other; the paper calls solution-orientation 'significant' and says only a few phenomena 'show an actual impact,' but the figures show overlapping intervals and the discussion itself uses 'hints' and 'presumably.' That is overstatement. More serious: the RQ2 models have an unclear unit of analysis. Design time is measured once per business requirement, while the predictors (white-box steps, location, explicit actors) are attributes of individual use cases, and requirements can contain several. The paper never states how the regression data were built. If each use case is a row, design time is duplicated for multi-use-case requirements and the credible intervals are too narrow. If each requirement is a row, the aggregation is unspecified and a requirement mixing UCs in the Business Requirements and Solution Proposal sections loses information. The authors concede in V-C3 that they did not consider hierarchical models; a requirement-level intercept is the standard fix. This needs to be resolved before the abstract's causal claim is credible. They also acknowledge unobserved confounders (author skill, stakeholder pressure), so the causal part should be framed as exploratory.\n\nMinor: there is a broken cross-reference ('we will discuss in ??') and the DOI is a placeholder. That is polish, not substance.\n\nThis paper deserves a serious referee. The descriptive RQ1 work is solid; RQ2 needs either re-estimation at the requirement level or a much more careful framing. My recommendation: send to peer review, ask for the RQ2 data construction to be clarified or the model re-run, and make sure the abstract matches the actual uncertainty.","headline":"Roughly: the descriptive adoption story is credible and worth reading, but the inferential RQ2 has an unresolved unit-of-analysis problem and the abstract oversells solution-orientation.","tokens_in":16868,"tokens_out":2803,"would_cite":true,"duration_ms":31657,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"After five years of industrial data, only one use-case quality defect—solution-orientation—showed a clear effect on downstream design time.","keywords":["use case descriptions","requirements engineering","requirements quality","industrial case study","causal inference","Bayesian data analysis","solution-orientation","design time"],"falsifier":"A re-analysis that adds author identity, requirement priority, and stakeholder pressure to the adjustment sets would falsify the solution-orientation claim if the estimated effect of white-box steps on design time disappears; likewise, a controlled experiment where the same requirements are written with and without white-box steps and design time is measured could settle the direction.","tokens_in":15946,"feed_emoji":"📋","tokens_out":8123,"duration_ms":83483,"temperature":0.7,"pith_summary":"Use case descriptions are a standard way to write functional requirements, but textbook advice on how to write good ones has rarely been tested against real projects. This paper studies five years of requirements from a globally distributed telecom company—1,188 business requirements containing 1,192 use cases—and asks whether the recommended quality rules actually mattered. The descriptive finding is that practice deviates from the textbooks: most use cases were titles rather than full templates, passive voice was common, and adoption grew only after active management promotion. The inferential finding is that, of all the quality heuristics tested, only solution-orientation showed a clear impact: use cases that describe system-internal white-box interactions, or that are placed in a \"Solution Proposal\" section, were associated with longer design time. The paper concludes that quality research on use cases should focus on the few factors that show observable effects rather than the full checklist.","feed_headline":"Only solution-oriented use cases measurably raised design time","feed_subtitle":"Five years of industrial data show white-box steps delayed design; most textbook quality rules showed no effect.","key_machinery":"The machinery is a combination of manual quality annotation and statistical causal inference. Template-style use cases were scored on the seven quality heuristics—coverage, cogent, coherent, consistent abstraction, consistent structure, consistent grammar, consideration of alternatives—operationalized as attributes like white-box steps, misplaced variations, coherence, and grammar consistency. Causal directed acyclic graphs (DAGs) make the assumed relationships explicit, and Bayesian regression models using backdoor adjustment estimate the average causal effect of each use-case property on design time, measured as days spent in the design stage. The central distinction carrying the main result is black-box versus white-box interaction level, which defines solution-orientation.","core_discovery":"The paper's central claim is that most established use-case quality recommendations did not correspond to measurable differences in practice, and that the exception—solution-orientation—worked in the direction textbooks warn about. Across 273 template-style use cases manually evaluated against a seven-criterion quality checklist, the number of white-box steps (steps describing system-internal interactions rather than user-system interactions) and the location of the use case in the \"Solution Proposal\" section both predicted longer time spent in the design stage. Other findings ran against the guidelines: use cases with more misplaced variations in the main scenario were associated with shorter design time, and most coherence or grammar measures had no detectable effect. The authors interpret this as evidence that use-case quality research should be steered toward empirically impactful phenomena.","pith_inferences":["A natural extension would be an intervention study where solution-oriented use cases are rewritten as black-box user-system interactions and design time is tracked; the paper's causal model predicts this would shorten the design stage.","The results imply that requirements quality checklists should be empirically validated against downstream activity outcomes before being adopted, since stylistic heuristics can be neutral or even inversely related to performance.","The adoption curve suggests that deliberate management promotion, not intrinsic usefulness alone, drove the format's spread; future adoption studies should measure promotion intensity as a variable.","The negative association between misplaced variations and design time could be explained by the cost of cross-referencing numbered extensions; a targeted experiment could test whether inline alternatives reduce cognitive load."],"forward_implications":["Adopting template-style use cases will not by itself shorten the design phase; the format's visible benefit in this data was for high-complexity requirements, where it produced the least variable design time.","Quality assurance efforts for use cases should concentrate on solution-orientation—removing white-box steps and keeping use cases in the requirements section—rather than on grammar and structural heuristics that showed no effect.","The textbook rule to keep alternative paths out of the main scenario is called into question: more misplaced variations were associated with shorter design time in the studied company.","For the case company, the single actionable requirements-engineering lever identified is use-case location: placing a use case in the Business Requirements section makes white-box steps much less likely at every complexity level."],"supporting_citations":[{"why":"It supplies the seven-criterion use-case quality checklist that the study operationalizes and tests.","marker":"[11]"},{"why":"It defines the basic use case template whose industrial adoption the study tracks.","marker":"[8]"},{"why":"It popularized the tabular template format and the user-level/system-level distinction behind solution-orientation.","marker":"[19]"},{"why":"It establishes the activity-based requirements quality view linking artifact quality to downstream activity performance.","marker":"[10]"},{"why":"It provides the Bayesian data analysis workflow used for estimation and model checking.","marker":"[42]"},{"why":"It provides the causal inference framework and the backdoor criterion used to identify adjustment sets.","marker":"[43]"},{"why":"It states the backdoor criterion used to derive which variables to control in each model.","marker":"[49]"}],"fun_headline_variants":["White-box use cases lengthen design time, defying guidelines","Most use-case quality rules show no real-world effect","Use-case guidelines miss mark, except for white-box steps","Most textbook use-case rules had no measurable impact"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The conclusions assume that the causal diagrams behind the statistical models capture all important common causes of use-case quality and design time; if unmeasured factors like author skill, requirement importance, or stakeholder pressure also drive both, the estimated effects may not be causal.","fun_headline_variants_meta":{"raw":{"variants":["White-box use cases lengthen design time, defying guidelines","Most use-case quality rules show no real-world effect","Use-case guidelines miss mark, except for white-box steps","Most textbook use-case rules had no measurable impact"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001053,"raw_usage":{"total_tokens":4412,"prompt_tokens":926,"completion_tokens":3486,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":542,"completion_tokens_details":{"reasoning_tokens":3421}},"tokens_in":542,"tokens_out":3486,"duration_ms":26475,"temperature":1.0,"reasoning_tokens":3421,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T00:35:37.324880+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A re-analysis that adds author identity, requirement priority, and stakeholder pressure to the adjustment sets would falsify the solution-orientation claim if the estimated effect of white-box steps on design time disappears; likewise, a controlled experiment where the same requirements are written with and without white-box steps and design time is measured could settle the direction.","supporting_citations":[{"cited_title":"Assessing the quality of use case descriptions,","cited_arxiv_id":null,"evidence_quote":"It supplies the seven-criterion use-case quality checklist that the study operationalizes and tests."},{"cited_title":"Basic use case template,","cited_arxiv_id":null,"evidence_quote":"It defines the basic use case template whose industrial adoption the study tracks."},{"cited_title":"Cockburn and L","cited_arxiv_id":null,"evidence_quote":"It popularized the tabular template format and the user-level/system-level distinction behind solution-orientation."},{"cited_title":"Requirements quality research: a harmonized theory, evaluation, and roadmap,","cited_arxiv_id":null,"evidence_quote":"It establishes the activity-based requirements quality view linking artifact quality to downstream activity performance."},{"cited_title":"McElreath, Statistical rethinking: A Bayesian course with examples in R and Stan","cited_arxiv_id":null,"evidence_quote":"It provides the Bayesian data analysis workflow used for estimation and model checking."},{"cited_title":"Causal inference,","cited_arxiv_id":null,"evidence_quote":"It states the backdoor criterion used to derive which variables to control in each model."}],"review_version":1}