{"id":"71048a6b-074c-4ffb-95e1-5f7d25395299","arxiv_id":"2501.11792","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Developers pick debugging strategies based on interacting contextual factors, with hypothesis testing as the default and codebase familiarity as a key determinant.","lead":"Researchers surveyed 35 developers and interviewed 16 experts to learn how they choose debugging strategies for hard web application bugs. They found that the choice depends on interacting contextual factors such as codebase familiarity, bug reproducibility, and organizational constraints, not just developer skill.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Interview prompts may have constructed the 'combinations of factors' finding; the central claim needs validation against unprompted behavior.","rationale":"The reader's weakest assumption (retrospective self-report recall bias) is real and acknowledged by the authors, but my concern is more specific: the semi-structured interview procedure itself may have produced the central 'combinations of factors' result by prompting participants with the exact strategy list and factor categories before asking them to reason about strategy selection. This is a construct-validity threat distinct from memory distortion: even with perfect recall, participants would be likely to produce multi-factor, strategy-switching narratives because the interviewer supplied those categories. The Figure 2 decision model amplifies this risk by encoding the prompted connections as deterministic rules without reporting how often each edge was mentioned or whether any edge reached consensus thresholds. Consequently, the central claim—that contextual factor combinations drive strategy choice—is not yet established as a property of real debugging behavior. This does not overturn the reader's conditional verdict; it reinforces it. The proposed re-coding of existing transcripts is a low-cost check that could distinguish prompted from spontaneous reasoning, and a brief think-aloud study would provide external validation. If both checks support the model, the claim would be substantially stronger; if not, the decision model should be presented as hypothesis-generating rather than descriptive.","tokens_in":16915,"tokens_out":3011,"duration_ms":34913,"concrete_test":"Re-code the Study 2 transcripts to mark each factor mention or strategy switch as either spontaneous (occurring before the interviewer introduced any strategy list or factor category) or prompted (occurring after the interviewer mentioned Table 3 or the factor categories). If a large majority of combination/switching statements occur only after prompting, the central claim is an artifact. Separately, run a small think-aloud study (n=6-8) where developers debug real or seeded web defects with no strategy list or factor categories provided, and compare the frequency and complexity of factor-based reasoning with the interview sample.","verdict_should_be":"UNCHANGED","load_bearing_attack":"In Section 3.2.3, the interview protocol gives participants a list of six strategies (Table 3), a list of contextual factor categories, and then asks them to select strategies and describe the factors they considered. This primes participants to articulate exactly the kind of multi-factor, strategy-switching reasoning that the paper reports as its central finding. The 'combinations of factors' and 'evolving throughout the debugging process' claims therefore may be artifacts of the elicitation procedure: participants were essentially handed the vocabulary of contextual factors and asked to connect them to strategies. Figure 2 then assembles these prompted connections into a deterministic decision model, but no per-edge counts, inter-rater reliability values, or behavioral validation are reported. The paper's own Section 5 concedes that the preliminary codebook 'might have influenced the responses' and that self-reports are subject to recall bias. Because the central claim is that real debugging decisions are driven by interacting contextual factors, the load-bearing premise is not merely that memory is accurate, but that the observed pattern would appear without the interviewer's scaffolding.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper investigates how developers choose debugging strategies for challenging web application defects through two complementary studies: a survey of 35 developers with varying expertise and semi-structured interviews with 16 expert developers. It identifies six categories of contextual factors, maps them to debugging strategies in Figure 2, and claims that strategy selection is driven by interacting combinations of factors that evolve during debugging, with hypothesis testing as a baseline and code familiarity/experience as decisive influences. The paper also derives implications for debugging tools and education.","tokens_in":17187,"tokens_out":4048,"duration_ms":43697,"significance":"If the central claims hold, the paper makes a useful contribution by moving beyond listing debugging strategies to modeling the contextual conditions under which developers select and switch strategies, with concrete implications for tool design and debugging education. The paper's strengths include its two-study design, systematic two-round coding with multiple coders, rich participant quotes, and a candid threats-to-validity section. However, the central empirical claim requires more direct evidence that the reported factor-strategy links are not artifacts of the elicitation protocol, and the decision model in Figure 2 would need stronger empirical grounding before it can support the paper's strongest conclusions.","major_comments":[{"comment":"The interview protocol gave participants the six strategies in Table 3 and the four contextual factor categories before asking them to choose strategies and describe the factors they considered. This scaffolding may have elicited exactly the kind of multi-factor, strategy-switching reasoning that the paper reports as its central finding. The manuscript does not report whether the same factor-strategy links appear in unprompted narratives, such as the open-ended accounts in Study 1 or the parts of Study 2 in which participants described recent experiences in their own words. Please provide such an analysis, or explicitly reframe the central claim as 'when prompted with these categories, developers report that...'.","section":"Section 3.2.3"},{"comment":"Figure 2 is labeled a 'Model of developer decision-making in web debugging strategy selection,' but no per-edge counts, inter-rater reliability values, or validation against observed debugging behavior are reported. The figure should either be supported with the number of participants whose transcripts support each edge and representative quotes, or relabeled as an illustrative synthesis rather than a validated decision model. In addition, several branches and edge labels (e.g., 'Deprecated?' with labels l and k) are unclear without a more detailed legend.","section":"Figure 2"},{"comment":"The abstract frames the contribution as covering 'different expertise levels,' and RQ1 asks about factors influencing strategy choice, but the Results section does not systematically compare strategy choices by expertise: Study 2 recruited only experts (8-38 years of experience) and Study 1 included many students with a median of 2 years. Claims such as 'experience and familiarity with the code are keys to making the correct decision' should be presented as participants' stated beliefs rather than as a demonstrated expertise effect.","section":"Abstract and Section 3"},{"comment":"The threats-to-validity section acknowledges recall bias, sample bias, and potential codebook influence, but these admitted limitations are not connected to the strength of the central claim that real debugging decisions are driven by interacting contextual factors. Because the model is derived from retrospective self-reports, the paper should state explicitly that the model represents reported reasoning and should describe how the mitigating evidence mentioned in Section 5 (e.g., references to GitHub commits or chat history) was used in the analysis.","section":"Section 5"}],"minor_comments":[{"comment":"The text 'Need to talk to stockholders or manager' should read 'stakeholders,' and the edge labels in the figure are referenced inconsistently in the text (e.g., 'Fig 2-i, 8' and 'Fig 2-d-k-NO').","section":"Figure 2"},{"comment":"The asterisk and underline notation for marking the provenance of factors is not consistently followed: some underlined factors, such as 'Complexity' in Table 5, are not discussed in the corresponding text, making the notation harder to follow.","section":"Tables 5-7"},{"comment":"The phrase 'double-blind analysis' is misleading given the collaborative open-coding process described in Section 3.2.4; please rephrase to describe the actual analysis procedure.","section":"Section 5"},{"comment":"The related work discussion would benefit from a closer comparison to Spinellis's bottom-up/top-down strategy-selection advice [47], since the paper's decision model partly overlaps with that distinction.","section":"Section 2.3"}],"recommendation":"major_revision","confidential_remarks":"The paper is not ready in its current form. The elicitation-priming concern is real and should be addressed head-on with additional analysis of unprompted data, or the central claim should be qualified accordingly. I do not see a fundamental flaw that requires rejection; the empirical material and qualitative analysis are valuable and capable of supporting a more carefully scoped contribution."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short take: this is a solid qualitative study that gives the debugging community a useful taxonomy of contextual factors, but the decision model in Figure 2 is a rational reconstruction from self-reports, not a validated model. The central claim about combinatorial factors is plausible and grounded in quotes, but the interview protocol partly primes the very combination language you end up reporting.\n\nWhat's new: prior work catalogued strategies and isolated success factors like code familiarity and defect complexity. This paper synthesizes those into a factor-strategy map specific to web debugging, and adds organizational context, project requirements, and individual traits as categories. The six-factor taxonomy in Tables 5–7 and the three new strategies in Table 4 are practical output. The two-round coding with disagreements resolved by scope discussion is standard and careful. The threats-to-validity section is unusually candid: they flag recall bias, sample bias, codebook influence, and the narrow domain.\n\nSoft spots: The biggest is that Study 2 hands participants a list of six strategies and four factor categories, then asks them to connect the dots. So the \"combinations of factors\" finding may be partly an artifact of the elicitation. The second phase does ask about real past experiences, which helps, but the ten scenarios are given, and the factor vocabulary is already on the table. Figure 2 is labeled a model of decision-making but no per-edge counts, inter-rater reliability, or behavioral validation support it. It's a plausible flow chart of reported reasoning, not a model that predicts choices. The paper's own limitations section concedes most of this, but the abstract and discussion outrun that caution.\n\nThat said, the central qualitative claim — developers say they weigh multiple contextual factors and switch strategies — is credible and does not overreach badly. The model is best read as a hypothesis for future work, not a validated result.\n\nWho should read it: HCI and software engineering researchers interested in debugging tools, education, or expertise. It's also a reasonable teaching example of qualitative methods. It deserves a serious referee, because the taxonomy is genuinely useful and the method, though imperfect, is transparent. The authors should be asked to soften the model's label, report per-edge saturation, and consider a validation study.\n\nRecommendation: send to peer review, conditional on revisions that clarify the exploratory nature of Figure 2 and the priming limitation.","headline":"Useful taxonomy of contextual factors for debugging strategy choice, but the Figure 2 decision model is a prudent reconstruction from primed self-reports, not a validated model.","tokens_in":17574,"tokens_out":2030,"would_cite":true,"duration_ms":20553,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Expert debugging strategy choice is driven by interacting contextual factors, with hypothesis testing as the baseline.","keywords":["debugging strategies","contextual factors","expert developers","web application defects","hypothesis testing","code familiarity","qualitative interviews","strategy selection"],"falsifier":"An observational study that records developers' actual debugging sessions and finds their strategy choices are unrelated to the defect and codebase factors listed here—for instance, the same habitual strategy is used across clear and unclear, familiar and unfamiliar contexts—would falsify the claimed factor-strategy links.","tokens_in":1395,"feed_emoji":"🐛","tokens_out":1594,"duration_ms":75244,"temperature":0.7,"pith_summary":"This paper argues that the way expert developers debug a challenging web application defect is not a matter of knowing a repertoire of strategies but of reading the situation: defect and codebase characteristics combine and interact, and the right choice changes as debugging proceeds. The authors establish this with two complementary studies—a survey of 35 developers and semi-structured interviews with 16 expert developers—and synthesize the results into a decision model for strategy selection. They find that hypothesis testing is the default starting point, while codebase familiarity and prior experience are the factors that most often determine whether a hypothesis can be formed and which supporting strategy will work. If the account is right, debugging competence includes a context-assessment skill that current teaching and tooling largely ignore, pointing to context-aware debugging tools and educational frameworks.","feed_headline":"Context, not just skill, steers how experts debug","feed_subtitle":"Two studies of 51 developers show defect and codebase cues jointly drive strategy choice—and training should follow.","key_machinery":"The load-bearing mechanism is the qualitative decision model in Figure 2, a factor-to-strategy mapping that shows how developers branch on checks such as whether a clear error message exists, whether the defect is reproducible, whether it is client-side, whether the developer has code access and codebase familiarity, whether the defect is user-specific or sporadic, and whether the codebase is small, deprecated, or familiar. The model is built by causal coding of interview transcripts, which locates and extracts the causal beliefs developers state between contextual factors and strategy choices. This machinery converts open-ended interviews into a compact branching structure that explains both initial strategy selection and mid-debugging switches.","core_discovery":"The paper's central discovery is that expert developers choose debugging strategies by weighing combinations of contextual factors, not by applying a single preferred method. Defect characteristics—clarity, reproducibility, and the root-related category covering network, data, configuration, and hardware—and codebase characteristics—familiarity, access, maintenance, technical stack, testability, and complexity—are the dominant influences, with organizational context, tool availability, individual traits, and project requirements also playing a role. Developers typically begin with hypothesis-test debugging and switch strategies as new information emerges; the proposed decision model in Figure 2 maps these factor checks to strategies such as backward-reasoning, forward-reasoning, simplification, binary-search, error-message debugging, and system-level checks. The paper also finds that experienced developers draw on strategies not previously documented and treat debugging as an occasion to improve code quality.","pith_inferences":["An implication the paper leaves implicit: if strategy selection is a context-reading skill, then an IDE feature that prompts developers to check factors like reproducibility, clear error message, and codebase familiarity before choosing a tactic could serve as both a teaching tool and a test of the model.","Beyond the paper's data, the same factor-strategy links could be tested in real time by instrumenting debugging sessions and comparing observed factor states against the Figure 2 branches, rather than relying on retrospective accounts.","A further extension: the growing use of AI code assistants may alter the familiarity factor, since developers can quickly learn unfamiliar code; whether that shifts their strategy choices away from forward-reasoning toward hypothesis-testing is a concrete empirical question."],"forward_implications":["Debugging education should teach context-assessment skills—evaluating reproducibility, clarity, codebase familiarity, and access—rather than only demonstrating individual strategies.","Debugging tools should be designed around problem contexts, because a tool that supports one strategy, such as backward-reasoning via a slicer, may be useless in contexts that call for simplification or binary-search.","The documented set of debugging strategies should be expanded beyond the six inherited from prior work to include system-level checks, external-resource consultation, and historical analysis with version-control bisection.","Because codebase familiarity and access so often decide whether hypothesis-testing and backward-reasoning are viable, lowering those barriers, through better code comprehension support, should change which strategies developers can effectively use.","Debugging is also a code-quality activity, so evaluations of debugging success should include maintainability outcomes, not only time-to-fix."],"supporting_citations":[{"why":"Early evidence that program comprehension shortens debugging time; anchors the paper's familiarity factor.","marker":"[13]"},{"why":"A foundational analysis of bug-location strategies; contributed to the strategy inventory used in interviews.","marker":"[21]"},{"why":"A modern debugging framework; supplies strategy definitions and parts of the contextual-factor taxonomy.","marker":"[47]"},{"why":"An experiment on how practitioners locate and fix bugs; grounds the hypothesis-testing, forward-reasoning, and backward-reasoning strategies.","marker":"[6]"},{"why":"Work on explicit programming strategies; motivates the study of strategy selection and the use of mentoring experience as an expertise criterion.","marker":"[24]"},{"why":"A study of contemporary debugging needs; supplies documented factors such as sporadic defects, compatibility, and environment.","marker":"[26]"},{"why":"A pocket guide to debugging; provides strategy definitions for simplification and binary-search used in the interview materials.","marker":"[12]"},{"why":"A study of novice debugging patterns; supports the familiarity and program-comprehension factor.","marker":"[2]"}],"fun_headline_variants":["Context combos, not just skill, steer expert debugging tactics","Why experts switch debug strategies: it's the situation, not the solution","Debugging choices hinge on layered cues, study of 51 shows","Expert debugging: context pairs decide the strategy, not experience alone","Challenging defects: experts pick tactics from context, not habit"],"cache_read_input_tokens":19840,"weakest_assumption_plain":"The paper's conclusions rest on retrospective self-reports from 35 survey respondents and 16 interviewees accurately capturing the reasoning that actually drove their debugging choices.","fun_headline_variants_meta":{"raw":{"variants":["Context combos, not just skill, steer expert debugging tactics","Why experts switch debug strategies: it's the situation, not the solution","Debugging choices hinge on layered cues, study of 51 shows","Expert debugging: context pairs decide the strategy, not experience alone","Challenging defects: experts pick tactics from context, not habit"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000244,"raw_usage":{"total_tokens":1498,"prompt_tokens":876,"completion_tokens":622,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":492,"completion_tokens_details":{"reasoning_tokens":532}},"tokens_in":492,"tokens_out":622,"duration_ms":6613,"temperature":1.0,"reasoning_tokens":532,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T17:51:13.598754+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"An observational study that records developers' actual debugging sessions and finds their strategy choices are unrelated to the defect and codebase factors listed here—for instance, the same habitual strategy is used across clear and unclear, familiar and unfamiliar contexts—would falsify the claimed factor-strategy links.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Early evidence that program comprehension shortens debugging time; anchors the paper's familiarity factor."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"A foundational analysis of bug-location strategies; contributed to the strategy inventory used in interviews."},{"cited_title":"Spinellis","cited_arxiv_id":null,"evidence_quote":"A modern debugging framework; supplies strategy definitions and parts of the contextual-factor taxonomy."},{"cited_title":"Böhme, E","cited_arxiv_id":null,"evidence_quote":"An experiment on how practitioners locate and fix bugs; grounds the hypothesis-testing, forward-reasoning, and backward-reasoning strategies."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Work on explicit programming strategies; motivates the study of strategy selection and the use of mentoring experience as an expertise criterion."},{"cited_title":"Layman, M","cited_arxiv_id":null,"evidence_quote":"A study of contemporary debugging needs; supplies documented factors such as sporadic defects, compatibility, and environment."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"A pocket guide to debugging; provides strategy definitions for simplification and binary-search used in the interview materials."},{"cited_title":"Ahmadzadeh, D","cited_arxiv_id":null,"evidence_quote":"A study of novice debugging patterns; supports the familiarity and program-comprehension factor."}],"review_version":1}