{"id":"2cdc0043-6d1d-4453-ac12-f0023ccadc4d","arxiv_id":"2412.14554","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"The paper reports 26 challenges in LLM-based software engineering, grouped into seven aspects, derived from a structured discussion among 24 academics and practitioners.","lead":"A workshop of 24 software engineering experts produced a list of 26 challenges for using large language models throughout the software lifecycle, from requirements to vulnerability management. The paper is a qualitative synthesis meant to guide research priorities in LLM4SE, but it is an expert opinion piece rather than a measured study.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Section 7 says 'seven challenges' while the abstract claims 26; without released transcripts and inter-rater reliability, the challenge list cannot be independently audited.","rationale":"The reader's weakest assumption identifies the representativeness and transparency of the qualitative study, and I agree that this is the most load-bearing issue. The strongest claim is essentially a numerical taxonomy: 26 key challenges across seven aspects. The paper does present 26 numbered challenges across seven subsections (Sections 4.1–4.7), so the claim is substantively instantiated. However, the derivation rests entirely on discussions among 24 participants, transcribed and coded primarily by one author, with no released raw data and no inter-rater reliability measure. The paper's own Section 6 acknowledges completeness threats. The internal inconsistencies (Section 4 intro: 'six aspects'; Section 7: 'seven challenges') are not fatal—they are likely wording errors—but they amplify the auditability problem, since the reader cannot determine which count is authoritative. I do not see evidence of fraud or misrepresentation; the paper cites a public repository and describes a recognizable qualitative procedure. The appropriate response is to keep the conditional verdict: accept the taxonomy as a useful expert roadmap only if the authors release the coding materials and correct the count inconsistencies. Thus the reader's verdict is unchanged.","tokens_in":34372,"tokens_out":4747,"duration_ms":35475,"concrete_test":"Request from the authors the anonymized transcripts and opinion cards for the six thematic sessions, and have two researchers who were not involved in the original study independently repeat the open card sorting described in Section 3, computing inter-rater agreement (e.g., Cohen's kappa) on the resulting challenge set. In parallel, independently enumerate the numbered challenges in Sections 4.1–4.7 and the aspect headings, and reconcile them with the abstract's '26 key challenges from seven aspects' and with the conclusion's 'seven challenges.' If the re-sorting produces a materially different list or the enumeration disagrees with the stated counts, the central claim fails.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Section 3 describes a qualitative coding process: the first author transcribes and open-codes the seminar discussions in NVivo, a second author verifies the opinion cards, and two authors sort cards into challenges. No inter-rater reliability statistic is reported, and the linked repository (footnote 1) contains session topics only, not transcripts or opinion cards. Section 6 concedes the 24-participant discussion may not cover all current LLM4SE challenges. These are acceptable limitations for a position paper, but they make the central claim—that the paper 'achieve[s] 26 key challenges from seven aspects'—hard to audit. The paper also contains internal inconsistencies about the same numbers: Section 4's introduction says the challenges are introduced 'from six aspects,' while the abstract and introduction say seven; Section 7's conclusion says 'We present seven challenges,' while the abstract claims 26 key challenges. If a reader cannot trust the counts in the paper's own text, the taxonomy's credibility is weakened independent of the qualitative method.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper reports a qualitative study of current challenges in LLM-based software engineering (LLM4SE). The authors organized a 24-participant seminar under the CCF Beautiful Lake Seminars, transcribed and open-coded the discussions in NVivo, and used open card sorting to derive 26 key challenges across seven aspects: software requirement and design, coding assistance, testing code generation, code review, software maintenance, software vulnerability management, and data, training, and evaluation. The paper also provides background material on the LLM4SE pipeline (data construction, fine-tuning, prompting, SE-specific LLMs) and a related-work survey organized by the same seven areas, and it concludes with a research roadmap and a threats-to-validity section.","tokens_in":34526,"tokens_out":2787,"duration_ms":24678,"significance":"If the challenge taxonomy is accepted, the paper offers a useful, expert-informed roadmap for LLM4SE research and practice, and it usefully spans the whole software development life cycle. The authors follow recognizable qualitative procedures, explicitly acknowledge completeness and representativeness threats in Section 6, and make session topics publicly available. The main value is as a position/roadmap paper rather than as a hypothesis-testing empirical study; its central contribution is the challenge list itself, which makes the consistency and auditability of that list the deciding factors for the paper's credibility.","major_comments":[{"comment":"The paper's central claim is the count and organization of challenges, but the manuscript states inconsistent numbers: the abstract and Section 1 claim 26 challenges from seven aspects; Section 4's introductory paragraph says the challenges are introduced 'from six aspects'; and Section 7 says 'We present seven challenges.' Because the taxonomy and its counts are the paper's main contribution, these inconsistencies must be reconciled and the final counts made unambiguous.","section":"Abstract, Section 4, Section 7"},{"comment":"The qualitative coding procedure is described at a high level (transcription, open coding in NVivo, verification by a second author, card sorting by two authors, review by two more), but no inter-rater reliability statistic, codebook, opinion cards, or raw transcripts are provided. The linked repository (footnote 1) contains only session topics. Without such an audit trail, a reader cannot independently verify that the reported 26 challenges and their grouping into seven aspects are supported by the discussed material rather than by the authors' post-hoc synthesis.","section":"Section 3"},{"comment":"The paper itself concedes that the challenge list may not cover all current LLM4SE challenges and that representativeness is addressed only by participant diversity. This is a reasonable limitation for a position paper, but the abstract's phrasing 'we achieve 26 key challenges from seven aspects' overstates the completeness of a list explicitly derived from 24 participants in a single seminar. The claims should be qualified to match the admitted scope, or additional evidence of saturation should be provided.","section":"Section 6"}],"minor_comments":[{"comment":"There are numerous typos and misspellings, including 'Technolgy' in the author affiliation, 'through discussion' for 'thorough discussion' in the abstract, 'techniuqes' in Section 2, 'specilizing' in Section 3, 'assitance' in Section 2.6, 'datasety' in the Section 4.4 summary, and 'to to understand' in Section 4.5. A careful proofreading pass is needed.","section":"Throughout"},{"comment":"The seminar dates are given as '19 Jan 2014 - 21 Jan 2024'; the year '2014' is a typo and should be '2024'.","section":"Section 3"},{"comment":"The citation 'CodeX [ ? ]' is an unresolved placeholder and must be corrected to the proper reference.","section":"Section 5.2"},{"comment":"The ACM Reference Format block contains placeholder text ('Make sure to enter the correct conference title', 'Conference acronym ’XX', 'XXXXXXX.XXXXXXX') and a template year of 2018; these should be replaced with the actual venue and metadata.","section":"ACM Reference Format"},{"comment":"The naming of models is inconsistent (e.g., 'CodeLlama' vs. 'CodeLLaMa' in Table 1 and text); please unify the spelling.","section":"Section 2.4"},{"comment":"Several entries in Table 1 have uncertain or missing values marked '—'; consider adding a note explaining what '—' means and whether the model size or pre-training data size is intentionally omitted.","section":"Table 1"}],"recommendation":"major_revision","confidential_remarks":"The paper's core contribution is a qualitative taxonomy, and the main concerns are internal consistency of the reported counts and the lack of an audit trail for the coding process; both are fixable within the manuscript's scope if the authors correct the numbers and add more methodological detail or explicitly reposition the paper as a position/roadmap article. The reference list includes many works by the authors' own research group, which is not itself a problem but may warrant a check that the related work is not over-weighted toward that group."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this is a workshop-synthesis position paper, not a new research result, and the 26-challenge taxonomy is a useful but loosely verified organizing device. It deserves a serious referee, but only if the authors fix a set of glaring internal inconsistencies and release the raw coding materials.\n\nWhat it does well: the paper is a rare structured attempt to capture where LLM4SE currently hurts. The qualitative method (transcription, open coding, card sorting) is standard for community-roadmap studies and the authors are transparent about the seminar structure and participant mix (24 people, 17 academic, 7 industry, six thematic sessions). The seven aspects mirror the SDLC plus data/training/eval, and the individual challenges (hallucination, benchmark saturation, data quality, review automation, etc.) are plausible and well cited. If you need a compact map of the field's open problems, this is a decent starting point.\n\nThe soft spots are mostly about trust in the packaging. The paper's own counts do not line up: the abstract and intro say 26 challenges from seven aspects, Section 4's intro says six aspects, and Section 7's conclusion says 'seven challenges.' There is also a date typo ('19 Jan 2014') in Section 3. None of these are fatal, but a reader should not have to reverse-engineer the structure. More substantively, the audit trail is thin: no transcripts, no opinion cards, no inter-rater reliability, and the linked repository only holds session topics. The authors admit in Section 6 that 24 participants may not cover all challenges, which is honest but underscores that the 26-challenge list is the curated opinion of a specific group. The mild self-citation pattern (several refs from the same labs) is worth noting but not disqualifying; it is a small community.\n\nThe bottom line: this is not a new mechanism or measurement, but it is a credible roadmap if you treat it as one group's structured synthesis. It deserves a serious referee. I would send it out, with a request for the authors to harmonize the counts, fix typos, and either release the coding artifacts or explicitly justify why they cannot.","headline":"Workshop-synthesis roadmap with a useful 26-challenge taxonomy, undermined by internal count inconsistencies and a thin audit trail.","tokens_in":35043,"tokens_out":2954,"would_cite":true,"duration_ms":25153,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper identifies 26 current challenges in applying large language models to software engineering, organized across seven stages of the development life cycle.","keywords":["large language models","software engineering","LLM4SE","software development life cycle","challenges","qualitative study","code generation","vulnerability management"],"falsifier":"Conduct an independent coding of the same seminar transcripts by researchers who did not attend the seminar and were not told the authors' categories; if their independently grouped challenges overlap with the authors' 26 in fewer than half of the items, the taxonomy is a product of the coding team rather than the discussion. Alternatively, a large survey of practitioners rating each of the 26 challenges could show that several are not widely felt.","tokens_in":34172,"feed_emoji":"🤖","tokens_out":6258,"duration_ms":52266,"temperature":0.7,"pith_summary":"This paper sets out to map where large language models still fall short when used in software engineering. Drawing on structured discussions among 24 researchers and practitioners, it identifies 26 challenges organized into seven areas: requirements and design, coding assistance, test generation, code review, maintenance, vulnerability management, and data, training, and evaluation. The authors' goal is to give the field a shared vocabulary and a research roadmap so that effort goes to the problems that actually slow down practice. If the list is right, it explains why current LLM-based tools are not yet trustworthy enough for end-to-end software development and where targeted research would pay off.","feed_headline":"26 challenges now shape LLM-assisted software engineering","feed_subtitle":"A 24-person expert discussion maps where code-generating models still fail, from requirements to vulnerability repair.","key_machinery":"The load-bearing structure of the paper is the challenge taxonomy itself: 26 challenges grouped into seven aspects of the software development life cycle plus model construction. Its companion mechanism is the qualitative coding pipeline, which transcribes six four-hour thematic sessions, converts the discussion into opinion cards through open coding, sorts the cards into candidate challenges, and has additional authors review the final set. The taxonomy is what carries the argument because the paper's contribution is not a new model or dataset but a stable inventory of pain points that future research can target and future evaluations can measure.","core_discovery":"The central claim is that LLM4SE is not blocked by a single bottleneck but by a recognizable constellation of 26 challenges, and that these challenges are stable enough to be named and grouped. The paper derives the list from face-to-face discussions at a three-day seminar, transcribing six thematic sessions and coding the material with a qualitative open-coding procedure followed by open card sorting and author verification. The resulting taxonomy spans the whole software development life cycle, from requirement prompts and domain knowledge, through hallucination, vulnerability injection, and project-level integration in code generation, to syntactic and semantic problems in generated tests, issue-specific code review, microservice dependency complexity in maintenance, vulnerability data scarcity, and unresolved data, training, and evaluation questions. The authors present these challenges as a roadmap: each named challenge points to a concrete research direction.","pith_inferences":["The taxonomy could double as a maturity checklist: a future survey could score each of the 26 challenges as solved, partially solved, or open, turning the qualitative list into a quantitative progress metric.","The absence of any priority ordering among the 26 challenges is a gap the authors do not fill; a ranking by industrial impact would require separate empirical work.","Some of the challenges are tied to current model limitations, such as context windows and training-data scarcity; if those constraints ease, the taxonomy would need revision, which suggests it should be treated as a snapshot rather than a fixed classification.","The same discussion method could be applied to adjacent areas, for example to map challenges in LLM-based DevOps or in using LLMs for embedded systems, where the specific challenge mix may differ."],"forward_implications":["If the taxonomy is correct, research on LLM-based code generation should prioritize hallucination control and project-level integration over further gains on simple benchmarks such as HumanEval and MBPP, whose scores the paper describes as near-limit.","If the taxonomy is correct, test generation research needs to attack syntactic compilability and semantic coverage together, including automatic mocking and assertion oracles.","If the taxonomy is correct, code review automation should be split by issue type and by industry versus open-source practice, rather than treated as one generic task.","If the taxonomy is correct, vulnerability management will not advance without new high-quality vulnerability explanation data and context-slicing methods that fit within LLM context windows.","If the taxonomy is correct, improving LLM4SE requires building evaluation frameworks that reproduce real development contexts instead of relying on benchmark scores."],"supporting_citations":[{"why":"The seminar event whose six thematic sessions and panel discussions generated the qualitative data analyzed in the paper.","marker":"[1]"},{"why":"The qualitative coding software used to conduct open coding and generate opinion cards from the transcribed discussions.","marker":"[38]"},{"why":"One of the two prior studies whose transcribing-and-coding procedure the paper follows for turning seminar discussions into coded opinion cards.","marker":"[62]"},{"why":"The other prior study whose transcribing-and-coding procedure the paper follows for the open-coding step.","marker":"[148]"},{"why":"The card-sorting study that supplies the open card sorting method used to group codes into candidate challenges.","marker":"[79]"}],"fun_headline_variants":["26 challenges revealed for LLM-powered software engineering","Expert panel maps 26 hurdles for LLM coding assistants","From design to fix: 26 challenges in LLM software engineering","LLM software tools face 26 named obstacles, new study says"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The whole list rests on the assumption that the views of 24 invited specialists, filtered through two authors' coding, stand in for the full set of challenges the field faces; the paper itself concedes the list may be incomplete.","fun_headline_variants_meta":{"raw":{"variants":["26 challenges revealed for LLM-powered software engineering","Expert panel maps 26 hurdles for LLM coding assistants","From design to fix: 26 challenges in LLM software engineering","LLM software tools face 26 named obstacles, new study says"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000606,"raw_usage":{"total_tokens":2824,"prompt_tokens":943,"completion_tokens":1881,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":559,"completion_tokens_details":{"reasoning_tokens":1812}},"tokens_in":559,"tokens_out":1881,"duration_ms":11004,"temperature":1.0,"reasoning_tokens":1812,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T12:07:09.216887+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Conduct an independent coding of the same seminar transcripts by researchers who did not attend the seminar and were not told the authors' categories; if their independently grouped challenges overlap with the authors' 26 in fewer than half of the items, the taxonomy is a product of the coding team rather than the discussion. Alternatively, a large survey of practitioners rating each of the 26 challenges could show that several are not widely felt.","supporting_citations":[],"review_version":1}