{"id":"de323b1f-a5c0-47bd-b4b9-9882b07e56de","arxiv_id":"2504.18858","paper_version":1,"verdict":"REJECT","confidence":"MODERATE","novelty_score":2.0,"correctness_risk":"high","formal_verification":"none","parameter_count":5,"one_line_summary":"A multivocal literature review finds ChatGPT's reported error rates range from single digits to over 80 percent depending on domain and task, yet its synthesized ranges are not backed by a released dataset.","lead":"This paper collects reported error rates of ChatGPT across healthcare, business, engineering, and software development, finding they vary widely by task and model version. It offers a cautionary map of where ChatGPT can be trusted, but its own numbers are partly unverifiable estimates.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The paper's quantitative core is not reproducible: heterogeneous performance metrics are converted to a single 'error rate' scale, and several headline SDLC-phase ranges are labelled 'estimated' without derivation from the cited studies.","rationale":"The reader's verdict of REJECT is appropriate, and the weakest assumption identified by the reader matches my own: heterogeneous metrics are treated as directly convertible to a single error-rate scale, and the 'estimated' ranges in the SDLC sections are not tied to traceable data. The paper's own limitations section explicitly acknowledges that consistent error rates had to be inferred, yet no inference rules or raw data are given. This is not an internal contradiction, but it is a fatal reproducibility gap for the paper's quantitative central claim. The qualitative takeaway—ChatGPT has non-negligible, task-dependent error rates and needs human oversight—is plausible and broadly consistent with the wider literature, so the paper may have some value as a narrative summary. However, the boxplots and specific percentages are presented as the main evidence-based contribution, and they cannot be verified from the manuscript. A focused data-extraction and source-verification check would settle whether the numbers are real measurements or author interpolations. Because the reader's rejection rests on exactly this unverifiable quantitative foundation, no change to the verdict is needed.","tokens_in":7862,"tokens_out":2776,"duration_ms":29490,"concrete_test":"Build a complete data-extraction table from the 32 references: for every quantitative claim in Sections 3 and 4, record the cited reference, the exact reported metric, the page or section where it appears, and the transformation used to obtain the claimed error rate. Then attempt to regenerate Figures 1 and 2 from that table. Specifically, retrieve the primary sources for claims such as 72%/77%/83% in healthcare [16], 53%/85% on CPA exams [19], 81.96%/87.5% on Python tasks [24], 52% on Stack Overflow [25], and the 'estimated' 5–20%, 10–30%, 30–50%, and 10–20% ranges in Sections 4.2–4.5 [5, 6, 7, 14, 25, 26, 27, 28]. If any figure is absent from the cited source, differs materially, or refers to a different metric, the synthesis is not reproducible and the quantitative ranges should not be reported as measurements.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central quantitative claim—domain- and phase-specific error ranges such as 5–20% for requirements/design and 10–50% for coding/testing/maintenance—rests on an ungrounded synthesis. Section 2.2 states that where accuracy was reported, error rates were inferred as (100% – accuracy), but the primary studies use fundamentally different metrics: exam pass rates, grade rubrics, code test-case pass rates, and even user acceptance ratings. These are not commensurable; a 47% error rate on a CPA exam is not the same kind of quantity as an 18% test-case failure rate. Pooling them into a single boxplot in Figures 1 and 2 creates distributions with no well-defined meaning. The SDLC-phase ranges are further weakened by explicit wording: Sections 4.2, 4.4, and 4.5 describe 'estimated error rates' (e.g., 'Estimated error rates in preliminary design suggestions ranged between 5% and 20%') without showing how those estimates were computed from the cited sources. The limitations section concedes that 'we had to infer consistent error rates in some cases,' but no mapping, extraction table, or dataset is provided, so the aggregate numbers cannot be checked. Some attributions also look unreliable: reference [14] is a multitask prompting paper, not a study of ChatGPT oversimplifying software designs, and [24] is cited for Python test-case success rates despite its title being about secure software development. If the primary numbers cannot be located in the cited references, the paper's quantitative contribution is unfalsifiable. The qualitative warning that ChatGPT is imperfect and needs human oversight is credible and consistent with prior evidence, but that message does not require the unsupported quantitative ranges that form the paper's stated contribution.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This manuscript reports a multivocal literature review (MLR) of ChatGPT error rates across broad domains and across software engineering lifecycle (SDLC) phases. The author compiles reported accuracy/error figures from academic and grey sources, converts accuracy to error rates where needed, groups results by domain (healthcare, business, economics, engineering, computer science) and SDLC phase (requirements, design, implementation, testing, maintenance), and visualizes the distributions in boxplots. The headline findings are that error rates are substantial and variable, for example 8–83% in healthcare, 5–20% in requirements/design, and 10–50% in implementation/testing/maintenance, with GPT-4 generally outperforming GPT-3.5. The conclusion recommends human oversight and validation before relying on ChatGPT outputs in professional settings.","tokens_in":8322,"tokens_out":4893,"duration_ms":45456,"significance":"The question addressed is timely and practically important, and the qualitative observation that LLM error rates vary by domain and task and that human oversight is needed is plausible and broadly consistent with the wider literature. If the quantitative synthesis were reliable, the paper could serve as a useful reference for practitioners deciding where to invest in human-in-the-loop review. However, the paper's contribution rests on a quantitative synthesis whose method is not reproducible and whose key ranges are partly based on unstated author estimates; the quantitative claims are not currently supported. The paper explicitly acknowledges several threats to validity in Section 5.2, which is commendable, but those limitations directly undermine the central numeric results rather than merely qualifying them.","major_comments":[{"comment":"The rule 'Where accuracy rather than error rates were reported, we inferred error rates as (100% – accuracy %)' is applied to studies whose metrics are not commensurable: exam pass rates, grade rubrics, code test-case pass rates, user acceptance ratings, and diagnostic accuracy. A 47% CPA-exam error rate is not the same kind of quantity as an 18% test-case failure rate or a 52% 'incorrect or partially incorrect' Stack Overflow answer rate, so pooling them into Figures 1 and 2 produces distributions without a well-defined interpretation. This is load-bearing because the headline domain- and phase-specific ranges are drawn from these pooled boxplots.","section":"§2.2"},{"comment":"The review is not reproducible as reported. The text does not provide the search strings, the databases and grey-literature sources with search dates, the number of records retrieved, screened, and included, the inclusion/exclusion decisions, or a data extraction table linking each data point to its source and metric. Without these, a reader cannot verify that the boxplots reflect the stated evidence base. This falls below the reporting standard expected for an MLR and for any quantitative synthesis.","section":"§2.1–2.2"},{"comment":"The central SDLC-phase ranges are introduced as 'estimated' without derivation: for example, 'Estimated error rates in preliminary design suggestions ranged between 5% and 20%' (§4.2), 'estimated error rates of around 10–30%' for testing (§4.4), and a code-review miss rate of 'approximately 30–50%' (§4.5). The cited paragraphs describe qualitative observations about oversimplification, missing boundary cases, and missed subtle flaws, but no calculation, mapping, or source-specific numbers are given in the text or in any supplementary table. Since these estimates are the quantitative basis of Figure 2 and of the conclusion that requirements/design are safer than implementation/testing/maintenance, the claim is not supported by the cited evidence as presented.","section":"§4.2, §4.4, §4.5"},{"comment":"Several reference assignments are not credible. Reference [14], cited in §2.2, §4.2, and §5.2, is titled 'Multitask Prompted Training Enables Zero-Shot Task Generalization' and does not appear to be a study of ChatGPT oversimplifying software designs; reference [24], titled 'Evaluating LLMs for Secure Software Development,' is cited in §3.5 and §4.3 for Python test-case success rates (81.96% and 87.5%). If these data points cannot be located in the cited sources, the quantitative claims drawn from them are unverifiable. The author should provide exact locations (table or figure numbers) or replace the citations.","section":"§3.5, §4.2, references"},{"comment":"The limitations section concedes that 'we had to infer consistent error rates in some cases' and that model versions were sometimes inferred from publication dates or context, but no sensitivity analysis or robustness check is provided. For a paper whose main output is numeric error ranges, the effect of these inferences and of the acknowledged reporting inconsistencies must be quantified or at least bounded; otherwise the ranges cannot be distinguished from the author's priors.","section":"§5.2"}],"minor_comments":[{"comment":"The reference list contains duplicate entries: [1] and [16], [2] and [18], [3] and [19], [4] and [20], [8] and [24], and [9] and [25] are the same works. This will confuse readers and suggests the list was not carefully curated.","section":"References"},{"comment":"Reference [15] (Tukey, Exploratory Data Analysis) is listed but not cited in the text; please cite it where boxplots are discussed or remove it.","section":"References"},{"comment":"The figures are not reproducible from the text alone: no data points, sample sizes, or per-study values are given. Even if the figures appear correctly in the compiled PDF, the paper should include a data table or online appendix listing the underlying values.","section":"Figures 1 and 2"},{"comment":"The abstract states 'Engineering tasks averaged 20–30%' while §3.4 reports 25% for GPT-4 on environmental engineering and 40–50% for GPT-3.5 on mechanical engineering; please reconcile or clarify which model and task set the 'average' covers.","section":"Abstract and §3.4"},{"comment":"The section heading appears as '2.2 2.2 Data Extraction and Synthesis' in the manuscript text; the duplicated number should be removed.","section":"§2.2"}],"recommendation":"reject","confidential_remarks":"The manuscript's central quantitative contribution depends on unsupported 'estimated' ranges and non-reproducible pooling of heterogeneous metrics, and the reference list contains duplicate and mismatched entries. A revision would require re-doing the synthesis with a full extraction table and corrected citations, which is beyond a normal revision cycle; I would not recommend major revision."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"What should I tell you about arXiv:2504.18858? It's a review paper that aggregates reported ChatGPT error rates across domains and software engineering lifecycle phases. The qualitative core is sensible and the high-level warning—don't fully trust ChatGPT without oversight—holds up. The quantitative narrative, though, is wobbly. Several headline ranges are labelled 'estimated' with no derivation, and the synthesis converts heterogeneous metrics into a single error-rate scale without enough justification. The limitations section honestly admits that error rates were inferred and model versions guessed, but the paper doesn't give you the underlying data to check any of it.\n\nWhat's genuinely useful: it collects a scattered set of evaluations into one place, groups them by domain and SDLC phase, and draws a distinction between early phases (requirements, design, with lower reported errors) and later phases (implementation, testing, maintenance, with wider variance). That framing is a reasonable organizer for practitioners trying to decide where to put human review. The boxplots are a decent visual summary, although the underlying numbers are shaky. The paper also correctly notes that model upgrades matter and that fluent output can mask incorrect answers.\n\nWhere it's soft: Section 2.2 says error = 100% - accuracy, but the source studies use pass rates, grades, test-case success, and user acceptance. Those aren't interchangeable. Sections 4.2, 4.4, and 4.5 give ranges like 5–20%, 10–30%, 30–50% with the word 'estimated' but no explanation of how those estimates were computed from the cited references. There's no data table, no search string, no screening count. The reference list also has misattributions: [14] is a multitask prompting paper, not evidence about ChatGPT oversimplifying designs, and [24] is about secure software development, not general Python test-case success. Those look like genuine citation errors, and they undercut confidence in the other attributions.\n\nNet: the message is right, but the numbers are not load-bearing in their current form. A reader who wants the qualitative takeaway can get it from the abstract. A reader who wants to rely on the specific ranges can't verify them. If the author were to resubmit with the extraction table, explicit mapping of metrics, and defense of the estimated ranges, the paper would be much stronger. As it stands, it's a useful map with some unreliable coordinates.\n\nMy recommendation: this deserves a serious referee, not a desk reject, but the referee should be told to focus on reproducibility. I would not cite it in its current form, but I'd bring it to the reading group to talk about what counts as evidence in LLM reviews.","headline":"A timely review with a sensible qualitative warning, but the quantitative error-rate synthesis is unverifiable as presented and needs heavy revision before it can be trusted.","tokens_in":8700,"tokens_out":2063,"would_cite":false,"duration_ms":18887,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"ChatGPT's error rates are non-negligible and task-dependent, ranging from about 5% in structured drafting to over 80% in rare-disease diagnosis, so full automation without human oversight is not yet safe.","keywords":["ChatGPT","error rates","large language models","multivocal literature review","software development lifecycle","reliability","human oversight"],"falsifier":"Re-run the synthesis from the raw scored outputs of the cited studies, recording each study's metric type (exact-match accuracy, pass rate, user acceptance) and recomputing errors without the uniform 100%-minus-accuracy shortcut; if the recomputed domain ranges differ materially from the reported 8% to 83% healthcare span or the 5% to 20% versus 10% to 50% lifecycle split, the headline figures collapse.","tokens_in":7661,"feed_emoji":"🤖","tokens_out":9787,"duration_ms":91112,"temperature":0.7,"pith_summary":"This paper is a synthesis of published and grey-literature measurements of ChatGPT's error rates. It aggregates studies across healthcare, business, economics, engineering, computer science, and software engineering, converting reported success rates into error rates and grouping them by domain and by software development phase. The central claim is that ChatGPT's errors are non-negligible and unevenly distributed: typical ranges run from roughly 8% to 83% in healthcare, about 5% to 20% in early design work, and 10% to 50% in coding, testing, and maintenance, with newer model versions consistently lowering but not eliminating errors. A fair reader should come away with a concrete map of where ChatGPT output can serve as a draft and where it must be checked by a human.","feed_headline":"ChatGPT's error rates range from 5% to 83%, review finds","feed_subtitle":"A synthesis across disciplines finds every software task still needs human validation.","key_machinery":"The carrying mechanism is the multivocal literature review itself, specifically its synthesis rule: each included study's reported accuracy, pass rate, or success rate is converted to an error rate by computing $(100\\% - \\text{accuracy})$, and the resulting data points are grouped into boxplots per domain and per software development lifecycle phase. Those boxplots carry the argument: they let the paper claim visible patterns in the spread and center of reported errors—narrow, lower boxes for requirements and design versus wider, higher boxes for coding, testing, and maintenance—rather than relying on any single benchmark.","core_discovery":"The paper's central discovery, stated as a synthesized empirical result, is that ChatGPT's reliability varies sharply with task, domain, and model version, and that no domain or software development lifecycle phase is error-free. The concrete numbers are: 28% error in GPT-3.5 clinical decision-making and 23% in GPT-4 final diagnosis, rising to 83% in rare-disease diagnosis; accounting exam errors falling from roughly 47% with GPT-3.5 to 15% with GPT-4; an economics midterm error rate dropping from 69% to 27%; coding success up to 87.5% yet 52% incorrect answers on real-world programming questions; and software engineering requirements and design phases at roughly 5% to 20% errors versus 10% to 50% for implementation, testing, and maintenance. The paper concludes that full reliance on ChatGPT without human oversight remains risky, especially in high-stakes settings.","pith_inferences":["A natural extension the paper does not spell out is a risk-tiering rule: allocate the heaviest human review to debugging, merge, and refactoring tasks, whose observed error ranges are roughly twice those of requirements drafting.","The conversion rule used here probably understates true errors on multiple-choice tasks where guessing or partial credit inflates accuracy, so the reported lower bounds should be read as optimistic in those settings.","If the version-upgrade trend continues, the same benchmarks should show GPT-5-era models pushing common-condition diagnostic errors below 20% while rare-disease and open-ended debugging errors stay above 30%; that is a testable prediction.","The divergence between objective correctness and user-rated quality reported in the underlying studies implies that error-rate tracking and user-trust tracking should be run side by side, not separately."],"forward_implications":["Treat every ChatGPT output as a draft that enters a validation pipeline, with review intensity scaled to the observed error range of the task.","Prioritize human review in healthcare diagnosis and in coding, testing, and maintenance, where the reported error ceilings are highest.","Version upgrades from GPT-3.5 to GPT-4 lower error rates but do not make any phase safe, so upgrade decisions should not replace oversight.","Structured, context-rich prompts are a measurable lever for reducing errors and should be treated as part of the engineering process.","Because fluent wrong answers are sometimes rated as high quality, user confidence in ChatGPT output is not a reliability signal."],"supporting_citations":[{"why":"Supplies the healthcare accuracy figures, including the 83% rare-disease diagnosis error ceiling.","marker":"[16]"},{"why":"Documents the accounting exam error drop from about 47% with GPT-3.5 to 15% with GPT-4.","marker":"[19]"},{"why":"Provides the economics midterm error rates of 69% for GPT-3.5 and 27% for GPT-4.","marker":"[20]"},{"why":"Reports Python programming success rates of 81.96% and 87.5% that anchor the coding-phase numbers.","marker":"[24]"},{"why":"Shows 52% incorrect or partially incorrect answers on real-world programming questions and user overtrust in fluent wrong answers.","marker":"[25]"},{"why":"Supplies the requirements-engineering claim that ChatGPT can draft usable requirements with occasional 5% to 15% error risk.","marker":"[5]"},{"why":"Provides the testing-phase estimate that generated unit tests are structurally correct but incomplete, around 10% to 30% errors.","marker":"[6]"},{"why":"Supports the design-phase evidence that ChatGPT helps with common architectures but misses constraints in specialized systems.","marker":"[7]"},{"why":"Documents debugging hallucinations and code-review misses that set the maintenance-phase risk range.","marker":"[27]"},{"why":"Establishes that ChatGPT behavior drifts over time, supporting the paper's call for continuous evaluation.","marker":"[32]"}],"fun_headline_variants":["ChatGPT errors range from 5% to 83% by task","Study: ChatGPT's error rates swing from 5% to 83%","ChatGPT still errs up to 83% in some tasks, review finds","Don't fully trust ChatGPT: errors up to 83% in review"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The synthesis assumes that every reported accuracy, pass rate, or success rate can be read as a direct error rate by subtracting it from 100%, and that the phase-level ranges labeled 'estimated' rest on the cited studies rather than the paper's own interpolation; if that conversion and those estimates are wrong, the boxplot ranges are not reliable measurements.","fun_headline_variants_meta":{"raw":{"variants":["ChatGPT errors range from 5% to 83% by task","Study: ChatGPT's error rates swing from 5% to 83%","ChatGPT still errs up to 83% in some tasks, review finds","Don't fully trust ChatGPT: errors up to 83% in review"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000772,"raw_usage":{"total_tokens":3489,"prompt_tokens":1090,"completion_tokens":2399,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":706,"completion_tokens_details":{"reasoning_tokens":2326}},"tokens_in":706,"tokens_out":2399,"duration_ms":18609,"temperature":1.0,"reasoning_tokens":2326,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T10:06:51.288566+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run the synthesis from the raw scored outputs of the cited studies, recording each study's metric type (exact-match accuracy, pass rate, user acceptance) and recomputing errors without the uniform 100%-minus-accuracy shortcut; if the recomputed domain ranges differ materially from the reported 8% to 83% healthcare span or the 5% to 20% versus 10% to 50% lifecycle split, the headline figures collapse.","supporting_citations":[{"cited_title":"Assessing the accuracy of GPT-3.5 and GPT-4 in clinical decision-making,","cited_arxiv_id":null,"evidence_quote":"Supplies the healthcare accuracy figures, including the 83% rare-disease diagnosis error ceiling."},{"cited_title":"ChatGPT’s accounting exam results reveal striking improvement,","cited_arxiv_id":null,"evidence_quote":"Documents the accounting exam error drop from about 47% with GPT-3.5 to 15% with GPT-4."},{"cited_title":"GPT-3.5 versus GPT-4 on my midterm exam,","cited_arxiv_id":null,"evidence_quote":"Provides the economics midterm error rates of 69% for GPT-3.5 and 27% for GPT-4."},{"cited_title":"Evaluating LLMs for Secure Software Development,","cited_arxiv_id":null,"evidence_quote":"Reports Python programming success rates of 81.96% and 87.5% that anchor the coding-phase numbers."},{"cited_title":"Amplitude modulation of acoustic waves in accelerating flows quantified using acoustic black and white hole analogues","cited_arxiv_id":"2308.00064","evidence_quote":"Shows 52% incorrect or partially incorrect answers on real-world programming questions and user overtrust in fluent wrong answers."},{"cited_title":"Exploring LLMs for Software Requirements Engineering,","cited_arxiv_id":null,"evidence_quote":"Supplies the requirements-engineering claim that ChatGPT can draft usable requirements with occasional 5% to 15% error risk."},{"cited_title":"Text-Blueprint: An Interactive Platform for Plan-based Conditional Generation","cited_arxiv_id":"2305.00034","evidence_quote":"Provides the testing-phase estimate that generated unit tests are structurally correct but incomplete, around 10% to 30% errors."},{"cited_title":"AI in Software Architecture: Opportunities and Risks,","cited_arxiv_id":null,"evidence_quote":"Supports the design-phase evidence that ChatGPT helps with common architectures but misses constraints in specialized systems."},{"cited_title":"How Is ChatGPT’s Behavior Changing Over Time?","cited_arxiv_id":null,"evidence_quote":"Establishes that ChatGPT behavior drifts over time, supporting the paper's call for continuous evaluation."}],"review_version":1}