REVIEW 3 major objections 6 minor 7 references
Assessing GPTZero's Accuracy in Identifying AI vs. Human-Written Essays
T0 review · 3 major / 6 minor · reviewed 2026-08-06 · deepseek-v4-flash
Pith's one-line read GPTZero almost always detects AI essays but misfires on human writing.
desk verdict Small, transparent GPTZero benchmark with a genuinely new length-axis, but the unstated binary threshold makes the headline accuracy claim unverifiable. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is GPTZero's continuous per-essay output: a percentage chance that the text was written by generative AI. The authors feed each essay into that system, sort the results into short, medium, and long length bands, average the percentages per band, and then reduce the output to binary 'Predicted as Human' versus 'Predicted as AI' labels in a confusion matrix. The argument compares those observed averages and counts against an ideal 0%-for-human, 100%-for-AI baseline. The binary labels in Table 3 require an implicit cutoff, which the paper does not state.
What would settle it
Re-run the same seventy-eight essays through GPTZero and record the reported percentage for each one; if setting the AI label threshold at 50%, 80%, or 95% changes the eight-false-positive count, then the paper's confusion matrix is an artifact of whatever default GPTZero used, not a stable finding.
Extended reading notes
Core claim
The central claim is that GPTZero can almost always correctly identify AI-generated texts, at roughly 90-99% accuracy, while its judgments on human-written essays fluctuate enough to produce eight false positives among fifty essays. The paper reports average GPTZero 'percent chance written by AI' scores of 99.17% for short AI essays, 97.00% for medium, and 98.83% for long, versus 35.56%, 10.29%, and 14.75% for the corresponding human essays. In the confusion matrix, all twenty-eight AI essays are classified as AI and forty-two of fifty human essays as human, yielding an overall error rate of 10.3% and a false-positive rate of 16%.
Load-bearing premise
The load-bearing premise is that GPTZero's hidden threshold for labelling an essay 'Predicted as AI' is fixed and consistent across all seventy-eight submissions, because the paper never states that cutoff.
Editorial extensions
If this is right
- If GPTZero's near-perfect detection of AI-generated essays is reliable, instructors can use it to catch unmodified ChatGPT-style submissions with high confidence.
- If the 16% false-positive rate on human essays holds, then in a class of fifty honest submissions about eight would be wrongly flagged, so a positive GPTZero result alone is not enough to accuse a student.
- Length does not fix the problem: short and long human essays drew more false positives than medium ones, so no simple word-count cutoff rescues detector precision.
- The authors' own suggested next step is to test mixed human-and-AI texts, because real student work is rarely either fully AI or fully human.
Reading between the lines
- Beyond the paper: the confusion matrix's 42/8 and 28/0 splits depend on an unreported GPTZero threshold; changing that threshold will move human essays across the line, so the specific numbers should not be treated as fixed detector properties.
- Beyond the paper: the high AI-detection rate likely reflects unmodified ChatGPT output; edited, paraphrased, or translated AI text would probably score lower, so real-world accuracy on polished submissions may be weaker than the 90-99% range suggests.
- Beyond the paper: a direct testable extension is to vary GPTZero's confidence cutoff on the same seventy-eight texts and report precision-recall curves, which would let educators choose a threshold that trades missed AI for fewer false accusations.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper reports an observational study of GPTZero's ability to distinguish AI-generated from human-written essays. The authors collected 28 essays generated with ChatGPT and 50 human-written student essays, submitted them to GPTZero, and recorded the resulting AI-percentage scores. They grouped essays into three length bins (short, medium, long), reported mean AI-percentages for each group (Table 1), and constructed a confusion matrix (Table 3) in which GPTZero classified all 28 AI essays correctly and misclassified 8 of 50 human essays as AI. The authors conclude that GPTZero is highly effective at detecting unmodified AI-generated text (around 90-99% accuracy) but unreliable for human-authored essays, and they caution educators against relying solely on such detectors.
Significance. If the findings are valid, they provide a useful, concrete data point on GPTZero's performance in an educational context, corroborating prior work (e.g., Walters 2023; Liu et al. 2024) that shows high detection of unmodified ChatGPT text but nontrivial false-positive rates for human writing. The study's strengths are its simple, reproducible methodology, its direct comparison of length categories, and its clear reporting of the confusion matrix and error rates. The paper also situates its results against a literature review of AI detector evaluations. However, the study's significance is limited by the small sample (78 essays), the absence of statistical uncertainty measures, and a load-bearing methodological omission (the unstated binary threshold used to convert continuous scores into classifications). These issues are fixable within the manuscript's scope, but they must be addressed before the central claims can be fully evaluated.
major comments (3)
- [Section 5, Table 3] The confusion matrix in Table 3 ('Predicted as Human' vs. 'Predicted as AI') requires converting GPTZero's continuous percentage scores into binary labels, but the paper never states the cutoff used for this conversion. The reported 42/8 split for human essays, the 16% false-positive rate, and the 10.3% overall error rate all depend on this hidden threshold. While the AI row is insensitive because all AI scores were at or above 97%, the human row is not; for instance, a threshold of 50% versus 90% would likely produce different false-positive counts. Please state the exact threshold (or provide the raw per-essay percentages) so that Table 3 and the derived error rates can be independently verified.
- [Section 5, Table 1 and text] There is a numeric inconsistency: the text states that 'for all AI categories, this chance never went below 98%,' but Table 1 reports 97.00% for the ChatGPT medium-length essays. Additionally, the conclusion's 'around 90-99% accuracy rate' does not match either the averaged AI-percentages (99.17%, 97.00%, 98.83%) or the confusion matrix accuracy (70/78 correct, 89.7%). Please reconcile these numbers and explicitly define how 'accuracy rate' is computed.
- [Section 5, Results and Conclusion] The paper's secondary claim about essay length—that short and long human essays produce more false positives than medium essays—is not supported by any statistical analysis. The results report only cell means without sample sizes per cell, standard deviations, confidence intervals, or significance tests. Consequently, the statements in the Conclusion that 'short and long human-generated texts... were found to be very inaccurate' and that 'medium length texts were found to be identified the most accurately' are not substantiated. Please provide per-cell counts, a measure of dispersion, and an appropriate test or conservative interpretation.
minor comments (6)
- [Section 3, Methodology] The length bins are defined with overlapping boundaries (0-100, 100-350, 350-800); please specify how exactly 100-word and 350-word essays are assigned.
- [Abstract and Section 5] The abstract states that AI essays were detected with '91-100% AI believed generation,' while Table 1 reports values between 97% and 99%. Please align these numbers and clarify which statistic is being summarized.
- [Section 4, References] Reference [7] is dated '1970' in the text; the chapter on machine-obfuscated plagiarism appears to be from a 2020 volume. Please correct the year.
- [Section 5, Results] The words 'exceptional' and 'dreadful' are informal; consider neutral phrasing such as 'highly accurate' and 'less reliable.'
- [Section 3, Methodology] The paper does not report the version of GPTZero used, the date of testing, or whether the free or paid tier was used; including this information is essential for reproducibility.
- [Section 6, Conclusion] The sentence 'the constant fluctuations in the data indicate that medium length texts cannot be accurately predicted every time' is vague; predictions are made for individual texts, not length groups, so please clarify the intended meaning.
Circularity Check
The AI-detection result is partly self-referential because the AI test set is ChatGPT output, the very class GPTZero is built to detect; the human false-positive finding remains independent.
-
self definitional
[Section 2 (Introduction) and Section 3 (Methodology); see also Table 3]
"we used GPTZero, made by the makers of ChatGPT. ... For the AI essays, they were generated using random prompts using ChatGPT."
The paper operationally defines 'AI generated texts' as ChatGPT-generated essays, and GPTZero is a detector produced by the same organization and designed to flag ChatGPT-style output. Consequently, the AI condition is the target distribution GPTZero was built to recognize, not an independent sample of 'AI text' in general. The near-perfect AI row in Table 3 (28/0) and the average scores in Table 1 (97.00-99.17%) therefore measure GPTZero recognizing its own intended input class by construction. The conclusion that 'GPTZero can almost always correctly identify AI generated texts, at around 90-99% accuracy rate' reduces, for this test set, to 'GPTZero identifies ChatGPT outputs,' which is its training objective.
full rationale
The paper contains no fitted parameter, no self-citation chain, and no derivation that is equivalent to its inputs. The only genuine circularity concern is the benchmark design: the AI essays were generated by ChatGPT, the same model family GPTZero targets, so the 100% AI-detection row in Table 3 is partly built into the definition of the test set. That is a self-definitional element, though the human-essay results are independent and provide real evidence about false positives. The unstated binary threshold used to convert GPTZero's continuous percentages into the 'Predicted as Human' vs. 'Predicted as AI' labels in Table 3 is a serious reproducibility gap, but it is not a circularity step because no threshold is fitted and reported as a prediction. Overall, partial circularity in the AI-detection claim, but independent content in the false-positive finding, warrants a score of 4 rather than a higher score.
Assumptions & free parameters
free parameters (3)
- Length bin cutoff at 100 words =
100 words
- Length bin cutoff at 350 words =
350 words
- Binary classification threshold for 'predicted as AI' =
not stated
assumptions (3)
- domain assumption Human essays were written without AI assistance.
- domain assumption ChatGPT-generated texts are representative of AI-generated essays that GPTZero should flag.
- domain assumption GPTZero's percentage output can be treated as a direct probability that a text was AI-written.
Cite this review
Pith. "Pith review of Assessing GPTZero's Accuracy in Identifying AI vs. Human-Written Essays." pith.science (2026). https://pith.science/paper/XVSBGYPQ
@misc{pith2026250623517,
author = {Pith},
title = {Pith review of: Assessing GPTZero's Accuracy in Identifying AI vs. Human-Written Essays},
year = {2026},
howpublished = {\url{https://pith.science/paper/XVSBGYPQ}},
note = {Machine review of arXiv:2506.23517}
}
read the original abstract
As the use of AI tools by students has become more prevalent, instructors have started using AI detection tools like GPTZero and QuillBot to detect AI written text. However, the reliability of these detectors remains uncertain. In our study, we focused mostly on the success rate of GPTZero, the most-used AI detector, in identifying AI-generated texts based on different lengths of randomly submitted essays: short (40-100 word count), medium (100-350 word count), and long (350-800 word count). We gathered a data set consisting of twenty-eight AI-generated papers and fifty human-written papers. With this randomized essay data, papers were individually plugged into GPTZero and measured for percentage of AI generation and confidence. A vast majority of the AI-generated papers were detected accurately (ranging from 91-100% AI believed generation), while the human generated essays fluctuated; there were a handful of false positives. These findings suggest that although GPTZero is effective at detecting purely AI-generated content, its reliability in distinguishing human-authored texts is limited. Educators should therefore exercise caution when relying solely on AI detection tools.
Reference graph
Works this paper leans on
-
[1]
Walters, W. H. (2023, January 1). The effectiveness of software designed to detect AI- generated writing: A comparison of 16 AI text detectors. De Gruyter. https://www.degruyter.com/document/doi/10.1515/opis-2022-0158/html
-
[2]
Liu, J. Q. J., Hui, K. T. K., Zoubi, F. A., Zhou, Z. Z. X., Samartzis, D., Yu, C. C. H., Chang, J. R., & Wong, A. Y. L. (2024, May 20). The Great Detectives: Humans versus AI detectors in catching large language model-generated medical writing - International Journal for Educational Integrity. BioMed Central. https://edintegrity.biomedcentral.com/articles...
-
[3]
Elkhatat, A. M., Elsaid, K., & Almeer, S. (2023, September 1). Evaluating the efficacy of AI content detection tools in differentiating between human and ai-generated text - international journal for educational integrity. BioMed Central. https://edintegrity.biomedcentral.com/articles/10.1007/s40979-023-00140-5
-
[4]
Popkov, A. A., & Barrett, T. S. (2024). AI vs Academia: Experimental study on AI text detectors’ accuracy in behavioral health academic writing. Accountability in Research. Advance online publication. https://doi.org/10.1080/08989621.2024.2331757
arXiv 2024
-
[5]
An Empirical Study of AI Generated Text Detection Tools
Akram, Arslan. An Empirical Study of AI Generated Text Detection Tools, arxiv.org/pdf/2310.01423. Accessed 1 Sept. 2024
work page Pith review arXiv 2024
-
[6]
Perkins, Mike, et al. “Detection of GPT-4 Generated Text in Higher Education: Combining Academic Judgement and Software to Identify Generative AI Tool Misuse - Journal of Academic Ethics.” SpringerLink, Springer Netherlands, 31 Oct. 2023, link.springer.com/article/10.1007/s10805-023-09492-6
-
[7]
Detecting Machine-Obfuscated Plagiarism
Foltýnek, Tomáš, et al. “Detecting Machine-Obfuscated Plagiarism.” SpringerLink, Springer International Publishing, 1 Jan. 1970, link.springer.com/chapter/10.1007/978-3-030-43687-2_68
Reviewed August 6, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.