REVIEW 4 major objections 4 minor 12 references
The Checking Problem: What must be true before AI ships in a regulated firm
T0 review · 4 major / 4 minor · reviewed 2026-08-03 · deepseek-v4-flash
Pith's one-line read AI pilots clear demonstrations, but 43.9% of those that pass cannot be placed in production in regulated firms, and the deciding factor is how much output a human must still check.
desk verdict A transparent, reproducible measurement of demo-to-production survival and review burden, with a few fixable mistakes in the headline claims. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
Two linked instruments. The first is the pair of acceptance bars: the demonstration bar (one run, one case, all correct) and the production bar (four measurable criteria: P1 accuracy ≥0.99 per element, P2 reproducibility of identical input/output, P3 groundedness of citations ≥0.95, P4 detectability of confidence AUC ≥0.70). The second is the review-burden costing, which computes the share of output a reviewer must inspect when triaging by stated confidence, with the threshold calibrated out of sample ('operational') rather than with hindsight ('oracle'). The asymmetry between the bars is the argument's load-bearing feature: P2, P3, P4 are invisible to a single demonstration.
What would settle it
Run the same six workflows with a panel of human reviewers who have full access to the documents and can use any signals they like, and compare their actual review patterns and missed errors against the paper's predicted 49% skip rate for governed configurations. If reviewers cannot safely skip half the output without exceeding a 1% residual error rate, or if more than 3 of 20 configurations break the tolerance when the calibrated threshold is applied to new cases, the central burden claim is falsified. Equivalently, a replication using a different confidence calibration procedure that fails t
Extended reading notes
Core claim
The paper's central claim is that a 'demonstration bar' (one correct run on one case) and a 'production bar' (sustained accuracy at least 0.99, reproducibility across repeats, grounded citations with threshold 0.95, and a confidence signal separating correct from incorrect elements with AUC at least 0.70) test different properties, and that the second tests properties the first cannot see. Measured across six workflows, four model families, three tool configurations and three repeats, 57 of 72 configurations cleared the demonstration bar and only 32 cleared the production bar, a survival rate of 56.1%. Among the 25 that failed production, accuracy was implicated 22 times but was the sole cau
Load-bearing premise
The 49% review-burden result assumes a reviewer triages solely by the model's stated confidence, calibrates the threshold on a pilot sample, and can safely skip the least-confident half of output while keeping the residual error under 1%; if real reviewers use other signals, or if the calibrated threshold does not transfer to unseen cases, the headline reduction is an artifact of the costing model.
Editorial extensions
If this is right
- If the claim is right, pilot acceptance criteria in regulated firms should include repeated runs and a check on whether cited passages exist, not just a single correct output.
- Vendors should be asked to report detectability (whether confidence separates correct from wrong) and groundedness, not just accuracy, because these determine review burden.
- A tool that states no confidence should be treated as requiring 100% manual review, regardless of its accuracy.
- Requiring citations and a stated confidence is a concrete, measurable control that cuts review burden roughly in half in most configurations while staying within a 1% residual error tolerance.
- Self-verification (a second read-through pass) does not improve accuracy or reproducibility and costs 2.3x latency; it only helps groundedness, so it is worth buying only when the exposure is an untraceable audit trail.
Reading between the lines
- If review burden is the binding constraint, the economics of AI deployment can be modeled as a coverage decision under a risk constraint, and a firm could compute a break-even review ratio below which a tool pays even at full inspection.
- The measured gap between demonstration and production suggests that many public and internal evaluations of language models, which typically report a single accuracy number, systematically overstate deployability; a standardized review-burden metric could be added to model cards.
- The finding that confidence signals are uninformative on summarization (AUC close to 0.5) points to a testable extension: tasks requiring composed prose may need confidence signals tied to claim-level sources rather than whole-output scores.
- A replication on real messy documents (scanned, reflowed tables) would likely widen the gap; the paper states its synthetic corpus is cleaner, so the measured gap is a lower bound.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes that enterprise AI programmes stall not because of model accuracy but because of the unmeasured cost of human review. It defines a two-bar instrument: a demonstration bar (one run, one case, all correct) and a production bar (accuracy, reproducibility, groundedness, and informative confidence). Six document-heavy financial workflows are run across four model families and three tool configurations, with three repeats each, yielding 5,093 scored output elements. The paper reports that 57 of 72 configurations pass the demonstration bar and 32 pass the production bar, and then estimates review burden, concluding that a plain tool requires 100% review, a governed tool (citations + confidence) reduces this to 49% while holding a 1% residual error tolerance in 17 of 20 configurations, and a self-verification tool reaches 44% while costing 2.3× latency. The central claims are the survival gap (43.9% of demo-passing configurations fail production) and the claim that review burden is measurable, reducible, and more consequential than raw accuracy.
Significance. If the measurement is sound, this is a valuable contribution. The paper attacks a widely quoted but rarely measured phenomenon, and it does so with several strengths: deterministic scoring for most workflows, out-of-sample calibration for the review-burden estimate, fully published code and corpus generation, real EDGAR filings for two workflows, and unusually candid limitations. The distinction between a demonstration bar and a production bar is operationally meaningful, and the idea of measuring the 'share of output a human must still check' is a genuinely useful framing for regulated deployments. However, the headline numbers rest on assumptions that are either unstated or internally inconsistent, and these issues must be resolved before the conclusions can be accepted.
major comments (4)
- [§2.1 vs §3.4] The demonstration-bar case selection is unspecified. Section 2.1 defines the bar as 'one run, one case, all elements correct,' but the paper never states how that one case is chosen for each configuration. Section 2.2 explicitly calls the demo input 'friendly,' yet no protocol (e.g., easiest case, random case, first case) is given. The headline survival rate of 56.1% (Section 4.1) and the claim that 43.9% of demo-passing configurations cannot be placed in production are directly contingent on this choice. Without a precise, reproducible rule for selecting the demonstration case, the central measurement cannot be interpreted or replicated.
- [§5.1 and Table 4] The operational review burden for C3 (0.435) is lower than the oracle lower bound (0.470), which is impossible under the paper's own definitions: the oracle is the minimum burden achievable under the 1% tolerance with hindsight. The likely cause is that the 'Operational' column averages all 20 configurations, including those where the tolerance failed (16 of 20 held for C3; 17 of 20 for C2). This conflates average burden with the burden actually achievable while satisfying the tolerance. The abstract and §5.2 claim 'reduces that to 49% while holding the residual error tolerance' — but the reported 49% is not restricted to the configurations where the tolerance held. The authors must report the average burden separately for configurations that held tolerance and reconcile the oracle/operational ordering.
- [Abstract and §5.2] The paper states that C3 is 'the only configuration that fails to hold the error tolerance' / 'the only configuration that breaks the error bar.' This is directly contradicted by Table 4, which shows tolerance held in 20/20 for C1, 17/20 for C2, and 16/20 for C3 — meaning C2 also fails in 3 configurations. This is not a subtle wording issue; it is a factual error in the abstract and in a key interpretive sentence. The conclusion that 'buying the most engineering did not buy the best outcome' may still be defensible, but the 'only configuration' claim must be corrected.
- [§5.1] The review-burden model assumes a human reviewer triages solely by the model's stated confidence and skips all output above a calibrated threshold. The paper calls the resulting 100% review for a no-confidence tool 'not a modelling artefact,' but it is exactly a modelling artefact: it follows from the assumption that confidence is the only triage signal. Real reviewers use domain judgment, spot checks, and other signals, and the paper provides no evidence that the assumed policy matches actual reviewer behavior. The headline reduction from 100% to 49% is therefore a scenario analysis under a strong, unvalidated assumption. The authors should either validate this policy or present it as an upper/lower bound with sensitivity analysis.
minor comments (4)
- [Table 4] The column header 'Oracle Operational Latency' is ambiguous; it should be split into separate columns with clear units. Also, the 'Holds' row counts are not visually tied to the averages, which contributed to the confusion in the major comment.
- [§4.3] The text reports W5 and W6 accuracy as 0.964 and 0.881, while Table 2 shows per-configuration values (e.g., 1.000, 0.941, 0.951). These text numbers appear to be means across configurations, but the paper should state this explicitly and ideally add a pooled column to Table 2.
- [§3.4] The paper says 'Three repeats per cell at temperature zero' but does not define 'cell.' It should clarify whether a cell is a configuration, a workflow-model pair, or a model-workflow-configuration combination.
- [Appendix A] The workflow briefs do not state the number of cases per workflow, which is needed to understand the demonstration bar's single-case selection. Adding case counts would improve reproducibility.
Circularity Check
No significant circularity: the central quantities are direct measurements or out-of-sample estimates, and no load-bearing self-citation or definitional reduction is present.
full rationale
The paper's central quantitative claims do not reduce to their inputs. The survival gap (57/72 demonstration vs 32/72 production; 43.9% of demonstration-passers failing production) is a direct count of measured outcomes against two explicitly stated bars. The headline review-burden reduction (C2 operational 0.488 vs C1 1.000) is not a fitted prediction: Section 5.1 states the operational estimate 'calibrates the threshold on half the cases and applies it to the held-out half,' i.e., out-of-sample, and Section 5.2 reports that tolerance held in only 17 of 20 C2 configurations, so the average is not smoothed over failures. The statement that a tool with no confidence signal requires 100% review is a definitional consequence of the paper's explicitly stated triage model ('A reviewer triages by stated confidence... Where a tool states no confidence, there is nothing to triage on'), not a hidden assumption slipped in as evidence for the 49% result; and the paper even labels it 'not a modelling artefact,' which is an overstatement, but the 49% figure stands on the held-out computation independently. There is no load-bearing self-citation: the reference list is entirely external to the author, and the paper explicitly credits the risk-coverage framework to Geifman & El-Yaniv rather than claiming it as novel ('This paper sits between three literatures... The contribution here is not the framework but its application'). No uniqueness theorem or imported ansatz from prior work by the same author is used. Section 8 candidly lists limitations (practitioner-elicited bars, cleaner corpus, small sample, residual ground-truth risk), which further indicates the derivation is not being protected by silence. Under the stated rules, this is a no-significant-circularity finding.
Assumptions & free parameters
free parameters (3)
- Production bar thresholds (P1, P3, P4) =
0.99 / 0.95 / 0.70
- Residual error tolerance for review burden =
1% of elements
- Operational confidence threshold (per configuration) =
not reported; calibrated on training half of cases
assumptions (4)
- domain assumption The four production-bar criteria and their numeric thresholds are the correct acceptance requirements for regulated deployment.
- domain assumption A reviewer can triage by stated confidence and safely skip output below a calibrated threshold up to a 1% residual error tolerance.
- domain assumption A single successful run on a single case is a meaningful operationalization of the demonstration bar.
- domain assumption Exact repetition at temperature zero is the right operationalization of reproducibility (P2).
Cite this review
Pith. "Pith review of The Checking Problem: What must be true before AI ships in a regulated firm." pith.science (2026). https://pith.science/paper/F4IW3VVB
@misc{pith2026260728666,
author = {Pith},
title = {Pith review of: The Checking Problem: What must be true before AI ships in a regulated firm},
year = {2026},
howpublished = {\url{https://pith.science/paper/F4IW3VVB}},
note = {Machine review of arXiv:2607.28666}
}
read the original abstract
Enterprise AI programmes stall at a rate that is widely quoted and poorly explained. This paper measures the mechanism. Six document-heavy workflows of the kind performed daily in regulated financial services were run across four model families and three tool configurations, three times each, producing 5,093 scored output elements across 72 configurations. Each configuration was assessed twice: against a demonstration bar, being a single correct run on a single case, and against a production bar requiring sustained accuracy, reproducibility across repeats, verifiable attribution, and a confidence signal that carries information. 57 of 72 configurations cleared the demonstration bar and 32 cleared the production bar, a survival rate of 56.1%. The paper then computes the review burden each configuration imposes, estimated out of sample rather than with hindsight. A tool that states no confidence requires review of 100% of its output, because it offers a reviewer no basis for triage. Requiring the tool to cite its sources and state a confidence reduces that to 49% while holding the residual error tolerance in 17 of 20 configurations. Adding a self-verification pass costs 2.3 times the latency of the plain configuration, reaches 44%, and is the only configuration that fails to hold the error tolerance. The practical implication is that the value of an AI workflow is set less by how often it is right than by how much of it a human must still check, and that the second property is measurable and rarely measured.
Reference graph
Works this paper leans on
-
[1]
The Impact of AI on Developer Productivity: Evidence from GitHub Copilot.arXiv (Cornell University)
Sida Peng, Eirini Kalliamvakou, Peter Cihon, Mert Demirer (2023). The Impact of AI on Developer Productivity: Evidence from GitHub Copilot.arXiv (Cornell University). arXiv:2302.06590
arXiv 2023
-
[2]
Shakked Noy, Whitney Zhang (2023). Experimental Evidence on the Productivity Effects of Generative Artificial Intelligence.SSRN Electronic Journal. DOI: 10.2139/ssrn.4375283
-
[3]
Generative AI at Work.The Quarterly Journal of Economics
Erik Brynjolfsson, Danielle Li, Lindsey Raymond (2025). Generative AI at Work.The Quarterly Journal of Economics. arXiv:2304.11771
arXiv 2025
-
[4]
Mollick, Hila Lifshitz-Assaf, Katherine Kellogg, Saran Rajendran et al
Fabrizio Dell’Acqua, Edward McFowland, Ethan R. Mollick, Hila Lifshitz-Assaf, Katherine Kellogg, Saran Rajendran et al. (2023). Navigating the Jagged Technological Frontier: Field Experimental Evidence of the Effects of AI on Knowledge Worker Productivity and Quality. SSRN Electronic Journal. DOI: 10.2139/ssrn.4573321
-
[5]
Chuan Guo, Geoff Pleiss, Yu Sun, Kilian Q. Weinberger (2017). On Calibration of Modern Neural Networks.arXiv (Cornell University). arXiv:1706.04599
arXiv 2017
-
[6]
Saurav Kadavath, Tom Conerly, Amanda Askell, Tom Henighan, Dawn Drain, Ethan Perez et al. (2022). Language Models (Mostly) Know What They Know.arXiv. arXiv:2207.05221
arXiv 2022
-
[7]
Selective Classification for Deep Neural Networks
Yonatan Geifman, Ran El-Yaniv (2017). Selective Classification for Deep Neural Networks. arXiv. arXiv:1705.08500
arXiv 2017
-
[8]
Hannah Rashkin, Vitaly Nikolaev, Matthew Lamm, Lora Aroyo, Michael Collins, Dipan- jan Das et al. (2023). Measuring Attribution in Natural Language Generation Models. Computational Linguistics. DOI: 10.1162/coli a 00490
doi:10.1162/coli 2023
Show all 12 references
-
[9]
Ziwei Ji, Nayeon Lee, Rita Frieske, Tiezheng Yu, Dan Su, Yan Xu et al. (2022). Survey of Hallucination in Natural Language Generation.arXiv. arXiv:2202.03629
2022 arXiv
-
[10]
Shahul Es, Jithin James, Luis Espinosa Anke, Steven Schockaert (2024). RAGAs: Auto- mated Evaluation of Retrieval Augmented Generation.Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics: System Demon- strations. arXiv:2309.15217
2024 arXiv
-
[11]
Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang et al. (2023). Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena.Ad- vances in Neural Information Processing Systems 36. arXiv:2306.05685
2023 arXiv
-
[12]
Scalable Runtime Governance for Agentic AI in Financial Services
Lukasz Szpruch, Agus Sudjianto, Tanveer Bhatti, Gary Ang (2026). Scalable Runtime Governance for Agentic AI in Financial Services. DOI: 10.2139/ssrn.6567199. 10
2026 doi
Reviewed August 3, 2026 · model on record in the stance chip above.
Discussion (0). Sign in to comment.