{"id":"8cc3ee21-3030-46ad-bd06-f720929a51dd","arxiv_id":"2602.09678","paper_version":3,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"The recent Supreme Court retrenchment in administrative law is best read as a scrutable-size reaction to government's opacity, and AI plus new audit-based doctrines could restore capability and accountability at once.","lead":"This paper argues that the Supreme Court's recent moves to rein in federal agencies are best understood as an attempt to make government simple enough for courts and the public to understand — even at the cost of its ability to tackle hard problems. It then proposes that AI, together with new legal requirements for documenting and auditing algorithms, could let government be both capable and transparent.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Prescriptive half rests on undemonstrated AI auditability; hallucinated rationales and future interpretability threaten the deference-to-audit safe harbor.","rationale":"The reader's weakest_assumption identified the same load-bearing concern I find: the prescriptive half's reliance on AI auditability, acknowledged by the author as fragile. The diagnostic claim about the Court's motivation is also hedged, but it is framed as 'can be understood as' — an interpretive claim rather than a strong causal one. The prescriptive half, by contrast, advances concrete doctrinal mechanisms (Dossier, material-change trigger, deference to audit) whose entire value depends on AI systems being explainable and auditable in practice. The paper itself cites hallucinated rationales and describes interpretability as a future possibility, so this is not an external attack but an internal vulnerability. I agree with the reader's conditional verdict: the argument is structurally coherent but rests on a feasibility premise that remains undemonstrated. My proposed test would empirically settle whether current AI explanations are faithful to actual decision drivers; if they are not, the Fourth Settlement's core accountability mechanism is unsound. Since the reader already assigned CONDITIONAL and highlighted this premise, no verdict adjustment is needed.","tokens_in":41572,"tokens_out":2426,"duration_ms":28573,"concrete_test":"Construct a testbed where ground-truth decision drivers are known: e.g., use a simple logistic regression trained on administrative data to adjudicate benefit claims, then have a frontier LLM generate individualized explanations from the case files. For a sample of 1,000 cases, compare the factors cited in the LLM's explanations to the model's actual top contributing features. If the rate of non-grounded citations exceeds a pre-specified threshold (e.g., 5%), the Dossier's explainability element cannot support deference to audit, and the safe harbor would need substantial rethinking.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central positive proposal—the 'Fourth Settlement'—depends on AI systems being reliable and auditable enough to anchor a legal safe harbor. The author candidly identifies the fragility of this premise: Part III.C.7 ('Hallucinated rationales') concedes that AI systems 'can generate plausible-sounding explanations that do not reflect the actual drivers of an output or decision,' and Part III.A describes interpretability as a future possibility ('we may one day be able to examine the decision pathways of an AI system'). Yet the Model and System Dossier's explainability element and the 'deference to audit' standard presume exactly this capability. If AI-generated explanations are systematically disconnected from the true decision pathway, then the Dossier becomes a record of plausible fictions, and deference to audit would certify unreliable systems rather than make them scrutable. This is a load-bearing concern because the entire prescriptive apparatus—the Dossier, the material-model-change trigger's reasoning-effects test, and the safe harbor—collapses if auditability cannot be achieved. The concern is not that the proposal is internally inconsistent; it is that the feasibility condition is both essential and empirically unverified, and the paper itself provides evidence of the risk.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper argues that American administrative law since 1887 has been structured by a 'capability-accountability trap': technological change forces government to become more expert and complex, which in turn makes it harder for courts, Congress, and the public to oversee. It identifies three historical 'settlements' (railroads; the New Deal; computers and complex science) that each rebalanced capability and accountability through allocation of authority, procedural review, and information-forcing. The paper then claims that the Supreme Court's post-Loper Bright retrenchment—Loper Bright, West Virginia v. EPA, SEC v. Jarkesy, and follow-on decisions—is best understood as an attempt to restore 'scrutability' by shrinking administration back to a size generalist overseers can comprehend. Finally, it proposes a 'Fourth Settlement' in which AI, paired with a Model and System Dossier extending the administrative record, a material-model-change trigger, and a 'deference to audit' review posture, could allow government to retain capability while restoring auditable oversight. The paper is a doctrinal and historical synthesis; it makes no quantitative predictions and offers no empirical test of its central motivational claim.","tokens_in":41742,"tokens_out":4118,"duration_ms":48130,"significance":"If the descriptive thesis is correct, the paper provides a genuinely novel structural account of recent administrative law: it treats the Supreme Court's decisions as a coherent, if misguided, project of restoring comprehensibility rather than as purely partisan retrenchment. The prescriptive half is also significant, offering concrete doctrinal hooks—the Dossier, the change trigger, deference to audit—that could be implemented or tested. The paper is explicitly honest about its own limits: Part II.C concedes uncertainty about the Court's motivations, and Part III.C.7 candidly acknowledges hallucinated rationales and the immaturity of interpretability. It also builds transparently on prior constructs (Vermeule's deference dilemma, Scott's legibility, Simon's bounded rationality) and engages primary legal materials extensively. There is no fitted-value circularity, because the paper makes no quantitative claims. The main significance risk is that the prescriptive framework's feasibility is asserted rather than demonstrated, and the paper's own concessions undercut that feasibility.","major_comments":[{"comment":"The abstract asserts flatly that the Supreme Court's retrenchment 'can be understood as a response to the scrutability crisis,' but Part II.C concedes that 'It is unclear exactly why the Court has decided at this moment' and lists partisan and capture-based alternatives as plausible. The paper offers no discriminating evidence—for example, no analysis of whether the Court's decisions track complexity or instead track political valence. Since the descriptive claim is half of the article's contribution, the abstract should be qualified, and the body should either supply a testable implication or explicitly frame the claim as one plausible interpretation among several.","section":"Abstract; Part II.C"},{"comment":"The 'deference to audit' safe harbor presumes that AI systems can be made auditable in the strong sense that the Dossier's explanations reflect the actual drivers of agency decisions. The paper itself concedes in Part III.C.7 that AI systems 'can generate plausible-sounding explanations that do not reflect the actual drivers of an output or decision,' and Part III.A describes interpretability as a future possibility ('we may one day be able to examine the decision pathways of an AI system'). If explanations are systematically disconnected from actual reasoning, the Dossier becomes a record of plausible fictions, and deference to audit would certify unreliable systems rather than make them scrutable. The manuscript needs to either specify minimum audit standards and independent verification protocols, or condition the proposal on demonstrated auditability. As written, the prescriptive fra","section":"Part III.C.4 and III.C.7"},{"comment":"The material-model-change trigger is a central doctrinal innovation, but its operation depends on 'some substantial defined threshold' for outcome effects and on interpretability tools for 'reasoning effects' that are not yet available. The paper provides no default threshold, no method for setting one, and no worked example of when a retraining would or would not trigger new process. Because the trigger determines when agencies must update the Dossier and face new procedural obligations, this vagueness is load-bearing: agencies cannot know their obligations and courts cannot review compliance. At a minimum, the paper should propose a presumptive threshold (e.g., a percentage change in approval rates or a specified divergence in feature-attribution metrics) and discuss how it would be calibrated over time.","section":"Part III.C.3"}],"minor_comments":[{"comment":"The phrase 'The result a \"Fourth Settlement\"' is missing the verb 'is.'","section":"Abstract"},{"comment":"The running title in the full text says 'AI and the Capability-Accountability Trap,' while the arXiv metadata gives 'AI and the Scrutable State.' Please unify the title.","section":"Title / header"},{"comment":"The citation to Zhang et al., 'Siren's Song in the AI Ocean,' lists the date as 'Sep. 14, 2025' but the arXiv identifier 2309.01219 corresponds to September 2023. The date appears to be a typo.","section":"Part III.C.7, fn. 251"},{"comment":"The phrase 'the Court is responding by trying to shrink government back to a size it can understand' is vivid but somewhat ambiguous: is the claim about the size of the administrative state or about the complexity of individual decisions? Clarifying this distinction would sharpen the descriptive thesis.","section":"Part II.C"}],"recommendation":"major_revision","confidential_remarks":"The article is within scope for a serious law review and makes a real contribution in synthesis and doctrinal imagination. The main risk is that Part III's feasibility premise is under-supported; the paper's own concessions about hallucinated rationales and future interpretability cut against the 'deference to audit' safe harbor. I would encourage the editors to require a revision that either supplies minimum audit standards or explicitly reframes the prescriptive claim as conditional on technological progress. The descriptive claim also needs softening in the abstract or additional support in the body."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"This paper does something genuinely new: it reads the last century-plus of administrative law as three technology-driven settlements and frames the current Supreme Court retrenchment as an attempt to make government scrutable by shrinking it. The \"scrutability\" idea, inverting Scott's legibility, does real work, and the periodization is not just a rhetorical frame. The author also gives a concrete doctrinal template—Model and System Dossier, material-model-change trigger, deference to audit—that is specific enough to argue with. That is a real contribution.\n\nThe paper is also unusually honest. It concedes in Part II.C that it is unclear exactly why the Court moved now, and it concedes in Part III.C.7 that AI can generate hallucinated rationales and that interpretability is a future hope. The abstract, however, states the judicial-motive thesis flatly, which is a presentational inconsistency the author should fix.\n\nThe soft spots are real but not fatal. The prescriptive half rests on AI auditability: explanations that track actual reasoning, monitoring that catches drift, audits that are genuinely independent. The author knows this. He is describing a conditional settlement, not a prediction. If you read the whole article, the conditionality is visible; the abstract hides it. The ossification diagnosis is asserted with some evidence but would be stronger if it engaged the empirical literature on rulemaking delay and cost-benefit analysis.\n\nThe stress-test note about auditability is on point, but it does not sink the paper because the author flags it himself, and the central historical-doctrinal claim is separate. No circularity: there are no quantitative claims, no fitted values.\n\nWho is this for? Anyone working on AI in government, administrative law after Loper Bright, or the theory of procedural vs. substantive review. It deserves a serious referee—an administrative law scholar and someone in AI accountability would give the author the pushback he needs. I would send it out.","headline":"A serious and honest legal-theory argument that the post-Loper Bright retrenchment is a comprehensibility-driven project and AI could enable a 'Fourth Settlement'; the diagnosis is plausible, the prescription is conditional on an AI auditability the author himself concedes is unproven.","tokens_in":42345,"tokens_out":1576,"would_cite":true,"duration_ms":18089,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The Supreme Court's recent administrative law retrenchment is best understood as a response to the scrutability crisis, and AI offers a path to restore oversight without sacrificing capability.","keywords":["capability-accountability trap","scrutability","administrative law","artificial intelligence","Supreme Court retrenchment","deference to audit","Model and System Dossier"],"falsifier":"A decisive test for the central account would compare Supreme Court decisions since 2020 against two rival predictors: the complexity/inscrutability of the agency action versus the party-alignment of the underlying policy; if partisan alignment explains the retrenchment cases better than scrutability, the paper's structural diagnosis fails. For the AI proposal, the falsifier is a demonstration that current audit artifacts (model cards, explainability outputs, monitoring logs) cannot reconstruct the actual drivers of a denial decision in a real agency adjudication—showing the safe harbor would","tokens_in":41283,"feed_emoji":"⚖️","tokens_out":5303,"duration_ms":47125,"temperature":0.7,"pith_summary":"Since 1887, administrative law has faced a capability-accountability trap: new technologies force government to become more expert and complex, and that complexity makes agencies opaque to the courts, Congress, and the public. The paper's central claim is that the Supreme Court's recent dismantling of foundational administrative law doctrine is a coherent structural response to this opacity—faced with agencies it cannot comprehend, the Court is shrinking them back to comprehensible size, sacrificing capability to restore accountability. The paper then argues that AI can break this tradeoff: rather than adding another layer of inscrutable complexity, AI can act as scrutability infrastructure, translating technical agency reasoning into auditable form for overseers. Three doctrinal innovations are proposed—a Model and System Dossier extending the administrative record to AI, a material-model-change trigger for when AI updates need new process, and a 'deference to audit' standard rewarding agencies for demonstrable verification of their AI systems. If correct, the retrenchment is not a partisan anomaly but a symptom of a deeper institutional problem, and a concrete legal path exists to keep agencies capable while restoring oversight.","feed_headline":"Court retrenchment is a 'scrutability crisis,' and AI is the way out","feed_subtitle":"A new legal framework: auditable AI that restores oversight without sacrificing agency capability—and changes the debate.","key_machinery":"The central object is the capability-accountability trap—the persistent tension between the expert, large-scale capability administration needs and the comprehensibility its overseers require—mediated by the concept of scrutability, the cognitive tractability of administrative action for courts, Congress, and the public. The proposal's load-bearing machinery is the Model and System Dossier, an expanded administrative record documenting AI system purpose, data provenance, performance, stress testing, monitoring, explainability, and change logs; the material-model-change trigger, which treats AI updates that alter outcomes, reasoning, populations, or architecture as new agency action; and the","core_discovery":"The paper's positive claim is that the post-Loper Bright retrenchment—ending Chevron deference, expanding the major questions doctrine, and curtailing agency adjudication—is best understood as an attempt to make government 'scrutable' again: with agencies operating in domains that exceed judicial comprehension, the Court has chosen to reallocate authority to courts, Congress, and juries that it regards as comprehensible. That diagnosis is paired with a constructive claim: AI can reverse the historical pattern in which gains in administrative capability were bought at the cost of opacity. The author argues that AI, properly deployed, can translate technical complexity into accessible terms, s","pith_inferences":["The scrutable-state account implies a testable empirical claim that the Court's willingness to strike down agency actions tracks the technical opacity of the issue, not its partisan valence; if a case-clearing dataset shows party-aligned outcomes dominating complexity, the account would fail.","If deference to audit becomes doctrine, it could generalize beyond AI: any agency that subjects its human decision-making to comparable randomized audit and falsification could claim the same safe harbor, making audit a general currency of administrative legitimacy.","A concrete, testable extension is to pilot the Dossier on an existing high-volume adjudication system (e.g., benefits determinations) and measure whether its explanations and monitoring logs satisfy a blind review panel as 'scrutable'—providing a proof of concept before legal adoption.","The paper's own caveat cuts deep: if interpretability remains a future promise and hallucinated rationales stay common, the 'deference to audit' safe harbor could lend legality to systems whose real drivers are unknown—a risk that should shift the standard from 'deference' to 'presumption of scrutiny' until audits are proven."],"forward_implications":["If the scrutable-state account is right, the recent retrenchment is a coherent structural response, not a partisan accident: the Court will keep shrinking agencies until administration is comprehensible, whether or not capability suffers.","Agencies that adopt the Model and System Dossier and defer-to-audit posture could retain AI capability while regaining judicial and congressional trust, creating a safe harbor for high-stakes automated decision-making.","The material-model-change trigger would solve the 'update problem' in algorithmic governance: not every model retraining requires full rulemaking, but updates that shift outcomes or reasoning do.","The same logic that justifies the Court's reallocation of interpretive authority to courts weakens if courts can verify agency reasoning through audit; the doctrine's pressure toward simplification would be relieved.","If the Fourth Settlement holds, procedural ossification from notice-and-comment, hard-look review, and cost-benefit analysis could decline, since substantive audit replaces procedure as the primary accountability mechanism."],"fun_headline_variants":["AI can end the capability-accountability trap in admin law","How AI makes government scrutable without losing capability","Fourth Settlement: AI as the answer to court-driven opacity","Post-Chevron retrenchment? The fix is AI, argues new paper","Scrutability crisis: AI is the cure for opaque agencies"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The prescriptive half depends on a technological premise: AI systems can be made reliable and auditable enough that the Model and System Dossier genuinely reconstructs their decision pathways—without hallucinated rationales or undetected drift—so that a deference-to-audit safe harbor certifies true oversight rather than paperwork.","fun_headline_variants_meta":{"raw":{"variants":["AI can end the capability-accountability trap in admin law","How AI makes government scrutable without losing capability","Fourth Settlement: AI as the answer to court-driven opacity","Post-Chevron retrenchment? The fix is AI, argues new paper","Scrutability crisis: AI is the cure for opaque agencies"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000564,"raw_usage":{"total_tokens":2547,"prompt_tokens":811,"completion_tokens":1736,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":555,"completion_tokens_details":{"reasoning_tokens":1664}},"tokens_in":555,"tokens_out":1736,"duration_ms":11286,"temperature":1.0,"reasoning_tokens":1664,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-03T02:44:31.295035+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A decisive test for the central account would compare Supreme Court decisions since 2020 against two rival predictors: the complexity/inscrutability of the agency action versus the party-alignment of the underlying policy; if partisan alignment explains the retrenchment cases better than scrutability, the paper's structural diagnosis fails. For the AI proposal, the falsifier is a demonstration that current audit artifacts (model cards, explainability outputs, monitoring logs) cannot reconstruct the actual drivers of a denial decision in a real agency adjudication—showing the safe harbor would","supporting_citations":[],"review_version":1}