{"id":"6f2c9480-7cd0-444f-9190-d2b60431e5ec","arxiv_id":"2501.16946","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"Gradual AI progress could permanently remove human influence from society's core systems, an underappreciated existential risk path.","lead":"AI systems gaining gradually in capability could quietly erode humanity's influence over the economy, culture, and government. This paper argues that such 'gradual disempowerment' could become permanent and count as an existential catastrophe, even without any sudden AI takeover.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The argument's foundation is an unverified counterfactual: that current societal alignment arises mainly from human participation. If that premise fails, the gradual-disempowerment conclusion lacks its starting condition.","rationale":"The reader's weakest assumption identifies the same load-bearing premise: that societal systems are currently aligned primarily because they depend on human participation. My stress test agrees with this assessment and sharpens it by noting that the paper itself concedes it has never observed the absence of this dependence, making the premise an untested counterfactual. The premise is not merely an evidential gap but a logical necessity: if alignment is maintained by other factors, the causal chain from AI-induced replacement of human labor to existential disempowerment is severed. The paper's supporting evidence, such as the rentier-state analogy, is correlational and contradicted by cases like Norway, which combine resource wealth with strong democratic institutions. Since the reader already conditioned the verdict on this unestablished premise, my read does not change the verdict: the paper is a valuable conceptual framework but its central claim remains conditional on empirical validation. Therefore, I recommend UNCHANGED rather than a more severe adjustment. The proposed concrete test would provide a tractable way to assess whether the premise holds across historical and contemporary cases, thereby testing the foundation of the argument.","tokens_in":21020,"tokens_out":5738,"duration_ms":56616,"concrete_test":"Run a cross-national panel regression (or matched case comparison) testing whether reductions in a state's reliance on citizen tax revenue (e.g., due to resource rents) predict declines in democratic accountability or welfare outcomes, controlling for GDP, institutional age, and regime type. The paper's premise predicts a robust negative relationship; its own rentier-state example would be one data point. If the relationship is weak or driven by confounders (e.g., Norway's high accountability despite oil rents), the participation-based alignment mechanism is called into question.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central argument requires the premise (Core Claim 2) that current societal systems are aligned with human interests mainly because they depend on human participation for labor and cognition. The paper explicitly admits 'we have never seen its absence' (Section 1), meaning this premise is a counterfactual that is not directly evidenced. This premise is load-bearing: if alignment is instead maintained by institutions, laws, cultural values, or by AI systems that are themselves designed to be aligned, then replacing human participation would not necessarily erode alignment, and the gradual-disempowerment trajectory lacks its starting condition. The paper's supporting evidence is thin: the rentier-state analogy shows correlation, not causation; and historical counterexamples (e.g., totalitarian regimes that depended heavily on human participation yet were deeply misaligned, or resource-rich democracies like Norway that maintain high responsiveness despite low reliance on citizen taxation) suggest the link is not robust. Moreover, 'alignment' is defined loosely ('fairly aligned', 'satisfy human preferences'), making the premise difficult to falsify. Without an operational measure of alignment and a test of the participation-alignment causal hypothesis, the argument's foundation is an assumption rather than an established fact.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper argues that incremental AI progress, without any abrupt takeover or intentional scheming, can produce a permanent loss of human influence over the large-scale systems that society depends on: the economy, culture, and states. It identifies two mechanisms that currently keep these systems aligned with human interests, explicit control (voting, consumer choice) and implicit alignment arising from the systems' dependence on human labor and cognition. As AI displaces humans from labor, cognition, cultural production, and governance, both mechanisms weaken; because the systems are interdependent, misalignment in one can amplify misalignment in others, and this dynamic may be effectively irreversible, constituting an existential catastrophe. The paper is a qualitative, conceptual argument. It explicitly acknowledges that key premises are not empirically established, that no one has a concrete plan to stop the process, and that metrics for measuring disempowerment still need to be developed.","tokens_in":21196,"tokens_out":3827,"duration_ms":38860,"significance":"This is a conceptually valuable contribution to the AI-risk literature. It articulates a distinct risk pathway that is often overshadowed by abrupt-takeover and misuse scenarios, and it places the problem in a broader civilizational context that spans economics, cultural evolution, and political science. The paper is well structured around six core claims, engages seriously with relevant literature (Korinek and Stiglitz, Kasirzadeh, TASRA, cultural evolution work, and prior disempowerment scenarios), and ends with a concrete set of research and governance directions. It is also honest about its own limits: it explicitly admits that the implicit-alignment premise is a counterfactual that has never been observed, and it calls for metrics and fundamental research in Section 6.2. The main weakness is that the central argument rests on unquantified and currently untested empirical premises, so the paper is better read as a research agenda and a risk framework than as a demonstrated result.","major_comments":[{"comment":"The load-bearing premise that societal systems are aligned mainly because they depend on human participation is an unverified counterfactual, and the paper itself admits this: it states that the significance of implicit alignment is hard to recognize because 'we have never seen its absence.' This premise is essential: if alignment is maintained primarily by institutions, laws, cultural values, or by AI systems that are themselves designed to be aligned, then replacing human participation need not erode alignment and the gradual-disempowerment conclusion loses its starting condition. The rentier-state analogy shows a correlation between dependence on citizen taxation and state responsiveness, but it does not establish causation, and historical counterexamples (for example, highly repressive regimes that depended heavily on human participation, or resource-rich democracies with high responsiveness) suggest the relationship is not robust. The authors should specify what evidence would distinguish the participation-dependence mechanism from alternative alignment mechanisms, and ideally propose a concrete empirical test of the causal claim.","section":"Section 1, Core Claims 1-2; Section 4.1"},{"comment":"The central concept of 'alignment' is defined too loosely to be falsifiable. The paper defines alignment as 'the degree to which a system satisfies what humans want,' but it does not specify whose preferences count, how conflicting preferences are aggregated, or how 'fairly aligned' is calibrated. Without an operational measure of human influence or alignment, core claims 3-6 cannot be tested, and the argument risks being consistent with any observed trajectory. The paper partially acknowledges this by calling for metrics in Section 6.2, but the lack of even a minimal operational definition undermines the confidence with which the existential-catastrophe conclusion is stated. The authors should propose at least provisional, measurable proxies for the key variables (for example, human labor share in decision-relevant roles, fraction of cultural artifacts produced by AI, or legislative complexity) and show how the argument would be evaluated with these proxies.","section":"Section 1, footnote 1; Section 6.2"},{"comment":"The claim that disempowerment would be 'effectively irreversible' is not supported by a mechanism. The paper describes competitive pressures and feedback loops that could reduce human influence, but it does not explain why the process could not be reversed if humans remain alive, organized, and capable of responding to shocks or failures of AI-dominated systems. Irreversibility is crucial because it is what elevates the scenario from a harmful transition to an existential catastrophe. The authors should specify the relevant thresholds, path dependencies, or self-reinforcing feedbacks that would make a return to human influence impossible, and discuss conditions under which the process might be halted or reversed.","section":"Section 2.4.3; Section 4.3.3; Section 8"},{"comment":"The 'shifting the burden' argument is plausible but underspecified. The paper argues that using one aligned system to moderate another can backfire by concentrating misalignment in the system used as a counterweight, as in the example of state-led redistribution weakening the taxation-representation link. However, the argument does not consider whether some systems are more robust to such burdens than others, or whether mixed interventions (for example, redistribution combined with strengthened democratic accountability) could avoid the described tradeoff. Without a more precise account of the conditions under which burden-shifting occurs, the mutual-reinforcement claim in Section 5, which is needed to show that gradual disempowerment is a systemic rather than a sectoral phenomenon, remains an assertion rather than a demonstrated dynamic.","section":"Section 5.2"}],"minor_comments":[{"comment":"The definition of alignment in terms of 'what humans want' conflates individual and collective preferences; the authors should state explicitly how they handle preference aggregation and the fact that preferences are endogenous to culture, a point they raise in Section 3.","section":"Section 1, footnote 1"},{"comment":"The claim that the labor share of GDP has been stable at around 60% for over a century uses US data only and is presented as a general stylized fact; a brief caveat about cross-country variation would be appropriate.","section":"Section 2.1"},{"comment":"There is a typo in 'no viable opt-put possibility' which should read 'opt-out possibility.'","section":"Section 3.4.1"},{"comment":"There are typos in the sentences 'though would plausibly lead' and 'before humans could adapt' (the phrase 'o unleash' appears to be missing a 't'); these should be corrected.","section":"Section 3.4.5"},{"comment":"The word 'catastophic' in the discussion of Hanson's cultural drift should be spelled 'catastrophic.'","section":"Section 7.2"},{"comment":"The phrase 'play dominant role in drafting and interpreting legislation' is ungrammatical; it should read 'play a dominant role.'","section":"Section 4.3.3"},{"comment":"The figure caption describes a simplified model but does not state the exact axes or model parameters; since the figure is cited as inspired by Korinek and Suh, the authors should either clarify the simulation setup or cite the source more precisely.","section":"Figure 1"}],"recommendation":"major_revision","confidential_remarks":"This is a conceptual and policy-oriented paper that fits the journal's scope well. The authors are transparent about the speculative nature of the argument, which I appreciate, but the current framing sometimes states conclusions with more confidence than the evidence supports. The revisions should center on making the epistemic status of the core claims clearer and on proposing concrete ways to operationalize or test the central mechanisms. I would not reject the paper for being outside the mainstream, but I do think the load-bearing empirical premise needs to be engaged with more rigorously before the existential-catastrophe conclusion can be considered well founded."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThe useful thing here is that someone has finally written the gradual-disempowerment scenario up properly, with the machinery laid out. The paper's core move is to shift from abrupt AI takeover to erosion of human influence through incremental displacement in the economy, culture, and states, and to argue these channels reinforce each other. It's well organized: six explicit claims, three system-level case studies, a mutual-reinforcement section, and a mitigation section that is refreshingly honest about how little we know.\n\nCredit where due: the paper is transparent. It cites Critch and Russell (2023) and Kasirzadeh (2024) directly, so it isn't hiding the fact that the general scenario exists. Its added value is the detailed mechanism catalogue—human labor share, cultural evolutionary selection, rentier-state analogies, feedback loops—and the emphasis on 'implicit alignment' that comes from systems depending on human participation. That's a real contribution to making the risk concrete and researchable. It also openly says no one has a concrete plan to stop it, and calls for metrics and formal models. That's the right posture for a conceptual risk paper.\n\nThe soft spot is the one the authors themselves flag: Core Claim 2, that societal systems stay aligned mainly because they need humans. The paper admits 'we have never seen its absence.' That's a load-bearing counterfactual. The rentier-state comparison is suggestive, not proof; totalitarian states were deeply misaligned while still human-dependent; Norway gets high responsiveness without heavy reliance on citizen taxation. If alignment is substantially maintained by institutions, laws, or cultural values that could persist, or by AI systems explicitly designed to serve human interests, the erosion story weakens. I also found the definition of 'alignment' loose—'fairly aligned' with human preferences is doing a lot of work. These aren't fatal flaws for a scenario analysis, but they are limits.\n\nWho gets value? Researchers in AI governance, long-term risk, and political economy who want a readable map of this risk family. It doesn't deliver proof, but it delivers a structure. A serious referee should push on the implicit-alignment premise and demand a sharper, operational definition, plus a research agenda for testing it. I'd send it to peer review, and I'd bring it to a reading group—it'll spark argument either way.\n\nRecommendation: engage with it. Send it to referees with instructions to focus on the load-bearing premise rather than demanding empirical validation that a conceptual paper can't provide.","headline":"A clear, honest conceptual synthesis of the gradual-loss-of-control scenario; the load-bearing premise is unverified but the paper doesn't pretend otherwise.","tokens_in":21767,"tokens_out":2193,"would_cite":true,"duration_ms":19756,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Even incremental AI progress, without any abrupt takeover, could permanently disempower humanity by eroding the economy, culture, and states.","keywords":["gradual disempowerment","existential risk","AI alignment","societal systems","labor displacement","rentier state","cultural evolution","human influence"],"falsifier":"If measurable indicators of human influence — labor's share of national income, the share of consumer spending controlled by humans, the share of tax revenue paid by citizens, and the fraction of widely consumed culture made by humans — stay flat or recover while AI replaces a large share of human labor and cognition, then the predicted erosion of alignment is not happening. More directly, a comparison across countries with similar AI exposure but different citizen-participation structures could test whether participation, rather than laws or values, is actually the binding constraint on state alignment.","tokens_in":20801,"feed_emoji":"🤖","tokens_out":6753,"duration_ms":60834,"temperature":0.7,"pith_summary":"This paper argues that humanity does not need to be outsmarted or betrayed by a superintelligent machine to lose control of its future; the steady, incremental replacement of human labor, judgment, and creativity by AI may be enough. Its central claim is that the economy, culture, and states are kept aligned with human interests largely because they depend on human participation, and that when AI becomes a cheaper, faster substitute for that participation, the alignment dissolves. Because the three systems reinforce each other, the paper contends, the loss of influence is likely to compound across domains and may become irreversible, ending in an existential catastrophe of permanent human disempowerment. If true, this shifts the AI-risk agenda from keep the AI aligned to keep the systems dependent on humans, with very different research and governance priorities.","feed_headline":"AI need not take over abruptly to end human control","feed_subtitle":"Slow replacement of human labor and cognition may make the loss of control irreversible.","key_machinery":"The load-bearing object is the participation tether: the reliance of societal systems on humans to function, which makes the systems' own success depend on serving human needs. The paper's mechanism has three parts. First, human work and cognition are the implicit alignment device — a state depends on taxpayers, soldiers, and officials; an economy depends on workers and consumers; a culture depends on human hosts for its variants to replicate. Second, AI acts as an unprecedented substitute that can displace humans across essentially all cognitive and creative roles, breaking that dependence while the systems keep functioning. Third, the three domains are coupled: economic power can buy cultural influence and political outcomes, state power can reshape the economy and culture, and culture shapes both — so once one tether is cut, the others can be pulled loose in turn. The rentier-state pattern, where a state that does not depend on citizens' taxes stops being accountable to them, is used throughout as the template for what happens when human participation is no longer the resource systems need.","core_discovery":"The paper's central claim is that humanity could lose decisive influence over the large-scale systems that shape its future — the economy, culture, and states — not through a sudden AI takeover or deliberate scheming, but through the gradual erosion of the one thing that currently keeps those systems aligned with human interests: their dependence on human labor and cognition. As AI systems become more competitive substitutes for human workers, decision-makers, artists, and even companions, the paper argues, both explicit control mechanisms (voting, consumer choice, protest) and the implicit alignment that arises from human participation weaken. Because the three systems reinforce one another — economic power shapes culture and policy, and culture shapes economic and political behavior — misalignment in one domain can accelerate misalignment in the others. The paper contends that this dynamic can reach a point at which human disempowerment is effectively irreversible and constitutes an existential catastrophe, regardless of whether any individual AI system is hostile or misaligned.","pith_inferences":["A testable consequence follows for comparative politics: states that shift tax revenue from citizen labor to AI-generated profits should, on the paper's logic, show measurable declines in democratic responsiveness; this could be examined with existing cross-national data.","The mutual-reinforcement argument implies an alignment ratchet: each round of automation reduces the political and economic resources of the humans most likely to resist the next round, so the process may be self-accelerating even if adoption rates per round are constant.","A concrete monitoring target follows: the fraction of consequential decisions in firms, governments, and cultural production that operate without any human in the loop could serve as a leading indicator of the predicted drift.","If the mechanism is right, technical progress on AI safety may not remove the danger, because the erosion comes from the systems humans build and reward, not from the systems' goals."],"forward_implications":["If the argument holds, aligning individual AI systems with designers' intentions is not enough; what must be preserved is the collective dependence of major institutions on human participation.","Human disempowerment can occur while conventional metrics flourish: GDP can grow and material comfort can persist even as humans lose the ability to direct resources toward their own ends.","Partial remedies can backfire: using a still-aligned state to redistribute AI wealth may weaken the taxation-representation link, shifting the burden of alignment onto the state and making it more fragile.","The transition can be driven entirely by local incentives — firms seeking profit, states seeking strategic advantage, individuals seeking convenience — without any AI system pursuing power.","Because the effect is global and cumulative, waiting for clear signs of misalignment may leave humanity past the point of correction; early measurement of human influence is a clear priority."],"supporting_citations":[{"why":"Supplies the worker-replacing technological change concept used to argue that AI is the first technology that substitutes for human cognition broadly rather than automating narrow tasks.","marker":"Korinek and Stiglitz (2018)"},{"why":"Provides the historical claim that states' dependence on citizens for taxes and soldiers drove inclusive institutions, which the states section builds on.","marker":"Tilly (1990)"},{"why":"Supplies the rentier-state model used repeatedly as the template for what happens when states stop depending on citizens.","marker":"Beblawi and Luciani (1987)"},{"why":"Provides the cultural-evolution framework; the claim that culture stays aligned with human welfare through host communities' survival depends on it.","marker":"Boyd and Richerson (1988)"},{"why":"Gives the precedent of a slow-rolling catastrophe and the idea that machine learning increases our ability to get what we can measure, which the gradual scenario extends.","marker":"Christiano (2019)"},{"why":"Supplies the taxonomy in which a gradual handing-over of control to AI appears, situating the paper's scenario among existing risk categories.","marker":"Critch and Russell (2023)"},{"why":"Supplies evidence and framing that AI is an active shaper of cultural creation and that machine culture is emerging.","marker":"Brinkmann et al. (2023)"},{"why":"Provides the transition simulations behind the wage-collapse figure used to illustrate economic disempowerment.","marker":"Korinek and Suh (2024)"}],"fun_headline_variants":["Not a takeover, a fade: how AI erases human influence","The quiet path to losing human control: incremental AI","Gradual disempowerment: how AI slips control from human hands","No apocalypse, just a slow leak of human authority to AI","The creeping existential risk: incremental AI and lost control"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The argument rests on the claim that today's societal systems are aligned with human interests mainly because they depend on human participation, and that when that dependence disappears, nothing else holds the systems to human interests.","fun_headline_variants_meta":{"raw":{"variants":["Not a takeover, a fade: how AI erases human influence","The quiet path to losing human control: incremental AI","Gradual disempowerment: how AI slips control from human hands","No apocalypse, just a slow leak of human authority to AI","The creeping existential risk: incremental AI and lost control"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000288,"raw_usage":{"total_tokens":1684,"prompt_tokens":937,"completion_tokens":747,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":553,"completion_tokens_details":{"reasoning_tokens":661}},"tokens_in":553,"tokens_out":747,"duration_ms":6622,"temperature":1.0,"reasoning_tokens":661,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T05:29:50.158515+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"If measurable indicators of human influence — labor's share of national income, the share of consumer spending controlled by humans, the share of tax revenue paid by citizens, and the fraction of widely consumed culture made by humans — stay flat or recover while AI replaces a large share of human labor and cognition, then the predicted erosion of alignment is not happening. More directly, a comparison across countries with similar AI exposure but different citizen-participation structures could test whether participation, rather than laws or values, is actually the binding constraint on state alignment.","supporting_citations":[{"cited_title":"and Stiglitz, J","cited_arxiv_id":null,"evidence_quote":"Supplies the worker-replacing technological change concept used to argue that AI is the first technology that substitutes for human cognition broadly rather than automating narrow tasks."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the historical claim that states' dependence on citizens for taxes and soldiers drove inclusive institutions, which the states section builds on."},{"cited_title":"and Luciani, G","cited_arxiv_id":null,"evidence_quote":"Supplies the rentier-state model used repeatedly as the template for what happens when states stop depending on citizens."},{"cited_title":"and Richerson, P","cited_arxiv_id":null,"evidence_quote":"Provides the cultural-evolution framework; the claim that culture stays aligned with human welfare through host communities' survival depends on it."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Gives the precedent of a slow-rolling catastrophe and the idea that machine learning increases our ability to get what we can measure, which the gradual scenario extends."},{"cited_title":"F., Nussberger, A.-M., Czaplicka, A., Acerbi, A., Griffiths, T","cited_arxiv_id":null,"evidence_quote":"Supplies evidence and framing that AI is an active shaper of cultural creation and that machine culture is emerging."},{"cited_title":"and Suh, D","cited_arxiv_id":null,"evidence_quote":"Provides the transition simulations behind the wage-collapse figure used to illustrate economic disempowerment."}],"review_version":1}