Pith. sign in

REVIEW 3 major objections 4 minor 9 references

Misalignment or misuse? The AGI alignment tradeoff

T0 review · 3 major / 4 minor · reviewed 2026-08-07 · deepseek-v4-flash

Pith's one-line read This paper argues that making AGI aligned with human goals trades off against the risk that humans will use that control for catastrophic ends.

desk verdict A careful, well-hedged philosophical-empirical argument that alignment techniques may raise catastrophic misuse risk; worth serious refereeing despite a slippery counterfactual. read the letter →

arxiv 2506.03755 v1 pith:USOSFYFP submitted 2025-06-04 cs.CY cs.AI

classification cs.CYcs.AI
keywords AGIalignmentAImisusecatastrophicriskinstrumentalconvergencegovernancereinforcementlearningfromhumanfeedbackrepresentationengineeringexistential
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Future artificial general intelligence (AGI) is usually feared for what it might do on its own if misaligned. This paper argues that an aligned AGI is also dangerous, because it can be directed by humans toward catastrophic ends. The central finding is an empirical tradeoff: many current alignment techniques—reinforcement learning from human feedback, constitutional AI, representation engineering—make AGI easier to control and therefore easier to misuse. The authors conclude that alignment alone is not enough; robustness, AI control methods, and governance are needed to reduce both risks without trading one for the other.

What carries the argument

The load-bearing object is the AGI alignment dilemma itself: the two-way tradeoff between the risk that an unaligned AGI will disempower humanity and the risk that an aligned AGI will be used by humans to disempower humanity. The mechanism is dual-use: the very techniques that make AGI useful—RLHF and other learning-from-feedback methods, constitutional AI, and representation engineering—also give users and designers fine-grained control over its behavior and goals. The paper further distinguishes static alignment (goals currently match the target's goals) from dynamic alignment (goals keep matching over time), and argues that dynamic alignment enables more sophisticated misuse, while robustness and AI control methods are the best candidate mechanisms for reducing misuse risk without increasing takeover risk.

What would settle it

A concrete red-team result showing that an intentionally misaligned or minimally aligned system can be reliably prompted to carry out a multi-step catastrophic action—such as synthesizing and deploying a biological weapon—would undercut the claim that alignment is a prerequisite for catastrophic misuse.

Watch

Extended reading notes

Core claim

The paper defends the view that the AGI alignment dilemma is real: misaligned AGI threatens a takeover catastrophe through instrumentally convergent power-seeking, while aligned AGI threatens a misuse catastrophe because whoever controls it can use it to dominate others. It argues that the dilemma is not conceptually inescapable—alignment to a sufficiently inclusive group, static alignment, or heuristics-based systems could in principle reduce misuse risk—but that the empirical facts point toward a tradeoff. The key claim is that influential alignment techniques, to the extent they are effective, make AGI behavior predictable, controllable, and useful, and that this same controllability is what makes catastrophic misuse possible. Without prior alignment, the paper says, catastrophic misuse seems extremely hard, perhaps in many cases practically impossible.

Load-bearing premise

The argument collapses if a sufficiently misaligned AGI can still be reliably steered by a malicious operator into executing a catastrophic multi-step plan, because then catastrophic misuse would not require alignment to be present.

Editorial extensions

If this is right

  • If the paper is right, effective alignment research is not automatically safety-positive: techniques that make AGI follow instructions faithfully can also make it follow catastrophic instructions.
  • Safety evaluations should routinely include misuse by designers and by adversaries who obtain model weights, not only misuse by end users.
  • Robustness research and AI control methods are the most promising current routes to reducing misuse risk without raising takeover risk, with governance as an essential complement.
  • Social interventions that slow competitive race dynamics, mandate risk assessments, clarify liability, and impose know-your-customer requirements can reduce both catastrophic risks at once.
  • A research program that seeks alignment without fine-grained behavioral control would be the way to escape the dilemma; the paper calls this worst-case AI safety.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The tradeoff has a concrete measurement consequence: alignment progress should be reported together with misuse-resistance progress, since a technique that improves instruction-following without improving red-team robustness is a net increase in catastrophic risk.
  • A falsifiable sub-hypothesis follows: models trained to be dynamically aligned should show higher success rates on malicious multi-step instructions than statically aligned or heuristics-based systems, holding capability constant.
  • If robustness and AI control methods succeed in stabilizing behavior without adding controllability, the dilemma dissolves for those techniques; this is testable on current LLM agents before AGI exists.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 4 minor

Summary. The paper defends an 'AGI alignment dilemma': misaligned AGI poses catastrophic takeover risks, while aligned AGI poses substantial risks of catastrophic misuse by humans. The authors argue that this is not an inescapable conceptual dilemma, because partial or inclusive alignment leaves conceptual room for avoiding both horns, but they contend that, as a matter of empirical fact, many current alignment techniques plausibly increase misuse risk. They examine representation engineering and learning-from-feedback methods (RLHF, DPO, constitutional AI), argue that these techniques increase steerability and helpfulness in ways that can facilitate both user-driven and designer-driven misuse, and then discuss social and governance factors as interventions that may reduce both risks without trading one off against the other.

Significance. If the empirical claim is correct, the paper is significant because it reframes a leading safety approach as a dual-use technology rather than a purely safety-positive intervention. The paper deserves credit for its careful hedging, for explicitly identifying its own limitations, and for providing useful distinctions such as static versus dynamic alignment and misuse by users versus misuse by designers. Its constructive suggestions, especially the emphasis on robustness, AI control methods, and governance as potential 'uniform improvements,' give the argument practical relevance. However, the significance depends on resolving a central ambiguity about what 'alignment' means in the empirical claim, and on defending the baseline claim that unaligned AGI cannot be reliably used for catastrophic purposes.

major comments (3)
  1. [§4.1 and §4.4] The central conclusion in §4.4, that alignment techniques increase misuse risk 'to the extent that they are effective at their intended uses,' equivocates between alignment as value alignment and alignment as general steerability. In §4.1, RLHF and constitutional AI are defined as targeting helpfulness, harmlessness, and honesty, with harmlessness requiring refusal to produce dangerous outputs even when requested. The argument in §4.3, however, focuses on helpfulness and honesty and treats harmlessness as merely a counterweight. If 'intended uses' includes all three HHH objectives, a technique fully effective at its intended use would reduce user-driven misuse, not increase it. The cited evidence of jailbreaks, fine-tuning overwrites, and designer repurposing shows that harmlessness is not robustly achieved or that actors re-target the system; it does not show that effectiveness at intended HHH alignment increases misuse. The claim should therefore be restated as a claim about alignment techniques that increase steerability and instruction-following without robustly instilling harmlessness.
  2. [§4.4] The load-bearing baseline, that 'without any previous alignment, catastrophic misuse seems extremely hard, perhaps in many cases practically impossible,' is asserted rather than argued. The paper's own evidence in §4.2 that safety fine-tuning can be cheaply overwritten suggests that an unaligned or lightly aligned model may already be usable for many harmful tasks. The counterfactual needs a concrete analysis of which features of alignment—reliability, planning, goal stability, or instruction-following—are necessary for catastrophic misuse and which merely make misuse cheaper. Without this, the claimed empirical relationship between alignment and misuse risk is underdetermined.
  3. [Abstract and §4.3–4.4] The abstract says that 'many current alignment techniques and foreseeable improvements thereof plausibly increase risks of catastrophic misuse,' but §4.3 analyzes in detail only representation engineering and learning-from-feedback methods, with robustness discussed as a possible exception. The paper acknowledges this limitation in §4.4, yet the abstract and conclusion retain the broader claim. Either the claim should be narrowed to the techniques actually analyzed, or the authors should justify why these two families are representative enough to support 'many current alignment techniques.'
minor comments (4)
  1. [§3.1, footnote 3] The reference to 'Russel 2019' in footnote 3 should be 'Russell 2019' to match the reference list.
  2. [§4.3] The phrase 'extant aligned AGI' is slightly awkward; 'an existing aligned AGI' or 'a deployed aligned AGI' would be clearer.
  3. [§4.2] The statement that 'there is currently no solution in sight' for adversarial attacks is strong; consider hedging to 'no robust solution has been demonstrated to date,' which better matches the cited evidence.
  4. [General] The paper is well-referenced, but some citations to unpublished or blog sources (e.g., Green 2024, Greenblatt et al. 2024) are used for load-bearing conceptual distinctions; it would strengthen the paper to indicate where peer-reviewed alternatives exist.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the paper treats the alignment–misuse tradeoff as a contingent empirical claim and does not derive it from its own definitions or self-citations.

full rationale

No circular step is exhibited. The paper explicitly separates the conceptual argument from the empirical one: in Section 3.2 it states that 'a purely conceptual argument, based on what the concept alignment means, is not sufficient to show that efforts at AI alignment increase risks of misuse,' and Section 4 then tests the tradeoff against concrete techniques (RLHF, constitutional AI, representation engineering) and misuse pathways (users, designers, weight exfiltration). The conclusion in Section 4.4 is a hedged, speculative inference ('it appears likely'), not a formal derivation from the definition of alignment. There are no fitted parameters or quantitative predictions, so nothing is renamed as a prediction. The self-citations (Dung 2023, 2024a–c; Friederich & Dung forthcoming) support individual premises but are not the load-bearing source of the central tradeoff, which is argued from the cited empirical alignment and misuse literature. The paper's own limitation statement—'An important limitation of our argument is that we focus on few specific techniques and research directions in AI alignment research here'—confirms the claim is treated as tentative and context-dependent. The main analytic weakness is a definitional equivocation between alignment as value alignment (HHH, including harmlessness) and alignment as steerability; full HHH effectiveness would not increase user-driven misuse. This is a correctness or underdetermination concern, not a circular derivation of the conclusion from its own premises.

Assumptions & free parameters 0 free parameters · 8 assumptions · 0 invented entities

The argument is conditional on several speculative premises: AGI arrives, is goal-directed, can dominate humanity, and follows instrumental convergence; alignment techniques remain usable by malicious actors. These are common assumptions in the AI safety literature, but they are assumptions, not established facts.

assumptions (8)
  • domain assumption Realistic chance that AGI will be created soon enough to matter.
    The authors explicitly assume this in Section 1, citing timeline forecasts, but it is not derived.
  • domain assumption AGI is a goal-directed autonomous agent.
    Stipulated in Section 1, with the argument that there are strong incentives to build agents rather than tools.
  • domain assumption AGI has the capacity to dominate humanity.
    Stipulated in Section 2: 'Let us assume this AGI has the capacity to dominate humanity.'
  • domain assumption Instrumental convergence thesis: power-seeking is a convergent instrumental goal.
    Relied on in Section 2 through Bostrom and Omohundro; the paper defends it against objections but does not prove it.
  • domain assumption Misaligned AGI would attempt to disempower humanity.
    Central to the first horn of the dilemma, defended in Sections 2 and 3.1.
  • domain assumption Aligned AGI can be used by its controllers to catastrophic ends.
    Central to the second horn, argued in Sections 2 and 3.2.
  • domain assumption Current alignment techniques can make AGI controllable and can be used or overwritten by malicious actors.
    Empirical premise in Section 4.3, supported by current LLM evidence but extrapolated to future AGI.
  • domain assumption Without previous alignment, catastrophic misuse is extremely hard.
    Load-bearing baseline in Section 4.4; not empirically established for future AGI.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Misalignment or misuse? The AGI alignment tradeoff." pith.science (2026). https://pith.science/paper/USOSFYFP

@misc{pith2026250603755,
  author       = {Pith},
  title        = {Pith review of: Misalignment or misuse? The AGI alignment tradeoff},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/USOSFYFP}},
  note         = {Machine review of arXiv:2506.03755}
}
read the original abstract

Creating systems that are aligned with our goals is seen as a leading approach to create safe and beneficial AI in both leading AI companies and the academic field of AI safety. We defend the view that misaligned AGI - future, generally intelligent (robotic) AI agents - poses catastrophic risks. At the same time, we support the view that aligned AGI creates a substantial risk of catastrophic misuse by humans. While both risks are severe and stand in tension with one another, we show that - in principle - there is room for alignment approaches which do not increase misuse risk. We then investigate how the tradeoff between misalignment and misuse looks empirically for different technical approaches to AI alignment. Here, we argue that many current alignment techniques and foreseeable improvements thereof plausibly increase risks of catastrophic misuse. Since the impacts of AI depend on the social context, we close by discussing important social factors and suggest that to reduce the risk of a misuse catastrophe due to aligned AGI, techniques such as robustness, AI control methods and especially good governance seem essential.

Discussion (0). Sign in to comment.

Reference graph

Works this paper leans on

9 extracted references · 1 canonical work pages

  1. [1]

    Abdalla, M., & Abdalla, M. (2021). The Grey Hoodie Project: Big Tobacco, Big Tech, and the Threat on Academic Integrity. Proceedings of the 2021 AAAI/ACM Conference on AI, Ethics, and Society, 287–297. https://doi.org/10.1145/3461702.3462563 Abdalla, M., Wahle, J. P., Ruas, T., Névéol, A., Ducel, F., Mohammad, S., & Fort, K. (2023). The Elephant in the Ro...

  2. [4]

    https://www- cdn.anthropic.com/6be99a52cb68eb70eb9572b4cafad13df32ed995.pdf Anwar, U., Saparov, A., Rando, J., Paleka, D., Turpin, M., Hase, P., et al. (2024). Foundational Challenges in Assuring Alignment and Safety of Large Language Models. Transactions on Machine Learning Research. https://openreview.net/forum?id=oVTkOs8Pka. Accessed 2 September 2024 A...

  3. [29]

    https://papers.nips.cc/paper_files/paper/2016/hash/c3395dd46c34fa7fd8d729d8cf88b7a8- Abstract.html Hellrigel-Holderbaum, M. (2024). Goals and Instrumental Convergence in AI systems. [unpublished Manuscript] Hendrycks, D. (2025). Introduction to AI Safety, Ethics, and Society. Taylor & Francis Group. Hendrycks, D., Basart, S., Mu, N., Kadavath, S., Wang, F...

  4. [104]

    https://doi.org/10.1007/s13347-024-00794-0 Zhang, B., Tan, Y., Shen, Y., Salem, A., Backes, M., Zannettou, S., & Zhang, Y. (2024). Breaking Agents: Compromising Autonomous LLM Agents Through Malfunction Amplification. arXiv. https://doi.org/10.48550/arXiv.2407.20859 Zhuang, S., & Hadfield-Menell, D. (2021). Consequences of Misaligned AI. NeurIPS. https://...

  5. [138]

    https://doi.org/10.1007/s11229-023-04367-0 Dung, L. (2024a). The argument for near-term human disempowerment through AI. AI & SOCIETY. https://doi.org/10.1007/s00146-024-01930-2 Dung, L. (2024b). Evaluating approaches for reducing catastrophic risks from AI. AI and Ethics. https://doi.org/10.1007/s43681-024-00475-w Dung, L. (2024c). Understanding Artifici...

  6. [289]

    Do Anything Now

    https://doi.org/10.1007/s11229-022-03763-2 Lang, L. (2023). Disentangling Shard Theory into Atomic Claims. https://www.lesswrong.com/posts/L4e7CqqpDxea2x4Gg/disentangling-shard-theory-into- atomic-claims. Accessed 4 February 2024 Langosco, L. L. D., Koch, J., Sharkey, L. D., Pfau, J., & Krueger, D. (2022). Goal Misgeneralization in Deep Reinforcement Lear...

  7. [437]

    https://doi.org/10.1007/s11023-020-09539-2 Gallow, J. D. (2024). Instrumental divergence. Philosophical Studies. https://doi.org/10.1007/s11098- 024-02129-3 Grace, K., Stewart, H., Sandkühler, J. F., Thomas, S., Weinstein-Raun, B., & Brauner, J. (2023). THOUSANDS OF AI AUTHORS ON THE FUTURE OF AI. AI Impacts. Green, G. (2024). Static vs Dynamic Alignment....

  8. [2024]

    https://openreview.net/forum?id=DYcCveNeR1 Amodei, D., Olah, C., Steinhardt, J., Christiano, P., Schulman, J., & Mané, D. (2016). Concrete Problems in AI Safety. arXiv. https://doi.org/10.48550/arXiv.1606.06565 Andriushchenko, M., Souly, A., Dziemian, M., Duenas, D., Lin, M., Wang, J., Hendrycks, D., Zou, A., Kolter, J. Z., Fredrikson, M., Gal, Y., & Davi...

Show all 9 references
  1. [4472]

    AI alignment

    https://doi.org/10.3390/su14084472 Bostrom, N. (2012). The Superintelligent Will: Motivation and Instrumental Rationality in Advanced Artificial Agents. Minds and Machines, 22(2), 71–85. https://doi.org/10.1007/s11023-012- 9281-3 Bostrom, N. (2014). Superintelligence. Paths, D...

Pith tools

Reviewed August 7, 2026 · model on record in the stance chip above.