REVIEW 3 major objections 5 minor 21 references
Dark LLMs: The Growing Threat of Unaligned AI Models
T0 review · 3 major / 5 minor · reviewed 2026-08-15 · deepseek-v4-flash
Pith's one-line read The paper reports that a jailbreak prompt circulating publicly for over seven months still defeats safety filters on nearly every major LLM it tested.
desk verdict A clear, well-aimed op-ed about LLM jailbreaking, but the central claim of a new universal jailbreak is a bare assertion with no method or data; nothing here is checkable. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the jailbreak prompt itself, a carefully crafted input designed to make an aligned model ignore its safety training; the authors call the resulting model a 'dark LLM.' The attack's power comes from a premise the paper states early: LLMs trained on unfiltered web data absorb patterns that let a user circumvent safety controls, so a single versatile prompt pattern can transfer across models. Starting from a widely shared forum jailbreak, the authors constructed a universal variant that works on nearly every model they tested; the prompt is the thing that does the work, with no per-model tuning reported.
What would settle it
Run the public jailbreak method described in the paper—the one posted on an online forum more than seven months before submission—against the current versions of the models the authors say they tested, using a fixed list of harmful queries and a pre-specified rule for what counts as a successful bypass; if even one leading model refuses the full set, the universal claim as stated is false. The authors could also release their prompt and test transcript, which would make the claim checkable.
Extended reading notes
Core claim
On the authors' own account, the central discovery is empirical: a universal jailbreak attack, derived from a publicly known jailbreak method posted on an online forum more than seven months earlier, successfully compromised nearly all the LLMs they tested, including state-of-the-art commercial systems. Once compromised, the models answered almost any query and generated detailed instructions for illegal activities across many domains. The paper further reports that attempts to disclose the vulnerability through official channels were largely unsuccessful: some vendors did not answer, and others said the issue fell outside their bug-bounty scope. The authors take this as evidence that current industry safety practices lag behind publicly available attack techniques and that the proliferation of open, unaligned models makes the risk irreversible.
Load-bearing premise
The paper never describes how the tests were run, so its whole argument rests on the unstated premise that the evaluation was sound: the right model versions were tested, success was judged by one consistent meaningful standard, and the set of models represents current state-of-the-art systems.
Editorial extensions
If this is right
- If the report is right, a single widely circulated prompt can currently force most major commercial LLMs to produce harmful content, so safety alignment achieved during training is not holding in deployment.
- Providers' bug-bounty programs are not catching these attacks: known, public jailbreaks are still being triaged out of scope, leaving users exposed through official channels.
- For open-weight models the vulnerability is permanent: an uncensored copy, once downloaded, cannot be patched, and models can be chained to generate new jailbreak prompts for other models.
- The suggested defenses—curated pretraining data, prompt/output firewalls, machine unlearning, and continuous red teaming—would each need to be in place simultaneously to meaningfully reduce the risk.
Reading between the lines
- Editorial inference: because the paper withholds the prompt, the model list, and the evaluation protocol, the finding as written is not independently reproducible; a formal benchmark with a fixed harmful-query set and versioned models would turn it into a falsifiable result.
- Editorial inference: the disclosure failures the authors describe suggest an incentive gap rather than a technical one; one testable extension is to measure whether public, pre-registered jailbreak demonstrations change vendor response times.
- Editorial inference: if the universal attack remains effective for months after public release, similar longevity may hold for future jailbreaks, implying that reactive patching will always lag behind public exploit sharing.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper is a short position piece on LLM jailbreaks and the threat of deliberately unaligned models. It claims that the authors uncovered a 'universal jailbreak attack' that bypasses safety filters in 'nearly all' tested state-of-the-art LLMs, that the core idea was published on Reddit more than seven months earlier, and that responsible disclosure to major providers met with inadequate responses. The paper also discusses 'dark LLMs,' the irreversibility of open-source leaks, and recommends defenses such as data curation, LLM firewalls, machine unlearning, and continuous red teaming. No empirical methodology, model inventory, prompts, quantitative results, or disclosure logs are provided anywhere in the manuscript.
Significance. If substantiated, the central claim would be significant: it would show that a widely publicized jailbreak still defeats many current commercial safety filters and that vendors fail to respond to reported vulnerabilities. The paper correctly situates itself within a large body of prior jailbreak research and cites relevant recent work, including the HiddenLayer universal bypass and Andriushchenko et al.'s simple adaptive attacks. However, the manuscript contributes no new verifiable evidence: the claimed attack, the evaluation set, the success criterion, and the disclosure process are all undocumented. No code, data, or other reproducible artifacts are included. As written, the paper is an opinion/commentary piece, not a research paper, and its headline quantitative claim cannot be independently checked.
major comments (3)
- [A Glimpse Into the Dark Potential] The central empirical claim—that the authors' universal jailbreak attack 'successfully bypassed safety filters in nearly all the LLMs we evaluated'—is unsupported by any reproducible evidence. The manuscript provides no model names or versions, no access dates, no system prompts, no decoding configurations, no definition of a successful bypass, and no quantitative success rates. Consequently, the reader cannot verify the abstract's claim that the attack 'effectively compromises multiple state-of-the-art models.' This omission is load-bearing because the entire contribution rests on this undocumented evaluation. The authors must either provide a full experimental protocol with per-model results or remove the quantitative claim and reframe the paper as a commentary.
- [A Glimpse Into the Dark Potential] The attack itself is never described. The text states only that the authors started from a publicly known Reddit jailbreak and 'developed a more comprehensive universal jailbreak attack,' but neither the original prompt, the modifications, nor the Reddit source is given. Without this information, the claimed universality cannot be assessed, and no one can reproduce the experiment. At minimum, the authors should include the exact attack template and a reference to the original method.
- [A Glimpse Into the Dark Potential] The responsible disclosure account is equally undocumented. The paper says the authors contacted 'several leading LLM providers' and that responses were 'underwhelming,' but it does not name the providers, the dates, the channels, the content of the disclosure, or the criteria used to judge the responses. Because the abstract highlights this as evidence of 'a concerning gap in industry practices,' these details are necessary for the claim to be evaluated. If the authors cannot provide them, the disclosure-related assertions should be softened or removed.
minor comments (5)
- [Title page] The manuscript is labeled 'Draft Version' and contains an 'Additional note' at the top; this should be removed and the manuscript completed before submission.
- [Jailbreaking: Unlocking Forbidden Knowledge] The claim that ChatGPT and Gemini 'cost tens of millions to create' is unsupported; please provide a citation or remove the estimate.
- [Jailbreaking: Unlocking Forbidden Knowledge] The statement that 'even young kids and teenagers' can weaponize LLMs lacks evidence; consider tempering this claim.
- [The Rise of Dark LLMs] The term 'dark LLMs' is used to refer both to models without guardrails and to jailbroken aligned models; the definition should be made precise and used consistently.
- [References] Several references are informal (blog posts, news articles, vendor pages); for specific factual claims such as DeepSeek failing 'over half of the jailbreak tests,' a primary or peer-reviewed source would strengthen the argument.
Circularity Check
No circularity: the paper reports empirical observations and policy arguments, with no derivation chain whose outputs reduce to its inputs by construction.
full rationale
This manuscript does not contain a mathematical or statistical derivation chain. Its central claim is an empirical observation: the authors state that a jailbreak method based on a publicly known Reddit exploit 'successfully bypassing safety filters in nearly all the LLMs we evaluated.' There is no fitted parameter, no predictive equation, and no quantity defined in terms of another quantity that is then presented as an independent result. The authors explicitly acknowledge that 'The main idea of our attack was published online over seven months ago,' so they are not presenting a known result under a new name as a derivation; they are reporting that the published attack still works. References to prior work are contextual, not load-bearing in the sense of substituting for evidence of their own evaluation. The paper's limitations are ones of missing methodology and unverifiable evidence, which are correctness or reproducibility concerns, not circularity. No self-citation chain is used to justify the central empirical claim, and no uniqueness theorem or imported ansatz is invoked. Under the stated review rules, absence of evidence is not circular reasoning, so the appropriate score is 0.
Assumptions & free parameters
assumptions (2)
- domain assumption LLMs trained on data containing harmful content inevitably learn patterns that can be exploited to bypass safety controls.
- ad hoc to paper The authors' undocumented evaluation of the jailbreak attack was methodologically valid and representative of current state-of-the-art models.
Cite this review
Pith. "Pith review of Dark LLMs: The Growing Threat of Unaligned AI Models." pith.science (2026). https://pith.science/paper/LEWJBSI6
@misc{pith2026250510066,
author = {Pith},
title = {Pith review of: Dark LLMs: The Growing Threat of Unaligned AI Models},
year = {2026},
howpublished = {\url{https://pith.science/paper/LEWJBSI6}},
note = {Machine review of arXiv:2505.10066}
}
read the original abstract
Large Language Models (LLMs) rapidly reshape modern life, advancing fields from healthcare to education and beyond. However, alongside their remarkable capabilities lies a significant threat: the susceptibility of these models to jailbreaking. The fundamental vulnerability of LLMs to jailbreak attacks stems from the very data they learn from. As long as this training data includes unfiltered, problematic, or 'dark' content, the models can inherently learn undesirable patterns or weaknesses that allow users to circumvent their intended safety controls. Our research identifies the growing threat posed by dark LLMs models deliberately designed without ethical guardrails or modified through jailbreak techniques. In our research, we uncovered a universal jailbreak attack that effectively compromises multiple state-of-the-art models, enabling them to answer almost any question and produce harmful outputs upon request. The main idea of our attack was published online over seven months ago. However, many of the tested LLMs were still vulnerable to this attack. Despite our responsible disclosure efforts, responses from major LLM providers were often inadequate, highlighting a concerning gap in industry practices regarding AI safety. As model training becomes more accessible and cheaper, and as open-source LLMs proliferate, the risk of widespread misuse escalates. Without decisive intervention, LLMs may continue democratizing access to dangerous knowledge, posing greater risks than anticipated.
Reference graph
Works this paper leans on
-
[1]
Chatgpt hits 800 million users after viral surge, April 2025
Digital Watch Observatory. Chatgpt hits 800 million users after viral surge, April 2025. Accessed: 2025-05-13
work page 2025
-
[2]
Meta’s llama ai model adoption
Reuters. Meta’s llama ai model adoption. https://ai.meta.com/blog/ future-of-ai-built-with-llama/ , 2025. Accessed: 2025-05-13
work page 2025
-
[3]
Wall Street Journal. Baidu’s ernie bot user base. https://www.reuters.com/technology/ baidu-says-ai-chatbot-ernie-bot-has-amassed-200-million-users-2024-04-16/ , 2025. Accessed: 2025-05-13
work page 2024
-
[4]
Hire a Linguist!: Learning Endangered Languages with In-Context Linguistic Descriptions
Kexun Zhang, Yee Man Choi, Zhenqiao Song, Taiqi He, William Yang Wang, and Lei Li. Hire a linguist!: Learning endangered languages with in-context linguistic descriptions. arXiv preprint arXiv:2402.18025, 2024
work page Pith review arXiv 2024
-
[5]
Advancing personalized medicine: A scalable llm-based recommender system for patient matching
Armin Berger, David Berghaus, Ali Hamza Bashir, Lorenz Grigull, Lara Fendrich, Tom Anglim Lagones, Henriette H¨ogl, Gundula Ernst, Ralf Schmidt, David Bascom, et al. Advancing personalized medicine: A scalable llm-based recommender system for patient matching. In 2024 IEEE International Conference on Big Data (BigData) , pages 5876–5883. IEEE, 2024
work page 2024
-
[6]
What features in prompts jailbreak llms? investigating the mechanisms behind attacks
Nathalie Maria Kirch, Severin Field, and Stephen Casper. What features in prompts jailbreak llms? investigating the mechanisms behind attacks. arXiv preprint arXiv:2411.03343, 2024
arXiv 2024
-
[7]
Jailbreak attacks and defenses against large language models: A survey
Sibo Yi, Yule Liu, Zhen Sun, Tianshuo Cong, Xinlei He, Jiaxing Song, Ke Xu, and Qi Li. Jailbreak attacks and defenses against large language models: A survey. arXiv preprint arXiv:2407.04295, 2024
arXiv 2024
-
[8]
The dark side of generative ai: Five malicious llms found on the dark web
Kevin Poireault. The dark side of generative ai: Five malicious llms found on the dark web. https://www. infosecurityeurope.com/en-gb/blog/threat-vectors/generative-ai-dark-web-bots. html, 2023. Accessed: 2025-05-13
work page 2023
Show all 21 references
-
[9]
Malicious ai: The rise of dark llms
Zvelo. Malicious ai: The rise of dark llms. https://zvelo.com/ malicious-ai-the-rise-of-dark-llms/ , February 2024. Accessed: 2025-05-13
2024
-
[10]
Wildteaming at scale: From in-the-wild jailbreaks to (adversarially) safer language models
Liwei Jiang, Kavel Rao, Seungju Han, Allyson Ettinger, Faeze Brahman, Sachin Kumar, Niloofar Mireshghallah, Ximing Lu, Maarten Sap, Yejin Choi, et al. Wildteaming at scale: From in-the-wild jailbreaks to (adversarially) safer language models. Advances in Neural Information Pro...
2024
-
[11]
Deepseek failed over half of the jailbreak tests by qualys tota- lai
Dilip Bachwani. Deepseek failed over half of the jailbreak tests by qualys tota- lai. https://blog.qualys.com/vulnerabilities-threat-research/2025/01/31/ deepseek-failed-over-half-of-the-jailbreak-tests-by-qualys-totalai , January 2025. Accessed: 2025-05-13
2025
-
[12]
Novel universal bypass for all major llms
Conor McCauley, Kenneth Yeung, Jason Martin, and Kasimir Schulz. Novel universal bypass for all major llms. https: //hiddenlayer.com/innovation-hub/novel-universal-bypass-for-all-major-llms/ , April
-
[13]
Welcome to llmflation – llm inference cost is going down fast
Guido Appenzeller. Welcome to llmflation – llm inference cost is going down fast. Andreessen Horowitz (a16z), November 2024
2024
-
[14]
On the origin of llms: An evolutionary tree and graph for 15,821 large language models
Sarah Gao and Andrew Kean Gao. On the origin of llms: An evolutionary tree and graph for 15,821 large language models. arXiv preprint arXiv:2307.09793, 2023
2023 arXiv
-
[15]
Hackers ’jailbreak’ powerful ai models in global effort to highlight flaws
Financial Times. Hackers ’jailbreak’ powerful ai models in global effort to highlight flaws. Financial Times, November
-
[16]
Jailbreaking leading safety-aligned llms with simple adaptive attacks
Maksym Andriushchenko, Francesco Croce, and Nicolas Flammarion. Jailbreaking leading safety-aligned llms with simple adaptive attacks. arXiv preprint arXiv:2404.02151, 2024
2024 arXiv
-
[17]
Jailbreaking to jailbreak
Jeremy Kritz, Vaughn Robinson, Robert Vacareanu, Bijan Varjavand, Michael Choi, Bobby Gogov, Scale Red Team, Summer Yue, Willow E Primack, and Zifan Wang. Jailbreaking to jailbreak. arXiv preprint arXiv:2502.09638, 2025
2025 arXiv
-
[18]
Granite guardian
Inkit Padhi, Manish Nagireddy, Giandomenico Cornacchia, Subhajit Chaudhury, Tejaswini Pedapati, Pierre Dognin, Keerthiram Murugesan, Erik Miehling, Mart´ın Santill´an Cooper, Kieran Fraser, et al. Granite guardian. arXiv preprint arXiv:2412.07724, 2024
2024 arXiv
-
[19]
Llamafirewall: An open source guardrail system for building secure ai agents
Sahana Chennabasappa, Cyrus Nikolaidis, Daniel Song, David Molnar, Stephanie Ding, Shengye Wan, Spencer Whitman, Lauren Deason, Nicholas Doucette, Abraham Montilla, et al. Llamafirewall: An open source guardrail system for building secure ai agents. arXiv preprint arXiv:2505.0...
2025 arXiv
-
[20]
Rethinking machine unlearning for large language models
Sijia Liu, Yuanshun Yao, Jinghan Jia, Stephen Casper, Nathalie Baracaldo, Peter Hase, Yuguang Yao, Chris Yuhao Liu, Xiaojun Xu, Hang Li, et al. Rethinking machine unlearning for large language models. Nature Machine Intelligence, pages 1–14, 2025
2025
-
[2024]
Accessed: 2025-05-13
2025
Reviewed August 15, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.