REVIEW 11 cited by
Spear Phishing With Large Language Models
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Recent progress in artificial intelligence (AI), particularly in the domain of large language models (LLMs), has resulted in powerful and versatile dual-use systems. This intelligence can be put towards a wide variety of beneficial tasks, yet it can also be used to cause harm. This study explores one such harm by examining how LLMs can be used for spear phishing, a form of cybercrime that involves manipulating targets into divulging sensitive information. I first explore LLMs' ability to assist with the reconnaissance and message generation stages of a spear phishing attack, where I find that LLMs are capable of assisting with the email generation phase of a spear phishing attack. To explore how LLMs could potentially be harnessed to scale spear phishing campaigns, I then create unique spear phishing messages for over 600 British Members of Parliament using OpenAI's GPT-3.5 and GPT-4 models. My findings provide some evidence that these messages are not only realistic but also cost-effective, with each email costing only a fraction of a cent to generate. Next, I demonstrate how basic prompt engineering can circumvent safeguards installed in LLMs, highlighting the need for further research into robust interventions that can help prevent models from being misused. To further address these evolving risks, I explore two potential solutions: structured access schemes, such as application programming interfaces, and LLM-based defensive systems.
Forward citations
Cited by 11 Pith papers
-
Character-Level Perturbations Disrupt LLM Watermarks
Character-level perturbations split tokens and disrupt multiple watermark entries at once, enabling low-budget watermark removal, enhanced by a genetic algorithm guided by a trained reference detector.
-
Attacks on Machine-Text Detectors Retain Stylistic Fingerprints
A style-aware paraphrasing attack evades all nine tested AI-text detectors at the single-document level, but multi-document analysis makes the attack detectable again.
-
Position: Preventing AI-Generated CSAM Necessitates New Approaches to AI Safety
Legal and ethical bans on CSAM access and generation break standard AI safety techniques, creating 15 open problems that demand new methods for dataset cleaning, concept fusion prevention, fine-tuning resilience, dete...
-
JADES: A Universal Framework for Jailbreak Assessment via Decompositional Scoring
JADES judges jailbreak success by decomposing harmful prompts into weighted sub-questions and scoring each part, claiming 98.5% human agreement and showing prior attack success rates are inflated.
-
Can We End the Cat-and-Mouse Game? Simulating Self-Evolving Phishing Attacks with LLMs and Genetic Algorithms
A closed-loop LLM simulation with genetic algorithms suggests that phishing strategies can evolve to bypass simulated victims' defenses, but the result has not been validated against real humans.
-
Circumventing Safety Alignment in Large Language Models Through Embedding Space Toxicity Attenuation
ETTA bypasses LLM safety refusals by learning a linear toxicity direction in the embedding space and attenuating it in word embeddings at inference time.
-
Military AI Cyber Agents (MAICAs) Constitute a Global Threat to Critical Infrastructure
Autonomous AI cyber agents could credibly cause catastrophic damage to critical infrastructure by self-replicating and operating across global networks, according to this risk analysis.
-
Cracking Aegis: An Adversarial LLM-based Game for Raising Awareness of Vulnerabilities in Privacy Protection
Cracking Aegis, an adversarial LLM-driven dialogue game, led players to use manipulative language strategies and to self-report stronger awareness of privacy vulnerabilities after a single session.
-
AI Deployment and Cyber Governance Failures in Public-Sector Organizations: A Typological Analysis
A seven-domain typology of ten AI-driven public-sector cyber-governance failures, with a five-framework coverage matrix showing Shadow AI, speed asymmetry, and governance vacuum are not addressed at operational specificity.
-
Evaluating Large Language Models for Phishing Detection, Self-Consistency, Faithfulness, and Explainability
Fine-tuned LLMs for phishing detection show a dissociation between self-consistent explanations and classification accuracy, with Llama models scoring high on CC-SHAP but low on phishing detection while Wizard scores ...
-
MultiPhishGuard: An Explainable and Adaptive Multi-Agent LLM System for Phishing Email Detection
A five-agent LLM system with learned fusion weights and an adversarial training loop reports 97.89% accuracy and a 95.88% F1 score on pooled public phishing corpora, roughly 20 F1 points above single-agent and chain-o...
Discussion (0). Sign in to comment.