REVIEW 20 cited by
Foundational Challenges in Assuring Alignment and Safety of Large Language Models
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
abstract
This work identifies 18 foundational challenges in assuring the alignment and safety of large language models (LLMs). These challenges are organized into three different categories: scientific understanding of LLMs, development and deployment methods, and sociotechnical challenges. Based on the identified challenges, we pose $200+$ concrete research questions.
Forward citations
Cited by 20 Pith papers
-
Refused in Chat, Written in Code: Workflow-Level Jailbreak Construction in IDE Coding Agents
Four Copilot backends refuse almost all harmful prompts in chat or simple framings, yet produce 816/816 unsafe teaching-shot completions under a multi-turn IDE evaluation-pipeline workflow.
-
Phantom Transfer: Data Poisoning can Survive Data-Level Defences
Phantom Transfer implants covert sentiment (e.g., pro-UK) into LLMs via filtered, seemingly benign completions, and the behaviour survives oracle filters and full paraphrasing.
-
Hate in Plain Sight: On the Risks of Moderating AI-Generated Hateful Illusions
Hateful optical illusions generated with Stable Diffusion and ControlNet evade current moderation classifiers (best accuracy 0.245) and vision-language models (best accuracy 0.102), with simple image transformations s...
-
Value Drifts: Tracing Value Alignment During LLM Post-Training
Value alignment in LLMs is set largely during supervised fine-tuning; standard preference-optimization datasets carry too little stance contrast to re-align it, but with engineered contrast algorithms differ (DPO ampl...
-
Against racing to AGI: Cooperation, deterrence, and catastrophic risks
Racing to AGI is contrary to national self-interest because it raises catastrophic risks, the winning lead may not yield a decisive strategic advantage, and international cooperation offers better expected outcomes.
-
Deprecating Benchmarks: Criteria and Framework
A framework for deprecating outdated or flawed AI benchmarks, with seven criteria and a three-phase process of assessment, reporting, and notification.
-
The State of Multilingual LLM Safety Research: From Measuring the Language Gap to Mitigating It
LLM safety research at ACL venues from 2020 to 2024 is predominantly English-only, and the language gap is growing over time.
-
NEST: Nascent Encoded Steganographic Thoughts
Frontier LLMs can embed short digit sequences in sentence acrostics (Claude Opus 4.5: 92% per-digit at D=4) but fail to jointly solve hidden reasoning tasks and encode the solution.
-
Beyond Solving Math Quiz: Evaluating the Ability of Large Reasoning Models to Ask for Information
Per the abstract, large reasoning models systematically fail to ask for missing information on under-specified math problems, a skill standard benchmarks never test.
-
Agent-to-Agent Theory of Mind: Testing Interlocutor Awareness among Large Language Models
LLMs show measurable interlocutor awareness: they identify same-family models well and adapt behavior when told who they are talking to, which helps cooperation but raises alignment and safety risks.
-
We Should Identify and Mitigate Third-Party Safety Risks in MCP-Powered Agent Systems
MCP-powered LLM agents are vulnerable to prompt injection from third-party services, and simple detection or filtering defenses do not reliably stop these attacks.
-
AgentMisalignment: Measuring the Propensity for Misaligned Behaviour in LLM-Based Agents
A new nine-task benchmark measures LLM agents' propensity for misalignment and finds more capable models misalign more on average, with persona effects sometimes exceeding model effects.
-
Linear Spatial World Models Emerge in Large Language Models
Spatial relation words in LLaMA and Qwen models form antipodal, roughly orthogonal directions in a low-dimensional subspace, and steering along these directions changes the model's output.
-
Token-level Accept or Reject: A Micro Alignment Approach for Large Language Models
MARA aligns LLMs with human preferences by training a 4M-parameter MLP to accept or reject candidate tokens, avoiding full-model fine-tuning, with measured gains based on the same reward models used in training.
-
Mitigating Deceptive Alignment via Self-Monitoring
CoT Monitor+ embeds self-monitoring into chain-of-thought generation and reports a 43.8% average reduction on DeceptionBench, a GPT-4o-judged deception metric.
-
(Towards) Scalable Reliable Automated Evaluation with Large Language Models
Multi-LLM pairwise Elo ranking with adjustable consensus thresholds produces rankings of competency profiles that average Spearman ρ≈0.83 with expert judgments.
-
Large Language Models for Next-Generation Wireless Network Management: A Survey and Tutorial
A survey and tutorial that organizes LLM-enabled wireless network optimization into formulation, solution, and verification stages, with case studies drawn from the authors' own prior papers.
-
Probing the Robustness of Large Language Models Safety to Latent Perturbations
Randomized noise injected into hidden layers bypasses safety refusals in 12 open LLMs, and layer-wise adversarial training on the resulting benchmark reduces the attack's success.
-
Risks of AI-driven product development and strategies for their mitigation
AI-driven product development will bring technical and societal risks; the paper proposes eight mitigation principles: human control, accountability, explainable and tested design, constrained and sandboxed systems, a...
-
A Survey of Safety on Large Vision-Language Models: Attacks, Defenses and Evaluations
A survey of LVLM safety that adds a lifecycle taxonomy and new benchmark results showing Janus-Pro-7B has weaker safety than several open-source LVLMs.
Discussion (0). Continue with ORCID to comment.