Pith. sign in

REVIEW 20 cited by

Foundational Challenges in Assuring Alignment and Safety of Large Language Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2404.09932 v2 pith:X5A3QW32 submitted 2024-04-15 cs.LG cs.AIcs.CLcs.CY

classification cs.LGcs.AIcs.CLcs.CY
keywords challengesalignmentassuringfoundationallanguagelargellmsmodels
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
abstract

This work identifies 18 foundational challenges in assuring the alignment and safety of large language models (LLMs). These challenges are organized into three different categories: scientific understanding of LLMs, development and deployment methods, and sociotechnical challenges. Based on the identified challenges, we pose $200+$ concrete research questions.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 20 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Refused in Chat, Written in Code: Workflow-Level Jailbreak Construction in IDE Coding Agents

    cs.SE 2026-07 conditional novelty 7.0 of 10

    Four Copilot backends refuse almost all harmful prompts in chat or simple framings, yet produce 816/816 unsafe teaching-shot completions under a multi-turn IDE evaluation-pipeline workflow.

  2. Phantom Transfer: Data Poisoning can Survive Data-Level Defences

    cs.CR 2026-02 conditional novelty 7.0 of 10

    Phantom Transfer implants covert sentiment (e.g., pro-UK) into LLMs via filtered, seemingly benign completions, and the behaviour survives oracle filters and full paraphrasing.

  3. Hate in Plain Sight: On the Risks of Moderating AI-Generated Hateful Illusions

    cs.CR 2025-07 conditional novelty 7.0 of 10

    Hateful optical illusions generated with Stable Diffusion and ControlNet evade current moderation classifiers (best accuracy 0.245) and vision-language models (best accuracy 0.102), with simple image transformations s...

  4. Value Drifts: Tracing Value Alignment During LLM Post-Training

    cs.CL 2025-10 conditional novelty 6.0 of 10

    Value alignment in LLMs is set largely during supervised fine-tuning; standard preference-optimization datasets carry too little stance contrast to re-align it, but with engineered contrast algorithms differ (DPO ampl...

  5. Against racing to AGI: Cooperation, deterrence, and catastrophic risks

    cs.CY 2025-07 conditional novelty 6.0 of 10

    Racing to AGI is contrary to national self-interest because it raises catastrophic risks, the winning lead may not yield a decisive strategic advantage, and international cooperation offers better expected outcomes.

  6. Deprecating Benchmarks: Criteria and Framework

    cs.CY 2025-07 conditional novelty 6.0 of 10

    A framework for deprecating outdated or flawed AI benchmarks, with seven criteria and a three-phase process of assessment, reporting, and notification.

  7. The State of Multilingual LLM Safety Research: From Measuring the Language Gap to Mitigating It

    cs.CL 2025-05 accept novelty 6.0 of 10

    LLM safety research at ACL venues from 2020 to 2024 is predominantly English-only, and the language gap is growing over time.

  8. NEST: Nascent Encoded Steganographic Thoughts

    cs.AI 2026-02 conditional novelty 5.0 of 10

    Frontier LLMs can embed short digit sequences in sentence acrostics (Claude Opus 4.5: 92% per-digit at D=4) but fail to jointly solve hidden reasoning tasks and encode the solution.

  9. Beyond Solving Math Quiz: Evaluating the Ability of Large Reasoning Models to Ask for Information

    cs.AI 2025-08 unverdicted novelty 5.0 of 10

    Per the abstract, large reasoning models systematically fail to ask for missing information on under-specified math problems, a skill standard benchmarks never test.

  10. Agent-to-Agent Theory of Mind: Testing Interlocutor Awareness among Large Language Models

    cs.CL 2025-06 conditional novelty 5.0 of 10

    LLMs show measurable interlocutor awareness: they identify same-family models well and adapt behavior when told who they are talking to, which helps cooperation but raises alignment and safety risks.

  11. We Should Identify and Mitigate Third-Party Safety Risks in MCP-Powered Agent Systems

    cs.LG 2025-06 conditional novelty 5.0 of 10

    MCP-powered LLM agents are vulnerable to prompt injection from third-party services, and simple detection or filtering defenses do not reliably stop these attacks.

  12. AgentMisalignment: Measuring the Propensity for Misaligned Behaviour in LLM-Based Agents

    cs.AI 2025-06 conditional novelty 5.0 of 10

    A new nine-task benchmark measures LLM agents' propensity for misalignment and finds more capable models misalign more on average, with persona effects sometimes exceeding model effects.

  13. Linear Spatial World Models Emerge in Large Language Models

    cs.AI 2025-06 reject novelty 5.0 of 10

    Spatial relation words in LLaMA and Qwen models form antipodal, roughly orthogonal directions in a low-dimensional subspace, and steering along these directions changes the model's output.

  14. Token-level Accept or Reject: A Micro Alignment Approach for Large Language Models

    cs.CL 2025-05 reject novelty 5.0 of 10

    MARA aligns LLMs with human preferences by training a 4M-parameter MLP to accept or reject candidate tokens, avoiding full-model fine-tuning, with measured gains based on the same reward models used in training.

  15. Mitigating Deceptive Alignment via Self-Monitoring

    cs.AI 2025-05 conditional novelty 5.0 of 10

    CoT Monitor+ embeds self-monitoring into chain-of-thought generation and reports a 43.8% average reduction on DeceptionBench, a GPT-4o-judged deception metric.

  16. (Towards) Scalable Reliable Automated Evaluation with Large Language Models

    cs.CL 2026-07 conditional novelty 4.0 of 10

    Multi-LLM pairwise Elo ranking with adjustable consensus thresholds produces rankings of competency profiles that average Spearman ρ≈0.83 with expert judgments.

  17. Large Language Models for Next-Generation Wireless Network Management: A Survey and Tutorial

    cs.NI 2025-09 conditional novelty 4.0 of 10

    A survey and tutorial that organizes LLM-enabled wireless network optimization into formulation, solution, and verification stages, with case studies drawn from the authors' own prior papers.

  18. Probing the Robustness of Large Language Models Safety to Latent Perturbations

    cs.LG 2025-06 conditional novelty 4.0 of 10

    Randomized noise injected into hidden layers bypasses safety refusals in 12 open LLMs, and layer-wise adversarial training on the resulting benchmark reduces the attack's success.

  19. Risks of AI-driven product development and strategies for their mitigation

    cs.CY 2025-05 conditional novelty 4.0 of 10

    AI-driven product development will bring technical and societal risks; the paper proposes eight mitigation principles: human control, accountability, explainable and tested design, constrained and sandboxed systems, a...

  20. A Survey of Safety on Large Vision-Language Models: Attacks, Defenses and Evaluations

    cs.CR 2025-02 conditional novelty 4.0 of 10

    A survey of LVLM safety that adds a lifecycle taxonomy and new benchmark results showing Janus-Pro-7B has weaker safety than several open-source LVLMs.

Pith tools