Pith. sign in

REVIEW 12 cited by

Alignment of Language Agents

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2103.14659 v1 pith:H2A2V5QW submitted 2021-03-26 cs.AI cs.LG

classification cs.AIcs.LG
keywords someagentsissueslanguagemisspecificationbehaviouraldiscusshumans
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

For artificial intelligence to be beneficial to humans the behaviour of AI agents needs to be aligned with what humans want. In this paper we discuss some behavioural issues for language agents, arising from accidental misspecification by the system designer. We highlight some ways that misspecification can occur and discuss some behavioural issues that could arise from misspecification, including deceptive or manipulative language, and review some approaches for avoiding these issues.

Discussion (0). Sign in to comment.

Forward citations

Cited by 12 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. A Scalable Approach to Evaluating Moral Sensitivity in LLMs

    cs.CY 2026-07 conditional novelty 6.5 of 10

    Under morally irrelevant noise, eight LLMs preserve the semantic content of identified moral features above calibrated floors, despite significant changes in feature counts.

  2. Safety Alignment Should Be Made More Than Just A Few Attention Heads

    cs.CR 2025-08 conditional novelty 5.0 of 10

    Safety-critical attention heads are few, jailbreak prompts lower their refusal-direction signal, and fine-tuning with head-level dropout spreads safety and improves robustness.

  3. Linearly Decoding Refused Knowledge in Aligned Language Models

    cs.CL 2025-06 conditional novelty 5.0 of 10

    Linear probes recover jailbreak-only answers from aligned models' hidden states, sometimes transfer from base models, and correlate with pairwise preference rankings.

  4. The Value of Gen-AI Conversations: A bottom-up Framework for AI Value Alignment

    cs.CY 2025-06 conditional novelty 5.0 of 10

    A qualitative case study of a career-counseling chatbot derives nine core values and 32 types of value misalignment from 187 ethically sensitive conversations, proposing a bottom-up complement to top-down AI alignment.

  5. Survival Games: Human-LLM Strategic Showdowns under Severe Resource Scarcity

    cs.HC 2025-05 reject novelty 5.0 of 10

    In a simulated survival game with two rule-based agents and one LLM-powered robot, DeepSeek models showed more detected unethical actions than OpenAI models, and jailbreak prompts sharply increased violations.

  6. Understanding How Value Neurons Shape the Generation of Specified Values in LLMs

    cs.CL 2025-05 conditional novelty 5.0 of 10

    ValueLocate identifies value neurons via activation probability differences between opposing value prompts, and amplifying or suppressing these neurons alters G-EVAL value scores in four LLMs.

  7. Kaleidoscope Gallery: Exploring Ethics and Generative AI Through Art

    cs.CY 2025-05 conditional novelty 5.0 of 10

    Ethics experts' definitions of five ethical theories, rendered as DALL-E 3 images and re-evaluated by the same experts, yield eight themes showing how morality, society, and learned associations shape and bias the mod...

  8. The Alignment Auditor: A Bayesian Framework for Verifying and Refining LLM Objectives

    cs.LG 2025-10 conditional novelty 4.0 of 10

    Bayesian IRL with sequential posterior updates can recover a usable toxicity-reduction reward from LLM demonstrations and reproduce ground-truth RLHF detoxification.

  9. Governable AI: Provable Safety Under Extreme Threat Models

    cs.AI 2025-08 reject novelty 4.0 of 10

    An architecture that places a digitally signed, tamper-proof rule checker between an AI and its actuators is claimed to guarantee safety, but the proof assumes the very rules it must supply.

  10. Private, Verifiable, and Auditable AI Systems

    cs.CR 2025-08 conditional novelty 4.0 of 10

    A thesis demonstrating partial prototypes for zk-verifiable model evaluation and privacy-preserving retrieval, and arguing these pieces can compose into end-to-end auditable AI systems.

  11. The Future of Continual Learning in the Era of Foundation Models: Three Key Directions

    cs.LG 2025-06 conditional novelty 4.0 of 10

    Continual learning should pivot from weight-update-based methods to continual compositionality and orchestration of foundation models and agents.

  12. Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle

    cs.CL 2025-09 conditional novelty 3.0 of 10

    A survey that maps reinforcement learning methods, datasets, benchmarks, and open-source tools across the full training lifecycle of large language models, focusing on verifiable-reward reasoning.

Pith tools