REVIEW 12 cited by
Alignment of Language Agents
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
For artificial intelligence to be beneficial to humans the behaviour of AI agents needs to be aligned with what humans want. In this paper we discuss some behavioural issues for language agents, arising from accidental misspecification by the system designer. We highlight some ways that misspecification can occur and discuss some behavioural issues that could arise from misspecification, including deceptive or manipulative language, and review some approaches for avoiding these issues.
Forward citations
Cited by 12 Pith papers
-
A Scalable Approach to Evaluating Moral Sensitivity in LLMs
Under morally irrelevant noise, eight LLMs preserve the semantic content of identified moral features above calibrated floors, despite significant changes in feature counts.
-
Safety Alignment Should Be Made More Than Just A Few Attention Heads
Safety-critical attention heads are few, jailbreak prompts lower their refusal-direction signal, and fine-tuning with head-level dropout spreads safety and improves robustness.
-
Linearly Decoding Refused Knowledge in Aligned Language Models
Linear probes recover jailbreak-only answers from aligned models' hidden states, sometimes transfer from base models, and correlate with pairwise preference rankings.
-
The Value of Gen-AI Conversations: A bottom-up Framework for AI Value Alignment
A qualitative case study of a career-counseling chatbot derives nine core values and 32 types of value misalignment from 187 ethically sensitive conversations, proposing a bottom-up complement to top-down AI alignment.
-
Survival Games: Human-LLM Strategic Showdowns under Severe Resource Scarcity
In a simulated survival game with two rule-based agents and one LLM-powered robot, DeepSeek models showed more detected unethical actions than OpenAI models, and jailbreak prompts sharply increased violations.
-
Understanding How Value Neurons Shape the Generation of Specified Values in LLMs
ValueLocate identifies value neurons via activation probability differences between opposing value prompts, and amplifying or suppressing these neurons alters G-EVAL value scores in four LLMs.
-
Kaleidoscope Gallery: Exploring Ethics and Generative AI Through Art
Ethics experts' definitions of five ethical theories, rendered as DALL-E 3 images and re-evaluated by the same experts, yield eight themes showing how morality, society, and learned associations shape and bias the mod...
-
The Alignment Auditor: A Bayesian Framework for Verifying and Refining LLM Objectives
Bayesian IRL with sequential posterior updates can recover a usable toxicity-reduction reward from LLM demonstrations and reproduce ground-truth RLHF detoxification.
-
Governable AI: Provable Safety Under Extreme Threat Models
An architecture that places a digitally signed, tamper-proof rule checker between an AI and its actuators is claimed to guarantee safety, but the proof assumes the very rules it must supply.
-
Private, Verifiable, and Auditable AI Systems
A thesis demonstrating partial prototypes for zk-verifiable model evaluation and privacy-preserving retrieval, and arguing these pieces can compose into end-to-end auditable AI systems.
-
The Future of Continual Learning in the Era of Foundation Models: Three Key Directions
Continual learning should pivot from weight-update-based methods to continual compositionality and orchestration of foundation models and agents.
-
Reinforcement Learning Meets Large Language Models: A Survey of Advancements and Applications Across the LLM Lifecycle
A survey that maps reinforcement learning methods, datasets, benchmarks, and open-source tools across the full training lifecycle of large language models, focusing on verifiable-reward reasoning.
Discussion (0). Sign in to comment.