REVIEW 25 cited by
Large Language Model Alignment: A Survey
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Recent years have witnessed remarkable progress made in large language models (LLMs). Such advancements, while garnering significant attention, have concurrently elicited various concerns. The potential of these models is undeniably vast; however, they may yield texts that are imprecise, misleading, or even detrimental. Consequently, it becomes paramount to employ alignment techniques to ensure these models to exhibit behaviors consistent with human values. This survey endeavors to furnish an extensive exploration of alignment methodologies designed for LLMs, in conjunction with the extant capability research in this domain. Adopting the lens of AI alignment, we categorize the prevailing methods and emergent proposals for the alignment of LLMs into outer and inner alignment. We also probe into salient issues including the models' interpretability, and potential vulnerabilities to adversarial attacks. To assess LLM alignment, we present a wide variety of benchmarks and evaluation methodologies. After discussing the state of alignment research for LLMs, we finally cast a vision toward the future, contemplating the promising avenues of research that lie ahead. Our aspiration for this survey extends beyond merely spurring research interests in this realm. We also envision bridging the gap between the AI alignment research community and the researchers engrossed in the capability exploration of LLMs for both capable and safe LLMs.
Forward citations
Cited by 25 Pith papers
-
When the Same Coefficients Reach Different Places: Asymmetric Realizability in Transplanting Tokenizers across Large Language Models
A 'breaker token' embedding can be inert in a donor LLM yet become a high-salience trigger after tokenizer transplant into a base LLM.
-
United Minds or Isolated Agents? Exploring Coordination of LLMs under Cognitive Load Theory
CoThinker, a multi-agent LLM framework inspired by human cognitive load theory, outperforms single-agent and debate baselines on reasoning-heavy benchmarks while underperforming on low-load instruction following.
-
Modeling Community Attitude through Reaction Tone: A Human-AI Collaborative Framework for Evaluating LLM Alignment with Linguistic Behaviors in Online Communities
Steering frontier LLMs with community identity fails to improve fidelity to authentic online reaction tones and attitudes, exposing a persistent realism gap.
-
Multi-Agent Debate Strategies: Survey, Taxonomy, and Challenges
A systematic review of 141 papers derives a three-axis taxonomy of multi-agent debate design (participants, interaction, agreement) and shows the field has converged on a narrow default pattern.
-
EPT Benchmark: Evaluation of Persian Trustworthiness in Large Language Models
A new Persian-Islamic trustworthiness benchmark ranks Claude highest and Qwen lowest across eight LLMs and finds safety is the weakest dimension.
-
Distributed AI Agents for Cognitive Underwater Robot Autonomy
UROSA controls underwater robots with distributed LLM/VLM agents, retrieval memory, and runtime code generation; feasibility is shown, but the claimed advantage over classical planners is not.
-
CALMA: A Process for Deriving Context-aligned Axes for Language Model Alignment
CALMA is a grounded-theory, participatory method for deriving community-specific language model alignment axes from open-ended user interactions and group discussion, piloted with two small groups.
-
Bradley-Terry and Multi-Objective Reward Modeling Are Complementary
Jointly training a Bradley-Terry preference head and a multi-attribute regression head on a shared embedding improves reward-model robustness to reward hacking and boosts multi-objective scoring performance.
-
Beyond Reactive Safety: Risk-Aware LLM Alignment via Long-Horizon Simulation
LLM-based long-horizon event simulation, used as a reward signal, is claimed to improve safety alignment and indirect-harm detection, but evaluation confounds simulation with the capability of the external projector model.
-
Disrupting Vision-Language Model-Driven Navigation Services via Adversarial Object Fusion
AdvOF crafts 3D adversarial objects that mislead VLM perception across multiple views and degrade VLN agent navigation success in simulation.
-
Position is Power: System Prompts as a Mechanism of Bias in Large Language Models (LLMs)
Placing demographic audience information in system prompts rather than user prompts shifts sentiment and ranking outputs across six commercial LLMs, but the design confounds position with instruction content.
-
Contrastive Weak-to-strong Generalization
Contrastive decoding between pre- and post-alignment weak models generates better supervision samples, improving weak-to-strong generalization on AlpacaEval2 and Arena-Hard.
-
Toward Preference-aligned Large Language Models via Residual-based Model Steering
Preference signals in LLM residual streams can be distilled into inference-time steering vectors that improve math and code benchmarks using only 100 preference pairs.
-
Re-Emergent Misalignment: How Narrow Fine-Tuning Erodes Safety Alignment in LLMs
Narrow fine-tuning on insecure code appears to erode prior safety alignment in Qwen2.5-Coder, with the misaligned model's internal activations moving back toward the base model.
-
Is Reasoning All You Need? Probing Bias in the Age of Reasoning Language Models
Using the CLEAR-Bias benchmark, the authors find that base language models are generally more robust to bias elicitation than CoT-prompted or reasoning-enabled models.
-
Perspective Dial: Measuring Perspective of Text and Guiding LLM Outputs
Perspective-Dial uses contrastive learning to build a perspective metric and greedy prompt optimization to steer LLM outputs toward a user-chosen viewpoint.
-
Understanding How University Guidelines Address Privacy and Security Issues of Generative AI in Academic Settings
Qualitative analysis of 46 university GenAI policy documents shows privacy and security concerns are acknowledged but inconsistently addressed, with vague terminology, reliance on existing frameworks, and limited conc...
-
Bridging Distribution Shift and AI Safety: Conceptual and Methodological Synergies
The paper proposes a one-to-one mapping between six causes of distribution shift and several AI safety issues, arguing for mutual method transfer through aligned definitions.
-
Salamandra Technical Report
Salamandra is an open, from-scratch multilingual LLM family with 2B, 7B, and 40B checkpoints, instruction-tuned variants, a vision proof-of-concept, and detailed evaluations across Iberian and European languages.
-
MARS: Margin and Semantic-Aware Data Augmentation for Reward Modeling
Concentrating synthetic preference paraphrases on low-margin pairs gives consistent but small reward-model and alignment gains in single-run experiments, while the abstract's semantic-aware, multi-benchmark claims are...
-
A Comprehensive Evaluation framework of Alignment Techniques for LLMs
The paper proposes a multi-dimensional framework to evaluate and compare LLM alignment techniques.
-
A Survey on Training-free Alignment of Large Language Models
A survey that catalogs and categorizes training-free LLM alignment methods into pre-decoding, in-decoding, and post-decoding, with a limited experimental comparison on one model.
-
Multi-level Value Alignment in Agentic AI Systems: Survey and Perspectives
A survey proposes a macro-meso-micro value framework for agentic AI alignment and maps applications, methods, and benchmarks onto it.
-
Large Language Models for EEG: A Comprehensive Survey and Taxonomy
A taxonomy and review of studies applying large language models to EEG signals, organized into four domains and three adaptation strategies.
-
Relative Bias: A Comparative Framework for Quantifying Bias in LLMs
A model is 'relatively biased' when its responses deviate from the consensus of a baseline LLM set, and this deviation can be scored by embedding distances or LLM judges plus equivalence tests.
Discussion (0). Continue with ORCID to comment.