Pith. sign in

REVIEW 25 cited by

Large Language Model Alignment: A Survey

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2309.15025 v1 pith:MBLCWCDM submitted 2023-09-26 cs.CL cs.AI

classification cs.CLcs.AI
keywords alignmentllmsresearchmodelssurveycapabilityexplorationlanguage
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Recent years have witnessed remarkable progress made in large language models (LLMs). Such advancements, while garnering significant attention, have concurrently elicited various concerns. The potential of these models is undeniably vast; however, they may yield texts that are imprecise, misleading, or even detrimental. Consequently, it becomes paramount to employ alignment techniques to ensure these models to exhibit behaviors consistent with human values. This survey endeavors to furnish an extensive exploration of alignment methodologies designed for LLMs, in conjunction with the extant capability research in this domain. Adopting the lens of AI alignment, we categorize the prevailing methods and emergent proposals for the alignment of LLMs into outer and inner alignment. We also probe into salient issues including the models' interpretability, and potential vulnerabilities to adversarial attacks. To assess LLM alignment, we present a wide variety of benchmarks and evaluation methodologies. After discussing the state of alignment research for LLMs, we finally cast a vision toward the future, contemplating the promising avenues of research that lie ahead. Our aspiration for this survey extends beyond merely spurring research interests in this realm. We also envision bridging the gap between the AI alignment research community and the researchers engrossed in the capability exploration of LLMs for both capable and safe LLMs.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 25 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. OpenAlex reports about 34 citations worldwide. Full citation record

  1. When the Same Coefficients Reach Different Places: Asymmetric Realizability in Transplanting Tokenizers across Large Language Models

    cs.LG 2025-12 conditional novelty 7.0 of 10

    A 'breaker token' embedding can be inert in a donor LLM yet become a high-salience trigger after tokenizer transplant into a base LLM.

  2. United Minds or Isolated Agents? Exploring Coordination of LLMs under Cognitive Load Theory

    cs.AI 2025-06 conditional novelty 7.0 of 10

    CoThinker, a multi-agent LLM framework inspired by human cognitive load theory, outperforms single-agent and debate baselines on reasoning-heavy benchmarks while underperforming on low-load instruction following.

  3. Modeling Community Attitude through Reaction Tone: A Human-AI Collaborative Framework for Evaluating LLM Alignment with Linguistic Behaviors in Online Communities

    cs.CL 2026-04 conditional novelty 6.5 of 10

    Steering frontier LLMs with community identity fails to improve fidelity to authentic online reaction tones and attitudes, exposing a persistent realism gap.

  4. Multi-Agent Debate Strategies: Survey, Taxonomy, and Challenges

    cs.SE 2026-07 accept novelty 6.0 of 10

    A systematic review of 141 papers derives a three-axis taxonomy of multi-agent debate design (participants, interaction, agreement) and shows the field has converged on a narrow default pattern.

  5. EPT Benchmark: Evaluation of Persian Trustworthiness in Large Language Models

    cs.CL 2025-09 conditional novelty 6.0 of 10

    A new Persian-Islamic trustworthiness benchmark ranks Claude highest and Qwen lowest across eight LLMs and finds safety is the weakest dimension.

  6. Distributed AI Agents for Cognitive Underwater Robot Autonomy

    cs.RO 2025-07 reject novelty 6.0 of 10

    UROSA controls underwater robots with distributed LLM/VLM agents, retrieval memory, and runtime code generation; feasibility is shown, but the claimed advantage over classical planners is not.

  7. CALMA: A Process for Deriving Context-aligned Axes for Language Model Alignment

    cs.CY 2025-07 conditional novelty 6.0 of 10

    CALMA is a grounded-theory, participatory method for deriving community-specific language model alignment axes from open-ended user interactions and group discussion, piloted with two small groups.

  8. Bradley-Terry and Multi-Objective Reward Modeling Are Complementary

    cs.LG 2025-07 conditional novelty 6.0 of 10

    Jointly training a Bradley-Terry preference head and a multi-attribute regression head on a shared embedding improves reward-model robustness to reward hacking and boosts multi-objective scoring performance.

  9. Beyond Reactive Safety: Risk-Aware LLM Alignment via Long-Horizon Simulation

    cs.AI 2025-06 reject novelty 6.0 of 10

    LLM-based long-horizon event simulation, used as a reward signal, is claimed to improve safety alignment and indirect-harm detection, but evaluation confounds simulation with the capability of the external projector model.

  10. Disrupting Vision-Language Model-Driven Navigation Services via Adversarial Object Fusion

    cs.CR 2025-05 conditional novelty 6.0 of 10

    AdvOF crafts 3D adversarial objects that mislead VLM perception across multiple views and degrade VLN agent navigation success in simulation.

  11. Position is Power: System Prompts as a Mechanism of Bias in Large Language Models (LLMs)

    cs.CY 2025-05 reject novelty 6.0 of 10

    Placing demographic audience information in system prompts rather than user prompts shifts sentiment and ranking outputs across six commercial LLMs, but the design confounds position with instruction content.

  12. Contrastive Weak-to-strong Generalization

    cs.CL 2025-10 conditional novelty 5.0 of 10

    Contrastive decoding between pre- and post-alignment weak models generates better supervision samples, improving weak-to-strong generalization on AlpacaEval2 and Arena-Hard.

  13. Toward Preference-aligned Large Language Models via Residual-based Model Steering

    cs.CL 2025-09 conditional novelty 5.0 of 10

    Preference signals in LLM residual streams can be distilled into inference-time steering vectors that improve math and code benchmarks using only 100 preference pairs.

  14. Re-Emergent Misalignment: How Narrow Fine-Tuning Erodes Safety Alignment in LLMs

    cs.LG 2025-07 conditional novelty 5.0 of 10

    Narrow fine-tuning on insecure code appears to erode prior safety alignment in Qwen2.5-Coder, with the misaligned model's internal activations moving back toward the base model.

  15. Is Reasoning All You Need? Probing Bias in the Age of Reasoning Language Models

    cs.CL 2025-07 conditional novelty 5.0 of 10

    Using the CLEAR-Bias benchmark, the authors find that base language models are generally more robust to bias elicitation than CoT-prompted or reasoning-enabled models.

  16. Perspective Dial: Measuring Perspective of Text and Guiding LLM Outputs

    cs.CL 2025-06 reject novelty 5.0 of 10

    Perspective-Dial uses contrastive learning to build a perspective metric and greedy prompt optimization to steer LLM outputs toward a user-chosen viewpoint.

  17. Understanding How University Guidelines Address Privacy and Security Issues of Generative AI in Academic Settings

    cs.HC 2025-06 conditional novelty 5.0 of 10

    Qualitative analysis of 46 university GenAI policy documents shows privacy and security concerns are acknowledged but inconsistently addressed, with vague terminology, reliance on existing frameworks, and limited conc...

  18. Bridging Distribution Shift and AI Safety: Conceptual and Methodological Synergies

    cs.LG 2025-05 conditional novelty 5.0 of 10

    The paper proposes a one-to-one mapping between six causes of distribution shift and several AI safety issues, arguing for mutual method transfer through aligned definitions.

  19. Salamandra Technical Report

    cs.CL 2025-02 conditional novelty 5.0 of 10

    Salamandra is an open, from-scratch multilingual LLM family with 2B, 7B, and 40B checkpoints, instruction-tuned variants, a vision proof-of-concept, and detailed evaluations across Iberian and European languages.

  20. MARS: Margin and Semantic-Aware Data Augmentation for Reward Modeling

    cs.LG 2026-02 reject novelty 4.0 of 10

    Concentrating synthetic preference paraphrases on low-margin pairs gives consistent but small reward-model and alignment gains in single-run experiments, while the abstract's semantic-aware, multi-benchmark claims are...

  21. A Comprehensive Evaluation framework of Alignment Techniques for LLMs

    cs.CL 2025-08 unverdicted novelty 4.0 of 10

    The paper proposes a multi-dimensional framework to evaluate and compare LLM alignment techniques.

  22. A Survey on Training-free Alignment of Large Language Models

    cs.CL 2025-08 conditional novelty 4.0 of 10

    A survey that catalogs and categorizes training-free LLM alignment methods into pre-decoding, in-decoding, and post-decoding, with a limited experimental comparison on one model.

  23. Multi-level Value Alignment in Agentic AI Systems: Survey and Perspectives

    cs.AI 2025-06 conditional novelty 4.0 of 10

    A survey proposes a macro-meso-micro value framework for agentic AI alignment and maps applications, methods, and benchmarks onto it.

  24. Large Language Models for EEG: A Comprehensive Survey and Taxonomy

    eess.SP 2025-06 conditional novelty 4.0 of 10

    A taxonomy and review of studies applying large language models to EEG signals, organized into four domains and three adaptation strategies.

  25. Relative Bias: A Comparative Framework for Quantifying Bias in LLMs

    cs.CL 2025-05 conditional novelty 4.0 of 10

    A model is 'relatively biased' when its responses deviate from the consensus of a baseline LLM set, and this deviation can be scored by embedding distances or LLM judges plus equivalence tests.

Pith tools