Pith. sign in

REVIEW 1 cited by

Stepwise Alignment for Constrained Language Model Policy Optimization

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2404.11049 v3 pith:WZLANFQE submitted 2024-04-17 cs.LG cs.AIcs.CL

classification cs.LGcs.AIcs.CL
keywords policysacpoalignmentoptimizationsafetylanguagealgorithmsconstrained
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Safety and trustworthiness are indispensable requirements for real-world applications of AI systems using large language models (LLMs). This paper formulates human value alignment as an optimization problem of the language model policy to maximize reward under a safety constraint, and then proposes an algorithm, Stepwise Alignment for Constrained Policy Optimization (SACPO). One key idea behind SACPO, supported by theory, is that the optimal policy incorporating reward and safety can be directly obtained from a reward-aligned policy. Building on this key idea, SACPO aligns LLMs step-wise with each metric while leveraging simple yet powerful alignment algorithms such as direct preference optimization (DPO). SACPO offers several advantages, including simplicity, stability, computational efficiency, and flexibility of algorithms and datasets. Under mild assumptions, our theoretical analysis provides the upper bounds on optimality and safety constraint violation. Our experimental results show that SACPO can fine-tune Alpaca-7B better than the state-of-the-art method in terms of both helpfulness and harmlessness.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Vulnerability Mitigation for Safety-Aligned Language Models via Debiasing

    cs.AI 2025-02 conditional novelty 6.0 of 10

    A decoding-time method called TSDI estimates and removes the context-free refusal bias caused by safety alignment, improving helpfulness while keeping safety.

Pith tools