Pith. sign in

REVIEW 2 cited by

Configurable Safety Tuning of Language Models with Synthetic Preference Data

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2404.00495 v1 pith:FZR2W743 submitted 2024-03-30 cs.CL cs.AI

classification cs.CLcs.AI
keywords safetyconfigurabledatapreferenceconfigurationslanguagellmsmethod
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

State-of-the-art language model fine-tuning techniques, such as Direct Preference Optimization (DPO), restrict user control by hard-coding predefined behaviors into the model. To address this, we propose a novel method, Configurable Safety Tuning (CST), that augments DPO using synthetic preference data to facilitate flexible safety configuration of LLMs at inference time. CST overcomes the constraints of vanilla DPO by introducing a system prompt specifying safety configurations, enabling LLM deployers to disable/enable safety preferences based on their need, just changing the system prompt. Our experimental evaluations indicate that CST successfully manages different safety configurations and retains the original functionality of LLMs, showing it is a robust method for configurable deployment. Data and models available at https://github.com/vicgalle/configurable-safety-tuning

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. MetaSC: Test-Time Safety Specification Optimization for Language Models

    cs.CL 2025-02 conditional novelty 6.0 of 10

    MetaSC improves language model safety by using a meta-critic to iteratively rewrite the safety specification that guides self-critique at inference time.

  2. Configurable Preference Tuning with Rubric-Guided Synthetic Data

    cs.CL 2025-06 conditional novelty 5.0 of 10

    CPT fine-tunes LLMs with DPO on rubric-guided synthetic preferences so that a system prompt can reconfigure output style at inference, with in-distribution accuracy gains over baselines.

Pith tools