REVIEW 3 cited by
Process for Adapting Language Models to Society (PALMS) with Values-Targeted Datasets
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Language models can generate harmful and biased outputs and exhibit undesirable behavior according to a given cultural context. We propose a Process for Adapting Language Models to Society (PALMS) with Values-Targeted Datasets, an iterative process to significantly change model behavior by crafting and fine-tuning on a dataset that reflects a predetermined set of target values. We evaluate our process using three metrics: quantitative metrics with human evaluations that score output adherence to a target value, toxicity scoring on outputs; and qualitative metrics analyzing the most common word associated with a given social category. Through each iteration, we add additional training dataset examples based on observed shortcomings from evaluations. PALMS performs significantly better on all metrics compared to baseline and control models for a broad range of GPT-3 language model sizes without compromising capability integrity. We find that the effectiveness of PALMS increases with model size. We show that significantly adjusting language model behavior is feasible with a small, hand-curated dataset.
Forward citations
Cited by 3 Pith papers
-
Moderating Harm: Benchmarking Large Language Models for Cyberbullying Detection in YouTube Comments
In a zero-shot benchmark on 5,080 YouTube comments, GPT-4.1 achieved the best F1 balance at 0.863, while Gemini favored recall and Claude favored precision.
-
Governance-as-a-Service: A Multi-Agent Framework for AI System Compliance and Policy Enforcement
An external policy-enforcement layer with a trust score claims to block risky AI-agent actions in simulations, but its core formula is inconsistent across the paper.
-
Alignment and Safety in Large Language Models: Safety Mechanisms, Training Paradigms, and Emerging Challenges
A broad survey of LLM alignment that catalogs objectives, benchmarks, SFT/RLHF/DPO methods, and safety challenges, without contributing new experimental or theoretical results.
Discussion (0). Sign in to comment.