REVIEW 20 cited by
Personalized Soups: Personalized Large Language Model Alignment via Post-hoc Parameter Merging
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
While Reinforcement Learning from Human Feedback (RLHF) aligns Large Language Models (LLMs) with general, aggregate human preferences, it is suboptimal for learning diverse, individual perspectives. In this work, we study Reinforcement Learning from Personalized Human Feedback (RLPHF) problem, wherein LLMs are aligned to multiple (sometimes conflicting) preferences by modeling alignment as a Multi-Objective Reinforcement Learning (MORL) problem. Compared to strong single-objective baselines, we show that we can achieve personalized alignment by decomposing preferences into multiple dimensions. These dimensions are defined based on personalizations that are declared as desirable by the user. In this work, we show that they can be efficiently trained independently in a distributed manner and combined effectively post-hoc through parameter merging. The code is available at https://github.com/joeljang/RLPHF.
Forward citations
Cited by 20 Pith papers
-
P-Check: Advancing Personalized Reward Model via Learning to Generate Dynamic Checklist
P-Check trains a checklist generator that produces query-specific, user-weighted evaluation criteria, improving LLM-judge reward accuracy on personalization benchmarks.
-
Cautious Context Steering for Language Model Personalization
CCS is a learned per-token gate for context steering that improves personalized generation on PRISM and four out-of-distribution benchmarks while avoiding a second forward pass per decoding step.
-
Don't Mix Rewards, Mix Policies: Policy Decomposition and Optimization for Multi-Reward RL
PRISM trains one positive policy per reward plus one global negative policy and merges their token logits, improving multi-reward RL for LLMs with inference-time controllability.
-
A Roadmap to Impactful Pluralistic Alignment Research
Pluralistic alignment research has produced no public evidence of adoption in deployed frontier models, so the field should focus on empirical justification, settled goals, and hill-climbable evaluations.
-
Instant Personalized Large Language Model Adaptation via Hypernetwork
A hypernetwork maps a user profile to LoRA adapter weights in a single forward pass, matching or beating per-user fine-tuning at a fraction of deployment cost.
-
The Alignment Veto: How Safety Training Suppresses Cultural Knowledge in LLMs
The full text builds the MENA Values benchmark (864 questions, 7 models) and reports that LLM cultural answers shift with language, decline with reasoning prompts, and hide strong internal preferences behind refusals—...
-
Simultaneous Multi-objective Alignment Across Verifiable and Non-verifiable Rewards
MAHALO aligns LLMs to multiple objectives in one model via per-objective action heads and PRM-guided decoding, improving math, value, and tutoring metrics jointly.
-
SharedRep-RLHF: A Shared Representation Approach to RLHF with Diverse Preferences
SharedRep-RLHF learns a shared preference representation across groups to improve worst-case reward estimates for minority annotators, but the theoretical guarantees are undermined by proof errors.
-
SynthesizeMe! Inducing Persona-Guided Prompts for Personalized Reward Models in LLMs
SynthesizeMe induces synthetic user personas from a few pairwise preferences and uses them to improve personalized LLM judging accuracy by a few points on a new benchmark.
-
Aligning VLM Assistants with Personalized Situated Cognition
The authors present PCogAlignBench, an 18k-sample benchmark of visual scenes with role-based users, and PCogAlign, a framework using a cognition-aware reward model to produce responses aligned with each user's roles.
-
Bounded Rationality for LLMs: Satisficing Alignment at Inference-Time
SITAlign is an inference-time constrained decoder that maximizes a primary reward while enforcing thresholds on secondary rewards, and it reports better primary-reward win-tie rates than weighted-objective decoding.
-
Multi-objective Large Language Model Alignment with Hierarchical Experts
HoE claims to align a single LLM to any preference vector over multiple objectives using training-free LoRA experts, lightweight trained routers, and nearest-neighbor preference routing.
-
Extended Inductive Reasoning for Personalized Preference Inference from Behavioral Signals
A 7B model trained with synthetic reasoning demonstrations plus reinforcement learning infers explicit user preference descriptions from behavioral signals, improving personalized response judging and generation.
-
Tuning-Free Personalized Alignment via Trial-Error-Explain In-Context Learning
TICL improves style personalization by iteratively adding model-generated negative examples and explanations to an in-context prompt, beating fine-tuned baselines in LLM-judged comparisons without any parameter updates.
-
LEMUR: Learning to Align with Multi-Objective Reinforcement Learning from Preference Feedback
LEMUR jointly learns a separate reward model for each teacher's preferences and uses them to train a population of multi-objective policies, beating baselines that merge feedback into one reward.
-
AMoPO: Adaptive Multi-objective Preference Optimization without Reward Models and Reference Models
AMoPO uses the model's own token probabilities to define Gaussian-sampled weights, combining per-dimension SimPO-style losses for reference-free multi-objective alignment.
-
MidPO: Dual Preference Optimization for Safety and Helpfulness in Large Language Models via a Mixture of Experts Framework
An MoE alignment pipeline combining two SPE-DPO-trained experts and a learned routing network reports better safety and helpfulness scores than existing dual-preference alignment baselines.
-
Do Implicit Personalization and Explicit Styles Conflict? PsPLUG: A Lightweight Plug-in for Balancing Personalization and Style in Customized LLMs
PsPLUG, a soft-prompt plug-in trained with style-conditioned preference pairs, preserves user identity under explicit style instructions and lets users tune personalization strength via an α scalar.
-
T-POP: Test-Time Personalization with Online Preference Feedback
T-POP uses dueling-bandit token selection to learn a reward function online from pairwise user feedback, enabling test-time personalization of a frozen LLM without fine-tuning.
-
AI Agent Behavioral Science
AI agents should be studied as behavioral entities shaped by context and interaction, not only as trained models.
Discussion (0). Continue with ORCID to comment.