Pith. sign in

REVIEW 1 cited by

UC-MOA: Utility-Conditioned Multi-Objective Alignment for Distributional Pareto-Optimality

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2503.10669 v2 pith:C442RACI submitted 2025-03-10 cs.CL cs.AI

classification cs.CLcs.AI
keywords alignmenthumanmodelsdistributionalllmsmulti-objectivenumericalpreferences
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Reinforcement Learning from Human Feedback (RLHF) has become a cornerstone for aligning large language models (LLMs) with human values. However, existing approaches struggle to capture the multi-dimensional, distributional nuances of human preferences. Methods such as RiC that directly inject raw reward values into prompts face significant numerical sensitivity issues--for instance, LLMs may fail to distinguish between 9.11 and 9.8--while alternatives like MORLHF, Rewarded Soups, and MODPO incur high computational costs by training multiple models. In this work, we introduce Utility-Conditioned Multi-Objective Alignment (UC-MOA), a novel framework that overcomes these limitations. Our approach leverages a diverse set of strictly increasing, non-linear utility functions to transform user-specified preferences into symbolic tokens, which are then used to condition a single LLM. This design not only mitigates numerical reasoning challenges but also substantially reduces training overhead, yielding models that achieve superior Pareto fronts and robust alignment across complex reward dimensions.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. ArGen: Auto-Regulation of Generative AI via GRPO and Policy-as-Code

    cs.CY 2025-09 conditional novelty 5.0 of 10

    A framework combining GRPO, LLM-as-judge rewards, and OPA-style policy checks improves a 1B medical assistant's domain-scope adherence by 70.9% in a 100-scenario LLM-judge evaluation.

Pith tools