Pith. sign in

REVIEW 6 cited by

What is the Alignment Objective of GRPO?

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2502.18548 v3 pith:FUFM5XAV submitted 2025-02-25 cs.LG cs.AI

classification cs.LGcs.AI
keywords preferenceaggregationgrpopolicyalgorithmpreferencesrewardmodel
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

In this note, we examine the aggregation of preferences achieved by the Group Policy Optimisation (GRPO) algorithm, a reinforcement learning method used to train advanced artificial intelligence models such as DeepSeek-R1-Zero and DeepSeekMath. The GRPO algorithm trains a policy using a reward preference model, which is computed by sampling a set of outputs for a given context, observing the corresponding rewards, and applying shift-and-scale normalisation to these reward values. Additionally, it incorporates a penalty function to discourage deviations from a reference policy. We present a framework that enables us to characterise the stationary policies of the GRPO algorithm. This analysis reveals that the aggregation of preferences differs fundamentally from standard logarithmic pooling, which is implemented by other approaches such as RLHF. The precise form of preference aggregation arises from the way the reward preference model is defined and from the penalty function, which we show to essentially correspond to the reverse Kullback-Leibler (KL) divergence between the aggregation policy and the reference policy. Interestingly, we demonstrate that for groups of size two, the reward preference model corresponds to pairwise comparison preferences, similar to those in other alignment methods based on pairwise comparison feedback. We provide explicit characterisations of the aggregate preference for binary questions, for groups of size two, and in the limit of large group size. This provides insights into the dependence of the aggregate preference on parameters such as the regularisation constant and the confidence margin of question answers. Finally, we discuss the aggregation of preferences obtained by modifying the GRPO algorithm to use direct KL divergence as the penalty or to use rewards without scale normalisation.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 6 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Small Agent Group is the Future of Digital Health

    cs.AI 2026-02 reject novelty 5.0 of 10

    A ten-agent group of 3B-4B LLMs with debate, retrieval, and role specialization beats 70B-72B single LLMs on clinical benchmarks while using less memory.

  2. Large Language Models Reasoning Abilities Under Non-Ideal Conditions After RL-Fine-Tuning

    cs.AI 2025-08 reject novelty 5.0 of 10

    RL fine-tuning of LLMs improves clean-benchmark accuracy while degrading accuracy under three injected-distractor evaluation scenarios, though one of the three scenarios contradicts the headline claim.

  3. Revisiting Group Relative Policy Optimization: Insights into On-Policy and Off-Policy Training

    cs.LG 2025-05 conditional novelty 5.0 of 10

    Off-policy GRPO, which estimates advantages from a slightly stale policy, matches or improves on-policy GRPO on math reasoning tasks and is supported by a new policy-improvement lower bound.

  4. Inference-Time Alignment Control for Diffusion Models with Reinforcement Learning Guidance

    cs.LG 2025-08 conditional novelty 4.0 of 10

    Blending a base diffusion model with its RL-finetuned version at sampling time lets users dial alignment strength, with the blend weight corresponding to the KL-regularization coefficient beta/w.

  5. LaViPlan : Language-Guided Visual Path Planning with RLVR

    cs.RO 2025-07 conditional novelty 4.0 of 10

    LaViPlan uses RLVR with GRPO and ADE/FDE rewards to fine-tune a 2B VLM for trajectory prediction, improving ADE/FDE on ROADWork and a normalized safety score on CODA-LM over supervised fine-tuning.

  6. Alignment and Safety in Large Language Models: Safety Mechanisms, Training Paradigms, and Emerging Challenges

    cs.AI 2025-07 reject novelty 1.0 of 10

    A broad survey of LLM alignment that catalogs objectives, benchmarks, SFT/RLHF/DPO methods, and safety challenges, without contributing new experimental or theoretical results.

Pith tools