Pith. sign in

REVIEW 7 cited by

Jointly Reinforcing Diversity and Quality in Language Model Generations

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2509.02534 v1 pith:6HJLP3I4 submitted 2025-09-02 cs.CL cs.LG

classification cs.CLcs.LG
keywords diversityqualitydarlingtaskscreativehigherjointlylanguage
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Post-training of Large Language Models (LMs) often prioritizes accuracy and helpfulness at the expense of diversity. This creates a tension: while post-training improves response quality, it also sharpens output distributions and reduces the range of ideas, limiting the usefulness of LMs in creative and exploratory tasks such as brainstorming, storytelling, or problem solving. We address this challenge with Diversity-Aware Reinforcement Learning (DARLING), a framework that jointly optimizes for response quality and semantic diversity. At its core, DARLING introduces a learned partition function to measure diversity beyond surface-level lexical variations. This diversity signal is then combined with a quality reward during online reinforcement learning, encouraging models to generate outputs that are both high-quality and distinct. Experiments across multiple model families and sizes show that DARLING generalizes to two regimes: non-verifiable tasks (instruction following and creative writing) and verifiable tasks (competition math). On five benchmarks in the first setting, DARLING consistently outperforms quality-only RL baselines, producing outputs that are simultaneously of higher quality and novelty. In the second setting, DARLING achieves higher pass@1 (solution quality) and pass@k (solution variety). Most strikingly, explicitly optimizing for diversity catalyzes exploration in online RL, which manifests itself as higher-quality responses.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 7 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Cast a Wider Net: Coordinated Pass@K Policy Optimization for Code Reasoning

    cs.CL 2026-05 unverdicted novelty 7.0 of 10

    CPPO coordinates K=4 distinct strategy tuples via planner and solver with multiplicative reward, improving pass@4 over baselines on APPS, CodeContests, and LiveCodeBench.

  2. The One-Word Census: Answer-Choice Conformity Across 44 Language Models

    cs.CL 2026-07 conditional novelty 6.0 of 10

    Forty-four language models asked to name one thing per category converge on the same modal answers far more than people do, with newest flagships most conformist and persona-tuned models most divergent.

  3. Vocabulary Dropout for Curriculum Diversity in LLM Co-Evolution

    cs.CL 2026-04 unverdicted novelty 6.0 of 10

    Vocabulary dropout prevents diversity collapse in LLM co-evolution by masking proposer logits, yielding average +4.4 point solver gains on mathematical reasoning benchmarks at 8B scale.

  4. Representation-Based Exploration for Language Models: From Test-Time to Post-Training

    cs.LG 2025-10 conditional novelty 6.0 of 10

    Representation-based elliptical bonuses improve inference-time and post-training pass@k for LLM reasoning, but the headline AIME result is tainted by validation/test overlap.

  5. Outcome-based Exploration for LLM Reasoning

    cs.LG 2025-09 conditional novelty 6.0 of 10

    Outcome-based exploration bonuses (UCB-Con and Batch) improve pass@1 and pass@32 for LLM math reasoning while slowing diversity collapse, supported by a bandit model with a strong generalization assumption.

  6. When Correct Solutions Repeat: Rarity-Aware Credit Redistribution for GRPO

    cs.AI 2026-08 conditional novelty 5.0 of 10

    Cue-GRPO reweights GRPO credit toward rare correct solution structures with deterministic cues, improving high-budget AIME pass@k on two 7-8B models.

  7. Optimizing Diversity and Quality through Base-Aligned Model Collaboration

    cs.CL 2025-11 conditional novelty 5.0 of 10

    At every token, a router switches between a base LLM and its aligned counterpart based on uncertainty and word type, improving the diversity-quality trade-off across open-ended generation tasks.

Pith tools