Pith. sign in

Paper Citation Record · LEDGER

Quality-constrained Entropy Maximization Policy Optimization for LLM Diversity

As of 15 August 2026, this Paper Citation Record lists 21 of 21 outbound references and 0 inbound Pith citation observations for arXiv:2602.15894.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2602.15894 v2

Coverage vector

measured 21 of 21 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-03T01:05:26.862251Z

measured 21 of 21 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

21 of 21 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved21
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 14262ad3-f793-40f6-82a3-2fdc91bf01c6 · outbound

This paper cites The Hyperfitting Phenomenon: Sharpening and Stabilizing LLMs for Open-Ended Text Generation.

Quality-constrained Entropy Maximization Policy Optimization for LLM Diversity The Hyperfitting Phenomenon: Sharpening and Stabilizing LLMs for Open-Ended Text Generation

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-03T01:05:25.931091Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T01:05:25.931091Z digest=sha256:a843e2214d805949af05d346f63b03e9a99d054ba69f3a246c38579e767ba35c

Observation 1b6c57f4-cddc-45ac-9548-d6f77e4e7b58 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Quality-constrained Entropy Maximization Policy Optimization for LLM Diversity DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-03T01:05:26.065231Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T01:05:26.065231Z digest=sha256:2a05fff6f8a1cccf2ca97964d6d344e7b9720ce808ed243ae40732ed312ff3b0

Observation 324fc313-9ff8-4574-be95-66ff29ff372a · outbound

This paper cites Benchmarking Linguistic Diversity of Large Language Models.

Quality-constrained Entropy Maximization Policy Optimization for LLM Diversity Benchmarking Linguistic Diversity of Large Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-03T01:05:26.105939Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T01:05:26.105939Z digest=sha256:72376169b157c1a5000b41acfa380f1754298919dba285b5745263bd9ad3ef76

Observation ab766472-f296-4cd2-aa58-0dad06b08328 · outbound

This paper cites Adam: A Method for Stochastic Optimization.

Quality-constrained Entropy Maximization Policy Optimization for LLM Diversity Adam: A Method for Stochastic Optimization

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-03T01:05:26.157112Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T01:05:26.157112Z digest=sha256:1a599f98eb49ff9dc889d595b3bb4be387a045dcee817600305282d2a021eacd

Observation de76183d-5471-4896-a0cb-a9876e4efa29 · outbound

This paper cites Llama 3.2: Revolutionizing edge ai and vision with open, customizable models.Meta AI Blog.

Quality-constrained Entropy Maximization Policy Optimization for LLM Diversity Llama 3.2: Revolutionizing edge ai and vision with open, customizable models.Meta AI Blog

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-03T01:05:26.269779Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T01:05:26.269779Z digest=sha256:76dd48fabbfd676ad0fcbadc4914668876fa9d361209c65e295a672f5fb94b4d

Observation 1cb8f421-bf9c-4558-a239-293ff200129e · outbound

This paper cites One fish, two fish, but not the whole sea: Alignment reduces language models' conceptual diversity.

Quality-constrained Entropy Maximization Policy Optimization for LLM Diversity One fish, two fish, but not the whole sea: Alignment reduces language models' conceptual diversity

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-03T01:05:26.300374Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T01:05:26.300374Z digest=sha256:5510af44b6a7301c64b8a85e517a6d71c508f3aa0bd8916b366ddc09dfc8ea22

Observation 59eb3f4c-8325-41e5-8839-808737117e40 · outbound

This paper cites EnTRPO: Trust Region Policy Optimization Method with Entropy Regularization.

Quality-constrained Entropy Maximization Policy Optimization for LLM Diversity EnTRPO: Trust Region Policy Optimization Method with Entropy Regularization

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-03T01:05:26.382864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T01:05:26.382864Z digest=sha256:df69d71da27d5b4d06c99a67c16aa8d4ddaaae0ea16ae9f02a81aa747d8e416f

Observation 8ba30e36-23e4-44aa-bc17-7e45b0c1344d · outbound

This paper cites Curiosity-Driven Reinforcement Learning from Human Feedback.

Quality-constrained Entropy Maximization Policy Optimization for LLM Diversity Curiosity-Driven Reinforcement Learning from Human Feedback

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-03T01:05:26.423214Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T01:05:26.423214Z digest=sha256:4b42395f4dc09a1bda2de48144801dc0d9b966f1771bf66dd39f902137d634ba

Observation 73a7d633-264e-4bb9-8e32-d8f4eb74b6e8 · outbound

This paper cites Zephyr: Direct Distillation of LM Alignment.

Quality-constrained Entropy Maximization Policy Optimization for LLM Diversity Zephyr: Direct Distillation of LM Alignment

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-03T01:05:26.526082Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T01:05:26.526082Z digest=sha256:3b9f9fe364a0c80cf4d2526b297c6eae98fd2aac97a537f22fb398fa3a5cdca3

Observation 96123e1f-eb81-4f60-b6d8-dddf9424ba72 · outbound

This paper cites Improving Diversity in Language Models: When Temperature Fails, Change the Loss.

Quality-constrained Entropy Maximization Policy Optimization for LLM Diversity Improving Diversity in Language Models: When Temperature Fails, Change the Loss

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-03T01:05:26.587811Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T01:05:26.587811Z digest=sha256:7243a5716ed07137f878d24fe2d66766eee946b2a609304a3b2ee808cdf3a524

Observation 1923558e-a596-44dd-848e-4e2abe2cbf80 · outbound

This paper cites Diversity-oriented data augmentation with large language models.arXiv preprint arXiv:2502.11671,.

Quality-constrained Entropy Maximization Policy Optimization for LLM Diversity Diversity-oriented data augmentation with large language models.arXiv preprint arXiv:2502.11671,

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-03T01:05:26.649957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T01:05:26.649957Z digest=sha256:dd278d2620e68fe09e64373ca2e878bc351ae0011a1bb54994b436cfa582ff1b

Observation 04f5bc96-1512-40fd-a245-7e7e8c4b39c6 · outbound

This paper cites Understanding the Effects of RLHF on the Quality and Detectability of LLM-Generated Texts.

Quality-constrained Entropy Maximization Policy Optimization for LLM Diversity Understanding the Effects of RLHF on the Quality and Detectability of LLM-Generated Texts

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-03T01:05:26.711004Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T01:05:26.711004Z digest=sha256:1c8b6e436f583396c8c0bc97e5a6ea69d85ecbc1b22021c9464ebbd11e11c869

Observation 621fba89-e0af-45d2-9940-4854dda28128 · outbound

This paper cites Qwen3 Technical Report.

Quality-constrained Entropy Maximization Policy Optimization for LLM Diversity Qwen3 Technical Report

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-03T01:05:26.772492Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T01:05:26.772492Z digest=sha256:ae69fae202214abc4a8c143870c24786851ec07d40bfa8662dda17383c00b238

Observation c6a94ba0-275f-4af2-a161-fbb381039acf · outbound

This paper cites Gvpo: Group variance policy optimization for large language model post-training.arXiv preprint arXiv:2504.19599,.

Quality-constrained Entropy Maximization Policy Optimization for LLM Diversity Gvpo: Group variance policy optimization for large language model post-training.arXiv preprint arXiv:2504.19599,

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-03T01:05:26.831986Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T01:05:26.831986Z digest=sha256:2ff91f1b7f011c47c6b54b2efd0d2b4e4812c9e6aca8c3b883d9fef966d31462

Observation fb2d0947-1140-4751-840c-0d1ce9197c10 · outbound

This paper cites Balancing Diversity and Risk in LLM Sampling: How to Select Your Method and Parameter for Open-Ended Text Generation.

Quality-constrained Entropy Maximization Policy Optimization for LLM Diversity Balancing Diversity and Risk in LLM Sampling: How to Select Your Method and Parameter for Open-Ended Text Generation

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-03T01:05:26.862251Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T01:05:26.862251Z digest=sha256:e78c6363f14dc3f01449ceffe21a482fff96ccb488bc336962342b6a10044237

Observation e13cc707-0d98-477d-9743-5a4d3dd6cc0b · outbound

This paper cites Qwen2.5 Technical Report.

Quality-constrained Entropy Maximization Policy Optimization for LLM Diversity Qwen2.5 Technical Report

Reference 1999

Resolution
unresolved
no resolver link, observed 2026-08-03T01:05:26.460098Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T01:05:26.460098Z digest=sha256:604f60e6b231f21b8b27e13a6f73aaef1c74d094bcb09ed83aba6afa6c4e40f8

Observation b65e7330-a004-4772-9c64-4fb025066304 · outbound

This paper cites Understanding the Effects of RLHF on LLM Generalisation and Diversity.

Quality-constrained Entropy Maximization Policy Optimization for LLM Diversity Understanding the Effects of RLHF on LLM Generalisation and Diversity

Reference 2014

Resolution
unresolved
no resolver link, observed 2026-08-03T01:05:26.197993Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T01:05:26.197993Z digest=sha256:04887dff9bfdbe589c1b6cef71d36dcd603a68878df0d43ef1ffc59784e2b7aa

Observation 3b5cd43f-b90e-4aa1-b8d1-b1246c6f8a49 · outbound

This paper cites Does Writing with Language Models Reduce Content Diversity?.

Quality-constrained Entropy Maximization Policy Optimization for LLM Diversity Does Writing with Language Models Reduce Content Diversity?

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-03T01:05:26.339411Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T01:05:26.339411Z digest=sha256:5df75d4cdff83b107130d3c3cf568a97ae5c7d4f2a26bff5ed85ca6ba3a44fa9

Observation 5860c44c-c0b3-4885-bd08-0a9d9dc6c7d8 · outbound

This paper cites Diverse Preference Optimization.

Quality-constrained Entropy Maximization Policy Optimization for LLM Diversity Diverse Preference Optimization

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-03T01:05:26.232865Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T01:05:26.232865Z digest=sha256:7239d40c61ca7f36f06a87195efafeb549a54a61197f5219eb2e88dbb1616f2a

Observation 2e40642a-8b54-48b6-9df4-79f0281f402b · outbound

This paper cites xverify: Efficient answer verifier for reasoning model evaluations.

Quality-constrained Entropy Maximization Policy Optimization for LLM Diversity xverify: Efficient answer verifier for reasoning model evaluations

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-03T01:05:25.973071Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T01:05:25.973071Z digest=sha256:62af7182c29130fe68b6d9d811d44cfbbbe2e22bd7b35887eab221cf390e28f6

Observation d0d23ebb-b93f-4398-8aea-a9fa06967d02 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Quality-constrained Entropy Maximization Policy Optimization for LLM Diversity Training Verifiers to Solve Math Word Problems

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-03T01:05:26.003252Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T01:05:26.003252Z digest=sha256:752d325ad5640a89ac5d1375e75ce9c4d8d63abcac618ad774d1cbc7d58b7371

Pith citing papers

No inbound Pith citation observations are available.