Pith. sign in

Paper Citation Record · LEDGER

On Advantage Estimates for Max@K Policy Gradients

As of 14 August 2026, this Paper Citation Record lists 75 of 75 outbound references and 3 inbound Pith citation observations for arXiv:2606.06080.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2606.06080 v1

Coverage vector

measured 75 of 75 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-06-28T02:21:57.143016Z

measured 78 of 78 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T00:49:54.921367Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

75 of 75 outbound references displayed

  • verified exact41
  • verified fuzzy0
  • unresolved30
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch3

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e9d87534-3b86-49b4-9d36-52120868d367 · outbound

This paper cites Finite-time analysis of the multiarmed bandit problem.Machine learning, 47(2):235–256, 2002.

On Advantage Estimates for Max@K Policy Gradients Finite-time analysis of the multiarmed bandit problem.Machine learning, 47(2):235–256, 2002

Reference 1

Resolution
unresolved
no resolver link, observed 2026-06-28T02:21:57.143016Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T02:21:57.143016Z digest=sha256:1203a53afcacda1e280a62fe9642a6ab78a1609f0ae5588c3f1f7c9fc4ebe801

Observation 543eebbd-3c98-410a-ac7e-bfae8c3bea4e · outbound

This paper cites The best of n worlds: Aligning reinforcement learning with best-of-n sampling via max@ k optimisation.

On Advantage Estimates for Max@K Policy Gradients The best of n worlds: Aligning reinforcement learning with best-of-n sampling via max@ k optimisation

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-07-02T12:06:56.456387Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-28T02:21:57.143016Z digest=sha256:d4dec25059de29e7920b4d3708dfb473e5f7a2447d457acd51c8412c651c3d22

Observation 7e4c3140-62f9-4b08-b63f-b3e808a90c04 · outbound

This paper cites Online Preference Alignment for Language Models via Count-based Exploration.

On Advantage Estimates for Max@K Policy Gradients Online Preference Alignment for Language Models via Count-based Exploration

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-07-02T12:06:56.459065Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-28T02:21:57.143016Z digest=sha256:8bb4163df6028e6dd777896b14b0d1e9fcbbb3f4862dd5f9b22426ac492f3e3e

Observation f0ec27f7-9f27-4d97-b816-ae168c473926 · outbound

This paper cites Post-training as reweighting: A stochastic view of reasoning trajectories in language models.

On Advantage Estimates for Max@K Policy Gradients Post-training as reweighting: A stochastic view of reasoning trajectories in language models

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T12:16:56.591850Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-28T02:21:57.143016Z digest=sha256:fdd9cbd2d92acd5314246e62c3deefd2f0ca20b9d91f24a9e645550a2aaafe39

Observation 2d3a9e82-3707-4987-b5c4-6cf763edcbcd · outbound

This paper cites Exploration by Random Network Distillation.

On Advantage Estimates for Max@K Policy Gradients Exploration by Random Network Distillation

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-07-02T12:06:56.450761Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-28T02:21:57.143016Z digest=sha256:2faba3d3d9f2799a3711a918fae91ce35b1010a6b368930d92bfe84f042fe790

Observation 35de7610-04bb-4d37-98cb-04f19d3a0d84 · outbound

This paper cites arXiv preprint arXiv:2510.15020 , year=.

On Advantage Estimates for Max@K Policy Gradients arXiv preprint arXiv:2510.15020 , year=

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T12:06:56.385892Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-28T02:21:57.143016Z digest=sha256:22e9500ed93f3d90d036c74681a7e60ed85fc0001bbc56a41d870b9ae3ad6b61

Observation 619fa3af-3bf9-43b9-b1cd-5cbf31564b8a · outbound

This paper cites Evaluating Large Language Models Trained on Code.

On Advantage Estimates for Max@K Policy Gradients Evaluating Large Language Models Trained on Code

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-07-02T12:16:56.588901Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-28T02:21:57.143016Z digest=sha256:63541baf84f59d08c2ac33a1852cc5f23658ebf2a705e735da45ab3d3972303f

Observation b02f5bbd-79a9-4cf7-8857-b3f6bccb225e · outbound

This paper cites Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models.

On Advantage Estimates for Max@K Policy Gradients Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-07-02T12:06:56.383237Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-28T02:21:57.143016Z digest=sha256:7d1b65d48879a08337f47a5a91bf0e86381cde6effecc645c7913da979c359ca

Observation eb6b85f0-3116-4476-ae8d-a83788160ada · outbound

This paper cites Reasoning with exploration: An entropy perspective.

On Advantage Estimates for Max@K Policy Gradients Reasoning with exploration: An entropy perspective

Reference 9

Resolution
verified exact
doi, observed 2026-06-28T02:31:30.232165Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-28T02:21:57.143016Z digest=sha256:4d7a0721e635c08e1fc2bf78e8da0980be8b3d87c20ed61f0eab2224481c976b

Observation 02b542aa-6423-49ae-b28c-45706bdeecf5 · outbound

This paper cites Deep reinforcement learning from human preferences.

On Advantage Estimates for Max@K Policy Gradients Deep reinforcement learning from human preferences

Reference 10

Resolution
unresolved
no resolver link, observed 2026-06-28T02:21:57.143016Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T02:21:57.143016Z digest=sha256:18c351765a368517cf831a0dc1d6c286b9842a887b3a2400df79981143390cb8

Observation 42035913-4b5e-4044-92b9-ac0aa32f408b · outbound

This paper cites Beyond variance reduction: Understanding the true impact of baselines on policy optimization.

On Advantage Estimates for Max@K Policy Gradients Beyond variance reduction: Understanding the true impact of baselines on policy optimization

Reference 11

Resolution
unresolved
no resolver link, observed 2026-06-28T02:21:57.143016Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T02:21:57.143016Z digest=sha256:3f6017c1bf546be76ad12a62dd91cb5737e3c89dece60edd6d6c8e31e512c92c

Observation 245a53bb-ddc2-425b-b236-3491c3f5ee93 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

On Advantage Estimates for Max@K Policy Gradients Training Verifiers to Solve Math Word Problems

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-07-02T12:06:56.425930Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-28T02:21:57.143016Z digest=sha256:ce18ce8782172b9a06c0181a98117b07e381c4262256f3a3b94539360d854a10

Observation 6b214efe-25b4-4ae3-833c-fbbe217175a7 · outbound

This paper cites The Entropy Mechanism of Reinforcement Learning for Reasoning Language Models.

On Advantage Estimates for Max@K Policy Gradients The Entropy Mechanism of Reinforcement Learning for Reasoning Language Models

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-07-02T12:16:56.585783Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-28T02:21:57.143016Z digest=sha256:fe9e619cda31c8974cf21f4ce0b9c99bcd23f570772f320a5207be79357ad00b

Observation 49c68fd0-b753-4d53-8b62-5e7babe84c66 · outbound

This paper cites Weight ensembling improves reasoning in language models.

On Advantage Estimates for Max@K Policy Gradients Weight ensembling improves reasoning in language models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-06-28T02:21:57.143016Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T02:21:57.143016Z digest=sha256:ea9305ae0a5ec4d702a46eb83b0dbacd6b830935109951e4c3b9e3ea99989dc9

Observation 3c8b4829-1361-4583-beb0-91682e013424 · outbound

This paper cites arXiv preprint arXiv:2505.17621 , year=.

On Advantage Estimates for Max@K Policy Gradients arXiv preprint arXiv:2505.17621 , year=

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-07-02T12:06:56.461761Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-28T02:21:57.143016Z digest=sha256:6981bac8a568f74f33baeef2120b0ae7a3ae392dc827b4be73236d9ae777f40f

Observation fd41b64d-5172-4cca-b828-31ffe3f6ec93 · outbound

This paper cites The Llama 3 Herd of Models.

On Advantage Estimates for Max@K Policy Gradients The Llama 3 Herd of Models

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-07-02T12:06:56.443870Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-28T02:21:57.143016Z digest=sha256:96eb6d5a6e682ed8588b726e6822145919688294de7d52ce20c21b06454cff53

Observation fdccaf20-1b01-4608-ae6f-ea37d8c6fe86 · outbound

This paper cites Variance reduction techniques for gradient estimates in reinforcement learning.Journal of Machine Learning Research, 5(Nov): 1471–1530, 2004.

On Advantage Estimates for Max@K Policy Gradients Variance reduction techniques for gradient estimates in reinforcement learning.Journal of Machine Learning Research, 5(Nov): 1471–1530, 2004

Reference 17

Resolution
unresolved
no resolver link, observed 2026-06-28T02:21:57.143016Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T02:21:57.143016Z digest=sha256:d03e0937b4059d1bf9cf5779836abb1b8e344e082b24712ac144253f498d03d6

Observation ea11b413-dd9c-4a29-b2db-4a27d8b283c9 · outbound

This paper cites MuProp: Unbiased Backpropagation for Stochastic Neural Networks.

On Advantage Estimates for Max@K Policy Gradients MuProp: Unbiased Backpropagation for Stochastic Neural Networks

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-07-02T12:06:56.448633Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-28T02:21:57.143016Z digest=sha256:9d6378cfdddafd301eea7279b115c0e8548c1c87dbf977fd964d48c6de8793eb

Observation 30aa64ea-53b9-42ce-bbe1-0fb09374700d · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

On Advantage Estimates for Max@K Policy Gradients DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-07-02T12:06:56.426380Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-28T02:21:57.143016Z digest=sha256:b579929441382abf8d850932b0ffa694023d056cf8fd0b9e8e02ed505c152c57

Observation 83954b2f-cd48-4257-adb2-0de7f8459dc4 · outbound

This paper cites Rewarding the Unlikely: Lifting GRPO Beyond Distribution Sharpening.

On Advantage Estimates for Max@K Policy Gradients Rewarding the Unlikely: Lifting GRPO Beyond Distribution Sharpening

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-07-02T12:06:56.428956Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-28T02:21:57.143016Z digest=sha256:b9c1a0cea5b5181fd1aada7f75baf36aa9f0da68d2cd7e5120411869d270e39c

Observation 8a335aa8-db5f-4693-82ea-ad6329f24100 · outbound

This paper cites Measuring Mathematical Problem Solving With the MATH Dataset.

On Advantage Estimates for Max@K Policy Gradients Measuring Mathematical Problem Solving With the MATH Dataset

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-07-02T12:06:56.433525Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-28T02:21:57.143016Z digest=sha256:f31c57282ec07fb1a5530693455b6cdb7372af1fd6a117f5256d8e7891605666

Observation c51556c9-30e1-4ed6-bf1e-22f81ba35f0b · outbound

This paper cites A class of statistics with asymptotically normal distribution.

On Advantage Estimates for Max@K Policy Gradients A class of statistics with asymptotically normal distribution

Reference 22

Resolution
unresolved
no resolver link, observed 2026-06-28T02:21:57.143016Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T02:21:57.143016Z digest=sha256:c1512c363d27c4a07107670611d94c3aeac428ff86b9a22e6ea975b17eda57bb

Observation 5aca0c07-0b64-4efe-a780-00e35b22a508 · outbound

This paper cites Emergent Slow Thinking in LLMs as Inverse Tree Freezing.

On Advantage Estimates for Max@K Policy Gradients Emergent Slow Thinking in LLMs as Inverse Tree Freezing

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-07-02T12:06:56.448464Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-28T02:21:57.143016Z digest=sha256:3e23530ff8cb717c97a1085ad27a6d9e6879aaeda59153a1acdd221715aff043

Observation 64be959f-a0ba-4b75-b248-ea764785db81 · outbound

This paper cites arXiv preprint arXiv:2509.25133 , year=.

On Advantage Estimates for Max@K Policy Gradients arXiv preprint arXiv:2509.25133 , year=

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-07-02T12:06:56.421477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-28T02:21:57.143016Z digest=sha256:2dadd8eb8759f6544500a68ba2db1f9655e24168cf0c766d482cf141685d42ad

Observation 5bab3412-5500-4e49-a263-ee8e2e57402a · outbound

This paper cites Adam: A Method for Stochastic Optimization.

On Advantage Estimates for Max@K Policy Gradients Adam: A Method for Stochastic Optimization

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-07-02T12:06:56.431368Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-28T02:21:57.143016Z digest=sha256:3aae046c47c1a183bdf536c72dc8e8585ce874f8956d11c1ef05ac2239da4b9e

Observation f8241fdf-c3de-44fd-9bfd-1bdff93c5bd8 · outbound

This paper cites Emergence of exploration in policy gradient reinforcement learning via resetting, 2023.

On Advantage Estimates for Max@K Policy Gradients Emergence of exploration in policy gradient reinforcement learning via resetting, 2023

Reference 26

Resolution
unresolved
no resolver link, observed 2026-06-28T02:21:57.143016Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T02:21:57.143016Z digest=sha256:cf0166d498f258b414a435ff123470c5aadcf7fa3230be78340bbb0005cf38b6

Observation f7c6c964-53e6-483f-96f9-7b9a743c4b83 · outbound

This paper cites Hwang, Jiangjiang Yang, Ronan Le Bras, Oyvind Tafjord, Christopher Wilhelm, Luca Soldaini, Noah A.

On Advantage Estimates for Max@K Policy Gradients Hwang, Jiangjiang Yang, Ronan Le Bras, Oyvind Tafjord, Christopher Wilhelm, Luca Soldaini, Noah A

Reference 27

Resolution
unresolved
no resolver link, observed 2026-06-28T02:21:57.143016Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T02:21:57.143016Z digest=sha256:0ff7f731e47e478bd1d2a9b6a1dc16842d5cda1941dd99f1f9ac23595d58118e

Observation 0a9baef8-00ab-4443-8d09-a9c68b841d2c · outbound

This paper cites Solving quantitative reasoning problems with language models.

On Advantage Estimates for Max@K Policy Gradients Solving quantitative reasoning problems with language models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-06-28T02:21:57.143016Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T02:21:57.143016Z digest=sha256:8eb50189f47e8c282e8564cfe523a344c3df5d8ea86d5fed6e49675896148c2e

Observation 566591fd-f894-43e7-9ac0-cc2a9b78f151 · outbound

This paper cites Jointly Reinforcing Diversity and Quality in Language Model Generations.

On Advantage Estimates for Max@K Policy Gradients Jointly Reinforcing Diversity and Quality in Language Model Generations

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-07-02T12:06:56.451211Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-28T02:21:57.143016Z digest=sha256:1bd87ae36ace8c95fd944d89ec74935db8a19195d259f9abd527ddeff09df419

Observation dee092f5-63ba-44e7-b934-48b10a46b4f9 · outbound

This paper cites Can llms guide their own exploration? gradient-guided reinforcement learning for llm reasoning.

On Advantage Estimates for Max@K Policy Gradients Can llms guide their own exploration? gradient-guided reinforcement learning for llm reasoning

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-07-02T12:06:56.453670Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-28T02:21:57.143016Z digest=sha256:9b43c4b2704984cb590119f987348311be99de8e3e1a783b786747962c92354c

Observation 70c5232d-eba7-4788-9178-9635a7407d29 · outbound

This paper cites Understanding r1-zero-like training: A critical perspective.

On Advantage Estimates for Max@K Policy Gradients Understanding r1-zero-like training: A critical perspective

Reference 31

Resolution
unresolved
no resolver link, observed 2026-06-28T02:21:57.143016Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T02:21:57.143016Z digest=sha256:3d11cb43471c2c27e7763fbb62d803f343072ad3481322274499a24838e29a58

Observation 740d59f0-ac71-408a-b8e8-e6c15fab080c · outbound

This paper cites RL squeezes, SFT expands: A comparative study of reasoning LLMs.

On Advantage Estimates for Max@K Policy Gradients RL squeezes, SFT expands: A comparative study of reasoning LLMs

Reference 32

Resolution
unresolved
no resolver link, observed 2026-06-28T02:21:57.143016Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T02:21:57.143016Z digest=sha256:8305ca92394f25d46b3eac0adb5a4063452b58dcb3a29fcca1c727e85e0bde57

Observation ac8ec50a-cae2-4cbe-b7dd-50064f44a405 · outbound

This paper cites The role of baselines in policy gradient optimization.Advances in Neural Information Processing Systems, 35:17818–17830, 2022.

On Advantage Estimates for Max@K Policy Gradients The role of baselines in policy gradient optimization.Advances in Neural Information Processing Systems, 35:17818–17830, 2022

Reference 33

Resolution
unresolved
no resolver link, observed 2026-06-28T02:21:57.143016Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T02:21:57.143016Z digest=sha256:f73ea713b7e5fc6f48a81f3fb87546ae64db753d644a8b63652bf6ada2686f72

Observation fdb5b8d2-32fc-4a8e-a7eb-a6c7ad61f758 · outbound

This paper cites Variational inference for monte carlo objectives.

On Advantage Estimates for Max@K Policy Gradients Variational inference for monte carlo objectives

Reference 34

Resolution
unresolved
no resolver link, observed 2026-06-28T02:21:57.143016Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T02:21:57.143016Z digest=sha256:ce896a75702a6ee40d4c8ca568c19e7d52b3e33b2bcc41373249b5b20078ae1b

Observation a487ccb8-4605-4075-bec3-882270cf3202 · outbound

This paper cites Asynchronous methods for deep reinforce- ment learning.

On Advantage Estimates for Max@K Policy Gradients Asynchronous methods for deep reinforce- ment learning

Reference 35

Resolution
unresolved
no resolver link, observed 2026-06-28T02:21:57.143016Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T02:21:57.143016Z digest=sha256:0fdc4e8f38cb8272dc9ae4eb44963f5f507ed3674301e21e5fac29d09ba252e0

Observation c3ff4086-40b9-4c50-8d35-d67f261fd71f · outbound

This paper cites Emergence of exploration in policy gradient reinforcement learning via retrying.

On Advantage Estimates for Max@K Policy Gradients Emergence of exploration in policy gradient reinforcement learning via retrying

Reference 36

Resolution
unresolved
no resolver link, observed 2026-06-28T02:21:57.143016Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T02:21:57.143016Z digest=sha256:5b41f564feeee3e973cf505c54cef033f2831bf13d719a98d85ae0c16a998b4f

Observation aaa432fd-26ee-46e7-90a0-daaa390eefc8 · outbound

This paper cites OpenAI o1 System Card.

On Advantage Estimates for Max@K Policy Gradients OpenAI o1 System Card

Reference 37

Resolution
verified exact
local_arxiv, observed 2026-07-02T12:06:56.440986Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-28T02:21:57.143016Z digest=sha256:d91fad81d743d1a49cfe4c30f6d3ac4f99412767ec3abe75fc70655aa16c00f9

Observation 068f44fa-383d-40bf-abad-a46d3f9217a8 · outbound

This paper cites Total stochastic gradient algorithms and applications in reinforcement learning.

On Advantage Estimates for Max@K Policy Gradients Total stochastic gradient algorithms and applications in reinforcement learning

Reference 38

Resolution
unresolved
no resolver link, observed 2026-06-28T02:21:57.143016Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T02:21:57.143016Z digest=sha256:b38e7904bbc84038a8c19e838f8e59584c285aff817436971096f10957b74614

Observation 7cc8d450-5098-41bb-ba1a-81b35467b508 · outbound

This paper cites A unified view of likelihood ratio and reparameterization gradients.

On Advantage Estimates for Max@K Policy Gradients A unified view of likelihood ratio and reparameterization gradients

Reference 39

Resolution
unresolved
no resolver link, observed 2026-06-28T02:21:57.143016Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T02:21:57.143016Z digest=sha256:c4acb5c4bc96ece470e298264d2f81e1f60210797a2be130f59d1ca8bce85601

Observation 5f33d517-d2b8-42e5-bed0-3e293bcb355f · outbound

This paper cites PIPPS: Flexible model- based policy search robust to the curse of chaos.

On Advantage Estimates for Max@K Policy Gradients PIPPS: Flexible model- based policy search robust to the curse of chaos

Reference 40

Resolution
unresolved
no resolver link, observed 2026-06-28T02:21:57.143016Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T02:21:57.143016Z digest=sha256:2cef84041f0f2109cd53c5cd6a685aa54423c13bb952be5f9e3ee19d130a9504

Observation ccc58204-e901-4294-bf34-6fc3d3bb8c40 · outbound

This paper cites Beyond the Sampled Token: Preserving Candidate Support in RLVR.

On Advantage Estimates for Max@K Policy Gradients Beyond the Sampled Token: Preserving Candidate Support in RLVR

Reference 41

Resolution
verified exact
local_arxiv, observed 2026-07-02T12:06:56.431196Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-28T02:21:57.143016Z digest=sha256:8e439a37be3fb684628381b71707cfbc4f8b2c9c4654dadb0141462547cd157a

Observation 80b04d8c-869a-437f-896e-e1f698062db6 · outbound

This paper cites Reinforcement learning of motor skills with policy gradients.

On Advantage Estimates for Max@K Policy Gradients Reinforcement learning of motor skills with policy gradients

Reference 42

Resolution
unresolved
no resolver link, observed 2026-06-28T02:21:57.143016Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T02:21:57.143016Z digest=sha256:3b95e263b5c92f50eb56145ac43e3608a26ae480d481e17157d00cdf921bf9ba

Observation feda5041-2619-462f-bc64-efacbe01732e · outbound

This paper cites Proximal Policy Optimization Algorithms.

On Advantage Estimates for Max@K Policy Gradients Proximal Policy Optimization Algorithms

Reference 43

Resolution
verified exact
local_arxiv, observed 2026-07-02T12:06:56.406133Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-28T02:21:57.143016Z digest=sha256:fd21ece09b0719a8cc386e9c8dd8463ddd9855f77f53ecef9ac93e86953f847a

Observation 978bf51c-3054-476a-bfff-6e3746dd43dd · outbound

This paper cites e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs.

On Advantage Estimates for Max@K Policy Gradients e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-07-02T12:06:56.433620Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-28T02:21:57.143016Z digest=sha256:f7cd146cd81834ae94f39f95187ec421f2591db021999972c3f6225de7d130a0

Observation 283fed06-b34b-4ef7-9aa5-28b5431b5e8f · outbound

This paper cites Rethinking Reflection in Pre-Training.

On Advantage Estimates for Max@K Policy Gradients Rethinking Reflection in Pre-Training

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-07-02T12:06:56.418746Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-28T02:21:57.143016Z digest=sha256:5aa80a0fd1b02a80f10c41985e105c7114f1f438c26c6766da3368cc9f7c5f4e

Observation 9c98f150-8c58-49ba-a853-dd94deba054c · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

On Advantage Estimates for Max@K Policy Gradients DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 46

Resolution
verified exact
local_arxiv, observed 2026-07-02T12:06:56.436184Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-28T02:21:57.143016Z digest=sha256:388bd9dcd19eb6d04d277fbe91d40a75e6d1bc162477f6a82cea7ecfc2f68d22

Observation a6ffc73c-cf59-4104-8cc2-e6ef512812c5 · outbound

This paper cites On entropy control in LLM-RL algorithms.

On Advantage Estimates for Max@K Policy Gradients On entropy control in LLM-RL algorithms

Reference 47

Resolution
unresolved
no resolver link, observed 2026-06-28T02:21:57.143016Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T02:21:57.143016Z digest=sha256:3c569b159d2be8079a5c4f87878f9c1cb65caac2a643d55cdf54931e947fb0a4

Observation 5f9f6c87-a403-485b-b3d9-44f4ae530620 · outbound

This paper cites HybridFlow: A Flexible and Efficient RLHF Framework.

On Advantage Estimates for Max@K Policy Gradients HybridFlow: A Flexible and Efficient RLHF Framework

Reference 48

Resolution
verified exact
local_arxiv, observed 2026-07-02T12:06:56.443553Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-28T02:21:57.143016Z digest=sha256:b05d7f016f36c4b79633cc68051a08f068f79bedf55cbae81c1cd6bc84a597fe

Observation 04a2bc2c-32e2-40d3-9bb4-51b9022ca6e9 · outbound

This paper cites Outcome-based Exploration for LLM Reasoning.

On Advantage Estimates for Max@K Policy Gradients Outcome-based Exploration for LLM Reasoning

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-07-02T12:06:56.413899Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-28T02:21:57.143016Z digest=sha256:1cc59f0362aeb3c30fd882f76537259ad2e9cd9a460eb75e4ae0d3a02d3b4713

Observation f3ea7472-d129-4b72-9155-eedbf8a6b1eb · outbound

This paper cites Kakade, Dean Foster, and Udaya Ghai.

On Advantage Estimates for Max@K Policy Gradients Kakade, Dean Foster, and Udaya Ghai

Reference 50

Resolution
unresolved
no resolver link, observed 2026-06-28T02:21:57.143016Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T02:21:57.143016Z digest=sha256:b9881ea957e1c83debf03ae1e70606ced623ad77b8c3841a053b8a7ab0413fe6

Observation 56ae5afe-71fd-4cb8-be8d-debda028fceb · outbound

This paper cites Optimizing language models for inference time objectives using reinforcement learning.

On Advantage Estimates for Max@K Policy Gradients Optimizing language models for inference time objectives using reinforcement learning

Reference 51

Resolution
unresolved
no resolver link, observed 2026-06-28T02:21:57.143016Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T02:21:57.143016Z digest=sha256:fdf050fc655bee847cc7f56a1e8c33134776a4fdbae74c6ddf24c14d0edaa5a2

Observation c9bdda4e-318e-41e5-a871-8b21caaf1d0a · outbound

This paper cites Rebar: Low-variance, unbiased gradient estimates for discrete latent variable models.Advances in Neural Information Processing Systems, 30, 2017.

On Advantage Estimates for Max@K Policy Gradients Rebar: Low-variance, unbiased gradient estimates for discrete latent variable models.Advances in Neural Information Processing Systems, 30, 2017

Reference 52

Resolution
unresolved
no resolver link, observed 2026-06-28T02:21:57.143016Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T02:21:57.143016Z digest=sha256:aa32dd2ca6afb9e149da15da961e096a9e197d7faaf4a3f92533f8dd19768ecb

Observation a4c882a5-141e-4c30-8a1d-7f1111eae931 · outbound

This paper cites Representation-Based Exploration for Language Models: From Test-Time to Post-Training.

On Advantage Estimates for Max@K Policy Gradients Representation-Based Exploration for Language Models: From Test-Time to Post-Training

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-07-16T02:22:28.823656Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-28T02:21:57.143016Z digest=sha256:e7d11a9d7a40119f8acac4584237c4e59e39b1ab51121779ea6e69d44cc2b090

Observation 04c33fe3-4eb1-414f-b0aa-153b880ba2e8 · outbound

This paper cites Pass@K policy optimization: Solving harder reinforcement learning problems.Advances in Neural Information Processing Systems, 38: 152416–152445, 2025.

On Advantage Estimates for Max@K Policy Gradients Pass@K policy optimization: Solving harder reinforcement learning problems.Advances in Neural Information Processing Systems, 38: 152416–152445, 2025

Reference 54

Resolution
unresolved
no resolver link, observed 2026-06-28T02:21:57.143016Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T02:21:57.143016Z digest=sha256:4d7220e3cd3308290d295ced4a6e683aa2efe81d933edf7a7264986acfc96abf

Observation 669a75df-30e2-4d7a-831d-0b8a5667faa5 · outbound

This paper cites OctoThinker: Mid-training Incentivizes Reinforcement Learning Scaling.

On Advantage Estimates for Max@K Policy Gradients OctoThinker: Mid-training Incentivizes Reinforcement Learning Scaling

Reference 55

Resolution
verified exact
arxiv_id, observed 2026-07-02T12:06:56.418404Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-28T02:21:57.143016Z digest=sha256:46fff2afada0a66ee16d42ac7a120781b2e62958bda24091b0fe411f79834936

Observation 6f6f9517-3815-4785-b5b3-3c9765a0d875 · outbound

This paper cites The Optimal Reward Baseline for Gradient-Based Reinforcement Learning.

On Advantage Estimates for Max@K Policy Gradients The Optimal Reward Baseline for Gradient-Based Reinforcement Learning

Reference 56

Resolution
verified exact
local_arxiv, observed 2026-07-02T12:06:56.445997Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-28T02:21:57.143016Z digest=sha256:43d022953e2ad8676a81cecca9237b115fb105356622f4cbd369844fabb837c8

Observation 195d58b3-f851-4598-8de3-4768f6b75741 · outbound

This paper cites Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs.

On Advantage Estimates for Max@K Policy Gradients Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMs

Reference 57

Resolution
verified exact
local_arxiv, observed 2026-07-02T12:06:56.454152Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-28T02:21:57.143016Z digest=sha256:778d9be70670e25bc0a34e8fcc721250951725779dc696c080b0d44614bdce91

Observation 59aaf811-b77f-4c93-81dc-59f67d2d8af0 · outbound

This paper cites Simple statistical gradient-following algorithms for connectionist reinforce- ment learning.Machine Learning, 8(3):229–256, May 1992.

On Advantage Estimates for Max@K Policy Gradients Simple statistical gradient-following algorithms for connectionist reinforce- ment learning.Machine Learning, 8(3):229–256, May 1992

Reference 58

Resolution
unresolved
no resolver link, observed 2026-06-28T02:21:57.143016Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T02:21:57.143016Z digest=sha256:36539d02127e42e7ee15dcd7f9b8b92f8e9253388239d5c1fea199b1157bd862

Observation 8792cc16-08d9-48c6-9f54-e21cc4766a3a · outbound

This paper cites Variance Reduction for Policy Gradient with Action-Dependent Factorized Baselines.

On Advantage Estimates for Max@K Policy Gradients Variance Reduction for Policy Gradient with Action-Dependent Factorized Baselines

Reference 59

Resolution
verified exact
local_arxiv, observed 2026-07-02T12:06:56.459203Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-28T02:21:57.143016Z digest=sha256:c3f8f8b723fb6a78caba92c3b688085c825c484c9469347036f0f3b5f0dd86a3

Observation 2e4e9f0b-bf9f-4747-b751-ac5e31ba3107 · outbound

This paper cites The invisible leash: Why rlvr may or may not escape its origin.

On Advantage Estimates for Max@K Policy Gradients The invisible leash: Why rlvr may or may not escape its origin

Reference 60

Resolution
verified exact
arxiv_id, observed 2026-07-02T12:06:56.438710Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-28T02:21:57.143016Z digest=sha256:322db033f050c0249bc069d792e2235b8a31692386abafee1c9ec4b8a93522b4

Observation b092c4ec-69cc-4936-a8eb-095eaa57e1d3 · outbound

This paper cites Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement.

On Advantage Estimates for Max@K Policy Gradients Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement

Reference 61

Resolution
verified exact
local_arxiv, observed 2026-07-02T12:06:56.438752Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-28T02:21:57.143016Z digest=sha256:9711f59b719097493458b74d16a0577e6ec9aa5ee189e10f9893adae758c5083

Observation 14dd87cb-18e1-40c3-ae63-0b100ed34c89 · outbound

This paper cites Qwen3 Technical Report.

On Advantage Estimates for Max@K Policy Gradients Qwen3 Technical Report

Reference 62

Resolution
verified exact
local_arxiv, observed 2026-07-02T12:06:56.456829Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-28T02:21:57.143016Z digest=sha256:26318212babdee6eb1d1ce2e97ba39e2237402d1c3c3a53a8824f3f8d5ab5436

Observation 51bec505-12cd-4214-814a-f74b1d3f61ef · outbound

This paper cites Qwen2.5 Technical Report.

On Advantage Estimates for Max@K Policy Gradients Qwen2.5 Technical Report

Reference 63

Resolution
verified exact
local_arxiv, observed 2026-07-02T12:06:56.446201Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-28T02:21:57.143016Z digest=sha256:1df9c0c83b36af55f90576eacd97a5de75f0320ac5bb820cc8929cb7219c674a

Observation d725b3bd-2cd2-4efe-beba-0877aaf7be3e · outbound

This paper cites arXiv preprint arXiv:2510.02172 , year=.

On Advantage Estimates for Max@K Policy Gradients arXiv preprint arXiv:2510.02172 , year=

Reference 64

Resolution
verified exact
arxiv_id, observed 2026-07-02T12:06:56.403210Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-28T02:21:57.143016Z digest=sha256:e72eb125cc59477838c5550e6a702fc5d725957f32ac8a8fa2e37271e1d22ce4

Observation 5f93d10b-abed-4841-b427-b5c443caddea · outbound

This paper cites Does reinforcement learning really incentivize reasoning capacity in LLMs beyond the base model? InThe Thirty-ninth Annual Conference on Neural Information Processing Systems, 2025.

On Advantage Estimates for Max@K Policy Gradients Does reinforcement learning really incentivize reasoning capacity in LLMs beyond the base model? InThe Thirty-ninth Annual Conference on Neural Information Processing Systems, 2025

Reference 65

Resolution
unresolved
no resolver link, observed 2026-06-28T02:21:57.143016Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T02:21:57.143016Z digest=sha256:e239df7506f55e7057d621036b782ae4966fafa468b4833d93094ceb82a7e730

Observation 826dfcea-4e34-4dc2-9e96-5a9b3c583ec9 · outbound

This paper cites On the interplay of pre-training, mid-training, and rl on reasoning language models, 2025 a.

On Advantage Estimates for Max@K Policy Gradients On the interplay of pre-training, mid-training, and rl on reasoning language models, 2025 a

Reference 66

Resolution
verified exact
arxiv_id, observed 2026-07-02T12:06:56.411254Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-28T02:21:57.143016Z digest=sha256:efa47b7c6f4c7c9e422c5688baf448ff864837cd2b7becb35132c3a73b1a7db8

Observation b436eb3d-c72a-4ac2-9093-dd0f0dd13dfb · outbound

This paper cites arXiv preprint arXiv:2509.25810 , year=.

On Advantage Estimates for Max@K Policy Gradients arXiv preprint arXiv:2509.25810 , year=

Reference 67

Resolution
verified exact
arxiv_id, observed 2026-07-02T12:06:56.413232Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-28T02:21:57.143016Z digest=sha256:b536b236a3945902d6a64e0353826b95e9d728313f6f015639a29f539defb36c

Observation 066f4021-50c9-41a4-b696-2c794db103c3 · outbound

This paper cites Echo chamber: Rl post-training amplifies behaviors learned in pretraining.

On Advantage Estimates for Max@K Policy Gradients Echo chamber: Rl post-training amplifies behaviors learned in pretraining

Reference 68

Resolution
unresolved
no resolver link, observed 2026-06-28T02:21:57.143016Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T02:21:57.143016Z digest=sha256:b75ff3ac6d00049a002172461cde91e21c84dce0c8ae902b9bf76638e4a6e29a

Observation 4c242cd4-ec7c-43f0-901d-3edb7e28c52e · outbound

This paper cites Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts.

On Advantage Estimates for Max@K Policy Gradients Act Only When It Pays: Efficient Reinforcement Learning for LLM Reasoning via Selective Rollouts

Reference 69

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T12:06:56.403920Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-28T02:21:57.143016Z digest=sha256:e309450f5bba3cf9ddbc4639a8203a182e605f89b297ed579e05347ef2c9ca25

Observation 024e228f-8c3a-40b3-93a7-ae7df8bbde22 · outbound

This paper cites First Return, Entropy-Eliciting Explore.

On Advantage Estimates for Max@K Policy Gradients First Return, Entropy-Eliciting Explore

Reference 70

Resolution
verified exact
arxiv_id, observed 2026-07-02T12:06:56.398549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-28T02:21:57.143016Z digest=sha256:5d2ac277fb355736bc3f445d872406b061ec523ae6a00145289a4269c09f35a5

Observation cc922107-1b70-4dd9-9ba4-9d7318dd7fe2 · outbound

This paper cites arXiv preprint arXiv:2509.15194 , year =.

On Advantage Estimates for Max@K Policy Gradients arXiv preprint arXiv:2509.15194 , year =

Reference 71

Resolution
verified exact
arxiv_id, observed 2026-07-02T12:06:56.400294Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-28T02:21:57.143016Z digest=sha256:235a30d18d1660d2d665258183629de69bad7b5c38d72762322ebc3060505b78

Observation 183e932e-84d4-4260-9a22-8f1071cb96e1 · outbound

This paper cites [29] employed a semantic diversity score with an external semantic comparator, and Tuyls et al.

On Advantage Estimates for Max@K Policy Gradients [29] employed a semantic diversity score with an external semantic comparator, and Tuyls et al

Reference 72

Resolution
unresolved
no resolver link, observed 2026-06-28T02:21:57.143016Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T02:21:57.143016Z digest=sha256:2658db37ec584b74a4d794eb3aee535cf3a4eafc715c7144b60d7fee1f2f0571

Observation f90062a8-11fd-4ccc-a88a-7c00d637077d · outbound

This paper cites Setlur et al.

On Advantage Estimates for Max@K Policy Gradients Setlur et al

Reference 73

Resolution
unresolved
no resolver link, observed 2026-06-28T02:21:57.143016Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T02:21:57.143016Z digest=sha256:a8d3792fdc61ea5f61eb0acd21c8be17e50755d454eb245238d7f4357bd69c4d

Observation 855ae7c3-8db9-4d06-9bfe-5b17296af384 · outbound

This paper cites " " Com pu te s bi no mi al c o e f f i c i e n t C (n , k ) in log - space.

On Advantage Estimates for Max@K Policy Gradients " " Com pu te s bi no mi al c o e f f i c i e n t C (n , k ) in log - space

Reference 74

Resolution
unresolved
no resolver link, observed 2026-06-28T02:21:57.143016Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T02:21:57.143016Z digest=sha256:fe23a6236bb1e9a7dd82dcab93042252147f81327b899a2831d2e821f2f89ccf

Observation 8fc7501e-7b95-47e5-986c-f736fdc0740b · outbound

This paper cites BoN mean.

On Advantage Estimates for Max@K Policy Gradients BoN mean

Reference 75

Resolution
malformed identifier
no resolver link, observed 2026-06-28T02:21:57.143016Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-28T02:21:57.143016Z digest=sha256:17d6d859eadae237612c0a50358304780a59494c3caf732b3273127031aa4e22

Pith citing papers

Observation 2ab39a61-7fdc-44b9-ba0c-266f239da232 · inbound

Rank-Conditioned Sample Reuse for the Plackett--Luce Best-of-$K$ Objective cites this paper.

Rank-Conditioned Sample Reuse for the Plackett--Luce Best-of-$K$ Objective On Advantage Estimates for Max@K Policy Gradients

Reference 6

Resolution
unresolved
no resolver link, observed 2026-07-14T06:44:16.198117Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T06:44:16.198117Z digest=sha256:4b80ba70e3ca93fe773b52c6f73f3b169e1c054c61a042181a3c92ef2cf357ea

Observation d3a013a1-75d5-4f88-a672-5f78cae46762 · inbound

Multi-Axis Max@K Reinforcement Learning for Representative Diversity in Text-to-Image Generation cites this paper.

Multi-Axis Max@K Reinforcement Learning for Representative Diversity in Text-to-Image Generation On Advantage Estimates for Max@K Policy Gradients

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-02T00:41:12.582293Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T00:41:12.582293Z digest=sha256:796fc9fa94a9d33f472c8c9483895b37de521f3aadbb538b12d35cf2a429a110

Observation 006768af-557e-43b9-92c8-62ce1a4519a5 · inbound

Stabilized Best-of-$K$ Training for Neural Combinatorial Optimization cites this paper.

Stabilized Best-of-$K$ Training for Neural Combinatorial Optimization On Advantage Estimates for Max@K Policy Gradients

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-04T00:49:54.921367Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T00:49:54.921367Z digest=sha256:c423fd4ad4c743173fa1626784ae6b1c8420815f3c1cf29fa92d4c50a6534611