Pith. sign in

Paper Citation Record · LEDGER

Beyond Negative Rollouts: Positive-Only Policy Optimization with Implicit Negative Gradients

As of 17 August 2026, this Paper Citation Record lists 66 of 66 outbound references and 1 inbound Pith citation observation for arXiv:2605.06650.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.06650 v1

Coverage vector

measured 66 of 66 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-08T09:53:32.077464Z

measured 67 of 67 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-13T02:39:02.891861Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

66 of 66 outbound references displayed

  • verified exact4
  • verified fuzzy26
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch36

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d563f731-ee17-492f-adc3-60d6701fe95d · outbound

This paper cites Attention is All you Need , url =.

Beyond Negative Rollouts: Positive-Only Policy Optimization with Implicit Negative Gradients Attention is All you Need , url =

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T15:42:36.525061Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-08T09:53:32.077464Z digest=sha256:fa69eee79fb6e748592c4831709bf40ef128f32bb6db2ca7edf6eb7817d32a2d

Observation f6da67ff-ca11-40d7-bada-24814a2c984f · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Beyond Negative Rollouts: Positive-Only Policy Optimization with Implicit Negative Gradients DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 2

Resolution
metadata mismatch
local_arxiv, observed 2026-05-11T20:16:08.343881Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-08T09:53:32.077464Z digest=sha256:88d8756d59cc2591053c4a5806927d740feaef8795b5f16a9f33ac47a3f097e3

Observation 86a922d8-98bd-428e-b34f-4ef70dc68b71 · outbound

This paper cites OpenAI o1 System Card.

Beyond Negative Rollouts: Positive-Only Policy Optimization with Implicit Negative Gradients OpenAI o1 System Card

Reference 3

Resolution
metadata mismatch
local_arxiv, observed 2026-05-11T20:16:08.377749Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-08T09:53:32.077464Z digest=sha256:0f2affc610111b7395180f436cc76b235ad1d3605a768bb40dacbe4f8c76fc69

Observation a817d273-cc3e-491a-9561-b70d8be59842 · outbound

This paper cites Advances in neural information processing systems , volume=.

Beyond Negative Rollouts: Positive-Only Policy Optimization with Implicit Negative Gradients Advances in neural information processing systems , volume=

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T15:42:36.512037Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-08T09:53:32.077464Z digest=sha256:c81a0ebdfdbfc57d9bd823e7e85d6e12ac7f3a8e085de63f618348962480110b

Observation 5492156d-34ea-4334-b71e-ce7cf998151a · outbound

This paper cites Advances in neural information processing systems , volume=.

Beyond Negative Rollouts: Positive-Only Policy Optimization with Implicit Negative Gradients Advances in neural information processing systems , volume=

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T15:42:36.521929Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-08T09:53:32.077464Z digest=sha256:e163501ab4510449a0acea65c04437ddc5b951aaa467e56c39949a755f5d50ae

Observation 3c078ff8-b9c2-44a7-9e6f-b381f2dca1bc · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Beyond Negative Rollouts: Positive-Only Policy Optimization with Implicit Negative Gradients Training Verifiers to Solve Math Word Problems

Reference 6

Resolution
metadata mismatch
local_arxiv, observed 2026-05-11T20:16:08.326347Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-08T09:53:32.077464Z digest=sha256:3247746f86dda4f6ebe188a2735d8c6a3d221c7f38a090f352bba57cf54a9c99

Observation 51c6398f-8033-40df-beab-b0ffba85eba8 · outbound

This paper cites Tulu 3: Pushing Frontiers in Open Language Model Post-Training.

Beyond Negative Rollouts: Positive-Only Policy Optimization with Implicit Negative Gradients Tulu 3: Pushing Frontiers in Open Language Model Post-Training

Reference 7

Resolution
metadata mismatch
local_arxiv, observed 2026-05-11T20:16:08.337590Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-08T09:53:32.077464Z digest=sha256:6d82ea7262bffa425a1367a50c6baadf9bfaee00627cd1720b7367de44f07ee7

Observation 593ed0bd-9acb-4c77-ae28-9b6873f7b051 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Beyond Negative Rollouts: Positive-Only Policy Optimization with Implicit Negative Gradients DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 8

Resolution
metadata mismatch
local_arxiv, observed 2026-05-11T20:16:08.292777Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-08T09:53:32.077464Z digest=sha256:1439648cf1a5ccad8bb4267289160b752a357fa7b9f05ce60698a77d92eaab5c

Observation 9248abcc-9976-4994-98f5-0595777b4539 · outbound

This paper cites Advances in neural information processing systems , volume=.

Beyond Negative Rollouts: Positive-Only Policy Optimization with Implicit Negative Gradients Advances in neural information processing systems , volume=

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T15:42:36.473621Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-08T09:53:32.077464Z digest=sha256:935fa8c369a31daf21ed549b55bebca60bdd0aa2a67c33cbf1d498d87aeb1bdb

Observation 56cd5bc5-c0cc-4ea0-9b85-11de1af7e445 · outbound

This paper cites Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?.

Beyond Negative Rollouts: Positive-Only Policy Optimization with Implicit Negative Gradients Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?

Reference 10

Resolution
metadata mismatch
local_arxiv, observed 2026-05-11T20:16:08.358561Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-08T09:53:32.077464Z digest=sha256:35ddede2af774d135771af15568856e624737a20c3a7e732d0f4b8c43c3f6e6f

Observation 4e263ac5-5bef-40c7-ac52-cb7c12d10a3f · outbound

This paper cites Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , pages=.

Beyond Negative Rollouts: Positive-Only Policy Optimization with Implicit Negative Gradients Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , pages=

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T15:42:36.477027Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-08T09:53:32.077464Z digest=sha256:7796a5507ec6223cc7bc829e795add8b0aa93d2d42a0878c65ade9b636255681

Observation 3c9b5981-774e-44ce-a4a0-e7dfe52009e7 · outbound

This paper cites Evaluating Large Language Models Trained on Code.

Beyond Negative Rollouts: Positive-Only Policy Optimization with Implicit Negative Gradients Evaluating Large Language Models Trained on Code

Reference 12

Resolution
metadata mismatch
local_arxiv, observed 2026-05-11T20:16:08.367958Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-08T09:53:32.077464Z digest=sha256:8f83d692c1e1d4addd3bdc0720aa2b8586574d4da92b22bc6ab042c0b946903f

Observation 66142db3-c4e0-40cc-b39b-25ad884102b2 · outbound

This paper cites Advances in neural information processing systems , volume=.

Beyond Negative Rollouts: Positive-Only Policy Optimization with Implicit Negative Gradients Advances in neural information processing systems , volume=

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T15:42:36.545713Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-08T09:53:32.077464Z digest=sha256:e7e56c8c17d6d6a6de5a2cf7cecb27a7912edc55fcd64a91292f0d16d77b3a15

Observation b4b67e7b-5a0f-49b3-b6bc-a54a00f0eadf · outbound

This paper cites Advances in neural information processing systems , volume=.

Beyond Negative Rollouts: Positive-Only Policy Optimization with Implicit Negative Gradients Advances in neural information processing systems , volume=

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T15:42:36.537124Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-08T09:53:32.077464Z digest=sha256:765063ccd517f1ae67c0a34337921697577b2046f1ce5932ef3b8b39a0e11600

Observation 7a98bb31-228e-48bd-9a9f-13665292225b · outbound

This paper cites International Conference on Machine Learning , pages=.

Beyond Negative Rollouts: Positive-Only Policy Optimization with Implicit Negative Gradients International Conference on Machine Learning , pages=

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T15:42:36.470246Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-08T09:53:32.077464Z digest=sha256:32427b611a513425903fb86281c7cb4830413d6ebcce113a8c6ebbb7a1779637

Observation 589a20ab-f1d4-4054-ba32-8f55f961d348 · outbound

This paper cites Advances in neural information processing systems , volume=.

Beyond Negative Rollouts: Positive-Only Policy Optimization with Implicit Negative Gradients Advances in neural information processing systems , volume=

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T15:42:36.463217Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-08T09:53:32.077464Z digest=sha256:786d099e340762b422b4915033ca9794b177461927d7349299278d68e30a26d2

Observation 84b6c866-8975-45e4-ae39-6deeedb4ee1a · outbound

This paper cites Understanding R1-Zero-Like Training: A Critical Perspective.

Beyond Negative Rollouts: Positive-Only Policy Optimization with Implicit Negative Gradients Understanding R1-Zero-Like Training: A Critical Perspective

Reference 17

Resolution
metadata mismatch
local_arxiv, observed 2026-05-11T20:16:08.394362Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-08T09:53:32.077464Z digest=sha256:beba212126c75e8c10989aaa31ed4fb6df3768ceefb263b3404a7509473cbf2d

Observation 32641ecf-e5b2-45a6-b888-b64751652a00 · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

Beyond Negative Rollouts: Positive-Only Policy Optimization with Implicit Negative Gradients DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 18

Resolution
metadata mismatch
local_arxiv, observed 2026-05-11T20:16:08.409100Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-08T09:53:32.077464Z digest=sha256:a95f227e3c422b2768de6fadffe17b32bde1c9235f5e5c70f4108dd5ee9d535f

Observation a8d96b7e-c706-4255-9cc9-3794f20afa88 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Beyond Negative Rollouts: Positive-Only Policy Optimization with Implicit Negative Gradients Proximal Policy Optimization Algorithms

Reference 19

Resolution
metadata mismatch
local_arxiv, observed 2026-05-11T20:16:08.390441Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-08T09:53:32.077464Z digest=sha256:d7b6f0b445fef32b0d52eb42903c295fc68e6050a30be8a8855856f9f49d42f9

Observation 5fb22f02-0e7e-4308-b751-47922b45e78d · outbound

This paper cites Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=.

Beyond Negative Rollouts: Positive-Only Policy Optimization with Implicit Negative Gradients Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T15:42:36.528130Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-08T09:53:32.077464Z digest=sha256:66033e5fbac04c01bc0140d297495d6b5302f7c91a0e5b30f373c65df8e4a104

Observation 87db5e5c-c036-49a1-b4cb-b4620fc3318b · outbound

This paper cites International conference on machine learning , pages=.

Beyond Negative Rollouts: Positive-Only Policy Optimization with Implicit Negative Gradients International conference on machine learning , pages=

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T15:42:36.495872Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-08T09:53:32.077464Z digest=sha256:17462cdcde3fb1dcf07fe0017bc0e48eb403cb64cf26e18650b66fc2dcc2c986

Observation a2223a68-6a78-4a49-91f7-af564b58edda · outbound

This paper cites Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , pages=.

Beyond Negative Rollouts: Positive-Only Policy Optimization with Implicit Negative Gradients Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , pages=

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T15:42:36.499682Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-08T09:53:32.077464Z digest=sha256:1cdf3580aad9b01c429c6df2c80b2f47f95f846582074b55beae55abfac68d94

Observation 47d0cb85-b182-4ba5-9d33-3254a44b3b6f · outbound

This paper cites Scaling Relationship on Learning Mathematical Reasoning with Large Language Models.

Beyond Negative Rollouts: Positive-Only Policy Optimization with Implicit Negative Gradients Scaling Relationship on Learning Mathematical Reasoning with Large Language Models

Reference 23

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T00:22:11.212173Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-08T09:53:32.077464Z digest=sha256:4a8b84ae4bf0dbbe8170591eb8661887ae8d642df341bc763273ed72cfaada43

Observation 3f7e34ca-08e9-4dd5-99db-6a47b3e6b156 · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

Beyond Negative Rollouts: Positive-Only Policy Optimization with Implicit Negative Gradients Advances in Neural Information Processing Systems , volume=

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T15:42:36.518390Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-08T09:53:32.077464Z digest=sha256:f416b24d82b1171699dc53fb96a535b61b6ae126320b94ed1b0a2dc04ddba0c6

Observation 359e0d48-73cc-4d37-a616-09c67fa5c344 · outbound

This paper cites Statistical Rejection Sampling Improves Preference Optimization.

Beyond Negative Rollouts: Positive-Only Policy Optimization with Implicit Negative Gradients Statistical Rejection Sampling Improves Preference Optimization

Reference 25

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T20:16:08.547916Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-08T09:53:32.077464Z digest=sha256:0887a90261fb7d6ae4e64501cf8cd30e8aa55d7a553b84d08af525cb521beab9

Observation 005b8e84-1296-4887-8e71-5e4308ac6468 · outbound

This paper cites Notion Blog , volume=.

Beyond Negative Rollouts: Positive-Only Policy Optimization with Implicit Negative Gradients Notion Blog , volume=

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T15:42:36.505941Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-08T09:53:32.077464Z digest=sha256:ee5c09ddbea93b21e2b0e8f67d13c35b27394e70ceef85c30ee0bbd41a120042

Observation 925289f0-5c84-4967-971c-26aa06694475 · outbound

This paper cites Hugging Face repository , volume=.

Beyond Negative Rollouts: Positive-Only Policy Optimization with Implicit Negative Gradients Hugging Face repository , volume=

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T15:42:36.540079Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-08T09:53:32.077464Z digest=sha256:ae53d74914430644d257c74aee04b186f5112b9e00a406671243105eb31a7947

Observation 001a5864-92dc-4b4d-99d8-687f67bf220f · outbound

This paper cites Measuring Mathematical Problem Solving With the MATH Dataset.

Beyond Negative Rollouts: Positive-Only Policy Optimization with Implicit Negative Gradients Measuring Mathematical Problem Solving With the MATH Dataset

Reference 28

Resolution
metadata mismatch
local_arxiv, observed 2026-05-11T20:16:08.469346Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-08T09:53:32.077464Z digest=sha256:14168f589dffe6cd0143838e2036657974eee765e3f4c3b9b2d8dccdb954b9cf

Observation 74c4afe3-03b1-4f85-805c-f622f8a712bc · outbound

This paper cites The twelfth international conference on learning representations , year=.

Beyond Negative Rollouts: Positive-Only Policy Optimization with Implicit Negative Gradients The twelfth international conference on learning representations , year=

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T15:42:36.515443Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-08T09:53:32.077464Z digest=sha256:89ebe5bbeecae0bdc637e9842ef7c4273291e1a9b0df803ae796042da69bf31e

Observation 4c85c11c-2945-4877-8a5d-87cd7376c18a · outbound

This paper cites Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=.

Beyond Negative Rollouts: Positive-Only Policy Optimization with Implicit Negative Gradients Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T15:42:36.492776Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-08T09:53:32.077464Z digest=sha256:482db1a79bd38423bb76b04c0a0ec143d0f8e1b81a05f362b20a3251787ff619

Observation aaab3ac3-7fec-4260-bf1b-0153868e32bf · outbound

This paper cites Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement.

Beyond Negative Rollouts: Positive-Only Policy Optimization with Implicit Negative Gradients Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-05-11T20:16:08.658561Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-08T09:53:32.077464Z digest=sha256:6959ae15d7b9168978d5207cd6d54c9e1de519466c36d554de10d7638965cad5

Observation 72eb94af-b7cb-4b09-96c8-d3787db1b213 · outbound

This paper cites The Llama 3 Herd of Models.

Beyond Negative Rollouts: Positive-Only Policy Optimization with Implicit Negative Gradients The Llama 3 Herd of Models

Reference 32

Resolution
metadata mismatch
local_arxiv, observed 2026-05-11T20:16:08.386791Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-08T09:53:32.077464Z digest=sha256:fa8d5dcabc83a701f0689a07de7e2845f957492f3d7ce2c7779ce4d38bf8189e

Observation 1ef36382-3470-4d22-8529-abf271353c6a · outbound

This paper cites Self-Consistency Improves Chain of Thought Reasoning in Language Models.

Beyond Negative Rollouts: Positive-Only Policy Optimization with Implicit Negative Gradients Self-Consistency Improves Chain of Thought Reasoning in Language Models

Reference 33

Resolution
metadata mismatch
local_arxiv, observed 2026-05-11T20:16:08.592423Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-08T09:53:32.077464Z digest=sha256:6429bda8845a4650ac0de808ebda4fcb3c9ab76735e20637d3499337b476e6f8

Observation 02dbe6b1-bf0a-4e23-8cbe-b3c469cd26f7 · outbound

This paper cites Advances in neural information processing systems , volume=.

Beyond Negative Rollouts: Positive-Only Policy Optimization with Implicit Negative Gradients Advances in neural information processing systems , volume=

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T15:42:36.466920Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-08T09:53:32.077464Z digest=sha256:38d2934348d61e7eab4c9ed2d869dc3e72c284b0fb14ee79f811fa39e448c2e9

Observation 8d04ccf6-3a44-4a86-997b-fb3b3144116a · outbound

This paper cites Proceedings of the 29th symposium on operating systems principles , pages=.

Beyond Negative Rollouts: Positive-Only Policy Optimization with Implicit Negative Gradients Proceedings of the 29th symposium on operating systems principles , pages=

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T15:42:36.503118Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-08T09:53:32.077464Z digest=sha256:f7f45d94f59d5ab052a6aa721fa083c0a9f4cffa001073b6e62b7e392916d0ff

Observation 0edb578f-77fe-4ded-8c68-bfd697ea2419 · outbound

This paper cites Biometrics bulletin , volume=.

Beyond Negative Rollouts: Positive-Only Policy Optimization with Implicit Negative Gradients Biometrics bulletin , volume=

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T15:42:36.542909Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-08T09:53:32.077464Z digest=sha256:373a2418ae7f8b91075b108a2d022315d5c92791cb0b195e1fb19e934466ceaa

Observation 8733d11d-aab6-4e6b-b01c-faa41c26832b · outbound

This paper cites Rewarding Progress: Scaling Automated Process Verifiers for LLM Reasoning.

Beyond Negative Rollouts: Positive-Only Policy Optimization with Implicit Negative Gradients Rewarding Progress: Scaling Automated Process Verifiers for LLM Reasoning

Reference 37

Resolution
metadata mismatch
arxiv_id, observed 2026-05-21T01:42:19.263401Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-08T09:53:32.077464Z digest=sha256:d709e3c45f82facabc77af25d83604c89bf592840fa040c0f03d15120f7fa397

Observation b79fdfe5-4ca3-4782-8ffa-ed070427ebe8 · outbound

This paper cites Process Reinforcement through Implicit Rewards.

Beyond Negative Rollouts: Positive-Only Policy Optimization with Implicit Negative Gradients Process Reinforcement through Implicit Rewards

Reference 38

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T20:23:31.624298Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-08T09:53:32.077464Z digest=sha256:120c4b46b7f3fb652335142cef530b5aa259f9600779c233306a180b9cfd5166

Observation 925e9021-0001-4d49-9fa3-ae8c68a979ab · outbound

This paper cites Stable and efficient single-rollout rl for multimodal reasoning.

Beyond Negative Rollouts: Positive-Only Policy Optimization with Implicit Negative Gradients Stable and efficient single-rollout rl for multimodal reasoning

Reference 39

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T20:16:08.484993Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-08T09:53:32.077464Z digest=sha256:ba15889d29de6b3099e845141556252c7e19a5ac4cab0bfffb89ba1f689a36ec

Observation 35cd58fa-c523-430e-b049-6bb58de43edf · outbound

This paper cites arXiv preprint arXiv:2602.20722 (2026) 3.

Beyond Negative Rollouts: Positive-Only Policy Optimization with Implicit Negative Gradients arXiv preprint arXiv:2602.20722 (2026) 3

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-11T20:16:08.382545Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-08T09:53:32.077464Z digest=sha256:cc6d8c38cb1793934ac8a400878160037ba8270599a445a92ae319b6be7a4c27

Observation d30608fa-cf65-4065-8f18-5ec835e56ba9 · outbound

This paper cites Group Sequence Policy Optimization.

Beyond Negative Rollouts: Positive-Only Policy Optimization with Implicit Negative Gradients Group Sequence Policy Optimization

Reference 41

Resolution
metadata mismatch
local_arxiv, observed 2026-05-11T20:16:08.574025Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-08T09:53:32.077464Z digest=sha256:c144961c61b79450a5b11f90474451d1a9ccff3cd25cc0e3bd9408f492045514

Observation 3674fc9c-03e3-47ad-b29c-9a3b3fee9d56 · outbound

This paper cites Gvpo: Group variance policy optimization for large language model post-training.

Beyond Negative Rollouts: Positive-Only Policy Optimization with Implicit Negative Gradients Gvpo: Group variance policy optimization for large language model post-training

Reference 42

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T20:16:08.498951Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-08T09:53:32.077464Z digest=sha256:35818473816192df43ea0fd9d2933ec16533f1a2375d67e4a655d242100db44a

Observation f8c02819-3dea-4cb8-95f6-d7e26226d67f · outbound

This paper cites Soft Adaptive Policy Optimization.

Beyond Negative Rollouts: Positive-Only Policy Optimization with Implicit Negative Gradients Soft Adaptive Policy Optimization

Reference 43

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T07:14:32.897338Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-08T09:53:32.077464Z digest=sha256:79966be6b218cb33bfb36d390acc2e93820bb97cc72a274265ec916966c7c021

Observation 547e9ef7-d3e0-4e37-800d-3d477049b579 · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

Beyond Negative Rollouts: Positive-Only Policy Optimization with Implicit Negative Gradients Advances in Neural Information Processing Systems , volume=

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T15:42:36.487879Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-08T09:53:32.077464Z digest=sha256:962a196bbe690999a885e653b043585cb885ad5043a3b15b4762b503dc216147

Observation 7bb4fe60-8670-4ed8-a45c-b20413df8fea · outbound

This paper cites MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts.

Beyond Negative Rollouts: Positive-Only Policy Optimization with Implicit Negative Gradients MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts

Reference 45

Resolution
metadata mismatch
local_arxiv, observed 2026-05-11T20:16:08.579199Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-08T09:53:32.077464Z digest=sha256:e2849927dbd8264153c120ebe76401bede76abef76c13cdaad018156c037848e

Observation cb57407c-d189-4e39-ba58-fc9475416c43 · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

Beyond Negative Rollouts: Positive-Only Policy Optimization with Implicit Negative Gradients Advances in Neural Information Processing Systems , volume=

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T15:42:36.534326Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-08T09:53:32.077464Z digest=sha256:c9ddaf3310318d03b0bb431c359619ef2337cee38b61325d5863d67a5f027bb6

Observation 61cfe08d-8965-4df7-86c5-18d40af3610e · outbound

This paper cites VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks.

Beyond Negative Rollouts: Positive-Only Policy Optimization with Implicit Negative Gradients VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks

Reference 47

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T09:36:04.819624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-08T09:53:32.077464Z digest=sha256:a784ee59dceca9938c7357495ee5c7ce4f6efd4bac2877676fba909303a7daf8

Observation d87f61ed-4dd0-4d57-a869-e9d4312d91d4 · outbound

This paper cites MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention.

Beyond Negative Rollouts: Positive-Only Policy Optimization with Implicit Negative Gradients MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention

Reference 48

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T09:28:16.586976Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-08T09:53:32.077464Z digest=sha256:b026d60c73d7e378a4b11508e1c8c1dd2895de78af4f74fc7cd196683145606e

Observation 314fdb32-8343-4a63-9f61-48325ea1d40e · outbound

This paper cites GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization.

Beyond Negative Rollouts: Positive-Only Policy Optimization with Implicit Negative Gradients GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization

Reference 49

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T05:31:56.190613Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-08T09:53:32.077464Z digest=sha256:12fce44227db7ca482dbdf84466cad070ec23f631a1b883db5193b12336433cf

Observation 290dd0d2-ab60-4a10-8da6-fc6d1dc2e3f9 · outbound

This paper cites Bapo: Stabilizing off-policy reinforcement learning for llms via balanced policy optimization with adaptive clipping.arXiv preprint arXiv:2510.18927.

Beyond Negative Rollouts: Positive-Only Policy Optimization with Implicit Negative Gradients Bapo: Stabilizing off-policy reinforcement learning for llms via balanced policy optimization with adaptive clipping.arXiv preprint arXiv:2510.18927

Reference 50

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T20:16:08.632113Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-08T09:53:32.077464Z digest=sha256:4cb4986c785300a5736df21baa0db032555eb5dd166e09dab352f44d1de018d2

Observation 48864e1d-7e00-47a9-847b-98d88e99a8ec · outbound

This paper cites Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing , pages=.

Beyond Negative Rollouts: Positive-Only Policy Optimization with Implicit Negative Gradients Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing , pages=

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T15:42:36.480467Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-08T09:53:32.077464Z digest=sha256:ab273e8937f7742d1540b81e2014cac5fc2d0a381e5c6c2f59a92a550305fefe

Observation fdfdd776-05fc-4677-8410-196eee7bdd3d · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

Beyond Negative Rollouts: Positive-Only Policy Optimization with Implicit Negative Gradients Advances in Neural Information Processing Systems , volume=

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T15:42:36.530885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-08T09:53:32.077464Z digest=sha256:4f6cf719cf53627279ff337d7aee9ec86bf2e9f7a1459a97f7d6ab8db45cadef

Observation c510675d-28a7-42dd-86b6-88d128064fba · outbound

This paper cites Zephyr: Direct Distillation of LM Alignment.

Beyond Negative Rollouts: Positive-Only Policy Optimization with Implicit Negative Gradients Zephyr: Direct Distillation of LM Alignment

Reference 53

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T10:13:57.632488Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-08T09:53:32.077464Z digest=sha256:3c0974bb94acbb67127e7646c4e6393234239076c60b2bae2a4b649b64dc11e0

Observation b8c432ad-a1c4-4228-8f84-5e9a210351c7 · outbound

This paper cites Beyond Human Data: Scaling Self-Training for Problem-Solving with Language Models.

Beyond Negative Rollouts: Positive-Only Policy Optimization with Implicit Negative Gradients Beyond Human Data: Scaling Self-Training for Problem-Solving with Language Models

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-05-11T20:16:08.428351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-08T09:53:32.077464Z digest=sha256:697c15c38341adb765b6d16a47d94dda1e4facc0aac7dfc912d7b09c138f87da

Observation 879b653f-a7ec-4d1d-96a0-6cccdbc684cc · outbound

This paper cites A Model Can Help Itself: Reward-Free Self-Training for LLM Reasoning.

Beyond Negative Rollouts: Positive-Only Policy Optimization with Implicit Negative Gradients A Model Can Help Itself: Reward-Free Self-Training for LLM Reasoning

Reference 55

Resolution
metadata mismatch
local_arxiv, observed 2026-05-11T20:16:08.418279Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-08T09:53:32.077464Z digest=sha256:5732364d455531b826bd40ee4fed713ff046410d653d590628f82762fd0e0026

Observation ba939c32-d000-47c9-ac83-58fb26b54e2f · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

Beyond Negative Rollouts: Positive-Only Policy Optimization with Implicit Negative Gradients Advances in Neural Information Processing Systems , volume=

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T15:42:36.483862Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-08T09:53:32.077464Z digest=sha256:f6006a2527f54e4a8f8926194067fab7216fa8e156d88b3f49400cf66dcb5ff5

Observation 5bfd387e-6cbb-4c70-a9bc-ce94c929ffbe · outbound

This paper cites The surprising effectiveness of negative reinforcement in llm reasoning.arXiv preprint arXiv:2506.01347.

Beyond Negative Rollouts: Positive-Only Policy Optimization with Implicit Negative Gradients The surprising effectiveness of negative reinforcement in llm reasoning.arXiv preprint arXiv:2506.01347

Reference 57

Resolution
verified exact
arxiv_id, observed 2026-05-11T20:16:08.531413Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-08T09:53:32.077464Z digest=sha256:2e2d2dee7ac913492258bcbe6fe566546921121d3b1249cc5693ffc0321faefe

Observation 7228c3f8-c4ae-447c-b9cb-859a9fdd0000 · outbound

This paper cites Reinforcement Learning for Reasoning in Large Language Models with One Training Example.

Beyond Negative Rollouts: Positive-Only Policy Optimization with Implicit Negative Gradients Reinforcement Learning for Reasoning in Large Language Models with One Training Example

Reference 58

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T19:51:05.238288Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-08T09:53:32.077464Z digest=sha256:60c2a0c9fd46a0380723f19f96ae80823fc2aee68c44a3763580e1d7862e4c67

Observation 23087d96-c9eb-4e1a-9177-39a534c8d45b · outbound

This paper cites Inftythink: Breaking the length limits of long-context reasoning in large language models.

Beyond Negative Rollouts: Positive-Only Policy Optimization with Implicit Negative Gradients Inftythink: Breaking the length limits of long-context reasoning in large language models

Reference 59

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T20:16:08.637931Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-08T09:53:32.077464Z digest=sha256:96b679ff8c8566834f3bd91e17d672e5fb11c09731f351ac62f5e105d99be83a

Observation e379a87a-3ac0-4490-ae20-ea4c30bdc293 · outbound

This paper cites Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters.

Beyond Negative Rollouts: Positive-Only Policy Optimization with Implicit Negative Gradients Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters

Reference 60

Resolution
metadata mismatch
local_arxiv, observed 2026-05-11T20:16:08.611703Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-08T09:53:32.077464Z digest=sha256:21f28a51385176ea1fa2566dddf7bd6d144f43a2c9f13c88cd7778f7a275c43d

Observation a41851d8-b134-4df4-8ec1-74d0c55a1f1b · outbound

This paper cites Learning beyond Teacher: Generalized On-Policy Distillation with Reward Extrapolation.

Beyond Negative Rollouts: Positive-Only Policy Optimization with Implicit Negative Gradients Learning beyond Teacher: Generalized On-Policy Distillation with Reward Extrapolation

Reference 61

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T05:11:07.754859Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-08T09:53:32.077464Z digest=sha256:1b2b746d762cd6fe8eb735768e56529fa7841630be5813c14b850a13439d62a4

Observation dbfb5099-ca7e-4e0a-9359-074699d578bc · outbound

This paper cites Clipo: Contrastive learning in policy optimization generalizes rlvr.

Beyond Negative Rollouts: Positive-Only Policy Optimization with Implicit Negative Gradients Clipo: Contrastive learning in policy optimization generalizes rlvr

Reference 62

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T20:16:08.537817Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-08T09:53:32.077464Z digest=sha256:be274f0ecc979d27cc11872708488d286a75aacd9910070b3ea87e492d0131d4

Observation 0745f993-d5b4-46db-bfb2-d6db30c84ef4 · outbound

This paper cites IEEE transactions on knowledge and data engineering , volume=.

Beyond Negative Rollouts: Positive-Only Policy Optimization with Implicit Negative Gradients IEEE transactions on knowledge and data engineering , volume=

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-05-26T15:42:36.509031Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-08T09:53:32.077464Z digest=sha256:27ae91c1d0fe805989d44238c731b3f21913932e6a2ca6510bd9f2cfc547d00e

Observation f1261673-5610-4d57-9cf7-0bf72264cf1e · outbound

This paper cites MathArena: Evaluating LLMs on Uncontaminated Math Competitions.

Beyond Negative Rollouts: Positive-Only Policy Optimization with Implicit Negative Gradients MathArena: Evaluating LLMs on Uncontaminated Math Competitions

Reference 64

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T00:10:14.938194Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-08T09:53:32.077464Z digest=sha256:505ac7dd44a64cffebf9524c79412cfadb8a7d6dbb96e5ebd5c27f3b33277d90

Observation 21ff4e1c-e768-4139-b5d8-986c02e676a9 · outbound

This paper cites Correcting the Mythos of KL-Regularization: Direct Alignment without Overoptimization via Chi-Squared Preference Optimization.

Beyond Negative Rollouts: Positive-Only Policy Optimization with Implicit Negative Gradients Correcting the Mythos of KL-Regularization: Direct Alignment without Overoptimization via Chi-Squared Preference Optimization

Reference 65

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T20:16:08.493642Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-08T09:53:32.077464Z digest=sha256:fde2abf5ed56aa68799fa154fc27a4c9a41d463eae50c07a0609866e4efd2358

Observation 3fdc76ff-37bc-4b45-8de3-fc3023dbb733 · outbound

This paper cites DeepSeek-Coder: When the Large Language Model Meets Programming -- The Rise of Code Intelligence.

Beyond Negative Rollouts: Positive-Only Policy Optimization with Implicit Negative Gradients DeepSeek-Coder: When the Large Language Model Meets Programming -- The Rise of Code Intelligence

Reference 66

Resolution
metadata mismatch
local_arxiv, observed 2026-05-11T20:16:08.553647Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-08T09:53:32.077464Z digest=sha256:dedad00926abc39b95ae3e44a862713a8afda9146ad8ebe050f416febebe35cd

Pith citing papers

Observation a4480743-638b-4363-9214-561a3a1f912c · inbound

Multimodal Reward Hacking in Reinforcement Learning cites this paper.

Multimodal Reward Hacking in Reinforcement Learning Beyond Negative Rollouts: Positive-Only Policy Optimization with Implicit Negative Gradients

Reference 33

Resolution
unresolved
no resolver link, observed 2026-07-13T02:39:02.891861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T02:39:02.891861Z digest=sha256:4140cf7b46356d67d83bb2fbb8c3a543410c0d2039eddea709cec3b507ff169d