Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-08T09:53:32.077464Z
Paper Citation Record · LEDGER
As of 17 August 2026, this Paper Citation Record lists 66 of 66 outbound references and 1 inbound Pith citation observation for arXiv:2605.06650.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-08T09:53:32.077464Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-07-13T02:39:02.891861Z
A source-named dated measurement, never combined with another source.
Source: cited_works
66 of 66 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation d563f731-ee17-492f-adc3-60d6701fe95d · outbound
Beyond Negative Rollouts: Positive-Only Policy Optimization with Implicit Negative Gradients Attention is All you Need , url =
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation f6da67ff-ca11-40d7-bada-24814a2c984f · outbound
Beyond Negative Rollouts: Positive-Only Policy Optimization with Implicit Negative Gradients DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 86a922d8-98bd-428e-b34f-4ef70dc68b71 · outbound
Beyond Negative Rollouts: Positive-Only Policy Optimization with Implicit Negative Gradients OpenAI o1 System Card
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation a817d273-cc3e-491a-9561-b70d8be59842 · outbound
Beyond Negative Rollouts: Positive-Only Policy Optimization with Implicit Negative Gradients Advances in neural information processing systems , volume=
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 5492156d-34ea-4334-b71e-ce7cf998151a · outbound
Beyond Negative Rollouts: Positive-Only Policy Optimization with Implicit Negative Gradients Advances in neural information processing systems , volume=
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 3c078ff8-b9c2-44a7-9e6f-b381f2dca1bc · outbound
Beyond Negative Rollouts: Positive-Only Policy Optimization with Implicit Negative Gradients Training Verifiers to Solve Math Word Problems
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 51c6398f-8033-40df-beab-b0ffba85eba8 · outbound
Beyond Negative Rollouts: Positive-Only Policy Optimization with Implicit Negative Gradients Tulu 3: Pushing Frontiers in Open Language Model Post-Training
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 593ed0bd-9acb-4c77-ae28-9b6873f7b051 · outbound
Beyond Negative Rollouts: Positive-Only Policy Optimization with Implicit Negative Gradients DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 9248abcc-9976-4994-98f5-0595777b4539 · outbound
Beyond Negative Rollouts: Positive-Only Policy Optimization with Implicit Negative Gradients Advances in neural information processing systems , volume=
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 56cd5bc5-c0cc-4ea0-9b85-11de1af7e445 · outbound
Beyond Negative Rollouts: Positive-Only Policy Optimization with Implicit Negative Gradients Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 4e263ac5-5bef-40c7-ac52-cb7c12d10a3f · outbound
Beyond Negative Rollouts: Positive-Only Policy Optimization with Implicit Negative Gradients Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , pages=
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 3c9b5981-774e-44ce-a4a0-e7dfe52009e7 · outbound
Beyond Negative Rollouts: Positive-Only Policy Optimization with Implicit Negative Gradients Evaluating Large Language Models Trained on Code
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 66142db3-c4e0-40cc-b39b-25ad884102b2 · outbound
Beyond Negative Rollouts: Positive-Only Policy Optimization with Implicit Negative Gradients Advances in neural information processing systems , volume=
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation b4b67e7b-5a0f-49b3-b6bc-a54a00f0eadf · outbound
Beyond Negative Rollouts: Positive-Only Policy Optimization with Implicit Negative Gradients Advances in neural information processing systems , volume=
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 7a98bb31-228e-48bd-9a9f-13665292225b · outbound
Beyond Negative Rollouts: Positive-Only Policy Optimization with Implicit Negative Gradients International Conference on Machine Learning , pages=
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 589a20ab-f1d4-4054-ba32-8f55f961d348 · outbound
Beyond Negative Rollouts: Positive-Only Policy Optimization with Implicit Negative Gradients Advances in neural information processing systems , volume=
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 84b6c866-8975-45e4-ae39-6deeedb4ee1a · outbound
Beyond Negative Rollouts: Positive-Only Policy Optimization with Implicit Negative Gradients Understanding R1-Zero-Like Training: A Critical Perspective
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 32641ecf-e5b2-45a6-b888-b64751652a00 · outbound
Beyond Negative Rollouts: Positive-Only Policy Optimization with Implicit Negative Gradients DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation a8d96b7e-c706-4255-9cc9-3794f20afa88 · outbound
Beyond Negative Rollouts: Positive-Only Policy Optimization with Implicit Negative Gradients Proximal Policy Optimization Algorithms
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 5fb22f02-0e7e-4308-b751-47922b45e78d · outbound
Beyond Negative Rollouts: Positive-Only Policy Optimization with Implicit Negative Gradients Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 87db5e5c-c036-49a1-b4cb-b4620fc3318b · outbound
Beyond Negative Rollouts: Positive-Only Policy Optimization with Implicit Negative Gradients International conference on machine learning , pages=
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation a2223a68-6a78-4a49-91f7-af564b58edda · outbound
Beyond Negative Rollouts: Positive-Only Policy Optimization with Implicit Negative Gradients Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , pages=
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 47d0cb85-b182-4ba5-9d33-3254a44b3b6f · outbound
Beyond Negative Rollouts: Positive-Only Policy Optimization with Implicit Negative Gradients Scaling Relationship on Learning Mathematical Reasoning with Large Language Models
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 3f7e34ca-08e9-4dd5-99db-6a47b3e6b156 · outbound
Beyond Negative Rollouts: Positive-Only Policy Optimization with Implicit Negative Gradients Advances in Neural Information Processing Systems , volume=
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 359e0d48-73cc-4d37-a616-09c67fa5c344 · outbound
Beyond Negative Rollouts: Positive-Only Policy Optimization with Implicit Negative Gradients Statistical Rejection Sampling Improves Preference Optimization
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 005b8e84-1296-4887-8e71-5e4308ac6468 · outbound
Beyond Negative Rollouts: Positive-Only Policy Optimization with Implicit Negative Gradients Notion Blog , volume=
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 925289f0-5c84-4967-971c-26aa06694475 · outbound
Beyond Negative Rollouts: Positive-Only Policy Optimization with Implicit Negative Gradients Hugging Face repository , volume=
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 001a5864-92dc-4b4d-99d8-687f67bf220f · outbound
Beyond Negative Rollouts: Positive-Only Policy Optimization with Implicit Negative Gradients Measuring Mathematical Problem Solving With the MATH Dataset
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 74c4afe3-03b1-4f85-805c-f622f8a712bc · outbound
Beyond Negative Rollouts: Positive-Only Policy Optimization with Implicit Negative Gradients The twelfth international conference on learning representations , year=
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 4c85c11c-2945-4877-8a5d-87cd7376c18a · outbound
Beyond Negative Rollouts: Positive-Only Policy Optimization with Implicit Negative Gradients Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation aaab3ac3-7fec-4260-bf1b-0153868e32bf · outbound
Beyond Negative Rollouts: Positive-Only Policy Optimization with Implicit Negative Gradients Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 72eb94af-b7cb-4b09-96c8-d3787db1b213 · outbound
Beyond Negative Rollouts: Positive-Only Policy Optimization with Implicit Negative Gradients The Llama 3 Herd of Models
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 1ef36382-3470-4d22-8529-abf271353c6a · outbound
Beyond Negative Rollouts: Positive-Only Policy Optimization with Implicit Negative Gradients Self-Consistency Improves Chain of Thought Reasoning in Language Models
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 02dbe6b1-bf0a-4e23-8cbe-b3c469cd26f7 · outbound
Beyond Negative Rollouts: Positive-Only Policy Optimization with Implicit Negative Gradients Advances in neural information processing systems , volume=
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 8d04ccf6-3a44-4a86-997b-fb3b3144116a · outbound
Beyond Negative Rollouts: Positive-Only Policy Optimization with Implicit Negative Gradients Proceedings of the 29th symposium on operating systems principles , pages=
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 0edb578f-77fe-4ded-8c68-bfd697ea2419 · outbound
Beyond Negative Rollouts: Positive-Only Policy Optimization with Implicit Negative Gradients Biometrics bulletin , volume=
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 8733d11d-aab6-4e6b-b01c-faa41c26832b · outbound
Beyond Negative Rollouts: Positive-Only Policy Optimization with Implicit Negative Gradients Rewarding Progress: Scaling Automated Process Verifiers for LLM Reasoning
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation b79fdfe5-4ca3-4782-8ffa-ed070427ebe8 · outbound
Beyond Negative Rollouts: Positive-Only Policy Optimization with Implicit Negative Gradients Process Reinforcement through Implicit Rewards
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 925e9021-0001-4d49-9fa3-ae8c68a979ab · outbound
Beyond Negative Rollouts: Positive-Only Policy Optimization with Implicit Negative Gradients Stable and efficient single-rollout rl for multimodal reasoning
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 35cd58fa-c523-430e-b049-6bb58de43edf · outbound
Beyond Negative Rollouts: Positive-Only Policy Optimization with Implicit Negative Gradients arXiv preprint arXiv:2602.20722 (2026) 3
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation d30608fa-cf65-4065-8f18-5ec835e56ba9 · outbound
Beyond Negative Rollouts: Positive-Only Policy Optimization with Implicit Negative Gradients Group Sequence Policy Optimization
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 3674fc9c-03e3-47ad-b29c-9a3b3fee9d56 · outbound
Beyond Negative Rollouts: Positive-Only Policy Optimization with Implicit Negative Gradients Gvpo: Group variance policy optimization for large language model post-training
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation f8c02819-3dea-4cb8-95f6-d7e26226d67f · outbound
Beyond Negative Rollouts: Positive-Only Policy Optimization with Implicit Negative Gradients Soft Adaptive Policy Optimization
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 547e9ef7-d3e0-4e37-800d-3d477049b579 · outbound
Beyond Negative Rollouts: Positive-Only Policy Optimization with Implicit Negative Gradients Advances in Neural Information Processing Systems , volume=
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 7bb4fe60-8670-4ed8-a45c-b20413df8fea · outbound
Beyond Negative Rollouts: Positive-Only Policy Optimization with Implicit Negative Gradients MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation cb57407c-d189-4e39-ba58-fc9475416c43 · outbound
Beyond Negative Rollouts: Positive-Only Policy Optimization with Implicit Negative Gradients Advances in Neural Information Processing Systems , volume=
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 61cfe08d-8965-4df7-86c5-18d40af3610e · outbound
Beyond Negative Rollouts: Positive-Only Policy Optimization with Implicit Negative Gradients VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation d87f61ed-4dd0-4d57-a869-e9d4312d91d4 · outbound
Beyond Negative Rollouts: Positive-Only Policy Optimization with Implicit Negative Gradients MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 314fdb32-8343-4a63-9f61-48325ea1d40e · outbound
Beyond Negative Rollouts: Positive-Only Policy Optimization with Implicit Negative Gradients GDPO: Group reward-Decoupled Normalization Policy Optimization for Multi-reward RL Optimization
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 290dd0d2-ab60-4a10-8da6-fc6d1dc2e3f9 · outbound
Beyond Negative Rollouts: Positive-Only Policy Optimization with Implicit Negative Gradients Bapo: Stabilizing off-policy reinforcement learning for llms via balanced policy optimization with adaptive clipping.arXiv preprint arXiv:2510.18927
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 48864e1d-7e00-47a9-847b-98d88e99a8ec · outbound
Beyond Negative Rollouts: Positive-Only Policy Optimization with Implicit Negative Gradients Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing , pages=
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation fdfdd776-05fc-4677-8410-196eee7bdd3d · outbound
Beyond Negative Rollouts: Positive-Only Policy Optimization with Implicit Negative Gradients Advances in Neural Information Processing Systems , volume=
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation c510675d-28a7-42dd-86b6-88d128064fba · outbound
Beyond Negative Rollouts: Positive-Only Policy Optimization with Implicit Negative Gradients Zephyr: Direct Distillation of LM Alignment
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation b8c432ad-a1c4-4228-8f84-5e9a210351c7 · outbound
Beyond Negative Rollouts: Positive-Only Policy Optimization with Implicit Negative Gradients Beyond Human Data: Scaling Self-Training for Problem-Solving with Language Models
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 879b653f-a7ec-4d1d-96a0-6cccdbc684cc · outbound
Beyond Negative Rollouts: Positive-Only Policy Optimization with Implicit Negative Gradients A Model Can Help Itself: Reward-Free Self-Training for LLM Reasoning
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation ba939c32-d000-47c9-ac83-58fb26b54e2f · outbound
Beyond Negative Rollouts: Positive-Only Policy Optimization with Implicit Negative Gradients Advances in Neural Information Processing Systems , volume=
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 5bfd387e-6cbb-4c70-a9bc-ce94c929ffbe · outbound
Beyond Negative Rollouts: Positive-Only Policy Optimization with Implicit Negative Gradients The surprising effectiveness of negative reinforcement in llm reasoning.arXiv preprint arXiv:2506.01347
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 7228c3f8-c4ae-447c-b9cb-859a9fdd0000 · outbound
Beyond Negative Rollouts: Positive-Only Policy Optimization with Implicit Negative Gradients Reinforcement Learning for Reasoning in Large Language Models with One Training Example
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 23087d96-c9eb-4e1a-9177-39a534c8d45b · outbound
Beyond Negative Rollouts: Positive-Only Policy Optimization with Implicit Negative Gradients Inftythink: Breaking the length limits of long-context reasoning in large language models
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation e379a87a-3ac0-4490-ae20-ea4c30bdc293 · outbound
Beyond Negative Rollouts: Positive-Only Policy Optimization with Implicit Negative Gradients Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation a41851d8-b134-4df4-8ec1-74d0c55a1f1b · outbound
Beyond Negative Rollouts: Positive-Only Policy Optimization with Implicit Negative Gradients Learning beyond Teacher: Generalized On-Policy Distillation with Reward Extrapolation
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation dbfb5099-ca7e-4e0a-9359-074699d578bc · outbound
Beyond Negative Rollouts: Positive-Only Policy Optimization with Implicit Negative Gradients Clipo: Contrastive learning in policy optimization generalizes rlvr
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 0745f993-d5b4-46db-bfb2-d6db30c84ef4 · outbound
Beyond Negative Rollouts: Positive-Only Policy Optimization with Implicit Negative Gradients IEEE transactions on knowledge and data engineering , volume=
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation f1261673-5610-4d57-9cf7-0bf72264cf1e · outbound
Beyond Negative Rollouts: Positive-Only Policy Optimization with Implicit Negative Gradients MathArena: Evaluating LLMs on Uncontaminated Math Competitions
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 21ff4e1c-e768-4139-b5d8-986c02e676a9 · outbound
Beyond Negative Rollouts: Positive-Only Policy Optimization with Implicit Negative Gradients Correcting the Mythos of KL-Regularization: Direct Alignment without Overoptimization via Chi-Squared Preference Optimization
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 3fdc76ff-37bc-4b45-8de3-fc3023dbb733 · outbound
Beyond Negative Rollouts: Positive-Only Policy Optimization with Implicit Negative Gradients DeepSeek-Coder: When the Large Language Model Meets Programming -- The Rise of Code Intelligence
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation a4480743-638b-4363-9214-561a3a1f912c · inbound
Multimodal Reward Hacking in Reinforcement Learning Beyond Negative Rollouts: Positive-Only Policy Optimization with Implicit Negative Gradients
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.