Pith. sign in

Paper Citation Record · LEDGER

On the Effect of Negative Gradient in Group Relative Deep Reinforcement Optimization

As of 17 August 2026, this Paper Citation Record lists 35 of 35 outbound references and 2 inbound Pith citation observations for arXiv:2505.18830.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.18830 v1

Coverage vector

measured 35 of 35 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:29:34.013794Z

measured 37 of 37 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T09:34:06.517190Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-11T15:16:09.090199Z

Reference resolution

35 of 35 outbound references displayed

  • verified exact1
  • verified fuzzy7
  • unresolved26
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 1c2ffe78-a13b-4ec9-8192-4f7d15c71d18 · outbound

This paper cites Gpg: A simple and strong reinforcement learning baseline for model reasoning.arXiv preprint arXiv:2504.02546, 2025.

On the Effect of Negative Gradient in Group Relative Deep Reinforcement Optimization Gpg: A simple and strong reinforcement learning baseline for model reasoning.arXiv preprint arXiv:2504.02546, 2025

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T14:29:30.082862Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:29:30.082862Z digest=sha256:8be7c5f6886e35b9c4c4197e05b229a90a98e69197df443a479c85d3d9c65c63

Observation 3519a9e7-04f0-4dbb-8464-61e670ba6ace · outbound

This paper cites DARE the Extreme: Revisiting Delta-Parameter Pruning For Fine-Tuned Models.

On the Effect of Negative Gradient in Group Relative Deep Reinforcement Optimization DARE the Extreme: Revisiting Delta-Parameter Pruning For Fine-Tuned Models

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-08-07T14:29:34.477945Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T14:29:30.182581Z digest=sha256:cf3e840678319d5007340c3b2ac0f527aa6ff553acbdda6931e3ae2553c58474

Observation b8da91cd-c47c-4176-8984-193b2cd7c071 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

On the Effect of Negative Gradient in Group Relative Deep Reinforcement Optimization DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T14:29:30.352756Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:29:30.352756Z digest=sha256:f7c81f96678452890ceea2ef23aacf448636553f8b79d640a27c3bfa7e49a289

Observation 0086d934-0735-4c00-b015-e426cf4b1793 · outbound

This paper cites OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems.

On the Effect of Negative Gradient in Group Relative Deep Reinforcement Optimization OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T14:29:30.472429Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:29:30.472429Z digest=sha256:50ce2f1acd310f3985ebfd7705249d69bbeaf952bdbeef7f22cdce8f07385bde

Observation ac04d9fc-c2c3-43dd-8890-51ecedb29804 · outbound

This paper cites The local elasticity of neural networks.

On the Effect of Negative Gradient in Group Relative Deep Reinforcement Optimization The local elasticity of neural networks

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:29:35.731705Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T14:29:30.598230Z digest=sha256:ffdde5dd34cdd642f2717bf8941f31b4608a910f6cace3548ae2e72bca1fa3cd

Observation 4e9a462d-48e0-493d-a2e0-c32b91e59503 · outbound

This paper cites Measuring mathematical problem solving with the math dataset.

On the Effect of Negative Gradient in Group Relative Deep Reinforcement Optimization Measuring mathematical problem solving with the math dataset

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T14:29:30.787987Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:29:30.787987Z digest=sha256:7e04f5c8c929370441b2e005754ca59c4020cccf689845951304bddf18741e71

Observation 01ee2673-2467-4663-a4dc-fe0a0a5ef0a9 · outbound

This paper cites Open-Reasoner-Zero: An Open Source Approach to Scaling Up Reinforcement Learning on the Base Model.

On the Effect of Negative Gradient in Group Relative Deep Reinforcement Optimization Open-Reasoner-Zero: An Open Source Approach to Scaling Up Reinforcement Learning on the Base Model

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T14:29:30.979686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:29:30.979686Z digest=sha256:b9127f2ca0bb62f119388cf74f66361f07ca9ebd6865dab31993d279542f069e

Observation e5e4e96a-22cc-4bcf-a91a-88ad472cf7c5 · outbound

This paper cites OpenAI o1 System Card.

On the Effect of Negative Gradient in Group Relative Deep Reinforcement Optimization OpenAI o1 System Card

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T14:29:31.111041Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:29:31.111041Z digest=sha256:37c259aa9671e4d1bf4fd4ed0e9eda1d83b5ee077539f6c867ab1cd0a465f41b

Observation 2e95c004-02d1-453c-956a-8e17c1eb7a31 · outbound

This paper cites Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning.

On the Effect of Negative Gradient in Group Relative Deep Reinforcement Optimization Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T14:29:31.259139Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:29:31.259139Z digest=sha256:b7de4ce612a0a31e7aea219c00905988ccaecd249b22ebd052bbd7706036b834

Observation 0213d9c1-098d-47b5-98e7-bc3263b303e9 · outbound

This paper cites Med-r1: Reinforce- ment learning for generalizable medical reasoning in vision-language models.arXiv preprint arXiv:2503.13939, 2025.

On the Effect of Negative Gradient in Group Relative Deep Reinforcement Optimization Med-r1: Reinforce- ment learning for generalizable medical reasoning in vision-language models.arXiv preprint arXiv:2503.13939, 2025

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T14:29:31.370676Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:29:31.370676Z digest=sha256:0c4bf31c479d59d9c9f0fd9199ee9598334c429ce1fc81bb14fa26be8a9cae2b

Observation 287758ef-d131-493f-a2a3-6fab4b88f346 · outbound

This paper cites Solving quantitative reasoning problems with language models.Advances in Neural Information Processing Systems, 35:3843–3857, 2022.

On the Effect of Negative Gradient in Group Relative Deep Reinforcement Optimization Solving quantitative reasoning problems with language models.Advances in Neural Information Processing Systems, 35:3843–3857, 2022

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T14:29:31.497301Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:29:31.497301Z digest=sha256:bb0c5358a916b4a107ff0606e1e9ee9dee17445d5f3182d0db84740a81d1f2ad

Observation 9138a9b8-c697-454e-9c8b-b2c04ae9b2f8 · outbound

This paper cites Understanding R1-Zero-Like Training: A Critical Perspective.

On the Effect of Negative Gradient in Group Relative Deep Reinforcement Optimization Understanding R1-Zero-Like Training: A Critical Perspective

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T14:29:31.672136Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:29:31.672136Z digest=sha256:eacb4bcc37c54141d044085c85fc13ddeaff0239988d2ae3ed6991feb7b46a2f

Observation 4bbda4f8-5219-4a63-b4d3-aca75d7097e0 · outbound

This paper cites Deepscaler: Surpassing o1-preview with a 1.5b model by scaling rl, 2025.

On the Effect of Negative Gradient in Group Relative Deep Reinforcement Optimization Deepscaler: Surpassing o1-preview with a 1.5b model by scaling rl, 2025

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T14:29:31.844838Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:29:31.844838Z digest=sha256:175c1a2ce85c2784525f2e3d02e463aae3aa7c9791cb83c4f98b7373a4a764f7

Observation 65cf08cb-4b81-423b-9bfe-765f73129ae8 · outbound

This paper cites Neural collapse with unconstrained features.

On the Effect of Negative Gradient in Group Relative Deep Reinforcement Optimization Neural collapse with unconstrained features

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:29:35.585539Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T14:29:32.001012Z digest=sha256:47e3b8ed0bca0afd411801449db9d9cb681ac6ee647ebf08fdc4cce057f9d07d

Observation cc9af7d2-8533-41df-8f1b-6f9a67183a9e · outbound

This paper cites Smaug: Fixing Failure Modes of Preference Optimisation with DPO-Positive.

On the Effect of Negative Gradient in Group Relative Deep Reinforcement Optimization Smaug: Fixing Failure Modes of Preference Optimisation with DPO-Positive

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T14:29:32.198096Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:29:32.198096Z digest=sha256:64b22f699d1d37a18bc11ca05db4f42e0eeca8f03907789a8aa4ee88a71b1756

Observation fd6b3d69-c9a3-4268-8dc1-340694b26bfb · outbound

This paper cites Direct preference optimization: Your language model is secretly a reward model.

On the Effect of Negative Gradient in Group Relative Deep Reinforcement Optimization Direct preference optimization: Your language model is secretly a reward model

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T14:29:32.334464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:29:32.334464Z digest=sha256:dde94058af1096850497e2c430bb2978c2b4b0a680a7d1ee05a15f660e6911d9

Observation d3f472fa-05a6-4984-b680-13d8fa5b52de · outbound

This paper cites Unintentional Unalignment: Likelihood Displacement in Direct Preference Optimization.

On the Effect of Negative Gradient in Group Relative Deep Reinforcement Optimization Unintentional Unalignment: Likelihood Displacement in Direct Preference Optimization

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T14:29:32.444781Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:29:32.444781Z digest=sha256:6182ec91af031ee9e7876f0e18db2b1e2d98c008dbf05a4bf61ff00f99b0f438

Observation cb0aa861-d8ce-4d6a-9d29-a536550589d8 · outbound

This paper cites Learning Dynamics of LLM Finetuning.

On the Effect of Negative Gradient in Group Relative Deep Reinforcement Optimization Learning Dynamics of LLM Finetuning

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T14:29:32.572317Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:29:32.572317Z digest=sha256:4d3fc4a3a0d114bd4d85f8dd98d37a6a0e3d93b0141c379097ea6f502560bf9b

Observation 41608882-f242-4ed3-8191-4bc586c69a70 · outbound

This paper cites Proximal Policy Optimization Algorithms.

On the Effect of Negative Gradient in Group Relative Deep Reinforcement Optimization Proximal Policy Optimization Algorithms

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T14:29:32.710722Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:29:32.710722Z digest=sha256:a472bbce8b0737430cac461fff3be8da0ab23a0a4db6e2a14dfccf44a2651f2a

Observation 726ede61-1ee1-403e-9235-575fa2c83438 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

On the Effect of Negative Gradient in Group Relative Deep Reinforcement Optimization DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T14:29:32.815490Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:29:32.815490Z digest=sha256:0229627a0e0821cdd685afad0a88101ec4708c0378b2d277ffdf4aa9ca793250

Observation 01b333f3-1f0b-4378-a7d7-4d48e3286691 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

On the Effect of Negative Gradient in Group Relative Deep Reinforcement Optimization Gemini: A Family of Highly Capable Multimodal Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T14:29:32.921726Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:29:32.921726Z digest=sha256:dc66e16abd797e003ab3a72a396b1deb7e1feaccbc33e4a4f7372d497000a625

Observation 96dff3bd-5d64-4f4f-ae3c-5372b0de278a · outbound

This paper cites Aime problem set 1983-2024, 2023.

On the Effect of Negative Gradient in Group Relative Deep Reinforcement Optimization Aime problem set 1983-2024, 2023

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:29:35.509164Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T14:29:33.000764Z digest=sha256:a3305936137a8c8e06c8e7b0f10c714988d34523ddfb8a07ae450957bd0e7468

Observation 21575429-1287-40a5-b8cd-f7afa7856141 · outbound

This paper cites Qwen2.5 Technical Report.

On the Effect of Negative Gradient in Group Relative Deep Reinforcement Optimization Qwen2.5 Technical Report

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T14:29:33.073364Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:29:33.073364Z digest=sha256:4f21697618527654abe79fff7884149ccb97916c6a87b0e87ee8ed8b983fbe43

Observation a15f090f-5325-4091-b01e-7fe9210c3d57 · outbound

This paper cites Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement.

On the Effect of Negative Gradient in Group Relative Deep Reinforcement Optimization Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T14:29:33.120361Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:29:33.120361Z digest=sha256:ed79826645dd791a8f309a7a1a75d5b6c31b062cd4f6eea3af7261560e12414f

Observation bd1349c0-d83a-4598-8710-ddbb685d6d35 · outbound

This paper cites Breaking the Softmax Bottleneck: A High-Rank RNN Language Model.

On the Effect of Negative Gradient in Group Relative Deep Reinforcement Optimization Breaking the Softmax Bottleneck: A High-Rank RNN Language Model

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T14:29:33.191566Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:29:33.191566Z digest=sha256:07f0f5e4780b956a15e93005ed09fa7a779dbbe5328b518a172afb09caece3c8

Observation 86e8509a-9d1f-4541-8271-238cb2c324a4 · outbound

This paper cites MetaMath: Bootstrap Your Own Mathematical Questions for Large Language Models.

On the Effect of Negative Gradient in Group Relative Deep Reinforcement Optimization MetaMath: Bootstrap Your Own Mathematical Questions for Large Language Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T14:29:33.267870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:29:33.267870Z digest=sha256:bee70258c54fb1adcd117f982e3ecd5ab6114c35c113e0e0a8789912998fb7f3

Observation 1e7a4dd1-a25f-4783-b71d-bac8905ba860 · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

On the Effect of Negative Gradient in Group Relative Deep Reinforcement Optimization DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T14:29:33.323315Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:29:33.323315Z digest=sha256:7d1de6c2dc652e79654ce658c30deab9b6399ad09e88dd24187327857bbfa2b4

Observation 5bab72b9-5519-4cfb-8f12-b78c64860908 · outbound

This paper cites Advancing LLM Reasoning Generalists with Preference Trees.

On the Effect of Negative Gradient in Group Relative Deep Reinforcement Optimization Advancing LLM Reasoning Generalists with Preference Trees

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T14:29:33.431718Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:29:33.431718Z digest=sha256:a1ee55050c0d3bc4312fde21cc34398f374bad4871dc2b8dcec4fb643d86397d

Observation b04807af-9df9-4bc3-b30c-5558edf2ef4f · outbound

This paper cites SimpleRL-Zoo: Investigating and Taming Zero Reinforcement Learning for Open Base Models in the Wild.

On the Effect of Negative Gradient in Group Relative Deep Reinforcement Optimization SimpleRL-Zoo: Investigating and Taming Zero Reinforcement Learning for Open Base Models in the Wild

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T14:29:33.492988Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:29:33.492988Z digest=sha256:87afec468422c6e12916db6e6d30e6be8e0f8d807356ac4aaed1f971bdb47cc8

Observation 4177ec77-2d36-4819-af22-815479f91d08 · outbound

This paper cites Implicit Geometry of Next-token Prediction: From Language Sparsity Patterns to Model Representations.

On the Effect of Negative Gradient in Group Relative Deep Reinforcement Optimization Implicit Geometry of Next-token Prediction: From Language Sparsity Patterns to Model Representations

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T14:29:33.608339Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:29:33.608339Z digest=sha256:1f26a614c5d21cbf46b39851583e4902aa96cc577497583a056b2370ddd6306c

Observation 93de0474-6ead-4896-9f95-87061801dad4 · outbound

This paper cites one commercially available ten-button lock may be opened by pressing – in any order – the correct five buttons.

On the Effect of Negative Gradient in Group Relative Deep Reinforcement Optimization one commercially available ten-button lock may be opened by pressing – in any order – the correct five buttons

Reference 31

Resolution
malformed identifier
raw_fallback, observed 2026-08-07T14:29:35.313069Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T14:29:33.696512Z digest=sha256:39e8d4aa613718038d7904cf80cbf17427f01f4c10d2d13258ce735c884cc157

Observation 37e0f6f9-5bc9-40c9-95f9-8c882721387e · outbound

This paper cites From the given graph, we can observe the following: - The roots off(x)are atx= 1 andx= 3.

On the Effect of Negative Gradient in Group Relative Deep Reinforcement Optimization From the given graph, we can observe the following: - The roots off(x)are atx= 1 andx= 3

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:29:35.094259Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T14:29:33.803416Z digest=sha256:53600253c0ef8e8ccd44c6699c64edc69557b1e8a60213008f17a91043956dbd

Observation 1ad0abfb-49ad-4c55-80ac-4b4a33f33c88 · outbound

This paper cites Therefore: - The roots of g(x)are also atx= 1 andx= 3.

On the Effect of Negative Gradient in Group Relative Deep Reinforcement Optimization Therefore: - The roots of g(x)are also atx= 1 andx= 3

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:29:34.956361Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T14:29:33.875782Z digest=sha256:a61912c5652e8d3909c3329850441237f9b593c6f49ff4044078cfcd07989d69

Observation 7ba4801a-f4b9-485c-8328-5cb5b74f5928 · outbound

This paper cites This simplifies to: f(x) = 0 The roots off(x)are atx= 1 andx= 3.

On the Effect of Negative Gradient in Group Relative Deep Reinforcement Optimization This simplifies to: f(x) = 0 The roots off(x)are atx= 1 andx= 3

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:29:34.870775Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T14:29:33.935421Z digest=sha256:572ccb7c7d68069edbe976ffea754546e627996db9efb5b085faafe5b44d9963

Observation d34fc82a-47eb-4b63-883c-158d98591680 · outbound

This paper cites This implies thatf(x)is an even function, and its graph is symmetric about the y-axis.

On the Effect of Negative Gradient in Group Relative Deep Reinforcement Optimization This implies thatf(x)is an even function, and its graph is symmetric about the y-axis

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:29:34.738641Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-07T14:29:34.013794Z digest=sha256:ebdf5dd24795cabc8b3c3394e2f17927f155f171162626ef262bbcccbbc15a89

Pith citing papers

Observation 151aac32-085e-47ea-a121-e760e2719af9 · inbound

Beyond the Sampled Token: Preserving Candidate Support in RLVR cites this paper.

Beyond the Sampled Token: Preserving Candidate Support in RLVR On the Effect of Negative Gradient in Group Relative Deep Reinforcement Optimization

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-04T09:34:06.517190Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T09:34:06.517190Z digest=sha256:5610e3030c845f240cedb0029efa61b3455eb70d7b4278460a67b03ad14069e4

Observation 89a73c6b-0b37-423f-a8fd-f01398c0874b · inbound

Odysseus: Scaling VLMs to 100+ Turn Decision-Making in Games via Reinforcement Learning cites this paper.

Odysseus: Scaling VLMs to 100+ Turn Decision-Making in Games via Reinforcement Learning On the Effect of Negative Gradient in Group Relative Deep Reinforcement Optimization

Reference 78

Resolution
verified exact
arxiv_id, observed 2026-05-11T15:16:09.093200Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-05-09T20:22:58.061772Z digest=sha256:acf122489c6643830c51336f6901852a150d0a5cd6754a232dfe5d4caca7e18d