Pith. sign in

Paper Citation Record · LEDGER

On the Effect of Negative Gradient in Group Relative Deep Reinforcement Optimization

As of 9 August 2026, this Paper Citation Record lists 35 of 35 outbound references and 2 inbound Pith citation observations for arXiv:2505.18830.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.18830 v1

Coverage vector

measured 35 of 35 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:29:34.013794Z

measured 37 of 37 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T09:34:06.517190Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-11T15:16:09.090199Z

Reference resolution

35 of 35 outbound references displayed

  • verified exact1
  • verified fuzzy7
  • unresolved26
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 1c2ffe78-a13b-4ec9-8192-4f7d15c71d18 · outbound

This paper cites Gpg: A simple and strong reinforcement learning baseline for model reasoning.arXiv preprint arXiv:2504.02546, 2025.

On the Effect of Negative Gradient in Group Relative Deep Reinforcement Optimization Gpg: A simple and strong reinforcement learning baseline for model reasoning.arXiv preprint arXiv:2504.02546, 2025

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T14:29:30.082862Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:29:30.082862Z digest=sha256:5a8be375f8a371ceec1d10aeee342dcaa32b79b18120e547da20b0e273d9d355

Observation 3519a9e7-04f0-4dbb-8464-61e670ba6ace · outbound

This paper cites DARE the Extreme: Revisiting Delta-Parameter Pruning For Fine-Tuned Models.

On the Effect of Negative Gradient in Group Relative Deep Reinforcement Optimization DARE the Extreme: Revisiting Delta-Parameter Pruning For Fine-Tuned Models

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-08-07T14:29:34.477945Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:29:30.182581Z digest=sha256:f7e85b816bc4a1dcca95ad9f47ee1b3325627a9110fdac2b1acb0b483d646014

Observation b8da91cd-c47c-4176-8984-193b2cd7c071 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

On the Effect of Negative Gradient in Group Relative Deep Reinforcement Optimization DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T14:29:30.352756Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:29:30.352756Z digest=sha256:159d6d42c1798ab5c7f606b2bbd390869ecbb5d21064a5983388c5e7a6e9f921

Observation 0086d934-0735-4c00-b015-e426cf4b1793 · outbound

This paper cites OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems.

On the Effect of Negative Gradient in Group Relative Deep Reinforcement Optimization OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T14:29:30.472429Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:29:30.472429Z digest=sha256:424ee58e2b87716791913fc51fb8894f02ca9b6f3380a78b0413243c6b70df4a

Observation ac04d9fc-c2c3-43dd-8890-51ecedb29804 · outbound

This paper cites The local elasticity of neural networks.

On the Effect of Negative Gradient in Group Relative Deep Reinforcement Optimization The local elasticity of neural networks

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:29:35.731705Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:29:30.598230Z digest=sha256:09d9654e0ffceaf5a072350de7aa459524fb3e56a093d7fc589b6e6a4321e941

Observation 4e9a462d-48e0-493d-a2e0-c32b91e59503 · outbound

This paper cites Measuring mathematical problem solving with the math dataset.

On the Effect of Negative Gradient in Group Relative Deep Reinforcement Optimization Measuring mathematical problem solving with the math dataset

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T14:29:30.787987Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:29:30.787987Z digest=sha256:be439266ee9ee6a74c28e3934172c91f0d133de18cd777fbae15cabc0e42535a

Observation 01ee2673-2467-4663-a4dc-fe0a0a5ef0a9 · outbound

This paper cites Open-Reasoner-Zero: An Open Source Approach to Scaling Up Reinforcement Learning on the Base Model.

On the Effect of Negative Gradient in Group Relative Deep Reinforcement Optimization Open-Reasoner-Zero: An Open Source Approach to Scaling Up Reinforcement Learning on the Base Model

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T14:29:30.979686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:29:30.979686Z digest=sha256:7786fbae8df28ffb1d73f5b1fb8236a90352db6744023922fff6fc855a9d09b1

Observation e5e4e96a-22cc-4bcf-a91a-88ad472cf7c5 · outbound

This paper cites OpenAI o1 System Card.

On the Effect of Negative Gradient in Group Relative Deep Reinforcement Optimization OpenAI o1 System Card

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T14:29:31.111041Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:29:31.111041Z digest=sha256:15ee4dfd7690a6ff6423a7fc228e6a14832fa333dd14bb8c51c9ae97688fb033

Observation 2e95c004-02d1-453c-956a-8e17c1eb7a31 · outbound

This paper cites Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning.

On the Effect of Negative Gradient in Group Relative Deep Reinforcement Optimization Search-R1: Training LLMs to Reason and Leverage Search Engines with Reinforcement Learning

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T14:29:31.259139Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:29:31.259139Z digest=sha256:5a0a8d218584bea86218872562f7500be4e9cd59b2a96f24b02c881248e7fcbc

Observation 0213d9c1-098d-47b5-98e7-bc3263b303e9 · outbound

This paper cites Med-r1: Reinforce- ment learning for generalizable medical reasoning in vision-language models.arXiv preprint arXiv:2503.13939, 2025.

On the Effect of Negative Gradient in Group Relative Deep Reinforcement Optimization Med-r1: Reinforce- ment learning for generalizable medical reasoning in vision-language models.arXiv preprint arXiv:2503.13939, 2025

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T14:29:31.370676Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:29:31.370676Z digest=sha256:665c7dcd792c3f55c37da61023c9d5cbcd76aaa8721518d38f0823c01269aa18

Observation 287758ef-d131-493f-a2a3-6fab4b88f346 · outbound

This paper cites Solving quantitative reasoning problems with language models.Advances in Neural Information Processing Systems, 35:3843–3857, 2022.

On the Effect of Negative Gradient in Group Relative Deep Reinforcement Optimization Solving quantitative reasoning problems with language models.Advances in Neural Information Processing Systems, 35:3843–3857, 2022

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T14:29:31.497301Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:29:31.497301Z digest=sha256:208bde4c0d31250dea0281d5ed31781185bf24f2a214d031f1718c9eca1280d4

Observation 9138a9b8-c697-454e-9c8b-b2c04ae9b2f8 · outbound

This paper cites Understanding R1-Zero-Like Training: A Critical Perspective.

On the Effect of Negative Gradient in Group Relative Deep Reinforcement Optimization Understanding R1-Zero-Like Training: A Critical Perspective

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T14:29:31.672136Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:29:31.672136Z digest=sha256:833b4707357d2784c6a1c7498650af9460044501d2365e6e17960e01fe90c399

Observation 4bbda4f8-5219-4a63-b4d3-aca75d7097e0 · outbound

This paper cites Deepscaler: Surpassing o1-preview with a 1.5b model by scaling rl, 2025.

On the Effect of Negative Gradient in Group Relative Deep Reinforcement Optimization Deepscaler: Surpassing o1-preview with a 1.5b model by scaling rl, 2025

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T14:29:31.844838Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:29:31.844838Z digest=sha256:8a4b891d5052390c200bd561afd9d71402e09d370f40f05ba57e02cdfa515d5e

Observation 65cf08cb-4b81-423b-9bfe-765f73129ae8 · outbound

This paper cites Neural collapse with unconstrained features.

On the Effect of Negative Gradient in Group Relative Deep Reinforcement Optimization Neural collapse with unconstrained features

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:29:35.585539Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:29:32.001012Z digest=sha256:21c3efd65b9bad756a3aa2186ace465e9bfa46a023fb9d2a27f6a74ab6a8f9ef

Observation cc9af7d2-8533-41df-8f1b-6f9a67183a9e · outbound

This paper cites Smaug: Fixing Failure Modes of Preference Optimisation with DPO-Positive.

On the Effect of Negative Gradient in Group Relative Deep Reinforcement Optimization Smaug: Fixing Failure Modes of Preference Optimisation with DPO-Positive

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T14:29:32.198096Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:29:32.198096Z digest=sha256:470c8807458098ce09ebde2952c777ded9e02ee8bc8a8934bfc07ddbdeeaa59a

Observation fd6b3d69-c9a3-4268-8dc1-340694b26bfb · outbound

This paper cites Direct preference optimization: Your language model is secretly a reward model.

On the Effect of Negative Gradient in Group Relative Deep Reinforcement Optimization Direct preference optimization: Your language model is secretly a reward model

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T14:29:32.334464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:29:32.334464Z digest=sha256:a8a2057e933559f459634567ad08db1c6ea62a053dce5a8df26a52cfb1d38af4

Observation d3f472fa-05a6-4984-b680-13d8fa5b52de · outbound

This paper cites Unintentional Unalignment: Likelihood Displacement in Direct Preference Optimization.

On the Effect of Negative Gradient in Group Relative Deep Reinforcement Optimization Unintentional Unalignment: Likelihood Displacement in Direct Preference Optimization

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T14:29:32.444781Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:29:32.444781Z digest=sha256:1775e31c4108d25c821edd0e447a7fdf6a2a311945f1e1a402bf0554327008e2

Observation cb0aa861-d8ce-4d6a-9d29-a536550589d8 · outbound

This paper cites Learning Dynamics of LLM Finetuning.

On the Effect of Negative Gradient in Group Relative Deep Reinforcement Optimization Learning Dynamics of LLM Finetuning

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T14:29:32.572317Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:29:32.572317Z digest=sha256:56b7c0af299d15bae76e4c55cf2cc4d0a149b431d7b6aadb777b7281a746c648

Observation 41608882-f242-4ed3-8191-4bc586c69a70 · outbound

This paper cites Proximal Policy Optimization Algorithms.

On the Effect of Negative Gradient in Group Relative Deep Reinforcement Optimization Proximal Policy Optimization Algorithms

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T14:29:32.710722Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:29:32.710722Z digest=sha256:abede2a4ef04a6192be97b003ff7ba0a11ec242776bc68db4dd1226879e292c6

Observation 726ede61-1ee1-403e-9235-575fa2c83438 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

On the Effect of Negative Gradient in Group Relative Deep Reinforcement Optimization DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T14:29:32.815490Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:29:32.815490Z digest=sha256:a811854f5842421b8c30e2b2d123ba0c743e012a121e3e7006860f7d6ac4d8ad

Observation 01b333f3-1f0b-4378-a7d7-4d48e3286691 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

On the Effect of Negative Gradient in Group Relative Deep Reinforcement Optimization Gemini: A Family of Highly Capable Multimodal Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T14:29:32.921726Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:29:32.921726Z digest=sha256:a9d87713615361833c0e95ae35fa12ac4fbaad635cce8e9ca793682c41d2db23

Observation 96dff3bd-5d64-4f4f-ae3c-5372b0de278a · outbound

This paper cites Aime problem set 1983-2024, 2023.

On the Effect of Negative Gradient in Group Relative Deep Reinforcement Optimization Aime problem set 1983-2024, 2023

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:29:35.509164Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:29:33.000764Z digest=sha256:ec0e4a0ebff4ee2495573766d3923fea34799580e1307added0fa227b512a5fe

Observation 21575429-1287-40a5-b8cd-f7afa7856141 · outbound

This paper cites Qwen2.5 Technical Report.

On the Effect of Negative Gradient in Group Relative Deep Reinforcement Optimization Qwen2.5 Technical Report

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T14:29:33.073364Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:29:33.073364Z digest=sha256:aeaaacc85282997a9582c60edb31a4ce59b0876c1b0b20b9d3199208211d0489

Observation a15f090f-5325-4091-b01e-7fe9210c3d57 · outbound

This paper cites Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement.

On the Effect of Negative Gradient in Group Relative Deep Reinforcement Optimization Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T14:29:33.120361Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:29:33.120361Z digest=sha256:d95222d01b37d97e687a39b6291dedd231155f2a76ae1985eda988c5cc5750c9

Observation bd1349c0-d83a-4598-8710-ddbb685d6d35 · outbound

This paper cites Breaking the Softmax Bottleneck: A High-Rank RNN Language Model.

On the Effect of Negative Gradient in Group Relative Deep Reinforcement Optimization Breaking the Softmax Bottleneck: A High-Rank RNN Language Model

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T14:29:33.191566Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:29:33.191566Z digest=sha256:dd347d3a85fcba44fc4401644d1d18b4fb45fd3b834227ac9f0c81709f82730d

Observation 86e8509a-9d1f-4541-8271-238cb2c324a4 · outbound

This paper cites MetaMath: Bootstrap Your Own Mathematical Questions for Large Language Models.

On the Effect of Negative Gradient in Group Relative Deep Reinforcement Optimization MetaMath: Bootstrap Your Own Mathematical Questions for Large Language Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T14:29:33.267870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:29:33.267870Z digest=sha256:57bba9a8421a3a2ef6ea85507b32bd9c1557ce5dd158e52b3c928ac1fc30792f

Observation 1e7a4dd1-a25f-4783-b71d-bac8905ba860 · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

On the Effect of Negative Gradient in Group Relative Deep Reinforcement Optimization DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T14:29:33.323315Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:29:33.323315Z digest=sha256:cd3455371a0142945b0019bcc3ea2adc7c5c564aeebdeebfcbeefb5720561fac

Observation 5bab72b9-5519-4cfb-8f12-b78c64860908 · outbound

This paper cites Advancing LLM Reasoning Generalists with Preference Trees.

On the Effect of Negative Gradient in Group Relative Deep Reinforcement Optimization Advancing LLM Reasoning Generalists with Preference Trees

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T14:29:33.431718Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:29:33.431718Z digest=sha256:5c598761ef2e0d423118c35846c9b78e1d0d0f14d8b1f60ce3055e36753966fc

Observation b04807af-9df9-4bc3-b30c-5558edf2ef4f · outbound

This paper cites SimpleRL-Zoo: Investigating and Taming Zero Reinforcement Learning for Open Base Models in the Wild.

On the Effect of Negative Gradient in Group Relative Deep Reinforcement Optimization SimpleRL-Zoo: Investigating and Taming Zero Reinforcement Learning for Open Base Models in the Wild

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T14:29:33.492988Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:29:33.492988Z digest=sha256:e14d52756d834462b5f142b38b97d544e4ac30170a0b657c41566a148717735a

Observation 4177ec77-2d36-4819-af22-815479f91d08 · outbound

This paper cites Implicit Geometry of Next-token Prediction: From Language Sparsity Patterns to Model Representations.

On the Effect of Negative Gradient in Group Relative Deep Reinforcement Optimization Implicit Geometry of Next-token Prediction: From Language Sparsity Patterns to Model Representations

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T14:29:33.608339Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:29:33.608339Z digest=sha256:3f40e3f194ffd54b81bed1bb2eb49f3579a7e174ecbea9bff047271308471941

Observation 93de0474-6ead-4896-9f95-87061801dad4 · outbound

This paper cites one commercially available ten-button lock may be opened by pressing – in any order – the correct five buttons.

On the Effect of Negative Gradient in Group Relative Deep Reinforcement Optimization one commercially available ten-button lock may be opened by pressing – in any order – the correct five buttons

Reference 31

Resolution
malformed identifier
raw_fallback, observed 2026-08-07T14:29:35.313069Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:29:33.696512Z digest=sha256:3054ce19419644b54ef7716a3689ce88fd7b4523132daaa2c078bee204848602

Observation 37e0f6f9-5bc9-40c9-95f9-8c882721387e · outbound

This paper cites From the given graph, we can observe the following: - The roots off(x)are atx= 1 andx= 3.

On the Effect of Negative Gradient in Group Relative Deep Reinforcement Optimization From the given graph, we can observe the following: - The roots off(x)are atx= 1 andx= 3

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:29:35.094259Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:29:33.803416Z digest=sha256:3052a6fe2ac80cfca286c6c6ed95eb504e113b1039ce88691f68d36d8e7a9c6c

Observation 1ad0abfb-49ad-4c55-80ac-4b4a33f33c88 · outbound

This paper cites Therefore: - The roots of g(x)are also atx= 1 andx= 3.

On the Effect of Negative Gradient in Group Relative Deep Reinforcement Optimization Therefore: - The roots of g(x)are also atx= 1 andx= 3

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:29:34.956361Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:29:33.875782Z digest=sha256:651dfd932b2c08188635a3a6f83cb06cea47e24d9930b917973adf3971c809e3

Observation 7ba4801a-f4b9-485c-8328-5cb5b74f5928 · outbound

This paper cites This simplifies to: f(x) = 0 The roots off(x)are atx= 1 andx= 3.

On the Effect of Negative Gradient in Group Relative Deep Reinforcement Optimization This simplifies to: f(x) = 0 The roots off(x)are atx= 1 andx= 3

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:29:34.870775Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:29:33.935421Z digest=sha256:33b60c9224fdfa0fea4b6d35e388d7e80479a6d4a41b5f794d04a3fac9fa3e1e

Observation d34fc82a-47eb-4b63-883c-158d98591680 · outbound

This paper cites This implies thatf(x)is an even function, and its graph is symmetric about the y-axis.

On the Effect of Negative Gradient in Group Relative Deep Reinforcement Optimization This implies thatf(x)is an even function, and its graph is symmetric about the y-axis

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:29:34.738641Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:29:34.013794Z digest=sha256:7f6f0509384a64c8d863bbe34e6db00c586b7d68f94ddb7ea5e6feebe7d344a8

Pith citing papers

Observation 151aac32-085e-47ea-a121-e760e2719af9 · inbound

Beyond the Sampled Token: Preserving Candidate Support in RLVR cites this paper.

Beyond the Sampled Token: Preserving Candidate Support in RLVR On the Effect of Negative Gradient in Group Relative Deep Reinforcement Optimization

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-04T09:34:06.517190Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T09:34:06.517190Z digest=sha256:f9e7f855361014aa1d2dbebfc2ef7117f92d88b18d8aa1a5607c25d80fb2e75d

Observation 89a73c6b-0b37-423f-a8fd-f01398c0874b · inbound

Odysseus: Scaling VLMs to 100+ Turn Decision-Making in Games via Reinforcement Learning cites this paper.

Odysseus: Scaling VLMs to 100+ Turn Decision-Making in Games via Reinforcement Learning On the Effect of Negative Gradient in Group Relative Deep Reinforcement Optimization

Reference 78

Resolution
verified exact
arxiv_id, observed 2026-05-11T15:16:09.093200Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-09T20:22:58.061772Z digest=sha256:181c6196721902f4f580c340aec5a742a7708ace3bef26ace954a1f336108b12