Pith. sign in

Paper Citation Record · LEDGER

From Trial-and-Error to Improvement: A Systematic Analysis of LLM Exploration Mechanisms in RLVR

As of 12 August 2026, this Paper Citation Record lists 33 of 33 outbound references and 3 inbound Pith citation observations for arXiv:2508.07534.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.07534 v2

Coverage vector

measured 33 of 33 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T22:03:28.188682Z

measured 36 of 36 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-18T00:02:24.352947Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-18T00:02:25.252362Z

Reference resolution

33 of 33 outbound references displayed

  • verified exact0
  • verified fuzzy5
  • unresolved28
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 322ba5f0-cb9c-49f0-8b58-317bba6f33f2 · outbound

This paper cites Decomposing the Entropy-Performance Exchange: The Missing Keys to Unlocking Effective Reinforcement Learning.

From Trial-and-Error to Improvement: A Systematic Analysis of LLM Exploration Mechanisms in RLVR Decomposing the Entropy-Performance Exchange: The Missing Keys to Unlocking Effective Reinforcement Learning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-05T22:03:28.053726Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T22:03:28.053726Z digest=sha256:47d8f6f7114c69138ce10d7460e44c8ad83a1f0d7d22e98829cd777236340087

Observation 2905d31a-65ad-473e-818e-f86ae11c2154 · outbound

This paper cites A Survey of Large Language Models.

From Trial-and-Error to Improvement: A Systematic Analysis of LLM Exploration Mechanisms in RLVR A Survey of Large Language Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-05T22:03:28.058500Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T22:03:28.058500Z digest=sha256:738c2fcf2734534aaf2aa15d39c89e166cc2c53f052df45776f1a16de7463229

Observation 17e3e04c-cc53-4260-ba52-1bdd66cd14ea · outbound

This paper cites Training language models to follow instructions with human feedback.

From Trial-and-Error to Improvement: A Systematic Analysis of LLM Exploration Mechanisms in RLVR Training language models to follow instructions with human feedback

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-05T22:03:28.063021Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T22:03:28.063021Z digest=sha256:54766fd3e1e88a4fa7dc94f4202d78edc814ef8e3dbe7a85964245b6a9711e72

Observation b421bde8-3511-4eac-92c2-d25e8ddbe07a · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

From Trial-and-Error to Improvement: A Systematic Analysis of LLM Exploration Mechanisms in RLVR DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-05T22:03:28.067351Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T22:03:28.067351Z digest=sha256:7ba9a03aa7e54af51e271ed94a23bff5da8a360ddb8b3ecb8c140fd9a9d1323f

Observation 9624f8e9-1b9d-4e39-bf3b-54a9be2b8fbc · outbound

This paper cites an unresolved cited work.

From Trial-and-Error to Improvement: A Systematic Analysis of LLM Exploration Mechanisms in RLVR Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-05T22:03:28.762818Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-05T22:03:28.071721Z digest=sha256:1867f2c6b001468244b9abf827597385646b96737a90d4af044a2357bfff00ec

Observation f4095a71-a846-4d77-bd8c-352760ee80b8 · outbound

This paper cites Tulu 3: Pushing Frontiers in Open Language Model Post-Training.

From Trial-and-Error to Improvement: A Systematic Analysis of LLM Exploration Mechanisms in RLVR Tulu 3: Pushing Frontiers in Open Language Model Post-Training

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-05T22:03:28.075990Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T22:03:28.075990Z digest=sha256:6ca7919e1f49e980afaabf3d565d77ddb978eb3a57f8e57c5df7d236d07b49e6

Observation 24bdff8e-b204-4b53-a0f5-c33bd7634669 · outbound

This paper cites Reasoning with Exploration: An Entropy Perspective.

From Trial-and-Error to Improvement: A Systematic Analysis of LLM Exploration Mechanisms in RLVR Reasoning with Exploration: An Entropy Perspective

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-05T22:03:28.081012Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T22:03:28.081012Z digest=sha256:ca8fe8996168669591060c728e05acc41cac4f300878169f686715e82f2d6176

Observation 011fa54c-4c8a-4470-a3b9-79a31d60fabc · outbound

This paper cites Exploration in deep reinforcement learning: A survey.

From Trial-and-Error to Improvement: A Systematic Analysis of LLM Exploration Mechanisms in RLVR Exploration in deep reinforcement learning: A survey

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-05T22:03:28.085936Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T22:03:28.085936Z digest=sha256:af1facce0ed9792eb8187d78638fadd3e5503a74cef0c1f09ae5be898d988504

Observation b4e661c4-d523-4537-b298-07d6ce18cea9 · outbound

This paper cites The Entropy Mechanism of Reinforcement Learning for Reasoning Language Models.

From Trial-and-Error to Improvement: A Systematic Analysis of LLM Exploration Mechanisms in RLVR The Entropy Mechanism of Reinforcement Learning for Reasoning Language Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-05T22:03:28.090177Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T22:03:28.090177Z digest=sha256:556efcc6c46dbfac1bf1be1683f3562b253401f3befaa0ddbba542aa6b97a568

Observation 67bce6db-ecda-40bc-9b2e-5cd71a80e58a · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

From Trial-and-Error to Improvement: A Systematic Analysis of LLM Exploration Mechanisms in RLVR DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-05T22:03:28.094332Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T22:03:28.094332Z digest=sha256:9a3593d0de46856ab079b41adb91141f4d4ac262060083b843c15359fe6522d3

Observation d6d691cb-a13d-487d-b724-d22111c95a82 · outbound

This paper cites The surprising effectiveness of negative reinforcement in llm reasoning.

From Trial-and-Error to Improvement: A Systematic Analysis of LLM Exploration Mechanisms in RLVR The surprising effectiveness of negative reinforcement in llm reasoning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-05T22:03:28.098747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T22:03:28.098747Z digest=sha256:9b982d8929d6745f85b3edb0f271dacc6dd55fe68636fe56d91b298e4eaa5472

Observation bd046845-72d6-4737-8501-5894b9ebd3da · outbound

This paper cites Math-shepherd: Verify and reinforce llms step-by-step without human annotations.

From Trial-and-Error to Improvement: A Systematic Analysis of LLM Exploration Mechanisms in RLVR Math-shepherd: Verify and reinforce llms step-by-step without human annotations

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T22:03:28.737764Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-05T22:03:28.103031Z digest=sha256:17b1c128f89e5b124e22600de1dff6a7a062b31e78e666c71b4fbdd006f2e571

Observation c2f1dcbc-69c2-4cb8-bdc3-a75f9d9db686 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

From Trial-and-Error to Improvement: A Systematic Analysis of LLM Exploration Mechanisms in RLVR DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-05T22:03:28.107150Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T22:03:28.107150Z digest=sha256:7ddafe4ed74c036deee705e4de696edf2e8b65c1fc3f76e84e61f2e6c9360754

Observation 50c3b81c-2c61-4c63-8326-dde21d646ebe · outbound

This paper cites Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?.

From Trial-and-Error to Improvement: A Systematic Analysis of LLM Exploration Mechanisms in RLVR Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-05T22:03:28.111517Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T22:03:28.111517Z digest=sha256:dead0324b36753dc14fd13447d8e785617b3ff369afea31c74687d45b8abef4d

Observation 0b894881-4190-409d-9d68-3950e4797a83 · outbound

This paper cites ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models.

From Trial-and-Error to Improvement: A Systematic Analysis of LLM Exploration Mechanisms in RLVR ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-05T22:03:28.116441Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T22:03:28.116441Z digest=sha256:5614910bd147138d4977809a962a726fc6787241f82ab99ede3ca03b5058e474

Observation 55ca83fb-bc87-4b29-8023-c78c84d38c92 · outbound

This paper cites Evaluating Large Language Models Trained on Code.

From Trial-and-Error to Improvement: A Systematic Analysis of LLM Exploration Mechanisms in RLVR Evaluating Large Language Models Trained on Code

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-05T22:03:28.120627Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T22:03:28.120627Z digest=sha256:beb7bcd0dba437b164e6e3bd91a8ea4f715ad5dcb358a247ebca5c687f542193

Observation c40f1373-96d7-47d6-bf88-a85fa7fb7c7d · outbound

This paper cites Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models.

From Trial-and-Error to Improvement: A Systematic Analysis of LLM Exploration Mechanisms in RLVR Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-05T22:03:28.124820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T22:03:28.124820Z digest=sha256:f670ea598b20ad03392eb64af5f297fc3faee5c549b34318b7a139022c838b37

Observation 55958e60-b658-4583-af0f-e5401c34048b · outbound

This paper cites an unresolved cited work.

From Trial-and-Error to Improvement: A Systematic Analysis of LLM Exploration Mechanisms in RLVR Unresolved cited work

Reference 18

Resolution
unresolved
raw_fallback, observed 2026-08-05T22:03:28.722600Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-05T22:03:28.128868Z digest=sha256:c3e4bd9763faec935f4e9552a2fcbd8416beae130b4d428b9745f4d76b7bd8a8

Observation a748fea4-02d5-468e-bc1d-8767a55ab860 · outbound

This paper cites Let's verify step by step.

From Trial-and-Error to Improvement: A Systematic Analysis of LLM Exploration Mechanisms in RLVR Let's verify step by step

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T22:03:28.709392Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-05T22:03:28.132945Z digest=sha256:fb13c54d4f49017ff8b9a32eabe184f1bbe61619a874c18d37e088de50edb05d

Observation 54180ce1-e458-43c7-9ce9-475e1ca386df · outbound

This paper cites Generative verifiers: Reward modeling as next-token prediction.

From Trial-and-Error to Improvement: A Systematic Analysis of LLM Exploration Mechanisms in RLVR Generative verifiers: Reward modeling as next-token prediction

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T22:03:28.693699Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-05T22:03:28.136547Z digest=sha256:cc45a1a2163b13216b5e10289e2c36a2623fce93f161cd1fedf67534455b5990

Observation 162588cb-f6b0-4e4e-8c70-8508c9f4bc9b · outbound

This paper cites The lessons of developing process reward models in mathematical reasoning.

From Trial-and-Error to Improvement: A Systematic Analysis of LLM Exploration Mechanisms in RLVR The lessons of developing process reward models in mathematical reasoning

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T22:03:28.679544Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-05T22:03:28.140300Z digest=sha256:e14469139a748f25a4e6aaaa5d936ef54676d05a15ed86366739223c33f17a35

Observation ea005420-8fac-45fb-be45-2bb5acf2547f · outbound

This paper cites Enhancing LLM Reasoning with Reward-guided Tree Search.

From Trial-and-Error to Improvement: A Systematic Analysis of LLM Exploration Mechanisms in RLVR Enhancing LLM Reasoning with Reward-guided Tree Search

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-05T22:03:28.144048Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T22:03:28.144048Z digest=sha256:ec6c002157268b042c945dcd4d6df8259fdaaad4b9e70c6762972e987c7b9cc0

Observation e7e5d277-296f-40f4-95ee-93a817267c03 · outbound

This paper cites Inference-time scaling for generalist reward modeling.

From Trial-and-Error to Improvement: A Systematic Analysis of LLM Exploration Mechanisms in RLVR Inference-time scaling for generalist reward modeling

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-05T22:03:28.147876Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T22:03:28.147876Z digest=sha256:eca34c3d1d593345a3f8b4c8ef2274ba2d34fe95dd237cc622c80625e70612ef

Observation f5b52097-3d64-4337-b9f5-4207d2bca587 · outbound

This paper cites Heimdall: test-time scaling on the generative verification.

From Trial-and-Error to Improvement: A Systematic Analysis of LLM Exploration Mechanisms in RLVR Heimdall: test-time scaling on the generative verification

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-05T22:03:28.151638Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T22:03:28.151638Z digest=sha256:1fd1681c4d8e133b81ec4b5db439a9b3d276a36908ce78dd9d13c3838c9955bc

Observation a5d8a914-54b5-435e-9846-bef7c95d6450 · outbound

This paper cites Towards Effective Code-Integrated Reasoning.

From Trial-and-Error to Improvement: A Systematic Analysis of LLM Exploration Mechanisms in RLVR Towards Effective Code-Integrated Reasoning

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-05T22:03:28.155704Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T22:03:28.155704Z digest=sha256:8e58a641cb8bbf8fec228a8f6074ae15db3333aaac498b8281228f282c7f65f9

Observation 4915f496-4c8b-48cd-a4fb-93f90b371a93 · outbound

This paper cites Beyond the 80/20 Rule: High-Entropy Minority Tokens Drive Effective Reinforcement Learning for LLM Reasoning.

From Trial-and-Error to Improvement: A Systematic Analysis of LLM Exploration Mechanisms in RLVR Beyond the 80/20 Rule: High-Entropy Minority Tokens Drive Effective Reinforcement Learning for LLM Reasoning

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-05T22:03:28.159653Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T22:03:28.159653Z digest=sha256:15fa6d45ba6c7ad9146f3559977e91955ca447f153ad129b51103ad844673692

Observation d8f702a4-4f9d-4541-ab53-b33821fc020f · outbound

This paper cites Stabilizing Knowledge, Promoting Reasoning: Dual-Token Constraints for RLVR.

From Trial-and-Error to Improvement: A Systematic Analysis of LLM Exploration Mechanisms in RLVR Stabilizing Knowledge, Promoting Reasoning: Dual-Token Constraints for RLVR

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-05T22:03:28.164240Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T22:03:28.164240Z digest=sha256:7293990a9f35d208b94d19fb4d61f7c4e85a78d4e0ef4fc126fdbdbc989407b7

Observation 8ab8723e-19b3-466a-9f43-335a577da71b · outbound

This paper cites Not everything is all you need: Toward low-redundant optimization for large language model alignment.

From Trial-and-Error to Improvement: A Systematic Analysis of LLM Exploration Mechanisms in RLVR Not everything is all you need: Toward low-redundant optimization for large language model alignment

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T22:03:28.664973Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-08-05T22:03:28.168268Z digest=sha256:9cdffed554a395ff3737467c2bc5a9a80de546d62bee40f74b46e11ac3348a2e

Observation 14e94b11-8757-47ad-a607-6a1e15eaeee5 · outbound

This paper cites Towards a Human-like Open-Domain Chatbot.

From Trial-and-Error to Improvement: A Systematic Analysis of LLM Exploration Mechanisms in RLVR Towards a Human-like Open-Domain Chatbot

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-05T22:03:28.172143Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T22:03:28.172143Z digest=sha256:80116232ae76ea26e1d172eea834fb552206dde5f06a554ea854cdc631b8e916

Observation 68da4352-c443-4b1c-972e-2e4d88fc095d · outbound

This paper cites Maximizing Confidence Alone Improves Reasoning.

From Trial-and-Error to Improvement: A Systematic Analysis of LLM Exploration Mechanisms in RLVR Maximizing Confidence Alone Improves Reasoning

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-05T22:03:28.176482Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T22:03:28.176482Z digest=sha256:3dc7331274f37a234d49ce2daca9ca6986168517258737343b6678b17cc51d5c

Observation 93f1999f-6b2b-455f-b6bb-c20fc753d306 · outbound

This paper cites Proximal Policy Optimization Algorithms.

From Trial-and-Error to Improvement: A Systematic Analysis of LLM Exploration Mechanisms in RLVR Proximal Policy Optimization Algorithms

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-05T22:03:28.180681Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T22:03:28.180681Z digest=sha256:127e1b350db7b6ee04619cfdffa1cd819cf8ebc88a4d7b7300c16b3e328e8628

Observation 017bbefa-64cf-4014-b67b-1bb109992938 · outbound

This paper cites SFT Memorizes, RL Generalizes: A Comparative Study of Foundation Model Post-training.

From Trial-and-Error to Improvement: A Systematic Analysis of LLM Exploration Mechanisms in RLVR SFT Memorizes, RL Generalizes: A Comparative Study of Foundation Model Post-training

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-05T22:03:28.184746Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T22:03:28.184746Z digest=sha256:275860c81ac1aa91ae0e0758110def754aae30408a2979d820bc069f2ef4fdad

Observation fca777cb-58e4-41ec-8d83-733f6ba0eeca · outbound

This paper cites Agentic Reinforced Policy Optimization.

From Trial-and-Error to Improvement: A Systematic Analysis of LLM Exploration Mechanisms in RLVR Agentic Reinforced Policy Optimization

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-05T22:03:28.188682Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T22:03:28.188682Z digest=sha256:3e92a21fd92123fdc3910d1a03748878ae4c5c66e6651c8510c35086e782391c

Pith citing papers

Observation 9a761aa7-9967-4ef2-81d7-1f3f78de71b1 · inbound

A Survey of Reinforcement Learning for Large Reasoning Models cites this paper.

A Survey of Reinforcement Learning for Large Reasoning Models From Trial-and-Error to Improvement: A Systematic Analysis of LLM Exploration Mechanisms in RLVR

Reference 106

Resolution
verified exact
arxiv_id, observed 2026-05-18T00:02:25.255020Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-18T00:02:24.352947Z digest=sha256:f3c85cd09d01afd56d8d1f4ccb0b692f8d73fb46d2b5d3883c602aeac85e4b7d

Observation ffad19ef-e6cf-4214-9ca7-868e40a51bff · inbound

ResRL: Boosting LLM Reasoning via Negative Sample Projection Residual Reinforcement Learning cites this paper.

ResRL: Boosting LLM Reasoning via Negative Sample Projection Residual Reinforcement Learning From Trial-and-Error to Improvement: A Systematic Analysis of LLM Exploration Mechanisms in RLVR

Reference 62

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T15:41:45.790129Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-09T19:26:57.596581Z digest=sha256:5f2c1f8af77f816895128810d29a038680576d63310f84e98c6b7b05b301a405

Observation 5c712013-bb5f-41fa-9ad7-1ab6b67256b3 · inbound

ResRL: Boosting LLM Reasoning via Negative Sample Projection Residual Reinforcement Learning cites this paper.

ResRL: Boosting LLM Reasoning via Negative Sample Projection Residual Reinforcement Learning From Trial-and-Error to Improvement: A Systematic Analysis of LLM Exploration Mechanisms in RLVR

Reference 62

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T03:55:56.685011Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-11T02:07:21.806345Z digest=sha256:3ac4d3898e09f7b7577ede747d9bd41ee24a42e82cf06371d2537a328b561a8c