Pith. sign in

Paper Citation Record · LEDGER

From Trial-and-Error to Improvement: A Systematic Analysis of LLM Exploration Mechanisms in RLVR

As of 7 August 2026, this Paper Citation Record lists 33 of 33 outbound references and 3 inbound Pith citation observations for arXiv:2508.07534.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.07534 v2

Coverage vector

measured 33 of 33 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T22:03:28.188682Z

measured 36 of 36 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-18T00:02:24.352947Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-18T00:02:25.252362Z

Reference resolution

33 of 33 outbound references displayed

  • verified exact0
  • verified fuzzy5
  • unresolved28
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 322ba5f0-cb9c-49f0-8b58-317bba6f33f2 · outbound

This paper cites Decomposing the Entropy-Performance Exchange: The Missing Keys to Unlocking Effective Reinforcement Learning.

From Trial-and-Error to Improvement: A Systematic Analysis of LLM Exploration Mechanisms in RLVR Decomposing the Entropy-Performance Exchange: The Missing Keys to Unlocking Effective Reinforcement Learning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-05T22:03:28.053726Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T22:03:28.053726Z digest=sha256:47d8f6f7114c69138ce10d7460e44c8ad83a1f0d7d22e98829cd777236340087

Observation 2905d31a-65ad-473e-818e-f86ae11c2154 · outbound

This paper cites A Survey of Large Language Models.

From Trial-and-Error to Improvement: A Systematic Analysis of LLM Exploration Mechanisms in RLVR A Survey of Large Language Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-05T22:03:28.058500Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T22:03:28.058500Z digest=sha256:738c2fcf2734534aaf2aa15d39c89e166cc2c53f052df45776f1a16de7463229

Observation 17e3e04c-cc53-4260-ba52-1bdd66cd14ea · outbound

This paper cites Training language models to follow instructions with human feedback.

From Trial-and-Error to Improvement: A Systematic Analysis of LLM Exploration Mechanisms in RLVR Training language models to follow instructions with human feedback

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-05T22:03:28.063021Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T22:03:28.063021Z digest=sha256:54766fd3e1e88a4fa7dc94f4202d78edc814ef8e3dbe7a85964245b6a9711e72

Observation b421bde8-3511-4eac-92c2-d25e8ddbe07a · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

From Trial-and-Error to Improvement: A Systematic Analysis of LLM Exploration Mechanisms in RLVR DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-05T22:03:28.067351Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T22:03:28.067351Z digest=sha256:7ba9a03aa7e54af51e271ed94a23bff5da8a360ddb8b3ecb8c140fd9a9d1323f

Observation 9624f8e9-1b9d-4e39-bf3b-54a9be2b8fbc · outbound

This paper cites an unresolved cited work.

From Trial-and-Error to Improvement: A Systematic Analysis of LLM Exploration Mechanisms in RLVR Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-05T22:03:28.762818Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-05T22:03:28.071721Z digest=sha256:665fc3545faa2573b84a253db2d8e3ef5dd2fa64b6b1555ab3cb314137858b9a

Observation f4095a71-a846-4d77-bd8c-352760ee80b8 · outbound

This paper cites Tulu 3: Pushing Frontiers in Open Language Model Post-Training.

From Trial-and-Error to Improvement: A Systematic Analysis of LLM Exploration Mechanisms in RLVR Tulu 3: Pushing Frontiers in Open Language Model Post-Training

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-05T22:03:28.075990Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T22:03:28.075990Z digest=sha256:1b9460000681a6468e2bceb3829a54e5272bcc5a9cfd9928efcdb6efe7c499e7

Observation 24bdff8e-b204-4b53-a0f5-c33bd7634669 · outbound

This paper cites Reasoning with Exploration: An Entropy Perspective.

From Trial-and-Error to Improvement: A Systematic Analysis of LLM Exploration Mechanisms in RLVR Reasoning with Exploration: An Entropy Perspective

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-05T22:03:28.081012Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T22:03:28.081012Z digest=sha256:ca8fe8996168669591060c728e05acc41cac4f300878169f686715e82f2d6176

Observation 011fa54c-4c8a-4470-a3b9-79a31d60fabc · outbound

This paper cites Exploration in deep reinforcement learning: A survey.

From Trial-and-Error to Improvement: A Systematic Analysis of LLM Exploration Mechanisms in RLVR Exploration in deep reinforcement learning: A survey

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-05T22:03:28.085936Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T22:03:28.085936Z digest=sha256:af1facce0ed9792eb8187d78638fadd3e5503a74cef0c1f09ae5be898d988504

Observation b4e661c4-d523-4537-b298-07d6ce18cea9 · outbound

This paper cites The Entropy Mechanism of Reinforcement Learning for Reasoning Language Models.

From Trial-and-Error to Improvement: A Systematic Analysis of LLM Exploration Mechanisms in RLVR The Entropy Mechanism of Reinforcement Learning for Reasoning Language Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-05T22:03:28.090177Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T22:03:28.090177Z digest=sha256:7f3dca8003f693173b31e9d55762fae0df44f557eb793b56fac5d90310b21934

Observation 67bce6db-ecda-40bc-9b2e-5cd71a80e58a · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

From Trial-and-Error to Improvement: A Systematic Analysis of LLM Exploration Mechanisms in RLVR DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-05T22:03:28.094332Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T22:03:28.094332Z digest=sha256:9a3593d0de46856ab079b41adb91141f4d4ac262060083b843c15359fe6522d3

Observation d6d691cb-a13d-487d-b724-d22111c95a82 · outbound

This paper cites The surprising effectiveness of negative reinforcement in llm reasoning.

From Trial-and-Error to Improvement: A Systematic Analysis of LLM Exploration Mechanisms in RLVR The surprising effectiveness of negative reinforcement in llm reasoning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-05T22:03:28.098747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T22:03:28.098747Z digest=sha256:9b982d8929d6745f85b3edb0f271dacc6dd55fe68636fe56d91b298e4eaa5472

Observation bd046845-72d6-4737-8501-5894b9ebd3da · outbound

This paper cites Math-shepherd: Verify and reinforce llms step-by-step without human annotations.

From Trial-and-Error to Improvement: A Systematic Analysis of LLM Exploration Mechanisms in RLVR Math-shepherd: Verify and reinforce llms step-by-step without human annotations

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T22:03:28.737764Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-05T22:03:28.103031Z digest=sha256:5f0adf80aa74da57e0078f104ef041e99870c4754e5c02edf8443dc1db87e0a8

Observation c2f1dcbc-69c2-4cb8-bdc3-a75f9d9db686 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

From Trial-and-Error to Improvement: A Systematic Analysis of LLM Exploration Mechanisms in RLVR DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-05T22:03:28.107150Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T22:03:28.107150Z digest=sha256:7ddafe4ed74c036deee705e4de696edf2e8b65c1fc3f76e84e61f2e6c9360754

Observation 50c3b81c-2c61-4c63-8326-dde21d646ebe · outbound

This paper cites Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?.

From Trial-and-Error to Improvement: A Systematic Analysis of LLM Exploration Mechanisms in RLVR Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-05T22:03:28.111517Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T22:03:28.111517Z digest=sha256:dead0324b36753dc14fd13447d8e785617b3ff369afea31c74687d45b8abef4d

Observation 0b894881-4190-409d-9d68-3950e4797a83 · outbound

This paper cites ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models.

From Trial-and-Error to Improvement: A Systematic Analysis of LLM Exploration Mechanisms in RLVR ProRL: Prolonged Reinforcement Learning Expands Reasoning Boundaries in Large Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-05T22:03:28.116441Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T22:03:28.116441Z digest=sha256:5614910bd147138d4977809a962a726fc6787241f82ab99ede3ca03b5058e474

Observation 55ca83fb-bc87-4b29-8023-c78c84d38c92 · outbound

This paper cites Evaluating Large Language Models Trained on Code.

From Trial-and-Error to Improvement: A Systematic Analysis of LLM Exploration Mechanisms in RLVR Evaluating Large Language Models Trained on Code

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-05T22:03:28.120627Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T22:03:28.120627Z digest=sha256:f5409ce58224fe13ff66616f49134cdc7c2d6eff4eab8fcebfd2880d10b7271e

Observation c40f1373-96d7-47d6-bf88-a85fa7fb7c7d · outbound

This paper cites Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models.

From Trial-and-Error to Improvement: A Systematic Analysis of LLM Exploration Mechanisms in RLVR Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-05T22:03:28.124820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T22:03:28.124820Z digest=sha256:f670ea598b20ad03392eb64af5f297fc3faee5c549b34318b7a139022c838b37

Observation 55958e60-b658-4583-af0f-e5401c34048b · outbound

This paper cites an unresolved cited work.

From Trial-and-Error to Improvement: A Systematic Analysis of LLM Exploration Mechanisms in RLVR Unresolved cited work

Reference 18

Resolution
unresolved
raw_fallback, observed 2026-08-05T22:03:28.722600Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-05T22:03:28.128868Z digest=sha256:f48156707e90f5b131da0b88125c52883c1eea93af998541799cf4274223c1cc

Observation a748fea4-02d5-468e-bc1d-8767a55ab860 · outbound

This paper cites Let's verify step by step.

From Trial-and-Error to Improvement: A Systematic Analysis of LLM Exploration Mechanisms in RLVR Let's verify step by step

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T22:03:28.709392Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-05T22:03:28.132945Z digest=sha256:1248d610aaaeaf082d2cdbbcc569ea33fe98e6ba662bc5ce4f6cae546e3df08e

Observation 54180ce1-e458-43c7-9ce9-475e1ca386df · outbound

This paper cites Generative verifiers: Reward modeling as next-token prediction.

From Trial-and-Error to Improvement: A Systematic Analysis of LLM Exploration Mechanisms in RLVR Generative verifiers: Reward modeling as next-token prediction

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T22:03:28.693699Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-05T22:03:28.136547Z digest=sha256:6adb8eb7de478afda9a2ce91f028717aa33ff2f74d813a9205bb7f1d4e8b7f2a

Observation 162588cb-f6b0-4e4e-8c70-8508c9f4bc9b · outbound

This paper cites The lessons of developing process reward models in mathematical reasoning.

From Trial-and-Error to Improvement: A Systematic Analysis of LLM Exploration Mechanisms in RLVR The lessons of developing process reward models in mathematical reasoning

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T22:03:28.679544Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-05T22:03:28.140300Z digest=sha256:093ee91ca613f97d0ebf7991cb71193ab718d29c3153978b6d3a63782924475d

Observation ea005420-8fac-45fb-be45-2bb5acf2547f · outbound

This paper cites Enhancing LLM Reasoning with Reward-guided Tree Search.

From Trial-and-Error to Improvement: A Systematic Analysis of LLM Exploration Mechanisms in RLVR Enhancing LLM Reasoning with Reward-guided Tree Search

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-05T22:03:28.144048Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T22:03:28.144048Z digest=sha256:f6eaa290e384168d7ba77fbe750bf264c97bc31c0c67b5ad6a0d411c511268ee

Observation e7e5d277-296f-40f4-95ee-93a817267c03 · outbound

This paper cites Inference-time scaling for generalist reward modeling.

From Trial-and-Error to Improvement: A Systematic Analysis of LLM Exploration Mechanisms in RLVR Inference-time scaling for generalist reward modeling

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-05T22:03:28.147876Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T22:03:28.147876Z digest=sha256:eca34c3d1d593345a3f8b4c8ef2274ba2d34fe95dd237cc622c80625e70612ef

Observation f5b52097-3d64-4337-b9f5-4207d2bca587 · outbound

This paper cites Heimdall: test-time scaling on the generative verification.

From Trial-and-Error to Improvement: A Systematic Analysis of LLM Exploration Mechanisms in RLVR Heimdall: test-time scaling on the generative verification

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-05T22:03:28.151638Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T22:03:28.151638Z digest=sha256:1fd1681c4d8e133b81ec4b5db439a9b3d276a36908ce78dd9d13c3838c9955bc

Observation a5d8a914-54b5-435e-9846-bef7c95d6450 · outbound

This paper cites Towards Effective Code-Integrated Reasoning.

From Trial-and-Error to Improvement: A Systematic Analysis of LLM Exploration Mechanisms in RLVR Towards Effective Code-Integrated Reasoning

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-05T22:03:28.155704Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T22:03:28.155704Z digest=sha256:a6509c6b123e0e08d3fe487d3d0809d09a74c3689d19a53fe7c487a2281739b7

Observation 4915f496-4c8b-48cd-a4fb-93f90b371a93 · outbound

This paper cites Beyond the 80/20 Rule: High-Entropy Minority Tokens Drive Effective Reinforcement Learning for LLM Reasoning.

From Trial-and-Error to Improvement: A Systematic Analysis of LLM Exploration Mechanisms in RLVR Beyond the 80/20 Rule: High-Entropy Minority Tokens Drive Effective Reinforcement Learning for LLM Reasoning

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-05T22:03:28.159653Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T22:03:28.159653Z digest=sha256:15fa6d45ba6c7ad9146f3559977e91955ca447f153ad129b51103ad844673692

Observation d8f702a4-4f9d-4541-ab53-b33821fc020f · outbound

This paper cites Stabilizing Knowledge, Promoting Reasoning: Dual-Token Constraints for RLVR.

From Trial-and-Error to Improvement: A Systematic Analysis of LLM Exploration Mechanisms in RLVR Stabilizing Knowledge, Promoting Reasoning: Dual-Token Constraints for RLVR

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-05T22:03:28.164240Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T22:03:28.164240Z digest=sha256:222e174a795ed2f4c6b8e0e74a8abea1d709252d7eb0f2fc4a9c554838c76f37

Observation 8ab8723e-19b3-466a-9f43-335a577da71b · outbound

This paper cites Not everything is all you need: Toward low-redundant optimization for large language model alignment.

From Trial-and-Error to Improvement: A Systematic Analysis of LLM Exploration Mechanisms in RLVR Not everything is all you need: Toward low-redundant optimization for large language model alignment

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T22:03:28.664973Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-05T22:03:28.168268Z digest=sha256:7608e8d6935f9e74f8a61fa3fcffcddbcd6b9afdac0744c1b2460f0398316a68

Observation 14e94b11-8757-47ad-a607-6a1e15eaeee5 · outbound

This paper cites Towards a Human-like Open-Domain Chatbot.

From Trial-and-Error to Improvement: A Systematic Analysis of LLM Exploration Mechanisms in RLVR Towards a Human-like Open-Domain Chatbot

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-05T22:03:28.172143Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T22:03:28.172143Z digest=sha256:80116232ae76ea26e1d172eea834fb552206dde5f06a554ea854cdc631b8e916

Observation 68da4352-c443-4b1c-972e-2e4d88fc095d · outbound

This paper cites Maximizing Confidence Alone Improves Reasoning.

From Trial-and-Error to Improvement: A Systematic Analysis of LLM Exploration Mechanisms in RLVR Maximizing Confidence Alone Improves Reasoning

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-05T22:03:28.176482Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T22:03:28.176482Z digest=sha256:39d80f3f165a312259cf8cb2ab44b4f1dd564b514df1214fe0f34a0592b16861

Observation 93f1999f-6b2b-455f-b6bb-c20fc753d306 · outbound

This paper cites Proximal Policy Optimization Algorithms.

From Trial-and-Error to Improvement: A Systematic Analysis of LLM Exploration Mechanisms in RLVR Proximal Policy Optimization Algorithms

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-05T22:03:28.180681Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T22:03:28.180681Z digest=sha256:127e1b350db7b6ee04619cfdffa1cd819cf8ebc88a4d7b7300c16b3e328e8628

Observation 017bbefa-64cf-4014-b67b-1bb109992938 · outbound

This paper cites SFT Memorizes, RL Generalizes: A Comparative Study of Foundation Model Post-training.

From Trial-and-Error to Improvement: A Systematic Analysis of LLM Exploration Mechanisms in RLVR SFT Memorizes, RL Generalizes: A Comparative Study of Foundation Model Post-training

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-05T22:03:28.184746Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T22:03:28.184746Z digest=sha256:11755a790445be32403200429a0dbee3f6b6c01beaad3d02337b1aebb4a90a92

Observation fca777cb-58e4-41ec-8d83-733f6ba0eeca · outbound

This paper cites Agentic Reinforced Policy Optimization.

From Trial-and-Error to Improvement: A Systematic Analysis of LLM Exploration Mechanisms in RLVR Agentic Reinforced Policy Optimization

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-05T22:03:28.188682Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T22:03:28.188682Z digest=sha256:3e92a21fd92123fdc3910d1a03748878ae4c5c66e6651c8510c35086e782391c

Pith citing papers

Observation 9a761aa7-9967-4ef2-81d7-1f3f78de71b1 · inbound

A Survey of Reinforcement Learning for Large Reasoning Models cites this paper.

A Survey of Reinforcement Learning for Large Reasoning Models From Trial-and-Error to Improvement: A Systematic Analysis of LLM Exploration Mechanisms in RLVR

Reference 106

Resolution
verified exact
arxiv_id, observed 2026-05-18T00:02:25.255020Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-18T00:02:24.352947Z digest=sha256:e3a4968f8bdd71542591caab37b0c6a9e2f741f186288c8bd5495b17a4563622

Observation ffad19ef-e6cf-4214-9ca7-868e40a51bff · inbound

ResRL: Boosting LLM Reasoning via Negative Sample Projection Residual Reinforcement Learning cites this paper.

ResRL: Boosting LLM Reasoning via Negative Sample Projection Residual Reinforcement Learning From Trial-and-Error to Improvement: A Systematic Analysis of LLM Exploration Mechanisms in RLVR

Reference 62

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T15:41:45.790129Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-09T19:26:57.596581Z digest=sha256:2592e278d3ae66fe46f06584a05e9a2389cdb837276fe32a89df2eb8ead48afc

Observation 5c712013-bb5f-41fa-9ad7-1ab6b67256b3 · inbound

ResRL: Boosting LLM Reasoning via Negative Sample Projection Residual Reinforcement Learning cites this paper.

ResRL: Boosting LLM Reasoning via Negative Sample Projection Residual Reinforcement Learning From Trial-and-Error to Improvement: A Systematic Analysis of LLM Exploration Mechanisms in RLVR

Reference 62

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T03:55:56.685011Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-11T02:07:21.806345Z digest=sha256:095d23fede220eef9667df5043abc382d745b1b9e221f2a715d7f5d3855e497c