Pith. sign in

Paper Citation Record · LEDGER

First Return, Entropy-Eliciting Explore

As of 9 August 2026, this Paper Citation Record lists 42 of 42 outbound references and 18 inbound Pith citation observations for arXiv:2507.07017.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.07017 v1

Coverage vector

measured 42 of 42 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T18:53:15.153126Z

measured 60 of 60 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 18 of 18 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T05:53:52.768850Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T09:09:43.539228Z

Reference resolution

42 of 42 outbound references displayed

  • verified exact0
  • verified fuzzy7
  • unresolved34
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 9467682a-6022-4bbf-91ce-59a20c95f063 · outbound

This paper cites Back to Basics: Revisiting REINFORCE Style Optimization for Learning from Human Feedback in LLMs.

First Return, Entropy-Eliciting Explore Back to Basics: Revisiting REINFORCE Style Optimization for Learning from Human Feedback in LLMs

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T18:53:11.406993Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:53:11.406993Z digest=sha256:25bfd335ff341a1ccc821c25044687f734054a6c4ec2785d10f5ae735ea76e21

Observation 1f96a7a7-cda7-4f5d-b81e-71a4897833e8 · outbound

This paper cites Using confidence bounds for exploitation-exploration trade-offs.JournalofMachineLearningResearch, 3(Nov):397–422, 2002.

First Return, Entropy-Eliciting Explore Using confidence bounds for exploitation-exploration trade-offs.JournalofMachineLearningResearch, 3(Nov):397–422, 2002

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:53:17.295629Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T18:53:11.509061Z digest=sha256:1bce120caf10c7ade571297ff6d71c9d125245e6656c394249d6c6b8d7ce6d11

Observation da3b65e4-0336-4c60-adc0-081ee30abc4b · outbound

This paper cites Finite-time analysis of the multiarmed bandit problem.

First Return, Entropy-Eliciting Explore Finite-time analysis of the multiarmed bandit problem

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T18:53:11.640190Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:53:11.640190Z digest=sha256:61162796fdd4eed58425c2e1d6d3a73d97e9c6442b3da67e82dc00e428eb1332

Observation f1a73968-6880-4e1f-955d-dbbf5321f161 · outbound

This paper cites Language models are few-shot learners.

First Return, Entropy-Eliciting Explore Language models are few-shot learners

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T18:53:11.762655Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:53:11.762655Z digest=sha256:4cebd62d469b125caf0197036316aa025f62465918cbbcf640da20e9492a80d9

Observation d58235a4-d96e-4d0e-888e-9f2403f8ca6c · outbound

This paper cites Do NOT Think That Much for 2+3=? On the Overthinking of o1-Like LLMs.

First Return, Entropy-Eliciting Explore Do NOT Think That Much for 2+3=? On the Overthinking of o1-Like LLMs

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T18:53:11.886340Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:53:11.886340Z digest=sha256:e35cbca5b73d447cfd869b1c8e297b7109fe828c95558be127e81dbcf651d31c

Observation 90fa350c-2be6-4f5b-9662-3d9b8c654482 · outbound

This paper cites Process Reinforcement through Implicit Rewards.

First Return, Entropy-Eliciting Explore Process Reinforcement through Implicit Rewards

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T18:53:11.982084Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:53:11.982084Z digest=sha256:88a291a9ffcd19d683477a499f427624ecc2aff9cacbe4751ad5e23f7a2f3e64

Observation 8df3f110-f943-4012-b73c-4441fdc1fff3 · outbound

This paper cites Assessing Validity of ICD-10 Administrative Data in Coding Comorbidities.

First Return, Entropy-Eliciting Explore Assessing Validity of ICD-10 Administrative Data in Coding Comorbidities

Reference 7

Resolution
metadata mismatch
local_arxiv, observed 2026-08-06T18:53:15.752569Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T18:53:12.073302Z digest=sha256:b32e3fc6af6c4fe0c3d42d5927437144754596d6eeaee3539b78acf6a56038a7

Observation 0bfd072f-e7fb-4c9a-b297-695dacd32624 · outbound

This paper cites S-GRPO: Early Exit via Reinforcement Learning in Reasoning Models.

First Return, Entropy-Eliciting Explore S-GRPO: Early Exit via Reinforcement Learning in Reasoning Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T18:53:12.135771Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:53:12.135771Z digest=sha256:e4a95a58859dc3f6d323ec5916e5e239183cbb78ded356014fb34874511d2411

Observation b2a8d39d-42a5-4abe-bf7d-575a39e8159f · outbound

This paper cites Go-Explore: a New Approach for Hard-Exploration Problems.

First Return, Entropy-Eliciting Explore Go-Explore: a New Approach for Hard-Exploration Problems

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T18:53:12.229578Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:53:12.229578Z digest=sha256:e3fa2fd457200ab30496130ace6ddfb7a63d7ef1a6abf28384d999820ff13c22

Observation dd151a39-067f-4f20-bc7f-7a61b8292e5d · outbound

This paper cites Stanley, and Jeff Clune.

First Return, Entropy-Eliciting Explore Stanley, and Jeff Clune

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T18:53:12.302180Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:53:12.302180Z digest=sha256:7882e6dc790776e68df319e6d15303a8dd837ed0d1ea19806d7e1628b5bf49a8

Observation b2697bc8-ad10-49f3-ac3c-cd673bc6c02c · outbound

This paper cites A Survey on Mathematical Reasoning and Optimization with Large Language Models.

First Return, Entropy-Eliciting Explore A Survey on Mathematical Reasoning and Optimization with Large Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T18:53:12.361332Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:53:12.361332Z digest=sha256:77595e0dfbabf8c1ed40a9f3f2cf74d90bbe8d3137bcb558e95c710dbe5bfb21

Observation 2b4d0b61-8c05-4d5c-b784-da600d060f2b · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

First Return, Entropy-Eliciting Explore DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T18:53:12.465006Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:53:12.465006Z digest=sha256:f6d385dd7e91df18b5aaa8396ff2eef95e7141e859343baa773ceafa3a65ed71

Observation 031a95bc-c452-4015-86f6-bdf98fce8502 · outbound

This paper cites Does Math Reasoning Improve General LLM Capabilities? Understanding Transferability of LLM Reasoning.

First Return, Entropy-Eliciting Explore Does Math Reasoning Improve General LLM Capabilities? Understanding Transferability of LLM Reasoning

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T18:53:12.561980Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:53:12.561980Z digest=sha256:cfe9e47d02fefd1902e2ba0fe8a9f2c88453bbb94abca8dd1e99108f53f53b3d

Observation 4fcc9aa5-d6d3-420c-90c1-532a17c17fe1 · outbound

This paper cites VinePPO: Refining Credit Assignment in RL Training of LLMs.

First Return, Entropy-Eliciting Explore VinePPO: Refining Credit Assignment in RL Training of LLMs

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T18:53:12.630333Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:53:12.630333Z digest=sha256:0d7b8187caf7e6ef5818feef578f33fd76c19c4eeba6e08c466a5e756681d8ef

Observation 89d42913-2226-4634-b56e-05c1653ec6c8 · outbound

This paper cites Temporal Sampling for Forgotten Reasoning in LLMs.

First Return, Entropy-Eliciting Explore Temporal Sampling for Forgotten Reasoning in LLMs

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T18:53:12.730408Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:53:12.730408Z digest=sha256:36dd0282ee3bc4b4407195ca3375ae8f01e869a9ca8f86a84d2502b782201a1a

Observation 12059643-acbf-418a-8c11-f1bc77026d36 · outbound

This paper cites Let’s verify step by step, 2023.

First Return, Entropy-Eliciting Explore Let’s verify step by step, 2023

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T18:53:12.795404Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:53:12.795404Z digest=sha256:7f9056451cf115e465b75916d10fc32c789b9e4362d7839147deb3c838297a31

Observation 41c2f940-5cb4-4352-8f6d-a1c4c6f0d3b6 · outbound

This paper cites Intelligent Go-Explore: Standing on the Shoulders of Giant Foundation Models.

First Return, Entropy-Eliciting Explore Intelligent Go-Explore: Standing on the Shoulders of Giant Foundation Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T18:53:12.860423Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:53:12.860423Z digest=sha256:a6118df38a0ded8ef32515dc000bda3d31093fe8365a231f149c44ccf9370c63

Observation cf567e20-b6ab-4828-be67-e36d925e32ff · outbound

This paper cites Improve Mathematical Reasoning in Language Models by Automated Process Supervision.

First Return, Entropy-Eliciting Explore Improve Mathematical Reasoning in Language Models by Automated Process Supervision

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T18:53:12.924944Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:53:12.924944Z digest=sha256:9c288411208770c784119f6508946bd8ef130b9ba5a0344b07fe00cc3eb66e7a

Observation 32634212-fcd9-4925-92e5-2cda5e921854 · outbound

This paper cites Tang, Manan Roongta, Colin Cai, Jeffrey Luo, Li Erran Li, Raluca Ada Popa, and Ion Stoica.

First Return, Entropy-Eliciting Explore Tang, Manan Roongta, Colin Cai, Jeffrey Luo, Li Erran Li, Raluca Ada Popa, and Ion Stoica

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T18:53:13.035225Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:53:13.035225Z digest=sha256:e571f43c147b247e775a230bbfd055b22aba684d6c4e5dbf240b6b5c4a02c36a

Observation 58b41610-eaed-4414-b7d1-3c3323f7478a · outbound

This paper cites MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention.

First Return, Entropy-Eliciting Explore MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T18:53:13.146511Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:53:13.146511Z digest=sha256:a27df62583775d36993b9e4f6181fde2825542e18f9464e0fa403e68ff4c3924

Observation 70828079-b558-4591-b755-4a9b04af1bef · outbound

This paper cites Training language models to follow instructions with human feedback.

First Return, Entropy-Eliciting Explore Training language models to follow instructions with human feedback

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:53:17.162881Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T18:53:13.202720Z digest=sha256:647c0577d86161b29b282bdbced8b91c2e72c577e29f9d9e69c4d6208460e2fa

Observation 3d113656-ccc7-4533-8eef-555599ea29de · outbound

This paper cites A Survey of Temporal Credit Assignment in Deep Reinforcement Learning.

First Return, Entropy-Eliciting Explore A Survey of Temporal Credit Assignment in Deep Reinforcement Learning

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T18:53:13.287355Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:53:13.287355Z digest=sha256:ec156ff25f2a1560ea00f296bc6415806b33c0d520a4b187d30dd27d1f61359c

Observation 472d729b-e22a-4390-ae03-d9cb47d9f45b · outbound

This paper cites Qwen2.5 Technical Report.

First Return, Entropy-Eliciting Explore Qwen2.5 Technical Report

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T18:53:13.385918Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:53:13.385918Z digest=sha256:4835a5410c0b32696b33aa77c0fd23f7ac406e9b4a6db68ae75dd4c9c8f03bb6

Observation d8d0e84f-522b-42c4-a44d-5483af15b854 · outbound

This paper cites Sequence level training with recurrent neural networks, 2016.

First Return, Entropy-Eliciting Explore Sequence level training with recurrent neural networks, 2016

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:53:17.005806Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T18:53:13.466582Z digest=sha256:d87dd25830dc7312ee8808e8d8527dc5d238abf7b66dc2a1850ab7c34324467b

Observation 04ff86ca-7e0c-4c22-8e62-472b22cd7c6a · outbound

This paper cites Proximal Policy Optimization Algorithms.

First Return, Entropy-Eliciting Explore Proximal Policy Optimization Algorithms

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T18:53:13.508181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:53:13.508181Z digest=sha256:d7bab74d3cbdd28a23691570ee3d57f8a1227ab446bb673dbc90a63126045f89

Observation 8be6ceca-94e6-444c-b349-47ec76ac118e · outbound

This paper cites High-Dimensional Continuous Control Using Generalized Advantage Estimation.

First Return, Entropy-Eliciting Explore High-Dimensional Continuous Control Using Generalized Advantage Estimation

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T18:53:13.561109Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:53:13.561109Z digest=sha256:26e7656c30719e163f5957ce14145afa331600c78ba9cb4c3682f0a402888ced

Observation 32a9bdeb-ac58-431b-a770-16c45ba907d3 · outbound

This paper cites Rewarding progress: Scaling automated process verifiers for llm reasoning,.

First Return, Entropy-Eliciting Explore Rewarding progress: Scaling automated process verifiers for llm reasoning,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:53:16.854979Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T18:53:13.648906Z digest=sha256:80aa213a54c05b0e0cee23059cc4414240b890d4a9dca653a7b58920e8cba807

Observation 12d153fa-9aaa-46a8-8017-aa1bf075aa24 · outbound

This paper cites Spurious rewards: Rethinking training signals in rlvr.

First Return, Entropy-Eliciting Explore Spurious rewards: Rethinking training signals in rlvr

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:53:16.702794Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T18:53:13.818787Z digest=sha256:b329c9b32b42da9b694e8b3597bca5ee62d4812ef8b8369f686ecbc76f4301e8

Observation b407463e-99a5-4f64-b345-a798cb0a5cf1 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

First Return, Entropy-Eliciting Explore DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T18:53:13.908046Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:53:13.908046Z digest=sha256:ae4bac7f171a300fb39ac010d5fb250526a6fd9235137ff2e6792d13c3dcbc6e

Observation 1628ed8d-f1db-4d53-8bdc-f2b7e54f2dc4 · outbound

This paper cites Hybridflow: A flexible and efficient rlhf framework.

First Return, Entropy-Eliciting Explore Hybridflow: A flexible and efficient rlhf framework

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T18:53:13.998381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:53:13.998381Z digest=sha256:cb9c5c6e496b672df0ae361043f2736edfc482b9b4e90bf108bd3321122d33b9

Observation 20959a98-1f1f-48db-88a7-50c6fca993bd · outbound

This paper cites MIT press Cambridge, 1998.

First Return, Entropy-Eliciting Explore MIT press Cambridge, 1998

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T18:53:14.057085Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:53:14.057085Z digest=sha256:2df69f26b5378fb7917f511d9948a8986771dc8945f7d2654bb2201fd6b1bfb9

Observation f561fe24-6207-4082-83ee-b7881fc157b1 · outbound

This paper cites Between mdps and semi-mdps: A framework for temporal abstraction in reinforcement learning.Artificial intelligence, 112(1–2):181–211, 1999.

First Return, Entropy-Eliciting Explore Between mdps and semi-mdps: A framework for temporal abstraction in reinforcement learning.Artificial intelligence, 112(1–2):181–211, 1999

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:53:16.366032Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T18:53:14.077103Z digest=sha256:49e46f8ce3cce2005d58d0b16717981ca4094710fee8fc447312a001d37d8808

Observation adba242a-f3f7-4a2a-86bf-1beaef2e5d1e · outbound

This paper cites Beyond the 80/20 rule: High-entropy minority tokens drive effective reinforcement learning for llm reasoning,.

First Return, Entropy-Eliciting Explore Beyond the 80/20 rule: High-entropy minority tokens drive effective reinforcement learning for llm reasoning,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:53:16.046672Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T18:53:14.168794Z digest=sha256:7d60691aafc417f66abd1b5336344b7f61672eae061f939578db7a60f05f84fb

Observation 8d1700c0-2ca1-4f5b-9e34-db0a7d90a11c · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models.Advancesin neural information processing systems, 35:24824–24837, 2022.

First Return, Entropy-Eliciting Explore Chain-of-thought prompting elicits reasoning in large language models.Advancesin neural information processing systems, 35:24824–24837, 2022

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T18:53:14.380943Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:53:14.380943Z digest=sha256:c725688cab52906835a460511a0a6d31d37ba617e720aac80348cf28148d827e

Observation e43dfba1-b086-4b45-ab7b-4d162700aceb · outbound

This paper cites A Comparative Study on Reasoning Patterns of OpenAI's o1 Model.

First Return, Entropy-Eliciting Explore A Comparative Study on Reasoning Patterns of OpenAI's o1 Model

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T18:53:14.480574Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:53:14.480574Z digest=sha256:25f307011e9114ce36e14ccf82f78860e73d69c968fabb06eed8f65d8b436760

Observation 1537cf48-c20c-4d91-8a72-da3bbe84d475 · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

First Return, Entropy-Eliciting Explore DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T18:53:14.618359Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:53:14.618359Z digest=sha256:0cff91056f8de6faa9ae41b73c2d028b4374a32b13cecaa5c61a4d35f2268f4a

Observation 2bcd14f6-b579-4a79-ae0e-9ebfb33ca07c · outbound

This paper cites VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks.

First Return, Entropy-Eliciting Explore VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T18:53:14.726477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:53:14.726477Z digest=sha256:d8a13b101534d9a3bc0e5374342a69e968bb0e445b6598d37c8203e06de0203c

Observation b8755a62-5c3b-40d9-a1db-1e0eec0da0c5 · outbound

This paper cites SimpleRL-Zoo: Investigating and Taming Zero Reinforcement Learning for Open Base Models in the Wild.

First Return, Entropy-Eliciting Explore SimpleRL-Zoo: Investigating and Taming Zero Reinforcement Learning for Open Base Models in the Wild

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T18:53:14.864096Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:53:14.864096Z digest=sha256:44874790c61ad460ce5600e75eaf27256a5bbce0ac7827c08090703ea8e61cfa

Observation 588b069e-cf15-4971-b6ae-7fdd8a2ffc51 · outbound

This paper cites SRPO: A Cross-Domain Implementation of Large-Scale Reinforcement Learning on LLM.

First Return, Entropy-Eliciting Explore SRPO: A Cross-Domain Implementation of Large-Scale Reinforcement Learning on LLM

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T18:53:15.010459Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:53:15.010459Z digest=sha256:321e8eef63d87c889c2fd6abda118fd1bb6af406b7d09c5d4adf2a66dacfb316

Observation 4534fe21-53e4-4202-91e8-dbccaba797ea · outbound

This paper cites Least-to-Most Prompting Enables Complex Reasoning in Large Language Models.

First Return, Entropy-Eliciting Explore Least-to-Most Prompting Enables Complex Reasoning in Large Language Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T18:53:15.153126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:53:15.153126Z digest=sha256:84796e9c3f4fb5433f9edcaabc3bfe4a86ef0c805a7168989a91e94466eae430

Observation d8a66a00-6f23-4a3f-ad4f-454fd58ab2dd · outbound

This paper cites Rewarding Progress: Scaling Automated Process Verifiers for LLM Reasoning.

First Return, Entropy-Eliciting Explore Rewarding Progress: Scaling Automated Process Verifiers for LLM Reasoning

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-06T18:53:13.735965Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:53:13.735965Z digest=sha256:9b0fe883f50467f0bc96022d0145f1e6353cd83118ed446a78b20462ef670059

Observation 41b217b6-13f5-40da-9b5e-1d3ab295520a · outbound

This paper cites Beyond the 80/20 Rule: High-Entropy Minority Tokens Drive Effective Reinforcement Learning for LLM Reasoning.

First Return, Entropy-Eliciting Explore Beyond the 80/20 Rule: High-Entropy Minority Tokens Drive Effective Reinforcement Learning for LLM Reasoning

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-06T18:53:14.276825Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:53:14.276825Z digest=sha256:1741cd102d470d36735b84d2fa41eaaefccca9ca1f02ed91a720e91d0f2480e4

Pith citing papers

Observation 7d8e9ecb-8f18-4653-b893-fc7ccde05aad · inbound

Know When to Explore: Difficulty-Aware Certainty as a Guide for LLM Reinforcement Learning cites this paper.

Know When to Explore: Difficulty-Aware Certainty as a Guide for LLM Reinforcement Learning First Return, Entropy-Eliciting Explore

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-05T14:23:54.061991Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:23:54.061991Z digest=sha256:3cebfb94a5b8929809553c5e24b280342dfc84c291f9dbf64415b2fb9d4f7f12

Observation 14b2af2a-9784-420d-848f-0382d14a4838 · inbound

Outcome-based Exploration for LLM Reasoning cites this paper.

Outcome-based Exploration for LLM Reasoning First Return, Entropy-Eliciting Explore

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-04T22:59:14.604039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:59:14.604039Z digest=sha256:3278be0c1336f514ffb564af3c57fe4a958f0c8de5c7cb38fc9a01db203257fc

Observation 31883257-a409-49e5-8385-e1ea6eae40f8 · inbound

Self-Reflective Generation at Test Time cites this paper.

Self-Reflective Generation at Test Time First Return, Entropy-Eliciting Explore

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-04T12:41:43.277273Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T12:41:43.277273Z digest=sha256:18a141e22d772ef789e14c804804980452901f8fbd8d0dae9f11cb4e8f248cdd

Observation 3e0d01b1-8879-442b-b095-39dc436c58a7 · inbound

Don't Tell the Answer, Truly Guide the Reasoning During RL Rollouts cites this paper.

Don't Tell the Answer, Truly Guide the Reasoning During RL Rollouts First Return, Entropy-Eliciting Explore

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-04T10:39:29.867992Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:39:29.867992Z digest=sha256:b2d375594e48a06c5fcdcf95eae7549f9f90fd90d4c481cf160c1394bd83ff64

Observation eaa1a40d-e8fb-4d0d-a1a4-f25b78d4160c · inbound

The Cancellation Hypothesis in Critic-Free RL: From Outcome Rewards to Token Credits cites this paper.

The Cancellation Hypothesis in Critic-Free RL: From Outcome Rewards to Token Credits First Return, Entropy-Eliciting Explore

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-12T08:01:29.024124Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-12T01:24:03.186413Z digest=sha256:27fc8b8f50d7193ccdbb40a3e768a73dd1d1a5d889bd575eeb2e93abf2789bfd

Observation f373a8cc-ce6f-4c16-8b6d-3b79152a151b · inbound

How You Begin is How You Reason: Driving Exploration in RLVR via Prefix-Tuned Priors cites this paper.

How You Begin is How You Reason: Driving Exploration in RLVR via Prefix-Tuned Priors First Return, Entropy-Eliciting Explore

Reference 52

Resolution
verified exact
arxiv_id, observed 2026-05-12T03:26:19.362669Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-12T03:25:04.955816Z digest=sha256:4f45992896d7c55f8faa410daf90efea0de6492bee251246b2a7101fde75ac7a

Observation 5f689fb8-57bb-4870-a57f-205763537f54 · inbound

Selective Off-Policy Reference Tuning with Plan Guidance cites this paper.

Selective Off-Policy Reference Tuning with Plan Guidance First Return, Entropy-Eliciting Explore

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-13T01:47:05.208530Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-13T01:28:18.615371Z digest=sha256:bb84a0a857b9776ceeef7e3efa1982eb70456d45e8f5384c4972d3baf24439eb

Observation 6617a73d-f319-48fa-873a-3f4d3d1feda7 · inbound

Selective Off-Policy Reference Tuning with Plan Guidance cites this paper.

Selective Off-Policy Reference Tuning with Plan Guidance First Return, Entropy-Eliciting Explore

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-14T21:22:59.327594Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-14T21:20:24.066520Z digest=sha256:4f9a9a3bb20b5208aa85f839f4c5161a5bdc6ecd7c5bedf72f7a2d43e0b69272

Observation 6962d0cf-3be1-46a4-9016-cd6295a98b42 · inbound

Reasoning Can Be Restored by Correcting a Few Decision Tokens cites this paper.

Reasoning Can Be Restored by Correcting a Few Decision Tokens First Return, Entropy-Eliciting Explore

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-19T20:57:46.813786Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-19T20:56:48.771058Z digest=sha256:e0ceb78d2d04eb09d045160c7281896f22e04947f43af085e0f301eb8c03dc24

Observation d9099c6b-26ae-4bd5-b61a-9248ff6573f5 · inbound

Trust Region On-Policy Distillation cites this paper.

Trust Region On-Policy Distillation First Return, Entropy-Eliciting Explore

Reference 132

Resolution
verified exact
arxiv_id, observed 2026-07-01T20:56:13.650396Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-28T17:38:50.313305Z digest=sha256:ab794bf5f4466c457ea30f6aa490256b355228a6c1678f0d2e9631b351dc21a7

Observation 024e228f-8c3a-40b3-93a7-ae7df8bbde22 · inbound

On Advantage Estimates for Max@K Policy Gradients cites this paper.

On Advantage Estimates for Max@K Policy Gradients First Return, Entropy-Eliciting Explore

Reference 70

Resolution
verified exact
arxiv_id, observed 2026-07-02T12:06:56.398549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-28T02:21:57.143016Z digest=sha256:c6a4e791b2bfac4377cfccb4e0acd164f94e9ec7a9a5198a614d36067c5bedf9

Observation 5523e626-d46b-4a60-8e80-73862cbeafab · inbound

OrderGrad: Optimizing Beyond the Mean with Order-Statistic Policy Gradient Estimation cites this paper.

OrderGrad: Optimizing Beyond the Mean with Order-Statistic Policy Gradient Estimation First Return, Entropy-Eliciting Explore

Reference 114

Resolution
verified exact
arxiv_id, observed 2026-07-02T12:16:56.967524Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-28T02:17:30.974692Z digest=sha256:d69a96820c9bdc3675b2e4d91ae035352091be884e036e96455333f08971e611

Observation 1254643d-d89f-4eb6-9f53-d3e6c4f518c3 · inbound

From Reasoning Traces to Reusable Modules: Understanding Compositional Generalization in Language Model Reasoning cites this paper.

From Reasoning Traces to Reusable Modules: Understanding Compositional Generalization in Language Model Reasoning First Return, Entropy-Eliciting Explore

Reference 123

Resolution
verified exact
arxiv_id, observed 2026-07-03T20:38:55.848699Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-27T01:13:11.483599Z digest=sha256:aec159e00d608a660403644ecbeb39ea3871ceebaa23d8dbb9199a73619fb036

Observation 3c977e00-2de3-42ac-b623-1fe22de75659 · inbound

Confidence Laundering in Agent Systems: Why Uncertainty Needs a Latent Carrier cites this paper.

Confidence Laundering in Agent Systems: Why Uncertainty Needs a Latent Carrier First Return, Entropy-Eliciting Explore

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-07-03T06:07:41.573469Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-27T12:53:05.153500Z digest=sha256:6a1fa3ea49932f09a00d46c5d9e7da7909bf117501fc23cdf63e8b873db63451

Observation b505331e-ed58-46a2-ba6d-ce078e8d855a · inbound

Modularized Reinforcement Learning on LLMs: From MDP Creation to Exploration and Learning cites this paper.

Modularized Reinforcement Learning on LLMs: From MDP Creation to Exploration and Learning First Return, Entropy-Eliciting Explore

Reference 276

Resolution
verified exact
arxiv_id, observed 2026-07-04T08:09:40.649189Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-26T12:15:08.304150Z digest=sha256:d31cf57409e1a2d060d531c88e980dc07bec0fc0183ef0c38ea4e8c85431f132

Observation 855d224b-6a44-4a97-ae37-fd7df83fc914 · inbound

What are Key Factors for Updates in RL for LLM Reasoning? cites this paper.

What are Key Factors for Updates in RL for LLM Reasoning? First Return, Entropy-Eliciting Explore

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-07-04T09:09:43.540812Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-26T10:24:53.245739Z digest=sha256:8409f4763346feb9f836ae0534f7809c5eea4d938ac57c299a4861926244d3b7

Observation 8331d4b8-dc11-4eda-b24c-c560ab2b1b34 · inbound

Beyond Entropy: Correctness-Aware Advantage Shaping via Contrastive Policy Optimization cites this paper.

Beyond Entropy: Correctness-Aware Advantage Shaping via Contrastive Policy Optimization First Return, Entropy-Eliciting Explore

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-02T01:39:10.853118Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:39:10.853118Z digest=sha256:0344e3bcf949682eebba5caaca651c4e417e254d5949a2b593f1221de1930393

Observation 8114cd00-53cd-4efb-baa9-9557b6121c75 · inbound

Optimizing What Policies Learn From: Recoverability-aware Rollout Intervention Learning cites this paper.

Optimizing What Policies Learn From: Recoverability-aware Rollout Intervention Learning First Return, Entropy-Eliciting Explore

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T05:53:52.768850Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:53:52.768850Z digest=sha256:afbcd101deea7c441375830e62c3d7b99291a765b09ca0f30dc148d101d37933