Pith. sign in

Paper Citation Record · LEDGER

Know When to Explore: Difficulty-Aware Certainty as a Guide for LLM Reinforcement Learning

As of 9 August 2026, this Paper Citation Record lists 41 of 41 outbound references and 7 inbound Pith citation observations for arXiv:2509.00125.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2509.00125 v1

Coverage vector

measured 41 of 41 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T14:23:54.829237Z

measured 48 of 48 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 7 of 7 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-03T03:33:43.695859Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T08:09:40.662689Z

Reference resolution

41 of 41 outbound references displayed

  • verified exact0
  • verified fuzzy16
  • unresolved25
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 0fae64bf-bc87-4977-87d0-02d7f71eaf7e · outbound

This paper cites Learning to reason with llms, 2024.

Know When to Explore: Difficulty-Aware Certainty as a Guide for LLM Reinforcement Learning Learning to reason with llms, 2024

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-05T14:23:52.012864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:23:52.012864Z digest=sha256:01a8e9724159eb48ea6dee978d0da0f0074ba8ac0eb4dc0922ec83ccf7b5785b

Observation 9202f34b-3cb7-442f-97a4-4eab7dba5831 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Know When to Explore: Difficulty-Aware Certainty as a Guide for LLM Reinforcement Learning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-05T14:23:52.044378Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:23:52.044378Z digest=sha256:622c68f40c5aef089758b75d996b703c7efb01e333cddfbd505429549d532009

Observation 4fd8135b-080c-4eec-bf07-a9e4ab3cbfe9 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Know When to Explore: Difficulty-Aware Certainty as a Guide for LLM Reinforcement Learning DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-05T14:23:52.097658Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:23:52.097658Z digest=sha256:285049220bc340b7246df7b6f4affe7aa8dec193c4b5ae7a4b93f0bf9963289b

Observation 34c9fb5e-b0bb-4a94-9879-00c3ce7d357a · outbound

This paper cites Qwen2.5 Technical Report.

Know When to Explore: Difficulty-Aware Certainty as a Guide for LLM Reinforcement Learning Qwen2.5 Technical Report

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-05T14:23:52.153387Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:23:52.153387Z digest=sha256:c93e0a7573709b7ac0558c9e6c0f139a2c59acad60fdec277087c78cb9c1bd79

Observation 82b20f8d-b312-4c59-b42d-59e511803bf2 · outbound

This paper cites Let’s verify step by step.

Know When to Explore: Difficulty-Aware Certainty as a Guide for LLM Reinforcement Learning Let’s verify step by step

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:23:58.685654Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T14:23:52.193635Z digest=sha256:4dad88418e94f0187571fbfe94b36f0a4f6281b46746ce0f5aaec8cc7eda3e60

Observation fb93fe45-4a55-4b3c-8e64-3b59a4fb0a79 · outbound

This paper cites Math-Shepherd: Verify and Reinforce LLMs Step-by-step without Human Annotations.

Know When to Explore: Difficulty-Aware Certainty as a Guide for LLM Reinforcement Learning Math-Shepherd: Verify and Reinforce LLMs Step-by-step without Human Annotations

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-05T14:23:52.252596Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:23:52.252596Z digest=sha256:4ad8812a5b6c716aba4f4a36e6be05b283ee08f6746a0eb2853ba2521dad7c24

Observation 3703a958-d47e-45d0-92bb-5518f6b39972 · outbound

This paper cites Scalable best-of-n selection for large language models via self-certainty.

Know When to Explore: Difficulty-Aware Certainty as a Guide for LLM Reinforcement Learning Scalable best-of-n selection for large language models via self-certainty

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-05T14:23:52.474556Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:23:52.474556Z digest=sha256:f408df69ed865c8aa86d1fc5ae795bb2ff616c08b15f42e50ddfe260c9d25547

Observation 21250019-abda-44ec-8a32-2a4851ebe388 · outbound

This paper cites The impact of intrinsic rewards on exploration in Reinforcement Learning.

Know When to Explore: Difficulty-Aware Certainty as a Guide for LLM Reinforcement Learning The impact of intrinsic rewards on exploration in Reinforcement Learning

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-05T14:23:52.574469Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:23:52.574469Z digest=sha256:404f3ae29165f44e57f2d52306b45e8b84396939518f3edb0d9ab367a67e8953

Observation ea1186f2-c4dc-4e56-a502-2c850b501999 · outbound

This paper cites Thompson sampling: An asymptotically optimal finite-time analysis.

Know When to Explore: Difficulty-Aware Certainty as a Guide for LLM Reinforcement Learning Thompson sampling: An asymptotically optimal finite-time analysis

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:23:58.434368Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T14:23:52.607689Z digest=sha256:33a9fab6e8b2d8987f64cc93d59bef4180a8405b20627b0b7b414d8b40a63f60

Observation 71f55c8d-7cfa-4b66-b4e4-a509853d5451 · outbound

This paper cites Thompson sampling for contextual bandits with linear payoffs.

Know When to Explore: Difficulty-Aware Certainty as a Guide for LLM Reinforcement Learning Thompson sampling for contextual bandits with linear payoffs

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:23:58.287010Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T14:23:52.684821Z digest=sha256:76c52c0ec197c451d50645c999fa9ecacaef7c9f3aa78faff6b0a216d3a56b0f

Observation 4060531c-1847-4585-b001-198f32c05dd6 · outbound

This paper cites Exploration-Exploitation in Constrained MDPs.

Know When to Explore: Difficulty-Aware Certainty as a Guide for LLM Reinforcement Learning Exploration-Exploitation in Constrained MDPs

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-05T14:23:52.737271Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:23:52.737271Z digest=sha256:796b1be8aa1805cba4569735a35af4f2f8b537f268ac8a3a2f497e0b3659aee9

Observation 822ae0fc-6db9-407b-be99-fa09a6948950 · outbound

This paper cites Why is posterior sampling better than optimism for reinforcement learning? In ICML, 2017.

Know When to Explore: Difficulty-Aware Certainty as a Guide for LLM Reinforcement Learning Why is posterior sampling better than optimism for reinforcement learning? In ICML, 2017

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:23:58.007796Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T14:23:52.809809Z digest=sha256:baa19bd215a93ab8bd2698374ed9ff04028d3240b5887b6f4a9500a4840345bb

Observation b1e214fe-a27f-4f1e-8e68-16d1520ea192 · outbound

This paper cites First return, then explore.

Know When to Explore: Difficulty-Aware Certainty as a Guide for LLM Reinforcement Learning First return, then explore

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:23:57.908440Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T14:23:52.860872Z digest=sha256:df65caab5ce46b877a9943c98d25d6f587142b7b4c33197b82e0269b71179dcf

Observation 5cfca35f-e4c9-4323-aab3-7ff3f6200c07 · outbound

This paper cites Automatic goal generation for reinforcement learning agents.

Know When to Explore: Difficulty-Aware Certainty as a Guide for LLM Reinforcement Learning Automatic goal generation for reinforcement learning agents

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:23:57.672989Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T14:23:52.915005Z digest=sha256:cc585930955c579336186646c6e3454e43e750a82a86c37ddae5fbbecfcec13d

Observation e0896c7a-5d24-49cc-99f5-bbf3454a21f9 · outbound

This paper cites MIT press Cambridge, 1998.

Know When to Explore: Difficulty-Aware Certainty as a Guide for LLM Reinforcement Learning MIT press Cambridge, 1998

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-05T14:23:52.986538Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:23:52.986538Z digest=sha256:3303e319fca73d691075d6d6210ef0381da02946ffbbc54c76b3836cc2893744

Observation 4b22afa6-dbbf-4e5e-a976-d323cbf41029 · outbound

This paper cites Adaptiveε-greedy exploration in reinforcement learning based on value differences.

Know When to Explore: Difficulty-Aware Certainty as a Guide for LLM Reinforcement Learning Adaptiveε-greedy exploration in reinforcement learning based on value differences

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:23:57.397962Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T14:23:53.029151Z digest=sha256:ee20c06b86cf7f23844516d9cf6119ef11e571ea4284f75763fc12e0ced2356a

Observation af537e40-7b37-47cf-8762-26f1ddde1bed · outbound

This paper cites # exploration: A study of count-based exploration for deep reinforcement learning.

Know When to Explore: Difficulty-Aware Certainty as a Guide for LLM Reinforcement Learning # exploration: A study of count-based exploration for deep reinforcement learning

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:23:57.233013Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T14:23:53.107117Z digest=sha256:7e401d998d8ccc551adccdf7c6a503a47b18de69040b205b764f5471f50eb593

Observation 625027fc-5657-457b-bc6a-a25ba0a7f7ea · outbound

This paper cites Deep exploration via bootstrapped dqn.

Know When to Explore: Difficulty-Aware Certainty as a Guide for LLM Reinforcement Learning Deep exploration via bootstrapped dqn

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:23:57.077614Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T14:23:53.159916Z digest=sha256:e0688b49537c615030fdd9881fb684d9a2d3a32a08e308d407873d14616b1c91

Observation 47b9055b-4018-4590-9ea4-cd41539fb231 · outbound

This paper cites Curious model-building control systems.

Know When to Explore: Difficulty-Aware Certainty as a Guide for LLM Reinforcement Learning Curious model-building control systems

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:23:56.843868Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T14:23:53.219109Z digest=sha256:b1fa70c233714066ace29e79bcc02ebd3d1c2bda9ca07d8db74b4a743b7cc7dd

Observation f64fc493-eefc-44ad-931a-d5febad69903 · outbound

This paper cites Curiosity-driven exploration by self-supervised prediction.

Know When to Explore: Difficulty-Aware Certainty as a Guide for LLM Reinforcement Learning Curiosity-driven exploration by self-supervised prediction

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:23:56.624749Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T14:23:53.288373Z digest=sha256:4ad72e3ccbdad15a0634a6ae522033f48d17f4232479238cc6dc9c20eebc10cc

Observation 071b3708-a6bb-4d24-9a98-dd9e66e6f994 · outbound

This paper cites Exploration by Random Network Distillation.

Know When to Explore: Difficulty-Aware Certainty as a Guide for LLM Reinforcement Learning Exploration by Random Network Distillation

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-05T14:23:53.349740Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:23:53.349740Z digest=sha256:2465ac86c51d4ee068ae9ac980587b889fba0dfa96772d7477dd204f618d7a2c

Observation b194bb3f-0f88-4cef-83ff-3c494a524657 · outbound

This paper cites Planning to explore via self-supervised world models.

Know When to Explore: Difficulty-Aware Certainty as a Guide for LLM Reinforcement Learning Planning to explore via self-supervised world models

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:23:56.400136Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T14:23:53.445822Z digest=sha256:ad275294bf7d0b50123ecf439c5d718dddee66f82dc12967877697b9d6862dd4

Observation 954ec9c2-13a1-461f-8f5b-5848e91dbe33 · outbound

This paper cites Vime: Variational information maximizing exploration.

Know When to Explore: Difficulty-Aware Certainty as a Guide for LLM Reinforcement Learning Vime: Variational information maximizing exploration

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:23:56.128094Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T14:23:53.513064Z digest=sha256:787ed2aceee311cf1b9cabc07bd69579ce9581f6ad8412ccd9e9c901d849f610

Observation b7cc4687-085e-40d0-bed4-45c222baf6b1 · outbound

This paper cites Diversity is all you need: Learning skills without a reward function.

Know When to Explore: Difficulty-Aware Certainty as a Guide for LLM Reinforcement Learning Diversity is all you need: Learning skills without a reward function

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:23:55.855227Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T14:23:53.586478Z digest=sha256:bef98ad01745bd55530e343c6360ced8630c28cb1b4e62df51ddc0a5e7dd070a

Observation d57a21b2-6589-4b9f-a40b-01630ae61e98 · outbound

This paper cites Ttrl: Test-time reinforcement learning.

Know When to Explore: Difficulty-Aware Certainty as a Guide for LLM Reinforcement Learning Ttrl: Test-time reinforcement learning

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:23:55.542844Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T14:23:53.655512Z digest=sha256:b0f980d5a8952c9e2e80069b0b9238532491cb7302475a0876ddba502fafcfb1

Observation 9420cb17-f685-4dd4-a6d8-48a6600f33fa · outbound

This paper cites Learning to Reason without External Rewards.

Know When to Explore: Difficulty-Aware Certainty as a Guide for LLM Reinforcement Learning Learning to Reason without External Rewards

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-05T14:23:53.726913Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:23:53.726913Z digest=sha256:09fde66e8535130ec3ea4fdfd913c0b3edde5385d4f3d574078a22b019f88ba3

Observation 747bd84f-aaff-4c49-8042-15dc9311dd0b · outbound

This paper cites Right Question is Already Half the Answer: Fully Unsupervised LLM Reasoning Incentivization.

Know When to Explore: Difficulty-Aware Certainty as a Guide for LLM Reinforcement Learning Right Question is Already Half the Answer: Fully Unsupervised LLM Reasoning Incentivization

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-05T14:23:53.805378Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:23:53.805378Z digest=sha256:830cc3ed65e9107629d80e752a0738ac17179fa6f2f2ce225e10c7f8f7589f00

Observation 91c343ed-07cd-4656-85ed-de4239c3f1c1 · outbound

This paper cites The Unreasonable Effectiveness of Entropy Minimization in LLM Reasoning.

Know When to Explore: Difficulty-Aware Certainty as a Guide for LLM Reinforcement Learning The Unreasonable Effectiveness of Entropy Minimization in LLM Reasoning

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-05T14:23:53.864174Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:23:53.864174Z digest=sha256:c1914f73ed5eaf551c9f1f0d3f056e6355bcb5ebeb322e19ba9b04a6ecebcf89

Observation 72352bc2-0201-46b3-9b37-71dd8e686483 · outbound

This paper cites One-shot Entropy Minimization.

Know When to Explore: Difficulty-Aware Certainty as a Guide for LLM Reinforcement Learning One-shot Entropy Minimization

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-05T14:23:53.913129Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:23:53.913129Z digest=sha256:6ab77dd5d0748b4001573da6f9eaadd6f04dd1c1c643edbb01b2e18a281f838e

Observation f2e1034f-eb4b-4ef7-9838-69608525c0c4 · outbound

This paper cites Beyond the 80/20 Rule: High-Entropy Minority Tokens Drive Effective Reinforcement Learning for LLM Reasoning.

Know When to Explore: Difficulty-Aware Certainty as a Guide for LLM Reinforcement Learning Beyond the 80/20 Rule: High-Entropy Minority Tokens Drive Effective Reinforcement Learning for LLM Reasoning

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-05T14:23:53.993294Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:23:53.993294Z digest=sha256:d2f25c76c6c288f622f68a5d7108e662b2fc82fb810dc59628ced6513e16bd30

Observation 7d8e9ecb-8f18-4653-b893-fc7ccde05aad · outbound

This paper cites First Return, Entropy-Eliciting Explore.

Know When to Explore: Difficulty-Aware Certainty as a Guide for LLM Reinforcement Learning First Return, Entropy-Eliciting Explore

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-05T14:23:54.061991Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:23:54.061991Z digest=sha256:3cebfb94a5b8929809553c5e24b280342dfc84c291f9dbf64415b2fb9d4f7f12

Observation f036f048-214b-4d9e-8a89-9046536ab732 · outbound

This paper cites Reasoning with Exploration: An Entropy Perspective.

Know When to Explore: Difficulty-Aware Certainty as a Guide for LLM Reinforcement Learning Reasoning with Exploration: An Entropy Perspective

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-05T14:23:54.165008Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:23:54.165008Z digest=sha256:aacbc1c39456a38f562c8bec96d046cefb66768fb3b195b83d7eb208d853d5e9

Observation 77947edf-38ea-4181-93e3-f094562744d5 · outbound

This paper cites Process Reinforcement through Implicit Rewards.

Know When to Explore: Difficulty-Aware Certainty as a Guide for LLM Reinforcement Learning Process Reinforcement through Implicit Rewards

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-05T14:23:54.263620Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:23:54.263620Z digest=sha256:6490381a67b9a26723ae0b009e945d55430b7f04b8802c7b3e665f832e88f735

Observation c944a4df-f17b-4859-8da6-eafd9de83834 · outbound

This paper cites Free Process Rewards without Process Labels.

Know When to Explore: Difficulty-Aware Certainty as a Guide for LLM Reinforcement Learning Free Process Rewards without Process Labels

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-05T14:23:54.346422Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:23:54.346422Z digest=sha256:875073fdde3a54e2ef9dc01707db449e578617448ed8a6ff8c7522da1f163dd8

Observation 6ace7843-b82d-4429-a83c-473e3d336d9d · outbound

This paper cites HybridFlow: A Flexible and Efficient RLHF Framework.

Know When to Explore: Difficulty-Aware Certainty as a Guide for LLM Reinforcement Learning HybridFlow: A Flexible and Efficient RLHF Framework

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-05T14:23:54.429131Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:23:54.429131Z digest=sha256:6c0a2f5a514d56875015fdbed7d94c49c5319a9181e97cb00dfa2b0ec8955e07

Observation 1e3a05f2-725d-472a-89fd-6cf0e83dd2eb · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

Know When to Explore: Difficulty-Aware Certainty as a Guide for LLM Reinforcement Learning DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-05T14:23:54.506644Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:23:54.506644Z digest=sha256:3b19351706fcbd92cae65e111f82b68de14a5acd9547f714d664a8e1c35808fd

Observation 7bcc21ec-a01b-4b73-b936-4c3833198956 · outbound

This paper cites Measuring Mathematical Problem Solving With the MATH Dataset.

Know When to Explore: Difficulty-Aware Certainty as a Guide for LLM Reinforcement Learning Measuring Mathematical Problem Solving With the MATH Dataset

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-05T14:23:54.563362Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:23:54.563362Z digest=sha256:55a715d8800d5607f46b4ee817dbbc6a5371fd3adab19843a0898747165eba67

Observation 5a84f266-0bad-4488-ae18-149dd2857216 · outbound

This paper cites The Entropy Mechanism of Reinforcement Learning for Reasoning Language Models.

Know When to Explore: Difficulty-Aware Certainty as a Guide for LLM Reinforcement Learning The Entropy Mechanism of Reinforcement Learning for Reasoning Language Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-05T14:23:54.656234Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:23:54.656234Z digest=sha256:beb4e4a83f9d85e8fef808a71c9db2c133e343b4e41969393fe73ce358c64c33

Observation 1f7d0dc9-5dd4-47be-8d47-40168e4891c3 · outbound

This paper cites Curriculum learning.

Know When to Explore: Difficulty-Aware Certainty as a Guide for LLM Reinforcement Learning Curriculum learning

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-05T14:23:54.766364Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:23:54.766364Z digest=sha256:c070d77c7714ab3d8ff2f4d54d907f94f09680ad52394815583c59a818e94a25

Observation 7ab1850b-1883-4ced-b777-0144ee09d814 · outbound

This paper cites Efficient memory management for large language model serving with pagedattention.

Know When to Explore: Difficulty-Aware Certainty as a Guide for LLM Reinforcement Learning Efficient memory management for large language model serving with pagedattention

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-05T14:23:54.804613Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:23:54.804613Z digest=sha256:ebd4d602b427cf86144328e492cdd47cc366ad457492dece6d0e521b57a2dd85

Observation 3480521a-824d-43ae-a716-80da712b82e7 · outbound

This paper cites Math-Verify: Math Verification Library.https://github.com/huggingface/math-verify, 2025.

Know When to Explore: Difficulty-Aware Certainty as a Guide for LLM Reinforcement Learning Math-Verify: Math Verification Library.https://github.com/huggingface/math-verify, 2025

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:23:55.242466Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T14:23:54.829237Z digest=sha256:35bbbf086f57d4887a15e1365c2b37d8be7add48b3e987ac116d0f3152715b88

Pith citing papers

Observation ccf703d3-489a-4cbe-adaf-5d000ca5b6bf · inbound

rePIRL: Learn PRM with Inverse RL for LLM Reasoning cites this paper.

rePIRL: Learn PRM with Inverse RL for LLM Reasoning Know When to Explore: Difficulty-Aware Certainty as a Guide for LLM Reinforcement Learning

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-21T13:14:10.971468Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-21T13:13:13.293921Z digest=sha256:e0a20c32bd116316f0339d1258483f054a213c8f0a023b11026455adf5abfa91

Observation 09b23c32-fb30-4c1a-a096-9dbebe7a4b2b · inbound

rePIRL: Learn PRM with Inverse RL for LLM Reasoning cites this paper.

rePIRL: Learn PRM with Inverse RL for LLM Reasoning Know When to Explore: Difficulty-Aware Certainty as a Guide for LLM Reinforcement Learning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-03T03:33:43.695859Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:33:43.695859Z digest=sha256:58fc69476843277c3ffda72ce3955364898865616491181bf958b8c9806ed48d

Observation 5c05e544-ae55-45ef-b807-6a1da0229726 · inbound

HTPO: Towards Exploration-Exploitation Balanced Policy Optimization via Hierarchical Token-level Objective Control cites this paper.

HTPO: Towards Exploration-Exploitation Balanced Policy Optimization via Hierarchical Token-level Objective Control Know When to Explore: Difficulty-Aware Certainty as a Guide for LLM Reinforcement Learning

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-12T00:51:14.596330Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-12T00:50:42.836549Z digest=sha256:50ac7d25f71f4d458ff2789cd9ed13a71dedf4fbae992e242970998c4f458660

Observation f4f5a4c4-890c-4235-bb9b-38ae12da4741 · inbound

Epistemic Uncertainty for Test-Time Discovery cites this paper.

Epistemic Uncertainty for Test-Time Discovery Know When to Explore: Difficulty-Aware Certainty as a Guide for LLM Reinforcement Learning

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-13T01:57:06.044731Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-13T01:52:41.192353Z digest=sha256:1501149eca673e1cb1fca3051840f73ff429aab4f4f08654d7728fee038dbd89

Observation efaff5c3-2482-45f9-8ded-02939fa1f10b · inbound

From Reasoning Traces to Reusable Modules: Understanding Compositional Generalization in Language Model Reasoning cites this paper.

From Reasoning Traces to Reusable Modules: Understanding Compositional Generalization in Language Model Reasoning Know When to Explore: Difficulty-Aware Certainty as a Guide for LLM Reinforcement Learning

Reference 147

Resolution
verified exact
arxiv_id, observed 2026-07-03T20:38:56.101152Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-27T01:13:11.483599Z digest=sha256:79738f7d5705769b13771dc80839271c6b81f664f3ad9d6d4fda72a5821cba12

Observation 09d9d026-81c7-47ef-bb80-983a28743583 · inbound

Formalizing Task-Space Complexity for Zero-Shot Generalization cites this paper.

Formalizing Task-Space Complexity for Zero-Shot Generalization Know When to Explore: Difficulty-Aware Certainty as a Guide for LLM Reinforcement Learning

Reference 69

Resolution
verified exact
arxiv_id, observed 2026-07-04T03:59:32.894403Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-26T17:28:55.356850Z digest=sha256:d2c9353731aa6ba3ea65a9cd424b27a5fe8ac0a9fcf8b9ba8daa7a29fb5a987e

Observation 7648875b-3f7a-4c46-af90-34ae6c91614b · inbound

Modularized Reinforcement Learning on LLMs: From MDP Creation to Exploration and Learning cites this paper.

Modularized Reinforcement Learning on LLMs: From MDP Creation to Exploration and Learning Know When to Explore: Difficulty-Aware Certainty as a Guide for LLM Reinforcement Learning

Reference 102

Resolution
verified exact
arxiv_id, observed 2026-07-04T08:09:40.664145Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-26T12:15:08.304150Z digest=sha256:45b0f61efe4405cee0d66d74042553098ff070283ca28a9b375470ee7508e105