Pith. sign in

Paper Citation Record · LEDGER

Know When to Explore: Difficulty-Aware Certainty as a Guide for LLM Reinforcement Learning

As of 20 August 2026, this Paper Citation Record lists 41 of 41 outbound references and 8 inbound Pith citation observations for arXiv:2509.00125.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2509.00125 v1

Coverage vector

measured 41 of 41 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T14:23:54.829237Z

measured 49 of 49 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 8 of 8 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-11T04:48:30.873942Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T08:09:40.662689Z

Reference resolution

41 of 41 outbound references displayed

  • verified exact0
  • verified fuzzy16
  • unresolved25
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 0fae64bf-bc87-4977-87d0-02d7f71eaf7e · outbound

This paper cites Learning to reason with llms, 2024.

Know When to Explore: Difficulty-Aware Certainty as a Guide for LLM Reinforcement Learning Learning to reason with llms, 2024

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-05T14:23:52.012864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:23:52.012864Z digest=sha256:0bb094ff42eccabaee94fcb23f6443082d4ba3d39c84722c280525b72c13e41d

Observation 9202f34b-3cb7-442f-97a4-4eab7dba5831 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Know When to Explore: Difficulty-Aware Certainty as a Guide for LLM Reinforcement Learning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-05T14:23:52.044378Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:23:52.044378Z digest=sha256:aae41ce79cc71d446969b34b339f266f556d621f0407940012b8aa26c0fe7704

Observation 4fd8135b-080c-4eec-bf07-a9e4ab3cbfe9 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Know When to Explore: Difficulty-Aware Certainty as a Guide for LLM Reinforcement Learning DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-05T14:23:52.097658Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:23:52.097658Z digest=sha256:eaf798695395ab5b9a46c795e865299a62e49131df3a084189edd40a06b7106d

Observation 34c9fb5e-b0bb-4a94-9879-00c3ce7d357a · outbound

This paper cites Qwen2.5 Technical Report.

Know When to Explore: Difficulty-Aware Certainty as a Guide for LLM Reinforcement Learning Qwen2.5 Technical Report

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-05T14:23:52.153387Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:23:52.153387Z digest=sha256:7e3f6d82ba9ef1d3c22aa478e95410355cba23de6943c051ae36c848c24d6550

Observation 82b20f8d-b312-4c59-b42d-59e511803bf2 · outbound

This paper cites Let’s verify step by step.

Know When to Explore: Difficulty-Aware Certainty as a Guide for LLM Reinforcement Learning Let’s verify step by step

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:23:58.685654Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T14:23:52.193635Z digest=sha256:5f3b0afa410eaad2dcd6bad9ef8e1d5b3c13e7c5f082fc36b994e61044745da5

Observation fb93fe45-4a55-4b3c-8e64-3b59a4fb0a79 · outbound

This paper cites Math-Shepherd: Verify and Reinforce LLMs Step-by-step without Human Annotations.

Know When to Explore: Difficulty-Aware Certainty as a Guide for LLM Reinforcement Learning Math-Shepherd: Verify and Reinforce LLMs Step-by-step without Human Annotations

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-05T14:23:52.252596Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:23:52.252596Z digest=sha256:9310c229982993bd2dc1c5239c5f09b03611f9295f579d4f28ea39610122033b

Observation 3703a958-d47e-45d0-92bb-5518f6b39972 · outbound

This paper cites Scalable best-of-n selection for large language models via self-certainty.

Know When to Explore: Difficulty-Aware Certainty as a Guide for LLM Reinforcement Learning Scalable best-of-n selection for large language models via self-certainty

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-05T14:23:52.474556Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:23:52.474556Z digest=sha256:3786a012dd0ae066af953b24c0a0557f96810528eb6803e5817fd4e22b85207e

Observation 21250019-abda-44ec-8a32-2a4851ebe388 · outbound

This paper cites The impact of intrinsic rewards on exploration in Reinforcement Learning.

Know When to Explore: Difficulty-Aware Certainty as a Guide for LLM Reinforcement Learning The impact of intrinsic rewards on exploration in Reinforcement Learning

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-05T14:23:52.574469Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:23:52.574469Z digest=sha256:5dfb85c8ecbb00feef887933d2b253cd54e85bb7dda7ee69e8e9037cce7e3565

Observation ea1186f2-c4dc-4e56-a502-2c850b501999 · outbound

This paper cites Thompson sampling: An asymptotically optimal finite-time analysis.

Know When to Explore: Difficulty-Aware Certainty as a Guide for LLM Reinforcement Learning Thompson sampling: An asymptotically optimal finite-time analysis

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:23:58.434368Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T14:23:52.607689Z digest=sha256:8b5d912a87cb7d2a063b7660ce83e3ddf0ddecb33c24c682f7353765b72b9b97

Observation 71f55c8d-7cfa-4b66-b4e4-a509853d5451 · outbound

This paper cites Thompson sampling for contextual bandits with linear payoffs.

Know When to Explore: Difficulty-Aware Certainty as a Guide for LLM Reinforcement Learning Thompson sampling for contextual bandits with linear payoffs

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:23:58.287010Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T14:23:52.684821Z digest=sha256:6d94fe030d602986f79f07a42e0bc59248f0ba27fc592f2cd85ca012d6eb87c1

Observation 4060531c-1847-4585-b001-198f32c05dd6 · outbound

This paper cites Exploration-Exploitation in Constrained MDPs.

Know When to Explore: Difficulty-Aware Certainty as a Guide for LLM Reinforcement Learning Exploration-Exploitation in Constrained MDPs

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-05T14:23:52.737271Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:23:52.737271Z digest=sha256:934c311a5b04013a6c08624119d9d922cf2fa08817a1c491e2d06aced6cd6b16

Observation 822ae0fc-6db9-407b-be99-fa09a6948950 · outbound

This paper cites Why is posterior sampling better than optimism for reinforcement learning? In ICML, 2017.

Know When to Explore: Difficulty-Aware Certainty as a Guide for LLM Reinforcement Learning Why is posterior sampling better than optimism for reinforcement learning? In ICML, 2017

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:23:58.007796Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T14:23:52.809809Z digest=sha256:fb32f79c5898bcbe5446e342d4540452b4795919dffb85b22ffda3f6a6c916fa

Observation b1e214fe-a27f-4f1e-8e68-16d1520ea192 · outbound

This paper cites First return, then explore.

Know When to Explore: Difficulty-Aware Certainty as a Guide for LLM Reinforcement Learning First return, then explore

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:23:57.908440Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T14:23:52.860872Z digest=sha256:91b5841c5617200921aec3cb120abf3e8ff80976d731cf06db5f4b5605c5d425

Observation 5cfca35f-e4c9-4323-aab3-7ff3f6200c07 · outbound

This paper cites Automatic goal generation for reinforcement learning agents.

Know When to Explore: Difficulty-Aware Certainty as a Guide for LLM Reinforcement Learning Automatic goal generation for reinforcement learning agents

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:23:57.672989Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T14:23:52.915005Z digest=sha256:d073706cc5c1cbaa2728a02a653af0f7cf25658a9bff2b8adf4d1a868448c562

Observation e0896c7a-5d24-49cc-99f5-bbf3454a21f9 · outbound

This paper cites MIT press Cambridge, 1998.

Know When to Explore: Difficulty-Aware Certainty as a Guide for LLM Reinforcement Learning MIT press Cambridge, 1998

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-05T14:23:52.986538Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:23:52.986538Z digest=sha256:fc2cd7890116973615545c3961a5d259de4f1ff286211f6ee962b27e8e3f627e

Observation 4b22afa6-dbbf-4e5e-a976-d323cbf41029 · outbound

This paper cites Adaptiveε-greedy exploration in reinforcement learning based on value differences.

Know When to Explore: Difficulty-Aware Certainty as a Guide for LLM Reinforcement Learning Adaptiveε-greedy exploration in reinforcement learning based on value differences

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:23:57.397962Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T14:23:53.029151Z digest=sha256:a3f6461938a791ad9f3feb06e30c60ac1e72d9bf8cee631e745c015586ccfa21

Observation af537e40-7b37-47cf-8762-26f1ddde1bed · outbound

This paper cites # exploration: A study of count-based exploration for deep reinforcement learning.

Know When to Explore: Difficulty-Aware Certainty as a Guide for LLM Reinforcement Learning # exploration: A study of count-based exploration for deep reinforcement learning

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:23:57.233013Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T14:23:53.107117Z digest=sha256:bafa935c1d749ef6c3db8c0e693d80058f5b4dd31018fefb47a8dc71e9829443

Observation 625027fc-5657-457b-bc6a-a25ba0a7f7ea · outbound

This paper cites Deep exploration via bootstrapped dqn.

Know When to Explore: Difficulty-Aware Certainty as a Guide for LLM Reinforcement Learning Deep exploration via bootstrapped dqn

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:23:57.077614Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T14:23:53.159916Z digest=sha256:189d5476c681c9fc240e10f5102ca2c07080688dc862829ea88e3dab6aebe73e

Observation 47b9055b-4018-4590-9ea4-cd41539fb231 · outbound

This paper cites Curious model-building control systems.

Know When to Explore: Difficulty-Aware Certainty as a Guide for LLM Reinforcement Learning Curious model-building control systems

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:23:56.843868Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T14:23:53.219109Z digest=sha256:2d7be9682a75b6dcc5b976f17939436ac38e8a05cb392d53ff2fd2508c8b86e1

Observation f64fc493-eefc-44ad-931a-d5febad69903 · outbound

This paper cites Curiosity-driven exploration by self-supervised prediction.

Know When to Explore: Difficulty-Aware Certainty as a Guide for LLM Reinforcement Learning Curiosity-driven exploration by self-supervised prediction

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:23:56.624749Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T14:23:53.288373Z digest=sha256:65dc2b891511e3d894e4f93cb4894323b61eaca31f8d35ac5d54b1006f6043b9

Observation 071b3708-a6bb-4d24-9a98-dd9e66e6f994 · outbound

This paper cites Exploration by Random Network Distillation.

Know When to Explore: Difficulty-Aware Certainty as a Guide for LLM Reinforcement Learning Exploration by Random Network Distillation

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-05T14:23:53.349740Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:23:53.349740Z digest=sha256:ded213d0238ac136ee0abeb0a06fcff4af0570f506055deac96cacb0cca73f4f

Observation b194bb3f-0f88-4cef-83ff-3c494a524657 · outbound

This paper cites Planning to explore via self-supervised world models.

Know When to Explore: Difficulty-Aware Certainty as a Guide for LLM Reinforcement Learning Planning to explore via self-supervised world models

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:23:56.400136Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T14:23:53.445822Z digest=sha256:92bc4ec01be5cf3c92910f21a1a38da780fe82b3d803e0e2eb1d5d68da0cae12

Observation 954ec9c2-13a1-461f-8f5b-5848e91dbe33 · outbound

This paper cites Vime: Variational information maximizing exploration.

Know When to Explore: Difficulty-Aware Certainty as a Guide for LLM Reinforcement Learning Vime: Variational information maximizing exploration

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:23:56.128094Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T14:23:53.513064Z digest=sha256:21417dfef445a413aacb2d56182e0c7e3dffd3325f62a227cc9bb3dd92d18419

Observation b7cc4687-085e-40d0-bed4-45c222baf6b1 · outbound

This paper cites Diversity is all you need: Learning skills without a reward function.

Know When to Explore: Difficulty-Aware Certainty as a Guide for LLM Reinforcement Learning Diversity is all you need: Learning skills without a reward function

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:23:55.855227Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T14:23:53.586478Z digest=sha256:49291e430c0f530e0b97fa0469f0d31ee7e071be6845f601ddc5dd7228f49370

Observation d57a21b2-6589-4b9f-a40b-01630ae61e98 · outbound

This paper cites Ttrl: Test-time reinforcement learning.

Know When to Explore: Difficulty-Aware Certainty as a Guide for LLM Reinforcement Learning Ttrl: Test-time reinforcement learning

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:23:55.542844Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T14:23:53.655512Z digest=sha256:39ef9fcab1bba5cbfb0731d551e505417d6a2c1f64fefa5c0555836972ef5798

Observation 9420cb17-f685-4dd4-a6d8-48a6600f33fa · outbound

This paper cites Learning to Reason without External Rewards.

Know When to Explore: Difficulty-Aware Certainty as a Guide for LLM Reinforcement Learning Learning to Reason without External Rewards

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-05T14:23:53.726913Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:23:53.726913Z digest=sha256:d448b042f29a5ff105a4ad23760f3ab8dcee3a268ca74ac679cc53c068a4971e

Observation 747bd84f-aaff-4c49-8042-15dc9311dd0b · outbound

This paper cites Right Question is Already Half the Answer: Fully Unsupervised LLM Reasoning Incentivization.

Know When to Explore: Difficulty-Aware Certainty as a Guide for LLM Reinforcement Learning Right Question is Already Half the Answer: Fully Unsupervised LLM Reasoning Incentivization

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-05T14:23:53.805378Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:23:53.805378Z digest=sha256:2fdfcbbcd1947385d8332bd358ce33c3782d6e9f1609e7ba10ced4c443d796d3

Observation 91c343ed-07cd-4656-85ed-de4239c3f1c1 · outbound

This paper cites The Unreasonable Effectiveness of Entropy Minimization in LLM Reasoning.

Know When to Explore: Difficulty-Aware Certainty as a Guide for LLM Reinforcement Learning The Unreasonable Effectiveness of Entropy Minimization in LLM Reasoning

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-05T14:23:53.864174Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:23:53.864174Z digest=sha256:456d21296e37de582bd33907499071991e93f6608600c6a7872c0f399b2e7845

Observation 72352bc2-0201-46b3-9b37-71dd8e686483 · outbound

This paper cites One-shot Entropy Minimization.

Know When to Explore: Difficulty-Aware Certainty as a Guide for LLM Reinforcement Learning One-shot Entropy Minimization

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-05T14:23:53.913129Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:23:53.913129Z digest=sha256:09a473e333781a7f62bd493500455751b0fd50a60bb30309bb01ab89407329d6

Observation f2e1034f-eb4b-4ef7-9838-69608525c0c4 · outbound

This paper cites Beyond the 80/20 Rule: High-Entropy Minority Tokens Drive Effective Reinforcement Learning for LLM Reasoning.

Know When to Explore: Difficulty-Aware Certainty as a Guide for LLM Reinforcement Learning Beyond the 80/20 Rule: High-Entropy Minority Tokens Drive Effective Reinforcement Learning for LLM Reasoning

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-05T14:23:53.993294Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:23:53.993294Z digest=sha256:27e63fbd4545fc972138255ae308f3f48bd1d10d0350fcbf6b8b75e87e71fed8

Observation 7d8e9ecb-8f18-4653-b893-fc7ccde05aad · outbound

This paper cites First Return, Entropy-Eliciting Explore.

Know When to Explore: Difficulty-Aware Certainty as a Guide for LLM Reinforcement Learning First Return, Entropy-Eliciting Explore

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-05T14:23:54.061991Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:23:54.061991Z digest=sha256:a4d19ddf5967461d91c66f599ab5d85157170934ba8b6a61c3c30009925bb17f

Observation f036f048-214b-4d9e-8a89-9046536ab732 · outbound

This paper cites Reasoning with Exploration: An Entropy Perspective.

Know When to Explore: Difficulty-Aware Certainty as a Guide for LLM Reinforcement Learning Reasoning with Exploration: An Entropy Perspective

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-05T14:23:54.165008Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:23:54.165008Z digest=sha256:d5b45cdd60ac32791e5d14c2b41aac7f45195db6c2b05b32e673d514d09fd879

Observation 77947edf-38ea-4181-93e3-f094562744d5 · outbound

This paper cites Process Reinforcement through Implicit Rewards.

Know When to Explore: Difficulty-Aware Certainty as a Guide for LLM Reinforcement Learning Process Reinforcement through Implicit Rewards

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-05T14:23:54.263620Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:23:54.263620Z digest=sha256:71e6c683b48a201b51767ae4170beadba51f56eb690869f04381f2523f392256

Observation c944a4df-f17b-4859-8da6-eafd9de83834 · outbound

This paper cites Free Process Rewards without Process Labels.

Know When to Explore: Difficulty-Aware Certainty as a Guide for LLM Reinforcement Learning Free Process Rewards without Process Labels

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-05T14:23:54.346422Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:23:54.346422Z digest=sha256:af7e0bd1a53a5639390ce432b0157d04244b3bd84c99ca8bc6d02a41f439c962

Observation 6ace7843-b82d-4429-a83c-473e3d336d9d · outbound

This paper cites HybridFlow: A Flexible and Efficient RLHF Framework.

Know When to Explore: Difficulty-Aware Certainty as a Guide for LLM Reinforcement Learning HybridFlow: A Flexible and Efficient RLHF Framework

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-05T14:23:54.429131Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:23:54.429131Z digest=sha256:1ee7d247cf38c70501e4490f4b59ed4200977d59f7635759c30415d65c717183

Observation 1e3a05f2-725d-472a-89fd-6cf0e83dd2eb · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

Know When to Explore: Difficulty-Aware Certainty as a Guide for LLM Reinforcement Learning DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-05T14:23:54.506644Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:23:54.506644Z digest=sha256:4f35915dc4f3289eb66e1d917783f01918f77e8c81fbfd2bbe5edec793d26a0d

Observation 7bcc21ec-a01b-4b73-b936-4c3833198956 · outbound

This paper cites Measuring Mathematical Problem Solving With the MATH Dataset.

Know When to Explore: Difficulty-Aware Certainty as a Guide for LLM Reinforcement Learning Measuring Mathematical Problem Solving With the MATH Dataset

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-05T14:23:54.563362Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:23:54.563362Z digest=sha256:340cccb2182d06ac46a13919c5f2ab166a5d9890dd1def75d753d8c6000c4721

Observation 5a84f266-0bad-4488-ae18-149dd2857216 · outbound

This paper cites The Entropy Mechanism of Reinforcement Learning for Reasoning Language Models.

Know When to Explore: Difficulty-Aware Certainty as a Guide for LLM Reinforcement Learning The Entropy Mechanism of Reinforcement Learning for Reasoning Language Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-05T14:23:54.656234Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:23:54.656234Z digest=sha256:2b47ef34580ed21c400e6f20d53fb6db1e3c6e564bc387209216018fa4baa858

Observation 1f7d0dc9-5dd4-47be-8d47-40168e4891c3 · outbound

This paper cites Curriculum learning.

Know When to Explore: Difficulty-Aware Certainty as a Guide for LLM Reinforcement Learning Curriculum learning

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-05T14:23:54.766364Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:23:54.766364Z digest=sha256:597cc3989e39a88115af4597e7785684015ec00b12c50a5e1ea4c3bf6855c0a9

Observation 7ab1850b-1883-4ced-b777-0144ee09d814 · outbound

This paper cites Efficient memory management for large language model serving with pagedattention.

Know When to Explore: Difficulty-Aware Certainty as a Guide for LLM Reinforcement Learning Efficient memory management for large language model serving with pagedattention

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-05T14:23:54.804613Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:23:54.804613Z digest=sha256:2b7cdaa73039ae65bf5abdd3985e8e19fc8230a8329c3f0d950ff31ea685e462

Observation 3480521a-824d-43ae-a716-80da712b82e7 · outbound

This paper cites Math-Verify: Math Verification Library.https://github.com/huggingface/math-verify, 2025.

Know When to Explore: Difficulty-Aware Certainty as a Guide for LLM Reinforcement Learning Math-Verify: Math Verification Library.https://github.com/huggingface/math-verify, 2025

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T14:23:55.242466Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-05T14:23:54.829237Z digest=sha256:0652b23e0341ad8a3473b60604e1be9daa643f0a98fe6403bc8ed663106756fa

Pith citing papers

Observation ccf703d3-489a-4cbe-adaf-5d000ca5b6bf · inbound

rePIRL: Learn PRM with Inverse RL for LLM Reasoning cites this paper.

rePIRL: Learn PRM with Inverse RL for LLM Reasoning Know When to Explore: Difficulty-Aware Certainty as a Guide for LLM Reinforcement Learning

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-21T13:14:10.971468Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-21T13:13:13.293921Z digest=sha256:3bbe7ab5bc7ae0b1857750fc60684da38ae41547ee194889367699f2b8820d56

Observation 09b23c32-fb30-4c1a-a096-9dbebe7a4b2b · inbound

rePIRL: Learn PRM with Inverse RL for LLM Reasoning cites this paper.

rePIRL: Learn PRM with Inverse RL for LLM Reasoning Know When to Explore: Difficulty-Aware Certainty as a Guide for LLM Reinforcement Learning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-03T03:33:43.695859Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:33:43.695859Z digest=sha256:8e77af5acec185963bd4af9a61abe6672d56cb761a4df57c7d9174ef6fcf3fa2

Observation 5c05e544-ae55-45ef-b807-6a1da0229726 · inbound

HTPO: Towards Exploration-Exploitation Balanced Policy Optimization via Hierarchical Token-level Objective Control cites this paper.

HTPO: Towards Exploration-Exploitation Balanced Policy Optimization via Hierarchical Token-level Objective Control Know When to Explore: Difficulty-Aware Certainty as a Guide for LLM Reinforcement Learning

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-12T00:51:14.596330Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-12T00:50:42.836549Z digest=sha256:a51ce47decface4237f9b9cb182b4c03a02738f6cf438098e16afc6c56d83b3f

Observation f4f5a4c4-890c-4235-bb9b-38ae12da4741 · inbound

Epistemic Uncertainty for Test-Time Discovery cites this paper.

Epistemic Uncertainty for Test-Time Discovery Know When to Explore: Difficulty-Aware Certainty as a Guide for LLM Reinforcement Learning

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-13T01:57:06.044731Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-13T01:52:41.192353Z digest=sha256:83e992224999b2e19ac9780c12d475322f8426f0f0b29953d8dfd17cc032dcf6

Observation efaff5c3-2482-45f9-8ded-02939fa1f10b · inbound

From Reasoning Traces to Reusable Modules: Understanding Compositional Generalization in Language Model Reasoning cites this paper.

From Reasoning Traces to Reusable Modules: Understanding Compositional Generalization in Language Model Reasoning Know When to Explore: Difficulty-Aware Certainty as a Guide for LLM Reinforcement Learning

Reference 147

Resolution
verified exact
arxiv_id, observed 2026-07-03T20:38:56.101152Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-06-27T01:13:11.483599Z digest=sha256:79cba775906bee5a8707cbbfc2e07b122fcaa0131711a1f45ee8c3fa2ef33aa3

Observation 09d9d026-81c7-47ef-bb80-983a28743583 · inbound

Formalizing Task-Space Complexity for Zero-Shot Generalization cites this paper.

Formalizing Task-Space Complexity for Zero-Shot Generalization Know When to Explore: Difficulty-Aware Certainty as a Guide for LLM Reinforcement Learning

Reference 69

Resolution
verified exact
arxiv_id, observed 2026-07-04T03:59:32.894403Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-06-26T17:28:55.356850Z digest=sha256:597e487ec868258b1794c1c4bed6869d33bac5aaa971daf4d22b09aab8a6917f

Observation 7648875b-3f7a-4c46-af90-34ae6c91614b · inbound

Modularized Reinforcement Learning on LLMs: From MDP Creation to Exploration and Learning cites this paper.

Modularized Reinforcement Learning on LLMs: From MDP Creation to Exploration and Learning Know When to Explore: Difficulty-Aware Certainty as a Guide for LLM Reinforcement Learning

Reference 102

Resolution
verified exact
arxiv_id, observed 2026-07-04T08:09:40.664145Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-06-26T12:15:08.304150Z digest=sha256:fe65c479e89eb68383913d38f4ff919662fd131e7aab79d8c2094b123f05cfdb

Observation 0da969a2-3a7c-41e3-a361-258f7d9934ab · inbound

Consilience for Verifier-Free Test-Time Scaling cites this paper.

Consilience for Verifier-Free Test-Time Scaling Know When to Explore: Difficulty-Aware Certainty as a Guide for LLM Reinforcement Learning

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T04:48:30.873942Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:48:30.873942Z digest=sha256:7c68da30194ca3c854194eb75b8a7002672b6ddb65058930a82221032d037ddb