Pith. sign in

Paper Citation Record · LEDGER

Outcome-based Exploration for LLM Reasoning

As of 18 August 2026, this Paper Citation Record lists 38 of 38 outbound references and 25 inbound Pith citation observations for arXiv:2509.06941.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2509.06941 v1

Coverage vector

measured 38 of 38 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-04T22:59:14.608524Z

measured 63 of 63 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 25 of 25 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T00:15:16.022969Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

38 of 38 outbound references displayed

  • verified exact0
  • verified fuzzy3
  • unresolved35
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

1
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 772784dd-2c91-4acc-97f1-975b46d09b0d · outbound

This paper cites 2:fort= 1,2,.

Outcome-based Exploration for LLM Reasoning 2:fort= 1,2,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T22:59:15.222855Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-04T22:59:14.608524Z digest=sha256:9ade6d41e4b110aa943081ee281fc32afe49f8481994574744aa57fef00d1872

Observation 288c2f09-c7eb-49ea-adb0-38ff0434e55b · outbound

This paper cites Dataset Reset Policy Optimization for RLHF.

Outcome-based Exploration for LLM Reasoning Dataset Reset Policy Optimization for RLHF

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-04T22:59:14.467898Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:59:14.467898Z digest=sha256:b0a4b70a1a1feeedd4a3a10c179daebbb54d563ade26ad98a8e3816f619a9262

Observation d36f566e-f01d-46dc-9179-90ccd70433a4 · outbound

This paper cites Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models.

Outcome-based Exploration for LLM Reasoning Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-04T22:59:14.473018Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:59:14.473018Z digest=sha256:0f5cd8e17880d9265f06ff21432bd9b1e6e3151ca887bc4eae7d45e31feee1ec

Observation 34ba54e2-b4ab-4ee0-846f-2aa771c9ce62 · outbound

This paper cites Reasoning with Exploration: An Entropy Perspective.

Outcome-based Exploration for LLM Reasoning Reasoning with Exploration: An Entropy Perspective

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-04T22:59:14.477484Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:59:14.477484Z digest=sha256:0b8cf9a78f14e02561113b230862c387c3a5383770a6507f94b82ab90a3aafde

Observation 3ce41058-e781-4281-a027-09c512726caa · outbound

This paper cites Weight ensembling improves reasoning in language models.arXiv preprint arXiv:2504.10478,.

Outcome-based Exploration for LLM Reasoning Weight ensembling improves reasoning in language models.arXiv preprint arXiv:2504.10478,

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-04T22:59:14.481689Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:59:14.481689Z digest=sha256:10321e283d5fbfca001d2e358473f47f63bf394db8e1915e3e36398d1c8d9ee4

Observation 4e0d72bb-28eb-42d3-b440-06304a84dbfb · outbound

This paper cites The Statistical Complexity of Interactive Decision Making.

Outcome-based Exploration for LLM Reasoning The Statistical Complexity of Interactive Decision Making

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-04T22:59:14.485864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:59:14.485864Z digest=sha256:20d2d38e039b6b05b11c8014e55501c39e47be05dcd38c7ae32047edd503cfb3

Observation cc6109ad-08bf-4aef-a875-997ebe7deb32 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Outcome-based Exploration for LLM Reasoning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-04T22:59:14.494351Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:59:14.494351Z digest=sha256:87ead6ec20dc15a7e0fb64b457c20399cedcc6c8958564f4388a926767426cde

Observation a506f41a-267b-49c8-a223-5e200306f839 · outbound

This paper cites Measuring Mathematical Problem Solving With the MATH Dataset.

Outcome-based Exploration for LLM Reasoning Measuring Mathematical Problem Solving With the MATH Dataset

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-04T22:59:14.498713Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:59:14.498713Z digest=sha256:b81b571be216c238cf1864a4e7b4ffb5a82ed647b36c87426f4fc0a4116567b0

Observation 359e9c1c-7d9d-440f-96fe-bbdbc4be1d7a · outbound

This paper cites OpenAI o1 System Card.

Outcome-based Exploration for LLM Reasoning OpenAI o1 System Card

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-04T22:59:14.503377Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:59:14.503377Z digest=sha256:16f8215b23cd4ebe9c19e3a9c39e96c0905a7b9b6f487f72d9bff0fdffeffae7

Observation 94662f91-9583-418f-89b4-06d067697620 · outbound

This paper cites Understanding the Effects of RLHF on LLM Generalisation and Diversity.

Outcome-based Exploration for LLM Reasoning Understanding the Effects of RLHF on LLM Generalisation and Diversity

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-04T22:59:14.507748Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:59:14.507748Z digest=sha256:94238e910eae6336f31ef84689bd9f305a6af871bb4ca2e4d0838fc9786c3e35

Observation 2c1a1a68-f682-4f72-ab3e-c7d1d3067f65 · outbound

This paper cites Jointly Reinforcing Diversity and Quality in Language Model Generations.

Outcome-based Exploration for LLM Reasoning Jointly Reinforcing Diversity and Quality in Language Model Generations

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-04T22:59:14.516795Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:59:14.516795Z digest=sha256:95d923c5a7250f03f14e22802d1209bc3cf9b0bfaf4f2ab431b4a7b53a4ae228

Observation 77a22718-0a52-4f80-9d4b-b5cd772f1e00 · outbound

This paper cites Approximating kl divergence, 2020.URL http://joschu.

Outcome-based Exploration for LLM Reasoning Approximating kl divergence, 2020.URL http://joschu

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T22:59:15.239298Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-04T22:59:14.530110Z digest=sha256:63d252b2652e27d0ee50c64e4f8ff41da8b4e9888d2bf3726fad85146fa99123

Observation 498f03d7-ccc7-4751-8b18-066d62061239 · outbound

This paper cites Can large reasoning models self-train?arXiv preprint arXiv:2505.21444,.

Outcome-based Exploration for LLM Reasoning Can large reasoning models self-train?arXiv preprint arXiv:2505.21444,

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-04T22:59:14.538465Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:59:14.538465Z digest=sha256:c49d5ca4a239944a78d1bf40f4d019b894b63d1b7e9cae77fabc65c749657570

Observation 26c7fea0-23d2-4572-910d-e602d65a79a6 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Outcome-based Exploration for LLM Reasoning DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-04T22:59:14.543111Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:59:14.543111Z digest=sha256:dab5dfdda79b66c64d1dbfa64755d4d9250e31333e38863217ae88b1c35c5d47

Observation 2daa3bfb-6428-44ca-a5e3-92d2dcbcebee · outbound

This paper cites HybridFlow: A Flexible and Efficient RLHF Framework.

Outcome-based Exploration for LLM Reasoning HybridFlow: A Flexible and Efficient RLHF Framework

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-04T22:59:14.547422Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:59:14.547422Z digest=sha256:d3568506d7a127d96451145ade26671918fd8e157d76bae45cffc39ab80c9607

Observation ae28b6aa-0078-4384-a72b-596394671fad · outbound

This paper cites Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters.

Outcome-based Exploration for LLM Reasoning Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-04T22:59:14.551823Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:59:14.551823Z digest=sha256:0ca0206d72ca173eb6d6441373551de032ee1b845767ffcef4d3483e277d0992

Observation 73d87d62-3600-436e-8f4c-b412c9fef2d6 · outbound

This paper cites Mind the Gap: Examining the Self-Improvement Capabilities of Large Language Models.

Outcome-based Exploration for LLM Reasoning Mind the Gap: Examining the Self-Improvement Capabilities of Large Language Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-04T22:59:14.556101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:59:14.556101Z digest=sha256:5916e981e2a91004c527eb189453db279d87fd6ace639f672bd8ea44fa3acc26

Observation fb93e9e2-8fda-416f-af68-8a2d11cd188d · outbound

This paper cites On a few pitfalls in KL divergence gradient estimation for RL.

Outcome-based Exploration for LLM Reasoning On a few pitfalls in KL divergence gradient estimation for RL

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-04T22:59:14.560552Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:59:14.560552Z digest=sha256:061b8b47d64eac2fea88940aa8471d485b0a793c9eb37c721518a1a8d7a89f7d

Observation 0d78bad3-7a1d-4923-8247-095ae001384d · outbound

This paper cites Optimizing Language Models for Inference Time Objectives using Reinforcement Learning.

Outcome-based Exploration for LLM Reasoning Optimizing Language Models for Inference Time Objectives using Reinforcement Learning

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-04T22:59:14.565055Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:59:14.565055Z digest=sha256:47941f666c640e571f2010f6e4b7da58a306211cc6f9ae0416983dbc57fd11b1

Observation ab716baf-2fff-4450-87df-83fb63bd4183 · outbound

This paper cites The invisible leash: Why rlvr may not escape its origin.arXiv preprint arXiv:2507.14843,.

Outcome-based Exploration for LLM Reasoning The invisible leash: Why rlvr may not escape its origin.arXiv preprint arXiv:2507.14843,

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-04T22:59:14.569218Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:59:14.569218Z digest=sha256:fdf30076266afc2369e3eec28d74c62eaa99064285b56df5ab293f8b983febef

Observation ee64cf97-aa98-409a-94f1-1a11ad286fef · outbound

This paper cites Exploratory Preference Optimization: Harnessing Implicit Q*-Approximation for Sample-Efficient RLHF.

Outcome-based Exploration for LLM Reasoning Exploratory Preference Optimization: Harnessing Implicit Q*-Approximation for Sample-Efficient RLHF

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-04T22:59:14.573190Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:59:14.573190Z digest=sha256:dca1c5d64ed874d7f34158dee5981472145e15d152e151a238bae16f8828f774

Observation ca620d9e-6f09-4fc7-bfbe-5a69903cbd8f · outbound

This paper cites Iterative Preference Learning from Human Feedback: Bridging Theory and Practice for RLHF under KL-Constraint.

Outcome-based Exploration for LLM Reasoning Iterative Preference Learning from Human Feedback: Bridging Theory and Practice for RLHF under KL-Constraint

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-04T22:59:14.577593Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:59:14.577593Z digest=sha256:bf15a80a1b3ed9bce706f541203943736c349846a92abee86be4d9aff9cf4e19

Observation 635fd568-f897-4a7b-80ec-6f02b0858d77 · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

Outcome-based Exploration for LLM Reasoning DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-04T22:59:14.585858Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:59:14.585858Z digest=sha256:cc02cd240654a8478879fc260ba7be58d7d3d03e28596ed44f9b83e2d425b2d0

Observation ac98f467-6180-45da-b120-1910ac200a53 · outbound

This paper cites Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?.

Outcome-based Exploration for LLM Reasoning Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-04T22:59:14.590415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:59:14.590415Z digest=sha256:d06d8e98efa78d8343b19e8a5b66397f401295a5168b1e1f70e428874a9bfacf

Observation 02725c34-f758-487b-923c-88331c46836e · outbound

This paper cites The Price of Format: Diversity Collapse in LLMs.

Outcome-based Exploration for LLM Reasoning The Price of Format: Diversity Collapse in LLMs

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-04T22:59:14.595531Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:59:14.595531Z digest=sha256:ce71199e72850b8bd7f005133af927d3fc9e13698c55d8a9d27a07c2352d80be

Observation 07e30d42-0401-4a3e-98d7-9f776f3d8e3e · outbound

This paper cites Self-Exploring Language Models: Active Preference Elicitation for Online Alignment.

Outcome-based Exploration for LLM Reasoning Self-Exploring Language Models: Active Preference Elicitation for Online Alignment

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-04T22:59:14.599880Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:59:14.599880Z digest=sha256:bb8175769cdcfb76eb04e0c4b207e8c3845a2cc22cc5cb94a3dae0b8e82a259c

Observation 14b2af2a-9784-420d-848f-0382d14a4838 · outbound

This paper cites First Return, Entropy-Eliciting Explore.

Outcome-based Exploration for LLM Reasoning First Return, Entropy-Eliciting Explore

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-04T22:59:14.604039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:59:14.604039Z digest=sha256:3ab2f7e0f166dd80d6762b414ccc15e3bfa71a209adfeade4cb09646e4889dd6

Observation 923ed1ad-05bf-484b-b01e-fd396b1f5ffd · outbound

This paper cites Optimistically optimistic exploration for provably efficient infinite- horizon reinforcement and imitation learning.arXiv preprint arXiv:2502.13900,.

Outcome-based Exploration for LLM Reasoning Optimistically optimistic exploration for provably efficient infinite- horizon reinforcement and imitation learning.arXiv preprint arXiv:2502.13900,

Reference 1996

Resolution
unresolved
no resolver link, observed 2026-08-04T22:59:14.521058Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:59:14.521058Z digest=sha256:9bee7df94166a75247e7b59dc4ac0e512f39d573f9cff098f20eceeb30cae975

Observation d6b198d2-3e16-4493-851b-f3f8ea1ea061 · outbound

This paper cites Exploration by Random Network Distillation.

Outcome-based Exploration for LLM Reasoning Exploration by Random Network Distillation

Reference 2002

Resolution
unresolved
no resolver link, observed 2026-08-04T22:59:14.458521Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:59:14.458521Z digest=sha256:d7b3b5b9416ca589a442ee21c08fb76beea52d8a52d1967a46281f413f563e89

Observation 8590e1fa-3675-4041-bd44-43f2f9dd1d6e · outbound

This paper cites Diverse Preference Optimization.

Outcome-based Exploration for LLM Reasoning Diverse Preference Optimization

Reference 2016

Resolution
unresolved
no resolver link, observed 2026-08-04T22:59:14.512402Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:59:14.512402Z digest=sha256:38ee8fce98be3165f4b12673b1a79151a1ffdc77ff4297d9bae015849e2769c3

Observation e2c308a8-ab1d-4242-83f7-f1ad2fa216db · outbound

This paper cites Value-Incentivized Preference Optimization: A Unified Approach to Online and Offline RLHF.

Outcome-based Exploration for LLM Reasoning Value-Incentivized Preference Optimization: A Unified Approach to Online and Offline RLHF

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-04T22:59:14.463200Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:59:14.463200Z digest=sha256:b79051a6334358d9bf0cf3b64f7f91ffd3cc158ee9a4e10f1b33289ca1e00307

Observation 5cc9057c-9b41-48ca-8c6c-47c9a99027ba · outbound

This paper cites Qwen2.5 Technical Report.

Outcome-based Exploration for LLM Reasoning Qwen2.5 Technical Report

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-04T22:59:14.581505Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:59:14.581505Z digest=sha256:41259f7ffb466d609191c3e1593b210fe2d2b6e41c6afc2138dbb13f0c4f6666

Observation 62e79bd9-5de4-4b91-b579-44a002b56daf · outbound

This paper cites e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs.

Outcome-based Exploration for LLM Reasoning e3: Learning to Explore Enables Extrapolation of Test-Time Compute for LLMs

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-04T22:59:14.534182Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:59:14.534182Z digest=sha256:21f79a61956ab367a7980dedfa000d5befe3c6dbcd7f180827f5886194316e36

Observation 4927a9db-db97-43b7-bc2c-0629bb70414b · outbound

This paper cites Navigate the unknown: Enhancing llm reasoning with intrinsic motivation guided exploration.arXiv preprint arXiv:2505.17621,.

Outcome-based Exploration for LLM Reasoning Navigate the unknown: Enhancing llm reasoning with intrinsic motivation guided exploration.arXiv preprint arXiv:2505.17621,

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-04T22:59:14.490216Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:59:14.490216Z digest=sha256:a753d9fe3987453b80b1a002898d74f38eab2b79c2c210dfa6b902f73f95f395

Observation d04a1e4f-1a34-405c-bc14-f9f4ca59e70d · outbound

This paper cites Attributing mode collapse in the fine-tuning of large language models.

Outcome-based Exploration for LLM Reasoning Attributing mode collapse in the fine-tuning of large language models

Reference 2022

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T22:59:15.254676Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-04T22:59:14.525912Z digest=sha256:42af6e9348295b5a42b763d7e94a28a75ff91034ce7a9b663afdd471c345f1f6

Observation 964e69e0-4b1d-4850-ae05-6aa530a6b73f · outbound

This paper cites Asymmetric rein- force for off-policy reinforcement learning: Balancing positive and negative rewards.arXiv preprint arXiv:2506.20520,.

Outcome-based Exploration for LLM Reasoning Asymmetric rein- force for off-policy reinforcement learning: Balancing positive and negative rewards.arXiv preprint arXiv:2506.20520,

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-04T22:59:14.443923Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:59:14.443923Z digest=sha256:70d2105b6adf8cd5629904c172cc54d3c52673d97a872837cc80bffb95b3e93b

Observation a61d8773-626d-46fe-b3bf-f7fa28628346 · outbound

This paper cites Online Preference Alignment for Language Models via Count-based Exploration.

Outcome-based Exploration for LLM Reasoning Online Preference Alignment for Language Models via Count-based Exploration

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-04T22:59:14.448812Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:59:14.448812Z digest=sha256:f6cba3f00a32cd431a7f12e83943ee8ec5edcdeec1b7325da32455900d7822e6

Observation b385f625-764e-405a-b34c-728dee68553d · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

Outcome-based Exploration for LLM Reasoning Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-04T22:59:14.454002Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:59:14.454002Z digest=sha256:f7bf6e25226cd12d626892299353a8c2403c001822894e8d07303dfc47d82ba6

Pith citing papers

Observation aba17bf7-aa41-4982-be33-c74970665f6d · inbound

Compute as Teacher: Turning Inference Compute Into Reference-Free Supervision cites this paper.

Compute as Teacher: Turning Inference Compute Into Reference-Free Supervision Outcome-based Exploration for LLM Reasoning

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-18T15:46:34.147476Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-18T15:45:09.730804Z digest=sha256:38bc3cafd30abd1acdac32ac56eee6c3b4ae5b43ff7f4ca25c095a86c3d117fc

Observation 93fc4f7e-c621-4a05-9be5-d73b16ef952b · inbound

Polychromic Objectives for Reinforcement Learning cites this paper.

Polychromic Objectives for Reinforcement Learning Outcome-based Exploration for LLM Reasoning

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-18T11:56:20.005263Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-18T11:54:29.955833Z digest=sha256:4d1f03fb4a6c241aba5d41b716c608a61274dcc05c225118141ba68fe7ea2a85

Observation d7bb58a8-660f-45d5-a3b7-88336b8b09c5 · inbound

Selective Expert Guidance for Effective and Diverse Exploration in Reinforcement Learning of LLMs cites this paper.

Selective Expert Guidance for Effective and Diverse Exploration in Reinforcement Learning of LLMs Outcome-based Exploration for LLM Reasoning

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-04T11:33:55.597960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:33:55.597960Z digest=sha256:3ebec13899fe84fcc3604be615d9f5e3fe7cf504d2687d7c703649f38f132ba6

Observation 47bb9eb1-86ab-484e-aecf-a73e478b05a3 · inbound

On the optimization dynamics of RLVR: Gradient gap and step size thresholds cites this paper.

On the optimization dynamics of RLVR: Gradient gap and step size thresholds Outcome-based Exploration for LLM Reasoning

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-18T08:36:07.367130Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-18T08:34:36.543874Z digest=sha256:9bdb3ad08310f6fee0f598fabf68288719e1b9e73521f443db9f19e8890cdbed

Observation f7150d8a-8e2d-4f20-95f4-cc2b6b036feb · inbound

Rethinking Entropy Interventions in RLVR: An Entropy Change Perspective cites this paper.

Rethinking Entropy Interventions in RLVR: An Entropy Change Perspective Outcome-based Exploration for LLM Reasoning

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-18T07:41:03.322829Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-18T07:37:12.501489Z digest=sha256:56a6945e905a7e09b77a91f96f966c2c568e7d95575839800eb225d3fc7aa441

Observation 8365833d-3a81-4807-8ea1-032cec879060 · inbound

Beyond the Sampled Token: Preserving Candidate Support in RLVR cites this paper.

Beyond the Sampled Token: Preserving Candidate Support in RLVR Outcome-based Exploration for LLM Reasoning

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-04T09:34:09.191451Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T09:34:09.191451Z digest=sha256:c714a9948d02a5ce222b67831b4ce6e15d79f3fcd3bfc78c3493c8095f4c8ddc

Observation 7f61e66d-647e-4ac4-9bb8-3c2b623fc48a · inbound

The Past Is Not Past: Memory-Enhanced Dynamic Reward Shaping cites this paper.

The Past Is Not Past: Memory-Enhanced Dynamic Reward Shaping Outcome-based Exploration for LLM Reasoning

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-11T10:16:05.109966Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T15:34:31.715954Z digest=sha256:332f7c1a62aa586dd52735ff03cf749aa1b59eb0fbcd416b53f175bba74f6233

Observation ec353b52-5182-4790-b97c-e9927c819038 · inbound

The Past Is Not Past: Memory-Enhanced Dynamic Reward Shaping cites this paper.

The Past Is Not Past: Memory-Enhanced Dynamic Reward Shaping Outcome-based Exploration for LLM Reasoning

Reference 26

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T15:35:32.810462Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-10T15:34:31.715954Z digest=sha256:c17530878209e1f71cd4a31131052e4700e552cc5fc45c4124b4f6bb45c3ee28

Observation bc8eb1c4-2005-4bcb-8c0a-d00995cf9619 · inbound

Uniform-Correct Policy Optimization: Breaking RLVR's Indifference to Diversity cites this paper.

Uniform-Correct Policy Optimization: Breaking RLVR's Indifference to Diversity Outcome-based Exploration for LLM Reasoning

Reference 36

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T15:26:15.307690Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-05-09T19:56:39.465133Z digest=sha256:b2ce47eff9f500b219f732174d3a410496304f88417b1ab167688bc726858a81

Observation 3a0da2d4-b844-4ad7-966f-e3e171b50b9c · inbound

Rebellious Student: Reversing Teacher Signals for Reasoning Exploration with Self-Distilled RLVR cites this paper.

Rebellious Student: Reversing Teacher Signals for Reasoning Exploration with Self-Distilled RLVR Outcome-based Exploration for LLM Reasoning

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-12T04:21:22.471904Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-12T04:20:19.940462Z digest=sha256:8e419bf370018e58f5059172ee561e95816cbe1c6c23a42200c1a23b1f448b77

Observation 9d1c9fdd-fa13-48a5-8ad1-5b2657b1e058 · inbound

Automated Reformulation of Robust Optimization via Memory-Augmented Large Language Models cites this paper.

Automated Reformulation of Robust Optimization via Memory-Augmented Large Language Models Outcome-based Exploration for LLM Reasoning

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-13T06:02:22.006754Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-13T06:01:15.799220Z digest=sha256:47393fe6ff84f08b7d60257ca0dbe9489d431459b2abd3deae3240fe882c7f6c

Observation 15b83eb1-f291-473d-bba4-8dbfa2f28860 · inbound

Beyond Mode Collapse: Distribution Matching for Diverse Reasoning cites this paper.

Beyond Mode Collapse: Distribution Matching for Diverse Reasoning Outcome-based Exploration for LLM Reasoning

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-20T05:33:04.041635Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-20T05:30:37.685873Z digest=sha256:c3ec35ddcac547ec092a907685bd911fe8a0e81839cc2aa4f38a0e69c08c3edd

Observation 6b0ad7f3-33f2-4b9e-a5a8-d3c0de4561b9 · inbound

Where Rollouts Begin: Low-Load, High-Leverage First-Token Diversification for RLVR cites this paper.

Where Rollouts Begin: Low-Load, High-Leverage First-Token Diversification for RLVR Outcome-based Exploration for LLM Reasoning

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-06-29T12:03:24.374582Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-29T11:55:13.135221Z digest=sha256:8d3b88e7057f28a9a4c5a90af648ce40030a1c3831b0d51f656cacc64bd79276

Observation 5c2f4873-9732-44ca-92cd-a2ca7728e82b · inbound

Exploiting Verification-Generation Gap: Test-Time Reinforcement Learning with Confidence-Conditioned Verification cites this paper.

Exploiting Verification-Generation Gap: Test-Time Reinforcement Learning with Confidence-Conditioned Verification Outcome-based Exploration for LLM Reasoning

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-07-02T02:26:27.241826Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-28T10:53:00.223228Z digest=sha256:f468ee9bbf7e3f6377292a57dae3f9bbc80e3bd7e23812e18381ae0f7ffa168e

Observation cfbde808-e086-48c3-a262-22ea20a23115 · inbound

Using Reward Uncertainty to Induce Diverse Behaviour in Reinforcement Learning cites this paper.

Using Reward Uncertainty to Induce Diverse Behaviour in Reinforcement Learning Outcome-based Exploration for LLM Reasoning

Reference 53

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T02:26:26.262089Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-06-28T10:59:44.092482Z digest=sha256:7ba9fdc1b2d84fe0630ba4011202b98ec7a53d979c92850da75e09485703022e

Observation 04a2bc2c-32e2-40d3-9bb4-51b9022ca6e9 · inbound

On Advantage Estimates for Max@K Policy Gradients cites this paper.

On Advantage Estimates for Max@K Policy Gradients Outcome-based Exploration for LLM Reasoning

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-07-02T12:06:56.413899Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-28T02:21:57.143016Z digest=sha256:3d28b3909bb6848272e136ad36f2f1390ba569a7567d0e4ea9110ff940c17780

Observation a8126bf8-e979-40d1-a9e8-ef3344887b16 · inbound

OrderGrad: Optimizing Beyond the Mean with Order-Statistic Policy Gradient Estimation cites this paper.

OrderGrad: Optimizing Beyond the Mean with Order-Statistic Policy Gradient Estimation Outcome-based Exploration for LLM Reasoning

Reference 90

Resolution
verified exact
arxiv_id, observed 2026-07-02T12:16:56.999981Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-28T02:17:30.974692Z digest=sha256:aae4d574b68ab1ef3d6798cd5721bfb4a7451b55f567fd63daed3547922af0f1

Observation d9b42875-8786-441d-b3f2-9035010c79b4 · inbound

Sample Where You Struggle: Sharpening Base Model Reasoning via Entropy-Guided Power Sampling cites this paper.

Sample Where You Struggle: Sharpening Base Model Reasoning via Entropy-Guided Power Sampling Outcome-based Exploration for LLM Reasoning

Reference 26

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T22:37:25.535668Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-06-27T18:48:37.544462Z digest=sha256:8e40243ac6803afe90740e1d7c7e27631b9e51805b449d2cf405cc2c37a1b900

Observation 460713b0-7beb-42ad-921e-862a533385bf · inbound

Reasoning or Memorization? Direction-Aware Diversity Exploration in LLM Reinforcement Learning cites this paper.

Reasoning or Memorization? Direction-Aware Diversity Exploration in LLM Reinforcement Learning Outcome-based Exploration for LLM Reasoning

Reference 10

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T04:47:38.177041Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-06-27T13:39:42.620290Z digest=sha256:04f3ca5bfdc496f8d713481319c4db125c8e9bd16dd29fead38d52b1214fa9d2

Observation 09bd86c5-8aec-4a16-8f1c-288d66fe7ce6 · inbound

When are likely answers right? On Sequence Probability and Correctness in LLMs cites this paper.

When are likely answers right? On Sequence Probability and Correctness in LLMs Outcome-based Exploration for LLM Reasoning

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-07-04T15:09:54.111481Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-06-26T01:55:09.626278Z digest=sha256:c3abef9fae19347972ee3b4f13bb0c6361333f3b855c71c58b9dac5739280951

Observation 9ae9caf7-aae5-4d94-965c-b29c28373d27 · inbound

Depth-Entropy Guided Sampling for Training-Free LLM Reasoning cites this paper.

Depth-Entropy Guided Sampling for Training-Free LLM Reasoning Outcome-based Exploration for LLM Reasoning

Reference 20

Resolution
unresolved
no resolver link, observed 2026-07-14T17:39:27.292075Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T17:39:27.292075Z digest=sha256:2c2cc846a7cbd4d33b3357d68b6ae7f9f989098987bbd0f9c9a389547a21c359

Observation a903df79-6c71-478c-beec-f75b19897c0a · inbound

It Takes 8 Tokens: Weak-to-Strong Off-Policy RL via Auxiliary Branches cites this paper.

It Takes 8 Tokens: Weak-to-Strong Off-Policy RL via Auxiliary Branches Outcome-based Exploration for LLM Reasoning

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-02T14:45:24.695248Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T14:45:24.695248Z digest=sha256:4af72c8e0d209b5784cd8e7e2dede1c05abd417c2317d9af9d4d2eeff3c6da92

Observation 48f9f5bd-daf5-4f16-aa10-d070be0d09a1 · inbound

Bridging Compute- and Data-Optimal Pretraining cites this paper.

Bridging Compute- and Data-Optimal Pretraining Outcome-based Exploration for LLM Reasoning

Reference 107

Resolution
unresolved
no resolver link, observed 2026-08-01T03:02:07.968437Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T03:02:07.968437Z digest=sha256:a2c8c0b612583d7eda28e11ee7782efb5d39cdd9f2807248fb7fe9d2f978e87f

Observation 93ca818a-4402-4f00-a1c5-b184bffdb329 · inbound

Instruction-Conditioned Exploration for Reinforcement Learning with Self-Distillation to an Unconditioned Policy cites this paper.

Instruction-Conditioned Exploration for Reinforcement Learning with Self-Distillation to an Unconditioned Policy Outcome-based Exploration for LLM Reasoning

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-04T15:25:49.290473Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T15:25:49.290473Z digest=sha256:b2d54778cacdbe91e442b89adf5a819107a110014ef83671c779ee4caa9e7ac3

Observation 0751ac2f-e549-4bfc-917f-80d0aead3b5c · inbound

Instruction-Conditioned Exploration for Reinforcement Learning with Self-Distillation to an Unconditioned Policy cites this paper.

Instruction-Conditioned Exploration for Reinforcement Learning with Self-Distillation to an Unconditioned Policy Outcome-based Exploration for LLM Reasoning

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T00:15:16.022969Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:15:16.022969Z digest=sha256:a5549b60372e84c0e734ee4cdd296e298096e0e1d58526732cd5ac0b9e6832de