Pith. sign in

Paper Citation Record · LEDGER

Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning

As of 14 August 2026, this Paper Citation Record lists 72 of 72 outbound references and 16 inbound Pith citation observations for arXiv:2502.06781.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.06781 v1

Coverage vector

measured 72 of 72 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-08T14:27:51.188077Z

measured 88 of 88 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 16 of 16 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:18:54.669121Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

72 of 72 outbound references displayed

  • verified exact1
  • verified fuzzy14
  • unresolved57
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 070b2498-fd8a-4ebf-8468-bee490322b3e · outbound

This paper cites Towards Large Reasoning Models: A Survey of Reinforced Reasoning with Large Language Models.

Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning Towards Large Reasoning Models: A Survey of Reinforced Reasoning with Large Language Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-08T14:27:50.952069Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:27:50.952069Z digest=sha256:b1e9dc01ebf61323d69213fa17a9bf58cfa90414a0df9a6dbf5e04e3c80342a1

Observation 00f685af-9b48-416e-b41c-e48ce97c90f6 · outbound

This paper cites Evaluation of openai o1: Opportunities and challenges of agi.

Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning Evaluation of openai o1: Opportunities and challenges of agi

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-08T14:27:50.956292Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:27:50.956292Z digest=sha256:49b2b93105496b348795c1b61b7ce812d253a0be13a4e37e1cb12f8cfb910c49

Observation e75f2754-1033-4cd8-b473-c3f72e91c319 · outbound

This paper cites Mathematical Language Models: A Survey.

Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning Mathematical Language Models: A Survey

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-08T14:27:50.960542Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:27:50.960542Z digest=sha256:d4dbb01bd01e158e7957c1c9491bf1c3ffc4a404cbf048d1b5fd8ab1b6009fe0

Observation 1d04e19b-9ea1-46b8-8400-d2cdbabb70e5 · outbound

This paper cites Learning mathematics with large language models: A comparative study with computer algebra systems and other tools.

Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning Learning mathematics with large language models: A comparative study with computer algebra systems and other tools

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:27:51.964175Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-08T14:27:50.964513Z digest=sha256:0677580b311df275a39d2da11a0fe8b38ca254fc6a6986bfe1c6e8ad20e0bf19

Observation dd449979-713d-4581-a4a4-02df42ae17fb · outbound

This paper cites InternLM-Math: Open Math Large Language Models Toward Verifiable Reasoning.

Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning InternLM-Math: Open Math Large Language Models Toward Verifiable Reasoning

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-08T14:27:50.968094Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:27:50.968094Z digest=sha256:9d54aeddf30733ce806b73e504b0768400867318b2261c9902387064bcbea7cf

Observation a579b9ce-63bc-4d3a-9636-af3be8dea892 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-08T14:27:50.971842Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:27:50.971842Z digest=sha256:61087239771ff6077890f5efe39b964012dbf81ff80a4be6321490c19b1120c1

Observation 629ab2a0-7625-4222-95cd-2447f843433c · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models.

Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning Chain-of-thought prompting elicits reasoning in large language models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-08T14:27:50.976015Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:27:50.976015Z digest=sha256:b0955eb5ea37e59e63ceb97f129c783ffd64e5b8e86caafad70c3c9937459916

Observation 1ff782f0-5b4d-4fae-b51b-d1721b242208 · outbound

This paper cites Self-Consistency Improves Chain of Thought Reasoning in Language Models.

Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning Self-Consistency Improves Chain of Thought Reasoning in Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-08T14:27:50.979230Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:27:50.979230Z digest=sha256:f5fbb58cd07589a59e23a17c60c6374b77988a70005718e9cd58c936b351632b

Observation d630a42d-af51-475d-86af-69e9aa1d0d52 · outbound

This paper cites Large language models are zero-shot reasoners.

Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning Large language models are zero-shot reasoners

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:27:51.948597Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-08T14:27:50.982287Z digest=sha256:93dc38bf59e3bb99d4477a07212d3f34446484bcaf99aeb5eb0beef71a61ca9e

Observation 8e5601dd-0f1a-4b45-ae29-df4c0c3fbd2d · outbound

This paper cites Learning to reason with llms.

Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning Learning to reason with llms

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-08T14:27:50.985001Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:27:50.985001Z digest=sha256:be6c2760a157534ebb70bde778641387105e8c1de0dde202ab703e2071a7eb7f

Observation 199ead34-9050-4d22-a51a-9b03c6f5914e · outbound

This paper cites O1 Replication Journey -- Part 2: Surpassing O1-preview through Simple Distillation, Big Progress or Bitter Lesson?.

Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning O1 Replication Journey -- Part 2: Surpassing O1-preview through Simple Distillation, Big Progress or Bitter Lesson?

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-08T14:27:50.987846Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:27:50.987846Z digest=sha256:9ec2f91c74bcfd3bc41057ba7d36d862d076373ca2c84a3cdd4806671e85ef1e

Observation 871245c2-a6a3-493c-a54d-f99196ba7204 · outbound

This paper cites Imitate, Explore, and Self-Improve: A Reproduction Report on Slow-thinking Reasoning Systems.

Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning Imitate, Explore, and Self-Improve: A Reproduction Report on Slow-thinking Reasoning Systems

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-08T14:27:50.990868Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:27:50.990868Z digest=sha256:93e3efe5f20f91b8ca39e61de6c61063ae5892244d60b032aa8c1123064dd821

Observation 2ce37bdf-1627-42da-af54-cbfded59f731 · outbound

This paper cites an unresolved cited work.

Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning Unresolved cited work

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-08T14:27:50.993788Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:27:50.993788Z digest=sha256:8a709d78ed73025a3968942d753050051e614a96392b432729c6bb7a23393ee8

Observation 292779e5-b639-4409-a57b-e2699deb09de · outbound

This paper cites DeepSeek-V3 Technical Report.

Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning DeepSeek-V3 Technical Report

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-08T14:27:50.997289Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:27:50.997289Z digest=sha256:30a175e3e2c96d469bfb59185ce55db77860319c49be275e496d6e551236bee8

Observation e2c1c0a8-5866-4a2f-9cb5-ad6154e5b334 · outbound

This paper cites Let's Verify Step by Step.

Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning Let's Verify Step by Step

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-08T14:27:51.000656Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:27:51.000656Z digest=sha256:f99a08aee0fe9776f4b2b4807e61e8ec8f2e5a5c992de6f41728d3bc6faca288

Observation 4e36297a-2947-4184-a310-4a514ddaf916 · outbound

This paper cites Kimi k1.5: Scaling Reinforcement Learning with LLMs.

Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning Kimi k1.5: Scaling Reinforcement Learning with LLMs

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-08T14:27:51.004122Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:27:51.004122Z digest=sha256:99dc7a6a202d5d406da7019a1ad3a05fa723a38fad242206b3b9e29eddc40209

Observation 8d9ae74e-9211-4336-bafe-8df79ea40d07 · outbound

This paper cites Math-shepherd: Verify and reinforce llms step-by-step without human annotations.

Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning Math-shepherd: Verify and reinforce llms step-by-step without human annotations

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:27:51.927158Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-08T14:27:51.007499Z digest=sha256:14c7f4c24c9f6d27b3703bd9e459ead3c0ee90219b983946e42c9f191d2d9dd2

Observation f8d070c7-893d-42df-86e7-3ed7693068de · outbound

This paper cites VinePPO: Refining Credit Assignment in RL Training of LLMs.

Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning VinePPO: Refining Credit Assignment in RL Training of LLMs

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-08T14:27:51.010658Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:27:51.010658Z digest=sha256:715b59ef52eba24a96ea505f36f830beb214d1def6ed5ad1ece46c44483947f3

Observation be090f1b-f199-43b0-baa9-6fbb70e31602 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning Proximal Policy Optimization Algorithms

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-08T14:27:51.014567Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:27:51.014567Z digest=sha256:ae08fe8318e728fe23fd9af450748f31f555e766e531e53e93a2450a159c15c1

Observation 2006b791-2a83-49d0-a968-4226a286814b · outbound

This paper cites Process Reinforcement through Implicit Rewards.

Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning Process Reinforcement through Implicit Rewards

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-08T14:27:51.017933Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:27:51.017933Z digest=sha256:5d8cf0369c7264bf73684e4a041df2bef71db25b77a778660a67dbf3b0847a88

Observation 64d9f09a-5f85-4fa4-88eb-10ec2f19afda · outbound

This paper cites Measuring Mathematical Problem Solving With the MATH Dataset.

Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning Measuring Mathematical Problem Solving With the MATH Dataset

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-08T14:27:51.021439Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:27:51.021439Z digest=sha256:0becbb5204844881b20501ea0620f279768eaad36d3fdcfc06e3c5de5d7a76b1

Observation 42494796-eabd-435e-9f73-f74ef2ad5a6c · outbound

This paper cites Qwq: Reflect deeply on the boundaries of the unknown, November 2024.

Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning Qwq: Reflect deeply on the boundaries of the unknown, November 2024

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-08T14:27:51.024829Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:27:51.024829Z digest=sha256:6615a38e819615d23fea1d028fe69020286594ca696969ecbf35c047ffd138bd

Observation 536d2d7d-acf4-4146-b2a1-09b8dfec53c0 · outbound

This paper cites Scaling Relationship on Learning Mathematical Reasoning with Large Language Models.

Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning Scaling Relationship on Learning Mathematical Reasoning with Large Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-08T14:27:51.027795Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:27:51.027795Z digest=sha256:30c0e725aff7b73f4106ba539f6f35c76e54556a7b6c98a946a3666df2e4e163

Observation e4969442-85fa-4f68-a47d-b11989dcd6b3 · outbound

This paper cites Scaling laws for reward model overoptimization.

Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning Scaling laws for reward model overoptimization

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-08T14:27:51.031167Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:27:51.031167Z digest=sha256:ad6e473888486745a0eaef73e9bd7b7f950ba3f628665d1de0b86d9203b35fe1

Observation 9b12b849-774d-42cf-9696-0417cfc1e1a3 · outbound

This paper cites Compositional preference models for aligning LMs.

Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning Compositional preference models for aligning LMs

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-08T14:27:51.034371Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:27:51.034371Z digest=sha256:e28b05f24ff5b91420ea831e18ee3b81f2dd8f11e9d8019e3fef8bf6356ed569

Observation 936b6233-ed02-4f62-887d-7961a625bba4 · outbound

This paper cites Tulu 3: Pushing Frontiers in Open Language Model Post-Training.

Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning Tulu 3: Pushing Frontiers in Open Language Model Post-Training

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-08T14:27:51.037751Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:27:51.037751Z digest=sha256:d5122e42a9b566abbf101b05578fe39e0510d0d52ea37f1f6e813c105415ff9e

Observation b0a1e810-bdc1-43c1-9dbe-8c19308aea81 · outbound

This paper cites Training language models to follow instructions with human feedback.

Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning Training language models to follow instructions with human feedback

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-08T14:27:51.041315Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:27:51.041315Z digest=sha256:427fc9bf9b37aa1e1c095ea4942bb914c6e3c521348c413b5722832ab2a1bd9c

Observation 60363d63-ef52-40a2-9268-446c007979b5 · outbound

This paper cites Measuring goodhart’s law.

Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning Measuring goodhart’s law

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:27:51.899764Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-08T14:27:51.044512Z digest=sha256:bf05997f68b6cb1fad24fd368fa1d2e5ca9d5002e1ea65e26db2698fe367e862

Observation a7892b10-f28a-4cf0-92e0-b61c143dd331 · outbound

This paper cites Training Language Models with Language Feedback at Scale.

Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning Training Language Models with Language Feedback at Scale

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-08T14:27:51.047668Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:27:51.047668Z digest=sha256:c1e017937c3053b57e8d4454b0be9878e4b83c7c2705bd2f43ca5cba478b7953

Observation 0c3fe5c1-4a8a-434b-865d-c3999a37b8c8 · outbound

This paper cites Reward Model Ensembles Help Mitigate Overoptimization.

Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning Reward Model Ensembles Help Mitigate Overoptimization

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-08T14:27:51.051139Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:27:51.051139Z digest=sha256:66951062ce8b2f31c122cbd250406e8078180328d4ef1946eb738f84c94193c7

Observation 9d9880a1-80ff-40c3-bc4f-2f9990694c83 · outbound

This paper cites BoNBoN Alignment for Large Language Models and the Sweetness of Best-of-n Sampling.

Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning BoNBoN Alignment for Large Language Models and the Sweetness of Best-of-n Sampling

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-08T14:27:51.054553Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:27:51.054553Z digest=sha256:7d3d913647ae05d902ca6d21934c49dc97823ee36b1601ca0d1e95aac5f5bc41

Observation 1d2dd4b9-7af2-4202-8fcf-f2333aee9d00 · outbound

This paper cites BOND: Aligning LLMs with Best-of-N Distillation.

Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning BOND: Aligning LLMs with Best-of-N Distillation

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-08T14:27:51.057910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:27:51.057910Z digest=sha256:5e524476b6b86a32c71e6bceb4631d8ad3c63390b050895ad75362b5e98f2054

Observation 761dece5-630e-4bd8-bafa-bbba8fe28237 · outbound

This paper cites An invariant form for the prior probability in estimation problems.Proceedings of the Royal Society of London.

Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning An invariant form for the prior probability in estimation problems.Proceedings of the Royal Society of London

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:27:51.890501Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-08T14:27:51.060751Z digest=sha256:5fb35cdbb389df64d5479a61e93cb945a363e3790307c72cc8daff34b680fb2d

Observation 829da88e-2bc1-4430-90a8-52240aabd898 · outbound

This paper cites Inference-aware fine-tuning for best-of-n sampling in large language models.

Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning Inference-aware fine-tuning for best-of-n sampling in large language models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-08T14:27:51.063480Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:27:51.063480Z digest=sha256:1fdcc5946f1250484d9d086814d87ef8821c2367e08914f82b88a85dcbc5b945

Observation 1c87d03d-98d3-4564-8e75-04f06a18433d · outbound

This paper cites an unresolved cited work.

Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning Unresolved cited work

Reference 35

Resolution
unresolved
raw_fallback, observed 2026-08-08T14:27:51.881237Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-08T14:27:51.066461Z digest=sha256:c7cae24410443b718e77e1a4f880f7fd769f9da43dd382529d6fb786b83e18e2

Observation 490b5225-c7e4-49b7-bef9-f6842dca9f64 · outbound

This paper cites From $r$ to $Q^*$: Your Language Model is Secretly a Q-Function.

Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning From $r$ to $Q^*$: Your Language Model is Secretly a Q-Function

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-08T14:27:51.069301Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:27:51.069301Z digest=sha256:fb4fae1ed9e1ad8a11dd54e621d33c8c47f115e13774a1563dbc9c4ff0b791d7

Observation ef4dc255-3e0b-4c4f-8de8-5f10a5aa7577 · outbound

This paper cites Inverse-Q*: Token Level Reinforcement Learning for Aligning Large Language Models Without Preference Data.

Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning Inverse-Q*: Token Level Reinforcement Learning for Aligning Large Language Models Without Preference Data

Reference 37

Resolution
verified exact
local_arxiv, observed 2026-08-08T14:27:51.354494Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-08T14:27:51.072172Z digest=sha256:f1031e2379b5dc5c893d028844fb4a3224b7cee127d93c451ad82475c45e6112

Observation b5ba40e0-c1a8-47d0-8eec-f3445f4f1487 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning Training Verifiers to Solve Math Word Problems

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-08T14:27:51.075427Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:27:51.075427Z digest=sha256:2b88e9ed3dc018377c85bb5d278472137a8149ad71e099cab881887b084d5de9

Observation 21c39724-a36f-4aff-aff7-5cf6d8cc758b · outbound

This paper cites Qwen2.5 Technical Report.

Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning Qwen2.5 Technical Report

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-08T14:27:51.079089Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:27:51.079089Z digest=sha256:070612bd54c3131782fd903934caa22ea1ccb3b927d36394a6601e0df079b496

Observation ac75e7e4-c444-4646-8d66-1dac7bdbfad3 · outbound

This paper cites Opendatalab: Empowering general artificial intelligence with open datasets, 2024.

Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning Opendatalab: Empowering general artificial intelligence with open datasets, 2024

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:27:51.872242Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-08T14:27:51.082523Z digest=sha256:18776f6b2a8fcb663462c62240c5c3aea664cbc21402bd47b64b0d806e9d8ccf

Observation d591aa9a-08ca-4501-89e3-220a28935053 · outbound

This paper cites Numinamath: The largest public dataset in ai4maths with 860k pairs of competition math problems and solutions.Hugging Face repository, 13:9, 2024.

Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning Numinamath: The largest public dataset in ai4maths with 860k pairs of competition math problems and solutions.Hugging Face repository, 13:9, 2024

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-08T14:27:51.085801Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:27:51.085801Z digest=sha256:7bf44552913190dd885d867bab775ed96c427fd9996d0fee48fb6dfe04856143

Observation 5a5c329a-4708-48b6-97c2-01a298ccd552 · outbound

This paper cites Hello GPT-4o, 2024.

Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning Hello GPT-4o, 2024

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-08T14:27:51.088874Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:27:51.088874Z digest=sha256:b45f9c63050671016bf70ec874f23966ead252fb4fe669f3ced1f3305639ba50

Observation 42c59246-87cd-4843-af78-27b2e239b241 · outbound

This paper cites Claude 3.5 sonnet, 2024.

Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning Claude 3.5 sonnet, 2024

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-08T14:27:51.092448Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:27:51.092448Z digest=sha256:5e29fa270b5245058abf9a848c5eb828a1bf4bbbae431156da38811d18159631

Observation 338d44e4-3e74-495e-8e48-25503ca63927 · outbound

This paper cites 7b model and 8k examples: Emerging reasoning with reinforcement learning is both effective and efficient.

Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning 7b model and 8k examples: Emerging reasoning with reinforcement learning is both effective and efficient

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:27:51.845013Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-08T14:27:51.095867Z digest=sha256:bb01bfee2f7fa49fc990693ec2ab3c6c2e4b76b729b0b947974dd97671a099bc

Observation 1593a911-7d7b-4aa5-ba85-16fd0f66797b · outbound

This paper cites rStar-Math: Small LLMs Can Master Math Reasoning with Self-Evolved Deep Thinking.

Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning rStar-Math: Small LLMs Can Master Math Reasoning with Self-Evolved Deep Thinking

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-08T14:27:51.099046Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:27:51.099046Z digest=sha256:f50e661a0257f56f7e3252a135b99f12ea10d13294f8bc9a37088805fa6288f0

Observation 089e71d6-00d3-40e5-99ae-c03ea1fd1b18 · outbound

This paper cites American invitational mathematics examination - aime.

Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning American invitational mathematics examination - aime

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:27:51.835134Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-08T14:27:51.102557Z digest=sha256:161785c6050ae67f8d466bac77081e8f3ccaf8af02728a3070f23b1b745bd4bd

Observation 6bf0ca30-56b9-4c38-8e8c-5330ed4184c9 · outbound

This paper cites Are Your LLMs Capable of Stable Reasoning?.

Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning Are Your LLMs Capable of Stable Reasoning?

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-08T14:27:51.106013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:27:51.106013Z digest=sha256:6779f4b89bf2fe5f25d87390571bdf7bbeeb4157d9dfedb6a97a6f7c7bf151e9

Observation 31d1c947-f4ef-4c66-85fe-d0e7361341e5 · outbound

This paper cites OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems.

Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-08T14:27:51.109531Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:27:51.109531Z digest=sha256:6bb052edd2e0e92d5d8792ea47f738ff1a095fa30f447aaa218f7948bcf17a4a

Observation a2fbfb90-8f35-4cb3-841a-923edbf8e8c1 · outbound

This paper cites Opencompass: A universal evaluation platform for foundation models.

Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning Opencompass: A universal evaluation platform for foundation models

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-08T14:27:51.112911Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:27:51.112911Z digest=sha256:1ec114d1f34623cc2066d6c62c94fdaa02dee911baf33bdfb45ebf3fa9ca22d9

Observation 92d69a59-5df4-483e-a0c9-aa54a3d0e38a · outbound

This paper cites Policy gradient meth- ods for reinforcement learning with function approximation.

Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning Policy gradient meth- ods for reinforcement learning with function approximation

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-08T14:27:51.116126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:27:51.116126Z digest=sha256:8afd1fc6d225e65fbb182130247d352f7eefeafd6db3118df022920591641624

Observation 8c06dc8d-24e4-4e83-830e-513fd2a8a2e0 · outbound

This paper cites Tree of thoughts: Deliberate problem solving with large language models.

Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning Tree of thoughts: Deliberate problem solving with large language models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-08T14:27:51.119909Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:27:51.119909Z digest=sha256:efaff28e8df26f890c4cd0f967b7337eb214dd7d9ed504412ee34a4936ec4c9e

Observation bb206363-c2c9-43d6-8c8d-729da4200dd5 · outbound

This paper cites Graph of thoughts: Solving elaborate problems with large language models.

Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning Graph of thoughts: Solving elaborate problems with large language models

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:27:51.808050Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-08T14:27:51.123115Z digest=sha256:cc11e40e246936491c7c51ecb9dbf486e028feeec9cde7454ae04aeb6327f0e7

Observation 9bfc999a-6fb0-4875-bee1-12abb2e7b0e4 · outbound

This paper cites MetaMath: Bootstrap Your Own Mathematical Questions for Large Language Models.

Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning MetaMath: Bootstrap Your Own Mathematical Questions for Large Language Models

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-08T14:27:51.126152Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:27:51.126152Z digest=sha256:45afc4167dde748b6359cbafc9a9832621257e18e39910dc08a1cbafc0e12b34

Observation cf905594-0d06-4889-9173-68a756270331 · outbound

This paper cites Augmenting Math Word Problems via Iterative Question Composing.

Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning Augmenting Math Word Problems via Iterative Question Composing

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-08T14:27:51.129269Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:27:51.129269Z digest=sha256:903db329800f9d6ecb5fc884caae82f2911233e92734cf20dc51b35a166aeead

Observation 9c17fa23-b188-44c1-88b5-7a3002de14d0 · outbound

This paper cites MuggleMath: Assessing the Impact of Query and Response Augmentation on Math Reasoning.

Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning MuggleMath: Assessing the Impact of Query and Response Augmentation on Math Reasoning

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-08T14:27:51.132425Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:27:51.132425Z digest=sha256:a4e7e738377082240a3501df29110121bc8fc9d028b0ff4353782a78d42977ae

Observation edd0dc10-fc9a-4e71-9e4e-f9acf6ecc1da · outbound

This paper cites Goat: Fine-tuned LLaMA Outperforms GPT-4 on Arithmetic Tasks.

Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning Goat: Fine-tuned LLaMA Outperforms GPT-4 on Arithmetic Tasks

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-08T14:27:51.135706Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:27:51.135706Z digest=sha256:1fd4e321ae72956d57e1d36c5d625c4be15c4bc2254cc5f2e239b25e2153b6cf

Observation a02c956f-67b4-4333-bf1e-e649ece2d74c · outbound

This paper cites MAmmoTH: Building Math Generalist Models through Hybrid Instruction Tuning.

Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning MAmmoTH: Building Math Generalist Models through Hybrid Instruction Tuning

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-08T14:27:51.138537Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:27:51.138537Z digest=sha256:80d4e320cb68cc9e44c642bea2c937436eee523a0ad9b74bdbf90c5d090fa37f

Observation ccde4491-3225-4820-94f6-b3dbd1c7d9ab · outbound

This paper cites SFT Memorizes, RL Generalizes: A Comparative Study of Foundation Model Post-training.

Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning SFT Memorizes, RL Generalizes: A Comparative Study of Foundation Model Post-training

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-08T14:27:51.141898Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:27:51.141898Z digest=sha256:942bf066d9c71617444441662d61f1a357bdb925f5022b46fd302a0196e27868

Observation 021f7021-240a-4ea6-b091-9b24fbe92bfc · outbound

This paper cites Can large language models reason and plan? Annals of the New York Academy of Sciences, 1534(1):15–18, 2024.

Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning Can large language models reason and plan? Annals of the New York Academy of Sciences, 1534(1):15–18, 2024

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:27:51.798417Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-08T14:27:51.145468Z digest=sha256:0d74de98f3ccc97df365ea42ed5f6d4b1ceeb7bbde52d79faae0711b7bcd53ed

Observation ab3810a8-8362-4285-8482-3cc08241ea9c · outbound

This paper cites WizardMath: Empowering Mathematical Reasoning for Large Language Models via Reinforced Evol-Instruct.

Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning WizardMath: Empowering Mathematical Reasoning for Large Language Models via Reinforced Evol-Instruct

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-08T14:27:51.148684Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:27:51.148684Z digest=sha256:9f78caed129be57cf67d8b7662ea06bf10964f91ce8ef86cdf5a5a95cca4b0d5

Observation 2c04f306-b46e-4265-ba4b-088a53fde1bc · outbound

This paper cites Beyond Human Data: Scaling Self-Training for Problem-Solving with Language Models.

Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning Beyond Human Data: Scaling Self-Training for Problem-Solving with Language Models

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-08T14:27:51.152051Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:27:51.152051Z digest=sha256:04bfb1fddcc0b7bcfa19ebfcd5723f0365acfe814ed8928349d7bc82baf49bcb

Observation 36414ed8-e16a-4d3a-9544-5ce4ca0b8136 · outbound

This paper cites Back to Basics: Revisiting REINFORCE Style Optimization for Learning from Human Feedback in LLMs.

Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning Back to Basics: Revisiting REINFORCE Style Optimization for Learning from Human Feedback in LLMs

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-08T14:27:51.155707Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:27:51.155707Z digest=sha256:52076fa756125cbef2e954eb6d91927d7e69f04824cd63baacf8dc7a1cfb4d0b

Observation db660ffc-25a4-402b-87d3-05a8d08395be · outbound

This paper cites You don’t need to re-generate the answer to the question because the standard answer has been given.

Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning You don’t need to re-generate the answer to the question because the standard answer has been given

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:27:51.787477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-08T14:27:51.159064Z digest=sha256:59fb0ddefb0d10017dd370470f5225cad3aa6969eb08d8f42059d8693f25935a

Observation 2717bfa1-4154-49af-a700-1ba82d031236 · outbound

This paper cites an unresolved cited work.

Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning Unresolved cited work

Reference 64

Resolution
unresolved
raw_fallback, observed 2026-08-08T14:27:51.777395Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-08T14:27:51.162379Z digest=sha256:e72947ac6929bca03008090ea96efbd30fb3b8a80620a333e0d57ed663fcfe36

Observation a0ff2ceb-134e-4708-947a-ab645d6c9128 · outbound

This paper cites As long as the answer is the same as the standard answer, it is enough.

Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning As long as the answer is the same as the standard answer, it is enough

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:27:51.767856Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-08T14:27:51.165691Z digest=sha256:7e2639f7b94051b4721d57c8526545d8aef9eb2378c22dc8f99baaaba2bd951a

Observation 995844ec-b558-4489-9c11-f784e2b4644a · outbound

This paper cites And some formulas are expressed in different ways, but they are equivalent and correct.

Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning And some formulas are expressed in different ways, but they are equivalent and correct

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:27:51.758301Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-08T14:27:51.168809Z digest=sha256:c5b7cef0375b4efc89766e4dc8334db26f4ca9735a16611039e4764f8d337898

Observation d9207753-5c50-4664-a353-bcd4a1e1babf · outbound

This paper cites A" or "B.

Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning A" or "B

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:27:51.748464Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-08T14:27:51.172027Z digest=sha256:d69b3c77891eb4a6ff468694de6155aee646c427e8b63810bff3d2683ee1dc5d

Observation 39f35344-4083-4e5e-8070-4a94cc29d417 · outbound

This paper cites an unresolved cited work.

Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning Unresolved cited work

Reference 68

Resolution
unresolved
raw_fallback, observed 2026-08-08T14:27:51.737774Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-08T14:27:51.175657Z digest=sha256:651dca7f7d48964aaec514df7da4aa6665c5657e68b12d5f7349a2e566678e93

Observation 644abfb2-54e4-4cb4-8385-84f39da5b153 · outbound

This paper cites an unresolved cited work.

Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning Unresolved cited work

Reference 69

Resolution
unresolved
raw_fallback, observed 2026-08-08T14:27:51.728145Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-08T14:27:51.178803Z digest=sha256:4cddcee04eacb987206db5560731665df107141fee787fabe68835b6f5b2d0bb

Observation 88c346e6-9494-44cb-ba24-e65330c9b28c · outbound

This paper cites an unresolved cited work.

Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning Unresolved cited work

Reference 70

Resolution
unresolved
raw_fallback, observed 2026-08-08T14:27:51.717965Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-08T14:27:51.181847Z digest=sha256:0260afcd908638391bbe0af940729ece504104629f01a728c1f0d6371dd84f66

Observation eb0d0544-e8d6-466c-85c9-256d984231ff · outbound

This paper cites an unresolved cited work.

Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning Unresolved cited work

Reference 71

Resolution
unresolved
raw_fallback, observed 2026-08-08T14:27:51.708176Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-08T14:27:51.185024Z digest=sha256:c0dfab34b4e05f414ad940fd3faebc096ff67ff0a76afefd2b76708ce8c1d0ff

Observation 0720b3e7-602e-4467-ac25-8c1517554c68 · outbound

This paper cites an unresolved cited work.

Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning Unresolved cited work

Reference 72

Resolution
unresolved
raw_fallback, observed 2026-08-08T14:27:51.698233Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-08T14:27:51.188077Z digest=sha256:cf742846b49aadebb1d2b09db4459cd080c891f8f3e6708ffc3fed33ab17d1ab

Pith citing papers

Observation 40a1032b-0f5d-4e52-8a77-687f5ff1de13 · inbound

Reinforcement Learning from Human Feedback cites this paper.

Reinforcement Learning from Human Feedback Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning

Reference 77

Resolution
verified exact
arxiv_id, observed 2026-05-22T19:32:01.345899Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-22T19:27:40.991325Z digest=sha256:a190d4bc4fe26c984fbe317e99753bc613cf87b2bd18fee93966e9a72b587d9c

Observation 6a317298-8807-4244-bf64-974ce588f354 · inbound

VARD: Efficient and Dense Fine-Tuning for Diffusion Models with Value-based RL cites this paper.

VARD: Efficient and Dense Fine-Tuning for Diffusion Models with Value-based RL Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T15:18:54.669121Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:18:54.669121Z digest=sha256:6be3a421b89ab050d0a203956de233c9ba8049c792c01d6b72d1a5902ee5f24a

Observation 04b7eaec-cacb-4cbb-85b3-b00ad11efe56 · inbound

RFTF: Reinforcement Fine-tuning for Embodied Agents with Temporal Feedback cites this paper.

RFTF: Reinforcement Fine-tuning for Embodied Agents with Temporal Feedback Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T14:10:12.939491Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:10:12.939491Z digest=sha256:0c399f5da5a9fabc524ff730c8b3a0f5cf2c75f8a7cb64c7591569e42b5f1206

Observation 747e0839-1ec3-49c4-b4f8-1e95d895c7b6 · inbound

Training LLMs for EHR-Based Reasoning Tasks via Reinforcement Learning cites this paper.

Training LLMs for EHR-Based Reasoning Tasks via Reinforcement Learning Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T12:40:09.933663Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:40:09.933663Z digest=sha256:8f6facb46b714e20a8f805df45f4021f58aad7aad53d23ff924e6ad3f1015f30

Observation 867e3e6f-9b30-47c1-b77b-cd076a11bcef · inbound

Response-Level Rewards Are All You Need for Online Reinforcement Learning in LLMs: A Mathematical Perspective cites this paper.

Response-Level Rewards Are All You Need for Online Reinforcement Learning in LLMs: A Mathematical Perspective Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:31.527587Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:31.527587Z digest=sha256:542c79b2531844d5e2adb0f1b317dda77a9bda20ccb42c12a7e9d3dc05353392

Observation 59a01924-d375-4698-afa9-7e01cd44e84c · inbound

Consistent Paths Lead to Truth: Self-Rewarding Reinforcement Learning for LLM Reasoning cites this paper.

Consistent Paths Lead to Truth: Self-Rewarding Reinforcement Learning for LLM Reasoning Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T05:09:44.200132Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:09:44.200132Z digest=sha256:83a0ff3bb7ef060826ec11eec901aac114aa879ac46f73a6aaee54d0741876dd

Observation 69dcfd01-d48e-4d21-85dd-be3372de2929 · inbound

Unleashing Embodied Task Planning Ability in LLMs via Reinforcement Learning cites this paper.

Unleashing Embodied Task Planning Ability in LLMs via Reinforcement Learning Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T21:55:01.685340Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:55:01.685340Z digest=sha256:e025de49491fd12a8a7be694008808cccf6d1f73b69030f0708a36d4424c5e04

Observation 1845ae5b-eaf4-43c4-a421-8b0b14f25fd2 · inbound

Test-Time Scaling with Reflective Generative Model cites this paper.

Test-Time Scaling with Reflective Generative Model Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T20:46:44.783885Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:46:44.783885Z digest=sha256:7feff76355e8dfce5864831ba3c33cd1b288ec314361ca3c5faebbb0924d8e14

Observation 9d18271c-db52-48d0-a8b3-0f038b000880 · inbound

A Technical Survey of Reinforcement Learning Techniques for Large Language Models cites this paper.

A Technical Survey of Reinforcement Learning Techniques for Large Language Models Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T19:59:28.348022Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:59:28.348022Z digest=sha256:549080b68088afd1e5f318625341178686bbd3de6b37ac5487e2229aff9c7983

Observation 40207b04-b788-474d-8ee3-8edcf2b3f3de · inbound

Alignment and Safety in Large Language Models: Safety Mechanisms, Training Paradigms, and Emerging Challenges cites this paper.

Alignment and Safety in Large Language Models: Safety Mechanisms, Training Paradigms, and Emerging Challenges Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning

Reference 198

Resolution
unresolved
no resolver link, observed 2026-08-06T14:13:06.979201Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:13:06.979201Z digest=sha256:33d1cdbb7762f56500273562f46ab37236e833ea0e79f015815b5ec3d7e3be6c

Observation 9eaa3cc7-49dd-4760-9e89-0b0c97512861 · inbound

Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models cites this paper.

Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-05T20:21:08.144794Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:21:08.144794Z digest=sha256:c55f84b04649dceedcc72bc1ce06fa128e277b7957c9c18501a6de50b10a72c5

Observation 36284a2e-89b8-4e03-ac8a-75c39a212196 · inbound

Uncertainty Under the Curve: A Sequence-Level Entropy Area Metric for Reasoning LLM cites this paper.

Uncertainty Under the Curve: A Sequence-Level Entropy Area Metric for Reasoning LLM Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-05T15:11:01.093701Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:11:01.093701Z digest=sha256:61b43b4dea4ec6477ed4c383f94e1c52816aeb7d3f7b379748b02b7d7e49e792

Observation da937472-cdc1-490a-beb9-4b79cd4fbc57 · inbound

Intern-S1-MO: Long-horizon Reasoning Agent for Olympiad?Level Mathematical Problem Solving cites this paper.

Intern-S1-MO: Long-horizon Reasoning Agent for Olympiad?Level Mathematical Problem Solving Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-04T06:40:24.891449Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:40:24.891449Z digest=sha256:746d49ffb7dece3f63e1589a7ad0a9ee3b71c00a2bb498fa3825832e769d412c

Observation cc15766a-f859-4eaa-9b3b-82060a08db9c · inbound

Reward Modeling for Reinforcement Learning-Based LLM Reasoning: Design, Challenges, and Evaluation cites this paper.

Reward Modeling for Reinforcement Learning-Based LLM Reasoning: Design, Challenges, and Evaluation Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-03T03:04:43.981554Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:04:43.981554Z digest=sha256:660d5f7410869f188906d78c6a27f77b30ee3c8640ac75ddf0ece02ede88c929

Observation 6e7d4cec-4529-43b9-bac8-67b44a0416a9 · inbound

Trust Region On-Policy Distillation cites this paper.

Trust Region On-Policy Distillation Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning

Reference 224

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T20:56:13.534551Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-06-28T17:38:50.313305Z digest=sha256:cde0eb26b1568b4aea870d7edf32585665e1c05c1331fd778f9fe5223bf8efa2

Observation 7bc710db-20a8-4c95-a0e0-c252eb3cb39f · inbound

The Periodic Table of LLM Reasoning: A Structured Survey of Reasoning Paradigms, Methods, and Failure Modes cites this paper.

The Periodic Table of LLM Reasoning: A Structured Survey of Reasoning Paradigms, Methods, and Failure Modes Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning

Reference 165

Resolution
verified exact
arxiv_id, observed 2026-06-27T13:00:56.109782Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-06-27T12:59:51.091008Z digest=sha256:47239691e381fca89fa884569602495501fa37beeaaa21bd5f04d50cd63a6868