Pith. sign in

Paper Citation Record · LEDGER

Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning

As of 8 August 2026, this Paper Citation Record lists 72 of 72 outbound references and 16 inbound Pith citation observations for arXiv:2502.06781.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.06781 v1

Coverage vector

measured 72 of 72 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-08T14:27:51.188077Z

measured 88 of 88 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 16 of 16 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:18:54.669121Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

72 of 72 outbound references displayed

  • verified exact1
  • verified fuzzy14
  • unresolved57
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 070b2498-fd8a-4ebf-8468-bee490322b3e · outbound

This paper cites Towards Large Reasoning Models: A Survey of Reinforced Reasoning with Large Language Models.

Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning Towards Large Reasoning Models: A Survey of Reinforced Reasoning with Large Language Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-08T14:27:50.952069Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:27:50.952069Z digest=sha256:aa2b45a7c2754fcf6ac1afccaf5e48e09df8ca735be121aa4a0d3006f64ba516

Observation 00f685af-9b48-416e-b41c-e48ce97c90f6 · outbound

This paper cites Evaluation of openai o1: Opportunities and challenges of agi.

Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning Evaluation of openai o1: Opportunities and challenges of agi

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-08T14:27:50.956292Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:27:50.956292Z digest=sha256:3b42de0f896ac0285e7a78545a757ef7ce6d4359445836d7f1c09d750f1fc643

Observation e75f2754-1033-4cd8-b473-c3f72e91c319 · outbound

This paper cites Mathematical Language Models: A Survey.

Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning Mathematical Language Models: A Survey

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-08T14:27:50.960542Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:27:50.960542Z digest=sha256:a146828e5b9d78b4832da94aa600cbbe3c65242b4ec962129956d959120ed141

Observation 1d04e19b-9ea1-46b8-8400-d2cdbabb70e5 · outbound

This paper cites Learning mathematics with large language models: A comparative study with computer algebra systems and other tools.

Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning Learning mathematics with large language models: A comparative study with computer algebra systems and other tools

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:27:51.964175Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T14:27:50.964513Z digest=sha256:a75f7a8fa2ceae734701b44d6073fccacd16f632e65c21c4eb42a05ad8e4bf5b

Observation dd449979-713d-4581-a4a4-02df42ae17fb · outbound

This paper cites InternLM-Math: Open Math Large Language Models Toward Verifiable Reasoning.

Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning InternLM-Math: Open Math Large Language Models Toward Verifiable Reasoning

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-08T14:27:50.968094Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:27:50.968094Z digest=sha256:3a4188873d489893ded411518931d47dba2c8d23ca97b515008b69bec969ba42

Observation a579b9ce-63bc-4d3a-9636-af3be8dea892 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-08T14:27:50.971842Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:27:50.971842Z digest=sha256:8cd73b38cc791d0d8e76ca361b5d5a9e8c7227fa30132c1097849e3da368a820

Observation 629ab2a0-7625-4222-95cd-2447f843433c · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models.

Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning Chain-of-thought prompting elicits reasoning in large language models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-08T14:27:50.976015Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:27:50.976015Z digest=sha256:e07dbba20c545e884cad1be84345f4684e5175b6bf80d328803795949f27e5d4

Observation 1ff782f0-5b4d-4fae-b51b-d1721b242208 · outbound

This paper cites Self-Consistency Improves Chain of Thought Reasoning in Language Models.

Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning Self-Consistency Improves Chain of Thought Reasoning in Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-08T14:27:50.979230Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:27:50.979230Z digest=sha256:bbbbb78b1716b8cf7d0a37f79ab0628134cea0741ca3f930ec462466dd7ee271

Observation d630a42d-af51-475d-86af-69e9aa1d0d52 · outbound

This paper cites Large language models are zero-shot reasoners.

Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning Large language models are zero-shot reasoners

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:27:51.948597Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T14:27:50.982287Z digest=sha256:e4665852b0c85ba913590a4470e3b5d5760f13a720e2cc210a47ae54fac91c0e

Observation 8e5601dd-0f1a-4b45-ae29-df4c0c3fbd2d · outbound

This paper cites Learning to reason with llms.

Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning Learning to reason with llms

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-08T14:27:50.985001Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:27:50.985001Z digest=sha256:277c259b7421d73dc27e8d9bd124ecc9a1fe207021bb30e831943f5ef0dd93b6

Observation 199ead34-9050-4d22-a51a-9b03c6f5914e · outbound

This paper cites O1 Replication Journey -- Part 2: Surpassing O1-preview through Simple Distillation, Big Progress or Bitter Lesson?.

Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning O1 Replication Journey -- Part 2: Surpassing O1-preview through Simple Distillation, Big Progress or Bitter Lesson?

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-08T14:27:50.987846Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:27:50.987846Z digest=sha256:a67e3eb142a746678bb8ec0d126591d3e7ef7cf7e04b6a2a8237a2bd00f74743

Observation 871245c2-a6a3-493c-a54d-f99196ba7204 · outbound

This paper cites Imitate, Explore, and Self-Improve: A Reproduction Report on Slow-thinking Reasoning Systems.

Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning Imitate, Explore, and Self-Improve: A Reproduction Report on Slow-thinking Reasoning Systems

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-08T14:27:50.990868Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:27:50.990868Z digest=sha256:cb9ce1c8a90bfd7fd4fdb2da6f008ba210aadd4f5db5cf79bbc6acd1e8619868

Observation 2ce37bdf-1627-42da-af54-cbfded59f731 · outbound

This paper cites an unresolved cited work.

Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning Unresolved cited work

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-08T14:27:50.993788Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:27:50.993788Z digest=sha256:4a3a275b6ef918ea1b240abdbed3fd99f68dea0e4ddc70f30630b41571d6f337

Observation 292779e5-b639-4409-a57b-e2699deb09de · outbound

This paper cites DeepSeek-V3 Technical Report.

Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning DeepSeek-V3 Technical Report

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-08T14:27:50.997289Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:27:50.997289Z digest=sha256:428562eebdc0db7b98c647ed5f84b2da44df392b6308a031ed2093bc592eaeec

Observation e2c1c0a8-5866-4a2f-9cb5-ad6154e5b334 · outbound

This paper cites Let's Verify Step by Step.

Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning Let's Verify Step by Step

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-08T14:27:51.000656Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:27:51.000656Z digest=sha256:a7368475dd7462c6c39e62eac8f241ef8699f6f727ec1b53c2eca6b945c46d1d

Observation 4e36297a-2947-4184-a310-4a514ddaf916 · outbound

This paper cites Kimi k1.5: Scaling Reinforcement Learning with LLMs.

Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning Kimi k1.5: Scaling Reinforcement Learning with LLMs

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-08T14:27:51.004122Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:27:51.004122Z digest=sha256:f42818e9dd42b7b0f141c3a70a35a1ede3073abb38f8b117e015fd6f099c411d

Observation 8d9ae74e-9211-4336-bafe-8df79ea40d07 · outbound

This paper cites Math-shepherd: Verify and reinforce llms step-by-step without human annotations.

Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning Math-shepherd: Verify and reinforce llms step-by-step without human annotations

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:27:51.927158Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T14:27:51.007499Z digest=sha256:5137ea3bf931869d2df53c0ae0d538121ed4355228c38cfdf5fb6bbbac12753e

Observation f8d070c7-893d-42df-86e7-3ed7693068de · outbound

This paper cites VinePPO: Refining Credit Assignment in RL Training of LLMs.

Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning VinePPO: Refining Credit Assignment in RL Training of LLMs

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-08T14:27:51.010658Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:27:51.010658Z digest=sha256:872637cad1667502e689c0a94dc8bb1f3292f864093b5f7479b16ce56d2908aa

Observation be090f1b-f199-43b0-baa9-6fbb70e31602 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning Proximal Policy Optimization Algorithms

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-08T14:27:51.014567Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:27:51.014567Z digest=sha256:c92a4db09cab75d5852a2e6449b7584624ac071f6eb7c3acf4c948acb5a3c72a

Observation 2006b791-2a83-49d0-a968-4226a286814b · outbound

This paper cites Process Reinforcement through Implicit Rewards.

Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning Process Reinforcement through Implicit Rewards

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-08T14:27:51.017933Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:27:51.017933Z digest=sha256:d890bb7373d91fda898d8537ffb89bbafe7e9cac8a015ce72e30818ab58ad611

Observation 64d9f09a-5f85-4fa4-88eb-10ec2f19afda · outbound

This paper cites Measuring Mathematical Problem Solving With the MATH Dataset.

Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning Measuring Mathematical Problem Solving With the MATH Dataset

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-08T14:27:51.021439Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:27:51.021439Z digest=sha256:f232e66ec9f7445dd2ac0a57e504c9605fd4e8ddc75d3841d72fb627aba35beb

Observation 42494796-eabd-435e-9f73-f74ef2ad5a6c · outbound

This paper cites Qwq: Reflect deeply on the boundaries of the unknown, November 2024.

Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning Qwq: Reflect deeply on the boundaries of the unknown, November 2024

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-08T14:27:51.024829Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:27:51.024829Z digest=sha256:91c6da2b5f6a510e1259fec3231bdad07a5f4c3b120da55d3d81d4396a92c71f

Observation 536d2d7d-acf4-4146-b2a1-09b8dfec53c0 · outbound

This paper cites Scaling Relationship on Learning Mathematical Reasoning with Large Language Models.

Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning Scaling Relationship on Learning Mathematical Reasoning with Large Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-08T14:27:51.027795Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:27:51.027795Z digest=sha256:1da595e81518aa962a3c8462cb28b85f76e3a89a7e40743f08ba63fb6676e54f

Observation e4969442-85fa-4f68-a47d-b11989dcd6b3 · outbound

This paper cites Scaling laws for reward model overoptimization.

Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning Scaling laws for reward model overoptimization

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-08T14:27:51.031167Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:27:51.031167Z digest=sha256:d54ee3ae6d733a3b171d7a9bd331d89a1fe3df7771ad3bcd09bcd2022d29ca19

Observation 9b12b849-774d-42cf-9696-0417cfc1e1a3 · outbound

This paper cites Compositional preference models for aligning LMs.

Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning Compositional preference models for aligning LMs

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-08T14:27:51.034371Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:27:51.034371Z digest=sha256:58ee0a4918842da30f73b2447661f5006b1e6442917060547dba64bc8e930d72

Observation 936b6233-ed02-4f62-887d-7961a625bba4 · outbound

This paper cites Tulu 3: Pushing Frontiers in Open Language Model Post-Training.

Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning Tulu 3: Pushing Frontiers in Open Language Model Post-Training

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-08T14:27:51.037751Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:27:51.037751Z digest=sha256:b27533bb57a9be8b60f228c6f4469ae2b9d7ff94263b107d42694df820285e35

Observation b0a1e810-bdc1-43c1-9dbe-8c19308aea81 · outbound

This paper cites Training language models to follow instructions with human feedback.

Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning Training language models to follow instructions with human feedback

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-08T14:27:51.041315Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:27:51.041315Z digest=sha256:aaf29c2db6c04d96600002cfd2fca1d6592343fc62fae4962a42c0cb8e966dea

Observation 60363d63-ef52-40a2-9268-446c007979b5 · outbound

This paper cites Measuring goodhart’s law.

Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning Measuring goodhart’s law

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:27:51.899764Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T14:27:51.044512Z digest=sha256:a909be9cf0147ffaf3b1473b81a9aec6815ea748807bedfa53178eeae18885c7

Observation a7892b10-f28a-4cf0-92e0-b61c143dd331 · outbound

This paper cites Training Language Models with Language Feedback at Scale.

Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning Training Language Models with Language Feedback at Scale

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-08T14:27:51.047668Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:27:51.047668Z digest=sha256:22b35f1db7f42b4eb43bbd3ad83145e6cdcd6e83d0b578114c9d5bb94929ec2a

Observation 0c3fe5c1-4a8a-434b-865d-c3999a37b8c8 · outbound

This paper cites Reward Model Ensembles Help Mitigate Overoptimization.

Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning Reward Model Ensembles Help Mitigate Overoptimization

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-08T14:27:51.051139Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:27:51.051139Z digest=sha256:b03e30001454b1c636146b105f6c6486ec9d98d73b4fe2f04f776f50414ed1ca

Observation 9d9880a1-80ff-40c3-bc4f-2f9990694c83 · outbound

This paper cites BoNBoN Alignment for Large Language Models and the Sweetness of Best-of-n Sampling.

Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning BoNBoN Alignment for Large Language Models and the Sweetness of Best-of-n Sampling

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-08T14:27:51.054553Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:27:51.054553Z digest=sha256:0ab6a08922a19d5541f9588e0b405c951808831bb9d91f7e7698c3443c15b725

Observation 1d2dd4b9-7af2-4202-8fcf-f2333aee9d00 · outbound

This paper cites BOND: Aligning LLMs with Best-of-N Distillation.

Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning BOND: Aligning LLMs with Best-of-N Distillation

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-08T14:27:51.057910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:27:51.057910Z digest=sha256:97af6db259d0034e16e8503fa2b0fdee84bcb78e95673e57e918ccd2e3c38e30

Observation 761dece5-630e-4bd8-bafa-bbba8fe28237 · outbound

This paper cites An invariant form for the prior probability in estimation problems.Proceedings of the Royal Society of London.

Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning An invariant form for the prior probability in estimation problems.Proceedings of the Royal Society of London

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:27:51.890501Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T14:27:51.060751Z digest=sha256:953f7e705bc8787b1d6b9e9906e71e8b414c2ee80954cdb669139d29ff17cf8f

Observation 829da88e-2bc1-4430-90a8-52240aabd898 · outbound

This paper cites Inference-aware fine-tuning for best-of-n sampling in large language models.

Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning Inference-aware fine-tuning for best-of-n sampling in large language models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-08T14:27:51.063480Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:27:51.063480Z digest=sha256:ec8a280e3e4d4c1eb55effeaae6434a21cbce43a94fa3c14e052a1e6822f5be2

Observation 1c87d03d-98d3-4564-8e75-04f06a18433d · outbound

This paper cites an unresolved cited work.

Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning Unresolved cited work

Reference 35

Resolution
unresolved
raw_fallback, observed 2026-08-08T14:27:51.881237Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T14:27:51.066461Z digest=sha256:b9af21a797fee9592cd99e06d9a091ad094b22bf54e8229d6771b49b93372519

Observation 490b5225-c7e4-49b7-bef9-f6842dca9f64 · outbound

This paper cites From $r$ to $Q^*$: Your Language Model is Secretly a Q-Function.

Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning From $r$ to $Q^*$: Your Language Model is Secretly a Q-Function

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-08T14:27:51.069301Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:27:51.069301Z digest=sha256:2e6b0bed1bb8fba29557b1646c6aa7fa9d898574ab895c8bf5b850334ac6cc36

Observation ef4dc255-3e0b-4c4f-8de8-5f10a5aa7577 · outbound

This paper cites Inverse-Q*: Token Level Reinforcement Learning for Aligning Large Language Models Without Preference Data.

Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning Inverse-Q*: Token Level Reinforcement Learning for Aligning Large Language Models Without Preference Data

Reference 37

Resolution
verified exact
local_arxiv, observed 2026-08-08T14:27:51.354494Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T14:27:51.072172Z digest=sha256:f3081d1ec33d14e9bc293c73add80fc9c8e5189589190311f3812733dcec4393

Observation b5ba40e0-c1a8-47d0-8eec-f3445f4f1487 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning Training Verifiers to Solve Math Word Problems

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-08T14:27:51.075427Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:27:51.075427Z digest=sha256:e3b9ad4510bdf7a1597bd8359a862873a9b3548580a3458067c30449c70ebfbe

Observation 21c39724-a36f-4aff-aff7-5cf6d8cc758b · outbound

This paper cites Qwen2.5 Technical Report.

Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning Qwen2.5 Technical Report

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-08T14:27:51.079089Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:27:51.079089Z digest=sha256:597791412e2b3ccf3f8f78dcc1544097f923fb481fb03ded1a5b866a08891f79

Observation ac75e7e4-c444-4646-8d66-1dac7bdbfad3 · outbound

This paper cites Opendatalab: Empowering general artificial intelligence with open datasets, 2024.

Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning Opendatalab: Empowering general artificial intelligence with open datasets, 2024

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:27:51.872242Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T14:27:51.082523Z digest=sha256:d6cf3dc5199855472b1292a05646c276558f5c4cde48dbf58fe7773440d600d9

Observation d591aa9a-08ca-4501-89e3-220a28935053 · outbound

This paper cites Numinamath: The largest public dataset in ai4maths with 860k pairs of competition math problems and solutions.Hugging Face repository, 13:9, 2024.

Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning Numinamath: The largest public dataset in ai4maths with 860k pairs of competition math problems and solutions.Hugging Face repository, 13:9, 2024

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-08T14:27:51.085801Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:27:51.085801Z digest=sha256:622c16069b16020f9cdccfb5eef1ef06ef9b41d96bafb6bae749557a3ced8f64

Observation 5a5c329a-4708-48b6-97c2-01a298ccd552 · outbound

This paper cites Hello GPT-4o, 2024.

Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning Hello GPT-4o, 2024

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-08T14:27:51.088874Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:27:51.088874Z digest=sha256:b71eb7b28c8929f759671c5d1b2d07f3bece1c83ace97f3c42dc28e8eb43cb63

Observation 42c59246-87cd-4843-af78-27b2e239b241 · outbound

This paper cites Claude 3.5 sonnet, 2024.

Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning Claude 3.5 sonnet, 2024

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-08T14:27:51.092448Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:27:51.092448Z digest=sha256:ee55f64d8dbbb2c7344bcfbefa272a9ea35e3d3e679fbf6b8ec9d42fe7c36e10

Observation 338d44e4-3e74-495e-8e48-25503ca63927 · outbound

This paper cites 7b model and 8k examples: Emerging reasoning with reinforcement learning is both effective and efficient.

Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning 7b model and 8k examples: Emerging reasoning with reinforcement learning is both effective and efficient

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:27:51.845013Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T14:27:51.095867Z digest=sha256:011335d854c2bb4f17acaa8c2ba984b60f20abdbac11fd60fcc7925f1508caa6

Observation 1593a911-7d7b-4aa5-ba85-16fd0f66797b · outbound

This paper cites rStar-Math: Small LLMs Can Master Math Reasoning with Self-Evolved Deep Thinking.

Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning rStar-Math: Small LLMs Can Master Math Reasoning with Self-Evolved Deep Thinking

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-08T14:27:51.099046Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:27:51.099046Z digest=sha256:6ca9016424177763148738d66d441fceae24b20f36e99911567f62688199f931

Observation 089e71d6-00d3-40e5-99ae-c03ea1fd1b18 · outbound

This paper cites American invitational mathematics examination - aime.

Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning American invitational mathematics examination - aime

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:27:51.835134Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T14:27:51.102557Z digest=sha256:71e30f2fedcf4d9f33decb0ce830076f4f204948546b6ad74c078c270a5cea32

Observation 6bf0ca30-56b9-4c38-8e8c-5330ed4184c9 · outbound

This paper cites Are Your LLMs Capable of Stable Reasoning?.

Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning Are Your LLMs Capable of Stable Reasoning?

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-08T14:27:51.106013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:27:51.106013Z digest=sha256:f31c1e94f6954d228754bfe2663f20cbabc8049bb73156b76c789440604c8059

Observation 31d1c947-f4ef-4c66-85fe-d0e7361341e5 · outbound

This paper cites OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems.

Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-08T14:27:51.109531Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:27:51.109531Z digest=sha256:d56dd2fa738d9ecd2bd7ff64e952b9c8f9352df1b7de64ac305b4439e2fb54b7

Observation a2fbfb90-8f35-4cb3-841a-923edbf8e8c1 · outbound

This paper cites Opencompass: A universal evaluation platform for foundation models.

Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning Opencompass: A universal evaluation platform for foundation models

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-08T14:27:51.112911Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:27:51.112911Z digest=sha256:431535b16ae6a6b585bb1e989d486d248b483a566028ba1aa9d6279ea57b2545

Observation 92d69a59-5df4-483e-a0c9-aa54a3d0e38a · outbound

This paper cites Policy gradient meth- ods for reinforcement learning with function approximation.

Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning Policy gradient meth- ods for reinforcement learning with function approximation

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-08T14:27:51.116126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:27:51.116126Z digest=sha256:c699cc1f1122583208232415e511cf2b67d0882d5ba10ab020b5139755318ba1

Observation 8c06dc8d-24e4-4e83-830e-513fd2a8a2e0 · outbound

This paper cites Tree of thoughts: Deliberate problem solving with large language models.

Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning Tree of thoughts: Deliberate problem solving with large language models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-08T14:27:51.119909Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:27:51.119909Z digest=sha256:1f9ef78594bd52e0740b7e7fcd3b2f60c871becd9769ad0f13fa8350cb58954e

Observation bb206363-c2c9-43d6-8c8d-729da4200dd5 · outbound

This paper cites Graph of thoughts: Solving elaborate problems with large language models.

Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning Graph of thoughts: Solving elaborate problems with large language models

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:27:51.808050Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T14:27:51.123115Z digest=sha256:e5cf2962c2b517f7f25d49486520ecbb2a2409c435fdcf82a3c4929949b57e9a

Observation 9bfc999a-6fb0-4875-bee1-12abb2e7b0e4 · outbound

This paper cites MetaMath: Bootstrap Your Own Mathematical Questions for Large Language Models.

Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning MetaMath: Bootstrap Your Own Mathematical Questions for Large Language Models

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-08T14:27:51.126152Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:27:51.126152Z digest=sha256:67aebf9977a190f8b251320d528105bf9407978516b7b484f0adc5ca89c645ac

Observation cf905594-0d06-4889-9173-68a756270331 · outbound

This paper cites Augmenting Math Word Problems via Iterative Question Composing.

Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning Augmenting Math Word Problems via Iterative Question Composing

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-08T14:27:51.129269Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:27:51.129269Z digest=sha256:c614b2e0a2443b20674816533c09410a3eb5702423d1db1bc67a935e94709b70

Observation 9c17fa23-b188-44c1-88b5-7a3002de14d0 · outbound

This paper cites MuggleMath: Assessing the Impact of Query and Response Augmentation on Math Reasoning.

Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning MuggleMath: Assessing the Impact of Query and Response Augmentation on Math Reasoning

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-08T14:27:51.132425Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:27:51.132425Z digest=sha256:beec82b471051a1b8b0ab958138d7f8e2f2d89d9b74d2c8813eecb62e46b66aa

Observation edd0dc10-fc9a-4e71-9e4e-f9acf6ecc1da · outbound

This paper cites Goat: Fine-tuned LLaMA Outperforms GPT-4 on Arithmetic Tasks.

Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning Goat: Fine-tuned LLaMA Outperforms GPT-4 on Arithmetic Tasks

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-08T14:27:51.135706Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:27:51.135706Z digest=sha256:5e483fb875c80a183ee8ca1bf39a60303e30808f217fae821364524360b2316a

Observation a02c956f-67b4-4333-bf1e-e649ece2d74c · outbound

This paper cites MAmmoTH: Building Math Generalist Models through Hybrid Instruction Tuning.

Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning MAmmoTH: Building Math Generalist Models through Hybrid Instruction Tuning

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-08T14:27:51.138537Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:27:51.138537Z digest=sha256:6976f0fd78b2d1e1ae20f5b760546dd6c8f15179fcc3dd4287ae05f492b698d7

Observation ccde4491-3225-4820-94f6-b3dbd1c7d9ab · outbound

This paper cites SFT Memorizes, RL Generalizes: A Comparative Study of Foundation Model Post-training.

Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning SFT Memorizes, RL Generalizes: A Comparative Study of Foundation Model Post-training

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-08T14:27:51.141898Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:27:51.141898Z digest=sha256:2892ba9d76b3e336d3145b3719f85df8a45471e692c269e8d6acb40696335ff5

Observation 021f7021-240a-4ea6-b091-9b24fbe92bfc · outbound

This paper cites Can large language models reason and plan? Annals of the New York Academy of Sciences, 1534(1):15–18, 2024.

Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning Can large language models reason and plan? Annals of the New York Academy of Sciences, 1534(1):15–18, 2024

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:27:51.798417Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T14:27:51.145468Z digest=sha256:defce3a0a915659d6cba2d505d68ccd1747ded7263d6aea1ea97c0797ca8de8b

Observation ab3810a8-8362-4285-8482-3cc08241ea9c · outbound

This paper cites WizardMath: Empowering Mathematical Reasoning for Large Language Models via Reinforced Evol-Instruct.

Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning WizardMath: Empowering Mathematical Reasoning for Large Language Models via Reinforced Evol-Instruct

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-08T14:27:51.148684Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:27:51.148684Z digest=sha256:fba19b91339a56cf249942011364dde723e5779f952671ca81cc0e1fcec7409f

Observation 2c04f306-b46e-4265-ba4b-088a53fde1bc · outbound

This paper cites Beyond Human Data: Scaling Self-Training for Problem-Solving with Language Models.

Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning Beyond Human Data: Scaling Self-Training for Problem-Solving with Language Models

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-08T14:27:51.152051Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:27:51.152051Z digest=sha256:34c24bf5e1e1f7672f9a5c740603f0765bdc2a236a878c9e5f127c8c728901ee

Observation 36414ed8-e16a-4d3a-9544-5ce4ca0b8136 · outbound

This paper cites Back to Basics: Revisiting REINFORCE Style Optimization for Learning from Human Feedback in LLMs.

Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning Back to Basics: Revisiting REINFORCE Style Optimization for Learning from Human Feedback in LLMs

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-08T14:27:51.155707Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T14:27:51.155707Z digest=sha256:baff091c78bfdd15b858cdcb7574fbab2663fbec746808deb0ac51cc6c69d534

Observation db660ffc-25a4-402b-87d3-05a8d08395be · outbound

This paper cites You don’t need to re-generate the answer to the question because the standard answer has been given.

Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning You don’t need to re-generate the answer to the question because the standard answer has been given

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:27:51.787477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T14:27:51.159064Z digest=sha256:3a7d6506aad5b56cfc753d41433c2d5ab61459ebfe60a20fd45779281eb64cf3

Observation 2717bfa1-4154-49af-a700-1ba82d031236 · outbound

This paper cites an unresolved cited work.

Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning Unresolved cited work

Reference 64

Resolution
unresolved
raw_fallback, observed 2026-08-08T14:27:51.777395Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T14:27:51.162379Z digest=sha256:339f241e4e34e6de497cf835ebf9d6c17bc47c09d6d3ce62aa9c6b233b50d807

Observation a0ff2ceb-134e-4708-947a-ab645d6c9128 · outbound

This paper cites As long as the answer is the same as the standard answer, it is enough.

Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning As long as the answer is the same as the standard answer, it is enough

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:27:51.767856Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T14:27:51.165691Z digest=sha256:fba83b23a3593709056bdb2de3d326ad3357d9af10aee848ca5e4b424f0782a2

Observation 995844ec-b558-4489-9c11-f784e2b4644a · outbound

This paper cites And some formulas are expressed in different ways, but they are equivalent and correct.

Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning And some formulas are expressed in different ways, but they are equivalent and correct

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:27:51.758301Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T14:27:51.168809Z digest=sha256:8da01a033339d5ffb65e4120c6bd81748cb5491849d26c48fd5a6eb947f80144

Observation d9207753-5c50-4664-a353-bcd4a1e1babf · outbound

This paper cites A" or "B.

Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning A" or "B

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T14:27:51.748464Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T14:27:51.172027Z digest=sha256:0c564760b3c896c712f7d9547bf842c7c32f06c085a4f3b0e86f7c6c8fda5808

Observation 39f35344-4083-4e5e-8070-4a94cc29d417 · outbound

This paper cites an unresolved cited work.

Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning Unresolved cited work

Reference 68

Resolution
unresolved
raw_fallback, observed 2026-08-08T14:27:51.737774Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T14:27:51.175657Z digest=sha256:08578ae7f4c48dd1ab9dc39a2c54ef463054d44054cd971df09f5ea9b9f1893d

Observation 644abfb2-54e4-4cb4-8385-84f39da5b153 · outbound

This paper cites an unresolved cited work.

Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning Unresolved cited work

Reference 69

Resolution
unresolved
raw_fallback, observed 2026-08-08T14:27:51.728145Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T14:27:51.178803Z digest=sha256:12a82747d589949935da88c3829e0e5ac3c2dae82eae90c881ce87fec09edd20

Observation 88c346e6-9494-44cb-ba24-e65330c9b28c · outbound

This paper cites an unresolved cited work.

Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning Unresolved cited work

Reference 70

Resolution
unresolved
raw_fallback, observed 2026-08-08T14:27:51.717965Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T14:27:51.181847Z digest=sha256:97a7e63f4ab41a74f9e1e2014599edd871646df7f38b3ad229327b3a09919604

Observation eb0d0544-e8d6-466c-85c9-256d984231ff · outbound

This paper cites an unresolved cited work.

Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning Unresolved cited work

Reference 71

Resolution
unresolved
raw_fallback, observed 2026-08-08T14:27:51.708176Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T14:27:51.185024Z digest=sha256:a245ea5eac39271040967c6c15b0c6fbbe1e58dd74f12ba0b91170e54926adbb

Observation 0720b3e7-602e-4467-ac25-8c1517554c68 · outbound

This paper cites an unresolved cited work.

Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning Unresolved cited work

Reference 72

Resolution
unresolved
raw_fallback, observed 2026-08-08T14:27:51.698233Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-08T14:27:51.188077Z digest=sha256:67c74170f57b780392110d15a2ecb1d5d237ec6c6995782860134468b038cb96

Pith citing papers

Observation 40a1032b-0f5d-4e52-8a77-687f5ff1de13 · inbound

Reinforcement Learning from Human Feedback cites this paper.

Reinforcement Learning from Human Feedback Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning

Reference 77

Resolution
verified exact
arxiv_id, observed 2026-05-22T19:32:01.345899Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-22T19:27:40.991325Z digest=sha256:59fb154010ab45b2c4c8a125bd3315b332ed0168463e44697133b01da6f5ec10

Observation 6a317298-8807-4244-bf64-974ce588f354 · inbound

VARD: Efficient and Dense Fine-Tuning for Diffusion Models with Value-based RL cites this paper.

VARD: Efficient and Dense Fine-Tuning for Diffusion Models with Value-based RL Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T15:18:54.669121Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:18:54.669121Z digest=sha256:fecf6d98789993f3a02fbb985331dd384c4fc165622e59149b9462debdf742fd

Observation 04b7eaec-cacb-4cbb-85b3-b00ad11efe56 · inbound

RFTF: Reinforcement Fine-tuning for Embodied Agents with Temporal Feedback cites this paper.

RFTF: Reinforcement Fine-tuning for Embodied Agents with Temporal Feedback Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T14:10:12.939491Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:10:12.939491Z digest=sha256:acacbd09f3912bfd2a8283a6fe6819e767463c6955c1bafca5932ca87a39784b

Observation 747e0839-1ec3-49c4-b4f8-1e95d895c7b6 · inbound

Training LLMs for EHR-Based Reasoning Tasks via Reinforcement Learning cites this paper.

Training LLMs for EHR-Based Reasoning Tasks via Reinforcement Learning Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T12:40:09.933663Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:40:09.933663Z digest=sha256:8b2a1bb7ea922074743d726f495c0160af9a93d33080020a87779f10537b5a40

Observation 867e3e6f-9b30-47c1-b77b-cd076a11bcef · inbound

Response-Level Rewards Are All You Need for Online Reinforcement Learning in LLMs: A Mathematical Perspective cites this paper.

Response-Level Rewards Are All You Need for Online Reinforcement Learning in LLMs: A Mathematical Perspective Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:31.527587Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:31.527587Z digest=sha256:81c86cfc6c7aebc5a25682bd89f9248d25c33d1d69eac0e0ab017bf575c47919

Observation 59a01924-d375-4698-afa9-7e01cd44e84c · inbound

Consistent Paths Lead to Truth: Self-Rewarding Reinforcement Learning for LLM Reasoning cites this paper.

Consistent Paths Lead to Truth: Self-Rewarding Reinforcement Learning for LLM Reasoning Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T05:09:44.200132Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:09:44.200132Z digest=sha256:aaa597645c6dc3b7a603ea4f32c78dc88fd8c628187e18d43c70e99baf913c7e

Observation 69dcfd01-d48e-4d21-85dd-be3372de2929 · inbound

Unleashing Embodied Task Planning Ability in LLMs via Reinforcement Learning cites this paper.

Unleashing Embodied Task Planning Ability in LLMs via Reinforcement Learning Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T21:55:01.685340Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:55:01.685340Z digest=sha256:4ca14508b859db84702f775ec0b60a4141d8dc85b9753ab73baacd1e64745e71

Observation 1845ae5b-eaf4-43c4-a421-8b0b14f25fd2 · inbound

Test-Time Scaling with Reflective Generative Model cites this paper.

Test-Time Scaling with Reflective Generative Model Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T20:46:44.783885Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:46:44.783885Z digest=sha256:556e52dea1ecc58bf5cca6a7000f1e8d71550bb0c75ef9348ab3b8e84d80ad8b

Observation 9d18271c-db52-48d0-a8b3-0f038b000880 · inbound

A Technical Survey of Reinforcement Learning Techniques for Large Language Models cites this paper.

A Technical Survey of Reinforcement Learning Techniques for Large Language Models Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T19:59:28.348022Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:59:28.348022Z digest=sha256:508388cf935996828b43ba366fb3d2abd7323ee65cccac407b62fde458d1b564

Observation 40207b04-b788-474d-8ee3-8edcf2b3f3de · inbound

Alignment and Safety in Large Language Models: Safety Mechanisms, Training Paradigms, and Emerging Challenges cites this paper.

Alignment and Safety in Large Language Models: Safety Mechanisms, Training Paradigms, and Emerging Challenges Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning

Reference 198

Resolution
unresolved
no resolver link, observed 2026-08-06T14:13:06.979201Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:13:06.979201Z digest=sha256:93059c6c48573993979b67e51e76ece2fc18e3b88139cc626546d5734f5440f8

Observation 9eaa3cc7-49dd-4760-9e89-0b0c97512861 · inbound

Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models cites this paper.

Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-05T20:21:08.144794Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:21:08.144794Z digest=sha256:a9a87b5ea38dee5e6c095495b2d5aa542a858a99c2efc7d14a2cecfed6efcfe8

Observation 36284a2e-89b8-4e03-ac8a-75c39a212196 · inbound

Uncertainty Under the Curve: A Sequence-Level Entropy Area Metric for Reasoning LLM cites this paper.

Uncertainty Under the Curve: A Sequence-Level Entropy Area Metric for Reasoning LLM Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-05T15:11:01.093701Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:11:01.093701Z digest=sha256:64a3243172127767181e249de54f417de06848c1a1df60c89adb2bc0c5e0388a

Observation da937472-cdc1-490a-beb9-4b79cd4fbc57 · inbound

Intern-S1-MO: Long-horizon Reasoning Agent for Olympiad?Level Mathematical Problem Solving cites this paper.

Intern-S1-MO: Long-horizon Reasoning Agent for Olympiad?Level Mathematical Problem Solving Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-04T06:40:24.891449Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:40:24.891449Z digest=sha256:0f486510fa01e3a139311503ba66b19c4abf9340f1cd22ea0c80bf1f33dde2bc

Observation cc15766a-f859-4eaa-9b3b-82060a08db9c · inbound

Reward Modeling for Reinforcement Learning-Based LLM Reasoning: Design, Challenges, and Evaluation cites this paper.

Reward Modeling for Reinforcement Learning-Based LLM Reasoning: Design, Challenges, and Evaluation Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-03T03:04:43.981554Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:04:43.981554Z digest=sha256:c46299d9af24a6cf89ef58a89e3d042a702a63a4229b47bca073759d5c9cacad

Observation 6e7d4cec-4529-43b9-bac8-67b44a0416a9 · inbound

Trust Region On-Policy Distillation cites this paper.

Trust Region On-Policy Distillation Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning

Reference 224

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T20:56:13.534551Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-28T17:38:50.313305Z digest=sha256:2af61be3826a368f2cac30a1fe26fea04e89f648c6f82ae83695e63d9b64a4cc

Observation 7bc710db-20a8-4c95-a0e0-c252eb3cb39f · inbound

The Periodic Table of LLM Reasoning: A Structured Survey of Reasoning Paradigms, Methods, and Failure Modes cites this paper.

The Periodic Table of LLM Reasoning: A Structured Survey of Reasoning Paradigms, Methods, and Failure Modes Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning

Reference 165

Resolution
verified exact
arxiv_id, observed 2026-06-27T13:00:56.109782Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-27T12:59:51.091008Z digest=sha256:ba17dcb83aae1ec0d2274ae1b7122357f90d0348e01720a3bedcba0477189d9a