Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-08T14:27:51.188077Z
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 72 of 72 outbound references and 16 inbound Pith citation observations for arXiv:2502.06781.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-08T14:27:51.188077Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T15:18:54.669121Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z
72 of 72 outbound references displayed
External citation measurements
0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z
Observation 070b2498-fd8a-4ebf-8468-bee490322b3e · outbound
Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning Towards Large Reasoning Models: A Survey of Reinforced Reasoning with Large Language Models
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 00f685af-9b48-416e-b41c-e48ce97c90f6 · outbound
Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning Evaluation of openai o1: Opportunities and challenges of agi
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e75f2754-1033-4cd8-b473-c3f72e91c319 · outbound
Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning Mathematical Language Models: A Survey
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1d04e19b-9ea1-46b8-8400-d2cdbabb70e5 · outbound
Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning Learning mathematics with large language models: A comparative study with computer algebra systems and other tools
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation dd449979-713d-4581-a4a4-02df42ae17fb · outbound
Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning InternLM-Math: Open Math Large Language Models Toward Verifiable Reasoning
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a579b9ce-63bc-4d3a-9636-af3be8dea892 · outbound
Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 629ab2a0-7625-4222-95cd-2447f843433c · outbound
Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning Chain-of-thought prompting elicits reasoning in large language models
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1ff782f0-5b4d-4fae-b51b-d1721b242208 · outbound
Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning Self-Consistency Improves Chain of Thought Reasoning in Language Models
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d630a42d-af51-475d-86af-69e9aa1d0d52 · outbound
Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning Large language models are zero-shot reasoners
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 8e5601dd-0f1a-4b45-ae29-df4c0c3fbd2d · outbound
Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning Learning to reason with llms
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 199ead34-9050-4d22-a51a-9b03c6f5914e · outbound
Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning O1 Replication Journey -- Part 2: Surpassing O1-preview through Simple Distillation, Big Progress or Bitter Lesson?
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 871245c2-a6a3-493c-a54d-f99196ba7204 · outbound
Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning Imitate, Explore, and Self-Improve: A Reproduction Report on Slow-thinking Reasoning Systems
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2ce37bdf-1627-42da-af54-cbfded59f731 · outbound
Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning Unresolved cited work
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 292779e5-b639-4409-a57b-e2699deb09de · outbound
Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning DeepSeek-V3 Technical Report
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e2c1c0a8-5866-4a2f-9cb5-ad6154e5b334 · outbound
Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning Let's Verify Step by Step
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4e36297a-2947-4184-a310-4a514ddaf916 · outbound
Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning Kimi k1.5: Scaling Reinforcement Learning with LLMs
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8d9ae74e-9211-4336-bafe-8df79ea40d07 · outbound
Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning Math-shepherd: Verify and reinforce llms step-by-step without human annotations
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation f8d070c7-893d-42df-86e7-3ed7693068de · outbound
Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning VinePPO: Refining Credit Assignment in RL Training of LLMs
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation be090f1b-f199-43b0-baa9-6fbb70e31602 · outbound
Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning Proximal Policy Optimization Algorithms
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2006b791-2a83-49d0-a968-4226a286814b · outbound
Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning Process Reinforcement through Implicit Rewards
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 64d9f09a-5f85-4fa4-88eb-10ec2f19afda · outbound
Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning Measuring Mathematical Problem Solving With the MATH Dataset
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 42494796-eabd-435e-9f73-f74ef2ad5a6c · outbound
Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning Qwq: Reflect deeply on the boundaries of the unknown, November 2024
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 536d2d7d-acf4-4146-b2a1-09b8dfec53c0 · outbound
Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning Scaling Relationship on Learning Mathematical Reasoning with Large Language Models
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e4969442-85fa-4f68-a47d-b11989dcd6b3 · outbound
Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning Scaling laws for reward model overoptimization
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9b12b849-774d-42cf-9696-0417cfc1e1a3 · outbound
Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning Compositional preference models for aligning LMs
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 936b6233-ed02-4f62-887d-7961a625bba4 · outbound
Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning Tulu 3: Pushing Frontiers in Open Language Model Post-Training
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b0a1e810-bdc1-43c1-9dbe-8c19308aea81 · outbound
Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning Training language models to follow instructions with human feedback
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 60363d63-ef52-40a2-9268-446c007979b5 · outbound
Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning Measuring goodhart’s law
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation a7892b10-f28a-4cf0-92e0-b61c143dd331 · outbound
Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning Training Language Models with Language Feedback at Scale
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0c3fe5c1-4a8a-434b-865d-c3999a37b8c8 · outbound
Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning Reward Model Ensembles Help Mitigate Overoptimization
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9d9880a1-80ff-40c3-bc4f-2f9990694c83 · outbound
Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning BoNBoN Alignment for Large Language Models and the Sweetness of Best-of-n Sampling
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1d2dd4b9-7af2-4202-8fcf-f2333aee9d00 · outbound
Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning BOND: Aligning LLMs with Best-of-N Distillation
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 761dece5-630e-4bd8-bafa-bbba8fe28237 · outbound
Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning An invariant form for the prior probability in estimation problems.Proceedings of the Royal Society of London
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 829da88e-2bc1-4430-90a8-52240aabd898 · outbound
Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning Inference-aware fine-tuning for best-of-n sampling in large language models
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1c87d03d-98d3-4564-8e75-04f06a18433d · outbound
Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning Unresolved cited work
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 490b5225-c7e4-49b7-bef9-f6842dca9f64 · outbound
Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning From $r$ to $Q^*$: Your Language Model is Secretly a Q-Function
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ef4dc255-3e0b-4c4f-8de8-5f10a5aa7577 · outbound
Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning Inverse-Q*: Token Level Reinforcement Learning for Aligning Large Language Models Without Preference Data
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation b5ba40e0-c1a8-47d0-8eec-f3445f4f1487 · outbound
Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning Training Verifiers to Solve Math Word Problems
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 21c39724-a36f-4aff-aff7-5cf6d8cc758b · outbound
Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning Qwen2.5 Technical Report
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ac75e7e4-c444-4646-8d66-1dac7bdbfad3 · outbound
Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning Opendatalab: Empowering general artificial intelligence with open datasets, 2024
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation d591aa9a-08ca-4501-89e3-220a28935053 · outbound
Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning Numinamath: The largest public dataset in ai4maths with 860k pairs of competition math problems and solutions.Hugging Face repository, 13:9, 2024
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5a5c329a-4708-48b6-97c2-01a298ccd552 · outbound
Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning Hello GPT-4o, 2024
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 42c59246-87cd-4843-af78-27b2e239b241 · outbound
Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning Claude 3.5 sonnet, 2024
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 338d44e4-3e74-495e-8e48-25503ca63927 · outbound
Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning 7b model and 8k examples: Emerging reasoning with reinforcement learning is both effective and efficient
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 1593a911-7d7b-4aa5-ba85-16fd0f66797b · outbound
Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning rStar-Math: Small LLMs Can Master Math Reasoning with Self-Evolved Deep Thinking
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 089e71d6-00d3-40e5-99ae-c03ea1fd1b18 · outbound
Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning American invitational mathematics examination - aime
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 6bf0ca30-56b9-4c38-8e8c-5330ed4184c9 · outbound
Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning Are Your LLMs Capable of Stable Reasoning?
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 31d1c947-f4ef-4c66-85fe-d0e7361341e5 · outbound
Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning OlympiadBench: A Challenging Benchmark for Promoting AGI with Olympiad-Level Bilingual Multimodal Scientific Problems
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a2fbfb90-8f35-4cb3-841a-923edbf8e8c1 · outbound
Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning Opencompass: A universal evaluation platform for foundation models
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 92d69a59-5df4-483e-a0c9-aa54a3d0e38a · outbound
Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning Policy gradient meth- ods for reinforcement learning with function approximation
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8c06dc8d-24e4-4e83-830e-513fd2a8a2e0 · outbound
Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning Tree of thoughts: Deliberate problem solving with large language models
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bb206363-c2c9-43d6-8c8d-729da4200dd5 · outbound
Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning Graph of thoughts: Solving elaborate problems with large language models
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 9bfc999a-6fb0-4875-bee1-12abb2e7b0e4 · outbound
Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning MetaMath: Bootstrap Your Own Mathematical Questions for Large Language Models
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cf905594-0d06-4889-9173-68a756270331 · outbound
Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning Augmenting Math Word Problems via Iterative Question Composing
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9c17fa23-b188-44c1-88b5-7a3002de14d0 · outbound
Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning MuggleMath: Assessing the Impact of Query and Response Augmentation on Math Reasoning
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation edd0dc10-fc9a-4e71-9e4e-f9acf6ecc1da · outbound
Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning Goat: Fine-tuned LLaMA Outperforms GPT-4 on Arithmetic Tasks
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a02c956f-67b4-4333-bf1e-e649ece2d74c · outbound
Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning MAmmoTH: Building Math Generalist Models through Hybrid Instruction Tuning
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ccde4491-3225-4820-94f6-b3dbd1c7d9ab · outbound
Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning SFT Memorizes, RL Generalizes: A Comparative Study of Foundation Model Post-training
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 021f7021-240a-4ea6-b091-9b24fbe92bfc · outbound
Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning Can large language models reason and plan? Annals of the New York Academy of Sciences, 1534(1):15–18, 2024
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation ab3810a8-8362-4285-8482-3cc08241ea9c · outbound
Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning WizardMath: Empowering Mathematical Reasoning for Large Language Models via Reinforced Evol-Instruct
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2c04f306-b46e-4265-ba4b-088a53fde1bc · outbound
Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning Beyond Human Data: Scaling Self-Training for Problem-Solving with Language Models
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 36414ed8-e16a-4d3a-9544-5ce4ca0b8136 · outbound
Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning Back to Basics: Revisiting REINFORCE Style Optimization for Learning from Human Feedback in LLMs
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation db660ffc-25a4-402b-87d3-05a8d08395be · outbound
Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning You don’t need to re-generate the answer to the question because the standard answer has been given
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 2717bfa1-4154-49af-a700-1ba82d031236 · outbound
Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning Unresolved cited work
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation a0ff2ceb-134e-4708-947a-ab645d6c9128 · outbound
Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning As long as the answer is the same as the standard answer, it is enough
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 995844ec-b558-4489-9c11-f784e2b4644a · outbound
Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning And some formulas are expressed in different ways, but they are equivalent and correct
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation d9207753-5c50-4664-a353-bcd4a1e1babf · outbound
Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning A" or "B
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 39f35344-4083-4e5e-8070-4a94cc29d417 · outbound
Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning Unresolved cited work
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 644abfb2-54e4-4cb4-8385-84f39da5b153 · outbound
Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning Unresolved cited work
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 88c346e6-9494-44cb-ba24-e65330c9b28c · outbound
Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning Unresolved cited work
Reference 70
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation eb0d0544-e8d6-466c-85c9-256d984231ff · outbound
Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning Unresolved cited work
Reference 71
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 0720b3e7-602e-4467-ac25-8c1517554c68 · outbound
Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning Unresolved cited work
Reference 72
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 40a1032b-0f5d-4e52-8a77-687f5ff1de13 · inbound
Reinforcement Learning from Human Feedback Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning
Reference 77
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 6a317298-8807-4244-bf64-974ce588f354 · inbound
VARD: Efficient and Dense Fine-Tuning for Diffusion Models with Value-based RL Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 04b7eaec-cacb-4cbb-85b3-b00ad11efe56 · inbound
RFTF: Reinforcement Fine-tuning for Embodied Agents with Temporal Feedback Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 747e0839-1ec3-49c4-b4f8-1e95d895c7b6 · inbound
Training LLMs for EHR-Based Reasoning Tasks via Reinforcement Learning Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 867e3e6f-9b30-47c1-b77b-cd076a11bcef · inbound
Response-Level Rewards Are All You Need for Online Reinforcement Learning in LLMs: A Mathematical Perspective Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 59a01924-d375-4698-afa9-7e01cd44e84c · inbound
Consistent Paths Lead to Truth: Self-Rewarding Reinforcement Learning for LLM Reasoning Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 69dcfd01-d48e-4d21-85dd-be3372de2929 · inbound
Unleashing Embodied Task Planning Ability in LLMs via Reinforcement Learning Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1845ae5b-eaf4-43c4-a421-8b0b14f25fd2 · inbound
Test-Time Scaling with Reflective Generative Model Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9d18271c-db52-48d0-a8b3-0f038b000880 · inbound
A Technical Survey of Reinforcement Learning Techniques for Large Language Models Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 40207b04-b788-474d-8ee3-8edcf2b3f3de · inbound
Alignment and Safety in Large Language Models: Safety Mechanisms, Training Paradigms, and Emerging Challenges Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning
Reference 198
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9eaa3cc7-49dd-4760-9e89-0b0c97512861 · inbound
Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 36284a2e-89b8-4e03-ac8a-75c39a212196 · inbound
Uncertainty Under the Curve: A Sequence-Level Entropy Area Metric for Reasoning LLM Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation da937472-cdc1-490a-beb9-4b79cd4fbc57 · inbound
Intern-S1-MO: Long-horizon Reasoning Agent for Olympiad?Level Mathematical Problem Solving Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cc15766a-f859-4eaa-9b3b-82060a08db9c · inbound
Reward Modeling for Reinforcement Learning-Based LLM Reasoning: Design, Challenges, and Evaluation Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6e7d4cec-4529-43b9-bac8-67b44a0416a9 · inbound
Trust Region On-Policy Distillation Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning
Reference 224
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 7bc710db-20a8-4c95-a0e0-c252eb3cb39f · inbound
The Periodic Table of LLM Reasoning: A Structured Survey of Reasoning Paradigms, Methods, and Failure Modes Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning
Reference 165
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.