Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T15:06:35.501572Z
Paper Citation Record · LEDGER
As of 19 August 2026, this Paper Citation Record lists 44 of 44 outbound references and 29 inbound Pith citation observations for arXiv:2505.16400.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T15:06:35.501572Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T11:04:14.050518Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z
44 of 44 outbound references displayed
External citation measurements
0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z
Observation bc0f1d88-e512-4685-a69c-d708aada8cb3 · outbound
AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning OpenCodeReasoning: Advancing Data Distillation for Competitive Coding
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cf4fbbea-9329-4e47-9ac9-32d3b8e00df8 · outbound
AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning Cosmos-Reason1: From Physical Common Sense To Embodied Reasoning
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 63df6452-c98d-437d-9275-daf942bf2c46 · outbound
AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning Matharena: Evaluatingllmsonuncontaminatedmathcompetitions,February2025
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 8d71d6b6-2bdf-41c3-9e4b-c08397732127 · outbound
AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning Llama-Nemotron: Efficient Reasoning Models.arXiv preprint arXiv:2505.00949, 2025
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 020157d3-fd39-4ff5-b0a8-5098219e32dc · outbound
AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning Bridging supervised learning and reinforcement learning in math reasoning.arXiv preprint, 2025
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 5190ac7e-046f-47b0-a9a9-66545e9cb3d2 · outbound
AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning Evaluating Large Language Models Trained on Code
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ffe16904-0259-44f0-ad08-81046c0cb338 · outbound
AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning Training Verifiers to Solve Math Word Problems
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 887e1314-3873-42ed-9526-052b06796094 · outbound
AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 652fbdd3-cbb5-4d55-a749-836820c87e0d · outbound
AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning The Llama 3 Herd of Models
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 81e55ba2-3080-4c5c-a19c-19e60495eb29 · outbound
AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning DeepSeek-Coder: When the Large Language Model Meets Programming -- The Rise of Code Intelligence
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 34d09bcb-f26b-439d-bf7f-56042ab63c9f · outbound
AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 245b7b6e-9de9-4fa0-81e2-ab6f86cd1c27 · outbound
AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning Skywork open reasoner series, 2025
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 0a914012-d692-4cc1-bae7-f1ee13803624 · outbound
AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning Measuring mathematical problem solving with the math dataset.NeurIPS, 2021
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 5c83f9db-dfb0-41f0-bbef-712e9d7a46e3 · outbound
AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning Open r1: A fully open reproduction of deepseek-r1, January 2025
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 90ee0e6b-52ff-40bd-ad18-0ff352523cd9 · outbound
AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0553564d-8612-4b45-8336-a6a2a33ca38f · outbound
AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning Adam: A Method for Stochastic Optimization
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 52356d23-5f99-4217-b9d9-d08978ed2ce8 · outbound
AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning Gonzalez, Hao Zhang, and Ion Stoica
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 77e21b3b-f8f9-48b9-8a42-71f26ff1a087 · outbound
AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning Numinamath
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation d6ef4e8f-f597-4ef6-b2de-afb965259a57 · outbound
AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning Let’s verify step by step
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 77f105be-63c4-4387-b8b9-392d2da950e6 · outbound
AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning DeepSeek-V3 Technical Report
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 62ea21e6-726c-4955-80ee-87ba1fac670b · outbound
AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning Skywork-Reward: Bag of Tricks for Reward Modeling in LLMs
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bca7eb50-4e78-44c8-b247-6e032ef6b77c · outbound
AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning Is your code generated by chatGPT really correct? rigorous evaluation of large language models for code generation
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 18a06f92-5396-49aa-895d-690b15053bc0 · outbound
AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning Evaluating language models for efficient code generation
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 943f7d5d-3417-4c90-a906-f22ca245594a · outbound
AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning AceMath: Advancing Frontier Math Reasoning with Post-Training and Reward Modeling
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f62916a8-3707-488a-871f-4ba17fc61ac8 · outbound
AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning Deepcoder: A fully open-source 14b coder at o3-mini level, 2025
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 64bde935-b262-40a0-9b46-8333690e42c4 · outbound
AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning Tang, Manan Roongta, Colin Cai, Jeffrey Luo, Li Erran Li, Raluca Ada Popa, and Ion Stoica
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 801014ec-d30e-4b72-a28d-a34a81e0fd50 · outbound
AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning Imitate, Explore, and Self-Improve: A Reproduction Report on Slow-thinking Reasoning Systems
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b3932edf-1f19-4485-917b-1965519b2f19 · outbound
AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning AIMO-2 Winning Solution: Building State-of-the-Art Mathematical Reasoning Models with OpenMathReasoning dataset
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f594eac6-4b05-4fe9-83b7-3d3ecbd2a1b2 · outbound
AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning Learning to reason with LLMs, 2024
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 5cfdd2fb-dfd5-430c-a1a6-72f17cf6f8a4 · outbound
AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning Qwen3, April 2025
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation c908a22d-9724-4234-acba-4d09f27f2be7 · outbound
AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning Qwq-32b: Embracing the power of reinforcement learning, March 2025
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 6dc5eecb-f298-4121-abf4-4b3f9ed98fcd · outbound
AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning Areal: Ant reasoning rl.https://github.com/inclusionAI/AReaL, 2025
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation ffa884c3-6476-46db-97a6-685fd284f4db · outbound
AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning Proximal Policy Optimization Algorithms
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 16f4defc-df20-453b-8f91-72d0605aa93a · outbound
AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 570359b4-1457-4bfd-a6db-e7e87a6ef781 · outbound
AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning HybridFlow: A Flexible and Efficient RLHF Framework
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e9c27ef8-6b48-43f1-904d-7af7db4955f0 · outbound
AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning Chain-of-thought prompting elicits reasoning in large language models.Advances in neural information processing systems, 35:24824–24837, 2022
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 51031af8-96fb-4a71-ae2a-093c23cca5ee · outbound
AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning Light-R1: Curriculum SFT, DPO and RL for Long COT from Scratch and Beyond
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0986e787-4d8d-4456-8c8c-e0b2d8b4c211 · outbound
AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning Simple statistical gradient-following algorithms for connectionist reinforcement learning.Machine learning, 8:229–256, 1992
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 48f6a565-6fa7-4c1f-b5f9-30faffd8f681 · outbound
AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning Qwen2.5 Technical Report
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 88905fb7-ae1a-483f-a07f-345628fd43d2 · outbound
AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 510d0643-00d6-4fef-aef8-1ab9ad7221ec · outbound
AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning DAPO: An Open-Source LLM Reinforcement Learning System at Scale
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 611726fe-6b2d-409d-8c1d-77e4328403d1 · outbound
AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ca8ca9ec-51a5-4129-a6ec-795d41618fb8 · outbound
AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning ACECODER: Acing Coder RL via Automated Test-Case Synthesis
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2b6e39db-4943-4156-866b-1525df2d61e7 · outbound
AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning r’s **, let’s break it down step by step. ### Step 1: Understand the Context - It seems you’re referring to the letter **
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation fe7e1c32-886c-47e8-8499-7d151890b587 · inbound
Flow-GRPO: Training Flow Matching Models via Online RL AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning
Reference 72
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 106c8006-b943-42f6-89be-79b3fc2895ba · inbound
MiMo-VL Technical Report AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c2edf7f1-ed2c-4a8f-80a6-07974803d571 · inbound
AceReason-Nemotron 1.1: Advancing Math and Code Reasoning through SFT and RL Synergy AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2a0e87ad-35f9-427c-8aae-3ec1f1dfbbb9 · inbound
Ring-lite: Scalable Reasoning via C3PO-Stabilized Reinforcement Learning for LLMs AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 162c4416-8a67-4ce8-9d07-a7313e31f590 · inbound
Squeeze the Soaked Sponge: Efficient Off-policy Reinforcement Finetuning for Large Language Model AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6712258d-8282-40f6-ad1f-31d0bc269b16 · inbound
CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7ea75ee1-7bc6-480f-8cec-9b1aa1fdea48 · inbound
Stabilizing Knowledge, Promoting Reasoning: Dual-Token Constraints for RLVR AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 9a6aba26-2b8e-416f-827a-99da7e9707bb · inbound
JT-Math: A Multi-Stage Framework for Advanced Mathematical Reasoning in Large Language Models AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 77c19d91-9dd0-49e4-bf54-bcd3b4dcf963 · inbound
PaVeRL-SQL: Text-to-SQL via Partial-Match Rewards and Verbal Reinforcement Learning AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2500b2d5-6c02-4916-99b1-335cb24575aa · inbound
A Survey of Reinforcement Learning for Large Reasoning Models AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning
Reference 74
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation cdc9989a-6d34-43aa-8d34-8c94d1d57ac4 · inbound
CE-GPPO: Coordinating Entropy via Gradient-Preserving Clipping Policy Optimization in Reinforcement Learning AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 3a280a8b-e968-43f2-8ec5-90bbbe690e94 · inbound
MENTOR: Reinforcement Learning via Flexible Teacher-Optimized Rewards for Tool-Use Distillation AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8b08caa5-d6d7-401f-be09-9adfd1e9b9a5 · inbound
OckBench: Measuring the Efficiency of LLM Reasoning AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 051254d6-6815-46c0-8514-79096d9ba67c · inbound
Entropy Ratio Clipping as a Soft Global Constraint for Stable Reinforcement Learning AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 604cc654-1611-494a-a9f5-b122ba7c947b · inbound
Efficient Reasoning on the Edge AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 30ceaf37-843f-45f8-8949-c6f54325ab76 · inbound
SUPERNOVA: Eliciting General Reasoning in LLMs with Reinforcement Learning on Natural Instructions AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 1eac8f59-c2c5-4a84-aa95-f9e582fcd24c · inbound
Characterizing Model-Native Skills AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 963255fd-4918-4ab8-a1be-4b5e861860c7 · inbound
Train Separately, Merge Together: Modular Post-Training with Mixture-of-Experts AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 993bcdb0-e451-4784-979e-d0e3b9e60613 · inbound
Understanding and Mitigating Spurious Signal Amplification in Test-Time Reinforcement Learning for Math Reasoning AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation edc3109a-db52-4c8f-b539-c362704567e8 · inbound
Delay, Plateau, or Collapse: Evaluating the Impact of Systematic Verification Error on RLVR AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 8cc6f019-6d8f-482f-924f-841fba0a9cdf · inbound
Towards Order Fairness: Mitigating LLMs Order Sensitivity through Dual Group Advantage Optimization AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 2c103f40-2966-4ba2-9c42-93d845d2059d · inbound
Scalable Token-Level Hallucination Detection in Large Language Models AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 4e0abeda-e4b8-4968-8653-cb0b3fd38774 · inbound
Respecting Self-Uncertainty in On-Policy Self-Distillation for Efficient LLM Reasoning AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 4d9f336e-c8ca-4339-b799-afc4c8a24a24 · inbound
When Self-Belief Misleads: Active Label Acquisition for Reinforcement Learning with Verifiable Rewards AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 49c626dd-add2-429b-a250-05b5da73c1f4 · inbound
RLVR Datasets and Where to Find Them: Tracing Data Lineage for Better Training Data AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation cae54978-a304-4bb7-9065-517d124851da · inbound
From Reasoning Traces to Reusable Modules: Understanding Compositional Generalization in Language Model Reasoning AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning
Reference 74
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation f080444e-7265-48f3-befc-cf754f9647fc · inbound
Beyond Entropy: Learning from Token-Level Distributional Deviations for LLM Reasoning AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 4afb7a0a-b48a-4c0a-ab97-a7d8e16d28ed · inbound
What are Key Factors for Updates in RL for LLM Reasoning? AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 70ba9640-73f3-4a19-bd19-009a3d40a001 · inbound
Ring-Zero: Scaling Zero RL to a Trillion Parameters for Emergent Reasoning AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.