Pith. sign in

Paper Citation Record · LEDGER

AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning

As of 19 August 2026, this Paper Citation Record lists 44 of 44 outbound references and 29 inbound Pith citation observations for arXiv:2505.16400.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.16400 v3

Coverage vector

measured 44 of 44 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:06:35.501572Z

measured 73 of 73 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 29 of 29 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:04:14.050518Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

44 of 44 outbound references displayed

  • verified exact0
  • verified fuzzy18
  • unresolved26
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation bc0f1d88-e512-4685-a69c-d708aada8cb3 · outbound

This paper cites OpenCodeReasoning: Advancing Data Distillation for Competitive Coding.

AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning OpenCodeReasoning: Advancing Data Distillation for Competitive Coding

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T15:06:35.285362Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:06:35.285362Z digest=sha256:02a81795a0a733203125e46b36f9ca2485440635278858e26926c42c2305bdd5

Observation cf4fbbea-9329-4e47-9ac9-32d3b8e00df8 · outbound

This paper cites Cosmos-Reason1: From Physical Common Sense To Embodied Reasoning.

AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning Cosmos-Reason1: From Physical Common Sense To Embodied Reasoning

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T15:06:35.291514Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:06:35.291514Z digest=sha256:028b2b4af86de3a487f6d22bb632b0a53d88487cb82732056e858a7ee1a16499

Observation 63df6452-c98d-437d-9275-daf942bf2c46 · outbound

This paper cites Matharena: Evaluatingllmsonuncontaminatedmathcompetitions,February2025.

AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning Matharena: Evaluatingllmsonuncontaminatedmathcompetitions,February2025

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:06:36.278571Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T15:06:35.296648Z digest=sha256:f6f2ee69031fe48b4563c4ca8fb5f2cce767388369788953a2c3b063362f541e

Observation 8d71d6b6-2bdf-41c3-9e4b-c08397732127 · outbound

This paper cites Llama-Nemotron: Efficient Reasoning Models.arXiv preprint arXiv:2505.00949, 2025.

AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning Llama-Nemotron: Efficient Reasoning Models.arXiv preprint arXiv:2505.00949, 2025

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T15:06:35.302290Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:06:35.302290Z digest=sha256:365fb420b77e04a1aaa296a13083962f0516784128e72690c61c15dfb5859bca

Observation 020157d3-fd39-4ff5-b0a8-5098219e32dc · outbound

This paper cites Bridging supervised learning and reinforcement learning in math reasoning.arXiv preprint, 2025.

AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning Bridging supervised learning and reinforcement learning in math reasoning.arXiv preprint, 2025

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:06:36.262756Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T15:06:35.307091Z digest=sha256:076866d797d2da20894ec27529ae4825402bf0e2cc6aeb0feed3a8cb98fdfabb

Observation 5190ac7e-046f-47b0-a9a9-66545e9cb3d2 · outbound

This paper cites Evaluating Large Language Models Trained on Code.

AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning Evaluating Large Language Models Trained on Code

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T15:06:35.311936Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:06:35.311936Z digest=sha256:1228f0df68b1ecdd5c1c5e23060d1312b61dc9d3feec3ef992c700e7bb5ab5db

Observation ffe16904-0259-44f0-ad08-81046c0cb338 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning Training Verifiers to Solve Math Word Problems

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T15:06:35.317188Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:06:35.317188Z digest=sha256:e0da2144979b204061c66a61a1f501a245099c3c5aa6617aeaba6e2b3fb360a1

Observation 887e1314-3873-42ed-9526-052b06796094 · outbound

This paper cites Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models.

AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning Omni-MATH: A Universal Olympiad Level Mathematic Benchmark For Large Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T15:06:35.322184Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:06:35.322184Z digest=sha256:d1a9bea5e0c356a11a0244ecf051c37bb72f235463e3178862758dc67724392d

Observation 652fbdd3-cbb5-4d55-a749-836820c87e0d · outbound

This paper cites The Llama 3 Herd of Models.

AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning The Llama 3 Herd of Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T15:06:35.327654Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:06:35.327654Z digest=sha256:6f840db6e6a7c2ed489831a0e8bb8e7be9d5cc6d3c33472430c285431789a6db

Observation 81e55ba2-3080-4c5c-a19c-19e60495eb29 · outbound

This paper cites DeepSeek-Coder: When the Large Language Model Meets Programming -- The Rise of Code Intelligence.

AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning DeepSeek-Coder: When the Large Language Model Meets Programming -- The Rise of Code Intelligence

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T15:06:35.333225Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:06:35.333225Z digest=sha256:30a4194c1481453040f0c80b4b7327acfd98d54a51b6b92d57f2f4a78ada9846

Observation 34d09bcb-f26b-439d-bf7f-56042ab63c9f · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T15:06:35.338695Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:06:35.338695Z digest=sha256:84fc4e385224042c9f11c9cb7d2f2c7e183e0079ce78f263d231d8156791feb1

Observation 245b7b6e-9de9-4fa0-81e2-ab6f86cd1c27 · outbound

This paper cites Skywork open reasoner series, 2025.

AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning Skywork open reasoner series, 2025

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:06:36.246484Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T15:06:35.346308Z digest=sha256:f121832f5fe90f0e2eb7a5f922d860f40c377de667977a78b56e0972632dac5a

Observation 0a914012-d692-4cc1-bae7-f1ee13803624 · outbound

This paper cites Measuring mathematical problem solving with the math dataset.NeurIPS, 2021.

AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning Measuring mathematical problem solving with the math dataset.NeurIPS, 2021

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:06:36.231267Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T15:06:35.351544Z digest=sha256:9a9d889b8dafb664046163fa1ec2b198310be71d0b0a983c2e5eb76b4037250d

Observation 5c83f9db-dfb0-41f0-bbef-712e9d7a46e3 · outbound

This paper cites Open r1: A fully open reproduction of deepseek-r1, January 2025.

AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning Open r1: A fully open reproduction of deepseek-r1, January 2025

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:06:36.216361Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T15:06:35.356186Z digest=sha256:addac456d9a7e46fac32da609de3995dcb6c8c28616a7788c8380c5b88f00449

Observation 90ee0e6b-52ff-40bd-ad18-0ff352523cd9 · outbound

This paper cites LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code.

AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T15:06:35.360560Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:06:35.360560Z digest=sha256:7ae1050d89b390cc90c16d8868507b1eef30ed0bd26ae1d39f4b66648aaf8fa3

Observation 0553564d-8612-4b45-8336-a6a2a33ca38f · outbound

This paper cites Adam: A Method for Stochastic Optimization.

AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning Adam: A Method for Stochastic Optimization

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T15:06:35.364966Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:06:35.364966Z digest=sha256:a0a52b9b997a1d2ae0c2f38dd87b4ece8aef88cce7acd661cba47ecf93db01aa

Observation 52356d23-5f99-4217-b9d9-d08978ed2ce8 · outbound

This paper cites Gonzalez, Hao Zhang, and Ion Stoica.

AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning Gonzalez, Hao Zhang, and Ion Stoica

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T15:06:35.369809Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:06:35.369809Z digest=sha256:bfa15286b292da80ecd5d43e545e031e9ec86bba0fd8bce782c80803a4a09bbc

Observation 77e21b3b-f8f9-48b9-8a42-71f26ff1a087 · outbound

This paper cites Numinamath.

AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning Numinamath

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:06:36.188890Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T15:06:35.374725Z digest=sha256:4ca655d07bec9fc001e1d07117dfb51fdd8b6238b33c7c1085bbd466584342fa

Observation d6ef4e8f-f597-4ef6-b2de-afb965259a57 · outbound

This paper cites Let’s verify step by step.

AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning Let’s verify step by step

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:06:36.171672Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T15:06:35.379388Z digest=sha256:d96ba048a56ada26b39444e7da51609dbb27e06936e55e38ecec2b9c03ba2606

Observation 77f105be-63c4-4387-b8b9-392d2da950e6 · outbound

This paper cites DeepSeek-V3 Technical Report.

AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning DeepSeek-V3 Technical Report

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T15:06:35.384699Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:06:35.384699Z digest=sha256:9da1593ce59bb4efe967d37defaa2c9492f16469b7cca71ff710596cb53d1ebc

Observation 62ea21e6-726c-4955-80ee-87ba1fac670b · outbound

This paper cites Skywork-Reward: Bag of Tricks for Reward Modeling in LLMs.

AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning Skywork-Reward: Bag of Tricks for Reward Modeling in LLMs

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T15:06:35.391101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:06:35.391101Z digest=sha256:711c77aba681b1ff968142c13bfeac6dd555c895f79e0954f7ba13605f8524a0

Observation bca7eb50-4e78-44c8-b247-6e032ef6b77c · outbound

This paper cites Is your code generated by chatGPT really correct? rigorous evaluation of large language models for code generation.

AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning Is your code generated by chatGPT really correct? rigorous evaluation of large language models for code generation

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:06:36.156142Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T15:06:35.396473Z digest=sha256:4e490d57bd963e519a09c0bad17ff588b23ef9d17b56c061e3f9bbf3a756b082

Observation 18a06f92-5396-49aa-895d-690b15053bc0 · outbound

This paper cites Evaluating language models for efficient code generation.

AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning Evaluating language models for efficient code generation

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:06:36.141039Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T15:06:35.401772Z digest=sha256:bfba8fa3b889fbcda6da852d9e0e1ec210788f5c574ea8f93b9daa4666510fe8

Observation 943f7d5d-3417-4c90-a906-f22ca245594a · outbound

This paper cites AceMath: Advancing Frontier Math Reasoning with Post-Training and Reward Modeling.

AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning AceMath: Advancing Frontier Math Reasoning with Post-Training and Reward Modeling

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T15:06:35.406551Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:06:35.406551Z digest=sha256:278c3bbd9b05c64a3b1713920f57cb5c5b2f3889f2225d3c6ed004f59a38d26e

Observation f62916a8-3707-488a-871f-4ba17fc61ac8 · outbound

This paper cites Deepcoder: A fully open-source 14b coder at o3-mini level, 2025.

AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning Deepcoder: A fully open-source 14b coder at o3-mini level, 2025

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:06:36.125432Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T15:06:35.412153Z digest=sha256:c155283642e459846dc10a173448723394963f5853cdb2d7543b8a353aaec8c6

Observation 64bde935-b262-40a0-9b46-8333690e42c4 · outbound

This paper cites Tang, Manan Roongta, Colin Cai, Jeffrey Luo, Li Erran Li, Raluca Ada Popa, and Ion Stoica.

AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning Tang, Manan Roongta, Colin Cai, Jeffrey Luo, Li Erran Li, Raluca Ada Popa, and Ion Stoica

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:06:36.109543Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T15:06:35.417224Z digest=sha256:291ae2cf0119812255b339abb8b41a1b9d14b7f591d22e28642cdb372dae2b7c

Observation 801014ec-d30e-4b72-a28d-a34a81e0fd50 · outbound

This paper cites Imitate, Explore, and Self-Improve: A Reproduction Report on Slow-thinking Reasoning Systems.

AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning Imitate, Explore, and Self-Improve: A Reproduction Report on Slow-thinking Reasoning Systems

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T15:06:35.422347Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:06:35.422347Z digest=sha256:58041ffcd03d441f207e4012a0440cd48a27ba97321b18905e6d2ab6d4e24814

Observation b3932edf-1f19-4485-917b-1965519b2f19 · outbound

This paper cites AIMO-2 Winning Solution: Building State-of-the-Art Mathematical Reasoning Models with OpenMathReasoning dataset.

AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning AIMO-2 Winning Solution: Building State-of-the-Art Mathematical Reasoning Models with OpenMathReasoning dataset

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T15:06:35.426645Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:06:35.426645Z digest=sha256:3d8a0b63ab58a29619e6508b11bf51a2261147aa54b8febea5e0e69ddfe9698f

Observation f594eac6-4b05-4fe9-83b7-3d3ecbd2a1b2 · outbound

This paper cites Learning to reason with LLMs, 2024.

AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning Learning to reason with LLMs, 2024

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:06:36.093590Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T15:06:35.432368Z digest=sha256:faf36c0646f591d907f10f021d39f7177fc3742e48362caae787175b0171f251

Observation 5cfdd2fb-dfd5-430c-a1a6-72f17cf6f8a4 · outbound

This paper cites Qwen3, April 2025.

AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning Qwen3, April 2025

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:06:36.077528Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T15:06:35.437045Z digest=sha256:403ccbeaaefc11fd1444f028c7b517b49ebdb36bf0443e2f270f18e7c278ab78

Observation c908a22d-9724-4234-acba-4d09f27f2be7 · outbound

This paper cites Qwq-32b: Embracing the power of reinforcement learning, March 2025.

AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning Qwq-32b: Embracing the power of reinforcement learning, March 2025

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:06:36.059219Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T15:06:35.441505Z digest=sha256:579039ac1eb0a6c4601781f6e74012bd03846532402a65a26f101c572b5f7138

Observation 6dc5eecb-f298-4121-abf4-4b3f9ed98fcd · outbound

This paper cites Areal: Ant reasoning rl.https://github.com/inclusionAI/AReaL, 2025.

AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning Areal: Ant reasoning rl.https://github.com/inclusionAI/AReaL, 2025

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:06:36.043379Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T15:06:35.446011Z digest=sha256:fa996aeeed46e9bb251810731d3859100bee246bfb4b0f2f8d84abe09f8badc5

Observation ffa884c3-6476-46db-97a6-685fd284f4db · outbound

This paper cites Proximal Policy Optimization Algorithms.

AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning Proximal Policy Optimization Algorithms

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T15:06:35.450509Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:06:35.450509Z digest=sha256:56ba291f80967864ec533a95f6b6a4f2425d3907cfc01e215ef5d97498b85823

Observation 16f4defc-df20-453b-8f91-72d0605aa93a · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T15:06:35.455071Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:06:35.455071Z digest=sha256:1eaaf9dbd11eebd19cf1bbf016fe8a412cc1ae8151e21d957b06a93b82ea1921

Observation 570359b4-1457-4bfd-a6db-e7e87a6ef781 · outbound

This paper cites HybridFlow: A Flexible and Efficient RLHF Framework.

AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning HybridFlow: A Flexible and Efficient RLHF Framework

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T15:06:35.459500Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:06:35.459500Z digest=sha256:134a3a9c022a360c5677c9607749e1247edd75d0c66b1a78074394766a7c653b

Observation e9c27ef8-6b48-43f1-904d-7af7db4955f0 · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models.Advances in neural information processing systems, 35:24824–24837, 2022.

AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning Chain-of-thought prompting elicits reasoning in large language models.Advances in neural information processing systems, 35:24824–24837, 2022

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:06:36.026419Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T15:06:35.464091Z digest=sha256:fc1657a8a97a6e9eb9b49516532a14376881d23ff09cb5dc67c0ff7f1834d053

Observation 51031af8-96fb-4a71-ae2a-093c23cca5ee · outbound

This paper cites Light-R1: Curriculum SFT, DPO and RL for Long COT from Scratch and Beyond.

AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning Light-R1: Curriculum SFT, DPO and RL for Long COT from Scratch and Beyond

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T15:06:35.469350Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:06:35.469350Z digest=sha256:e51a14de53c4dac0c22411df84f1e96586747adc3defb35656c41001e9b1099d

Observation 0986e787-4d8d-4456-8c8c-e0b2d8b4c211 · outbound

This paper cites Simple statistical gradient-following algorithms for connectionist reinforcement learning.Machine learning, 8:229–256, 1992.

AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning Simple statistical gradient-following algorithms for connectionist reinforcement learning.Machine learning, 8:229–256, 1992

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:06:36.009629Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T15:06:35.474413Z digest=sha256:a7c7a8fff2f3489e01deb457ef7fbead5d665da9d6db555a69c4deff7a66d865

Observation 48f6a565-6fa7-4c1f-b5f9-30faffd8f681 · outbound

This paper cites Qwen2.5 Technical Report.

AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning Qwen2.5 Technical Report

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T15:06:35.478748Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:06:35.478748Z digest=sha256:8f0c79f5d139d9e37ea72a11a93835cb39b80c8d1222bb5bffbba86e9a6aa0d4

Observation 88905fb7-ae1a-483f-a07f-345628fd43d2 · outbound

This paper cites Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement.

AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning Qwen2.5-Math Technical Report: Toward Mathematical Expert Model via Self-Improvement

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T15:06:35.483074Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:06:35.483074Z digest=sha256:bb993fcf6eaa3e51c7bf5de5008f75c1acff9d974a3dac5cf4724dbe4575c732

Observation 510d0643-00d6-4fef-aef8-1ab9ad7221ec · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T15:06:35.487503Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:06:35.487503Z digest=sha256:e11bf580fe81f5dd4181129b0811a57b8a2c0075ff6717e15d343925eb69cc31

Observation 611726fe-6b2d-409d-8c1d-77e4328403d1 · outbound

This paper cites Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?.

AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T15:06:35.492037Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:06:35.492037Z digest=sha256:17b52a9adfc329a99f34f180d6c76c77412355c683c26cd693a8601bf7e8a3e9

Observation ca8ca9ec-51a5-4129-a6ec-795d41618fb8 · outbound

This paper cites ACECODER: Acing Coder RL via Automated Test-Case Synthesis.

AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning ACECODER: Acing Coder RL via Automated Test-Case Synthesis

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T15:06:35.496885Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:06:35.496885Z digest=sha256:94389613acaa91678bc18bd5924f3cc1bdda0aff4da5033f2060e7e1935c2107

Observation 2b6e39db-4943-4156-866b-1525df2d61e7 · outbound

This paper cites r’s **, let’s break it down step by step. ### Step 1: Understand the Context - It seems you’re referring to the letter **.

AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning r’s **, let’s break it down step by step. ### Step 1: Understand the Context - It seems you’re referring to the letter **

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T15:06:35.992091Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T15:06:35.501572Z digest=sha256:2672988380047b8490aede212b807f203fe0fc590a0890f7cc90fedf973885b4

Pith citing papers

Observation fe7e1c32-886c-47e8-8499-7d151890b587 · inbound

Flow-GRPO: Training Flow Matching Models via Online RL cites this paper.

Flow-GRPO: Training Flow Matching Models via Online RL AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning

Reference 72

Resolution
verified exact
arxiv_id, observed 2026-05-11T18:45:16.862524Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-11T18:45:16.641012Z digest=sha256:84077d25876cd7dec79000fe31c41b83feeeace36ae663ab522f15156d58c255

Observation 106c8006-b943-42f6-89be-79b3fc2895ba · inbound

MiMo-VL Technical Report cites this paper.

MiMo-VL Technical Report AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T11:04:14.050518Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:04:14.050518Z digest=sha256:9ad42459142b4272334f978e19000047c5e1352f42726174e874b8ead4d71325

Observation c2edf7f1-ed2c-4a8f-80a6-07974803d571 · inbound

AceReason-Nemotron 1.1: Advancing Math and Code Reasoning through SFT and RL Synergy cites this paper.

AceReason-Nemotron 1.1: Advancing Math and Code Reasoning through SFT and RL Synergy AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T00:41:33.940080Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:41:33.940080Z digest=sha256:50804d3b26e82a9c136d182f1fcca5b171cccb57ebd149b2eea633aef512eb41

Observation 2a0e87ad-35f9-427c-8aae-3ec1f1dfbbb9 · inbound

Ring-lite: Scalable Reasoning via C3PO-Stabilized Reinforcement Learning for LLMs cites this paper.

Ring-lite: Scalable Reasoning via C3PO-Stabilized Reinforcement Learning for LLMs AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T00:14:22.519933Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:14:22.519933Z digest=sha256:c4d23941515b410170eca48de52e73d69a8774c6e7c534ca816a9028a6ed94c0

Observation 162c4416-8a67-4ce8-9d07-a7313e31f590 · inbound

Squeeze the Soaked Sponge: Efficient Off-policy Reinforcement Finetuning for Large Language Model cites this paper.

Squeeze the Soaked Sponge: Efficient Off-policy Reinforcement Finetuning for Large Language Model AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T18:57:38.595495Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:57:38.595495Z digest=sha256:19a7b74b06117222a98d98f35ee6594c8f908b1ff9563e8c3a64d64d6eb077eb

Observation 6712258d-8282-40f6-ad1f-31d0bc269b16 · inbound

CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks cites this paper.

CodeJudgeBench: Benchmarking LLM-as-a-Judge for Coding Tasks AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T17:32:30.634578Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:32:30.634578Z digest=sha256:470cb5c8c927ee1522a203d7cea959622d129d41753a9f8ef4c49e18bff42126

Observation 7ea75ee1-7bc6-480f-8cec-9b1aa1fdea48 · inbound

Stabilizing Knowledge, Promoting Reasoning: Dual-Token Constraints for RLVR cites this paper.

Stabilizing Knowledge, Promoting Reasoning: Dual-Token Constraints for RLVR AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-21T23:24:26.284917Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-21T23:20:45.685446Z digest=sha256:b489097d19563216464b7125a2406091b07157d1234abfefcc5e2b1393cf25a6

Observation 9a6aba26-2b8e-416f-827a-99da7e9707bb · inbound

JT-Math: A Multi-Stage Framework for Advanced Mathematical Reasoning in Large Language Models cites this paper.

JT-Math: A Multi-Stage Framework for Advanced Mathematical Reasoning in Large Language Models AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T14:11:51.663587Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:11:51.663587Z digest=sha256:11ec3e8bc91586b7e8ec0ed65ab45493b14956b803c8194616c5b26e9a0f188e

Observation 77c19d91-9dd0-49e4-bf54-bcd3b4dcf963 · inbound

PaVeRL-SQL: Text-to-SQL via Partial-Match Rewards and Verbal Reinforcement Learning cites this paper.

PaVeRL-SQL: Text-to-SQL via Partial-Match Rewards and Verbal Reinforcement Learning AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-04T22:48:32.672774Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:48:32.672774Z digest=sha256:36cf319668ac577821c913bc078b62afbaf9f43c41bd490dc9c4edb7e3ff5ee9

Observation 2500b2d5-6c02-4916-99b1-335cb24575aa · inbound

A Survey of Reinforcement Learning for Large Reasoning Models cites this paper.

A Survey of Reinforcement Learning for Large Reasoning Models AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning

Reference 74

Resolution
verified exact
arxiv_id, observed 2026-05-18T00:02:25.126142Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-18T00:02:24.352947Z digest=sha256:852698396bcacedbdedb0fd566e18aba80d775f4c00a37c03caba38b3b734c0b

Observation cdc9989a-6d34-43aa-8d34-8c94d1d57ac4 · inbound

CE-GPPO: Coordinating Entropy via Gradient-Preserving Clipping Policy Optimization in Reinforcement Learning cites this paper.

CE-GPPO: Coordinating Entropy via Gradient-Preserving Clipping Policy Optimization in Reinforcement Learning AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-18T15:11:32.270447Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-18T15:06:33.991929Z digest=sha256:90f8d5776619b4a509cc71bf5ed67f35c04099ec825bdc5e538688555eb6934e

Observation 3a280a8b-e968-43f2-8ec5-90bbbe690e94 · inbound

MENTOR: Reinforcement Learning via Flexible Teacher-Optimized Rewards for Tool-Use Distillation cites this paper.

MENTOR: Reinforcement Learning via Flexible Teacher-Optimized Rewards for Tool-Use Distillation AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-04T08:55:58.550417Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T08:55:58.550417Z digest=sha256:5b7b12798a46c9828c47797c69e7c6f95d7e60037f9969039e36726dd3d85cd9

Observation 8b08caa5-d6d7-401f-be09-9adfd1e9b9a5 · inbound

OckBench: Measuring the Efficiency of LLM Reasoning cites this paper.

OckBench: Measuring the Efficiency of LLM Reasoning AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-03T23:27:36.778635Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T23:27:36.778635Z digest=sha256:d85438ce9c69355a95d4cff78c70aa4094c2f3dc94179ef783cd3c67069cf5fd

Observation 051254d6-6815-46c0-8514-79096d9ba67c · inbound

Entropy Ratio Clipping as a Soft Global Constraint for Stable Reinforcement Learning cites this paper.

Entropy Ratio Clipping as a Soft Global Constraint for Stable Reinforcement Learning AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-17T00:58:46.311837Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-17T00:53:51.251900Z digest=sha256:fe0612bda7410abb7e86dcc58a13b055bc0afe0821524781741cf5ef83bf9d77

Observation 604cc654-1611-494a-a9f5-b122ba7c947b · inbound

Efficient Reasoning on the Edge cites this paper.

Efficient Reasoning on the Edge AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning

Reference 28

Resolution
unresolved
no resolver link, observed 2026-07-13T23:28:12.790404Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T23:28:12.790404Z digest=sha256:170101e352d24c168a41c94fed0f6a8aca8c579f8185b93e56033ee5487b83a9

Observation 30ceaf37-843f-45f8-8949-c6f54325ab76 · inbound

SUPERNOVA: Eliciting General Reasoning in LLMs with Reinforcement Learning on Natural Instructions cites this paper.

SUPERNOVA: Eliciting General Reasoning in LLMs with Reinforcement Learning on Natural Instructions AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-11T07:11:01.629864Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-10T17:16:29.695213Z digest=sha256:538c208bf2e039769a95a9a14058beec5697c17202038618b3f2d1ac3377a7d3

Observation 1eac8f59-c2c5-4a84-aa95-f9e582fcd24c · inbound

Characterizing Model-Native Skills cites this paper.

Characterizing Model-Native Skills AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning

Reference 41

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T06:06:19.468811Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-10T05:42:49.694715Z digest=sha256:700f942beb2c4bcefa84b14a088ce634140b36e077595e2a627de2cc79786ca0

Observation 963255fd-4918-4ab8-a1be-4b5e861860c7 · inbound

Train Separately, Merge Together: Modular Post-Training with Mixture-of-Experts cites this paper.

Train Separately, Merge Together: Modular Post-Training with Mixture-of-Experts AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-10T05:56:11.305553Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-10T05:52:28.822723Z digest=sha256:3175910de2f12332cd6f14b6139ab6372b570a64d275fb17916af1b5240a5fc5

Observation 993bcdb0-e451-4784-979e-d0e3b9e60613 · inbound

Understanding and Mitigating Spurious Signal Amplification in Test-Time Reinforcement Learning for Math Reasoning cites this paper.

Understanding and Mitigating Spurious Signal Amplification in Test-Time Reinforcement Learning for Math Reasoning AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T14:16:07.278352Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-09T22:55:19.471307Z digest=sha256:de3ee8f44bd3b446f702aac5dab298a391eb1e95925e0d0e051d25b9ab5e46c2

Observation edc3109a-db52-4c8f-b539-c362704567e8 · inbound

Delay, Plateau, or Collapse: Evaluating the Impact of Systematic Verification Error on RLVR cites this paper.

Delay, Plateau, or Collapse: Evaluating the Impact of Systematic Verification Error on RLVR AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-10T23:45:53.101099Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-10T18:52:52.969408Z digest=sha256:e5a17af33d7b8171fd93bc2933459818679aab9b081e741dd25aba3abd62db59

Observation 8cc6f019-6d8f-482f-924f-841fba0a9cdf · inbound

Towards Order Fairness: Mitigating LLMs Order Sensitivity through Dual Group Advantage Optimization cites this paper.

Towards Order Fairness: Mitigating LLMs Order Sensitivity through Dual Group Advantage Optimization AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning

Reference 38

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T07:37:29.841604Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-05-13T07:32:58.404947Z digest=sha256:57d87d259ea668ab5a43c855756aa4a48242c9fbbb73f93bed326f84adc8d2c4

Observation 2c103f40-2966-4ba2-9c42-93d845d2059d · inbound

Scalable Token-Level Hallucination Detection in Large Language Models cites this paper.

Scalable Token-Level Hallucination Detection in Large Language Models AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-13T05:52:22.421954Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-13T05:49:23.534294Z digest=sha256:2d0a3b9a1f6de4ea0b08223df3f7f9b4851edccc476cb26192f477d4268dcdba

Observation 4e0abeda-e4b8-4968-8653-cb0b3fd38774 · inbound

Respecting Self-Uncertainty in On-Policy Self-Distillation for Efficient LLM Reasoning cites this paper.

Respecting Self-Uncertainty in On-Policy Self-Distillation for Efficient LLM Reasoning AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-14T19:52:52.411736Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-05-14T19:51:17.595707Z digest=sha256:9ad27154c99864615705bba1b8d60ddffb303de9f497ead77396a0cedb0947c7

Observation 4d9f336e-c8ca-4339-b799-afc4c8a24a24 · inbound

When Self-Belief Misleads: Active Label Acquisition for Reinforcement Learning with Verifiable Rewards cites this paper.

When Self-Belief Misleads: Active Label Acquisition for Reinforcement Learning with Verifiable Rewards AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-06-29T23:04:01.322678Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-29T22:58:46.313028Z digest=sha256:281d418ed00b21e8a1b9fbc4247f42e9fac6af7c1699158aea78d9aa6b115039

Observation 49c626dd-add2-429b-a250-05b5da73c1f4 · inbound

RLVR Datasets and Where to Find Them: Tracing Data Lineage for Better Training Data cites this paper.

RLVR Datasets and Where to Find Them: Tracing Data Lineage for Better Training Data AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-06-29T19:43:55.053657Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-06-29T19:34:12.081362Z digest=sha256:14899d851e7c87d730d21e98558259dc7e400fbbe79683cb643df2816dda8995

Observation cae54978-a304-4bb7-9065-517d124851da · inbound

From Reasoning Traces to Reusable Modules: Understanding Compositional Generalization in Language Model Reasoning cites this paper.

From Reasoning Traces to Reusable Modules: Understanding Compositional Generalization in Language Model Reasoning AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning

Reference 74

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T20:38:55.942362Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-06-27T01:13:11.483599Z digest=sha256:7eccf7f96d6df0714267f3f50a5048eacbaa9b9ba90a1ebdc2aaba069093cc95

Observation f080444e-7265-48f3-befc-cf754f9647fc · inbound

Beyond Entropy: Learning from Token-Level Distributional Deviations for LLM Reasoning cites this paper.

Beyond Entropy: Learning from Token-Level Distributional Deviations for LLM Reasoning AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-07-04T03:49:30.159779Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-26T17:37:43.856000Z digest=sha256:562431b63194a93b07887c520d41d9e88d12f8542049e83d2eed53d5b2132932

Observation 4afb7a0a-b48a-4c0a-ab97-a7d8e16d28ed · inbound

What are Key Factors for Updates in RL for LLM Reasoning? cites this paper.

What are Key Factors for Updates in RL for LLM Reasoning? AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning

Reference 19

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T09:09:43.581253Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-06-26T10:24:53.245739Z digest=sha256:3cc364e265416ed5043c5935ac9f130b6754bcd33f8f66e419247f2434b9728a

Observation 70ba9640-73f3-4a19-bd19-009a3d40a001 · inbound

Ring-Zero: Scaling Zero RL to a Trillion Parameters for Emergent Reasoning cites this paper.

Ring-Zero: Scaling Zero RL to a Trillion Parameters for Emergent Reasoning AceReason-Nemotron: Advancing Math and Code Reasoning through Reinforcement Learning

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-02T06:37:36.445162Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T06:37:36.445162Z digest=sha256:7ef69877ff96e310ec2c9267393a0d9184619c908764b1c09d25a1cd6d6502c1