Pith. sign in

Paper Citation Record · LEDGER

The Climb Carves Wisdom Deeper Than the Summit: On the Noisy Rewards in Learning to Reason

As of 10 August 2026, this Paper Citation Record lists 38 of 38 outbound references and 7 inbound Pith citation observations for arXiv:2505.22653.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.22653 v1

Coverage vector

measured 38 of 38 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:08:26.418982Z

measured 45 of 45 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 7 of 7 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T21:21:41.900632Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T06:06:40.819064Z

Reference resolution

38 of 38 outbound references displayed

  • verified exact0
  • verified fuzzy17
  • unresolved21
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 336360d1-d394-4743-b0ae-7be6a657ddb8 · outbound

This paper cites Rethinking reflection in pre-training, 2025.

The Climb Carves Wisdom Deeper Than the Summit: On the Noisy Rewards in Learning to Reason Rethinking reflection in pre-training, 2025

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T13:08:20.774171Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:08:20.774171Z digest=sha256:a86cdae426f615d5efce6e825133ccf4bb3d8fd3f39721e2d507bba63ba969fa

Observation cfee5b61-bfa2-41b1-8ec9-ebeaef0ecf13 · outbound

This paper cites Math- arena: Evaluating llms on uncontaminated math competitions, February 2025.

The Climb Carves Wisdom Deeper Than the Summit: On the Noisy Rewards in Learning to Reason Math- arena: Evaluating llms on uncontaminated math competitions, February 2025

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:08:30.141674Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:08:20.882511Z digest=sha256:dde2fd557be7cb4049e9255d9c79662c8f7d74f1a115f6bc58600d5b053b2208

Observation 998bfda0-0c08-46be-aa8c-b9ec8d061431 · outbound

This paper cites Do not think that much for 2+3=? on the overthinking of o1-like llms, 2025.

The Climb Carves Wisdom Deeper Than the Summit: On the Noisy Rewards in Learning to Reason Do not think that much for 2+3=? on the overthinking of o1-like llms, 2025

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T13:08:20.971903Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:08:20.971903Z digest=sha256:32638f393fdb9f2a35f5fe2f3ba7d16f38a539d63cd2ad4493c5b8d644108745

Observation dd4d2aa7-3ebd-4178-84e4-761d8bc8e0d7 · outbound

This paper cites The accuracy paradox in RLHF: When better reward models don‘t yield better language models.

The Climb Carves Wisdom Deeper Than the Summit: On the Noisy Rewards in Learning to Reason The accuracy paradox in RLHF: When better reward models don‘t yield better language models

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:08:29.793026Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:08:21.046797Z digest=sha256:18f82d23140a13463dd6105a6969a222cfc00ed1807467b788bc05b881d4fcbb

Observation 5cc1211e-9c9b-4cbe-a5b8-4b980717d800 · outbound

This paper cites Fortify the shortest stave in attention: Enhancing context awareness of large language models for effective tool use.

The Climb Carves Wisdom Deeper Than the Summit: On the Noisy Rewards in Learning to Reason Fortify the shortest stave in attention: Enhancing context awareness of large language models for effective tool use

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:08:29.579980Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:08:21.149444Z digest=sha256:db3520ff8cbbcf9a9f51273e5a9d75f61767e21291933a54e6b3bd714e170c17

Observation fa6a6ec4-aa68-44e1-ae75-ea2c1ad58338 · outbound

This paper cites Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning, 2025.

The Climb Carves Wisdom Deeper Than the Summit: On the Noisy Rewards in Learning to Reason Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning, 2025

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T13:08:21.245093Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:08:21.245093Z digest=sha256:73a273c57176769968dc1a91df68f6b62903164ca9b18c64dbb53267ccf51321

Observation 2fbaa476-27b7-4e7f-a4ab-9f973e655974 · outbound

This paper cites Q*: Improving multi-step reasoning for LLMs with deliberative planning, 2024.

The Climb Carves Wisdom Deeper Than the Summit: On the Noisy Rewards in Learning to Reason Q*: Improving multi-step reasoning for LLMs with deliberative planning, 2024

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:08:29.438633Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:08:21.723778Z digest=sha256:b0cf68096882243c0db4f3ed2b0d35106189984bb150565cbae38e5be9119912

Observation 3881cd0d-3e5b-43ba-989f-bbf78fe41ef5 · outbound

This paper cites Fleiss et al.

The Climb Carves Wisdom Deeper Than the Summit: On the Noisy Rewards in Learning to Reason Fleiss et al

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:08:29.324041Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:08:22.439678Z digest=sha256:1e31a626b752e52ccb5d39502b8e5bc16de5177ba85ae804543e0c0bd63453bf

Observation 2bca5e9a-1439-47f8-aaf6-7f87adf5a74b · outbound

This paper cites Angelopoulos, Jiantao Jiao, Banghua Zhu, Joseph E.

The Climb Carves Wisdom Deeper Than the Summit: On the Noisy Rewards in Learning to Reason Angelopoulos, Jiantao Jiao, Banghua Zhu, Joseph E

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:08:29.144749Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:08:22.712573Z digest=sha256:1f218a75711b1b24f67749f23d657ae11be5e0cb2ea347b8145f2ac88e92344d

Observation 64d7804c-fb4f-41bd-872c-ca23e7ddd1b6 · outbound

This paper cites an unresolved cited work.

The Climb Carves Wisdom Deeper Than the Summit: On the Noisy Rewards in Learning to Reason Unresolved cited work

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T13:08:22.798068Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:08:22.798068Z digest=sha256:c688a628040ee31774aaec3e4897bd441b9d03f95b4b21dfe82db43fc3ea29d0

Observation bd96b21b-2804-4f67-adb6-7cb18f974884 · outbound

This paper cites Training large language models to reason in a continuous latent space, 2024.

The Climb Carves Wisdom Deeper Than the Summit: On the Noisy Rewards in Learning to Reason Training large language models to reason in a continuous latent space, 2024

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:08:28.939096Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:08:22.889241Z digest=sha256:e4a2ceb28dd48ddd23dc73eae7ca2b09c6ba52bc57a870b39e832727c76fc02a

Observation 635020b9-ad11-4fdd-af1a-d3fb2e407619 · outbound

This paper cites Measuring mathematical problem solving with the MATH dataset.

The Climb Carves Wisdom Deeper Than the Summit: On the Noisy Rewards in Learning to Reason Measuring mathematical problem solving with the MATH dataset

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T13:08:22.944756Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:08:22.944756Z digest=sha256:c46159bd27dabbf2aec05d1b5b6c08e769ee6c73a2b691e546924ca2fb7b21dc

Observation ee38baa9-bbe4-4dfd-88bb-1f454b229ffb · outbound

This paper cites Open-reasoner-zero: An open source approach to scaling up reinforcement learning on the base model, 2025.

The Climb Carves Wisdom Deeper Than the Summit: On the Noisy Rewards in Learning to Reason Open-reasoner-zero: An open source approach to scaling up reinforcement learning on the base model, 2025

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T13:08:23.015073Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:08:23.015073Z digest=sha256:16f464f24f494a018b51a5af5cfd1f0ca49455fe7b3c38b38e435a64dcf004a7

Observation 378ba01a-8ced-4322-ad0c-92e08e401e32 · outbound

This paper cites Human-centric dialog training via offline reinforcement learning.

The Climb Carves Wisdom Deeper Than the Summit: On the Noisy Rewards in Learning to Reason Human-centric dialog training via offline reinforcement learning

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:08:28.733015Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:08:23.127945Z digest=sha256:0cfe42fc48c7d18382ab8d9e98f903beede4fa7b85fb17bea02a9b06efe69cdb

Observation cca45cb9-7636-447b-b605-72d834896154 · outbound

This paper cites Smith, and Hannaneh Hajishirzi.

The Climb Carves Wisdom Deeper Than the Summit: On the Noisy Rewards in Learning to Reason Smith, and Hannaneh Hajishirzi

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:08:28.526514Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:08:23.187072Z digest=sha256:8d67f6b3ae946c9e52b3d0a90a9214e1a94c50e3c68b8508aeda0b842bbe9dd9

Observation a2855bad-a891-4e71-8de2-9b2da73099ff · outbound

This paper cites Skywork-reward: Bag of tricks for reward modeling in llms, 2024.

The Climb Carves Wisdom Deeper Than the Summit: On the Noisy Rewards in Learning to Reason Skywork-reward: Bag of tricks for reward modeling in llms, 2024

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:08:28.347936Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:08:23.332664Z digest=sha256:fab07a836da4e0568da1e1e3f4d6680779c0dcb6cae41bbed9d8c2f78ac3e6db

Observation 9c252d4a-6eba-44fc-8507-e4a48c7916b7 · outbound

This paper cites Liu, Kevin Lin, John Hewitt, Ashwin Paranjape, Michele Bevilacqua, Fabio Petroni, and Percy Liang.

The Climb Carves Wisdom Deeper Than the Summit: On the Noisy Rewards in Learning to Reason Liu, Kevin Lin, John Hewitt, Ashwin Paranjape, Michele Bevilacqua, Fabio Petroni, and Percy Liang

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T13:08:23.476989Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:08:23.476989Z digest=sha256:1ddebf45bd6d685f4b35129a95f573decdcf00b4da98481e569da8ff34c101a8

Observation 14cce726-7ded-4d3d-8c40-62737ca8a218 · outbound

This paper cites RM-bench: Benchmarking reward models of language models with subtlety and style.

The Climb Carves Wisdom Deeper Than the Summit: On the Noisy Rewards in Learning to Reason RM-bench: Benchmarking reward models of language models with subtlety and style

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:08:28.032968Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:08:23.564975Z digest=sha256:28eb4f7fc457c3981fc65977197ac038f9720820e13bbe0ea14dc0399c0ccdf4

Observation 8f5711fa-ead4-4460-b00e-d306e59b5d64 · outbound

This paper cites The llama 3 herd of models, 2024.

The Climb Carves Wisdom Deeper Than the Summit: On the Noisy Rewards in Learning to Reason The llama 3 herd of models, 2024

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T13:08:23.664904Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:08:23.664904Z digest=sha256:20e0463a3c77b9362f4d7eee17b72c3bb3ff19ed5c1ce788d9215446b4e7e809

Observation 08f18d02-a882-4f7a-be83-c2d6097de2ab · outbound

This paper cites s1: Simple test-time scaling, 2025.

The Climb Carves Wisdom Deeper Than the Summit: On the Noisy Rewards in Learning to Reason s1: Simple test-time scaling, 2025

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T13:08:23.784375Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:08:23.784375Z digest=sha256:3fd67d20e18400566e7e82d7f96fadcff1ba1415b8af821fc487b0e23fef0f8b

Observation 69dee594-ee21-4766-be72-8080c5d2b93d · outbound

This paper cites Webgpt: Browser-assisted question-answering with human feedback, 2022.

The Climb Carves Wisdom Deeper Than the Summit: On the Noisy Rewards in Learning to Reason Webgpt: Browser-assisted question-answering with human feedback, 2022

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T13:08:24.224696Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:08:24.224696Z digest=sha256:fc2270fe16fc4d41036dfddf47f415457b93cf1fe08020514c211102aa037142

Observation 10d11159-eb51-492b-9e73-18ce08a42abe · outbound

This paper cites an unresolved cited work.

The Climb Carves Wisdom Deeper Than the Summit: On the Noisy Rewards in Learning to Reason Unresolved cited work

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T13:08:24.363161Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:08:24.363161Z digest=sha256:c24a08d2235cb921d41804ca6148e681f82c1e4faf077dbb7b6158ea016dbafd

Observation 59acf32a-c8fb-4bb1-ae9e-927664cdeaf9 · outbound

This paper cites Tinyzero.

The Climb Carves Wisdom Deeper Than the Summit: On the Noisy Rewards in Learning to Reason Tinyzero

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T13:08:24.515080Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:08:24.515080Z digest=sha256:f866cfec2ecb55d10831cc3af86c5d25188b642bd2e3341a931e5df9bff171d5

Observation 7d818213-3f27-4f53-bf0e-9dcb03f1d974 · outbound

This paper cites Lee, and Sanjeev Arora.

The Climb Carves Wisdom Deeper Than the Summit: On the Noisy Rewards in Learning to Reason Lee, and Sanjeev Arora

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:08:27.645811Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:08:24.675041Z digest=sha256:c0eae9ba405b556b00fef9b61a1a700fcad43692321dbb5e5976bef4bc1bb2cd

Observation f037cab5-2d76-46dc-a77e-88e8f44d0696 · outbound

This paper cites an unresolved cited work.

The Climb Carves Wisdom Deeper Than the Summit: On the Noisy Rewards in Learning to Reason Unresolved cited work

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T13:08:24.812651Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:08:24.812651Z digest=sha256:c507efe7cbc42cc30052147a983865243d2cfcde3feaeac79d7baf9411e8ac1d

Observation 9181ac7a-da72-464e-939f-190848bf5bc9 · outbound

This paper cites High- dimensional continuous control using generalized advantage estimation, 2018.

The Climb Carves Wisdom Deeper Than the Summit: On the Noisy Rewards in Learning to Reason High- dimensional continuous control using generalized advantage estimation, 2018

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T13:08:24.984971Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:08:24.984971Z digest=sha256:55de92d88eaf6977d48397222e3daeb7a71dd8684253471eb49369d6494eab8b

Observation 11c5364e-0c72-46a3-aeb0-994d299d28bb · outbound

This paper cites Proximal policy optimization algorithms, 2017.

The Climb Carves Wisdom Deeper Than the Summit: On the Noisy Rewards in Learning to Reason Proximal policy optimization algorithms, 2017

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T13:08:25.215091Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:08:25.215091Z digest=sha256:a14b94c83bff34ebb1424f12efbb36e662dcbd49505f3b52f17de65a8d13dbd0

Observation f7c7bc53-22d7-49cc-9177-88a313d57cbe · outbound

This paper cites HybridFlow: A Flexible and Efficient RLHF Framework.

The Climb Carves Wisdom Deeper Than the Summit: On the Noisy Rewards in Learning to Reason HybridFlow: A Flexible and Efficient RLHF Framework

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T13:08:25.364234Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:08:25.364234Z digest=sha256:3e3e59505c79aec30be0f7eac9472cf3c80891a41740ede6f47764782b4a0401

Observation 4272dc37-bbc5-4ce1-aef3-aed663720f48 · outbound

This paper cites Kimi k1.5: Scaling reinforcement learning with llms, 2025.

The Climb Carves Wisdom Deeper Than the Summit: On the Noisy Rewards in Learning to Reason Kimi k1.5: Scaling reinforcement learning with llms, 2025

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T13:08:25.504871Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:08:25.504871Z digest=sha256:b85830e0011be2e9cd4c04d50f2c248b953aadad104e6e55c16d032255671dc2

Observation 26abeed2-a18d-420c-9afd-e7d1a79adca9 · outbound

This paper cites Dedicated feedback and edit models empower inference- time scaling for open-ended general-domain tasks, 2025.

The Climb Carves Wisdom Deeper Than the Summit: On the Noisy Rewards in Learning to Reason Dedicated feedback and edit models empower inference- time scaling for open-ended general-domain tasks, 2025

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:08:27.328579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:08:25.654842Z digest=sha256:3193f040303a07e4026ec55b436812b7d36cef8e41147f7d51606d7fad7da5f7

Observation 7b6fdd6e-43c2-4eb5-ae7c-4740d996a3fd · outbound

This paper cites Rethinking reward model evaluation: Are we barking up the wrong tree? InThe Thirteenth International Conference on Learning Representations, 2025.

The Climb Carves Wisdom Deeper Than the Summit: On the Noisy Rewards in Learning to Reason Rethinking reward model evaluation: Are we barking up the wrong tree? InThe Thirteenth International Conference on Learning Representations, 2025

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:08:27.187371Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:08:25.754909Z digest=sha256:e227427da3d7a20b94b9cb3d63503f7ca2bc25b195e65118a209f6ff1609b831

Observation bea22113-c115-4ecd-812f-1e937335b71a · outbound

This paper cites Qwen2.5 Technical Report.

The Climb Carves Wisdom Deeper Than the Summit: On the Noisy Rewards in Learning to Reason Qwen2.5 Technical Report

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T13:08:25.864911Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:08:25.864911Z digest=sha256:b3e61daeebbea3455476d7179a61d9c8869c5d790d09cf3029114fdc9a4b226e

Observation ea6018df-d49a-4ef0-bbf8-b858be14404c · outbound

This paper cites Demystifying long chain-of-thought reasoning in llms, 2025.

The Climb Carves Wisdom Deeper Than the Summit: On the Noisy Rewards in Learning to Reason Demystifying long chain-of-thought reasoning in llms, 2025

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T13:08:25.962237Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:08:25.962237Z digest=sha256:0fb33dba60cc710a4d8d4cde7c95fdb5f15537d55327f492346645e659cd79b8

Observation a4b3b8f4-9604-4a7b-b586-b5d512a31f90 · outbound

This paper cites Does reinforcement learning really incentivize reasoning capacity in llms beyond the base model?, 2025.

The Climb Carves Wisdom Deeper Than the Summit: On the Noisy Rewards in Learning to Reason Does reinforcement learning really incentivize reasoning capacity in llms beyond the base model?, 2025

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T13:08:26.027884Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:08:26.027884Z digest=sha256:bd7a67fbec013f6855f5aceedd3612563213bf5a6d5707d1ff81efb59c793952

Observation 9eb3e823-7347-465f-93b5-749aa9c69481 · outbound

This paper cites ReST- MCTS*: LLM self-training via process reward guided tree search.

The Climb Carves Wisdom Deeper Than the Summit: On the Noisy Rewards in Learning to Reason ReST- MCTS*: LLM self-training via process reward guided tree search

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T13:08:26.063576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:08:26.063576Z digest=sha256:958cb49245f88bfa15abf4f17a46d8a221f2394736bf7918191b0e6bc68ed3f5

Observation 7f4098c4-46f3-4d53-a169-68d533074504 · outbound

This paper cites Found in the middle: How language models use long contexts better via plug-and-play positional encoding, 2024.

The Climb Carves Wisdom Deeper Than the Summit: On the Noisy Rewards in Learning to Reason Found in the middle: How language models use long contexts better via plug-and-play positional encoding, 2024

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:08:27.018577Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:08:26.156253Z digest=sha256:13e981e68d46bfbc502ca7f0e27ee8a8a96cd9b7893ba83c0d0df1679b013c23

Observation 1396fc1a-f01e-4061-8aa8-07a35ce0d174 · outbound

This paper cites Rmb: Comprehensively benchmarking reward models in llm alignment, 2025.

The Climb Carves Wisdom Deeper Than the Summit: On the Noisy Rewards in Learning to Reason Rmb: Comprehensively benchmarking reward models in llm alignment, 2025

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:08:26.853533Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:08:26.285176Z digest=sha256:c55f1414a26f1ea81d2ba69492bd7f93d128a78acff0c718323d706f0d5758c7

Observation a4ece631-7ac0-429d-88eb-33ba21a03c33 · outbound

This paper cites Assistant:␣<think>.

The Climb Carves Wisdom Deeper Than the Summit: On the Noisy Rewards in Learning to Reason Assistant:␣<think>

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:08:26.655809Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T13:08:26.418982Z digest=sha256:c741ec9e9ab3095617cd076da271084145d02c376528ae7b149c5ac9e78cb61f

Pith citing papers

Observation 2c78ad0a-163e-4860-87d2-2b82f402346b · inbound

ASTRO: Teaching Language Models to Reason by Reflecting and Backtracking In-Context cites this paper.

ASTRO: Teaching Language Models to Reason by Reflecting and Backtracking In-Context The Climb Carves Wisdom Deeper Than the Summit: On the Noisy Rewards in Learning to Reason

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T21:21:41.900632Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T21:21:41.900632Z digest=sha256:7417ffa33986953ef5bdbd7668d512e96710ce1d0e40a10a301e5746c618b59e

Observation c1e0f7d7-0436-40f6-9d38-34d50fec9880 · inbound

StepHint: Multi-level Stepwise Hints Enhance Reinforcement Learning to Reason cites this paper.

StepHint: Multi-level Stepwise Hints Enhance Reinforcement Learning to Reason The Climb Carves Wisdom Deeper Than the Summit: On the Noisy Rewards in Learning to Reason

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T20:26:46.857127Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:26:46.857127Z digest=sha256:af94f21fbaeaaee25fd767581c996eb5504ac891fd25be6a81346992f46ca710

Observation 44a3cdc7-cb6d-4f98-8329-29478b16b20e · inbound

Mirage or Method? How Model-Task Alignment Induces Divergent RL Conclusions cites this paper.

Mirage or Method? How Model-Task Alignment Induces Divergent RL Conclusions The Climb Carves Wisdom Deeper Than the Summit: On the Noisy Rewards in Learning to Reason

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-05T14:35:50.973056Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:35:50.973056Z digest=sha256:aab8a381db15a449be06d508135a8ac3416b1c72a6f2858b561982ed1b6a3a51

Observation 108954cf-c305-4ec2-8327-6849d809a05d · inbound

Delay, Plateau, or Collapse: Evaluating the Impact of Systematic Verification Error on RLVR cites this paper.

Delay, Plateau, or Collapse: Evaluating the Impact of Systematic Verification Error on RLVR The Climb Carves Wisdom Deeper Than the Summit: On the Noisy Rewards in Learning to Reason

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-10T23:45:53.070918Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-10T18:52:52.969408Z digest=sha256:dd89a568998747f6f9526480f169e7f83a3cd146164a7619857c2e8dd65c00d3

Observation 8ebcca75-35d3-4ee1-87a0-82b75a9e79f3 · inbound

Experience Sharing in Mutual Reinforcement Learning for Heterogeneous Language Models cites this paper.

Experience Sharing in Mutual Reinforcement Learning for Heterogeneous Language Models The Climb Carves Wisdom Deeper Than the Summit: On the Noisy Rewards in Learning to Reason

Reference 57

Resolution
verified exact
arxiv_id, observed 2026-05-11T04:00:55.164977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-05-11T02:02:41.411795Z digest=sha256:be5c4bdeacf0e73101a846b72184c62c690b2435f65d3d1424280e36c5caa0d8

Observation b8673b48-961e-4c8e-a55e-11795d86cadd · inbound

Trust Region On-Policy Distillation cites this paper.

Trust Region On-Policy Distillation The Climb Carves Wisdom Deeper Than the Summit: On the Noisy Rewards in Learning to Reason

Reference 99

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T20:56:13.695124Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-06-28T17:38:50.313305Z digest=sha256:ef4b3858e5910b3d6bf518d056ceceaf7754c2d88f17f1218f747a3ee116701a

Observation 7a0e6b5b-8f8d-4cde-9b40-2fdff5e0c9b4 · inbound

GeoMin: Data-Efficient Semi-Supervised RLVR via Geometric Distribution Modeling cites this paper.

GeoMin: Data-Efficient Semi-Supervised RLVR via Geometric Distribution Modeling The Climb Carves Wisdom Deeper Than the Summit: On the Noisy Rewards in Learning to Reason

Reference 51

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T06:06:40.820599Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-06-28T07:45:43.320339Z digest=sha256:abc50cd5c11aa24f23e4586bca9e328aa44340f12452834a2b88ab4afc9aed50