Pith. sign in

Paper Citation Record · LEDGER

Rethinking RL Evaluation: Can Benchmarks Truly Reveal Failures of RL Methods?

As of 17 August 2026, this Paper Citation Record lists 25 of 25 outbound references and 1 inbound Pith citation observation for arXiv:2510.10541.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2510.10541 v2

Coverage vector

measured 25 of 25 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-04T10:20:20.407043Z

measured 26 of 26 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-25T21:15:07.735336Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-04T19:30:07.857950Z

Reference resolution

25 of 25 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved25
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 4b1a2395-b57b-4d82-90e2-ecbfe2981f0c · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Rethinking RL Evaluation: Can Benchmarks Truly Reveal Failures of RL Methods? Training Verifiers to Solve Math Word Problems

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-04T10:20:18.047448Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:20:18.047448Z digest=sha256:44bf1cedbebf2124af777647c7efb90f529ad3f4b1f4c1ad42cdd82be1bd8904

Observation 3ed62fb8-7d48-44dc-98be-f3aef646042c · outbound

This paper cites Shortcut learning in deep neural networks.

Rethinking RL Evaluation: Can Benchmarks Truly Reveal Failures of RL Methods? Shortcut learning in deep neural networks

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-04T10:20:18.121512Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:20:18.121512Z digest=sha256:762d65e4578d85253ae755c348656f24318e4b940a3d0268f09edaf627211cff

Observation 9a346905-9ddc-414e-bf1d-1a8080810ceb · outbound

This paper cites Measuring massive multitask language understanding.

Rethinking RL Evaluation: Can Benchmarks Truly Reveal Failures of RL Methods? Measuring massive multitask language understanding

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-04T10:20:18.270251Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:20:18.270251Z digest=sha256:fee8fcbefc50d52b1ecaf0c0f6ce0e1ed4211fc7d262e9a819bcd877d01b0cb9

Observation 56a3c550-052d-4348-b875-16894bba678d · outbound

This paper cites Adversarial Examples for Evaluating Reading Comprehension Systems.

Rethinking RL Evaluation: Can Benchmarks Truly Reveal Failures of RL Methods? Adversarial Examples for Evaluating Reading Comprehension Systems

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-04T10:20:18.361180Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:20:18.361180Z digest=sha256:c2395a76d2f52c03238c3e67ece1ce9e60c0c4da16af69c3eb755d0eada54c7f

Observation a6fa7254-93ae-453e-b732-8cbf73c1e452 · outbound

This paper cites Solving Quantitative Reasoning Problems with Language Models.

Rethinking RL Evaluation: Can Benchmarks Truly Reveal Failures of RL Methods? Solving Quantitative Reasoning Problems with Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-04T10:20:18.458587Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:20:18.458587Z digest=sha256:ba81499a83c8b8b792b7c7011bddb2895258d77286e254b333be221760f26925

Observation 433ff859-2a4e-4de6-b343-f9853686de99 · outbound

This paper cites Think or not think: A study of explicit thinking in rule-based visual reinforcement fine-tuning.

Rethinking RL Evaluation: Can Benchmarks Truly Reveal Failures of RL Methods? Think or not think: A study of explicit thinking in rule-based visual reinforcement fine-tuning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-04T10:20:18.524173Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:20:18.524173Z digest=sha256:7d5293ca39a182b7f21be05c7bf22b6972f726d5105fe4e390285ec0d4a2abef

Observation ac220ade-97f4-47ec-ae94-9b163ee48c29 · outbound

This paper cites Let's verify step by step.

Rethinking RL Evaluation: Can Benchmarks Truly Reveal Failures of RL Methods? Let's verify step by step

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-04T10:20:18.590115Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:20:18.590115Z digest=sha256:9077fe8ba552b71fd878f1b6ac7065ffc58cfdaa4a891ef47d2e8b4e4195a5f3

Observation 0ec04a5c-12af-419c-a04b-721bdc6b9759 · outbound

This paper cites Deepscaler: Surpassing o1-preview with a 1.5 b model by scaling rl, 2025.

Rethinking RL Evaluation: Can Benchmarks Truly Reveal Failures of RL Methods? Deepscaler: Surpassing o1-preview with a 1.5 b model by scaling rl, 2025

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-04T10:20:18.665265Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:20:18.665265Z digest=sha256:e6c42ceb5821e07ac9a40ee4dc9b9fe853e893cc77c11c6d2b8a40ae692eb77d

Observation cdbde81f-bb83-4466-8f7c-4181ea1a7913 · outbound

This paper cites Training language models to follow instructions with human feedback.

Rethinking RL Evaluation: Can Benchmarks Truly Reveal Failures of RL Methods? Training language models to follow instructions with human feedback

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-04T10:20:18.758528Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:20:18.758528Z digest=sha256:3a72bab663e45ac3e215c83610168939a949ac7b0c2d5da484c0e75b7c4c0931

Observation ebcd2427-c0b6-4803-bf06-5750abe77b37 · outbound

This paper cites Curriculum reinforcement learning from easy to hard tasks improves llm reasoning.

Rethinking RL Evaluation: Can Benchmarks Truly Reveal Failures of RL Methods? Curriculum reinforcement learning from easy to hard tasks improves llm reasoning

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-04T10:20:18.878037Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:20:18.878037Z digest=sha256:a0eba2403bbc023785667a55c353d5ad3f3cbb86463b7c5d3e868f269f1c368f

Observation db9d3758-fb1d-40ca-a4a0-5cb6970d7162 · outbound

This paper cites Direct preference optimization: Your language model is secretly a reward model.

Rethinking RL Evaluation: Can Benchmarks Truly Reveal Failures of RL Methods? Direct preference optimization: Your language model is secretly a reward model

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-04T10:20:18.993321Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:20:18.993321Z digest=sha256:94186912b2faf2d8bd480e8f31cf4fec9430cf5e511157eb0d87309de4015529

Observation 983eeed1-c37b-46be-919d-6036b7315bb1 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Rethinking RL Evaluation: Can Benchmarks Truly Reveal Failures of RL Methods? Proximal Policy Optimization Algorithms

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-04T10:20:19.123246Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:20:19.123246Z digest=sha256:6b8e37e521bc8ad692a9f3cefeef33c50ea40a961dc77087740c0ad8ee5fb664

Observation 747ce24d-dcb6-4120-835c-625de9a36168 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Rethinking RL Evaluation: Can Benchmarks Truly Reveal Failures of RL Methods? DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-04T10:20:19.245701Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:20:19.245701Z digest=sha256:843dca3b8ff1df67eb14717014d1d47e99e37cbd4c96fd9ac320b0851eb07877

Observation 47e73c85-0fb2-4f3b-a262-063dcc4ff5a1 · outbound

This paper cites HEAD-QA: A Healthcare Dataset for Complex Reasoning.

Rethinking RL Evaluation: Can Benchmarks Truly Reveal Failures of RL Methods? HEAD-QA: A Healthcare Dataset for Complex Reasoning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-04T10:20:19.327712Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:20:19.327712Z digest=sha256:e632abe6e310e682467b786e203666c4d3371c8c991c0da75e27f84ffe129c0c

Observation e01159f6-6375-4e28-975d-3c8078019ee7 · outbound

This paper cites Self-Consistency Improves Chain of Thought Reasoning in Language Models.

Rethinking RL Evaluation: Can Benchmarks Truly Reveal Failures of RL Methods? Self-Consistency Improves Chain of Thought Reasoning in Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-04T10:20:19.428770Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:20:19.428770Z digest=sha256:fbc29a4f8dfecf1e611821a33e1212609c4263afcf992816388da0c8c7c8bc42

Observation 5f1baf7b-1d70-40de-a153-53e171474b09 · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models.

Rethinking RL Evaluation: Can Benchmarks Truly Reveal Failures of RL Methods? Chain-of-thought prompting elicits reasoning in large language models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-04T10:20:19.588462Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:20:19.588462Z digest=sha256:e4f1e7f8de13613e6670494980ebfb836f087f5857cb1f99b2d73aea6d5bf42c

Observation 9c7ad0c6-f9cd-4d50-9f3f-88f64a9e3845 · outbound

This paper cites Tree of Thoughts: Deliberate Problem Solving with Large Language Models.

Rethinking RL Evaluation: Can Benchmarks Truly Reveal Failures of RL Methods? Tree of Thoughts: Deliberate Problem Solving with Large Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-04T10:20:19.721870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:20:19.721870Z digest=sha256:89b9ba516c6a97002e16c525df8c7facd5b86d69502bd6f86ccb2b25662d27eb

Observation e7fc6145-cb8b-4ef3-9e68-4cf5b78681d8 · outbound

This paper cites Frame: Feedback-refined agent methodology for enhancing medical research insights.

Rethinking RL Evaluation: Can Benchmarks Truly Reveal Failures of RL Methods? Frame: Feedback-refined agent methodology for enhancing medical research insights

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-04T10:20:19.865669Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:20:19.865669Z digest=sha256:90456a2c846b68d97e470578f99c3eb836eaad6323a1362e8bb17695ab4327b6

Observation d3795470-4370-46d5-9301-1677b24c1bc2 · outbound

This paper cites SimpleRL-Zoo: Investigating and Taming Zero Reinforcement Learning for Open Base Models in the Wild.

Rethinking RL Evaluation: Can Benchmarks Truly Reveal Failures of RL Methods? SimpleRL-Zoo: Investigating and Taming Zero Reinforcement Learning for Open Base Models in the Wild

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-04T10:20:19.991179Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:20:19.991179Z digest=sha256:6cb5e082cfed41842a4d26f35c0b75285eb45e9e92697d771882e508a833b098

Observation 664256ce-62d2-4b7a-9711-f17a87d68503 · outbound

This paper cites LlamaFactory: Unified Efficient Fine-Tuning of 100+ Language Models.

Rethinking RL Evaluation: Can Benchmarks Truly Reveal Failures of RL Methods? LlamaFactory: Unified Efficient Fine-Tuning of 100+ Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-04T10:20:20.043540Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:20:20.043540Z digest=sha256:e15249668bfb80d6468a4db17958ea75e0da6c5f9155e77b14b7f48c5b01690f

Observation 85245822-8158-4d77-a0bb-7676a8d9b49a · outbound

This paper cites Easyr1: An efficient, scalable, multi-modality rl training framework.

Rethinking RL Evaluation: Can Benchmarks Truly Reveal Failures of RL Methods? Easyr1: An efficient, scalable, multi-modality rl training framework

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-04T10:20:20.091347Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:20:20.091347Z digest=sha256:7adcd7a64519c987afe58a314532532a7c434b61a4aa5ff84f37cae3a7a0dbaf

Observation af54b007-148c-4ebc-bc7c-6ce0efd556a9 · outbound

This paper cites write newline.

Rethinking RL Evaluation: Can Benchmarks Truly Reveal Failures of RL Methods? write newline

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-04T10:20:20.207076Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:20:20.207076Z digest=sha256:7c16c60b1235dd224bfc64defebdc46717053d2a508e93df290cfdf1f70487c9

Observation 32270c72-0952-4a5f-aea5-d7edcf434404 · outbound

This paper cites @esa (Ref.

Rethinking RL Evaluation: Can Benchmarks Truly Reveal Failures of RL Methods? @esa (Ref

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-04T10:20:20.302554Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:20:20.302554Z digest=sha256:345c4367dbcf9b79d1007fb78b8fe2f0c2753998f1d820d9543845b7a51d0ff3

Observation 700a9b75-6981-47c6-ba3c-aea8b9ed55e9 · outbound

This paper cites an unresolved cited work.

Rethinking RL Evaluation: Can Benchmarks Truly Reveal Failures of RL Methods? Unresolved cited work

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-04T10:20:20.377723Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:20:20.377723Z digest=sha256:14f090f506574fb87ca79ee9ed52138d8cdcc756c4bfff1df7dffd14186b753a

Observation ce937af4-a397-4290-9587-029f73a6da46 · outbound

This paper cites an unresolved cited work.

Rethinking RL Evaluation: Can Benchmarks Truly Reveal Failures of RL Methods? Unresolved cited work

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-04T10:20:20.407043Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T10:20:20.407043Z digest=sha256:4fd1ef232142396c23fde69132723be5fcd6f239b69d34d30cc6991419044987

Pith citing papers

Observation a8fbd985-0893-40d4-a625-a2a3988894ea · inbound

Beyond One-Size-Fits-All: Diagnosis-Driven Online Reinforcement Learning with Offline Priors cites this paper.

Beyond One-Size-Fits-All: Diagnosis-Driven Online Reinforcement Learning with Offline Priors Rethinking RL Evaluation: Can Benchmarks Truly Reveal Failures of RL Methods?

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-07-04T19:30:07.859170Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-06-25T21:15:07.735336Z digest=sha256:282849fd2570158eaa8f7609940eb8045f3d73b81fda3d0eb17f2083f16119d0