Pith. sign in

Paper Citation Record · LEDGER

Towards High Data Efficiency in Reinforcement Learning with Verifiable Reward

As of 11 August 2026, this Paper Citation Record lists 21 of 21 outbound references and 7 inbound Pith citation observations for arXiv:2509.01321.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2509.01321 v1

Coverage vector

measured 21 of 21 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T12:45:41.982954Z

measured 28 of 28 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00

measured 7 of 7 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-29T14:49:39.188220Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T03:09:29.035824Z

Reference resolution

21 of 21 outbound references displayed

  • verified exact0
  • verified fuzzy7
  • unresolved14
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 88bfa3fa-7790-4cb2-a1f6-177f484a5ec9 · outbound

This paper cites This objective aims to select a subset Y that is both diverse (as captured by det(SY)) and influential (as promoted by the product of weights ∏i∈Y wi).

Towards High Data Efficiency in Reinforcement Learning with Verifiable Reward This objective aims to select a subset Y that is both diverse (as captured by det(SY)) and influential (as promoted by the product of weights ∏i∈Y wi)

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:45:42.198696Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-05T12:45:41.966040Z digest=sha256:07fd76acef0ecabfb0f9926bce608b3f40e37d9c8705b3f5e795d6476dadbd0b

Observation 5a34c283-42b9-4ecd-8940-9542f172ac93 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Towards High Data Efficiency in Reinforcement Learning with Verifiable Reward DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-05T12:45:41.919862Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:45:41.919862Z digest=sha256:5e74ba527732890b148aae603510b2e7614a30b59bd67da9105ead44a107b7a0

Observation c075a77f-a5b2-4ff2-bfd0-c722d676654b · outbound

This paper cites Determinantal point processes for machine learning.

Towards High Data Efficiency in Reinforcement Learning with Verifiable Reward Determinantal point processes for machine learning

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-05T12:45:41.926767Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:45:41.926767Z digest=sha256:d169c11c1ce98717cd3790af38e584b0e4bb379b6b6b6db1d5ef5393dbdb4f47

Observation c9397d96-dbf1-4a87-929b-45ee2b2021b6 · outbound

This paper cites GPQA: A Graduate-Level Google-Proof Q&A Benchmark.

Towards High Data Efficiency in Reinforcement Learning with Verifiable Reward GPQA: A Graduate-Level Google-Proof Q&A Benchmark

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-05T12:45:41.937712Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:45:41.937712Z digest=sha256:1be0ea008fa5dc0d6ab1ec8831c7f346fedcc90b814db39648a5aaadade36a79

Observation 0bd7a51a-db29-4631-b72a-09e361584711 · outbound

This paper cites Kimi k1.5: Scaling Reinforcement Learning with LLMs.

Towards High Data Efficiency in Reinforcement Learning with Verifiable Reward Kimi k1.5: Scaling Reinforcement Learning with LLMs

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-05T12:45:41.944746Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:45:41.944746Z digest=sha256:afa6f7cddeab991c028f494725d5a2120abca431045787a679b8d28b5272fff3

Observation d1159417-894d-428f-a0b1-3afcd9f06bff · outbound

This paper cites Beyond the 80/20 Rule: High-Entropy Minority Tokens Drive Effective Reinforcement Learning for LLM Reasoning.

Towards High Data Efficiency in Reinforcement Learning with Verifiable Reward Beyond the 80/20 Rule: High-Entropy Minority Tokens Drive Effective Reinforcement Learning for LLM Reasoning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-05T12:45:41.948669Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:45:41.948669Z digest=sha256:671b729d0d4036f212ad6c217273aa6d3ea796418eeb3a9eaec565ae47eba14f

Observation a9dcc236-a80b-4850-bcbb-1a067acc5307 · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

Towards High Data Efficiency in Reinforcement Learning with Verifiable Reward DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-05T12:45:41.952290Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:45:41.952290Z digest=sha256:78b7bf32bdf7bd2f30802fcf26e55a1e5e02441d5b003ece267797f535cbeaf9

Observation 43746aa5-ec9c-4df8-8e5c-4f0de72a7011 · outbound

This paper cites HARP: A challenging human-annotated math reasoning benchmark.

Towards High Data Efficiency in Reinforcement Learning with Verifiable Reward HARP: A challenging human-annotated math reasoning benchmark

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-05T12:45:41.956023Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:45:41.956023Z digest=sha256:c0fce67f38f42cbf1c17b52b2d6d19660686ff43d388441859286ca9c5acb776

Observation bcf8172e-f661-4089-b7e5-131ca1103dfb · outbound

This paper cites VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks.

Towards High Data Efficiency in Reinforcement Learning with Verifiable Reward VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-05T12:45:41.959563Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:45:41.959563Z digest=sha256:0de7f14a82eb780ec853b1c92b4c66b68375fd26fd01dcaeb60570458c5b4a1a

Observation 5e0a894f-8988-445e-8468-2d97eeda04ba · outbound

This paper cites A Survey of Large Language Models.

Towards High Data Efficiency in Reinforcement Learning with Verifiable Reward A Survey of Large Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-05T12:45:41.962651Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:45:41.962651Z digest=sha256:843e7aaba728bc2decac2b5001ed1c37c6c13d8c8a8bb684e8335854c3c39bb5

Observation 48f3f476-4dc4-4040-9d57-7bfa4cf05fbe · outbound

This paper cites We train DeepSeek-R1-Distill-Qwen-7B and DeepSeek-R1-Distill-Llama-8B on 64×H200 GPUs, and Qwen2.5-Math-7B on 32×H200 GPUs.

Towards High Data Efficiency in Reinforcement Learning with Verifiable Reward We train DeepSeek-R1-Distill-Qwen-7B and DeepSeek-R1-Distill-Llama-8B on 64×H200 GPUs, and Qwen2.5-Math-7B on 32×H200 GPUs

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:45:42.188191Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-05T12:45:41.969785Z digest=sha256:c3f29c2403314993b75abaadf4aab6bb1e951dd7f9f29c1f17e613f678972a5b

Observation 6e6c5ecd-519d-4fa9-952d-eb28c627c9b0 · outbound

This paper cites We follow Zheng et al.

Towards High Data Efficiency in Reinforcement Learning with Verifiable Reward We follow Zheng et al

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:45:42.158494Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-05T12:45:41.979652Z digest=sha256:22cc88f8b02435ea90d6ea537d759dc9c9f1bbbbb4660b1fe07f2c84cd00473e

Observation c3b2bdac-a2c9-44be-987f-c4d4608ca16d · outbound

This paper cites • Random: Randomly samples data from the training set.

Towards High Data Efficiency in Reinforcement Learning with Verifiable Reward • Random: Randomly samples data from the training set

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:45:42.148614Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-05T12:45:41.982954Z digest=sha256:e7b5e072b8466190bfd64eeae99249369a676929d0cb50ba4844744e09a693b7

Observation ceba9f09-7548-451a-a022-897a36409df8 · outbound

This paper cites Similar to Yue et al.

Towards High Data Efficiency in Reinforcement Learning with Verifiable Reward Similar to Yue et al

Reference 256

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:45:42.178162Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-05T12:45:41.973389Z digest=sha256:4c7f5928d18b5bb503dc67266f93ac60e8dd593f8d0f362a24d7f425b44ad758

Observation c3b6a25e-7df7-48b1-9668-7093fca649c4 · outbound

This paper cites Reasoning with Exploration: An Entropy Perspective.

Towards High Data Efficiency in Reinforcement Learning with Verifiable Reward Reasoning with Exploration: An Entropy Perspective

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-05T12:45:41.916058Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:45:41.916058Z digest=sha256:982c9c695240156892936190bcb39a421aacf02fafb9c457c78fb9534c4a872c

Observation 05558ae2-4940-4aa7-8ec2-355ffabdf59b · outbound

This paper cites User: \n [question] \n Please reason step by step, and put your final answer within \boxed{}. \n \n Assistant:.

Towards High Data Efficiency in Reinforcement Learning with Verifiable Reward User: \n [question] \n Please reason step by step, and put your final answer within \boxed{}. \n \n Assistant:

Reference 2019

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:45:42.168425Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-05T12:45:41.976582Z digest=sha256:a62b9b0813df0655c74761c50578de9c87b3f7e879561027d661c2432f1bee4d

Observation 5bda1edf-bfb7-4388-b30d-6e23b180f3e0 · outbound

This paper cites OpenAI o1 System Card.

Towards High Data Efficiency in Reinforcement Learning with Verifiable Reward OpenAI o1 System Card

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-05T12:45:41.923179Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:45:41.923179Z digest=sha256:f66ea3a2542694035867fde73c5e83b730989c7115201112de7fa9bc453fc2ea

Observation 2940da6b-fcee-47ff-9612-fdcb03b71617 · outbound

This paper cites LearnAlign: Data Selection for LLM Reinforcement Learning with Improved Gradient Alignment.

Towards High Data Efficiency in Reinforcement Learning with Verifiable Reward LearnAlign: Data Selection for LLM Reinforcement Learning with Improved Gradient Alignment

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-05T12:45:41.930273Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:45:41.930273Z digest=sha256:ec310d296f823fa26df5f7fb22d7ef565c4d3a57b1b0c17efa43f996e245fa85

Observation 6deb6c8d-754b-4b7d-ab59-6d44bd4c19f8 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Towards High Data Efficiency in Reinforcement Learning with Verifiable Reward DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-05T12:45:41.941078Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:45:41.941078Z digest=sha256:e592a76ae7947b6609bd90fb9fb6f50867f93571f682e47359cce70b8f0ce6c9

Observation b113e1de-a27c-4a7e-9a9b-8bb9bd634cc6 · outbound

This paper cites Understanding R1-Zero-Like Training: A Critical Perspective.

Towards High Data Efficiency in Reinforcement Learning with Verifiable Reward Understanding R1-Zero-Like Training: A Critical Perspective

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-05T12:45:41.933948Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T12:45:41.933948Z digest=sha256:2166cd8a564fa8df4fa7e2a1e0cf3bad49186f38db6e1779ffcdc57f71f01394

Observation fa6945b6-7936-49d9-b632-7454cdcb75f3 · outbound

This paper cites Zachary Ankner, Cody Blakeney, Kartik Sreenivasan, Max Marion, Matthew L.

Towards High Data Efficiency in Reinforcement Learning with Verifiable Reward Zachary Ankner, Cody Blakeney, Kartik Sreenivasan, Max Marion, Matthew L

Reference 2025

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T12:45:42.208226Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-05T12:45:41.912268Z digest=sha256:073299e70e910291e89317e2c4192d91dea61638605032c40e5a727cf76ee43c

Pith citing papers

Observation c7f55504-31e6-49ef-94f0-d257b98496ea · inbound

Beyond Majority Voting: Towards Fine-grained and More Reliable Reward Signal for Test-Time Reinforcement Learning cites this paper.

Beyond Majority Voting: Towards Fine-grained and More Reliable Reward Signal for Test-Time Reinforcement Learning Towards High Data Efficiency in Reinforcement Learning with Verifiable Reward

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T21:53:35.117871Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-16T21:53:30.911773Z digest=sha256:beab0af748a16201c8abe9155f128aa893cb16f72be267c42c6f167f33542c7b

Observation 0017b414-69ab-4557-ba87-51182b5626e4 · inbound

DARE: Difficulty-Adaptive Reinforcement Learning with Co-Evolved Difficulty Estimation cites this paper.

DARE: Difficulty-Adaptive Reinforcement Learning with Co-Evolved Difficulty Estimation Towards High Data Efficiency in Reinforcement Learning with Verifiable Reward

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-05-12T03:16:19.312129Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-12T03:12:49.428954Z digest=sha256:278788227d52215509f2c9bb18702b9385cb939b65c42c7920058a7f9bbf3558

Observation 3a17c0a3-cb51-4d50-becd-eea6adcefb54 · inbound

IRDS: Interpretable RLVR Data Selection via Verifier-Coupled Sparse Autoencoder Coverage cites this paper.

IRDS: Interpretable RLVR Data Selection via Verifier-Coupled Sparse Autoencoder Coverage Towards High Data Efficiency in Reinforcement Learning with Verifiable Reward

Reference 5

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T14:53:31.190838Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-29T14:49:39.188220Z digest=sha256:cc0a184e223229dd85404ae2abd0374c0a38865eaa01749fb3cc2cf88c2f2e72

Observation 30d0e72f-8d3b-44f4-9bf3-76c60791f499 · inbound

RLVR without Ineffective Samples: Group Prioritized Off-Policy Optimization for LLM Reasoning cites this paper.

RLVR without Ineffective Samples: Group Prioritized Off-Policy Optimization for LLM Reasoning Towards High Data Efficiency in Reinforcement Learning with Verifiable Reward

Reference 50

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T21:16:13.366003Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-28T17:25:50.758630Z digest=sha256:bb1b9bad4e1cd95c7e23f12801f8f3ae8eb0849b9ca9194551f7b28e70999d92

Observation 0d6c5df4-001e-4297-abea-637ad4718b16 · inbound

Smart Picks in the Dark: Towards Efficient RLVR for Reasoning via Tracing Metacognitive Pivots cites this paper.

Smart Picks in the Dark: Towards Efficient RLVR for Reasoning via Tracing Metacognitive Pivots Towards High Data Efficiency in Reinforcement Learning with Verifiable Reward

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-07-02T07:26:46.260794Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-28T06:55:09.927034Z digest=sha256:a337fcb1f9c4796132ee821eaee2066a96e52df825f443a74794d13ab9ed7a23

Observation 642a7c29-2225-48ce-a3f3-600d24e1f83a · inbound

Reasoning Arena: Trace Tournaments When Verifiable Rewards Fall Short cites this paper.

Reasoning Arena: Trace Tournaments When Verifiable Rewards Fall Short Towards High Data Efficiency in Reinforcement Learning with Verifiable Reward

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-07-03T00:17:29.544665Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-27T17:18:03.459289Z digest=sha256:51c2dd6e44c801382e993f5d8bb4f896baab6fc39ba6a42b423c87ef6d4df3c0

Observation c639ed77-de36-4ef0-ae45-59cbff164918 · inbound

Manifold Bandits: Bayesian Curriculum Learning over the Latent Geometry of Large Language Models cites this paper.

Manifold Bandits: Bayesian Curriculum Learning over the Latent Geometry of Large Language Models Towards High Data Efficiency in Reinforcement Learning with Verifiable Reward

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-07-04T03:09:29.037593Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-26T18:29:47.613424Z digest=sha256:74296f3342d459be098fbbc9d54f591ded5446c7a9ba877ec3135df818f8e8f7