Pith. sign in

Paper Citation Record · LEDGER

Optimizing Chain-of-Thought Reasoners via Gradient Variance Minimization in Rejection Sampling and RL

As of 15 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 9 inbound Pith citation observations for arXiv:2505.02391.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.02391 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 9 of 9 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 9 of 9 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T06:01:49.664298Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T06:06:41.070540Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation bccae710-a182-47cd-93e1-f582b97def4c · inbound

SIGMA: Refining Large Language Model Reasoning via Sibling-Guided Monte Carlo Augmentation cites this paper.

SIGMA: Refining Large Language Model Reasoning via Sibling-Guided Monte Carlo Augmentation Optimizing Chain-of-Thought Reasoners via Gradient Variance Minimization in Rejection Sampling and RL

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T06:01:49.664298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T06:01:49.664298Z digest=sha256:aafd96a50f14291e45ded5a698aa12f6cf12ae98bd988492b10ae1ded7c74779

Observation 41bb5f15-b3ab-47ba-9182-7d978c723933 · inbound

Variance-Aware Baselines and Adaptive Learning Rates for Reinforcement Learning with Verifiable Rewards cites this paper.

Variance-Aware Baselines and Adaptive Learning Rates for Reinforcement Learning with Verifiable Rewards Optimizing Chain-of-Thought Reasoners via Gradient Variance Minimization in Rejection Sampling and RL

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-03T19:38:23.430181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T19:38:23.430181Z digest=sha256:d30e93d21370b061170ffe8b6a95d59d9cc01991f965dff5114595a30f877345

Observation f0549afe-c503-48f6-b347-d34c7c8c723c · inbound

Variance-Aware Baselines and Adaptive Learning Rates for Reinforcement Learning with Verifiable Rewards cites this paper.

Variance-Aware Baselines and Adaptive Learning Rates for Reinforcement Learning with Verifiable Rewards Optimizing Chain-of-Thought Reasoners via Gradient Variance Minimization in Rejection Sampling and RL

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-04T06:47:16.900557Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:47:16.900557Z digest=sha256:886e8c3a9454a3fc63c0dcff00b281770e9af8624183ab9c605401a6de7017a4

Observation c27f7bb9-1692-4506-9202-d39226132c92 · inbound

A Task-Centric Theory for Iterative Self-Improvement with Easy-to-Hard Curricula cites this paper.

A Task-Centric Theory for Iterative Self-Improvement with Easy-to-Hard Curricula Optimizing Chain-of-Thought Reasoners via Gradient Variance Minimization in Rejection Sampling and RL

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-03T02:43:38.039514Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T02:43:38.039514Z digest=sha256:458deaf4260d4f58393ba125d7a68f023560155863c0b4a2cba4e31673e68981

Observation 5063b41d-8853-4fbb-88d2-cb6914f75e65 · inbound

VI-CuRL: Stabilizing Verifier-Independent RL Reasoning via Confidence-Guided Variance Reduction cites this paper.

VI-CuRL: Stabilizing Verifier-Independent RL Reasoning via Confidence-Guided Variance Reduction Optimizing Chain-of-Thought Reasoners via Gradient Variance Minimization in Rejection Sampling and RL

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-25T07:26:41.631562Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-25T07:26:35.767179Z digest=sha256:dccd9a49beee54c9f669e1af3d17820d88cb9043d40216213d9ad78dab03929c

Observation 3e1a0d99-f6cc-429c-96a5-c48079f6e7c8 · inbound

Your Model Diversity, Not Method, Determines Reasoning Strategy cites this paper.

Your Model Diversity, Not Method, Determines Reasoning Strategy Optimizing Chain-of-Thought Reasoners via Gradient Variance Minimization in Rejection Sampling and RL

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-11T10:56:02.867497Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-05-10T15:16:30.448400Z digest=sha256:927bdd657b70448ff9b105fa7dbbe2348a0e9df19aeffce72df87fbb6caeaf8b

Observation 9f7167e3-4397-4ab1-b6f7-9864df108db4 · inbound

Rethinking Importance Sampling in LLM Policy Optimization: A Cumulative Token Perspective cites this paper.

Rethinking Importance Sampling in LLM Policy Optimization: A Cumulative Token Perspective Optimizing Chain-of-Thought Reasoners via Gradient Variance Minimization in Rejection Sampling and RL

Reference 11

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T04:30:55.487341Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-05-11T01:19:58.448344Z digest=sha256:bd1589bbcac77d6235e7a2172233d0871edff1def8b6bf0b3fdada807e953df3

Observation b6111562-bf36-46ae-bf57-8811dabd805c · inbound

Rollout-Level Advantage-Prioritized Experience Replay for GRPO cites this paper.

Rollout-Level Advantage-Prioritized Experience Replay for GRPO Optimizing Chain-of-Thought Reasoners via Gradient Variance Minimization in Rejection Sampling and RL

Reference 72

Resolution
verified exact
arxiv_id, observed 2026-07-02T06:06:41.072433Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-06-28T07:43:21.284574Z digest=sha256:1381f163bb2409c1a80aef9ddab279be82909632ab6563d1bc8ceeb7111e368c

Observation 514e9eb2-f1de-49df-ad60-8c9559b9b20c · inbound

Distill Where You Fail: Recovering Learning Signals of Negative RL-Groups from Adaptive Teacher Guidance cites this paper.

Distill Where You Fail: Recovering Learning Signals of Negative RL-Groups from Adaptive Teacher Guidance Optimizing Chain-of-Thought Reasoners via Gradient Variance Minimization in Rejection Sampling and RL

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-05T00:22:11.865683Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T00:22:11.865683Z digest=sha256:4d1bdeca1dc5f939b9b8eeaaeb4b82ac2a192361f3c8f82d6cbb6f3e5843b0f6