Pith. sign in

Paper Citation Record · LEDGER

CVPO: Enhancing LLM Reinforcement Learning Reasoning via Value-Variance Adaptation and Dynamic Curriculum Learning

As of 15 August 2026, this Paper Citation Record lists 18 of 18 outbound references and 0 inbound Pith citation observations for arXiv:2608.03068.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.03068 v1

Coverage vector

measured 18 of 18 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-08T01:03:50.900234Z

measured 18 of 18 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

18 of 18 outbound references displayed

  • verified exact1
  • verified fuzzy1
  • unresolved16
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 198e8273-1157-42ea-87a4-f1de19bf9c72 · outbound

This paper cites DeepSeek-V3 Technical Report.

CVPO: Enhancing LLM Reinforcement Learning Reasoning via Value-Variance Adaptation and Dynamic Curriculum Learning DeepSeek-V3 Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-08T01:03:50.820158Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T01:03:50.820158Z digest=sha256:ba1587189131df59ca751b6cd996a4b444bfdfa9e326e1f3e64534a7a01f1b4d

Observation f49a1024-cbde-4c97-a3d1-15074d109f82 · outbound

This paper cites Open-Reasoner-Zero: An Open Source Approach to Scaling Up Reinforcement Learning on the Base Model.

CVPO: Enhancing LLM Reinforcement Learning Reasoning via Value-Variance Adaptation and Dynamic Curriculum Learning Open-Reasoner-Zero: An Open Source Approach to Scaling Up Reinforcement Learning on the Base Model

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-08T01:03:50.835692Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T01:03:50.835692Z digest=sha256:6b07015e0d48f4415fafbebe7259fe160a9252882bd5296ece484b1926aa388c

Observation c6307409-bfff-417d-96c1-a50c2b545310 · outbound

This paper cites Mihir Prabhudesai, Lili Chen, Alex Ippoliti, Katerina Fragkiadaki, Hao Liu, and Deepak Pathak.

CVPO: Enhancing LLM Reinforcement Learning Reasoning via Value-Variance Adaptation and Dynamic Curriculum Learning Mihir Prabhudesai, Lili Chen, Alex Ippoliti, Katerina Fragkiadaki, Hao Liu, and Deepak Pathak

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-08T01:03:50.840766Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T01:03:50.840766Z digest=sha256:875511be47708e3d5583954598325629093527ce58c0f8878ea905cd1f68465f

Observation d7e09ead-8a90-4a9e-95fc-49075a756c03 · outbound

This paper cites Maximizing Confidence Alone Improves Reasoning.

CVPO: Enhancing LLM Reinforcement Learning Reasoning via Value-Variance Adaptation and Dynamic Curriculum Learning Maximizing Confidence Alone Improves Reasoning

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-08T01:03:50.846084Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T01:03:50.846084Z digest=sha256:fc5f67f47ea2e195e46121b55773ae5c9cb8a68c3792d0d38cd67180b5503ff7

Observation 4bb7b8b2-740b-4747-ab40-041cd0009fed · outbound

This paper cites an unresolved cited work.

CVPO: Enhancing LLM Reinforcement Learning Reasoning via Value-Variance Adaptation and Dynamic Curriculum Learning Unresolved cited work

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-08T01:03:50.851038Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T01:03:50.851038Z digest=sha256:23735a36db5c4742c6d5e903c3cf3da57201d14277dab75419ded9932fbae9e1

Observation c48d44f1-d719-44e6-8d60-4d963273726d · outbound

This paper cites Offline Regularised Reinforcement Learning for Large Language Models Alignment.

CVPO: Enhancing LLM Reinforcement Learning Reasoning via Value-Variance Adaptation and Dynamic Curriculum Learning Offline Regularised Reinforcement Learning for Large Language Models Alignment

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-08T01:03:50.856286Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T01:03:50.856286Z digest=sha256:7248778e425d6620873ec28ac6f570a5042d511bfe5715373e71ae652993a458

Observation d7ab1b69-45a5-4627-893c-40d3a1179c0b · outbound

This paper cites Spurious Rewards: Rethinking Training Signals in RLVR.

CVPO: Enhancing LLM Reinforcement Learning Reasoning via Value-Variance Adaptation and Dynamic Curriculum Learning Spurious Rewards: Rethinking Training Signals in RLVR

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-08T01:03:50.867049Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T01:03:50.867049Z digest=sha256:b79edd6ea33843ae9c456278ccc47c7b8aab47c6eff269e6871894cf29e5fbcf

Observation 06828e25-d770-4794-91aa-04f2467a6c68 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

CVPO: Enhancing LLM Reinforcement Learning Reasoning via Value-Variance Adaptation and Dynamic Curriculum Learning DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-08T01:03:50.872022Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T01:03:50.872022Z digest=sha256:a01ec3cbee4e0266f17c733b0f9e132c0fdbd732b0b9f8c7260fd63f1ac14cb7

Observation bb9eb6a6-0b0a-470f-be7b-34ee3c3bd211 · outbound

This paper cites Efficient Reinforcement Finetuning via Adaptive Curriculum Learning.

CVPO: Enhancing LLM Reinforcement Learning Reasoning via Value-Variance Adaptation and Dynamic Curriculum Learning Efficient Reinforcement Finetuning via Adaptive Curriculum Learning

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-08T01:03:50.877512Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T01:03:50.877512Z digest=sha256:a28ddba364b0b16b1804d35fb32ef4f510804dc5f898e4098701b330dd055064

Observation 656f73ed-543a-4fa8-9343-6d6f4884cfd9 · outbound

This paper cites DGRO: Enhancing LLM Reasoning via Exploration-Exploitation Control and Reward Variance Management.

CVPO: Enhancing LLM Reinforcement Learning Reasoning via Value-Variance Adaptation and Dynamic Curriculum Learning DGRO: Enhancing LLM Reasoning via Exploration-Exploitation Control and Reward Variance Management

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-08-08T01:03:51.005031Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-08T01:03:50.882610Z digest=sha256:86e362b582d8afb7263271ad2116c554ce59c61896b731214aeed190f9cbb35f

Observation 04e254a3-3d8f-4309-b802-dd0885ef1db6 · outbound

This paper cites Kimi k1.5: Scaling Reinforcement Learning with LLMs.

CVPO: Enhancing LLM Reinforcement Learning Reasoning via Value-Variance Adaptation and Dynamic Curriculum Learning Kimi k1.5: Scaling Reinforcement Learning with LLMs

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-08T01:03:50.887036Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T01:03:50.887036Z digest=sha256:50880d0ddf4205629e9ad0760dfb0c9dc9ce308bd1d0db2db28d70999661343a

Observation 234fcb07-92af-4807-996e-7e34e791be50 · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

CVPO: Enhancing LLM Reinforcement Learning Reasoning via Value-Variance Adaptation and Dynamic Curriculum Learning DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-08T01:03:50.891503Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T01:03:50.891503Z digest=sha256:3bb00aa7f9d6ae6b877a1a37c1e2cba65927b656aca44d5f4aed7f70cc47a5c2

Observation 47b2cca1-ef7f-429d-a4b9-b4f81fd44a93 · outbound

This paper cites What's Behind PPO's Collapse in Long-CoT? Value Optimization Holds the Secret.

CVPO: Enhancing LLM Reinforcement Learning Reasoning via Value-Variance Adaptation and Dynamic Curriculum Learning What's Behind PPO's Collapse in Long-CoT? Value Optimization Holds the Secret

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-08T01:03:50.895805Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T01:03:50.895805Z digest=sha256:eb0e87a5c4ff2c496230439678c5ce51ae945bd6c96ed7f61589c5278a874cd7

Observation f64dcd2a-c888-4b80-90f1-c759daa5af2f · outbound

This paper cites VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks.

CVPO: Enhancing LLM Reinforcement Learning Reasoning via Value-Variance Adaptation and Dynamic Curriculum Learning VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-08T01:03:50.900234Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T01:03:50.900234Z digest=sha256:0f76c61aad218131242dc28c4a00e134b18a8da0316b61ffd45affc5643bef2c

Observation 53510328-38fc-4371-ab95-f428ae84eb73 · outbound

This paper cites Proximal Policy Optimization Algorithms.

CVPO: Enhancing LLM Reinforcement Learning Reasoning via Value-Variance Adaptation and Dynamic Curriculum Learning Proximal Policy Optimization Algorithms

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-08T01:03:50.862158Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T01:03:50.862158Z digest=sha256:a250314d682890c582c6f4b597dd6874342827d0dfe7109087cf37c84c1b7f15

Observation f05c8805-548d-48d0-8a4b-158208d944e1 · outbound

This paper cites InProceedings of the Neural Information Processing Systems Track on Datasets and Benchmarks 1, NeurIPS Datasets and Benchmarks 2021, December 2021, virtual.

CVPO: Enhancing LLM Reinforcement Learning Reasoning via Value-Variance Adaptation and Dynamic Curriculum Learning InProceedings of the Neural Information Processing Systems Track on Datasets and Benchmarks 1, NeurIPS Datasets and Benchmarks 2021, December 2021, virtual

Reference 2021

Resolution
verified fuzzy
raw_fallback, observed 2026-08-08T01:03:51.347930Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-08T01:03:50.825233Z digest=sha256:e3ac8073aec407423774f453d6338999a360b8604bf18f00bfa8e815a19c2144

Observation 471fe5c0-bde5-4e7e-a082-1e7ca1a9f9f6 · outbound

This paper cites MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention.

CVPO: Enhancing LLM Reinforcement Learning Reasoning via Value-Variance Adaptation and Dynamic Curriculum Learning MiniMax-M1: Scaling Test-Time Compute Efficiently with Lightning Attention

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-08T01:03:50.814168Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T01:03:50.814168Z digest=sha256:bb12820d115f5bfda3db527a8c60f5719cc7f34ddcb47c718e6e56124afbea6f

Observation da67af90-fa16-46a0-a5c0-349c22f72ca6 · outbound

This paper cites REINFORCE++: Stabilizing Critic-Free Policy Optimization with Global Advantage Normalization.

CVPO: Enhancing LLM Reinforcement Learning Reasoning via Value-Variance Adaptation and Dynamic Curriculum Learning REINFORCE++: Stabilizing Critic-Free Policy Optimization with Global Advantage Normalization

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-08T01:03:50.829866Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T01:03:50.829866Z digest=sha256:ebe80182d3c64b880bb1c2295d67e21c730549989ac31c4a95030b2e6b2d48eb

Pith citing papers

No inbound Pith citation observations are available.