Pith. sign in

Paper Citation Record · LEDGER

CEPO: RLVR Self-Distillation using Contrastive Evidence Policy Optimization

As of 16 August 2026, this Paper Citation Record lists 29 of 29 outbound references and 2 inbound Pith citation observations for arXiv:2605.19436.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.19436 v1

Coverage vector

measured 29 of 29 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-20T07:25:34.637803Z

measured 31 of 31 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-11T04:17:44.132452Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-11T04:17:44.759376Z

Reference resolution

29 of 29 outbound references displayed

  • verified exact18
  • verified fuzzy10
  • unresolved1
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation cd16c89e-b10c-4460-8175-6f7dd9798967 · outbound

This paper cites Qwen3-VL Technical Report.

CEPO: RLVR Self-Distillation using Contrastive Evidence Policy Optimization Qwen3-VL Technical Report

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-20T07:28:06.692342Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-20T07:25:34.637803Z digest=sha256:907760d424460028e0662cd3c405ec9a96834b6ce1059a4031129b68b5838920

Observation 9f7f6f2a-cbd1-4c62-8605-8652dff3c5b9 · outbound

This paper cites Enhancing reinforcement learning with dense rewards from language model critic.

CEPO: RLVR Self-Distillation using Contrastive Evidence Policy Optimization Enhancing reinforcement learning with dense rewards from language model critic

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T07:28:07.789041Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-20T07:25:34.637803Z digest=sha256:de66f3bc2a4767712dac3ff140b7aca487c711bc74d006b6dab60370637a1557

Observation de406f7d-30ce-4ab4-9bfa-dee53628e503 · outbound

This paper cites Hdpo: Hybrid distillation policy optimization via privileged self-distillation.

CEPO: RLVR Self-Distillation using Contrastive Evidence Policy Optimization Hdpo: Hybrid distillation policy optimization via privileged self-distillation

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-20T07:28:06.732172Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-20T07:25:34.637803Z digest=sha256:985c74291d71f80f0d24278f91bd258af4dd4ec54160530bc0080b6710dfb0ed

Observation 49e1f607-f397-46ee-be6e-07acc6a6a873 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

CEPO: RLVR Self-Distillation using Contrastive Evidence Policy Optimization DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-05-20T07:28:06.715382Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-20T07:25:34.637803Z digest=sha256:dc0695bd9b9a67b424a4e8e0bb96928c45a32ff158c2e8547f535c17331bfca7

Observation 3bba8b69-66f5-4828-818a-8142788905ab · outbound

This paper cites Segment policy optimization: Ef- fective segment-level credit assignment in rl for large language models.

CEPO: RLVR Self-Distillation using Contrastive Evidence Policy Optimization Segment policy optimization: Ef- fective segment-level credit assignment in rl for large language models

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-20T07:28:06.696139Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-20T07:25:34.637803Z digest=sha256:8df640a1d8ee08f70c8adfaad516ef18b07403fe47955c45db7c3dc5aa0e94aa

Observation bfb0cdab-fa69-4677-b91b-e458441e5908 · outbound

This paper cites Reinforcement Learning via Self-Distillation.

CEPO: RLVR Self-Distillation using Contrastive Evidence Policy Optimization Reinforcement Learning via Self-Distillation

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-05-20T07:28:06.740802Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-20T07:25:34.637803Z digest=sha256:13c9359fd7b8768e02a253c358a82bad1c2d7473e2e1e1feb37aa150eba5dcf4

Observation d41cce22-37b6-41db-8958-43dea5070073 · outbound

This paper cites Vineppo: Refining credit assignment in rl training of llms.

CEPO: RLVR Self-Distillation using Contrastive Evidence Policy Optimization Vineppo: Refining credit assignment in rl training of llms

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T07:28:07.806585Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-20T07:25:34.637803Z digest=sha256:c65f9794c881e3278dd0642416b489d1c9ebe9d0de8aac077bebe17fc2231333

Observation f752aaaa-ec60-4f3c-8bca-5ec713ef0259 · outbound

This paper cites Efficient memory management for large language model serving with pagedattention.

CEPO: RLVR Self-Distillation using Contrastive Evidence Policy Optimization Efficient memory management for large language model serving with pagedattention

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T07:28:07.803751Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-20T07:25:34.637803Z digest=sha256:15d722a8a71ca166b6d000770237e5348cefab425b3971da84e2f0d8bf8a2cf1

Observation fb9b350e-582e-426b-b3b5-7c252a10b441 · outbound

This paper cites Let’s verify step by step.

CEPO: RLVR Self-Distillation using Contrastive Evidence Policy Optimization Let’s verify step by step

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T07:28:07.801519Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-20T07:25:34.637803Z digest=sha256:86e48ae169fba8689d353578a42c82bd0966fb3992c402972b048f6177efebec

Observation 9437ccef-8cb4-4beb-b315-9edb91f04ed3 · outbound

This paper cites Decoupled Weight Decay Regularization.

CEPO: RLVR Self-Distillation using Contrastive Evidence Policy Optimization Decoupled Weight Decay Regularization

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-05-20T07:28:06.743892Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-20T07:25:34.637803Z digest=sha256:24aad42b17a06de3c8fcb2b68f8c98e82b4f67cfbfd5a1ff7024b1f238b34e1a

Observation ef5f2040-11c3-4095-ad24-8bf13606edf0 · outbound

This paper cites Inter-gps: Interpretable geometry problem solving with formal language and symbolic reasoning.

CEPO: RLVR Self-Distillation using Contrastive Evidence Policy Optimization Inter-gps: Interpretable geometry problem solving with formal language and symbolic reasoning

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T07:28:07.808483Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-20T07:25:34.637803Z digest=sha256:fb2a8bc893d6206fffd11cb1f66879c7e56d76ef2c9a2e426f0c728f4e223a87

Observation e0793516-e4ec-483a-b45b-b9f87305e1cf · outbound

This paper cites Privileged Information Distillation for Language Models.

CEPO: RLVR Self-Distillation using Contrastive Evidence Policy Optimization Privileged Information Distillation for Language Models

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-22T08:25:38.085126Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-20T07:25:34.637803Z digest=sha256:5be15f964ae26fbcec6f8d3ed012bd3d2038d4eafa2081357bce9865e31bca2d

Observation 6b6b04f9-7334-44dc-8466-e12506ac3660 · outbound

This paper cites an unresolved cited work.

CEPO: RLVR Self-Distillation using Contrastive Evidence Policy Optimization Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-05-20T07:28:07.810293Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-20T07:25:34.637803Z digest=sha256:822f431db285b1b727994de3fe05090eb36ab51c1ebafa82dc7be03dc70b7e1d

Observation 813132db-5d51-4a2d-8414-50ecf1838102 · outbound

This paper cites Direct preference optimization: Your language model is secretly a reward model.

CEPO: RLVR Self-Distillation using Contrastive Evidence Policy Optimization Direct preference optimization: Your language model is secretly a reward model

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T07:28:07.797670Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-20T07:25:34.637803Z digest=sha256:4ec183f0e7449ed5152496058968f172c11c2775d21fa7e33bff1c8fdaa554d7

Observation 3653660e-4cd6-40b9-8842-ee81e9705f96 · outbound

This paper cites Proximal Policy Optimization Algorithms.

CEPO: RLVR Self-Distillation using Contrastive Evidence Policy Optimization Proximal Policy Optimization Algorithms

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-05-20T07:28:06.712482Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-20T07:25:34.637803Z digest=sha256:abd265670285e0c157376e95645ada957c57790ec36d2cb430eda57f1a1a2360

Observation c424760a-73d0-4362-8a38-f3972fd3a327 · outbound

This paper cites Rewarding Progress: Scaling Automated Process Verifiers for LLM Reasoning.

CEPO: RLVR Self-Distillation using Contrastive Evidence Policy Optimization Rewarding Progress: Scaling Automated Process Verifiers for LLM Reasoning

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-21T01:42:19.263401Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-20T07:25:34.637803Z digest=sha256:a75be2b5fc954f6f37c5ec9ce27dccc45b8f462b85a8c0daa2e090e7f0a26fa0

Observation 74ae5e96-c6ca-4e43-9f32-8705243c0aad · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

CEPO: RLVR Self-Distillation using Contrastive Evidence Policy Optimization DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-05-20T07:28:06.699451Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-20T07:25:34.637803Z digest=sha256:24f09cde22c0753020b1e5686ba05e10c503d649dcd5a5e24aceae179f4c65e9

Observation 49ba4685-b6b3-42f3-8f79-ec2a17ab8e83 · outbound

This paper cites Measuring multimodal mathematical reasoning with math-vision dataset.

CEPO: RLVR Self-Distillation using Contrastive Evidence Policy Optimization Measuring multimodal mathematical reasoning with math-vision dataset

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T07:28:07.792830Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-20T07:25:34.637803Z digest=sha256:3a573bd83d4d41906f284440810d9db7e281f6e03a46c0b54fbfe697f1a985ea

Observation 1dc53f8c-65dc-4127-818e-bb63134e56ae · outbound

This paper cites LogicVista: Multimodal LLM Logical Reasoning Benchmark in Visual Contexts.

CEPO: RLVR Self-Distillation using Contrastive Evidence Policy Optimization LogicVista: Multimodal LLM Logical Reasoning Benchmark in Visual Contexts

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-05-20T07:28:06.738133Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-20T07:25:34.637803Z digest=sha256:838e26094fa5bf4506cfabf87218664fbec97f94e626c51be3a4c77293207a8e

Observation c6c886dc-109c-49d2-ab17-458f03f002c4 · outbound

This paper cites Qwen3 Technical Report.

CEPO: RLVR Self-Distillation using Contrastive Evidence Policy Optimization Qwen3 Technical Report

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-05-20T07:28:06.735030Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-20T07:25:34.637803Z digest=sha256:791b5f5ffd11627318ae97aca70336390c75f94dc4d4f6bdd0ef107f162c81c7

Observation f4444fb1-8423-45b8-915b-697e7349ccef · outbound

This paper cites Self-Distilled RLVR.

CEPO: RLVR Self-Distillation using Contrastive Evidence Policy Optimization Self-Distilled RLVR

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-05-20T07:28:06.709083Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-20T07:25:34.637803Z digest=sha256:2252724c2de3546d5a4c6d68bdfd93975ccae47b7ffe1556f6c3c5663bbf8868

Observation c2f87e32-de1c-4df6-8d41-f21c45c57c44 · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

CEPO: RLVR Self-Distillation using Contrastive Evidence Policy Optimization DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-05-20T07:28:06.718341Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-20T07:25:34.637803Z digest=sha256:0ee02d0ac6df6a3aa04cb18d18d0a4ad2b06facedef21ef3f84b32afd110c2c8

Observation 1b06546e-8fd5-4bf0-889e-d5238403854e · outbound

This paper cites Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi.

CEPO: RLVR Self-Distillation using Contrastive Evidence Policy Optimization Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T07:28:07.795477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-20T07:25:34.637803Z digest=sha256:7ec82cc4935d99c98b6e3be95d4a02340f94f574f6da31086ef84a4af8f0361e

Observation cf422d56-d3d4-4fd6-814f-79948df646ba · outbound

This paper cites From Reasoning to Agentic: Credit Assignment in Reinforcement Learning for Large Language Models.

CEPO: RLVR Self-Distillation using Contrastive Evidence Policy Optimization From Reasoning to Agentic: Credit Assignment in Reinforcement Learning for Large Language Models

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-05-20T07:28:06.747246Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-20T07:25:34.637803Z digest=sha256:5dcc1ad2d7d457a6b0abe0884439df07bc9a1e8d4ab79beb59a5673dc5a757ac

Observation fb4dd984-5cab-4c60-b6d5-7b8ee616348b · outbound

This paper cites Lmms-eval: Reality check on the evaluation of large multimodal models.

CEPO: RLVR Self-Distillation using Contrastive Evidence Policy Optimization Lmms-eval: Reality check on the evaluation of large multimodal models

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T07:28:07.799533Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-20T07:25:34.637803Z digest=sha256:cf06a0063f4881a7eaecaca69c5bad2d0d12dcbc8b4a81fa6b55eed874274326

Observation 19b192e4-371e-4a1b-ae75-3129d962a124 · outbound

This paper cites Self-Distilled Reasoner: On-Policy Self-Distillation for Large Language Models.

CEPO: RLVR Self-Distillation using Contrastive Evidence Policy Optimization Self-Distilled Reasoner: On-Policy Self-Distillation for Large Language Models

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-05-20T07:28:06.706111Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-20T07:25:34.637803Z digest=sha256:e936d534707febf97a20ad6a1d672d19cc8bd54f07647c5ccc09a582a5c3102c

Observation 268fae63-85c0-4afc-8e47-1e13dcb335b2 · outbound

This paper cites PyTorch FSDP: Experiences on Scaling Fully Sharded Data Parallel.

CEPO: RLVR Self-Distillation using Contrastive Evidence Policy Optimization PyTorch FSDP: Experiences on Scaling Fully Sharded Data Parallel

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-05-20T07:28:06.702664Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-20T07:25:34.637803Z digest=sha256:f41cc8cf107ca027deb39a618ff2925ee7cade33e16cf2a976e3550d671e4da8

Observation 8c601782-78d5-49c8-bcf3-7c09989aa339 · outbound

This paper cites EasyR1: An efficient, scalable, multi-modality RL training framework.

CEPO: RLVR Self-Distillation using Contrastive Evidence Policy Optimization EasyR1: An efficient, scalable, multi-modality RL training framework

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-05-20T07:28:07.790952Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-20T07:25:34.637803Z digest=sha256:3a0501a8d8a68772c28234e396015cea456f0576a502d3f5ed52eb02fee8f76c

Observation 94f731f4-8d62-4817-aa77-2f3d071c6fc5 · outbound

This paper cites DynaMath: A Dynamic Visual Benchmark for Evaluating Mathematical Reasoning Robustness of Vision Language Models.

CEPO: RLVR Self-Distillation using Contrastive Evidence Policy Optimization DynaMath: A Dynamic Visual Benchmark for Evaluating Mathematical Reasoning Robustness of Vision Language Models

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-20T07:28:06.725569Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-20T07:25:34.637803Z digest=sha256:9ccfc5134040c2aef42ef2aff748e72ea926c98b99745a1bebea626372300461

Pith citing papers

Observation e331cbeb-d053-4a43-b835-8a72ecaea650 · inbound

H$^2$SD: Hybrid Hindsight Self-Distillation cites this paper.

H$^2$SD: Hybrid Hindsight Self-Distillation CEPO: RLVR Self-Distillation using Contrastive Evidence Policy Optimization

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-01T13:52:48.791895Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T13:52:48.791895Z digest=sha256:ff790f888f2cbaaf86ac5e0661e9638b4f04409170e51713045103d12a8a35ed

Observation cc1008c9-cd83-4da0-8597-d3369fc394c9 · inbound

Perception Before Supervision: Self-Contained Visual Distillation from Counterfactual Blind Spots cites this paper.

Perception Before Supervision: Self-Contained Visual Distillation from Counterfactual Blind Spots CEPO: RLVR Self-Distillation using Contrastive Evidence Policy Optimization

Reference 8

Resolution
metadata mismatch
local_arxiv, observed 2026-08-11T04:17:44.764322Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T04:17:44.132452Z digest=sha256:aaa27a8fd4d0e86551892f69c1215e011b1725529409f12efd1315842c1360d4