Pith. sign in

Paper Citation Record · LEDGER

Step-level Value Preference Optimization for Mathematical Reasoning

As of 16 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 19 inbound Pith citation observations for arXiv:2406.10858.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2406.10858 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 19 of 19 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 19 of 19 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T11:40:27.581679Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-20T09:42:04.234036Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation ded0495b-0a64-467b-9aa0-caf6eaca9810 · inbound

Agent Q: Advanced Reasoning and Learning for Autonomous AI Agents cites this paper.

Agent Q: Advanced Reasoning and Learning for Autonomous AI Agents Step-level Value Preference Optimization for Mathematical Reasoning

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-20T09:42:04.237687Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-05-20T09:41:59.979595Z digest=sha256:9768ce1c91a98efeda54d187d418945f64924dfca7e028ce002044aad5f6e67b

Observation 29b33817-5798-4864-aa5d-e1d24d5fec9b · inbound

Progressive Multimodal Reasoning via Active Retrieval cites this paper.

Progressive Multimodal Reasoning via Active Retrieval Step-level Value Preference Optimization for Mathematical Reasoning

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T11:55:10.172389Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:55:10.172389Z digest=sha256:dd1c8f0d567b8da7e9bf29c4cb7ab918f6b76a6fcb5eb7d6837291d332fd4e92

Observation 60a9534b-3660-473d-bc99-93e66693ff89 · inbound

Improving Multi-Step Reasoning Abilities of Large Language Models with Direct Advantage Policy Optimization cites this paper.

Improving Multi-Step Reasoning Abilities of Large Language Models with Direct Advantage Policy Optimization Step-level Value Preference Optimization for Mathematical Reasoning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T04:56:22.377628Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:56:22.377628Z digest=sha256:0397b5e10e308b32cd1ae59012734d13c3500d1b42c069c7a05ce6acbd2fe7d3

Observation 547c4f5f-6ffd-45f1-9c91-fac1ba3bd1f3 · inbound

Do NOT Think That Much for 2+3=? On the Overthinking of o1-Like LLMs cites this paper.

Do NOT Think That Much for 2+3=? On the Overthinking of o1-Like LLMs Step-level Value Preference Optimization for Mathematical Reasoning

Reference 88

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T15:51:29.518054Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-05-13T15:51:29.022336Z digest=sha256:4562771afe76a4e45e6f3788560178b9cbd025eb1ed4f67ada6a741b1789ef41

Observation e7183681-11f1-4634-a67a-e1a621731655 · inbound

BoostStep: Boosting mathematical capability of Large Language Models via improved single-step reasoning cites this paper.

BoostStep: Boosting mathematical capability of Large Language Models via improved single-step reasoning Step-level Value Preference Optimization for Mathematical Reasoning

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-10T21:57:50.695810Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T21:57:50.695810Z digest=sha256:d8f35fe5e3931aa90a5941cc1541089c8bc397773732cc78d72ddb961a407e95

Observation 6f125f57-580a-479a-98b2-a7d853c4db5d · inbound

Towards Large Reasoning Models: A Survey of Reinforced Reasoning with Large Language Models cites this paper.

Towards Large Reasoning Models: A Survey of Reinforced Reasoning with Large Language Models Step-level Value Preference Optimization for Mathematical Reasoning

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-15T21:20:59.490615Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-15T21:20:59.128986Z digest=sha256:038d9167067cfd73694b2fbbaa4c25c1b19e04cbc6946a7e559c113df9446429

Observation 472214e9-0d6c-4a4b-9b67-3c20c0f517b1 · inbound

Reveal the Mystery of DPO: The Connection between DPO and RL Algorithms cites this paper.

Reveal the Mystery of DPO: The Connection between DPO and RL Algorithms Step-level Value Preference Optimization for Mathematical Reasoning

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-09T06:04:01.595054Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T06:04:01.595054Z digest=sha256:75906735440d963541946ff8ce94a55f7016d010fdde257c74adc0ff0feb11cd

Observation b40e4971-2e58-423d-a19f-11be2b6f416e · inbound

R1-VL: Learning to Reason with Multimodal Large Language Models via Step-wise Group Relative Policy Optimization cites this paper.

R1-VL: Learning to Reason with Multimodal Large Language Models via Step-wise Group Relative Policy Optimization Step-level Value Preference Optimization for Mathematical Reasoning

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-16T15:04:22.805429Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-16T15:04:22.690503Z digest=sha256:00789ddeb4a0d4114947ab5458018e7c1270f2d712ee3885fae7e63a3e4d4d23

Observation f0f1e1f3-92a8-43a5-b4ed-aa5dc82d590b · inbound

Efficient Pretraining Length Scaling cites this paper.

Efficient Pretraining Length Scaling Step-level Value Preference Optimization for Mathematical Reasoning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-16T11:40:27.581679Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:40:27.581679Z digest=sha256:c3d45c5d155a24cb4ceb25cb763e0d5726d4e9336fe9c8cc93c2c6f4e0a80f9c

Observation a4d0c8dd-2a5e-4fbf-ab85-a3e4745f2f97 · inbound

Qwen Look Again: Guiding Vision-Language Reasoning Models to Re-attention Visual Information cites this paper.

Qwen Look Again: Guiding Vision-Language Reasoning Models to Re-attention Visual Information Step-level Value Preference Optimization for Mathematical Reasoning

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T12:46:56.966984Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:46:56.966984Z digest=sha256:b4f09b12f238ce70dcc9dbe20954f4aa3e8dfa649f2425ff152cb4662ed0c0f2

Observation 5d3c20b9-ec80-4818-b818-4671909126f3 · inbound

AI Agent Behavioral Science cites this paper.

AI Agent Behavioral Science Step-level Value Preference Optimization for Mathematical Reasoning

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T11:00:53.398869Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:00:53.398869Z digest=sha256:2095056d4d29784eb663879de1adb0507be34ec9188599a82529af5c0d5682fb

Observation 6749814d-8a1e-4756-9cbc-f9935e3b72cf · inbound

CheMatAgent: Enhancing LLMs for Chemistry and Materials Science through Tree-Search Based Tool Learning cites this paper.

CheMatAgent: Enhancing LLMs for Chemistry and Materials Science through Tree-Search Based Tool Learning Step-level Value Preference Optimization for Mathematical Reasoning

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T05:36:38.444011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:36:38.444011Z digest=sha256:3ae09e6a721bbaf04f6b4cea55e8c6daf951c6cf8794c407429204ad4f8e087f

Observation 81238a6e-2b0e-48b9-9012-9cd10c027c1c · inbound

DuaShepherd: Integrating Stepwise Correctness and Potential Rewards for Mathematical Reasoning cites this paper.

DuaShepherd: Integrating Stepwise Correctness and Potential Rewards for Mathematical Reasoning Step-level Value Preference Optimization for Mathematical Reasoning

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T23:35:35.544750Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:35:35.544750Z digest=sha256:ccbb18cc05a27abaf6eb4076b69c1b925a41b8d8fbb8aa261c860baa7d3c8139

Observation 183659f9-af00-406b-ad67-4fea6c4eb65c · inbound

SpikingMamba: Towards Energy-Efficient Large Language Models via Knowledge Distillation from Mamba cites this paper.

SpikingMamba: Towards Energy-Efficient Large Language Models via Knowledge Distillation from Mamba Step-level Value Preference Optimization for Mathematical Reasoning

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-18T09:46:12.439895Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-18T09:44:53.290259Z digest=sha256:33d7b31044dea8c6a275cd67973428bfabdbcc428049939414d4661575e2ce92

Observation 54ce838a-b3af-4aa0-ae8c-c181090fcbff · inbound

Unlocking Exploration in RLVR: Uncertainty-aware Advantage Shaping for Deeper Reasoning cites this paper.

Unlocking Exploration in RLVR: Uncertainty-aware Advantage Shaping for Deeper Reasoning Step-level Value Preference Optimization for Mathematical Reasoning

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-18T07:21:04.436629Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-18T07:20:01.505216Z digest=sha256:338d804a08e5f8d13607179fe2a3934c17800f3d93bdd6173ac23bf25b996f1d

Observation eec809d2-b0ab-4299-903e-b35cfb1d93e6 · inbound

APCD: Adaptive Path-Contrastive Decoding for Reliable Large Language Model Generation cites this paper.

APCD: Adaptive Path-Contrastive Decoding for Reliable Large Language Model Generation Step-level Value Preference Optimization for Mathematical Reasoning

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-05-12T05:26:25.191736Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-05-12T05:22:25.475956Z digest=sha256:50c4da45faf5351ade8060bc53e401125de34921a768aca8425c327b439d45e3

Observation 3dd2b738-438b-4692-8fc8-673e0334ace8 · inbound

Towards Order Fairness: Mitigating LLMs Order Sensitivity through Dual Group Advantage Optimization cites this paper.

Towards Order Fairness: Mitigating LLMs Order Sensitivity through Dual Group Advantage Optimization Step-level Value Preference Optimization for Mathematical Reasoning

Reference 44

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T07:37:29.952562Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-05-13T07:32:58.404947Z digest=sha256:8b0ecff8cadd8f34c9a82d6cbe532a4483156f4b9178615f82ce0d3186347893

Observation 7080c4b6-4ef2-4eaf-87ae-ec831fef7d5b · inbound

Multi-Turn On-Policy Distillation with Prefix Replay cites this paper.

Multi-Turn On-Policy Distillation with Prefix Replay Step-level Value Preference Optimization for Mathematical Reasoning

Reference 75

Resolution
unresolved
no resolver link, observed 2026-07-11T13:53:36.775836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T13:53:36.775836Z digest=sha256:b59a2b9f2e5c336ab7750b52cc1536b1b4a1f800ba99e49b987402df9a4b2e9d

Observation a9acd4b2-4ebd-42bf-90a9-62837c84ed50 · inbound

Multi-Turn On-Policy Distillation with Prefix Replay cites this paper.

Multi-Turn On-Policy Distillation with Prefix Replay Step-level Value Preference Optimization for Mathematical Reasoning

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-02T08:40:39.974766Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:40:39.974766Z digest=sha256:6a044b5bdba72212952b0e87a0fcef596d9d136055a6fa09131ed23945383dd1