Pith. sign in

Paper Citation Record · LEDGER

Step-level Value Preference Optimization for Mathematical Reasoning

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 15 inbound Pith citation observations for arXiv:2406.10858.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2406.10858 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 15 of 15 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 15 of 15 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-09T06:04:01.595054Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-20T09:42:04.234036Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation ded0495b-0a64-467b-9aa0-caf6eaca9810 · inbound

Agent Q: Advanced Reasoning and Learning for Autonomous AI Agents cites this paper.

Agent Q: Advanced Reasoning and Learning for Autonomous AI Agents Step-level Value Preference Optimization for Mathematical Reasoning

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-20T09:42:04.237687Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-20T09:41:59.979595Z digest=sha256:4bc75aa92233d84ea75cd40541d526a6bd1bdc0d5e84a90f23a9ee4b9505021d

Observation 547c4f5f-6ffd-45f1-9c91-fac1ba3bd1f3 · inbound

Do NOT Think That Much for 2+3=? On the Overthinking of o1-Like LLMs cites this paper.

Do NOT Think That Much for 2+3=? On the Overthinking of o1-Like LLMs Step-level Value Preference Optimization for Mathematical Reasoning

Reference 88

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T15:51:29.518054Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-13T15:51:29.022336Z digest=sha256:10e48dc8307fffa4b5fc53de162b498fed1cc5a16847490af3444a723e36b5b5

Observation 6f125f57-580a-479a-98b2-a7d853c4db5d · inbound

Towards Large Reasoning Models: A Survey of Reinforced Reasoning with Large Language Models cites this paper.

Towards Large Reasoning Models: A Survey of Reinforced Reasoning with Large Language Models Step-level Value Preference Optimization for Mathematical Reasoning

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-15T21:20:59.490615Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-15T21:20:59.128986Z digest=sha256:4740244de9efddd7d0d06a8927660de7b28195424e0c806b3af542e0e6edf67f

Observation 472214e9-0d6c-4a4b-9b67-3c20c0f517b1 · inbound

Reveal the Mystery of DPO: The Connection between DPO and RL Algorithms cites this paper.

Reveal the Mystery of DPO: The Connection between DPO and RL Algorithms Step-level Value Preference Optimization for Mathematical Reasoning

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-09T06:04:01.595054Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T06:04:01.595054Z digest=sha256:0d96ed1aff9f3eb78d237f2ea2db735e5a8bec8ab02b5fe0d7f41877f0ca541b

Observation b40e4971-2e58-423d-a19f-11be2b6f416e · inbound

R1-VL: Learning to Reason with Multimodal Large Language Models via Step-wise Group Relative Policy Optimization cites this paper.

R1-VL: Learning to Reason with Multimodal Large Language Models via Step-wise Group Relative Policy Optimization Step-level Value Preference Optimization for Mathematical Reasoning

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-16T15:04:22.805429Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-16T15:04:22.690503Z digest=sha256:25d44610743e54c54f9daf0de4cb9c467a70985a3b2d4564d244618149b62f85

Observation a4d0c8dd-2a5e-4fbf-ab85-a3e4745f2f97 · inbound

Qwen Look Again: Guiding Vision-Language Reasoning Models to Re-attention Visual Information cites this paper.

Qwen Look Again: Guiding Vision-Language Reasoning Models to Re-attention Visual Information Step-level Value Preference Optimization for Mathematical Reasoning

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T12:46:56.966984Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:46:56.966984Z digest=sha256:98f14408cf47006a8c71c4a4adc6b568838ea30ec1b05cbba2034b0592f62d99

Observation 5d3c20b9-ec80-4818-b818-4671909126f3 · inbound

AI Agent Behavioral Science cites this paper.

AI Agent Behavioral Science Step-level Value Preference Optimization for Mathematical Reasoning

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T11:00:53.398869Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:00:53.398869Z digest=sha256:c7cdba589cacb30325f67e0f5d34b52dd2bd1a586123c50c67a654e34a8c7e64

Observation 6749814d-8a1e-4756-9cbc-f9935e3b72cf · inbound

CheMatAgent: Enhancing LLMs for Chemistry and Materials Science through Tree-Search Based Tool Learning cites this paper.

CheMatAgent: Enhancing LLMs for Chemistry and Materials Science through Tree-Search Based Tool Learning Step-level Value Preference Optimization for Mathematical Reasoning

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T05:36:38.444011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:36:38.444011Z digest=sha256:2fd5537830602e7cb68d9878a4ad395f08151840e18fbae735df3ad943bf2b41

Observation 81238a6e-2b0e-48b9-9012-9cd10c027c1c · inbound

DuaShepherd: Integrating Stepwise Correctness and Potential Rewards for Mathematical Reasoning cites this paper.

DuaShepherd: Integrating Stepwise Correctness and Potential Rewards for Mathematical Reasoning Step-level Value Preference Optimization for Mathematical Reasoning

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T23:35:35.544750Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:35:35.544750Z digest=sha256:fce8c119bd5b4755023de757af4ced0a3a9a5869d3e04501617ab752156eb217

Observation 183659f9-af00-406b-ad67-4fea6c4eb65c · inbound

SpikingMamba: Towards Energy-Efficient Large Language Models via Knowledge Distillation from Mamba cites this paper.

SpikingMamba: Towards Energy-Efficient Large Language Models via Knowledge Distillation from Mamba Step-level Value Preference Optimization for Mathematical Reasoning

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-18T09:46:12.439895Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-18T09:44:53.290259Z digest=sha256:42b395d0ca8dbf24655606794ee15d551be1de610dd8abbf0fc8aeeffb95896c

Observation 54ce838a-b3af-4aa0-ae8c-c181090fcbff · inbound

Unlocking Exploration in RLVR: Uncertainty-aware Advantage Shaping for Deeper Reasoning cites this paper.

Unlocking Exploration in RLVR: Uncertainty-aware Advantage Shaping for Deeper Reasoning Step-level Value Preference Optimization for Mathematical Reasoning

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-18T07:21:04.436629Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-18T07:20:01.505216Z digest=sha256:11968f76339b578bdbe6e4846b46beab1765168746d518a2887bcdbb02b4e89b

Observation eec809d2-b0ab-4299-903e-b35cfb1d93e6 · inbound

APCD: Adaptive Path-Contrastive Decoding for Reliable Large Language Model Generation cites this paper.

APCD: Adaptive Path-Contrastive Decoding for Reliable Large Language Model Generation Step-level Value Preference Optimization for Mathematical Reasoning

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-05-12T05:26:25.191736Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-12T05:22:25.475956Z digest=sha256:049b2107bef9806fc0c7e7c137990ea59531a81f8be08172d545636aca3d9158

Observation 3dd2b738-438b-4692-8fc8-673e0334ace8 · inbound

Towards Order Fairness: Mitigating LLMs Order Sensitivity through Dual Group Advantage Optimization cites this paper.

Towards Order Fairness: Mitigating LLMs Order Sensitivity through Dual Group Advantage Optimization Step-level Value Preference Optimization for Mathematical Reasoning

Reference 44

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T07:37:29.952562Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-13T07:32:58.404947Z digest=sha256:5621aa3ba37f0cbddeca99201279a3eb8684f5334fc5d920cba096ef2ca45a3a

Observation 7080c4b6-4ef2-4eaf-87ae-ec831fef7d5b · inbound

Multi-Turn On-Policy Distillation with Prefix Replay cites this paper.

Multi-Turn On-Policy Distillation with Prefix Replay Step-level Value Preference Optimization for Mathematical Reasoning

Reference 75

Resolution
unresolved
no resolver link, observed 2026-07-11T13:53:36.775836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T13:53:36.775836Z digest=sha256:569657b76718c02cc929bce7848613c59d82cf9b0198c4bebda531246501e34a

Observation a9acd4b2-4ebd-42bf-90a9-62837c84ed50 · inbound

Multi-Turn On-Policy Distillation with Prefix Replay cites this paper.

Multi-Turn On-Policy Distillation with Prefix Replay Step-level Value Preference Optimization for Mathematical Reasoning

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-02T08:40:39.974766Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:40:39.974766Z digest=sha256:3a14e8e1eb5a95e1b0c5e37f2f2528aa8a7b8f454d74ebd0a85bfcfcbaf924ca