Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 15 inbound Pith citation observations for arXiv:2406.10858.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-09T06:04:01.595054Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-20T09:42:04.234036Z
0 of 0 outbound references displayed
External citation measurements
No source-named external measurement is stored.
No outbound reference observations are available for this paper version.
Observation ded0495b-0a64-467b-9aa0-caf6eaca9810 · inbound
Agent Q: Advanced Reasoning and Learning for Autonomous AI Agents Step-level Value Preference Optimization for Mathematical Reasoning
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 547c4f5f-6ffd-45f1-9c91-fac1ba3bd1f3 · inbound
Do NOT Think That Much for 2+3=? On the Overthinking of o1-Like LLMs Step-level Value Preference Optimization for Mathematical Reasoning
Reference 88
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 6f125f57-580a-479a-98b2-a7d853c4db5d · inbound
Towards Large Reasoning Models: A Survey of Reinforced Reasoning with Large Language Models Step-level Value Preference Optimization for Mathematical Reasoning
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 472214e9-0d6c-4a4b-9b67-3c20c0f517b1 · inbound
Reveal the Mystery of DPO: The Connection between DPO and RL Algorithms Step-level Value Preference Optimization for Mathematical Reasoning
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b40e4971-2e58-423d-a19f-11be2b6f416e · inbound
R1-VL: Learning to Reason with Multimodal Large Language Models via Step-wise Group Relative Policy Optimization Step-level Value Preference Optimization for Mathematical Reasoning
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation a4d0c8dd-2a5e-4fbf-ab85-a3e4745f2f97 · inbound
Qwen Look Again: Guiding Vision-Language Reasoning Models to Re-attention Visual Information Step-level Value Preference Optimization for Mathematical Reasoning
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5d3c20b9-ec80-4818-b818-4671909126f3 · inbound
AI Agent Behavioral Science Step-level Value Preference Optimization for Mathematical Reasoning
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6749814d-8a1e-4756-9cbc-f9935e3b72cf · inbound
CheMatAgent: Enhancing LLMs for Chemistry and Materials Science through Tree-Search Based Tool Learning Step-level Value Preference Optimization for Mathematical Reasoning
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 81238a6e-2b0e-48b9-9012-9cd10c027c1c · inbound
DuaShepherd: Integrating Stepwise Correctness and Potential Rewards for Mathematical Reasoning Step-level Value Preference Optimization for Mathematical Reasoning
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 183659f9-af00-406b-ad67-4fea6c4eb65c · inbound
SpikingMamba: Towards Energy-Efficient Large Language Models via Knowledge Distillation from Mamba Step-level Value Preference Optimization for Mathematical Reasoning
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 54ce838a-b3af-4aa0-ae8c-c181090fcbff · inbound
Unlocking Exploration in RLVR: Uncertainty-aware Advantage Shaping for Deeper Reasoning Step-level Value Preference Optimization for Mathematical Reasoning
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation eec809d2-b0ab-4299-903e-b35cfb1d93e6 · inbound
APCD: Adaptive Path-Contrastive Decoding for Reliable Large Language Model Generation Step-level Value Preference Optimization for Mathematical Reasoning
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 3dd2b738-438b-4692-8fc8-673e0334ace8 · inbound
Towards Order Fairness: Mitigating LLMs Order Sensitivity through Dual Group Advantage Optimization Step-level Value Preference Optimization for Mathematical Reasoning
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.
Observation 7080c4b6-4ef2-4eaf-87ae-ec831fef7d5b · inbound
Multi-Turn On-Policy Distillation with Prefix Replay Step-level Value Preference Optimization for Mathematical Reasoning
Reference 75
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a9acd4b2-4ebd-42bf-90a9-62837c84ed50 · inbound
Multi-Turn On-Policy Distillation with Prefix Replay Step-level Value Preference Optimization for Mathematical Reasoning
Reference 76
Source-reported events for the cited work
Unavailable: canonical work link unavailable.