Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-04T13:44:42.889956Z
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 31 of 31 outbound references and 0 inbound Pith citation observations for arXiv:2608.02149.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-04T13:44:42.889956Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
31 of 31 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 0d5209f4-e2fb-458e-992f-b0a770131b5b · outbound
Beyond the Mean: Multi-Moment Policy Optimization for LLM Reasoning DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6d2b32d9-87b1-4bba-84ed-fc3f4ef25269 · outbound
Beyond the Mean: Multi-Moment Policy Optimization for LLM Reasoning Advances in Neural Information Processing Systems , volume=
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 53186b75-65a6-4686-bc35-4da457bbb35e · outbound
Beyond the Mean: Multi-Moment Policy Optimization for LLM Reasoning Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 90afc498-6351-41c4-9777-6ca0d0fdedbd · outbound
Beyond the Mean: Multi-Moment Policy Optimization for LLM Reasoning Advances in Neural Information Processing Systems , volume=
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 323d704e-600f-4940-865d-198f5a908e14 · outbound
Beyond the Mean: Multi-Moment Policy Optimization for LLM Reasoning Beyond the Sampled Token: Preserving Candidate Support in RLVR
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 729733ae-9987-4e8c-aca0-688845e6a566 · outbound
Beyond the Mean: Multi-Moment Policy Optimization for LLM Reasoning Machine learning , volume=
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ce21e18e-e14c-41af-a3f4-b49e12f096d5 · outbound
Beyond the Mean: Multi-Moment Policy Optimization for LLM Reasoning arXiv preprint arXiv:2602.02710 , year=
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8ab93098-64be-4742-aa57-a9f4fd05c4fd · outbound
Beyond the Mean: Multi-Moment Policy Optimization for LLM Reasoning Proximal Policy Optimization Algorithms
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1dbc430a-0bee-4cfc-8b31-1a00ca4d3f7d · outbound
Beyond the Mean: Multi-Moment Policy Optimization for LLM Reasoning Transactions of the American Mathematical Society , volume=
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a75913b7-3fc9-4ce1-ba32-f567b45c445b · outbound
Beyond the Mean: Multi-Moment Policy Optimization for LLM Reasoning Statistics & probability letters , volume=
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 867b1ec8-7068-4172-a7ee-f67adb1efc2b · outbound
Beyond the Mean: Multi-Moment Policy Optimization for LLM Reasoning Canadian Journal of Mathematics , volume=
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9f0d7f77-bff0-4442-9fec-17be97ef1551 · outbound
Beyond the Mean: Multi-Moment Policy Optimization for LLM Reasoning arXiv preprint arXiv:2601.18779 , year=
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 451ef461-c10a-47a3-8426-8130383d516c · outbound
Beyond the Mean: Multi-Moment Policy Optimization for LLM Reasoning arXiv preprint arXiv:2602.21189 , year=
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 22a0bc3d-0cdd-49bb-a5aa-83258e00ca71 · outbound
Beyond the Mean: Multi-Moment Policy Optimization for LLM Reasoning Qwen3 Technical Report
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 511f50f6-471c-4206-bb31-8a581321acdf · outbound
Beyond the Mean: Multi-Moment Policy Optimization for LLM Reasoning Proceedings of the Twentieth European Conference on Computer Systems , pages=
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ed68086f-f966-48d3-8a3f-f15ffe3fafd5 · outbound
Beyond the Mean: Multi-Moment Policy Optimization for LLM Reasoning Measuring Mathematical Problem Solving With the MATH Dataset
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c376b6d3-0b92-41b0-a44d-bc5e95ab786f · outbound
Beyond the Mean: Multi-Moment Policy Optimization for LLM Reasoning 2025 , howpublished =
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 50d1c619-e285-4906-a690-0e82a7612cc1 · outbound
Beyond the Mean: Multi-Moment Policy Optimization for LLM Reasoning Beyond Mode Collapse: Distribution Matching for Diverse Reasoning
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e18ef7a0-5499-4bff-8516-6b38465b8313 · outbound
Beyond the Mean: Multi-Moment Policy Optimization for LLM Reasoning Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eb3425a5-3729-4f41-a508-4b9dfa134d56 · outbound
Beyond the Mean: Multi-Moment Policy Optimization for LLM Reasoning Beyond Benchmarks: MathArena as an Evaluation Platform for Mathematics with LLMs
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dd81640a-f32f-4349-8699-8fd4c34fdef8 · outbound
Beyond the Mean: Multi-Moment Policy Optimization for LLM Reasoning Tulu 3: Pushing Frontiers in Open Language Model Post-Training
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 76bb9b35-0e89-426d-aa2d-8445e1964eca · outbound
Beyond the Mean: Multi-Moment Policy Optimization for LLM Reasoning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 601c01fc-f827-46be-80ab-f4ef28a8bd95 · outbound
Beyond the Mean: Multi-Moment Policy Optimization for LLM Reasoning Information Gain-based Rollout Policy Optimization: An Adaptive Tree-Structured Rollout Approach for Multi-Turn LLM Agents
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c5e6ef5a-aab2-45de-8d24-68d3f06d044f · outbound
Beyond the Mean: Multi-Moment Policy Optimization for LLM Reasoning arXiv preprint arXiv:2509.25133 , year=
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1616428c-6970-468b-a01c-d481697f6cfd · outbound
Beyond the Mean: Multi-Moment Policy Optimization for LLM Reasoning Findings of the Association for Computational Linguistics: ACL 2026 , pages=
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 33f4a1b7-835d-41e6-a3a1-1ad31fd9f8fa · outbound
Beyond the Mean: Multi-Moment Policy Optimization for LLM Reasoning arXiv preprint arXiv:2509.15207 , year=
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ab809372-8108-4b96-a612-56dd41d4016c · outbound
Beyond the Mean: Multi-Moment Policy Optimization for LLM Reasoning arXiv preprint arXiv:2509.26209 , year=
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f465618d-2642-4c03-99f4-85b6e8c5f630 · outbound
Beyond the Mean: Multi-Moment Policy Optimization for LLM Reasoning TreePO: Bridging the Gap of Policy Optimization and Efficacy and Inference Efficiency with Heuristic Tree-based Modeling
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 74c03c94-b637-4681-8f32-77dc7cceaec1 · outbound
Beyond the Mean: Multi-Moment Policy Optimization for LLM Reasoning Springer Series in Statistics ( , year=
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d97f1f9c-884e-43d8-b2bb-2ee87fa170d5 · outbound
Beyond the Mean: Multi-Moment Policy Optimization for LLM Reasoning International conference on machine learning , pages=
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d3f7cfcf-88a9-4fe6-9cbb-5f4d42b469ec · outbound
Beyond the Mean: Multi-Moment Policy Optimization for LLM Reasoning High-Dimensional Continuous Control Using Generalized Advantage Estimation
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.