Pith. sign in

Paper Citation Record · LEDGER

Beyond the Mean: Multi-Moment Policy Optimization for LLM Reasoning

As of 8 August 2026, this Paper Citation Record lists 31 of 31 outbound references and 0 inbound Pith citation observations for arXiv:2608.02149.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.02149 v1

Coverage vector

measured 31 of 31 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-04T13:44:42.889956Z

measured 31 of 31 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

31 of 31 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved31
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 0d5209f4-e2fb-458e-992f-b0a770131b5b · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Beyond the Mean: Multi-Moment Policy Optimization for LLM Reasoning DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-04T13:44:40.090481Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T13:44:40.090481Z digest=sha256:42b64c6fdb0e2696fd38d37b8c57835b48b3d4a511464b31d23dd02b0d27c5d0

Observation 6d2b32d9-87b1-4bba-84ed-fc3f4ef25269 · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

Beyond the Mean: Multi-Moment Policy Optimization for LLM Reasoning Advances in Neural Information Processing Systems , volume=

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-04T13:44:40.125359Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T13:44:40.125359Z digest=sha256:e0b18bcb9f7ee335d407c59d3309d5d38755a3f838545da7f1cb8975eff64107

Observation 53186b75-65a6-4686-bc35-4da457bbb35e · outbound

This paper cites Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models.

Beyond the Mean: Multi-Moment Policy Optimization for LLM Reasoning Pass@k Training for Adaptively Balancing Exploration and Exploitation of Large Reasoning Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-04T13:44:40.174330Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T13:44:40.174330Z digest=sha256:a7bb0fb3767b60e66ab1bd59eb809f4f8206b48e643b75b7a0d7078784c1911a

Observation 90afc498-6351-41c4-9777-6ca0d0fdedbd · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

Beyond the Mean: Multi-Moment Policy Optimization for LLM Reasoning Advances in Neural Information Processing Systems , volume=

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-04T13:44:40.218780Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T13:44:40.218780Z digest=sha256:43c59f8abccddeb5c80f8d7778622ecd1e48f4f760c9fb5464307376bf5c562b

Observation 323d704e-600f-4940-865d-198f5a908e14 · outbound

This paper cites Beyond the Sampled Token: Preserving Candidate Support in RLVR.

Beyond the Mean: Multi-Moment Policy Optimization for LLM Reasoning Beyond the Sampled Token: Preserving Candidate Support in RLVR

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-04T13:44:40.295338Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T13:44:40.295338Z digest=sha256:53b045d4e50e1284b171799a1d2e41488fd9a52c6d3e7726fef185702cc57185

Observation 729733ae-9987-4e8c-aca0-688845e6a566 · outbound

This paper cites Machine learning , volume=.

Beyond the Mean: Multi-Moment Policy Optimization for LLM Reasoning Machine learning , volume=

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-04T13:44:40.372508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T13:44:40.372508Z digest=sha256:c64ea636c757ee9ad348d3348e79088dec61600cb5a11f0325349fa766bf2a28

Observation ce21e18e-e14c-41af-a3f4-b49e12f096d5 · outbound

This paper cites arXiv preprint arXiv:2602.02710 , year=.

Beyond the Mean: Multi-Moment Policy Optimization for LLM Reasoning arXiv preprint arXiv:2602.02710 , year=

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-04T13:44:40.376806Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T13:44:40.376806Z digest=sha256:a1159b459afb1fdbedf0c270aa706753ab8e96a3f6003e9378627da2e055842a

Observation 8ab93098-64be-4742-aa57-a9f4fd05c4fd · outbound

This paper cites Proximal Policy Optimization Algorithms.

Beyond the Mean: Multi-Moment Policy Optimization for LLM Reasoning Proximal Policy Optimization Algorithms

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-04T13:44:40.380372Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T13:44:40.380372Z digest=sha256:7799042ffd3d1bc5f0c1347467f17be43f1eb4344d7db4033787e7784416f669

Observation 1dbc430a-0bee-4cfc-8b31-1a00ca4d3f7d · outbound

This paper cites Transactions of the American Mathematical Society , volume=.

Beyond the Mean: Multi-Moment Policy Optimization for LLM Reasoning Transactions of the American Mathematical Society , volume=

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-04T13:44:40.424271Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T13:44:40.424271Z digest=sha256:f6d5a7b16835c3e624362d76c78a96050ff4546387e2b0620c7b8f632e933d2c

Observation a75913b7-3fc9-4ce1-ba32-f567b45c445b · outbound

This paper cites Statistics & probability letters , volume=.

Beyond the Mean: Multi-Moment Policy Optimization for LLM Reasoning Statistics & probability letters , volume=

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-04T13:44:40.546665Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T13:44:40.546665Z digest=sha256:e6c6d01d8a5739154c697cf237b89ce124cba65bed9fa75632b0e0028fed45de

Observation 867b1ec8-7068-4172-a7ee-f67adb1efc2b · outbound

This paper cites Canadian Journal of Mathematics , volume=.

Beyond the Mean: Multi-Moment Policy Optimization for LLM Reasoning Canadian Journal of Mathematics , volume=

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-04T13:44:40.655152Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T13:44:40.655152Z digest=sha256:ef9d5d2f4b4403ca6c272963cd24bd10510d038e9948500e2884540f0df108e8

Observation 9f0d7f77-bff0-4442-9fec-17be97ef1551 · outbound

This paper cites arXiv preprint arXiv:2601.18779 , year=.

Beyond the Mean: Multi-Moment Policy Optimization for LLM Reasoning arXiv preprint arXiv:2601.18779 , year=

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-04T13:44:40.813207Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T13:44:40.813207Z digest=sha256:a2ee7c57887309172708f079fa0a11c0374037b089ec1b7ca85dd71d26de3f95

Observation 451ef461-c10a-47a3-8426-8130383d516c · outbound

This paper cites arXiv preprint arXiv:2602.21189 , year=.

Beyond the Mean: Multi-Moment Policy Optimization for LLM Reasoning arXiv preprint arXiv:2602.21189 , year=

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-04T13:44:40.932110Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T13:44:40.932110Z digest=sha256:e1abc5ce600aba48feb65bef78a1494b45d7636a4a6e0d386ee1256d44d36033

Observation 22a0bc3d-0cdd-49bb-a5aa-83258e00ca71 · outbound

This paper cites Qwen3 Technical Report.

Beyond the Mean: Multi-Moment Policy Optimization for LLM Reasoning Qwen3 Technical Report

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-04T13:44:41.090156Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T13:44:41.090156Z digest=sha256:7daa32aecbfe2286683665e85430d48ac15c1ca986151146856947473500b04f

Observation 511f50f6-471c-4206-bb31-8a581321acdf · outbound

This paper cites Proceedings of the Twentieth European Conference on Computer Systems , pages=.

Beyond the Mean: Multi-Moment Policy Optimization for LLM Reasoning Proceedings of the Twentieth European Conference on Computer Systems , pages=

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-04T13:44:41.231524Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T13:44:41.231524Z digest=sha256:9a97d6772b0c32d21b1e23360e6545439aa7d587c003bce9afe1c0fb6a49f1a7

Observation ed68086f-f966-48d3-8a3f-f15ffe3fafd5 · outbound

This paper cites Measuring Mathematical Problem Solving With the MATH Dataset.

Beyond the Mean: Multi-Moment Policy Optimization for LLM Reasoning Measuring Mathematical Problem Solving With the MATH Dataset

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-04T13:44:41.310367Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T13:44:41.310367Z digest=sha256:d2ccee0f0426437ecf53327cb2c6ce0406eeaf497e25e10f1b42815610aba670

Observation c376b6d3-0b92-41b0-a44d-bc5e95ab786f · outbound

This paper cites 2025 , howpublished =.

Beyond the Mean: Multi-Moment Policy Optimization for LLM Reasoning 2025 , howpublished =

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-04T13:44:41.515791Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T13:44:41.515791Z digest=sha256:3f4f981b14c6cf7a6f3cb4b2f4f7a7a6113d5be18cce42bf9cd7f25742900810

Observation 50d1c619-e285-4906-a690-0e82a7612cc1 · outbound

This paper cites Beyond Mode Collapse: Distribution Matching for Diverse Reasoning.

Beyond the Mean: Multi-Moment Policy Optimization for LLM Reasoning Beyond Mode Collapse: Distribution Matching for Diverse Reasoning

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-04T13:44:41.612646Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T13:44:41.612646Z digest=sha256:0f558c04799bd3520eccd355cf203ec4c2b5d4b10e0331aea7e33778f54c8b8f

Observation e18ef7a0-5499-4bff-8516-6b38465b8313 · outbound

This paper cites Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=.

Beyond the Mean: Multi-Moment Policy Optimization for LLM Reasoning Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) , pages=

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-04T13:44:41.707471Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T13:44:41.707471Z digest=sha256:7c8abffe5fe8d17db8d492c85c04fc0bcbd9d3df9daf2ff3dd0fe81d69121405

Observation eb3425a5-3729-4f41-a508-4b9dfa134d56 · outbound

This paper cites Beyond Benchmarks: MathArena as an Evaluation Platform for Mathematics with LLMs.

Beyond the Mean: Multi-Moment Policy Optimization for LLM Reasoning Beyond Benchmarks: MathArena as an Evaluation Platform for Mathematics with LLMs

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-04T13:44:41.794435Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T13:44:41.794435Z digest=sha256:62a456975447a8996adc7832621a0d57d31ff469fb4f54dc0905f3f42701e1b0

Observation dd81640a-f32f-4349-8699-8fd4c34fdef8 · outbound

This paper cites Tulu 3: Pushing Frontiers in Open Language Model Post-Training.

Beyond the Mean: Multi-Moment Policy Optimization for LLM Reasoning Tulu 3: Pushing Frontiers in Open Language Model Post-Training

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-04T13:44:41.925937Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T13:44:41.925937Z digest=sha256:9a79ca9288a72b13a22b74d6bc3ce5a18e7420e00586ca5ff29f25f26037e995

Observation 76bb9b35-0e89-426d-aa2d-8445e1964eca · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Beyond the Mean: Multi-Moment Policy Optimization for LLM Reasoning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-04T13:44:42.094603Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T13:44:42.094603Z digest=sha256:24d9de373f55efa112ad625b845ea22b12630969ba3fed477d7b0b03eebd7165

Observation 601c01fc-f827-46be-80ab-f4ef28a8bd95 · outbound

This paper cites Information Gain-based Rollout Policy Optimization: An Adaptive Tree-Structured Rollout Approach for Multi-Turn LLM Agents.

Beyond the Mean: Multi-Moment Policy Optimization for LLM Reasoning Information Gain-based Rollout Policy Optimization: An Adaptive Tree-Structured Rollout Approach for Multi-Turn LLM Agents

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-04T13:44:42.254343Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T13:44:42.254343Z digest=sha256:acf25a6b04257ad24dfd82beda444b66c8cb6cfcb3f73e5595084e92c176cb78

Observation c5e6ef5a-aab2-45de-8d24-68d3f06d044f · outbound

This paper cites arXiv preprint arXiv:2509.25133 , year=.

Beyond the Mean: Multi-Moment Policy Optimization for LLM Reasoning arXiv preprint arXiv:2509.25133 , year=

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-04T13:44:42.441366Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T13:44:42.441366Z digest=sha256:38e2b25486e604363afb0f404cc7f8da54c18503af05ea4ac4b10a57738d6e3d

Observation 1616428c-6970-468b-a01c-d481697f6cfd · outbound

This paper cites Findings of the Association for Computational Linguistics: ACL 2026 , pages=.

Beyond the Mean: Multi-Moment Policy Optimization for LLM Reasoning Findings of the Association for Computational Linguistics: ACL 2026 , pages=

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-04T13:44:42.515466Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T13:44:42.515466Z digest=sha256:c80bedd86204f28c47552c0c25c4199e76ca7843c7d5fbeed1057798b45970ea

Observation 33f4a1b7-835d-41e6-a3a1-1ad31fd9f8fa · outbound

This paper cites arXiv preprint arXiv:2509.15207 , year=.

Beyond the Mean: Multi-Moment Policy Optimization for LLM Reasoning arXiv preprint arXiv:2509.15207 , year=

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-04T13:44:42.573566Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T13:44:42.573566Z digest=sha256:ac84a2146ee8d377256186b7cab6915700fae1492c6791d2b5532d4240647bfa

Observation ab809372-8108-4b96-a612-56dd41d4016c · outbound

This paper cites arXiv preprint arXiv:2509.26209 , year=.

Beyond the Mean: Multi-Moment Policy Optimization for LLM Reasoning arXiv preprint arXiv:2509.26209 , year=

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-04T13:44:42.659322Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T13:44:42.659322Z digest=sha256:3e2a6d69494cf6d7439267b5e61f6d63fbfa9dc8a245da708531d2f78d3fa536

Observation f465618d-2642-4c03-99f4-85b6e8c5f630 · outbound

This paper cites TreePO: Bridging the Gap of Policy Optimization and Efficacy and Inference Efficiency with Heuristic Tree-based Modeling.

Beyond the Mean: Multi-Moment Policy Optimization for LLM Reasoning TreePO: Bridging the Gap of Policy Optimization and Efficacy and Inference Efficiency with Heuristic Tree-based Modeling

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-04T13:44:42.721995Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T13:44:42.721995Z digest=sha256:2d919302bfb6ee5907123fe4f05e55f082eb90c32d1209051528b8e05cd1c224

Observation 74c03c94-b637-4681-8f32-77dc7cceaec1 · outbound

This paper cites Springer Series in Statistics ( , year=.

Beyond the Mean: Multi-Moment Policy Optimization for LLM Reasoning Springer Series in Statistics ( , year=

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-04T13:44:42.764654Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T13:44:42.764654Z digest=sha256:7fd10a73e081fd9a98ac930f7a55aa41e91cb9b7a8abf7e6ab65264f15f61df2

Observation d97f1f9c-884e-43d8-b2bb-2ee87fa170d5 · outbound

This paper cites International conference on machine learning , pages=.

Beyond the Mean: Multi-Moment Policy Optimization for LLM Reasoning International conference on machine learning , pages=

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-04T13:44:42.817011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T13:44:42.817011Z digest=sha256:612fd5040bf4c0ba7082b5e8eb01949373e83db6f64f9cc50af124cb7c18cd5f

Observation d3f7cfcf-88a9-4fe6-9cbb-5f4d42b469ec · outbound

This paper cites High-Dimensional Continuous Control Using Generalized Advantage Estimation.

Beyond the Mean: Multi-Moment Policy Optimization for LLM Reasoning High-Dimensional Continuous Control Using Generalized Advantage Estimation

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-04T13:44:42.889956Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T13:44:42.889956Z digest=sha256:6ff06cb629f3710fe79655ced0dde57631f9077f118cd1a94bd1537d43312469

Pith citing papers

No inbound Pith citation observations are available.