Pith. sign in

Paper Citation Record · LEDGER

HPO: Hysteretic Policy Optimization for Stable and Efficient Training under Sparse-Reward Regime

As of 22 August 2026, this Paper Citation Record lists 23 of 23 outbound references and 0 inbound Pith citation observations for arXiv:2605.30201.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.30201 v1

Coverage vector

measured 23 of 23 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-06-29T09:01:03.618498Z

measured 23 of 23 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

23 of 23 outbound references displayed

  • verified exact3
  • verified fuzzy0
  • unresolved20
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e39a9985-dd35-4723-b520-3d5480392985 · outbound

This paper cites Enhancing reinforcement learning with dense rewards from language model critic.

HPO: Hysteretic Policy Optimization for Stable and Efficient Training under Sparse-Reward Regime Enhancing reinforcement learning with dense rewards from language model critic

Reference 1

Resolution
unresolved
no resolver link, observed 2026-06-29T09:01:03.618498Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T09:01:03.618498Z digest=sha256:1dcf63443ed2af4ad4d7e3f1e1171fa286969084f60a3d19b7dc95f417848740

Observation 69caea79-9fe4-46ea-9d42-fd4eff773694 · outbound

This paper cites an unresolved cited work.

HPO: Hysteretic Policy Optimization for Stable and Efficient Training under Sparse-Reward Regime Unresolved cited work

Reference 2

Resolution
unresolved
no resolver link, observed 2026-06-29T09:01:03.618498Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T09:01:03.618498Z digest=sha256:1124ee50361819884be235948935dc7184c12a137826ff2856ee6cd79e0dc3df

Observation 611dcfd6-8825-4c0b-97a5-d13de040fd84 · outbound

This paper cites Zhang, Han Bao, Hanwei Xu, Haocheng Wang, Haowei Zhang, Honghui Ding, Huajian Xin, Huazuo Gao, Hui Li, Hui Qu, J.

HPO: Hysteretic Policy Optimization for Stable and Efficient Training under Sparse-Reward Regime Zhang, Han Bao, Hanwei Xu, Haocheng Wang, Haowei Zhang, Honghui Ding, Huajian Xin, Huazuo Gao, Hui Li, Hui Qu, J

Reference 3

Resolution
unresolved
no resolver link, observed 2026-06-29T09:01:03.618498Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T09:01:03.618498Z digest=sha256:a32d21f878f8578ddff263c7f64ed45353deb4a82c5acb1b370183b7df4aeeb5

Observation 7d5d61ed-08a6-4266-9fb1-27a59694138a · outbound

This paper cites DenseGRPO: From sparse to dense reward for flow matching model alignment.

HPO: Hysteretic Policy Optimization for Stable and Efficient Training under Sparse-Reward Regime DenseGRPO: From sparse to dense reward for flow matching model alignment

Reference 4

Resolution
unresolved
no resolver link, observed 2026-06-29T09:01:03.618498Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T09:01:03.618498Z digest=sha256:125e4b354dc8f75bf3b08dfc0367fda12adff6bd240d009b49ad574e8547f99e

Observation 39fc10bb-581c-4451-91b6-d2f5bce7668c · outbound

This paper cites RAFT: Reward ranked finetuning for generative foundation model alignment.Transactions on Machine Learning Research, 2023.

HPO: Hysteretic Policy Optimization for Stable and Efficient Training under Sparse-Reward Regime RAFT: Reward ranked finetuning for generative foundation model alignment.Transactions on Machine Learning Research, 2023

Reference 5

Resolution
unresolved
no resolver link, observed 2026-06-29T09:01:03.618498Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T09:01:03.618498Z digest=sha256:fb7f7005956badd9d92661f51066ba2f443b07207da3ea85a5ce99e1bf9e2eed

Observation bab408fa-a595-4120-b294-c0ef882a71b7 · outbound

This paper cites Re- wardmap: Tackling sparse rewards in fine-grained visual reasoning via multi-stage reinforce- ment learning.

HPO: Hysteretic Policy Optimization for Stable and Efficient Training under Sparse-Reward Regime Re- wardmap: Tackling sparse rewards in fine-grained visual reasoning via multi-stage reinforce- ment learning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-06-29T09:01:03.618498Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T09:01:03.618498Z digest=sha256:9fdd01033678569d86a6e7c6c1a991dda07be547d03b2f7351fa2d837da21646

Observation aacf218f-d7d6-47a5-929d-39514e07cbd5 · outbound

This paper cites an unresolved cited work.

HPO: Hysteretic Policy Optimization for Stable and Efficient Training under Sparse-Reward Regime Unresolved cited work

Reference 7

Resolution
unresolved
no resolver link, observed 2026-06-29T09:01:03.618498Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T09:01:03.618498Z digest=sha256:5504e4b042871be20b549cf94d02d4cd7daecf2b66fe035c5eb56182df27c3c3

Observation 09a5a2da-c1e2-455c-84b2-a5f5cd768f49 · outbound

This paper cites Soft adaptive policy optimization, 2025.

HPO: Hysteretic Policy Optimization for Stable and Efficient Training under Sparse-Reward Regime Soft adaptive policy optimization, 2025

Reference 8

Resolution
unresolved
no resolver link, observed 2026-06-29T09:01:03.618498Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T09:01:03.618498Z digest=sha256:e18ba2f01cc0e36ed335f9a5e04d96431dff7de0556740720d4a9100e07af49d

Observation bdf835ff-720b-4f43-ac3d-6abf490b13b4 · outbound

This paper cites Rewarding the unlikely: Lifting grpo beyond distribution sharpening, 2025.

HPO: Hysteretic Policy Optimization for Stable and Efficient Training under Sparse-Reward Regime Rewarding the unlikely: Lifting grpo beyond distribution sharpening, 2025

Reference 9

Resolution
unresolved
no resolver link, observed 2026-06-29T09:01:03.618498Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T09:01:03.618498Z digest=sha256:9e3b66d6d92490f3e36152b91ff54e6193b1833dc2bb945431b4acb400d7468f

Observation 1f0f7608-ec8c-4634-959c-e421be09efc5 · outbound

This paper cites Reinforcement Learning via Self-Distillation.

HPO: Hysteretic Policy Optimization for Stable and Efficient Training under Sparse-Reward Regime Reinforcement Learning via Self-Distillation

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-06-29T09:03:15.687060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-29T09:01:03.618498Z digest=sha256:b091847a98e48abbfb7edb776b400878ea19b7691ce3597238a0d64894c097d5

Observation a82861da-32fc-4f52-9bb3-d870baa9158a · outbound

This paper cites Understanding r1-zero-like training: A critical perspective.

HPO: Hysteretic Policy Optimization for Stable and Efficient Training under Sparse-Reward Regime Understanding r1-zero-like training: A critical perspective

Reference 11

Resolution
unresolved
no resolver link, observed 2026-06-29T09:01:03.618498Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T09:01:03.618498Z digest=sha256:a5e422b1d0f932a64bcd93fc6f292296c8a3f7316f7c10c8a162b8a7fb79dc4b

Observation 90001c7c-6a19-4a2c-9d22-7af35690bd14 · outbound

This paper cites Hysteretic Q-Learning: An Algorithm for Decentralized Reinforcement Learning in Cooperative Multi-agent Teams.

HPO: Hysteretic Policy Optimization for Stable and Efficient Training under Sparse-Reward Regime Hysteretic Q-Learning: An Algorithm for Decentralized Reinforcement Learning in Cooperative Multi-agent Teams

Reference 12

Resolution
unresolved
no resolver link, observed 2026-06-29T09:01:03.618498Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T09:01:03.618498Z digest=sha256:35fb258eca7eb0c9d9929c05d5a07093f0104cd12639d0ecdfcec3a4ead6f307

Observation ee0fb757-c0d5-4099-b7ea-692ddbb5cc69 · outbound

This paper cites an unresolved cited work.

HPO: Hysteretic Policy Optimization for Stable and Efficient Training under Sparse-Reward Regime Unresolved cited work

Reference 13

Resolution
unresolved
no resolver link, observed 2026-06-29T09:01:03.618498Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T09:01:03.618498Z digest=sha256:032cf84a22e819565a55c1703cbfc9450903bf9c9c725851cbfb1a3dc574a2de

Observation f9c2ff86-5807-4dd2-867f-535ac000ceef · outbound

This paper cites Tinyzero.

HPO: Hysteretic Policy Optimization for Stable and Efficient Training under Sparse-Reward Regime Tinyzero

Reference 14

Resolution
unresolved
no resolver link, observed 2026-06-29T09:01:03.618498Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T09:01:03.618498Z digest=sha256:68cee398836790dad58d95652b7ee5c69f90737e8f36b11b4c4a4bfb836d279e

Observation 79636ff0-2e79-4f0f-a5f1-4bf6fb3644df · outbound

This paper cites Direct preference optimization: Your language model is secretly a reward model.

HPO: Hysteretic Policy Optimization for Stable and Efficient Training under Sparse-Reward Regime Direct preference optimization: Your language model is secretly a reward model

Reference 15

Resolution
unresolved
no resolver link, observed 2026-06-29T09:01:03.618498Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T09:01:03.618498Z digest=sha256:10622e7d4ca4ff5ca020749738ac1988ef9b09a983ded5236950b9ea0137448f

Observation 17f6492f-5404-4f03-9f86-2988ac74c19c · outbound

This paper cites Reasoning Language Models for Root Cause Analysis in 5G Wireless Networks.

HPO: Hysteretic Policy Optimization for Stable and Efficient Training under Sparse-Reward Regime Reasoning Language Models for Root Cause Analysis in 5G Wireless Networks

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-06-29T09:03:15.687531Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-29T09:01:03.618498Z digest=sha256:64f62d716a8d6aa6e529e8c76c84f3af33c87badd051af7663a65ae44e29e168

Observation 88713596-c662-4a69-bfc0-466991631b9e · outbound

This paper cites Proximal Policy Optimization Algorithms.

HPO: Hysteretic Policy Optimization for Stable and Efficient Training under Sparse-Reward Regime Proximal Policy Optimization Algorithms

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-06-29T09:03:15.690088Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-29T09:01:03.618498Z digest=sha256:ae0c4f27679b807b5f77df40e6232c758836d346cc701dfaf2bb0cd30bbc9d90

Observation b576e0ec-3515-4803-9bf4-6852d2adfd61 · outbound

This paper cites an unresolved cited work.

HPO: Hysteretic Policy Optimization for Stable and Efficient Training under Sparse-Reward Regime Unresolved cited work

Reference 18

Resolution
unresolved
no resolver link, observed 2026-06-29T09:01:03.618498Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T09:01:03.618498Z digest=sha256:622fa8305174fd3138560d799cda9c33687ddfdba2a5284b50cba8a29aaee321

Observation 4839da91-6396-4c61-a9ed-26222be91207 · outbound

This paper cites Hybridflow: A flexible and efficient rlhf framework.

HPO: Hysteretic Policy Optimization for Stable and Efficient Training under Sparse-Reward Regime Hybridflow: A flexible and efficient rlhf framework

Reference 19

Resolution
unresolved
no resolver link, observed 2026-06-29T09:01:03.618498Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T09:01:03.618498Z digest=sha256:afaec6b7e0615a86e60975cca65c5870eb01a92b5e5a5d4f1e3e7d4b7eb3c798

Observation f3b0af17-308b-461f-88e6-68848bae52c0 · outbound

This paper cites A minimalist approach to llm reasoning: from rejection sampling to reinforce, 2025.

HPO: Hysteretic Policy Optimization for Stable and Efficient Training under Sparse-Reward Regime A minimalist approach to llm reasoning: from rejection sampling to reinforce, 2025

Reference 20

Resolution
unresolved
no resolver link, observed 2026-06-29T09:01:03.618498Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T09:01:03.618498Z digest=sha256:6d7a53562153d515dc2023df82d2f051757aac93e4910686941fa287af1dccc7

Observation 3ad9e361-91c0-4369-aa80-4692c14bd898 · outbound

This paper cites Qwen3 technical report, 2025.

HPO: Hysteretic Policy Optimization for Stable and Efficient Training under Sparse-Reward Regime Qwen3 technical report, 2025

Reference 21

Resolution
unresolved
no resolver link, observed 2026-06-29T09:01:03.618498Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T09:01:03.618498Z digest=sha256:2151c211632b186e48dcdd825d62663a86553bae7ee740ba1bef3beab88c6d44

Observation 29e2724e-02d8-4372-9a41-891c4beb1d01 · outbound

This paper cites Dapo: An open-source llm reinforcement learning system at scale, 2025.

HPO: Hysteretic Policy Optimization for Stable and Efficient Training under Sparse-Reward Regime Dapo: An open-source llm reinforcement learning system at scale, 2025

Reference 22

Resolution
unresolved
no resolver link, observed 2026-06-29T09:01:03.618498Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T09:01:03.618498Z digest=sha256:579486131693458e60f46675e82793024675fa35cf8f893e1a0a326ff4a988ff

Observation ff42b07c-5c05-402d-a176-6895a276780d · outbound

This paper cites Group sequence policy optimization, 2025.

HPO: Hysteretic Policy Optimization for Stable and Efficient Training under Sparse-Reward Regime Group sequence policy optimization, 2025

Reference 23

Resolution
unresolved
no resolver link, observed 2026-06-29T09:01:03.618498Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-29T09:01:03.618498Z digest=sha256:00ebae495f6b5ab4b82849edd4b84927af4fd7343040a88b6ca88ea137f566ef

Pith citing papers

No inbound Pith citation observations are available.