Pith. sign in

Paper Citation Record · LEDGER

Hint-Guided Diversified Policy Optimization for LLM Reasoning

As of 23 August 2026, this Paper Citation Record lists 35 of 35 outbound references and 0 inbound Pith citation observations for arXiv:2606.03021.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2606.03021 v2

Coverage vector

measured 35 of 35 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-02T12:35:53.841154Z

measured 35 of 35 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-23T06:30:58.430688+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

35 of 35 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved34
  • parse uncertain1
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 9cd82d61-b56b-4d8f-9f13-34bd64e80f03 · outbound

This paper cites It should only elaborate on the high-level strategies and concepts, without going into specific calculations.

Hint-Guided Diversified Policy Optimization for LLM Reasoning It should only elaborate on the high-level strategies and concepts, without going into specific calculations

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-02T12:35:52.052898Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T12:35:52.052898Z digest=sha256:3a3e96dfc94622c1f55eaa3b1a5c034925e66e0b8c955d848bc3798369f55601

Observation 385afc16-3699-4125-8320-8e9d49d717e7 · outbound

This paper cites an unresolved cited work.

Hint-Guided Diversified Policy Optimization for LLM Reasoning Unresolved cited work

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-02T12:35:52.105117Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T12:35:52.105117Z digest=sha256:fca1baf6e98f6a843b85a4b46a912458994748ff4d99522ba7b571eae6e95cd9

Observation 899a701b-d902-4f49-be45-ba8587e7d2b1 · outbound

This paper cites an unresolved cited work.

Hint-Guided Diversified Policy Optimization for LLM Reasoning Unresolved cited work

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-02T12:35:52.185053Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T12:35:52.185053Z digest=sha256:0cb447380ceb1459343972e874cc975db8a8390bf8761c144bc1f904fa685882

Observation 979abb64-bdab-4e0d-b2b6-67b019b00cd8 · outbound

This paper cites Qwen3 Technical Report.

Hint-Guided Diversified Policy Optimization for LLM Reasoning Qwen3 Technical Report

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-02T12:35:51.310647Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T12:35:51.310647Z digest=sha256:c68becc70597d8263dd88a319615b5dea6bb2a7234f6f75f577301ca97cf9043

Observation b7129faf-0e80-44ad-910e-eebc887b4ae0 · outbound

This paper cites propose-select-think.

Hint-Guided Diversified Policy Optimization for LLM Reasoning propose-select-think

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-02T12:35:52.835670Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T12:35:52.835670Z digest=sha256:b13cbb9613e62acc040ea239d90030b156340351b85b11b721a0e47a14b49262

Observation 76bdf2ed-787b-4446-a334-e03c072b8e29 · outbound

This paper cites Yes” as a measure of the similarity between the two candi- date solutions. As shown by “HDPO (LLM-Div).

Hint-Guided Diversified Policy Optimization for LLM Reasoning Yes” as a measure of the similarity between the two candi- date solutions. As shown by “HDPO (LLM-Div)

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-02T12:35:51.700318Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T12:35:51.700318Z digest=sha256:d5d352312d4578efa5d78c73dfa4a67b7a38f998651c25056f83dd841bac34ff

Observation 39731c6f-2d28-4915-b94e-3bc739c303af · outbound

This paper cites an unresolved cited work.

Hint-Guided Diversified Policy Optimization for LLM Reasoning Unresolved cited work

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-02T12:35:51.845328Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T12:35:51.845328Z digest=sha256:b6f82ef481d2e1bf826145563acb61107f55bdd8eb33f66939585ed054c23c3f

Observation dd16bfac-7cb3-42c3-ad1b-fe3cd3f88465 · outbound

This paper cites ex- plore–evaluate–select.

Hint-Guided Diversified Policy Optimization for LLM Reasoning ex- plore–evaluate–select

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-02T12:35:51.995434Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T12:35:51.995434Z digest=sha256:78d325c3c0dfd09016a9dc92135b3b059b1e477e7ea76107b85655383c13cd7e

Observation 5e99ce3d-65a6-453e-933c-ba1c6eb4c932 · outbound

This paper cites [1]”, “[2].

Hint-Guided Diversified Policy Optimization for LLM Reasoning [1]”, “[2]

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-02T12:35:52.245160Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T12:35:52.245160Z digest=sha256:71c851c376b8846aa50033abf6550899ce6139d8c109c0d1aff73ef8861c6505

Observation 6d28b855-43b2-407a-a798-1051481b37e7 · outbound

This paper cites an unresolved cited work.

Hint-Guided Diversified Policy Optimization for LLM Reasoning Unresolved cited work

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-02T12:35:52.311408Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T12:35:52.311408Z digest=sha256:1d9530fa220a49c9c40ad4332043f36501ccc01da0e9dfb5cd5058fc3eb621ca

Observation ef730c80-1d06-4cc6-9e13-5492557f0045 · outbound

This paper cites Yes", otherwise output.

Hint-Guided Diversified Policy Optimization for LLM Reasoning Yes", otherwise output

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-02T12:35:52.362940Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T12:35:52.362940Z digest=sha256:ecf4c04c646de7c64a7b43a17de7155aa459cd17c6444b30fc33b241f6f803bf

Observation c9fa484b-de64-4e03-9d2c-d9ef92b202b8 · outbound

This paper cites an unresolved cited work.

Hint-Guided Diversified Policy Optimization for LLM Reasoning Unresolved cited work

Reference 15

Resolution
parse uncertain
no resolver link, observed 2026-08-02T12:35:52.414340Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T12:35:52.414340Z digest=sha256:7ed9aea3dc6492caefb8992031831a2e22c20fd1e89f112df9646bc7f506e113

Observation 2a052d0c-7bf0-476b-92d9-46860eae661c · outbound

This paper cites an unresolved cited work.

Hint-Guided Diversified Policy Optimization for LLM Reasoning Unresolved cited work

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-02T12:35:52.465863Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T12:35:52.465863Z digest=sha256:29baeb8b57e2bf45ebce7b19da26d2af155080bfb72a44a02ae3105415c4b149

Observation d53f7e89-9203-4e0a-8659-4df956b38678 · outbound

This paper cites We need to find the values ofa,b, andcthat maximize|a|+|b|+|c|while satisfying these constraints.

Hint-Guided Diversified Policy Optimization for LLM Reasoning We need to find the values ofa,b, andcthat maximize|a|+|b|+|c|while satisfying these constraints

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-02T12:35:52.513747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T12:35:52.513747Z digest=sha256:c1cd8ffcbe28db59b8b1f9d77597f78d5fae773831410df5e83f559832005115

Observation 9c57d91c-3805-4377-ba7f-785502f9e381 · outbound

This paper cites Then express a, b, c in terms of these values and use linear programming or symmetry arguments to maximize |a| + |b| + |c|.

Hint-Guided Diversified Policy Optimization for LLM Reasoning Then express a, b, c in terms of these values and use linear programming or symmetry arguments to maximize |a| + |b| + |c|

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-02T12:35:52.591365Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T12:35:52.591365Z digest=sha256:9a4c361e3d6151552c3be1d662409b6c64ea85aa6f952af097c4263233df0d9d

Observation 9beee6ca-e8d8-48be-b696-a85acde1f562 · outbound

This paper cites Apply the method of Lagrange multipliers to maximize the linear functional |a| + |b| + |c| subject to the quadratic constraint.

Hint-Guided Diversified Policy Optimization for LLM Reasoning Apply the method of Lagrange multipliers to maximize the linear functional |a| + |b| + |c| subject to the quadratic constraint

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-02T12:35:52.653407Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T12:35:52.653407Z digest=sha256:f717b68cc25e5fc0068fcfc6efd07bf1be8b8f553a8559fc7663abd7baddc9f3

Observation 1248710d-c51b-4cf2-8ab1-450984e3ebb0 · outbound

This paper cites However, this may miss the global maximum if the optimal polynomial is not symmetric or has non-zero a and b.

Hint-Guided Diversified Policy Optimization for LLM Reasoning However, this may miss the global maximum if the optimal polynomial is not symmetric or has non-zero a and b

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-02T12:35:52.702950Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T12:35:52.702950Z digest=sha256:41232bf22d1ba3974898b5ce0169480f5ef51d391b2d89f8dbc2586f11c46c54

Observation 3428ce2f-5445-452a-b208-d0731abe43b4 · outbound

This paper cites Scale and shift the Chebyshev polynomial to satisfy the bound|P(x)| ≤1and compute the coefficients to find the maximum of |a| + |b| + |c|.

Hint-Guided Diversified Policy Optimization for LLM Reasoning Scale and shift the Chebyshev polynomial to satisfy the bound|P(x)| ≤1and compute the coefficients to find the maximum of |a| + |b| + |c|

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-02T12:35:52.756629Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T12:35:52.756629Z digest=sha256:3f2b282445d5a09bdba907396e99942d6fffeeccc9abc9b000e4d65b91e15faf

Observation d4a4c7f8-5cfd-49b1-88d5-6bc9f4879966 · outbound

This paper cites - PointPis 4 units away from the circle, so the distance fromPto the centerOis6 + 4 = 10.

Hint-Guided Diversified Policy Optimization for LLM Reasoning - PointPis 4 units away from the circle, so the distance fromPto the centerOis6 + 4 = 10

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-02T12:35:52.893352Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T12:35:52.893352Z digest=sha256:0e596950d65ec76f9b7ff5eff4ed0e54a0a52b31d4b79f650430fea7110836b7

Observation 24ba2916-f45f-4821-9ce0-eadac1e024ca · outbound

This paper cites But sincePis 4 units away from the circle and AB is parallel to ← →OP, the perpendicular distance fromOto ABmust be 4 (as 8 would place AB outside the circle).

Hint-Guided Diversified Policy Optimization for LLM Reasoning But sincePis 4 units away from the circle and AB is parallel to ← →OP, the perpendicular distance fromOto ABmust be 4 (as 8 would place AB outside the circle)

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-02T12:35:52.981079Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T12:35:52.981079Z digest=sha256:46aad6a89f1f28999dd97d3a9eb3b6aa977421db64995af52871a6467fe7af1c

Observation b2be02ac-7311-420d-92d5-4a9a777144d8 · outbound

This paper cites - The chord ABis parallel to the x-axis and 2 units below the x-axis (since the distance from Oto ABis 4).

Hint-Guided Diversified Policy Optimization for LLM Reasoning - The chord ABis parallel to the x-axis and 2 units below the x-axis (since the distance from Oto ABis 4)

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-02T12:35:53.047009Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T12:35:53.047009Z digest=sha256:2756d533ec7018fec570192987dbae06d65af38d372fe6c0129489ec8ab45f22

Observation 5a7f3fdb-33e5-41a6-acbe-b62e5ad9d97d · outbound

This paper cites an unresolved cited work.

Hint-Guided Diversified Policy Optimization for LLM Reasoning Unresolved cited work

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-02T12:35:53.106699Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T12:35:53.106699Z digest=sha256:bd8d0c756ef760fda1c20e67b4b730f84783b913e9003f37f816270232733b61

Observation 0dcc36f0-25ae-4a07-bf5c-4094777d030c · outbound

This paper cites 19 Case 2 (Generation Model: Qwen2.5-Math-7B-HDPO) Question: CircleOhas radius 6.

Hint-Guided Diversified Policy Optimization for LLM Reasoning 19 Case 2 (Generation Model: Qwen2.5-Math-7B-HDPO) Question: CircleOhas radius 6

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-02T12:35:53.291528Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T12:35:53.291528Z digest=sha256:1214f0ae71347877b7e6862d22cc87c5b76ecea563fd157e91bd7aea917ec761

Observation a930312e-37f1-4013-9a76-38a24a41f4a8 · outbound

This paper cites SinceABis parallel toOP and the distance between them is 2, the perpendicular distance fromOtoABis either4 + 2 = 6 or4−2 = 2.

Hint-Guided Diversified Policy Optimization for LLM Reasoning SinceABis parallel toOP and the distance between them is 2, the perpendicular distance fromOtoABis either4 + 2 = 6 or4−2 = 2

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-02T12:35:53.450558Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T12:35:53.450558Z digest=sha256:ac141d627fc46e952f7f88c4003c3b6942864caaadc2b0f5c50707ffd88f57ed

Observation a760fd0a-70d2-482e-9e34-cfe6b241c814 · outbound

This paper cites an unresolved cited work.

Hint-Guided Diversified Policy Optimization for LLM Reasoning Unresolved cited work

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-02T12:35:53.506946Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T12:35:53.506946Z digest=sha256:709fe0e3013ba4ccedc4038be9e4108b46831a0e12e528c91cb863715258654f

Observation 4d786531-c111-4ad6-bd7c-a381c23304cc · outbound

This paper cites Since chordABis parallel to ← →OP, it is horizontal, and the distance betweenAB and ← →OPis 2, soABis either aty= 2ory=−2.

Hint-Guided Diversified Policy Optimization for LLM Reasoning Since chordABis parallel to ← →OP, it is horizontal, and the distance betweenAB and ← →OPis 2, soABis either aty= 2ory=−2

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-02T12:35:53.617668Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T12:35:53.617668Z digest=sha256:94ce8f1660ba31cd4403c5841482f45f3fc618c7c3d62b51ccad3db82c22fbb7

Observation cdb55b2d-a1c6-44a6-b6c4-f3d6fbd63325 · outbound

This paper cites propose-select-think.

Hint-Guided Diversified Policy Optimization for LLM Reasoning propose-select-think

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-02T12:35:53.677733Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T12:35:53.677733Z digest=sha256:29b7431f26943c6ce11248ce5b8e8326e740c3bb136103f7acad18aef3adca3f

Observation 1545a31a-f893-4671-aacc-0b9ec204f355 · outbound

This paper cites Then apply the sum of cosine series formula for angles in arithmetic sequence, simplifying the resulting expression using symmetry and periodicity of the cosine function.

Hint-Guided Diversified Policy Optimization for LLM Reasoning Then apply the sum of cosine series formula for angles in arithmetic sequence, simplifying the resulting expression using symmetry and periodicity of the cosine function

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-02T12:35:53.745864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T12:35:53.745864Z digest=sha256:6bf64e120e2a89fd1b0b31283b12a930de313860bdb5c22077c58e397c56cbbe

Observation 69b98ad9-4d0a-47e5-9252-bb491351ac34 · outbound

This paper cites However, this approach lacks precision and relies on approximation, making it unsuitable for exact computation.

Hint-Guided Diversified Policy Optimization for LLM Reasoning However, this approach lacks precision and relies on approximation, making it unsuitable for exact computation

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-02T12:35:53.793195Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T12:35:53.793195Z digest=sha256:9bab7fac407930f8930683a8a695858dc2005f27ad37e8421c0581b4af4697e1

Observation d77077b0-1efc-4674-897d-4c46714e7023 · outbound

This paper cites an unresolved cited work.

Hint-Guided Diversified Policy Optimization for LLM Reasoning Unresolved cited work

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-02T12:35:53.820126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T12:35:53.820126Z digest=sha256:79496796723eba735ac903e0e9a37d8441c5b8cce2d9b082c33a4a2dcbd6b4f3

Observation 83cc1e74-017d-4047-a9aa-072a95d4cd11 · outbound

This paper cites </Candidate Solutions> <selected>[1]</selected> <thinking> We are given the sum: sin2 4◦ + sin2 8◦ + sin2 12◦ +· · ·+ sin2 176◦ This is a sum ofsin 2 θforθ= 4k ◦ wherek= 1,2,.

Hint-Guided Diversified Policy Optimization for LLM Reasoning </Candidate Solutions> <selected>[1]</selected> <thinking> We are given the sum: sin2 4◦ + sin2 8◦ + sin2 12◦ +· · ·+ sin2 176◦ This is a sum ofsin 2 θforθ= 4k ◦ wherek= 1,2,

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-02T12:35:53.841154Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T12:35:53.841154Z digest=sha256:4b59343ae632253772391cceabf4651b5c45d555489ac9e93d628195cef38e41

Observation 86cde73d-8e00-4f3c-9df8-f0a67f849b41 · outbound

This paper cites GPQA: A Graduate-Level Google-Proof Q&A Benchmark.

Hint-Guided Diversified Policy Optimization for LLM Reasoning GPQA: A Graduate-Level Google-Proof Q&A Benchmark

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-02T12:35:50.958736Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T12:35:50.958736Z digest=sha256:05c848f1e8d8c11454817aa762ceac9f96c1ceeb734939e9ac37019a449cf797

Observation 8a118376-6d2a-4c97-be81-c204f113c223 · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

Hint-Guided Diversified Policy Optimization for LLM Reasoning DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-02T12:35:51.506873Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T12:35:51.506873Z digest=sha256:9211d0f65f7cab803c7ee1f1a7048299c7dcbe1b8184af7239f294558f1dc719

Observation 5f2bc28b-a410-497d-b18d-f30fbda8d347 · outbound

This paper cites Learning to Reason under Off-Policy Guidance.

Hint-Guided Diversified Policy Optimization for LLM Reasoning Learning to Reason under Off-Policy Guidance

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-02T12:35:51.091898Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T12:35:51.091898Z digest=sha256:a5109a1827c48f9a6b96d6ad44755677e2688561817b0ff2aba9f2633230ed61

Observation 1db52df6-2413-495b-b2a4-73745f4af1a5 · outbound

This paper cites Kimi k1.5: Scaling Reinforcement Learning with LLMs.

Hint-Guided Diversified Policy Optimization for LLM Reasoning Kimi k1.5: Scaling Reinforcement Learning with LLMs

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-02T12:35:51.048915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T12:35:51.048915Z digest=sha256:b4bdf54bbde15771b3886e711f44cfbdaca3950ceaa2e8365be4aa9ab1fbc970

Pith citing papers

No inbound Pith citation observations are available.