Pith. sign in

Paper Citation Record · LEDGER

Multi-turn Reinforcement Learning from Preference Human Feedback

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 6 inbound Pith citation observations for arXiv:2405.14655.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2405.14655 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 6 of 6 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 6 of 6 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-09T18:18:09.280582Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-17T12:04:10.473003Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation be71c4ed-012f-4b6b-9f99-37eee45a950b · inbound

Training Language Models to Self-Correct via Reinforcement Learning cites this paper.

Training Language Models to Self-Correct via Reinforcement Learning Multi-turn Reinforcement Learning from Preference Human Feedback

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-05-17T12:04:10.476073Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-17T12:04:10.210508Z digest=sha256:221ae86f79fab1af9836c6292f595c3930ae0543f8200b5c98046709afa8fd5f

Observation b51c2a23-0629-41b7-baf2-ebe83f2b8bfd · inbound

Towards Large Reasoning Models: A Survey of Reinforced Reasoning with Large Language Models cites this paper.

Towards Large Reasoning Models: A Survey of Reinforced Reasoning with Large Language Models Multi-turn Reinforcement Learning from Preference Human Feedback

Reference 128

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T21:20:59.256774Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-15T21:20:59.128986Z digest=sha256:8a1735af9c6dc1ec2f66aa7303d1646d85eb690aaa76f2ea28256538b3d384be

Observation 65733ea7-e156-4cf6-a895-3d651428df4c · inbound

CollabLLM: From Passive Responders to Active Collaborators cites this paper.

CollabLLM: From Passive Responders to Active Collaborators Multi-turn Reinforcement Learning from Preference Human Feedback

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-09T18:18:09.280582Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T18:18:09.280582Z digest=sha256:cd0750605cb36d245b32153b72fe2a36632a18d9ca8cc87f82d790930d92c826

Observation 06c03cc1-5efd-4109-ad8b-30f0011fe3dc · inbound

Response-Level Rewards Are All You Need for Online Reinforcement Learning in LLMs: A Mathematical Perspective cites this paper.

Response-Level Rewards Are All You Need for Online Reinforcement Learning in LLMs: A Mathematical Perspective Multi-turn Reinforcement Learning from Preference Human Feedback

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T11:32:30.406644Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:32:30.406644Z digest=sha256:0e777f0d79eb5ba700f658712936f3b2abd493f90a9500cf58675279f942faf0

Observation be2dfc51-47b6-4556-814c-b0588fa2897e · inbound

PAG: Multi-Turn Reinforced LLM Self-Correction with Policy as Generative Verifier cites this paper.

PAG: Multi-Turn Reinforced LLM Self-Correction with Policy as Generative Verifier Multi-turn Reinforcement Learning from Preference Human Feedback

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T04:39:35.492469Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:39:35.492469Z digest=sha256:a82f000b20a60a46049299bd381bf2df2e8ce18648bd86bb52f70ef93b5f1e16

Observation 3327ff83-8f88-4819-8567-34e633969dde · inbound

Robust pid sliding mode control for dc servo motor speed control cites this paper.

Robust pid sliding mode control for dc servo motor speed control Multi-turn Reinforcement Learning from Preference Human Feedback

Reference 2010

Resolution
unresolved
no resolver link, observed 2026-08-05T23:42:04.428118Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T23:42:04.428118Z digest=sha256:64ef88f2a804f8873aef2bf3c20ec80667a93decf8b696754baaaa1d18fa6226