Pith. sign in

Paper Citation Record · LEDGER

P-Check: Advancing Personalized Reward Model via Learning to Generate Dynamic Checklist

As of 17 August 2026, this Paper Citation Record lists 8 of 8 outbound references and 1 inbound Pith citation observation for arXiv:2601.02986.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2601.02986 v3

Coverage vector

measured 8 of 8 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-03T12:29:44.180267Z

measured 9 of 9 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-10T13:58:53.430492Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-05-10T14:00:28.429589Z

Reference resolution

8 of 8 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved6
  • parse uncertain1
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d251bfe8-e455-4f79-8f50-b301e2fcc9b7 · outbound

This paper cites an unresolved cited work.

P-Check: Advancing Personalized Reward Model via Learning to Generate Dynamic Checklist Unresolved cited work

Reference 2

Resolution
parse uncertain
no resolver link, observed 2026-08-03T12:29:44.124830Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T12:29:44.124830Z digest=sha256:8440788b85105791de7daf3e16d6f559cc40635c3e6f38cada04b96b6165ae4e

Observation 241219b3-cab4-4018-924b-1d4a01c8c11f · outbound

This paper cites Use the checklist as feedback: incorporate relevant criteria while preserving correct and helpful content.

P-Check: Advancing Personalized Reward Model via Learning to Generate Dynamic Checklist Use the checklist as feedback: incorporate relevant criteria while preserving correct and helpful content

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-03T12:29:44.180267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T12:29:44.180267Z digest=sha256:d25d4e44ea1ecab2c21fa70fc76be240bdb5211d28995a749eaf505e8f0011ae

Observation 7c94dc5b-54a7-489e-9f51-e93e7a6c470a · outbound

This paper cites StepWiser: Stepwise Generative Judges for Wiser Reasoning.

P-Check: Advancing Personalized Reward Model via Learning to Generate Dynamic Checklist StepWiser: Stepwise Generative Judges for Wiser Reasoning

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-03T12:29:43.974215Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T12:29:43.974215Z digest=sha256:27dc6bae24f3920551c4ff19f3b7d54e59d54fa9b78d7fadef56e33256216a8e

Observation 2ddb89cf-e45f-4530-9a3b-1dea7a706d62 · outbound

This paper cites Scaling Relationship on Learning Mathematical Reasoning with Large Language Models.

P-Check: Advancing Personalized Reward Model via Learning to Generate Dynamic Checklist Scaling Relationship on Learning Mathematical Reasoning with Large Language Models

Reference 6

Resolution
malformed identifier
no resolver link, observed 2026-08-03T12:29:44.062804Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T12:29:44.062804Z digest=sha256:c08dc1d36b4f4c03706bb83561056c2c78df76a85588b2eeb141ec4f852a91a2

Observation d6a81e8f-f4af-48cb-860e-13fef9e080e3 · outbound

This paper cites Anisha Gunjal, Anthony Wang, Elaine Lau, Vaskar Nath, Yunzhong He, Bing Liu, and Sean Hendryx.

P-Check: Advancing Personalized Reward Model via Learning to Generate Dynamic Checklist Anisha Gunjal, Anthony Wang, Elaine Lau, Vaskar Nath, Yunzhong He, Bing Liu, and Sean Hendryx

Reference 1993

Resolution
unresolved
no resolver link, observed 2026-08-03T12:29:43.686297Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T12:29:43.686297Z digest=sha256:85f1259d9f61a0d154e9c0512f3296842feeb31789a9e5a927692edf189955d9

Observation 92c0a982-70b6-4323-94a8-1cc5954d08a9 · outbound

This paper cites Personalized Soups: Personalized Large Language Model Alignment via Post-hoc Parameter Merging.

P-Check: Advancing Personalized Reward Model via Learning to Generate Dynamic Checklist Personalized Soups: Personalized Large Language Model Alignment via Post-hoc Parameter Merging

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-03T12:29:43.847480Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T12:29:43.847480Z digest=sha256:c7467281dd2b4700eb52dcfc73bb0e576b5c7f2e79fa55e9fb8e3d9ec89980a2

Observation e32ca7d8-4c0b-4a10-93a1-8148a62a2950 · outbound

This paper cites Efficient Memory Management for Large Language Model Serving with PagedAttention.

P-Check: Advancing Personalized Reward Model via Learning to Generate Dynamic Checklist Efficient Memory Management for Large Language Model Serving with PagedAttention

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-03T12:29:43.908771Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T12:29:43.908771Z digest=sha256:202ed341a82b60227429daac1a20cce1a6d0ec8c7541899980f23f4031f5ba1a

Observation a747276c-e1b0-41d0-81a2-c2a8913f7025 · outbound

This paper cites Rubrics as Rewards: Reinforcement Learning Beyond Verifiable Domains.

P-Check: Advancing Personalized Reward Model via Learning to Generate Dynamic Checklist Rubrics as Rewards: Reinforcement Learning Beyond Verifiable Domains

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-03T12:29:43.786052Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T12:29:43.786052Z digest=sha256:6df8a7f74b32ae09d3d71ebfb4df165315c8a81ee85bbe1ff42348f4c75273ff

Pith citing papers

Observation 2e14b05e-f5bf-4458-ab3e-8169f130eb29 · inbound

Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges cites this paper.

Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges P-Check: Advancing Personalized Reward Model via Learning to Generate Dynamic Checklist

Reference 150

Resolution
verified exact
local_arxiv, observed 2026-05-10T14:00:28.432183Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-05-10T13:58:53.430492Z digest=sha256:74bd8ea87e6819f7e94176b9e46300b9451ffb24975cad2f1131ebbcedb71ead