Pith. sign in

Paper Citation Record · LEDGER

P-Check: Advancing Personalized Reward Model via Learning to Generate Dynamic Checklist

As of 10 August 2026, this Paper Citation Record lists 8 of 8 outbound references and 1 inbound Pith citation observation for arXiv:2601.02986.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2601.02986 v3

Coverage vector

measured 8 of 8 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-03T12:29:44.180267Z

measured 9 of 9 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-10T13:58:53.430492Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-05-10T14:00:28.429589Z

Reference resolution

8 of 8 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved6
  • parse uncertain1
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d251bfe8-e455-4f79-8f50-b301e2fcc9b7 · outbound

This paper cites an unresolved cited work.

P-Check: Advancing Personalized Reward Model via Learning to Generate Dynamic Checklist Unresolved cited work

Reference 2

Resolution
parse uncertain
no resolver link, observed 2026-08-03T12:29:44.124830Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T12:29:44.124830Z digest=sha256:6affa34b2cd38b6d6ecd4bfcabe3c167ed310b960d413c8e08f25ed09f88ec04

Observation 241219b3-cab4-4018-924b-1d4a01c8c11f · outbound

This paper cites Use the checklist as feedback: incorporate relevant criteria while preserving correct and helpful content.

P-Check: Advancing Personalized Reward Model via Learning to Generate Dynamic Checklist Use the checklist as feedback: incorporate relevant criteria while preserving correct and helpful content

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-03T12:29:44.180267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T12:29:44.180267Z digest=sha256:d332cad04ca04dd199bcf7682a528ed2fcbe4b4f03bfa7ea43a01aa704cec43f

Observation 7c94dc5b-54a7-489e-9f51-e93e7a6c470a · outbound

This paper cites StepWiser: Stepwise Generative Judges for Wiser Reasoning.

P-Check: Advancing Personalized Reward Model via Learning to Generate Dynamic Checklist StepWiser: Stepwise Generative Judges for Wiser Reasoning

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-03T12:29:43.974215Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T12:29:43.974215Z digest=sha256:809d058556b42b7a70466616f87249d4992f86252951551c040870cf52823adb

Observation 2ddb89cf-e45f-4530-9a3b-1dea7a706d62 · outbound

This paper cites Scaling Relationship on Learning Mathematical Reasoning with Large Language Models.

P-Check: Advancing Personalized Reward Model via Learning to Generate Dynamic Checklist Scaling Relationship on Learning Mathematical Reasoning with Large Language Models

Reference 6

Resolution
malformed identifier
no resolver link, observed 2026-08-03T12:29:44.062804Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T12:29:44.062804Z digest=sha256:91b96d4e40042964942cb26a396ad128a58c508de8b10d7388afadc591ee9be7

Observation d6a81e8f-f4af-48cb-860e-13fef9e080e3 · outbound

This paper cites Anisha Gunjal, Anthony Wang, Elaine Lau, Vaskar Nath, Yunzhong He, Bing Liu, and Sean Hendryx.

P-Check: Advancing Personalized Reward Model via Learning to Generate Dynamic Checklist Anisha Gunjal, Anthony Wang, Elaine Lau, Vaskar Nath, Yunzhong He, Bing Liu, and Sean Hendryx

Reference 1993

Resolution
unresolved
no resolver link, observed 2026-08-03T12:29:43.686297Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T12:29:43.686297Z digest=sha256:31c155be5c82dae14b8ba850190d56b102d0c063f274a3d0fd9859e5c183b146

Observation 92c0a982-70b6-4323-94a8-1cc5954d08a9 · outbound

This paper cites Personalized Soups: Personalized Large Language Model Alignment via Post-hoc Parameter Merging.

P-Check: Advancing Personalized Reward Model via Learning to Generate Dynamic Checklist Personalized Soups: Personalized Large Language Model Alignment via Post-hoc Parameter Merging

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-03T12:29:43.847480Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T12:29:43.847480Z digest=sha256:59539c589d7f7fd2f0dd80c46a10aa339aee2a95539eb79a114779bedd0ceb63

Observation e32ca7d8-4c0b-4a10-93a1-8148a62a2950 · outbound

This paper cites Efficient Memory Management for Large Language Model Serving with PagedAttention.

P-Check: Advancing Personalized Reward Model via Learning to Generate Dynamic Checklist Efficient Memory Management for Large Language Model Serving with PagedAttention

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-03T12:29:43.908771Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T12:29:43.908771Z digest=sha256:eb9f1c2c96ef980e62260ed25b008c5e471ded21f2e86fb8fac97903074018fd

Observation a747276c-e1b0-41d0-81a2-c2a8913f7025 · outbound

This paper cites Rubrics as Rewards: Reinforcement Learning Beyond Verifiable Domains.

P-Check: Advancing Personalized Reward Model via Learning to Generate Dynamic Checklist Rubrics as Rewards: Reinforcement Learning Beyond Verifiable Domains

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-03T12:29:43.786052Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T12:29:43.786052Z digest=sha256:05506b09b466623f05a59f69e2c187051930a201955c9fa91ddb977c7d62cc50

Pith citing papers

Observation 2e14b05e-f5bf-4458-ab3e-8169f130eb29 · inbound

Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges cites this paper.

Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges P-Check: Advancing Personalized Reward Model via Learning to Generate Dynamic Checklist

Reference 150

Resolution
verified exact
local_arxiv, observed 2026-05-10T14:00:28.432183Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T13:58:53.430492Z digest=sha256:dcf16cabe8e9a14ff3904e5fa848cac5258ae75623065bdebc857401ba900b85