Pith. sign in

Paper Citation Record · LEDGER

WorldPM: Scaling Human Preference Modeling

As of 15 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 9 inbound Pith citation observations for arXiv:2505.10527.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.10527 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 9 of 9 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 9 of 9 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T05:49:45.723481Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T00:07:27.940098Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation c6887f61-aa82-4a10-9f0e-16a07b78309b · inbound

RewardDance: Reward Scaling in Visual Generation cites this paper.

RewardDance: Reward Scaling in Visual Generation WorldPM: Scaling Human Preference Modeling

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-04T20:09:00.504202Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:09:00.504202Z digest=sha256:6d81183534da0c748b33caa527a0881e9a050f0990fb7d78be2b1c1a992c07e3

Observation f423c084-1aae-4047-82a9-ee19963d9dc0 · inbound

AI Can Learn Scientific Taste cites this paper.

AI Can Learn Scientific Taste WorldPM: Scaling Human Preference Modeling

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-02T18:14:49.961366Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T18:14:49.961366Z digest=sha256:8616f71e2839179557ed12096168add5d202ccb380bdc31fac4795bc37442887

Observation c7f984f2-8052-45cb-abbe-747d8cfe1a63 · inbound

Beyond Overlap Metrics: Rewarding Reasoning and Preferences for Faithful Multi-Role Dialogue Summarization cites this paper.

Beyond Overlap Metrics: Rewarding Reasoning and Preferences for Faithful Multi-Role Dialogue Summarization WorldPM: Scaling Human Preference Modeling

Reference 9

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T06:51:46.328497Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-05-10T06:47:34.623340Z digest=sha256:0da9b93d9c9e6276dcaa77d193eb7ba4972ce9b82a533d0193c88000c826c101

Observation d8ebd339-9560-40c2-9808-0fb0e0b52c20 · inbound

Leveraging Verifier-Based Reinforcement Learning in Image Editing cites this paper.

Leveraging Verifier-Based Reinforcement Learning in Image Editing WorldPM: Scaling Human Preference Modeling

Reference 55

Resolution
verified exact
arxiv_id, observed 2026-05-12T10:06:27.446483Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-07T08:00:33.307429Z digest=sha256:ec8a423aed242645d41d6e8914cda2c2d8bd8e0fb3dbb2089ad28dabb4280681

Observation b54261bb-8481-4f67-9bc9-04a6944e3572 · inbound

Leveraging Verifier-Based Reinforcement Learning in Image Editing cites this paper.

Leveraging Verifier-Based Reinforcement Learning in Image Editing WorldPM: Scaling Human Preference Modeling

Reference 55

Resolution
verified exact
arxiv_id, observed 2026-05-21T09:14:05.932676Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-21T09:11:02.183133Z digest=sha256:a35d0cf64fa542dfed119402761d880ab7fefcac588f905f3a66f9fc92354310

Observation 2619ad0c-fc3c-45c9-8374-4ed6ab95b580 · inbound

RewardHarness: Self-Evolving Agentic Post-Training cites this paper.

RewardHarness: Self-Evolving Agentic Post-Training WorldPM: Scaling Human Preference Modeling

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-12T08:31:23.995255Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-12T01:05:25.567581Z digest=sha256:f54921804c4befb971f6a6c59ffd9c9e03d207ca2a6277ae971e50e7576c733f

Observation 7808e38c-29ba-4e74-a93b-4aad4a8ebd20 · inbound

Z-Reward: Beyond Scalar Rewards by Internalizing Reasoning into Score Distributions cites this paper.

Z-Reward: Beyond Scalar Rewards by Internalizing Reasoning into Score Distributions WorldPM: Scaling Human Preference Modeling

Reference 54

Resolution
verified exact
arxiv_id, observed 2026-07-03T00:07:27.941638Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-27T17:30:57.001021Z digest=sha256:993109776952fe52d21071f1241179a72bb8342cef5323be4c92b826dd05667e

Observation 2d25d118-c7e0-45f1-8ef5-4658c4a9af16 · inbound

Z-Reward: Beyond Scalar Rewards by Internalizing Reasoning into Score Distributions cites this paper.

Z-Reward: Beyond Scalar Rewards by Internalizing Reasoning into Score Distributions WorldPM: Scaling Human Preference Modeling

Reference 54

Resolution
unresolved
no resolver link, observed 2026-07-15T10:53:37.186361Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T10:53:37.186361Z digest=sha256:2c353402d57165fbcb97f57f32fea24d60bef81ed1079e81cc7c12d295d63d83

Observation e834e349-dfe1-4c35-b6a7-699bc4d1b514 · inbound

RRC: Unlocking Generative Reward Models in LLM Reinforcement Learning via Ranking-Based Reward Construction cites this paper.

RRC: Unlocking Generative Reward Models in LLM Reinforcement Learning via Ranking-Based Reward Construction WorldPM: Scaling Human Preference Modeling

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:45.723481Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:49:45.723481Z digest=sha256:4c708b81bc20ec5398993fb93b41cf1744dfefa0766975c84dfa3e61d69e7f84