Pith. sign in

Paper Citation Record · LEDGER

West-of-N: Synthetic Preferences for Self-Improving Reward Models

As of 22 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 11 inbound Pith citation observations for arXiv:2401.12086.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2401.12086 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 11 of 11 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 11 of 11 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T00:21:59.244898Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T13:20:39.606465Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation c40431ed-2060-4019-bc40-aef5516d718d · inbound

Piecing It All Together: Verifying Multi-Hop Multimodal Claims cites this paper.

Piecing It All Together: Verifying Multi-Hop Multimodal Claims West-of-N: Synthetic Preferences for Self-Improving Reward Models

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-12T20:34:03.843013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T20:34:03.843013Z digest=sha256:0254a905182b2e114ec55bb9479914eb9a005749daab79880406a5bc918231a2

Observation f83f81e4-53fb-4b58-bc24-dc87258681b0 · inbound

Self-Generated Critiques Boost Reward Modeling for Language Models cites this paper.

Self-Generated Critiques Boost Reward Modeling for Language Models West-of-N: Synthetic Preferences for Self-Improving Reward Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-12T12:58:30.027648Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T12:58:30.027648Z digest=sha256:81067cef8b0f95d0a9232e354a235ddcdc5c50d4ffc948795044dc9938948d48

Observation 5a8d4734-9aa9-4cac-9a4e-e10413807cee · inbound

Self-Improvement in Language Models: The Sharpening Mechanism cites this paper.

Self-Improvement in Language Models: The Sharpening Mechanism West-of-N: Synthetic Preferences for Self-Improving Reward Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-12T00:11:14.744562Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T00:11:14.744562Z digest=sha256:7da9117620a4ec4c95b8056d14274c48ec41b1a663b6364b971bd37cc4300974

Observation cbaa0a07-b350-4c6d-8652-9aa6c7fa4965 · inbound

Surveying the Effects of Quality, Diversity, and Complexity in Synthetic Data From Large Language Models cites this paper.

Surveying the Effects of Quality, Diversity, and Complexity in Synthetic Data From Large Language Models West-of-N: Synthetic Preferences for Self-Improving Reward Models

Reference 145

Resolution
unresolved
no resolver link, observed 2026-08-11T22:57:01.895819Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T22:57:01.895819Z digest=sha256:940ffcf37cdd853f72c07213fb20aea5867be803acc1bea31d5c28a6a729f9a3

Observation b00e01a3-5dd9-49aa-8af0-8751ab8a689b · inbound

R.I.P.: Better Models by Survival of the Fittest Prompts cites this paper.

R.I.P.: Better Models by Survival of the Fittest Prompts West-of-N: Synthetic Preferences for Self-Improving Reward Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-09T23:02:23.371686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T23:02:23.371686Z digest=sha256:c5a7e250167c280cb7956127071b466096d15cbac1744e203f14d6cb916ed1a5

Observation 41958aec-42c1-4d19-abc4-64c44ef6fdf9 · inbound

RICo: Refined In-Context Contribution for Automatic Instruction-Tuning Data Selection cites this paper.

RICo: Refined In-Context Contribution for Automatic Instruction-Tuning Data Selection West-of-N: Synthetic Preferences for Self-Improving Reward Models

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-15T23:12:38.970109Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T23:12:38.970109Z digest=sha256:8a2b917803e05384738bc1230b651df363608d5796d9d41e2370f40362bb4efa

Observation 50a0298e-512c-4696-87af-53bfcd238469 · inbound

Mutual-Taught for Co-adapting Policy and Reward Models cites this paper.

Mutual-Taught for Co-adapting Policy and Reward Models West-of-N: Synthetic Preferences for Self-Improving Reward Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-15T20:53:39.688797Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:53:39.688797Z digest=sha256:0d9a1e058114d8fc08983e8733bbd8e532798711b288e86098226e1ab4c56bae

Observation 139a865c-3c99-47d3-bdc5-87120610f228 · inbound

Enterprise Large Language Model Evaluation Benchmark cites this paper.

Enterprise Large Language Model Evaluation Benchmark West-of-N: Synthetic Preferences for Self-Improving Reward Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T22:56:33.472641Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:56:33.472641Z digest=sha256:fabcfefbd8019fceaa649d78c732b2d303aa4b452050c81fe264436771bfd74b

Observation d924b7eb-7f25-42fe-836d-3b7a7fe23925 · inbound

Bridging Offline and Online Reinforcement Learning for LLMs cites this paper.

Bridging Offline and Online Reinforcement Learning for LLMs West-of-N: Synthetic Preferences for Self-Improving Reward Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T22:28:07.646010Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:28:07.646010Z digest=sha256:19aed2fa04141c80d9891f1f324f5c04282cae455b305c4b2a608b6a15834059

Observation b5cc09f4-cc04-4bc2-9819-1bbb87653000 · inbound

A Survey on Open Dataset Search in the LLM Era: Retrospectives and Perspectives cites this paper.

A Survey on Open Dataset Search in the LLM Era: Retrospectives and Perspectives West-of-N: Synthetic Preferences for Self-Improving Reward Models

Reference 115

Resolution
verified exact
local_arxiv, observed 2026-08-05T13:20:39.610566Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-05T13:20:39.171623Z digest=sha256:a54b25091387eaf26c43e717e9edfa15ef77c2e49b2d5dd472d8493eb0be393d

Observation f2bc5e41-e5d7-48bc-b5c2-c480fdcea8f8 · inbound

Preference Tree Optimization: Enhancing Goal-Oriented Dialogue with Look-Ahead Simulations cites this paper.

Preference Tree Optimization: Enhancing Goal-Oriented Dialogue with Look-Ahead Simulations West-of-N: Synthetic Preferences for Self-Improving Reward Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-16T00:21:59.244898Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-16T00:21:59.244898Z digest=sha256:07b44cb970344f0e6669a896386a5568a88e14ee3917a44c867d6f5e599e0fc0