Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 22 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 11 inbound Pith citation observations for arXiv:2401.12086.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-16T00:21:59.244898Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-05T13:20:39.606465Z
0 of 0 outbound references displayed
External citation measurements
No source-named external measurement is stored.
No outbound reference observations are available for this paper version.
Observation c40431ed-2060-4019-bc40-aef5516d718d · inbound
Piecing It All Together: Verifying Multi-Hop Multimodal Claims West-of-N: Synthetic Preferences for Self-Improving Reward Models
Reference 2022
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f83f81e4-53fb-4b58-bc24-dc87258681b0 · inbound
Self-Generated Critiques Boost Reward Modeling for Language Models West-of-N: Synthetic Preferences for Self-Improving Reward Models
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5a8d4734-9aa9-4cac-9a4e-e10413807cee · inbound
Self-Improvement in Language Models: The Sharpening Mechanism West-of-N: Synthetic Preferences for Self-Improving Reward Models
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cbaa0a07-b350-4c6d-8652-9aa6c7fa4965 · inbound
Surveying the Effects of Quality, Diversity, and Complexity in Synthetic Data From Large Language Models West-of-N: Synthetic Preferences for Self-Improving Reward Models
Reference 145
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b00e01a3-5dd9-49aa-8af0-8751ab8a689b · inbound
R.I.P.: Better Models by Survival of the Fittest Prompts West-of-N: Synthetic Preferences for Self-Improving Reward Models
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 41958aec-42c1-4d19-abc4-64c44ef6fdf9 · inbound
RICo: Refined In-Context Contribution for Automatic Instruction-Tuning Data Selection West-of-N: Synthetic Preferences for Self-Improving Reward Models
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 50a0298e-512c-4696-87af-53bfcd238469 · inbound
Mutual-Taught for Co-adapting Policy and Reward Models West-of-N: Synthetic Preferences for Self-Improving Reward Models
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 139a865c-3c99-47d3-bdc5-87120610f228 · inbound
Enterprise Large Language Model Evaluation Benchmark West-of-N: Synthetic Preferences for Self-Improving Reward Models
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d924b7eb-7f25-42fe-836d-3b7a7fe23925 · inbound
Bridging Offline and Online Reinforcement Learning for LLMs West-of-N: Synthetic Preferences for Self-Improving Reward Models
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b5cc09f4-cc04-4bc2-9819-1bbb87653000 · inbound
A Survey on Open Dataset Search in the LLM Era: Retrospectives and Perspectives West-of-N: Synthetic Preferences for Self-Improving Reward Models
Reference 115
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation f2bc5e41-e5d7-48bc-b5c2-c480fdcea8f8 · inbound
Preference Tree Optimization: Enhancing Goal-Oriented Dialogue with Look-Ahead Simulations West-of-N: Synthetic Preferences for Self-Improving Reward Models
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.