Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 15 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 41 inbound Pith citation observations for arXiv:2309.06657.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-15T20:40:33.689723Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-03T01:17:30.660177Z
0 of 0 outbound references displayed
External citation measurements
No source-named external measurement is stored.
No outbound reference observations are available for this paper version.
Observation c824ec30-1705-4773-92cb-e167098feb8e · inbound
RED: Unleashing Token-Level Rewards from Holistic Feedback via Reward Redistribution Statistical Rejection Sampling Improves Preference Optimization
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c33e5668-b295-4b15-9dba-830b5acee434 · inbound
Enhancing the Reasoning Ability of Multimodal Large Language Models via Mixed Preference Optimization Statistical Rejection Sampling Improves Preference Optimization
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation b530734b-bf51-43c2-bb55-b2345dbb6d04 · inbound
BPO: Towards Balanced Preference Optimization between Knowledge Breadth and Depth in Alignment Statistical Rejection Sampling Improves Preference Optimization
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c5b7fa9b-491e-42b9-b450-131bf43bd4af · inbound
Reward Modeling with Ordinal Feedback: Wisdom of the Crowd Statistical Rejection Sampling Improves Preference Optimization
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 47aab478-dada-4d51-968b-e8001c8c2eef · inbound
DSTC: Direct Preference Learning with Only Self-Generated Tests and Code to Improve Code LMs Statistical Rejection Sampling Improves Preference Optimization
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 32855b2e-1994-4440-8ade-2291b1ab1ee9 · inbound
Video-Text Dataset Construction from Multi-AI Feedback: Promoting Weak-to-Strong Preference Learning for Video Large Language Models Statistical Rejection Sampling Improves Preference Optimization
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8e461d49-9e9e-40c5-a340-1ab0792071c5 · inbound
Mind the Gap: Examining the Self-Improvement Capabilities of Large Language Models Statistical Rejection Sampling Improves Preference Optimization
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 648433e8-5327-4257-bc40-6a092fb608d9 · inbound
Hybrid Preference Optimization for Alignment: Provably Faster Convergence Rates by Combining Offline Preferences with Online Exploration Statistical Rejection Sampling Improves Preference Optimization
Reference 2022
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6a4a4bff-479e-4ce9-a2de-474314a45475 · inbound
The Superalignment of Superhuman Intelligence with Large Language Models Statistical Rejection Sampling Improves Preference Optimization
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 94e74565-4152-452b-a67b-9fc9f26d2469 · inbound
How to Synthesize Text Data without Model Collapse? Statistical Rejection Sampling Improves Preference Optimization
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0ef3e392-0d32-4c33-b50e-afea7fff831b · inbound
CRPO: Confidence-Reward Driven Preference Optimization for Machine Translation Statistical Rejection Sampling Improves Preference Optimization
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3d3bb5d8-2270-436a-ab72-33a6e358a940 · inbound
Clear Preferences Leave Traces: Reference Model-Guided Sampling for Preference Learning Statistical Rejection Sampling Improves Preference Optimization
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9f60a0c9-e24e-48f1-9307-f3c5565b4fa2 · inbound
Reviving The Classics: Active Reward Modeling in Large Language Model Alignment Statistical Rejection Sampling Improves Preference Optimization
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f0aa428f-f6e7-43bd-a7e9-c9c24fce7851 · inbound
CALM: Co-evolution of Algorithms and Language Model for Automatic Heuristic Design Statistical Rejection Sampling Improves Preference Optimization
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 10740686-15de-458a-8a08-82ef54fc1c73 · inbound
Learning a Pessimistic Reward Model in RLHF Statistical Rejection Sampling Improves Preference Optimization
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9d2fe920-b1a1-4ac1-acf1-4f5e463db2ba · inbound
Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Statistical Rejection Sampling Improves Preference Optimization
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 41e00320-9fdb-472b-a391-488796e832a0 · inbound
Modeling and Optimizing User Preferences in AI Copilots: A Comprehensive Survey and Taxonomy Statistical Rejection Sampling Improves Preference Optimization
Reference 104
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5e9a7c7f-6d8b-48bb-a13c-f3dd10c93cb9 · inbound
Thompson Sampling in Online RLHF with General Function Approximation Statistical Rejection Sampling Improves Preference Optimization
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8d9f8622-e1c9-452a-a71b-742cf211ddd4 · inbound
Smoothed Preference Optimization via ReNoise Inversion for Aligning Diffusion Models with Varied Human Preferences Statistical Rejection Sampling Improves Preference Optimization
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4c055b56-2701-4ae8-8fbb-2fc9eba799b1 · inbound
ReCUT: Balancing Reasoning Length and Accuracy in LLMs via Stepwise Trails and Preference Optimization Statistical Rejection Sampling Improves Preference Optimization
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d68b4f96-d8a1-4c78-a8f5-74ec1ec203de · inbound
Bridging Offline and Online Reinforcement Learning for LLMs Statistical Rejection Sampling Improves Preference Optimization
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8345db02-524c-42cc-9e93-99dfba1bf963 · inbound
Inverse Reinforcement Learning Meets Large Language Model Post-Training: Basics, Advances, and Opportunities Statistical Rejection Sampling Improves Preference Optimization
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 092382b6-2b8c-42d5-9650-e3796259d9ac · inbound
Temporal Self-Rewarding Language Models: Decoupling Chosen-Rejected via Past-Future Statistical Rejection Sampling Improves Preference Optimization
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b8b226be-80c6-4ec8-90d4-8fd3d7beda10 · inbound
PKG-DPO: Optimizing Domain-Specific AI systems with Physics Knowledge Graphs and Direct Preference Optimization Statistical Rejection Sampling Improves Preference Optimization
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 72979ff6-f349-44ed-a7d6-d56854c786b7 · inbound
Do LLM Modules Generalize? A Study on Motion Generation for Autonomous Driving Statistical Rejection Sampling Improves Preference Optimization
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0ea8dbbf-0c3d-4d0c-b1c7-fb18f039d4e7 · inbound
SHE: Stepwise Hybrid Examination Reinforcement Learning Framework for E-commerce Search Relevance Statistical Rejection Sampling Improves Preference Optimization
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 0fdf0e4a-ba7e-46c6-9975-7e81f73d13b9 · inbound
Beyond Importance Sampling: Rejection-Gated Policy Optimization Statistical Rejection Sampling Improves Preference Optimization
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation a2a2cfc4-c3ad-45dd-89c0-3094a38f3b29 · inbound
Reasoning Structure Matters for Safety Alignment of Reasoning Models Statistical Rejection Sampling Improves Preference Optimization
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 37f56f6a-d22a-466c-b5aa-bcf3033f0566 · inbound
HiPO: Hierarchical Preference Optimization for Adaptive Reasoning in LLMs Statistical Rejection Sampling Improves Preference Optimization
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 50a5ae45-660c-4b1b-b769-58288a06f74b · inbound
Supplement Generation Training for Enhancing Agentic Task Performance Statistical Rejection Sampling Improves Preference Optimization
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 8f1a64bc-0b13-4256-9573-312652e6000e · inbound
Efficient Preference Poisoning Attack on Offline RLHF Statistical Rejection Sampling Improves Preference Optimization
Reference 80
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 359e0d48-73cc-4d37-a616-09c67fa5c344 · inbound
Beyond Negative Rollouts: Positive-Only Policy Optimization with Implicit Negative Gradients Statistical Rejection Sampling Improves Preference Optimization
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 8070190b-b727-4a74-b125-30ae1a237423 · inbound
Response Time Enhances Alignment with Heterogeneous Preferences Statistical Rejection Sampling Improves Preference Optimization
Reference 166
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 562b6676-0d06-48a8-beef-5182bb4a7395 · inbound
TokenRatio: Principled Token-Level Preference Optimization via Ratio Matching Statistical Rejection Sampling Improves Preference Optimization
Reference 112
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 8df7b4cb-a5d7-49bb-8496-79b847b3a2a6 · inbound
TokenRatio: Principled Token-Level Preference Optimization via Ratio Matching Statistical Rejection Sampling Improves Preference Optimization
Reference 112
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 83ac0f78-88ac-4817-9194-7682d793bd13 · inbound
Gradient-Guided Reward Optimization for Inference-time Alignment Statistical Rejection Sampling Improves Preference Optimization
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 83c5d7a1-e3bf-44bd-afd3-774f0fea94e4 · inbound
BV-Blend: Uncertainty-Weighted Historical Baselines for Stable Critic-Free RL with Verifiable Rewards Statistical Rejection Sampling Improves Preference Optimization
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation d44c5aba-3f08-4bb6-b0d3-ca20f5df1148 · inbound
Multi-Turn On-Policy Distillation with Prefix Replay Statistical Rejection Sampling Improves Preference Optimization
Reference 261
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8c2d654c-71e1-4dc5-b084-4cc6390e2569 · inbound
Multi-Turn On-Policy Distillation with Prefix Replay Statistical Rejection Sampling Improves Preference Optimization
Reference 262
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4bdf339d-0117-4b99-ad3b-f719cd719c69 · inbound
Test-Time Scaling via Error Localization Statistical Rejection Sampling Improves Preference Optimization
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a7c89481-c50e-4fb7-bd88-1b64a93cb4d1 · inbound
Beyond Post-Hoc Temperature Scaling: Bilevel Optimization for LLM Calibration Statistical Rejection Sampling Improves Preference Optimization
Reference 164
Source-reported events for the cited work
Unavailable: canonical work link unavailable.