Pith. sign in

Paper Citation Record · LEDGER

Statistical Rejection Sampling Improves Preference Optimization

As of 15 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 41 inbound Pith citation observations for arXiv:2309.06657.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2309.06657 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 41 of 41 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 41 of 41 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T20:40:33.689723Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T01:17:30.660177Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation c824ec30-1705-4773-92cb-e167098feb8e · inbound

RED: Unleashing Token-Level Rewards from Holistic Feedback via Reward Redistribution cites this paper.

RED: Unleashing Token-Level Rewards from Holistic Feedback via Reward Redistribution Statistical Rejection Sampling Improves Preference Optimization

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-12T21:50:34.729198Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T21:50:34.729198Z digest=sha256:4007d2c094a4c1005f2c162c5ddedcc6294e6ff2cfe48fce014eddd57e3d6e29

Observation c33e5668-b295-4b15-9dba-830b5acee434 · inbound

Enhancing the Reasoning Ability of Multimodal Large Language Models via Mixed Preference Optimization cites this paper.

Enhancing the Reasoning Ability of Multimodal Large Language Models via Mixed Preference Optimization Statistical Rejection Sampling Improves Preference Optimization

Reference 55

Resolution
verified exact
arxiv_id, observed 2026-05-16T09:16:17.289989Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-16T09:16:17.150383Z digest=sha256:14a96d038c1105ed19f03263e0536d4446064da2c07acd691480562f271e62b8

Observation b530734b-bf51-43c2-bb55-b2345dbb6d04 · inbound

BPO: Towards Balanced Preference Optimization between Knowledge Breadth and Depth in Alignment cites this paper.

BPO: Towards Balanced Preference Optimization between Knowledge Breadth and Depth in Alignment Statistical Rejection Sampling Improves Preference Optimization

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-12T19:14:36.991382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:14:36.991382Z digest=sha256:f142c55bf6c79e92f36dc267afe9db6a1a9ce0722bcdaf743c4e347c3ddf6f29

Observation c5b7fa9b-491e-42b9-b450-131bf43bd4af · inbound

Reward Modeling with Ordinal Feedback: Wisdom of the Crowd cites this paper.

Reward Modeling with Ordinal Feedback: Wisdom of the Crowd Statistical Rejection Sampling Improves Preference Optimization

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-12T17:15:47.885158Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T17:15:47.885158Z digest=sha256:6d7b73583e3892faea9ec99e4d312f58513afc5afb0a3ec9ed8b3df65d9c6849

Observation 47aab478-dada-4d51-968b-e8001c8c2eef · inbound

DSTC: Direct Preference Learning with Only Self-Generated Tests and Code to Improve Code LMs cites this paper.

DSTC: Direct Preference Learning with Only Self-Generated Tests and Code to Improve Code LMs Statistical Rejection Sampling Improves Preference Optimization

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-12T17:05:15.145501Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T17:05:15.145501Z digest=sha256:415cb17b195a7dcb47cfbcececfc94ff7660245dd57534a64fb7d487a591c03f

Observation 32855b2e-1994-4440-8ade-2291b1ab1ee9 · inbound

Video-Text Dataset Construction from Multi-AI Feedback: Promoting Weak-to-Strong Preference Learning for Video Large Language Models cites this paper.

Video-Text Dataset Construction from Multi-AI Feedback: Promoting Weak-to-Strong Preference Learning for Video Large Language Models Statistical Rejection Sampling Improves Preference Optimization

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-12T13:31:10.616988Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:31:10.616988Z digest=sha256:8dc0b9629dbbfcada2766eec55a2b8baeb0f5072d7ef4a8b28cc0998dc5ff116

Observation 8e461d49-9e9e-40c5-a340-1ab0792071c5 · inbound

Mind the Gap: Examining the Self-Improvement Capabilities of Large Language Models cites this paper.

Mind the Gap: Examining the Self-Improvement Capabilities of Large Language Models Statistical Rejection Sampling Improves Preference Optimization

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-11T23:15:28.677019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T23:15:28.677019Z digest=sha256:b18220656efee7ebcea40d032ebde1747ee0c3fbcef759cb7c49253240654612

Observation 648433e8-5327-4257-bc40-6a092fb608d9 · inbound

Hybrid Preference Optimization for Alignment: Provably Faster Convergence Rates by Combining Offline Preferences with Online Exploration cites this paper.

Hybrid Preference Optimization for Alignment: Provably Faster Convergence Rates by Combining Offline Preferences with Online Exploration Statistical Rejection Sampling Improves Preference Optimization

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-11T15:55:20.371083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:55:20.371083Z digest=sha256:c04916053adad31a0c8bc801a3fd1eefb2cf02084abd1517d0ea11f12110474b

Observation 6a4a4bff-479e-4ce9-a2de-474314a45475 · inbound

The Superalignment of Superhuman Intelligence with Large Language Models cites this paper.

The Superalignment of Superhuman Intelligence with Large Language Models Statistical Rejection Sampling Improves Preference Optimization

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T15:18:17.773756Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T15:18:17.773756Z digest=sha256:687f798ad8e2a847cdb573d69bf13d3ea71e611e9b9c1bdc72a12c79b2875807

Observation 94e74565-4152-452b-a67b-9fc9f26d2469 · inbound

How to Synthesize Text Data without Model Collapse? cites this paper.

How to Synthesize Text Data without Model Collapse? Statistical Rejection Sampling Improves Preference Optimization

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-11T12:05:48.496034Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T12:05:48.496034Z digest=sha256:84842cc201e5c8df410e24022fca4f7b953eec1fc3467bb2dc84b6f14069dd28

Observation 0ef3e392-0d32-4c33-b50e-afea7fff831b · inbound

CRPO: Confidence-Reward Driven Preference Optimization for Machine Translation cites this paper.

CRPO: Confidence-Reward Driven Preference Optimization for Machine Translation Statistical Rejection Sampling Improves Preference Optimization

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-10T15:33:49.325078Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T15:33:49.325078Z digest=sha256:802609115cd2a102b860913ce122e2c3cf8057cc2f40eadce34a22f3cca11aee

Observation 3d3bb5d8-2270-436a-ab72-33a6e358a940 · inbound

Clear Preferences Leave Traces: Reference Model-Guided Sampling for Preference Learning cites this paper.

Clear Preferences Leave Traces: Reference Model-Guided Sampling for Preference Learning Statistical Rejection Sampling Improves Preference Optimization

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-10T14:39:22.630705Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:39:22.630705Z digest=sha256:3f564d69788480788df9caf20dba7f19cc49f3bc93207b9b4f4cb514d62949a3

Observation 9f60a0c9-e24e-48f1-9307-f3c5565b4fa2 · inbound

Reviving The Classics: Active Reward Modeling in Large Language Model Alignment cites this paper.

Reviving The Classics: Active Reward Modeling in Large Language Model Alignment Statistical Rejection Sampling Improves Preference Optimization

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-09T11:47:17.565775Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T11:47:17.565775Z digest=sha256:9b832577d828d9d09d771e52a75d73ad73285746708565332a45ad9194b8e335

Observation f0aa428f-f6e7-43bd-a7e9-c9c24fce7851 · inbound

CALM: Co-evolution of Algorithms and Language Model for Automatic Heuristic Design cites this paper.

CALM: Co-evolution of Algorithms and Language Model for Automatic Heuristic Design Statistical Rejection Sampling Improves Preference Optimization

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-15T20:40:33.689723Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:40:33.689723Z digest=sha256:4cfd6c01fb4393f6d37228f2108bfe898745ab61f041888ba36ef0f5f116b339

Observation 10740686-15de-458a-8a08-82ef54fc1c73 · inbound

Learning a Pessimistic Reward Model in RLHF cites this paper.

Learning a Pessimistic Reward Model in RLHF Statistical Rejection Sampling Improves Preference Optimization

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T14:01:05.396920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:01:05.396920Z digest=sha256:460ea26b048e48c87d8b3c36a94be3bc96c43a673e1bd399c17d1f08d561980f

Observation 9d2fe920-b1a1-4ac1-acf1-4f5e463db2ba · inbound

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment cites this paper.

Square$\chi$PO: Differentially Private and Robust $\chi^2$-Preference Optimization in Offline Direct Alignment Statistical Rejection Sampling Improves Preference Optimization

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T13:44:57.881471Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:44:57.881471Z digest=sha256:6c180a5d1cd709071a53919ee1aa9b0f0c969c162a734309abe4a9f3f7f22e6e

Observation 41e00320-9fdb-472b-a391-488796e832a0 · inbound

Modeling and Optimizing User Preferences in AI Copilots: A Comprehensive Survey and Taxonomy cites this paper.

Modeling and Optimizing User Preferences in AI Copilots: A Comprehensive Survey and Taxonomy Statistical Rejection Sampling Improves Preference Optimization

Reference 104

Resolution
unresolved
no resolver link, observed 2026-08-07T13:24:30.648329Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:24:30.648329Z digest=sha256:a5810ca6c75b750888a53b829b81df77c18efdd6bb762b6802782f08050c3ffd

Observation 5e9a7c7f-6d8b-48bb-a13c-f3dd10c93cb9 · inbound

Thompson Sampling in Online RLHF with General Function Approximation cites this paper.

Thompson Sampling in Online RLHF with General Function Approximation Statistical Rejection Sampling Improves Preference Optimization

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T12:43:48.431444Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:43:48.431444Z digest=sha256:a5d946e9ad536c490314230aa951124901d2a887e2e1046a1285bd1fea21f5f6

Observation 8d9f8622-e1c9-452a-a71b-742cf211ddd4 · inbound

Smoothed Preference Optimization via ReNoise Inversion for Aligning Diffusion Models with Varied Human Preferences cites this paper.

Smoothed Preference Optimization via ReNoise Inversion for Aligning Diffusion Models with Varied Human Preferences Statistical Rejection Sampling Improves Preference Optimization

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:57.308215Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:26:57.308215Z digest=sha256:0e3b8274d1ac108faebbc7013dbbe9db6075ff9eeba73478b7d1421c677d6c88

Observation 4c055b56-2701-4ae8-8fbb-2fc9eba799b1 · inbound

ReCUT: Balancing Reasoning Length and Accuracy in LLMs via Stepwise Trails and Preference Optimization cites this paper.

ReCUT: Balancing Reasoning Length and Accuracy in LLMs via Stepwise Trails and Preference Optimization Statistical Rejection Sampling Improves Preference Optimization

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T04:24:02.421271Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:24:02.421271Z digest=sha256:30f3ac7abd8b630f27e0bd7358ca99f1cc4a57bf57fff5d7dbcb2e4cd83c5924

Observation d68b4f96-d8a1-4c78-a8f5-74ec1ec203de · inbound

Bridging Offline and Online Reinforcement Learning for LLMs cites this paper.

Bridging Offline and Online Reinforcement Learning for LLMs Statistical Rejection Sampling Improves Preference Optimization

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T22:28:07.163268Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:28:07.163268Z digest=sha256:b2eac8e7a1ee0d9c3e06d182cacf5f59c920d59103a1a258978f1dc0fafecef9

Observation 8345db02-524c-42cc-9e93-99dfba1bf963 · inbound

Inverse Reinforcement Learning Meets Large Language Model Post-Training: Basics, Advances, and Opportunities cites this paper.

Inverse Reinforcement Learning Meets Large Language Model Post-Training: Basics, Advances, and Opportunities Statistical Rejection Sampling Improves Preference Optimization

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-06T16:34:25.092262Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:34:25.092262Z digest=sha256:587e1d1c70cc0c737664b3b1aa7023d66d399aa292fcfde078ddfd8306c75175

Observation 092382b6-2b8c-42d5-9650-e3796259d9ac · inbound

Temporal Self-Rewarding Language Models: Decoupling Chosen-Rejected via Past-Future cites this paper.

Temporal Self-Rewarding Language Models: Decoupling Chosen-Rejected via Past-Future Statistical Rejection Sampling Improves Preference Optimization

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-05T23:06:49.645169Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:06:49.645169Z digest=sha256:e026a775be428d46c6f29a2e74344fb8f9e0ec8131445c172462382cdc72ee37

Observation b8b226be-80c6-4ec8-90d4-8fd3d7beda10 · inbound

PKG-DPO: Optimizing Domain-Specific AI systems with Physics Knowledge Graphs and Direct Preference Optimization cites this paper.

PKG-DPO: Optimizing Domain-Specific AI systems with Physics Knowledge Graphs and Direct Preference Optimization Statistical Rejection Sampling Improves Preference Optimization

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-05T16:30:13.422573Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:30:13.422573Z digest=sha256:9e0055354d25106b2c7150b930dd4985d64a9de8bc60b94b65a1ccbc8c7b76f2

Observation 72979ff6-f349-44ed-a7d6-d56854c786b7 · inbound

Do LLM Modules Generalize? A Study on Motion Generation for Autonomous Driving cites this paper.

Do LLM Modules Generalize? A Study on Motion Generation for Autonomous Driving Statistical Rejection Sampling Improves Preference Optimization

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-05T11:30:39.841578Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T11:30:39.841578Z digest=sha256:a271387477603f509272559aaaa8b2196340f446e063d0997c84fa8cd9a97229

Observation 0ea8dbbf-0c3d-4d0c-b1c7-fb18f039d4e7 · inbound

SHE: Stepwise Hybrid Examination Reinforcement Learning Framework for E-commerce Search Relevance cites this paper.

SHE: Stepwise Hybrid Examination Reinforcement Learning Framework for E-commerce Search Relevance Statistical Rejection Sampling Improves Preference Optimization

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-18T09:36:11.230544Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-18T09:33:57.037949Z digest=sha256:b6a11f8ce8456253b60a92cb87188109a1b01e9d4a49e51428222b6227c33d4c

Observation 0fdf0e4a-ba7e-46c6-9975-7e81f73d13b9 · inbound

Beyond Importance Sampling: Rejection-Gated Policy Optimization cites this paper.

Beyond Importance Sampling: Rejection-Gated Policy Optimization Statistical Rejection Sampling Improves Preference Optimization

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-10T11:30:18.778593Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-10T11:27:31.464362Z digest=sha256:b72187b16e3606b80297aba569653e06046495b45da05cf1e0613470fc628f90

Observation a2a2cfc4-c3ad-45dd-89c0-3094a38f3b29 · inbound

Reasoning Structure Matters for Safety Alignment of Reasoning Models cites this paper.

Reasoning Structure Matters for Safety Alignment of Reasoning Models Statistical Rejection Sampling Improves Preference Optimization

Reference 42

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T12:56:04.115722Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-05-10T02:36:57.093584Z digest=sha256:0fede938aeb79585d37353d08c45ebfd539b262cc9da298b64672a3ce0b4e816

Observation 37f56f6a-d22a-466c-b5aa-bcf3033f0566 · inbound

HiPO: Hierarchical Preference Optimization for Adaptive Reasoning in LLMs cites this paper.

HiPO: Hierarchical Preference Optimization for Adaptive Reasoning in LLMs Statistical Rejection Sampling Improves Preference Optimization

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-10T00:54:48.584130Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-10T00:51:37.506096Z digest=sha256:27f5e2b8e1272ff67fce54499b826a183aad0e4869488ab15ad50d57813b1173

Observation 50a5ae45-660c-4b1b-b769-58288a06f74b · inbound

Supplement Generation Training for Enhancing Agentic Task Performance cites this paper.

Supplement Generation Training for Enhancing Agentic Task Performance Statistical Rejection Sampling Improves Preference Optimization

Reference 46

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T00:59:49.531092Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-05-10T00:58:27.655909Z digest=sha256:e0b5bb135a7fbc264c5b7298176482b487cdb77abcdd11b529f7e3f359898bd2

Observation 8f1a64bc-0b13-4256-9573-312652e6000e · inbound

Efficient Preference Poisoning Attack on Offline RLHF cites this paper.

Efficient Preference Poisoning Attack on Offline RLHF Statistical Rejection Sampling Improves Preference Optimization

Reference 80

Resolution
metadata mismatch
arxiv_id, observed 2026-05-09T05:50:27.173052Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-05-08T19:29:25.000361Z digest=sha256:c8a0a439af3f4415eccb9ed27f3227f91b42ec671734afd09440abce4f37f11f

Observation 359e0d48-73cc-4d37-a616-09c67fa5c344 · inbound

Beyond Negative Rollouts: Positive-Only Policy Optimization with Implicit Negative Gradients cites this paper.

Beyond Negative Rollouts: Positive-Only Policy Optimization with Implicit Negative Gradients Statistical Rejection Sampling Improves Preference Optimization

Reference 25

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T20:16:08.547916Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-05-08T09:53:32.077464Z digest=sha256:7258ba755522758a1919ab2ba6b556c056bd45daeaed9fb6d072a579d57c3cae

Observation 8070190b-b727-4a74-b125-30ae1a237423 · inbound

Response Time Enhances Alignment with Heterogeneous Preferences cites this paper.

Response Time Enhances Alignment with Heterogeneous Preferences Statistical Rejection Sampling Improves Preference Optimization

Reference 166

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T04:45:59.782062Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-05-11T01:04:26.288913Z digest=sha256:de22237da24b088db09558d49556eea45017f310099e6e56907481ce31b275b4

Observation 562b6676-0d06-48a8-beef-5182bb4a7395 · inbound

TokenRatio: Principled Token-Level Preference Optimization via Ratio Matching cites this paper.

TokenRatio: Principled Token-Level Preference Optimization via Ratio Matching Statistical Rejection Sampling Improves Preference Optimization

Reference 112

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T04:57:17.284521Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-05-13T04:55:55.013900Z digest=sha256:9097b3d1cbf57671ce30d43d3282d41e2bfd0f5efdf648ebeabef938a1ae509b

Observation 8df7b4cb-a5d7-49bb-8496-79b847b3a2a6 · inbound

TokenRatio: Principled Token-Level Preference Optimization via Ratio Matching cites this paper.

TokenRatio: Principled Token-Level Preference Optimization via Ratio Matching Statistical Rejection Sampling Improves Preference Optimization

Reference 112

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T05:45:06.667802Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-05-15T05:41:10.714594Z digest=sha256:8e2f877bf4d47e7f01ce85d75c5e36eb432074380e759b8e1cd363aaab09279b

Observation 83ac0f78-88ac-4817-9194-7682d793bd13 · inbound

Gradient-Guided Reward Optimization for Inference-time Alignment cites this paper.

Gradient-Guided Reward Optimization for Inference-time Alignment Statistical Rejection Sampling Improves Preference Optimization

Reference 29

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T01:17:30.661524Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-06-27T16:43:04.984866Z digest=sha256:dadd139ce1a6a8ef3a29b48a9d6450322817ec69ccd00de482157722be791964

Observation 83c5d7a1-e3bf-44bd-afd3-774f0fea94e4 · inbound

BV-Blend: Uncertainty-Weighted Historical Baselines for Stable Critic-Free RL with Verifiable Rewards cites this paper.

BV-Blend: Uncertainty-Weighted Historical Baselines for Stable Critic-Free RL with Verifiable Rewards Statistical Rejection Sampling Improves Preference Optimization

Reference 52

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T12:44:40.158482Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-06-30T10:07:39.554999Z digest=sha256:652db67fb0b1a68e70d92c7ed790e4e5c3d673fb84637fd503aea98e91de46bf

Observation d44c5aba-3f08-4bb6-b0d3-ca20f5df1148 · inbound

Multi-Turn On-Policy Distillation with Prefix Replay cites this paper.

Multi-Turn On-Policy Distillation with Prefix Replay Statistical Rejection Sampling Improves Preference Optimization

Reference 261

Resolution
unresolved
no resolver link, observed 2026-07-11T13:53:36.775836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T13:53:36.775836Z digest=sha256:690fd90f701ded8ba250cfd705e8a98180a55abd623cc966546aa282afc029fb

Observation 8c2d654c-71e1-4dc5-b084-4cc6390e2569 · inbound

Multi-Turn On-Policy Distillation with Prefix Replay cites this paper.

Multi-Turn On-Policy Distillation with Prefix Replay Statistical Rejection Sampling Improves Preference Optimization

Reference 262

Resolution
unresolved
no resolver link, observed 2026-08-02T08:41:02.899448Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T08:41:02.899448Z digest=sha256:215c11f1e635d9e0737d737e7911e5163ace69c4f14745bd148853fb3b89ad79

Observation 4bdf339d-0117-4b99-ad3b-f719cd719c69 · inbound

Test-Time Scaling via Error Localization cites this paper.

Test-Time Scaling via Error Localization Statistical Rejection Sampling Improves Preference Optimization

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-01T07:28:20.944852Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T07:28:20.944852Z digest=sha256:2b6d2d8340cdde406052082e075085618155f6cafc27ea32e7e29ab3b01c05a2

Observation a7c89481-c50e-4fb7-bd88-1b64a93cb4d1 · inbound

Beyond Post-Hoc Temperature Scaling: Bilevel Optimization for LLM Calibration cites this paper.

Beyond Post-Hoc Temperature Scaling: Bilevel Optimization for LLM Calibration Statistical Rejection Sampling Improves Preference Optimization

Reference 164

Resolution
unresolved
no resolver link, observed 2026-08-15T14:33:57.530737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T14:33:57.530737Z digest=sha256:5eb72e66aa94df8ba617cc27c653ec9df301707d88c27d20aee491e745df71a2