Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-10T23:21:34.194345Z
Paper Citation Record · LEDGER
As of 16 August 2026, this Paper Citation Record lists 26 of 26 outbound references and 1 inbound Pith citation observation for arXiv:2412.20996.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-10T23:21:34.194345Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T10:54:07.795586Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-07T10:54:08.296887Z
26 of 26 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 3564f12d-8753-4209-a3c5-1e83fba7a1ce · outbound
Plug-and-Play Training Framework for Preference Optimization Unresolved cited work
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation a6b4d70d-78c5-4441-b785-8677fbc71a73 · outbound
Plug-and-Play Training Framework for Preference Optimization Unresolved cited work
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 15a1c3a6-f2bf-446e-bd22-95f5aae24996 · outbound
Plug-and-Play Training Framework for Preference Optimization Large Language Monkeys: Scaling Inference Compute with Repeated Sampling
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 59b29907-41e9-47b2-8867-0504f47bb51d · outbound
Plug-and-Play Training Framework for Preference Optimization Evaluating Large Language Models Trained on Code
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e6eb9c67-cb41-4bb6-823e-2f022e4af616 · outbound
Plug-and-Play Training Framework for Preference Optimization Christiano, Jan Leike, Tom B
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bda95619-6b25-49b0-93fc-dcf667f81195 · outbound
Plug-and-Play Training Framework for Preference Optimization Training Verifiers to Solve Math Word Problems
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b9db6b3c-c69a-4f0e-bb39-3694e933b3f3 · outbound
Plug-and-Play Training Framework for Preference Optimization Unresolved cited work
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4a9fd596-47fd-4510-95c5-ac298639d5bf · outbound
Plug-and-Play Training Framework for Preference Optimization ChatGLM: A Family of Large Language Models from GLM-130B to GLM-4 All Tools
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 14bdd9a8-efc5-43c5-a674-38c1d6e609b8 · outbound
Plug-and-Play Training Framework for Preference Optimization Unresolved cited work
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 52a55dc2-aa65-4924-be4b-b5b8ef5595ed · outbound
Plug-and-Play Training Framework for Preference Optimization Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen - Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ca2d97d6-65c3-449a-a8e3-ee73b60c5396 · outbound
Plug-and-Play Training Framework for Preference Optimization Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a17fe85c-4a1b-49b8-a4d2-e93bef8b1e43 · outbound
Plug-and-Play Training Framework for Preference Optimization Hashimoto
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3bb3c2b1-e36b-4716-b104-ac03abfa729e · outbound
Plug-and-Play Training Framework for Preference Optimization Unresolved cited work
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b1241d34-40ee-45e3-9f0d-0d6b37bec68d · outbound
Plug-and-Play Training Framework for Preference Optimization Unresolved cited work
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 3e8f2b64-e06a-4f87-b95e-65de95d3453b · outbound
Plug-and-Play Training Framework for Preference Optimization SimPO: Simple Preference Optimization with a Reference-Free Reward
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9bde83f2-9ad4-4195-88e5-facb4cf8b364 · outbound
Plug-and-Play Training Framework for Preference Optimization Smaug: Fixing Failure Modes of Preference Optimisation with DPO-Positive
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9bc54500-0cbe-4546-a9e8-2b0c65b3e808 · outbound
Plug-and-Play Training Framework for Preference Optimization Advantage-Weighted Regression: Simple and Scalable Off-Policy Reinforcement Learning
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 49b7ebf0-4907-404b-a9a7-618928c2db44 · outbound
Plug-and-Play Training Framework for Preference Optimization Unresolved cited work
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 19657ecd-2bd3-40df-b998-4af41fcecbe5 · outbound
Plug-and-Play Training Framework for Preference Optimization Manning, Stefano Ermon, and Chelsea Finn
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9a3acaa4-fc68-48dc-8028-907020d36a94 · outbound
Plug-and-Play Training Framework for Preference Optimization Proximal Policy Optimization Algorithms
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a0ab056c-8d66-43d0-a617-e3feb2b29c60 · outbound
Plug-and-Play Training Framework for Preference Optimization Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 405f1f26-4d29-43a2-a344-a82199705812 · outbound
Plug-and-Play Training Framework for Preference Optimization Le, Ed H
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a6e3c076-9dc9-4fa4-a83a-532397b0be8f · outbound
Plug-and-Play Training Framework for Preference Optimization Qwen2 Technical Report
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4fb18886-77a4-47b7-a14e-9cf8ea65bc92 · outbound
Plug-and-Play Training Framework for Preference Optimization Kwok, Zhenguo Li, Adrian Weller, and Weiyang Liu
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0cea60b1-0e76-4b86-acf1-d5eba7e78544 · outbound
Plug-and-Play Training Framework for Preference Optimization online" 'onlinestring :=
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dec7390c-6d40-4df8-85e1-03edd2e6e097 · outbound
Plug-and-Play Training Framework for Preference Optimization write newline
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1d65934a-13c2-4695-8d3f-7e4726f1ae05 · inbound
Progressive Mastery: Customized Curriculum Learning with Guided Prompting for Mathematical Reasoning Plug-and-Play Training Framework for Preference Optimization
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.