Pith. sign in

Paper Citation Record · LEDGER

Plug-and-Play Training Framework for Preference Optimization

As of 16 August 2026, this Paper Citation Record lists 26 of 26 outbound references and 1 inbound Pith citation observation for arXiv:2412.20996.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.20996 v1

Coverage vector

measured 26 of 26 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T23:21:34.194345Z

measured 27 of 27 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T10:54:07.795586Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-07T10:54:08.296887Z

Reference resolution

26 of 26 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved26
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 3564f12d-8753-4209-a3c5-1e83fba7a1ce · outbound

This paper cites an unresolved cited work.

Plug-and-Play Training Framework for Preference Optimization Unresolved cited work

Reference 1

Resolution
unresolved
raw_fallback, observed 2026-08-10T23:21:34.589737Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-10T23:21:34.044741Z digest=sha256:37afd2586916ac067fb344caa2b60e932232bc9e5f86e0fcf0a22abd06bd7390

Observation a6b4d70d-78c5-4441-b785-8677fbc71a73 · outbound

This paper cites an unresolved cited work.

Plug-and-Play Training Framework for Preference Optimization Unresolved cited work

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:34.052324Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:34.052324Z digest=sha256:e4406398731d8aa1e207b0de55d6b0272706d7c3adf60ca32e6279a56d5edaf4

Observation 15a1c3a6-f2bf-446e-bd22-95f5aae24996 · outbound

This paper cites Large Language Monkeys: Scaling Inference Compute with Repeated Sampling.

Plug-and-Play Training Framework for Preference Optimization Large Language Monkeys: Scaling Inference Compute with Repeated Sampling

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:34.059249Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:34.059249Z digest=sha256:9eb5210d7508fbc48436b45e3b4bf56781f3f4662369403fdc92365cc41bf40e

Observation 59b29907-41e9-47b2-8867-0504f47bb51d · outbound

This paper cites Evaluating Large Language Models Trained on Code.

Plug-and-Play Training Framework for Preference Optimization Evaluating Large Language Models Trained on Code

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:34.066593Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:34.066593Z digest=sha256:0d80c958e6ec0a4ebf481afc1216f488fb8d16a6e7ff9af400a23e4a0abbc055

Observation e6eb9c67-cb41-4bb6-823e-2f022e4af616 · outbound

This paper cites Christiano, Jan Leike, Tom B.

Plug-and-Play Training Framework for Preference Optimization Christiano, Jan Leike, Tom B

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:34.072127Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:34.072127Z digest=sha256:38127f5a0cec2c02ee4433cd27efb8fa6e5e76eca81a70f4b890091f23c31978

Observation bda95619-6b25-49b0-93fc-dcf667f81195 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Plug-and-Play Training Framework for Preference Optimization Training Verifiers to Solve Math Word Problems

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:34.077261Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:34.077261Z digest=sha256:2f56f5ea1276abc0735fe65443c5386079dfd4941110d09082f2302459783fce

Observation b9db6b3c-c69a-4f0e-bb39-3694e933b3f3 · outbound

This paper cites an unresolved cited work.

Plug-and-Play Training Framework for Preference Optimization Unresolved cited work

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:34.084160Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:34.084160Z digest=sha256:bf4c11fbc169d386c977aefd78353104b11898ed35fff1f174176bcdab7f32bc

Observation 4a9fd596-47fd-4510-95c5-ac298639d5bf · outbound

This paper cites ChatGLM: A Family of Large Language Models from GLM-130B to GLM-4 All Tools.

Plug-and-Play Training Framework for Preference Optimization ChatGLM: A Family of Large Language Models from GLM-130B to GLM-4 All Tools

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:34.090842Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:34.090842Z digest=sha256:c74c058bceffe6763000aedd1ed7fb6c7a2e19a4e4486ef3fa07edc6cc9c01e8

Observation 14bdd9a8-efc5-43c5-a674-38c1d6e609b8 · outbound

This paper cites an unresolved cited work.

Plug-and-Play Training Framework for Preference Optimization Unresolved cited work

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:34.097336Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:34.097336Z digest=sha256:4706785659ac5850229f522a60d3ec174e0ad58be1b2c9cdb9ccb1afb2a74122

Observation 52a55dc2-aa65-4924-be4b-b5b8ef5595ed · outbound

This paper cites Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen - Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen.

Plug-and-Play Training Framework for Preference Optimization Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen - Zhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:34.102620Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:34.102620Z digest=sha256:663e3c112e9c46f526b7653969a3e6255840690979a49d700381933d3f5809f6

Observation ca2d97d6-65c3-449a-a8e3-ee73b60c5396 · outbound

This paper cites Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs.

Plug-and-Play Training Framework for Preference Optimization Step-DPO: Step-wise Preference Optimization for Long-chain Reasoning of LLMs

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:34.108960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:34.108960Z digest=sha256:38066dbd99476c4b5ebb8cf2736a5aacffa7e687acfb0c0f1714839cd0ae7446

Observation a17fe85c-4a1b-49b8-a4d2-e93bef8b1e43 · outbound

This paper cites Hashimoto.

Plug-and-Play Training Framework for Preference Optimization Hashimoto

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:34.115144Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:34.115144Z digest=sha256:8a0ccf823775a0b26e348b28a92dafa6045d6207fd7104e7b7bc3a7b4b7ea540

Observation 3bb3c2b1-e36b-4716-b104-ac03abfa729e · outbound

This paper cites an unresolved cited work.

Plug-and-Play Training Framework for Preference Optimization Unresolved cited work

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:34.120444Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:34.120444Z digest=sha256:21adcc0e3c6857799fcfb4f3635ded446a091e50ebeaf753fe8e2fea595ef5c3

Observation b1241d34-40ee-45e3-9f0d-0d6b37bec68d · outbound

This paper cites an unresolved cited work.

Plug-and-Play Training Framework for Preference Optimization Unresolved cited work

Reference 14

Resolution
unresolved
raw_fallback, observed 2026-08-10T23:21:34.509937Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-10T23:21:34.125691Z digest=sha256:0a75c94a095ffb235d6b0802e6dadd8b9652c9908aba851a574af8d2b6be62fe

Observation 3e8f2b64-e06a-4f87-b95e-65de95d3453b · outbound

This paper cites SimPO: Simple Preference Optimization with a Reference-Free Reward.

Plug-and-Play Training Framework for Preference Optimization SimPO: Simple Preference Optimization with a Reference-Free Reward

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:34.130362Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:34.130362Z digest=sha256:93a9b90f3cb44ca57cab8fa8d0615a852404a2ba7c1b857f20de7e1bee18008a

Observation 9bde83f2-9ad4-4195-88e5-facb4cf8b364 · outbound

This paper cites Smaug: Fixing Failure Modes of Preference Optimisation with DPO-Positive.

Plug-and-Play Training Framework for Preference Optimization Smaug: Fixing Failure Modes of Preference Optimisation with DPO-Positive

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:34.136271Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:34.136271Z digest=sha256:f9b3e06d724d2d4cf910664e2994efde1c6088c9a97ce20e8ef04d9ff3f5fd71

Observation 9bc54500-0cbe-4546-a9e8-2b0c65b3e808 · outbound

This paper cites Advantage-Weighted Regression: Simple and Scalable Off-Policy Reinforcement Learning.

Plug-and-Play Training Framework for Preference Optimization Advantage-Weighted Regression: Simple and Scalable Off-Policy Reinforcement Learning

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:34.141759Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:34.141759Z digest=sha256:817984937701bf78b63d8467eaff49bf07e8ba3780f0516342375f1777ec1626

Observation 49b7ebf0-4907-404b-a9a7-618928c2db44 · outbound

This paper cites an unresolved cited work.

Plug-and-Play Training Framework for Preference Optimization Unresolved cited work

Reference 18

Resolution
unresolved
raw_fallback, observed 2026-08-10T23:21:34.493314Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-10T23:21:34.146667Z digest=sha256:deee9e19c37d25ffbd79b633f76ce5f6ee85b3f564bd7e02a20155a4472998de

Observation 19657ecd-2bd3-40df-b998-4af41fcecbe5 · outbound

This paper cites Manning, Stefano Ermon, and Chelsea Finn.

Plug-and-Play Training Framework for Preference Optimization Manning, Stefano Ermon, and Chelsea Finn

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:34.152861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:34.152861Z digest=sha256:7dee52520d3c0c88d9315011353c8e0eb8cfd75554f06873bbd5e9d08820d9e5

Observation 9a3acaa4-fc68-48dc-8028-907020d36a94 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Plug-and-Play Training Framework for Preference Optimization Proximal Policy Optimization Algorithms

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:34.158147Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:34.158147Z digest=sha256:275b00ff051d335bf27d7b83e60e0c4531d79302cc1921d60290572c233da271

Observation a0ab056c-8d66-43d0-a617-e3feb2b29c60 · outbound

This paper cites Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters.

Plug-and-Play Training Framework for Preference Optimization Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:34.164247Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:34.164247Z digest=sha256:511918d68efe3ec91840660d97361f45c6aad270fad2b1e8a3575cff9d948d52

Observation 405f1f26-4d29-43a2-a344-a82199705812 · outbound

This paper cites Le, Ed H.

Plug-and-Play Training Framework for Preference Optimization Le, Ed H

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:34.169349Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:34.169349Z digest=sha256:5a6e992352d226b6f3f21f59596242f921d631ba2178cbc73fdc26b13335cdb6

Observation a6e3c076-9dc9-4fa4-a83a-532397b0be8f · outbound

This paper cites Qwen2 Technical Report.

Plug-and-Play Training Framework for Preference Optimization Qwen2 Technical Report

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:34.174862Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:34.174862Z digest=sha256:0854938577fd59920d03bbd1ab43a2c903962b1714c277dd19730d6f9e2895cb

Observation 4fb18886-77a4-47b7-a14e-9cf8ea65bc92 · outbound

This paper cites Kwok, Zhenguo Li, Adrian Weller, and Weiyang Liu.

Plug-and-Play Training Framework for Preference Optimization Kwok, Zhenguo Li, Adrian Weller, and Weiyang Liu

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:34.181641Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:34.181641Z digest=sha256:42d610ec30328d7c5b625bd9165a90df28a2dc3257f8c86b21a5c119ddc0c168

Observation 0cea60b1-0e76-4b86-acf1-d5eba7e78544 · outbound

This paper cites online" 'onlinestring :=.

Plug-and-Play Training Framework for Preference Optimization online" 'onlinestring :=

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:34.188212Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:34.188212Z digest=sha256:92ead938d40312e167f18b8b65be5e4db6ebc2c04d554af57315dd7582a83114

Observation dec7390c-6d40-4df8-85e1-03edd2e6e097 · outbound

This paper cites write newline.

Plug-and-Play Training Framework for Preference Optimization write newline

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-10T23:21:34.194345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T23:21:34.194345Z digest=sha256:926462cd70b86c889a097e66deea17ad30a54d3b2a11bf6871c7991c493b8df6

Pith citing papers

Observation 1d65934a-13c2-4695-8d3f-7e4726f1ae05 · inbound

Progressive Mastery: Customized Curriculum Learning with Guided Prompting for Mathematical Reasoning cites this paper.

Progressive Mastery: Customized Curriculum Learning with Guided Prompting for Mathematical Reasoning Plug-and-Play Training Framework for Preference Optimization

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-08-07T10:54:08.303337Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-07T10:54:07.795586Z digest=sha256:78b13b6ef7460c259dd99dcaf4ae1885827bfa778e495b9e9bf482cba4aa8ed2