Pith. sign in

Paper Citation Record · LEDGER

Efficient Online RFT with Plug-and-Play LLM Judges: Unlocking State-of-the-Art Performance

As of 19 August 2026, this Paper Citation Record lists 42 of 42 outbound references and 0 inbound Pith citation observations for arXiv:2506.05748.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.05748 v1

Coverage vector

measured 42 of 42 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T10:19:36.208057Z

measured 42 of 42 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

42 of 42 outbound references displayed

  • verified exact1
  • verified fuzzy19
  • unresolved20
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch2

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation bfc05f2a-2cd1-423e-b517-028202c5982b · outbound

This paper cites Each category of algorithms presents unique benefits and constraints, rendering their integrated application beneficial in real -world scenarios.

Efficient Online RFT with Plug-and-Play LLM Judges: Unlocking State-of-the-Art Performance Each category of algorithms presents unique benefits and constraints, rendering their integrated application beneficial in real -world scenarios

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:19:41.116682Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T10:19:32.904699Z digest=sha256:34de245af74bfce740e5ebfc331bf668890181cabea42797ab18faba0efda1ef

Observation 4bf31566-4f24-4634-8570-42b5186bb4d4 · outbound

This paper cites LLM-as- a-Judge.

Efficient Online RFT with Plug-and-Play LLM Judges: Unlocking State-of-the-Art Performance LLM-as- a-Judge

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:19:41.070207Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T10:19:32.952718Z digest=sha256:ce40201f92d379973cc28d2bbece49f1b8b69431f930d67ea0eb230e544a1585

Observation c4e8b6e8-0188-4312-a89d-4c30f929e8c7 · outbound

This paper cites score" field in [-1, 1] and a short.

Efficient Online RFT with Plug-and-Play LLM Judges: Unlocking State-of-the-Art Performance score" field in [-1, 1] and a short

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:19:40.876633Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T10:19:33.128347Z digest=sha256:6200e0b9361648730762955c372a6c15092d78a48f9a7b8998a4666548ded411

Observation 9c517a11-173d-4123-849e-a78414a33087 · outbound

This paper cites 𝑏𝑒𝑡𝑡𝑒𝑟":.

Efficient Online RFT with Plug-and-Play LLM Judges: Unlocking State-of-the-Art Performance 𝑏𝑒𝑡𝑡𝑒𝑟":

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:19:40.636409Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T10:19:33.304988Z digest=sha256:b72845104fa1a3cae491c76b5b1672b0979e949715f58706484a9217b060d3b4

Observation 0394a0ad-67b0-4824-8573-f5df187a27c4 · outbound

This paper cites be funnier.

Efficient Online RFT with Plug-and-Play LLM Judges: Unlocking State-of-the-Art Performance be funnier

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:19:40.765243Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T10:19:33.229860Z digest=sha256:04d7bfa5fa62635cac18348f2ee34b90e44c020624375f018c85d83b1d13e4c5

Observation 21804ec0-a5c9-4bad-a29b-1e5668030674 · outbound

This paper cites plug -and-play.

Efficient Online RFT with Plug-and-Play LLM Judges: Unlocking State-of-the-Art Performance plug -and-play

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:19:40.366764Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T10:19:33.434466Z digest=sha256:bb99984dc476136c4df0c40c49ca7d2d1e41f02468ccb5492a2b693628131d5c

Observation ed1ee598-6ad1-4a86-b403-bab2b8c2cefc · outbound

This paper cites Which answer is better? Return ‘A’ or ‘B’ only.

Efficient Online RFT with Plug-and-Play LLM Judges: Unlocking State-of-the-Art Performance Which answer is better? Return ‘A’ or ‘B’ only

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:19:40.499258Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T10:19:33.363323Z digest=sha256:a9519b0fe5466ac8a5b978a8518677088cb8a571e80d476a6f546a9f5826a6ca

Observation 2966fd8d-7156-4a75-887e-f9bfc63802e3 · outbound

This paper cites an unresolved cited work.

Efficient Online RFT with Plug-and-Play LLM Judges: Unlocking State-of-the-Art Performance Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:19:39.929275Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T10:19:33.586566Z digest=sha256:353774fb6142422dfba7651f2bbaca01cb89461c2bee68059a3efc94fae8e671

Observation 07f5747b-7a14-49e1-9a5b-c3a911624fd7 · outbound

This paper cites A” or “B.

Efficient Online RFT with Plug-and-Play LLM Judges: Unlocking State-of-the-Art Performance A” or “B

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:19:40.160618Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T10:19:33.507669Z digest=sha256:eb91f5392963494a1ae3476c2d8889c9e1adce7733221904ee73705e98762453

Observation 008d4bbd-9253-4423-838d-373d19c01c03 · outbound

This paper cites Online and Offline Reinforcement Learning by Planning with a Learned Model,.

Efficient Online RFT with Plug-and-Play LLM Judges: Unlocking State-of-the-Art Performance Online and Offline Reinforcement Learning by Planning with a Learned Model,

Reference 10

Resolution
metadata mismatch
raw_fallback, observed 2026-08-07T10:19:37.491599Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T10:19:34.462170Z digest=sha256:1f7e26e2c7a824cb767a8afd69f2691dbf153efc3a3782f5d6caa5ee6e5b953e

Observation 8bb93105-9cad-4e16-aae9-7de035cd3255 · outbound

This paper cites Direct Preference Optimization: Your Language Model is Secretly a Reward Model Oral,.

Efficient Online RFT with Plug-and-Play LLM Judges: Unlocking State-of-the-Art Performance Direct Preference Optimization: Your Language Model is Secretly a Reward Model Oral,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:19:39.631521Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T10:19:33.659848Z digest=sha256:6c40f94d75b46e0013eac46ccc6599c1f784c24d0458737b9c9779c6624b82e2

Observation 4230bc2a-447e-413d-b868-637db43f9f7d · outbound

This paper cites A Survey of Reinforcement Learning from Human Feedback,.

Efficient Online RFT with Plug-and-Play LLM Judges: Unlocking State-of-the-Art Performance A Survey of Reinforcement Learning from Human Feedback,

Reference 12

Resolution
verified exact
doi, observed 2026-08-07T10:19:36.981902Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T10:19:33.727935Z digest=sha256:3553d6c52e21caf476959962328eb0f7654365e3233926faa9f16791ac9e9864

Observation 52201413-e62d-4376-99d4-f540ac16dbc2 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Efficient Online RFT with Plug-and-Play LLM Judges: Unlocking State-of-the-Art Performance DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T10:19:33.848792Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:19:33.848792Z digest=sha256:c0d176906844bef4d3bbd1dfd70b5a36a59143149952ee45204fa6b6d0661a00

Observation 650048c6-1110-4069-93d1-b38e3de8259a · outbound

This paper cites Security and Privacy Challenges of Large Language Models: A Survey,.

Efficient Online RFT with Plug-and-Play LLM Judges: Unlocking State-of-the-Art Performance Security and Privacy Challenges of Large Language Models: A Survey,

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T10:19:33.926564Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:19:33.926564Z digest=sha256:e21dfb6fba9aa364b450d530c32f8c41d3c28dfb94e02d8db03e5f14a0f6a4d4

Observation 9f9e4aee-6743-4532-8a83-70b6c4ab6cb9 · outbound

This paper cites A Survey on Hallucination in Large Language Models: Principles, Taxonomy, Challenges, and Open Questions,.

Efficient Online RFT with Plug-and-Play LLM Judges: Unlocking State-of-the-Art Performance A Survey on Hallucination in Large Language Models: Principles, Taxonomy, Challenges, and Open Questions,

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T10:19:34.046186Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:19:34.046186Z digest=sha256:0fbc5c6e4397d86a0c831b769a5b977175440716d99eea7dbb5d2c42c251d4a8

Observation 9a466f73-917c-48b2-9f60-e20479759f8d · outbound

This paper cites Qwen2.5 Technical Report.

Efficient Online RFT with Plug-and-Play LLM Judges: Unlocking State-of-the-Art Performance Qwen2.5 Technical Report

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T10:19:34.135241Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:19:34.135241Z digest=sha256:439dc81703ea9bf8678443a214db30bf0146c8e197793755d8a61d5bf428f0b5

Observation 008311f0-f88e-4e7a-a34e-34c079fed166 · outbound

This paper cites Self-rewarding language models,.

Efficient Online RFT with Plug-and-Play LLM Judges: Unlocking State-of-the-Art Performance Self-rewarding language models,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:19:39.400464Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T10:19:34.252515Z digest=sha256:eaba3f9293491c5a46fad172ac7c3cb3cd931468d5394bf411d7523c968d63ac

Observation d4f62995-bf80-46ad-bb75-e2a1664aeb75 · outbound

This paper cites RLAIF: Scaling Reinforcement Learning from Human Feedback with AI Feedback,.

Efficient Online RFT with Plug-and-Play LLM Judges: Unlocking State-of-the-Art Performance RLAIF: Scaling Reinforcement Learning from Human Feedback with AI Feedback,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:19:39.116623Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T10:19:34.294065Z digest=sha256:a7068bc620e22186f235cf4ea6bbc2d133192d2089e7e40d451604437cc16650

Observation 3cdbe7f3-4c87-41ac-8be6-f70283a8170b · outbound

This paper cites Evaluating Text -to-Visual Generation with Image -to-Text Generation,.

Efficient Online RFT with Plug-and-Play LLM Judges: Unlocking State-of-the-Art Performance Evaluating Text -to-Visual Generation with Image -to-Text Generation,

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T10:19:34.415108Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:19:34.415108Z digest=sha256:d8bd4646c1e824b1d0ab075fb0227366a673881e271ea403a44a685da9a9a499

Observation 2e689453-8921-4ec3-85a3-8d38b9adadda · outbound

This paper cites more proficient.

Efficient Online RFT with Plug-and-Play LLM Judges: Unlocking State-of-the-Art Performance more proficient

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:19:41.034025Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T10:19:33.043123Z digest=sha256:f9519868b5943b8bd34589efe1fd8b901b9f6bdcff394a97d53cfad188905814

Observation 50ef29fd-e9e9-4a45-bdb5-a57243e42712 · outbound

This paper cites Training Language Models to Follow Instructions with Human Feedback,.

Efficient Online RFT with Plug-and-Play LLM Judges: Unlocking State-of-the-Art Performance Training Language Models to Follow Instructions with Human Feedback,

Reference 21

Resolution
metadata mismatch
raw_fallback, observed 2026-08-07T10:19:37.289842Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T10:19:34.523346Z digest=sha256:96d0832d591dd5b84774a61636e45e0394952792e5c7a1958db99dbe6cb1c8cb

Observation b3f198b2-d25f-4a79-afe5-066bbddecdb4 · outbound

This paper cites Reinforcement Learning Enhanced LLMs: A Survey.

Efficient Online RFT with Plug-and-Play LLM Judges: Unlocking State-of-the-Art Performance Reinforcement Learning Enhanced LLMs: A Survey

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T10:19:34.611764Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:19:34.611764Z digest=sha256:fe4b48b076ff7bbae0dea76cc1479d44f1e570112aeeafd57f32f036746b4a9d

Observation 27d2466d-933e-4b88-ad92-f2b3074f42d3 · outbound

This paper cites Survey on Large Language Model -Enhanced Reinforcement Learning: Concept, Taxonomy, and Methods,.

Efficient Online RFT with Plug-and-Play LLM Judges: Unlocking State-of-the-Art Performance Survey on Large Language Model -Enhanced Reinforcement Learning: Concept, Taxonomy, and Methods,

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T10:19:34.727234Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:19:34.727234Z digest=sha256:8680cb4198d95aa763fae679956decb8317a5dc3ca27532d30a607807ef0e731

Observation 74ae2aea-c188-4db3-8f8c-e0aa6aeca16a · outbound

This paper cites On-Device Qwen2.5: Efficient LLM Inference with Model Compression and Hardware Acceleration.

Efficient Online RFT with Plug-and-Play LLM Judges: Unlocking State-of-the-Art Performance On-Device Qwen2.5: Efficient LLM Inference with Model Compression and Hardware Acceleration

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T10:19:34.793339Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:19:34.793339Z digest=sha256:fa8c58a722a1e6076b13a6e6590f6ee9ffbafcf2a248853fb9b119ae98949603

Observation 7ed344d8-cc99-4521-8a1f-236530c45dda · outbound

This paper cites Human-like Summarization Evaluation with ChatGPT.

Efficient Online RFT with Plug-and-Play LLM Judges: Unlocking State-of-the-Art Performance Human-like Summarization Evaluation with ChatGPT

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T10:19:34.827221Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:19:34.827221Z digest=sha256:9c05fd0bb0335d2362d588aabe43828530fb99d25b503122e2dfbdabd50c1fad

Observation 2d96a1e9-2f2a-4270-9eba-9578b9740e94 · outbound

This paper cites Justice or Prejudice? Quantifying Biases in LLM-as-a-Judge.

Efficient Online RFT with Plug-and-Play LLM Judges: Unlocking State-of-the-Art Performance Justice or Prejudice? Quantifying Biases in LLM-as-a-Judge

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T10:19:34.917509Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:19:34.917509Z digest=sha256:b61e4371677c212aef858993241f350951f1fb7040f2a2689dee9ce0a16a48f0

Observation be93dd90-2e75-4460-9faa-018aaa2009e0 · outbound

This paper cites LLM-as-a-Judge & Reward Model: What They Can and Cannot Do.

Efficient Online RFT with Plug-and-Play LLM Judges: Unlocking State-of-the-Art Performance LLM-as-a-Judge & Reward Model: What They Can and Cannot Do

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T10:19:34.977810Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:19:34.977810Z digest=sha256:f8271194387967b13a54423c889a0008b8b2da1719622251f616f4fcb151f805

Observation b0a10ede-dbbe-42c8-b22f-a4d359ab22be · outbound

This paper cites RLAIF vs. RLHF: scaling reinforcement learning from human feedback with AI feedback,.

Efficient Online RFT with Plug-and-Play LLM Judges: Unlocking State-of-the-Art Performance RLAIF vs. RLHF: scaling reinforcement learning from human feedback with AI feedback,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:19:38.869784Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T10:19:35.034720Z digest=sha256:73f5074104ebab4eb4f78ffb1b5f9eddfe26d083e2ae7961d61f71dd1c930184

Observation fd153a62-f658-4573-8f88-c7e1ac52c102 · outbound

This paper cites Large Language Models Can Self -Improve,.

Efficient Online RFT with Plug-and-Play LLM Judges: Unlocking State-of-the-Art Performance Large Language Models Can Self -Improve,

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T10:19:35.096249Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:19:35.096249Z digest=sha256:208e1c42c248c6da23c58b7e5e8ac438fa967f17d7f54f56f1c9a5d2f24b3084

Observation 8bac4ca2-9414-4a58-9b45-926eb8f81762 · outbound

This paper cites Advancing Large Language Model Attribution through Self-Improving,.

Efficient Online RFT with Plug-and-Play LLM Judges: Unlocking State-of-the-Art Performance Advancing Large Language Model Attribution through Self-Improving,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:19:38.691907Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T10:19:35.180761Z digest=sha256:7f11ad41d22f8e1625c28c971062e8c03e845bda8d76d4e7e1fbd6e10d3e1845

Observation 7c9c9731-28a4-45d0-914a-91332507933e · outbound

This paper cites Systematic Evaluation of LLM-as-a-Judge in LLM Alignment Tasks: Explainable Metrics and Diverse Prompt Templates.

Efficient Online RFT with Plug-and-Play LLM Judges: Unlocking State-of-the-Art Performance Systematic Evaluation of LLM-as-a-Judge in LLM Alignment Tasks: Explainable Metrics and Diverse Prompt Templates

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T10:19:35.386305Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:19:35.386305Z digest=sha256:d784c1630e48edad2436fcd3a547a2a66475c560ece27fae1b40594c6a846071

Observation 3ed7257e-2d9d-42ea-b02f-56b633518d3b · outbound

This paper cites Can LLM be a Personalized Judge?,.

Efficient Online RFT with Plug-and-Play LLM Judges: Unlocking State-of-the-Art Performance Can LLM be a Personalized Judge?,

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T10:19:35.466565Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:19:35.466565Z digest=sha256:9f52c9291c214edc26b5d26058c6cb149a8c3a33b668230522acebaef909c849

Observation 485880b4-b338-414d-b015-92fa35054ec6 · outbound

This paper cites ReST-MCTS*: LLM Self- Training via Process Reward Guided Tree Search,.

Efficient Online RFT with Plug-and-Play LLM Judges: Unlocking State-of-the-Art Performance ReST-MCTS*: LLM Self- Training via Process Reward Guided Tree Search,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:19:38.522228Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T10:19:35.548712Z digest=sha256:3b94115f6f9defe5efbcec81219ca140b7e790e1526e05d60ddaf4291c606d98

Observation 56da1a36-06d1-4eee-a34c-2ce65173a5a5 · outbound

This paper cites Self-Play Preference Optimization for Language Model Alignment,.

Efficient Online RFT with Plug-and-Play LLM Judges: Unlocking State-of-the-Art Performance Self-Play Preference Optimization for Language Model Alignment,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:19:38.280897Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T10:19:35.634935Z digest=sha256:6f564de7337109b6382283bbe8581b84af68b65de1facb5a50645e2ed8e4c4b6

Observation 599d70aa-640b-4ce5-836e-8eaff13c67e2 · outbound

This paper cites Training language models to follow instructions with human feedback,.

Efficient Online RFT with Plug-and-Play LLM Judges: Unlocking State-of-the-Art Performance Training language models to follow instructions with human feedback,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:19:37.785037Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T10:19:35.826010Z digest=sha256:4f802384ebd6766f5a70226ef5b280d4e523f8c2ebd59a0709de07166da00c18

Observation 3bcf55f9-03af-4a07-bc39-bd4c46f401cb · outbound

This paper cites Constitutional AI: Harmlessness from AI Feedback.

Efficient Online RFT with Plug-and-Play LLM Judges: Unlocking State-of-the-Art Performance Constitutional AI: Harmlessness from AI Feedback

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T10:19:35.928792Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:19:35.928792Z digest=sha256:0817e58609e8acd8ef3478547128206347c48bb7745177aab85cc24e1dd3e4cd

Observation b18f9567-20f9-477a-9a02-cfd1064c7b01 · outbound

This paper cites RLAIF vs. RLHF: Scaling Reinforcement Learning from Human Feedback with AI Feedback.

Efficient Online RFT with Plug-and-Play LLM Judges: Unlocking State-of-the-Art Performance RLAIF vs. RLHF: Scaling Reinforcement Learning from Human Feedback with AI Feedback

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T10:19:35.995780Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:19:35.995780Z digest=sha256:0fd96c3c10656cedb461fa8591f7e61bf6a77ed84c1ea4f89545874877a9ba27

Observation 99d79685-81e9-4773-abbd-48ea36bbf58a · outbound

This paper cites Efficient Memory Management for Large Language Model Serving with PagedAttention.

Efficient Online RFT with Plug-and-Play LLM Judges: Unlocking State-of-the-Art Performance Efficient Memory Management for Large Language Model Serving with PagedAttention

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T10:19:36.068399Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:19:36.068399Z digest=sha256:0b7e7a7afccb166f592cdfb9ec95eccf907cead0d7780730597dcb8121c67b21

Observation e97fd2f1-6d6e-44c7-9cc1-9fcf64545a28 · outbound

This paper cites RewardBench: Evaluating Reward Models for Language Modeling.

Efficient Online RFT with Plug-and-Play LLM Judges: Unlocking State-of-the-Art Performance RewardBench: Evaluating Reward Models for Language Modeling

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T10:19:36.134562Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:19:36.134562Z digest=sha256:e338a360bd0660a5ee3afae2bcc13bc9a287b07c923d82c12e3352043414d35d

Observation 3be384a4-a619-4978-9f80-881a09551d7f · outbound

This paper cites Iterative Reasoning Preference Optimization,.

Efficient Online RFT with Plug-and-Play LLM Judges: Unlocking State-of-the-Art Performance Iterative Reasoning Preference Optimization,

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:19:37.604055Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T10:19:36.208057Z digest=sha256:d5f5b50a47b6d5c3d79dcba2792b6abe8d145c30a5427f197bfeb834ebf20bc5

Observation 16b69448-1360-4ae8-affb-5f829c56681e · outbound

This paper cites Available: https://neurips.cc/virtual/2024/108142.

Efficient Online RFT with Plug-and-Play LLM Judges: Unlocking State-of-the-Art Performance Available: https://neurips.cc/virtual/2024/108142

Reference 2024

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:19:38.000462Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T10:19:35.719817Z digest=sha256:2cb115e6fa2f69b213bc921061e9bf6454c6b5be30d374aa2b83959693e9023c

Observation b797fe01-8b85-4336-b0b0-c115757d42f8 · outbound

This paper cites an unresolved cited work.

Efficient Online RFT with Plug-and-Play LLM Judges: Unlocking State-of-the-Art Performance Unresolved cited work

Reference 3836

Resolution
unresolved
no resolver link, observed 2026-08-07T10:19:35.287341Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:19:35.287341Z digest=sha256:abac8da506a4a342b0af90c04b9391540acee2c28544ce7fa385db5a90fc69a5

Pith citing papers

No inbound Pith citation observations are available.