Pith. sign in

Paper Citation Record · LEDGER

Efficient Online RFT with Plug-and-Play LLM Judges: Unlocking State-of-the-Art Performance

As of 9 August 2026, this Paper Citation Record lists 42 of 42 outbound references and 0 inbound Pith citation observations for arXiv:2506.05748.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.05748 v1

Coverage vector

measured 42 of 42 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T10:19:36.208057Z

measured 42 of 42 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

42 of 42 outbound references displayed

  • verified exact1
  • verified fuzzy19
  • unresolved20
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch2

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation bfc05f2a-2cd1-423e-b517-028202c5982b · outbound

This paper cites Each category of algorithms presents unique benefits and constraints, rendering their integrated application beneficial in real -world scenarios.

Efficient Online RFT with Plug-and-Play LLM Judges: Unlocking State-of-the-Art Performance Each category of algorithms presents unique benefits and constraints, rendering their integrated application beneficial in real -world scenarios

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:19:41.116682Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T10:19:32.904699Z digest=sha256:b7c9c89b0704a6537e6b79ebbb1b2bd793ee46d23fe0c4e92400d4d040643f80

Observation 4bf31566-4f24-4634-8570-42b5186bb4d4 · outbound

This paper cites LLM-as- a-Judge.

Efficient Online RFT with Plug-and-Play LLM Judges: Unlocking State-of-the-Art Performance LLM-as- a-Judge

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:19:41.070207Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T10:19:32.952718Z digest=sha256:dc8c54be2909bde92ce4bc0cb0271952ab9b3a29863b2835ce33c5b144ed3908

Observation c4e8b6e8-0188-4312-a89d-4c30f929e8c7 · outbound

This paper cites score" field in [-1, 1] and a short.

Efficient Online RFT with Plug-and-Play LLM Judges: Unlocking State-of-the-Art Performance score" field in [-1, 1] and a short

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:19:40.876633Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T10:19:33.128347Z digest=sha256:8b6f8030a2bbe5446bd73cbd8db0972121e732c8f5fe782eb946e9f2d38c0f6a

Observation 9c517a11-173d-4123-849e-a78414a33087 · outbound

This paper cites 𝑏𝑒𝑡𝑡𝑒𝑟":.

Efficient Online RFT with Plug-and-Play LLM Judges: Unlocking State-of-the-Art Performance 𝑏𝑒𝑡𝑡𝑒𝑟":

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:19:40.636409Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T10:19:33.304988Z digest=sha256:477394032760c14d91b39623acea4ca9bc4f182f5c0e3ffce5bada8ee88e13be

Observation 0394a0ad-67b0-4824-8573-f5df187a27c4 · outbound

This paper cites be funnier.

Efficient Online RFT with Plug-and-Play LLM Judges: Unlocking State-of-the-Art Performance be funnier

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:19:40.765243Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T10:19:33.229860Z digest=sha256:fbd302873452a22ae7ff4ee6930fc93d6f9a4268b52c03f09410895da39f2532

Observation 21804ec0-a5c9-4bad-a29b-1e5668030674 · outbound

This paper cites plug -and-play.

Efficient Online RFT with Plug-and-Play LLM Judges: Unlocking State-of-the-Art Performance plug -and-play

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:19:40.366764Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T10:19:33.434466Z digest=sha256:52747b8fb43b5480a26fc655070f41714a175ea4862fcf1cf890cdcccc6100cd

Observation ed1ee598-6ad1-4a86-b403-bab2b8c2cefc · outbound

This paper cites Which answer is better? Return ‘A’ or ‘B’ only.

Efficient Online RFT with Plug-and-Play LLM Judges: Unlocking State-of-the-Art Performance Which answer is better? Return ‘A’ or ‘B’ only

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:19:40.499258Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T10:19:33.363323Z digest=sha256:ef3159750682e1b4c498a77dd4e5292bc11def15c0f25b774e042151927c8faa

Observation 2966fd8d-7156-4a75-887e-f9bfc63802e3 · outbound

This paper cites an unresolved cited work.

Efficient Online RFT with Plug-and-Play LLM Judges: Unlocking State-of-the-Art Performance Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:19:39.929275Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T10:19:33.586566Z digest=sha256:2454a2ead80ca6cf1f2f31192de47af868637fe1312f3e9f986f4d04d9d6ce71

Observation 07f5747b-7a14-49e1-9a5b-c3a911624fd7 · outbound

This paper cites A” or “B.

Efficient Online RFT with Plug-and-Play LLM Judges: Unlocking State-of-the-Art Performance A” or “B

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:19:40.160618Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T10:19:33.507669Z digest=sha256:f76019a4323a86690094b9796ce10a7aa6e8bf5a4c8b6a29e5d07529998d5598

Observation 008d4bbd-9253-4423-838d-373d19c01c03 · outbound

This paper cites Online and Offline Reinforcement Learning by Planning with a Learned Model,.

Efficient Online RFT with Plug-and-Play LLM Judges: Unlocking State-of-the-Art Performance Online and Offline Reinforcement Learning by Planning with a Learned Model,

Reference 10

Resolution
metadata mismatch
raw_fallback, observed 2026-08-07T10:19:37.491599Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T10:19:34.462170Z digest=sha256:db1d61a17f1b1e7c1ede0bde55b78df5597948a3aae7848f7910d39f81da124d

Observation 8bb93105-9cad-4e16-aae9-7de035cd3255 · outbound

This paper cites Direct Preference Optimization: Your Language Model is Secretly a Reward Model Oral,.

Efficient Online RFT with Plug-and-Play LLM Judges: Unlocking State-of-the-Art Performance Direct Preference Optimization: Your Language Model is Secretly a Reward Model Oral,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:19:39.631521Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T10:19:33.659848Z digest=sha256:a310050dc38ae7d669491a0f89ff01c3825cc976f0b6d0d934b2d68b7ee63766

Observation 4230bc2a-447e-413d-b868-637db43f9f7d · outbound

This paper cites A Survey of Reinforcement Learning from Human Feedback,.

Efficient Online RFT with Plug-and-Play LLM Judges: Unlocking State-of-the-Art Performance A Survey of Reinforcement Learning from Human Feedback,

Reference 12

Resolution
verified exact
doi, observed 2026-08-07T10:19:36.981902Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T10:19:33.727935Z digest=sha256:3717f7c425e8a24b53ce61528e4b440cceaa8904aeb466d9c669f3f25777bb4e

Observation 52201413-e62d-4376-99d4-f540ac16dbc2 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Efficient Online RFT with Plug-and-Play LLM Judges: Unlocking State-of-the-Art Performance DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T10:19:33.848792Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:19:33.848792Z digest=sha256:815e60ac33ee0af4f788f56dd643b864d874337c13f8336ad4f0ac86a40d47fc

Observation 650048c6-1110-4069-93d1-b38e3de8259a · outbound

This paper cites Security and Privacy Challenges of Large Language Models: A Survey,.

Efficient Online RFT with Plug-and-Play LLM Judges: Unlocking State-of-the-Art Performance Security and Privacy Challenges of Large Language Models: A Survey,

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T10:19:33.926564Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:19:33.926564Z digest=sha256:054465cf808601a3e4c82fcd80c5276dc44bd60f491f327a309e52aca2807d8e

Observation 9f9e4aee-6743-4532-8a83-70b6c4ab6cb9 · outbound

This paper cites A Survey on Hallucination in Large Language Models: Principles, Taxonomy, Challenges, and Open Questions,.

Efficient Online RFT with Plug-and-Play LLM Judges: Unlocking State-of-the-Art Performance A Survey on Hallucination in Large Language Models: Principles, Taxonomy, Challenges, and Open Questions,

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T10:19:34.046186Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:19:34.046186Z digest=sha256:6a516ba8d4870efe93bc398822c313544f6fb0554bc588f19da1339101e0a691

Observation 9a466f73-917c-48b2-9f60-e20479759f8d · outbound

This paper cites Qwen2.5 Technical Report.

Efficient Online RFT with Plug-and-Play LLM Judges: Unlocking State-of-the-Art Performance Qwen2.5 Technical Report

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T10:19:34.135241Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:19:34.135241Z digest=sha256:54623994d062a83aa231d7d42e16f86f4707129921ac37b7fce9fee1d7e4f698

Observation 008311f0-f88e-4e7a-a34e-34c079fed166 · outbound

This paper cites Self-rewarding language models,.

Efficient Online RFT with Plug-and-Play LLM Judges: Unlocking State-of-the-Art Performance Self-rewarding language models,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:19:39.400464Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T10:19:34.252515Z digest=sha256:0e17d6e384674a726c8266cb9cfa5ccb665cf3f328e411294cd970b9770af4fd

Observation d4f62995-bf80-46ad-bb75-e2a1664aeb75 · outbound

This paper cites RLAIF: Scaling Reinforcement Learning from Human Feedback with AI Feedback,.

Efficient Online RFT with Plug-and-Play LLM Judges: Unlocking State-of-the-Art Performance RLAIF: Scaling Reinforcement Learning from Human Feedback with AI Feedback,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:19:39.116623Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T10:19:34.294065Z digest=sha256:a87862509d208a9a4d2485855300b66464655118cf1422d65b8347582fc5f3d8

Observation 3cdbe7f3-4c87-41ac-8be6-f70283a8170b · outbound

This paper cites Evaluating Text -to-Visual Generation with Image -to-Text Generation,.

Efficient Online RFT with Plug-and-Play LLM Judges: Unlocking State-of-the-Art Performance Evaluating Text -to-Visual Generation with Image -to-Text Generation,

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T10:19:34.415108Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:19:34.415108Z digest=sha256:03ac90bbde812f5b30072ab1cd4d35542fd83bb5ada98049348286ebe7c2db78

Observation 2e689453-8921-4ec3-85a3-8d38b9adadda · outbound

This paper cites more proficient.

Efficient Online RFT with Plug-and-Play LLM Judges: Unlocking State-of-the-Art Performance more proficient

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:19:41.034025Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T10:19:33.043123Z digest=sha256:4d11420a028e7c8bdfb0e081e52457e5603a5e8ec825ec87020a794bf2304935

Observation 50ef29fd-e9e9-4a45-bdb5-a57243e42712 · outbound

This paper cites Training Language Models to Follow Instructions with Human Feedback,.

Efficient Online RFT with Plug-and-Play LLM Judges: Unlocking State-of-the-Art Performance Training Language Models to Follow Instructions with Human Feedback,

Reference 21

Resolution
metadata mismatch
raw_fallback, observed 2026-08-07T10:19:37.289842Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T10:19:34.523346Z digest=sha256:245f6e13340e564cb0a64fcb81c2d0080069950afd24358b831e497bf0ec1b84

Observation b3f198b2-d25f-4a79-afe5-066bbddecdb4 · outbound

This paper cites Reinforcement Learning Enhanced LLMs: A Survey.

Efficient Online RFT with Plug-and-Play LLM Judges: Unlocking State-of-the-Art Performance Reinforcement Learning Enhanced LLMs: A Survey

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T10:19:34.611764Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:19:34.611764Z digest=sha256:22f0318ce1c9520b06b409583ce5b40e223414e5396d8559d73dc67913cdc7f6

Observation 27d2466d-933e-4b88-ad92-f2b3074f42d3 · outbound

This paper cites Survey on Large Language Model -Enhanced Reinforcement Learning: Concept, Taxonomy, and Methods,.

Efficient Online RFT with Plug-and-Play LLM Judges: Unlocking State-of-the-Art Performance Survey on Large Language Model -Enhanced Reinforcement Learning: Concept, Taxonomy, and Methods,

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T10:19:34.727234Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:19:34.727234Z digest=sha256:8ca74c285e2f109c6275911cf22d4d573b6488d1d8ee2d9cd666bc9f64db7c27

Observation 74ae2aea-c188-4db3-8f8c-e0aa6aeca16a · outbound

This paper cites On-Device Qwen2.5: Efficient LLM Inference with Model Compression and Hardware Acceleration.

Efficient Online RFT with Plug-and-Play LLM Judges: Unlocking State-of-the-Art Performance On-Device Qwen2.5: Efficient LLM Inference with Model Compression and Hardware Acceleration

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T10:19:34.793339Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:19:34.793339Z digest=sha256:800b49ebe1168208163abc24b96d409db14f1441c374a443c47ae49e4479ac78

Observation 7ed344d8-cc99-4521-8a1f-236530c45dda · outbound

This paper cites Human-like Summarization Evaluation with ChatGPT.

Efficient Online RFT with Plug-and-Play LLM Judges: Unlocking State-of-the-Art Performance Human-like Summarization Evaluation with ChatGPT

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T10:19:34.827221Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:19:34.827221Z digest=sha256:6a90482731ef6f149296274f581b74ceb4571fee3782e2048d9e2f15488e9752

Observation 2d96a1e9-2f2a-4270-9eba-9578b9740e94 · outbound

This paper cites Justice or Prejudice? Quantifying Biases in LLM-as-a-Judge.

Efficient Online RFT with Plug-and-Play LLM Judges: Unlocking State-of-the-Art Performance Justice or Prejudice? Quantifying Biases in LLM-as-a-Judge

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T10:19:34.917509Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:19:34.917509Z digest=sha256:743365b485dee1f4def5657d6493f63c4776137ace90978444ebf8a6ec0a95cb

Observation be93dd90-2e75-4460-9faa-018aaa2009e0 · outbound

This paper cites LLM-as-a-Judge & Reward Model: What They Can and Cannot Do.

Efficient Online RFT with Plug-and-Play LLM Judges: Unlocking State-of-the-Art Performance LLM-as-a-Judge & Reward Model: What They Can and Cannot Do

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T10:19:34.977810Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:19:34.977810Z digest=sha256:73bc24b8ad40290428827fe7ee044e90cc1decaaa795b3c756f71a2211fb16a0

Observation b0a10ede-dbbe-42c8-b22f-a4d359ab22be · outbound

This paper cites RLAIF vs. RLHF: scaling reinforcement learning from human feedback with AI feedback,.

Efficient Online RFT with Plug-and-Play LLM Judges: Unlocking State-of-the-Art Performance RLAIF vs. RLHF: scaling reinforcement learning from human feedback with AI feedback,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:19:38.869784Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T10:19:35.034720Z digest=sha256:8d47249ff579c7f43c9ec3b62535f33b303d248a8745c4158f46e73b508fc563

Observation fd153a62-f658-4573-8f88-c7e1ac52c102 · outbound

This paper cites Large Language Models Can Self -Improve,.

Efficient Online RFT with Plug-and-Play LLM Judges: Unlocking State-of-the-Art Performance Large Language Models Can Self -Improve,

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T10:19:35.096249Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:19:35.096249Z digest=sha256:06a4bbc19068a6a3890b6cb928fbecc7fdb575bb5e260becb3094ca03d617509

Observation 8bac4ca2-9414-4a58-9b45-926eb8f81762 · outbound

This paper cites Advancing Large Language Model Attribution through Self-Improving,.

Efficient Online RFT with Plug-and-Play LLM Judges: Unlocking State-of-the-Art Performance Advancing Large Language Model Attribution through Self-Improving,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:19:38.691907Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T10:19:35.180761Z digest=sha256:d772dd6f27fc31ad718e3931d78baec951485588a10fdcb155ec2ebf96a8a286

Observation 7c9c9731-28a4-45d0-914a-91332507933e · outbound

This paper cites Systematic Evaluation of LLM-as-a-Judge in LLM Alignment Tasks: Explainable Metrics and Diverse Prompt Templates.

Efficient Online RFT with Plug-and-Play LLM Judges: Unlocking State-of-the-Art Performance Systematic Evaluation of LLM-as-a-Judge in LLM Alignment Tasks: Explainable Metrics and Diverse Prompt Templates

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T10:19:35.386305Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:19:35.386305Z digest=sha256:cad7bac6d6ef32674a6b888a336a0973539c440f41a8940e026751fb739c60a8

Observation 3ed7257e-2d9d-42ea-b02f-56b633518d3b · outbound

This paper cites Can LLM be a Personalized Judge?,.

Efficient Online RFT with Plug-and-Play LLM Judges: Unlocking State-of-the-Art Performance Can LLM be a Personalized Judge?,

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T10:19:35.466565Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:19:35.466565Z digest=sha256:635fb177e055fc9dc50f65a213b1f5a51c24ac1775e6763b4cd17685522cbca0

Observation 485880b4-b338-414d-b015-92fa35054ec6 · outbound

This paper cites ReST-MCTS*: LLM Self- Training via Process Reward Guided Tree Search,.

Efficient Online RFT with Plug-and-Play LLM Judges: Unlocking State-of-the-Art Performance ReST-MCTS*: LLM Self- Training via Process Reward Guided Tree Search,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:19:38.522228Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T10:19:35.548712Z digest=sha256:5e5596faa3ce621690770268d6b803005ba4f92c8f978212b832354e217c8255

Observation 56da1a36-06d1-4eee-a34c-2ce65173a5a5 · outbound

This paper cites Self-Play Preference Optimization for Language Model Alignment,.

Efficient Online RFT with Plug-and-Play LLM Judges: Unlocking State-of-the-Art Performance Self-Play Preference Optimization for Language Model Alignment,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:19:38.280897Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T10:19:35.634935Z digest=sha256:cb79223c79a2bec28b5f3f034762e85ebefa1eca124ebc55f2f22e3164276d71

Observation 599d70aa-640b-4ce5-836e-8eaff13c67e2 · outbound

This paper cites Training language models to follow instructions with human feedback,.

Efficient Online RFT with Plug-and-Play LLM Judges: Unlocking State-of-the-Art Performance Training language models to follow instructions with human feedback,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:19:37.785037Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T10:19:35.826010Z digest=sha256:254ae939f04c1627670ee565cdc30f484efe40eda90a8b321af33fdb0561aae0

Observation 3bcf55f9-03af-4a07-bc39-bd4c46f401cb · outbound

This paper cites Constitutional AI: Harmlessness from AI Feedback.

Efficient Online RFT with Plug-and-Play LLM Judges: Unlocking State-of-the-Art Performance Constitutional AI: Harmlessness from AI Feedback

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T10:19:35.928792Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:19:35.928792Z digest=sha256:0683d151839519e409df05fed5c3bb07b43b108e9006ee50494b2976583ca83a

Observation b18f9567-20f9-477a-9a02-cfd1064c7b01 · outbound

This paper cites RLAIF vs. RLHF: Scaling Reinforcement Learning from Human Feedback with AI Feedback.

Efficient Online RFT with Plug-and-Play LLM Judges: Unlocking State-of-the-Art Performance RLAIF vs. RLHF: Scaling Reinforcement Learning from Human Feedback with AI Feedback

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T10:19:35.995780Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:19:35.995780Z digest=sha256:84d2154876681183957e99bcf897b0f33405f46185aab73fd6b567d81cba9634

Observation 99d79685-81e9-4773-abbd-48ea36bbf58a · outbound

This paper cites Efficient Memory Management for Large Language Model Serving with PagedAttention.

Efficient Online RFT with Plug-and-Play LLM Judges: Unlocking State-of-the-Art Performance Efficient Memory Management for Large Language Model Serving with PagedAttention

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T10:19:36.068399Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:19:36.068399Z digest=sha256:a72629a3335236a87e8311447f1019b142b9d33f4ca32aa3239967abe5d8c1c8

Observation e97fd2f1-6d6e-44c7-9cc1-9fcf64545a28 · outbound

This paper cites RewardBench: Evaluating Reward Models for Language Modeling.

Efficient Online RFT with Plug-and-Play LLM Judges: Unlocking State-of-the-Art Performance RewardBench: Evaluating Reward Models for Language Modeling

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T10:19:36.134562Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:19:36.134562Z digest=sha256:be7c2459c6492aa83363b7a417e077476e23e400e8b11d4267809682638c873a

Observation 3be384a4-a619-4978-9f80-881a09551d7f · outbound

This paper cites Iterative Reasoning Preference Optimization,.

Efficient Online RFT with Plug-and-Play LLM Judges: Unlocking State-of-the-Art Performance Iterative Reasoning Preference Optimization,

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:19:37.604055Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T10:19:36.208057Z digest=sha256:be13634344ab9742f675b1ca2dde53000aafd87586e59a966b51e184c12b25e2

Observation 16b69448-1360-4ae8-affb-5f829c56681e · outbound

This paper cites Available: https://neurips.cc/virtual/2024/108142.

Efficient Online RFT with Plug-and-Play LLM Judges: Unlocking State-of-the-Art Performance Available: https://neurips.cc/virtual/2024/108142

Reference 2024

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:19:38.000462Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T10:19:35.719817Z digest=sha256:0704fa44eea510c2f9195dcf9ac6ea77f721bde1c13eccaceb4f709f0fe29950

Observation b797fe01-8b85-4336-b0b0-c115757d42f8 · outbound

This paper cites an unresolved cited work.

Efficient Online RFT with Plug-and-Play LLM Judges: Unlocking State-of-the-Art Performance Unresolved cited work

Reference 3836

Resolution
unresolved
no resolver link, observed 2026-08-07T10:19:35.287341Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:19:35.287341Z digest=sha256:297516423981d03a4ceeec6b5f542fb3210133e2d051d61625d6d2752e8cb59e

Pith citing papers

No inbound Pith citation observations are available.