Pith. sign in

Paper Citation Record · LEDGER

Aligning Frozen LLMs by Reinforcement Learning: An Iterative Reweight-then-Optimize Approach

As of 19 August 2026, this Paper Citation Record lists 58 of 58 outbound references and 0 inbound Pith citation observations for arXiv:2506.17828.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.17828 v2

Coverage vector

measured 58 of 58 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T19:08:14.574756Z

measured 58 of 58 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

58 of 58 outbound references displayed

  • verified exact0
  • verified fuzzy9
  • unresolved49
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation ee611819-3f7e-465d-8532-0b70afc8d4e7 · outbound

This paper cites GPT-4 Technical Report.

Aligning Frozen LLMs by Reinforcement Learning: An Iterative Reweight-then-Optimize Approach GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-15T19:08:14.202527Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:08:14.202527Z digest=sha256:71d6f4d25c09e596238ba23e230f4e5a0be34d3a1b50adfc5501a326dce91ff0

Observation 50cbe863-13a2-45e1-a89f-22ae5b59ea57 · outbound

This paper cites Reinforcement learning: Theory and algorithms.

Aligning Frozen LLMs by Reinforcement Learning: An Iterative Reweight-then-Optimize Approach Reinforcement learning: Theory and algorithms

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:08:15.795787Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T19:08:14.211866Z digest=sha256:2c87eade756200ddf55211c70085554aa523d1310b73a6308645ef761f3aef19

Observation 744ad2b3-92b3-4073-a463-75b8c62d8368 · outbound

This paper cites Back to Basics: Revisiting REINFORCE Style Optimization for Learning from Human Feedback in LLMs.

Aligning Frozen LLMs by Reinforcement Learning: An Iterative Reweight-then-Optimize Approach Back to Basics: Revisiting REINFORCE Style Optimization for Learning from Human Feedback in LLMs

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T19:08:14.218608Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:08:14.218608Z digest=sha256:c8845dd64ca5af922b0ed3f7f6bd79f4c9984956b12a1c07bb979ae7664efed5

Observation d84473d0-5e33-4456-af2e-9c13dd60aa38 · outbound

This paper cites A General Theoretical Paradigm to Understand Learning from Human Preferences.

Aligning Frozen LLMs by Reinforcement Learning: An Iterative Reweight-then-Optimize Approach A General Theoretical Paradigm to Understand Learning from Human Preferences

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-15T19:08:14.225388Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:08:14.225388Z digest=sha256:8dc88b0a58c190a71c4f78008fda6a4f24df9789fd735cc36bf6b79278d41573

Observation a4f7cb5e-e5a5-433d-a631-1ae0fe61e43f · outbound

This paper cites Qwen Technical Report.

Aligning Frozen LLMs by Reinforcement Learning: An Iterative Reweight-then-Optimize Approach Qwen Technical Report

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-15T19:08:14.232737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:08:14.232737Z digest=sha256:8b3192f992293bf8bb3d324084391f7e9b60f22a51f6292aa3415ab1282980b9

Observation 81b0bc1d-1216-4e58-8df1-c1690ef2a97d · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

Aligning Frozen LLMs by Reinforcement Learning: An Iterative Reweight-then-Optimize Approach Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T19:08:14.239876Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:08:14.239876Z digest=sha256:a6b676fa47eb1f633315a6b1d38f1f9cd6f79d6268fbb08195a222e58c3d25af

Observation f1a7550d-660c-4f82-afc0-44c3f8f2f27a · outbound

This paper cites A survey of monte carlo tree search methods.IEEE Transactions on Computational Intelligence and AI in games, 4(1):1–43, 2012.

Aligning Frozen LLMs by Reinforcement Learning: An Iterative Reweight-then-Optimize Approach A survey of monte carlo tree search methods.IEEE Transactions on Computational Intelligence and AI in games, 4(1):1–43, 2012

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T19:08:14.245892Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:08:14.245892Z digest=sha256:5eec5766d290030484f11922a354fabe4b90e42b92de6ccbc147091e5dcfbc7d

Observation e2610822-92da-4561-a7b9-a48eb13a532f · outbound

This paper cites Transfer Q Star: Principled Decoding for LLM Alignment.

Aligning Frozen LLMs by Reinforcement Learning: An Iterative Reweight-then-Optimize Approach Transfer Q Star: Principled Decoding for LLM Alignment

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T19:08:14.252424Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:08:14.252424Z digest=sha256:48b5e3cd0885862708943966235322e8705c7229ce5e491e021869545c9f1018

Observation a3997ffc-3c70-4c65-8f76-584313a894f9 · outbound

This paper cites Dataset Reset Policy Optimization for RLHF.

Aligning Frozen LLMs by Reinforcement Learning: An Iterative Reweight-then-Optimize Approach Dataset Reset Policy Optimization for RLHF

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T19:08:14.259906Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:08:14.259906Z digest=sha256:53772fd52d9f8d7f8c5f0c275a2ceea1f8a06ce278447aa5a6b1d135f7985e91

Observation dd6790e4-3450-4c44-b3ef-67a9dd4cc1d6 · outbound

This paper cites Stop summation: Min-form credit assignment is all process reward model needs for reasoning.arXiv preprint arXiv:2504.15275, 2025.

Aligning Frozen LLMs by Reinforcement Learning: An Iterative Reweight-then-Optimize Approach Stop summation: Min-form credit assignment is all process reward model needs for reasoning.arXiv preprint arXiv:2504.15275, 2025

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T19:08:14.267848Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:08:14.267848Z digest=sha256:87bac6a1bf33c2be327b6a144d6f062eb4e4fdea0f68121ae794632e705c9929

Observation 4c16a787-d391-4da9-94da-d5f1daf4f6d3 · outbound

This paper cites Ultrafeedback: Boosting language models with high-quality feedback.

Aligning Frozen LLMs by Reinforcement Learning: An Iterative Reweight-then-Optimize Approach Ultrafeedback: Boosting language models with high-quality feedback

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T19:08:14.274186Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:08:14.274186Z digest=sha256:3f97983c9d0cd275b84e82ce9be292a180cbaa18c77b74f648528401d38ae342

Observation 79b5424d-9322-432d-889b-622567018dc3 · outbound

This paper cites RAFT: Reward rAnked FineTuning for Generative Foundation Model Alignment.

Aligning Frozen LLMs by Reinforcement Learning: An Iterative Reweight-then-Optimize Approach RAFT: Reward rAnked FineTuning for Generative Foundation Model Alignment

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-15T19:08:14.280121Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:08:14.280121Z digest=sha256:ebd6e9d86793f42c417c0bab14a981c03ad17576b24cd66aff7f892930ad532f

Observation c1fcae9a-7bbd-4ac7-99fc-69b3a4ae3da9 · outbound

This paper cites Rlhf workflow: From reward modeling to online rlhf, 2024.

Aligning Frozen LLMs by Reinforcement Learning: An Iterative Reweight-then-Optimize Approach Rlhf workflow: From reward modeling to online rlhf, 2024

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-15T19:08:14.286087Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:08:14.286087Z digest=sha256:c3a617f343d11575beae688d6bb27f5490fd82911409a5317896e16bd4109437

Observation e719762f-c6b8-431c-99fb-b21406c357b9 · outbound

This paper cites Length-Controlled AlpacaEval: A Simple Way to Debias Automatic Evaluators.

Aligning Frozen LLMs by Reinforcement Learning: An Iterative Reweight-then-Optimize Approach Length-Controlled AlpacaEval: A Simple Way to Debias Automatic Evaluators

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-15T19:08:14.291859Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:08:14.291859Z digest=sha256:73c920c173341d3f718850fba6e65e32c8e2a81e5c4db46a1a0da05328a223d3

Observation 2fc67dec-5e38-4bc8-9f2f-9125c1729522 · outbound

This paper cites Alpacafarm: A simulation framework for methods that learn from human feedback.Advances in Neural Information Processing Systems, 36:30039–30069, 2023.

Aligning Frozen LLMs by Reinforcement Learning: An Iterative Reweight-then-Optimize Approach Alpacafarm: A simulation framework for methods that learn from human feedback.Advances in Neural Information Processing Systems, 36:30039–30069, 2023

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:08:15.742734Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T19:08:14.297478Z digest=sha256:b59637bc8f4e85e9054daed5aabcef82ee312a78c847664e89af0b32a5f8d50a

Observation 7284d86f-6f80-4f70-baec-d5b67ccac989 · outbound

This paper cites Stop Regressing: Training Value Functions via Classification for Scalable Deep RL.

Aligning Frozen LLMs by Reinforcement Learning: An Iterative Reweight-then-Optimize Approach Stop Regressing: Training Value Functions via Classification for Scalable Deep RL

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-15T19:08:14.303868Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:08:14.303868Z digest=sha256:ca6864fa79ce214db7656a55c1deffd6118ae7cea4a411393aaca33afb064119

Observation dedde5e3-c863-47fd-a951-c8cbe2efcce2 · outbound

This paper cites Scaling laws for reward model overoptimization.

Aligning Frozen LLMs by Reinforcement Learning: An Iterative Reweight-then-Optimize Approach Scaling laws for reward model overoptimization

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-15T19:08:14.311009Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:08:14.311009Z digest=sha256:861448980347977c8f296e4cae4f338f735650ed06a14316b76be0eec41e99a8

Observation 89640b75-cedc-4d30-a588-1a24a1e32584 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Aligning Frozen LLMs by Reinforcement Learning: An Iterative Reweight-then-Optimize Approach DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-15T19:08:14.317597Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:08:14.317597Z digest=sha256:091cd7749699d0150de582fda89775b4f5cd06b4e71cc27f4bd698b1ea7faaf9

Observation 36272bd8-8096-4d01-9205-4825fe02531c · outbound

This paper cites Value Augmented Sampling for Language Model Alignment and Personalization.

Aligning Frozen LLMs by Reinforcement Learning: An Iterative Reweight-then-Optimize Approach Value Augmented Sampling for Language Model Alignment and Personalization

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-15T19:08:14.323757Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:08:14.323757Z digest=sha256:a348fbf9898637f2711772ca8b11ae0751e40410532b801787d5c3ac22a6489d

Observation b1e45707-df66-4644-855f-3f21030122be · outbound

This paper cites Inverse preference learning: Preference-based rl without a reward function.

Aligning Frozen LLMs by Reinforcement Learning: An Iterative Reweight-then-Optimize Approach Inverse preference learning: Preference-based rl without a reward function

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:08:15.712208Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T19:08:14.330108Z digest=sha256:3cf38ae459b5d81732eea36d72a56eb66ce8e6c80468b1cd72971ac69b982937

Observation b43fc796-254a-4e5c-ae5d-95b7debcbc94 · outbound

This paper cites Deal: Decoding-time alignment for large language models.

Aligning Frozen LLMs by Reinforcement Learning: An Iterative Reweight-then-Optimize Approach Deal: Decoding-time alignment for large language models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-15T19:08:14.335859Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:08:14.335859Z digest=sha256:c24f362580a9bbbaa0b6b47fd73016cbde7839e46d0cf9b03a8d9f00613c7813

Observation bded4ea0-0e61-4bd1-9592-25d177f931d7 · outbound

This paper cites The N+ Implementation Details of RLHF with PPO: A Case Study on TL;DR Summarization.

Aligning Frozen LLMs by Reinforcement Learning: An Iterative Reweight-then-Optimize Approach The N+ Implementation Details of RLHF with PPO: A Case Study on TL;DR Summarization

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-15T19:08:14.341846Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:08:14.341846Z digest=sha256:6bcb10cf5a3e53dc70056232309fd22a493ae23451699fc929f20e325b79e246

Observation 877e0763-f7ec-4693-8cf1-43712e2fe91b · outbound

This paper cites LLM-Blender: Ensembling Large Language Models with Pairwise Ranking and Generative Fusion.

Aligning Frozen LLMs by Reinforcement Learning: An Iterative Reweight-then-Optimize Approach LLM-Blender: Ensembling Large Language Models with Pairwise Ranking and Generative Fusion

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-15T19:08:14.347674Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:08:14.347674Z digest=sha256:d4a793a599675554aa383b73931f48493f73dd0fa283fb3129103a48bc279368

Observation 2c522dfd-c841-4e6d-82de-c703bdd4596f · outbound

This paper cites A natural policy gradient.Advances in neural information processing systems, 14, 2001.

Aligning Frozen LLMs by Reinforcement Learning: An Iterative Reweight-then-Optimize Approach A natural policy gradient.Advances in neural information processing systems, 14, 2001

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-15T19:08:14.355893Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:08:14.355893Z digest=sha256:f8af2e23dac0b8f0afce65d15c3023e26e2e555e2291617d349f151cf76bb74a

Observation 0db04d10-d604-4cc7-a8b4-31080f448bea · outbound

This paper cites ARGS: Alignment as Reward-Guided Search.

Aligning Frozen LLMs by Reinforcement Learning: An Iterative Reweight-then-Optimize Approach ARGS: Alignment as Reward-Guided Search

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-15T19:08:14.361484Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:08:14.361484Z digest=sha256:19ce4dbd472b04a3a80d3f533b6455ac044cb5e1121af71f53b1b7aac303b176

Observation c8aee78b-af36-4b8f-812f-3d206e6dae75 · outbound

This paper cites Aligning Large Language Models with Representation Editing: A Control Perspective.

Aligning Frozen LLMs by Reinforcement Learning: An Iterative Reweight-then-Optimize Approach Aligning Large Language Models with Representation Editing: A Control Perspective

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-15T19:08:14.368583Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:08:14.368583Z digest=sha256:f54da9130d08a7e85110401cae8818648f0b6efdcc6d65543c67fd4bee06dd9e

Observation a8c9afe5-3b56-498f-9129-1c8a0280e79a · outbound

This paper cites Optimization Issues in KL-Constrained Approximate Policy Iteration.

Aligning Frozen LLMs by Reinforcement Learning: An Iterative Reweight-then-Optimize Approach Optimization Issues in KL-Constrained Approximate Policy Iteration

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-15T19:08:14.376134Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:08:14.376134Z digest=sha256:35f5ea29a9604e78e63ec7d291f76cfc458e8dc73eec4338a6ef290f51b7abe6

Observation c418a9ce-776a-47e1-9aec-d7e8708e6590 · outbound

This paper cites Cascade Reward Sampling for Efficient Decoding-Time Alignment.

Aligning Frozen LLMs by Reinforcement Learning: An Iterative Reweight-then-Optimize Approach Cascade Reward Sampling for Efficient Decoding-Time Alignment

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-15T19:08:14.383059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:08:14.383059Z digest=sha256:1e7bd7e0d225b654317d523c99a8d60af2caf09b262f8a1c5f4c21e4279cc1ad

Observation 14f6de3c-da9f-46e4-b8b5-55009934b19f · outbound

This paper cites Alpacaeval: An automatic evaluator of instruction-following models, 2023.

Aligning Frozen LLMs by Reinforcement Learning: An Iterative Reweight-then-Optimize Approach Alpacaeval: An automatic evaluator of instruction-following models, 2023

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-15T19:08:14.389432Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:08:14.389432Z digest=sha256:3357499516c26426409d09fdacf1c9f741349473f7fdf7ce2a9c0f57466346b6

Observation 5f93dc1f-f411-4cc7-a57d-744f866108d0 · outbound

This paper cites Let’s verify step by step.

Aligning Frozen LLMs by Reinforcement Learning: An Iterative Reweight-then-Optimize Approach Let’s verify step by step

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-15T19:08:14.395816Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:08:14.395816Z digest=sha256:4c37a137be813a5f9d20189b8318fb498fdc98169a968df4cfa5f02a6d877466

Observation b7b2f6b1-c043-4007-97d4-cac7186fb370 · outbound

This paper cites Inference-Time Language Model Alignment via Integrated Value Guidance.

Aligning Frozen LLMs by Reinforcement Learning: An Iterative Reweight-then-Optimize Approach Inference-Time Language Model Alignment via Integrated Value Guidance

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-15T19:08:14.402192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:08:14.402192Z digest=sha256:a3b226397c21aac207960ec8e9abafea905d7ff44dad2fbbec05fa95b3a85f07

Observation 3ed8fc66-0ea9-4060-8d1c-948040777f5c · outbound

This paper cites Controlled Decoding from Language Models.

Aligning Frozen LLMs by Reinforcement Learning: An Iterative Reweight-then-Optimize Approach Controlled Decoding from Language Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-15T19:08:14.408376Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:08:14.408376Z digest=sha256:1d5fac348e9ab82ed7d3dd0d1ec0c761121c66a31ab77c0e9647382c94461082

Observation 87f5d692-b0d3-47fb-b503-187c38a5abe7 · outbound

This paper cites WebGPT: Browser-assisted question-answering with human feedback.

Aligning Frozen LLMs by Reinforcement Learning: An Iterative Reweight-then-Optimize Approach WebGPT: Browser-assisted question-answering with human feedback

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-15T19:08:14.414389Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:08:14.414389Z digest=sha256:e7d727489bc5b59903675468ac58d9c1d8782df2dd61a165f8d70f747a1bc5c8

Observation f95e0cbb-e569-4a80-9640-34b721cd5cb6 · outbound

This paper cites Training language models to follow instructions with human feedback.Advances in Neural Information Processing Systems, 35:27730–27744, 2022.

Aligning Frozen LLMs by Reinforcement Learning: An Iterative Reweight-then-Optimize Approach Training language models to follow instructions with human feedback.Advances in Neural Information Processing Systems, 35:27730–27744, 2022

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-15T19:08:14.420029Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:08:14.420029Z digest=sha256:70d0ad7c3d612fec5e8d33dbc82ab4fed11421978df4f5ada099f942dbe408b7

Observation ea6f7082-3051-4f71-9b59-378044cee4f9 · outbound

This paper cites TreeBoN: Enhancing Inference-Time Alignment with Speculative Tree-Search and Best-of-N Sampling.

Aligning Frozen LLMs by Reinforcement Learning: An Iterative Reweight-then-Optimize Approach TreeBoN: Enhancing Inference-Time Alignment with Speculative Tree-Search and Best-of-N Sampling

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-15T19:08:14.426493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:08:14.426493Z digest=sha256:0fb5e36294755c5cec212e7fef0889524505c560951ce5a040648e041ef83aa8

Observation f814bcac-3f49-4600-8840-29fd9fe3bdbc · outbound

This paper cites From $r$ to $Q^*$: Your Language Model is Secretly a Q-Function.

Aligning Frozen LLMs by Reinforcement Learning: An Iterative Reweight-then-Optimize Approach From $r$ to $Q^*$: Your Language Model is Secretly a Q-Function

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-15T19:08:14.432916Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:08:14.432916Z digest=sha256:c7a0feaf447111e42e4bcbf223e218b1ab29123a51ad76a65d9f9479907533f0

Observation ee815ad4-78de-4461-a074-aa116b52515c · outbound

This paper cites Direct preference optimization: Your language model is secretly a reward model.Advances in Neural Information Processing Systems, 36, 2024.

Aligning Frozen LLMs by Reinforcement Learning: An Iterative Reweight-then-Optimize Approach Direct preference optimization: Your language model is secretly a reward model.Advances in Neural Information Processing Systems, 36, 2024

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-15T19:08:14.439663Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:08:14.439663Z digest=sha256:29d9e422eeec27b3a14974bbf68ddfbb52e3d774e139e30c610be0f4012883ab

Observation 9b9875b7-c12e-4415-a1b4-fdc0f9ec9713 · outbound

This paper cites Trust Region Policy Optimization.

Aligning Frozen LLMs by Reinforcement Learning: An Iterative Reweight-then-Optimize Approach Trust Region Policy Optimization

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-15T19:08:14.445386Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:08:14.445386Z digest=sha256:83b614e5a188eb81717747dbf74487d154768e4bd9c4bea3149c0c89a9bc3818

Observation e7cf550e-8b3a-44d2-8867-db4b01af33c8 · outbound

This paper cites High-Dimensional Continuous Control Using Generalized Advantage Estimation.

Aligning Frozen LLMs by Reinforcement Learning: An Iterative Reweight-then-Optimize Approach High-Dimensional Continuous Control Using Generalized Advantage Estimation

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-15T19:08:14.450938Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:08:14.450938Z digest=sha256:79aeab6df5121b26f541b0a3629e12f148290f8de302f6fd9495fd80f0cab8ce

Observation 7637d207-7007-465c-97e8-b52741fe1e5f · outbound

This paper cites Proximal Policy Optimization Algorithms.

Aligning Frozen LLMs by Reinforcement Learning: An Iterative Reweight-then-Optimize Approach Proximal Policy Optimization Algorithms

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-15T19:08:14.456949Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:08:14.456949Z digest=sha256:6f9d3036137120addc2d844c9cee33e740627b0e6f86431e74fb809329a6cef5

Observation f0d89a27-f388-4cf0-a80b-f992c655430f · outbound

This paper cites Adaptive trust region policy optimization: Global convergence and faster rates for regularized mdps.

Aligning Frozen LLMs by Reinforcement Learning: An Iterative Reweight-then-Optimize Approach Adaptive trust region policy optimization: Global convergence and faster rates for regularized mdps

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:08:15.636908Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T19:08:14.462585Z digest=sha256:0d3a32b275869d8ace54f0bd313cc3e111896ef4c9d35bfcf841d1b545eb2037

Observation f9e936f6-3c4e-4d46-bac2-8476db04cf37 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Aligning Frozen LLMs by Reinforcement Learning: An Iterative Reweight-then-Optimize Approach DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-15T19:08:14.468474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:08:14.468474Z digest=sha256:b822a498934ce222a3f051d5e2f83d21675bbff8783b384da75542be7b15a1fa

Observation e2b8e3f7-fe23-462d-825a-cd7375a0d7cf · outbound

This paper cites Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters.

Aligning Frozen LLMs by Reinforcement Learning: An Iterative Reweight-then-Optimize Approach Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-15T19:08:14.475991Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:08:14.475991Z digest=sha256:8298e9c49bf9c137e36e4dbd2c5e837123589f799ee4c44a870c1a330e99ceff

Observation 199d94c2-2516-4795-abfe-d4d2ee6e611e · outbound

This paper cites Hybrid RL: Using Both Offline and Online Data Can Make RL Efficient.

Aligning Frozen LLMs by Reinforcement Learning: An Iterative Reweight-then-Optimize Approach Hybrid RL: Using Both Offline and Online Data Can Make RL Efficient

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-15T19:08:14.482273Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:08:14.482273Z digest=sha256:bcbd173d945a2af4a3117bcc29382b68531f8359b81ab8be88edece27ad2cb8e

Observation 3492c259-5465-469b-a9c4-d284785595eb · outbound

This paper cites Learning to summarize with human feedback.Advances in Neural Information Processing Systems, 33:3008–3021, 2020.

Aligning Frozen LLMs by Reinforcement Learning: An Iterative Reweight-then-Optimize Approach Learning to summarize with human feedback.Advances in Neural Information Processing Systems, 33:3008–3021, 2020

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-15T19:08:14.488903Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:08:14.488903Z digest=sha256:29c278610d39c423ecd23036dc07e32e0349ff3e90e8622c6fff5de5e967ddd2

Observation bdc53125-8e08-49f7-80c8-b67f95e85451 · outbound

This paper cites On the convergence rates of policy gradient methods.Journal of Machine Learning Research, 23(282):1–36, 2022.

Aligning Frozen LLMs by Reinforcement Learning: An Iterative Reweight-then-Optimize Approach On the convergence rates of policy gradient methods.Journal of Machine Learning Research, 23(282):1–36, 2022

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-15T19:08:14.495629Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:08:14.495629Z digest=sha256:576b5597e8f533c8291324b2444fd0db55bc0bd59d547a6b6b8c45c0236dcdee

Observation 5dfb5fc4-6623-405c-9e18-0c96c5b66eea · outbound

This paper cites Monte Carlo Tree Search Boosts Reasoning via Iterative Preference Learning.

Aligning Frozen LLMs by Reinforcement Learning: An Iterative Reweight-then-Optimize Approach Monte Carlo Tree Search Boosts Reasoning via Iterative Preference Learning

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-15T19:08:14.502275Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:08:14.502275Z digest=sha256:6ce8d1d37213c5eb2cac2ffce064fa1fb06fdf86fb28b3ba1a4e6df22828a0b2

Observation 806b563c-2708-49fb-bfe1-a9a2e1ecfcbe · outbound

This paper cites Iterative preference learning from human feedback: Bridging theory and practice for rlhf under kl- constraint.

Aligning Frozen LLMs by Reinforcement Learning: An Iterative Reweight-then-Optimize Approach Iterative preference learning from human feedback: Bridging theory and practice for rlhf under kl- constraint

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:08:15.594361Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T19:08:14.509171Z digest=sha256:146be1581c33c7cf2e274162763c5f17240f1a7464955afb3294f39f8f9cc70d

Observation 16303d39-c57b-4079-bc76-5640df34b622 · outbound

This paper cites Is DPO Superior to PPO for LLM Alignment? A Comprehensive Study.

Aligning Frozen LLMs by Reinforcement Learning: An Iterative Reweight-then-Optimize Approach Is DPO Superior to PPO for LLM Alignment? A Comprehensive Study

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-15T19:08:14.515292Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:08:14.515292Z digest=sha256:ff843ba406e5a39acdf5e3b89fa22fe300996db436dad5a817f3bdfee0a9d936

Observation c741543d-1997-4ca3-bf7b-67c408694c59 · outbound

This paper cites GenARM: Reward Guided Generation with Autoregressive Reward Model for Test-time Alignment.

Aligning Frozen LLMs by Reinforcement Learning: An Iterative Reweight-then-Optimize Approach GenARM: Reward Guided Generation with Autoregressive Reward Model for Test-time Alignment

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-15T19:08:14.522546Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:08:14.522546Z digest=sha256:8f5ef591c81b2cdf4688f65d35dd8453c06b88278a58d9ba1879be1ec78d23f9

Observation d26b7b2a-f374-4450-9ec2-af2f58b57e35 · outbound

This paper cites FUDGE: Controlled Text Generation With Future Discriminators.

Aligning Frozen LLMs by Reinforcement Learning: An Iterative Reweight-then-Optimize Approach FUDGE: Controlled Text Generation With Future Discriminators

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-15T19:08:14.529371Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:08:14.529371Z digest=sha256:56ec52b911612bd543258b64faa5882e256a9f5b460695814780ed4c63467de6

Observation b5696eaf-536f-45ae-8674-e416f20989c5 · outbound

This paper cites Llm alignment through successive policy re-weighting (spr).

Aligning Frozen LLMs by Reinforcement Learning: An Iterative Reweight-then-Optimize Approach Llm alignment through successive policy re-weighting (spr)

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:08:15.574252Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T19:08:14.535725Z digest=sha256:d29ca606dd1cad4e610039f7c43ef27ede305cb7b2004aba3aed87aec024d358

Observation ae25fa4c-a574-4b60-9738-0fd31626139b · outbound

This paper cites Weak-to-Strong Search: Align Large Language Models via Searching over Small Language Models.

Aligning Frozen LLMs by Reinforcement Learning: An Iterative Reweight-then-Optimize Approach Weak-to-Strong Search: Align Large Language Models via Searching over Small Language Models

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-15T19:08:14.541348Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T19:08:14.541348Z digest=sha256:a5a9b3c2c81e62598eeb08b17fcf720277c66bf776e7ba9c478ebcf1c0b40592

Observation 70baa008-3fdb-4b7c-b6ef-c1acef246b2e · outbound

This paper cites Starling-7b: Improving llm helpfulness & harmlessness with rlaif.2023, 2023.

Aligning Frozen LLMs by Reinforcement Learning: An Iterative Reweight-then-Optimize Approach Starling-7b: Improving llm helpfulness & harmlessness with rlaif.2023, 2023

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:08:15.553701Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T19:08:14.546908Z digest=sha256:fd21cde4b37c2c32b00c05367e50ad745f124292f0d225ca1c0b2a91cf271389

Observation 7ec3185e-4a7d-4051-88ce-a8b0f99235ff · outbound

This paper cites However, optimizing these algorithms for peak performance requires substantial effort and resources, which are often beyond the reach of the open-source community.

Aligning Frozen LLMs by Reinforcement Learning: An Iterative Reweight-then-Optimize Approach However, optimizing these algorithms for peak performance requires substantial effort and resources, which are often beyond the reach of the open-source community

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:08:15.533244Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T19:08:14.553567Z digest=sha256:5b6a2a242f165d0ffd7843dd53b5595ce9f53c42178c0167ca7d206d9dac7aeb

Observation edf01e89-e595-4e80-acd8-c18e1853083e · outbound

This paper cites an unresolved cited work.

Aligning Frozen LLMs by Reinforcement Learning: An Iterative Reweight-then-Optimize Approach Unresolved cited work

Reference 56

Resolution
unresolved
raw_fallback, observed 2026-08-15T19:08:15.514386Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T19:08:14.560736Z digest=sha256:acb01348ca333c7bf7db570e76f1ee5fd969842346b48f6a21977ffe5f04cda3

Observation 8b0a2116-d1f3-4836-ab67-a5de1c6dad5d · outbound

This paper cites an unresolved cited work.

Aligning Frozen LLMs by Reinforcement Learning: An Iterative Reweight-then-Optimize Approach Unresolved cited work

Reference 57

Resolution
unresolved
raw_fallback, observed 2026-08-15T19:08:15.496566Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T19:08:14.567968Z digest=sha256:d457690737c8563ecbd5a1323c02497f77771ebe42f4f54b5cfe31faaecebfe2

Observation c279e0b0-8cb8-4173-8dee-b642b7fa5d73 · outbound

This paper cites A"or "B"to indicate your choice. Your response should use the format: Comparison: <one-sentence comparison and explanation> Preferred: <.

Aligning Frozen LLMs by Reinforcement Learning: An Iterative Reweight-then-Optimize Approach A"or "B"to indicate your choice. Your response should use the format: Comparison: <one-sentence comparison and explanation> Preferred: <

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T19:08:15.477168Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T19:08:14.574756Z digest=sha256:5ab7b894284621f019cdf60b7593ec3ca3ca2783aeb453a4921cce4df4683407

Pith citing papers

No inbound Pith citation observations are available.