Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T19:08:14.574756Z
Paper Citation Record · LEDGER
As of 19 August 2026, this Paper Citation Record lists 58 of 58 outbound references and 0 inbound Pith citation observations for arXiv:2506.17828.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T19:08:14.574756Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
58 of 58 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation ee611819-3f7e-465d-8532-0b70afc8d4e7 · outbound
Aligning Frozen LLMs by Reinforcement Learning: An Iterative Reweight-then-Optimize Approach GPT-4 Technical Report
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 50cbe863-13a2-45e1-a89f-22ae5b59ea57 · outbound
Aligning Frozen LLMs by Reinforcement Learning: An Iterative Reweight-then-Optimize Approach Reinforcement learning: Theory and algorithms
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 744ad2b3-92b3-4073-a463-75b8c62d8368 · outbound
Aligning Frozen LLMs by Reinforcement Learning: An Iterative Reweight-then-Optimize Approach Back to Basics: Revisiting REINFORCE Style Optimization for Learning from Human Feedback in LLMs
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d84473d0-5e33-4456-af2e-9c13dd60aa38 · outbound
Aligning Frozen LLMs by Reinforcement Learning: An Iterative Reweight-then-Optimize Approach A General Theoretical Paradigm to Understand Learning from Human Preferences
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a4f7cb5e-e5a5-433d-a631-1ae0fe61e43f · outbound
Aligning Frozen LLMs by Reinforcement Learning: An Iterative Reweight-then-Optimize Approach Qwen Technical Report
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 81b0bc1d-1216-4e58-8df1-c1690ef2a97d · outbound
Aligning Frozen LLMs by Reinforcement Learning: An Iterative Reweight-then-Optimize Approach Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f1a7550d-660c-4f82-afc0-44c3f8f2f27a · outbound
Aligning Frozen LLMs by Reinforcement Learning: An Iterative Reweight-then-Optimize Approach A survey of monte carlo tree search methods.IEEE Transactions on Computational Intelligence and AI in games, 4(1):1–43, 2012
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e2610822-92da-4561-a7b9-a48eb13a532f · outbound
Aligning Frozen LLMs by Reinforcement Learning: An Iterative Reweight-then-Optimize Approach Transfer Q Star: Principled Decoding for LLM Alignment
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a3997ffc-3c70-4c65-8f76-584313a894f9 · outbound
Aligning Frozen LLMs by Reinforcement Learning: An Iterative Reweight-then-Optimize Approach Dataset Reset Policy Optimization for RLHF
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dd6790e4-3450-4c44-b3ef-67a9dd4cc1d6 · outbound
Aligning Frozen LLMs by Reinforcement Learning: An Iterative Reweight-then-Optimize Approach Stop summation: Min-form credit assignment is all process reward model needs for reasoning.arXiv preprint arXiv:2504.15275, 2025
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4c16a787-d391-4da9-94da-d5f1daf4f6d3 · outbound
Aligning Frozen LLMs by Reinforcement Learning: An Iterative Reweight-then-Optimize Approach Ultrafeedback: Boosting language models with high-quality feedback
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 79b5424d-9322-432d-889b-622567018dc3 · outbound
Aligning Frozen LLMs by Reinforcement Learning: An Iterative Reweight-then-Optimize Approach RAFT: Reward rAnked FineTuning for Generative Foundation Model Alignment
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c1fcae9a-7bbd-4ac7-99fc-69b3a4ae3da9 · outbound
Aligning Frozen LLMs by Reinforcement Learning: An Iterative Reweight-then-Optimize Approach Rlhf workflow: From reward modeling to online rlhf, 2024
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e719762f-c6b8-431c-99fb-b21406c357b9 · outbound
Aligning Frozen LLMs by Reinforcement Learning: An Iterative Reweight-then-Optimize Approach Length-Controlled AlpacaEval: A Simple Way to Debias Automatic Evaluators
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2fc67dec-5e38-4bc8-9f2f-9125c1729522 · outbound
Aligning Frozen LLMs by Reinforcement Learning: An Iterative Reweight-then-Optimize Approach Alpacafarm: A simulation framework for methods that learn from human feedback.Advances in Neural Information Processing Systems, 36:30039–30069, 2023
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 7284d86f-6f80-4f70-baec-d5b67ccac989 · outbound
Aligning Frozen LLMs by Reinforcement Learning: An Iterative Reweight-then-Optimize Approach Stop Regressing: Training Value Functions via Classification for Scalable Deep RL
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dedde5e3-c863-47fd-a951-c8cbe2efcce2 · outbound
Aligning Frozen LLMs by Reinforcement Learning: An Iterative Reweight-then-Optimize Approach Scaling laws for reward model overoptimization
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 89640b75-cedc-4d30-a588-1a24a1e32584 · outbound
Aligning Frozen LLMs by Reinforcement Learning: An Iterative Reweight-then-Optimize Approach DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 36272bd8-8096-4d01-9205-4825fe02531c · outbound
Aligning Frozen LLMs by Reinforcement Learning: An Iterative Reweight-then-Optimize Approach Value Augmented Sampling for Language Model Alignment and Personalization
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b1e45707-df66-4644-855f-3f21030122be · outbound
Aligning Frozen LLMs by Reinforcement Learning: An Iterative Reweight-then-Optimize Approach Inverse preference learning: Preference-based rl without a reward function
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation b43fc796-254a-4e5c-ae5d-95b7debcbc94 · outbound
Aligning Frozen LLMs by Reinforcement Learning: An Iterative Reweight-then-Optimize Approach Deal: Decoding-time alignment for large language models
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bded4ea0-0e61-4bd1-9592-25d177f931d7 · outbound
Aligning Frozen LLMs by Reinforcement Learning: An Iterative Reweight-then-Optimize Approach The N+ Implementation Details of RLHF with PPO: A Case Study on TL;DR Summarization
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 877e0763-f7ec-4693-8cf1-43712e2fe91b · outbound
Aligning Frozen LLMs by Reinforcement Learning: An Iterative Reweight-then-Optimize Approach LLM-Blender: Ensembling Large Language Models with Pairwise Ranking and Generative Fusion
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2c522dfd-c841-4e6d-82de-c703bdd4596f · outbound
Aligning Frozen LLMs by Reinforcement Learning: An Iterative Reweight-then-Optimize Approach A natural policy gradient.Advances in neural information processing systems, 14, 2001
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0db04d10-d604-4cc7-a8b4-31080f448bea · outbound
Aligning Frozen LLMs by Reinforcement Learning: An Iterative Reweight-then-Optimize Approach ARGS: Alignment as Reward-Guided Search
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c8aee78b-af36-4b8f-812f-3d206e6dae75 · outbound
Aligning Frozen LLMs by Reinforcement Learning: An Iterative Reweight-then-Optimize Approach Aligning Large Language Models with Representation Editing: A Control Perspective
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a8c9afe5-3b56-498f-9129-1c8a0280e79a · outbound
Aligning Frozen LLMs by Reinforcement Learning: An Iterative Reweight-then-Optimize Approach Optimization Issues in KL-Constrained Approximate Policy Iteration
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c418a9ce-776a-47e1-9aec-d7e8708e6590 · outbound
Aligning Frozen LLMs by Reinforcement Learning: An Iterative Reweight-then-Optimize Approach Cascade Reward Sampling for Efficient Decoding-Time Alignment
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 14f6de3c-da9f-46e4-b8b5-55009934b19f · outbound
Aligning Frozen LLMs by Reinforcement Learning: An Iterative Reweight-then-Optimize Approach Alpacaeval: An automatic evaluator of instruction-following models, 2023
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5f93dc1f-f411-4cc7-a57d-744f866108d0 · outbound
Aligning Frozen LLMs by Reinforcement Learning: An Iterative Reweight-then-Optimize Approach Let’s verify step by step
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b7b2f6b1-c043-4007-97d4-cac7186fb370 · outbound
Aligning Frozen LLMs by Reinforcement Learning: An Iterative Reweight-then-Optimize Approach Inference-Time Language Model Alignment via Integrated Value Guidance
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3ed8fc66-0ea9-4060-8d1c-948040777f5c · outbound
Aligning Frozen LLMs by Reinforcement Learning: An Iterative Reweight-then-Optimize Approach Controlled Decoding from Language Models
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 87f5d692-b0d3-47fb-b503-187c38a5abe7 · outbound
Aligning Frozen LLMs by Reinforcement Learning: An Iterative Reweight-then-Optimize Approach WebGPT: Browser-assisted question-answering with human feedback
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f95e0cbb-e569-4a80-9640-34b721cd5cb6 · outbound
Aligning Frozen LLMs by Reinforcement Learning: An Iterative Reweight-then-Optimize Approach Training language models to follow instructions with human feedback.Advances in Neural Information Processing Systems, 35:27730–27744, 2022
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ea6f7082-3051-4f71-9b59-378044cee4f9 · outbound
Aligning Frozen LLMs by Reinforcement Learning: An Iterative Reweight-then-Optimize Approach TreeBoN: Enhancing Inference-Time Alignment with Speculative Tree-Search and Best-of-N Sampling
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f814bcac-3f49-4600-8840-29fd9fe3bdbc · outbound
Aligning Frozen LLMs by Reinforcement Learning: An Iterative Reweight-then-Optimize Approach From $r$ to $Q^*$: Your Language Model is Secretly a Q-Function
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ee815ad4-78de-4461-a074-aa116b52515c · outbound
Aligning Frozen LLMs by Reinforcement Learning: An Iterative Reweight-then-Optimize Approach Direct preference optimization: Your language model is secretly a reward model.Advances in Neural Information Processing Systems, 36, 2024
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9b9875b7-c12e-4415-a1b4-fdc0f9ec9713 · outbound
Aligning Frozen LLMs by Reinforcement Learning: An Iterative Reweight-then-Optimize Approach Trust Region Policy Optimization
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e7cf550e-8b3a-44d2-8867-db4b01af33c8 · outbound
Aligning Frozen LLMs by Reinforcement Learning: An Iterative Reweight-then-Optimize Approach High-Dimensional Continuous Control Using Generalized Advantage Estimation
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7637d207-7007-465c-97e8-b52741fe1e5f · outbound
Aligning Frozen LLMs by Reinforcement Learning: An Iterative Reweight-then-Optimize Approach Proximal Policy Optimization Algorithms
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f0d89a27-f388-4cf0-a80b-f992c655430f · outbound
Aligning Frozen LLMs by Reinforcement Learning: An Iterative Reweight-then-Optimize Approach Adaptive trust region policy optimization: Global convergence and faster rates for regularized mdps
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation f9e936f6-3c4e-4d46-bac2-8476db04cf37 · outbound
Aligning Frozen LLMs by Reinforcement Learning: An Iterative Reweight-then-Optimize Approach DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e2b8e3f7-fe23-462d-825a-cd7375a0d7cf · outbound
Aligning Frozen LLMs by Reinforcement Learning: An Iterative Reweight-then-Optimize Approach Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 199d94c2-2516-4795-abfe-d4d2ee6e611e · outbound
Aligning Frozen LLMs by Reinforcement Learning: An Iterative Reweight-then-Optimize Approach Hybrid RL: Using Both Offline and Online Data Can Make RL Efficient
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3492c259-5465-469b-a9c4-d284785595eb · outbound
Aligning Frozen LLMs by Reinforcement Learning: An Iterative Reweight-then-Optimize Approach Learning to summarize with human feedback.Advances in Neural Information Processing Systems, 33:3008–3021, 2020
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bdc53125-8e08-49f7-80c8-b67f95e85451 · outbound
Aligning Frozen LLMs by Reinforcement Learning: An Iterative Reweight-then-Optimize Approach On the convergence rates of policy gradient methods.Journal of Machine Learning Research, 23(282):1–36, 2022
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5dfb5fc4-6623-405c-9e18-0c96c5b66eea · outbound
Aligning Frozen LLMs by Reinforcement Learning: An Iterative Reweight-then-Optimize Approach Monte Carlo Tree Search Boosts Reasoning via Iterative Preference Learning
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 806b563c-2708-49fb-bfe1-a9a2e1ecfcbe · outbound
Aligning Frozen LLMs by Reinforcement Learning: An Iterative Reweight-then-Optimize Approach Iterative preference learning from human feedback: Bridging theory and practice for rlhf under kl- constraint
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 16303d39-c57b-4079-bc76-5640df34b622 · outbound
Aligning Frozen LLMs by Reinforcement Learning: An Iterative Reweight-then-Optimize Approach Is DPO Superior to PPO for LLM Alignment? A Comprehensive Study
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c741543d-1997-4ca3-bf7b-67c408694c59 · outbound
Aligning Frozen LLMs by Reinforcement Learning: An Iterative Reweight-then-Optimize Approach GenARM: Reward Guided Generation with Autoregressive Reward Model for Test-time Alignment
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d26b7b2a-f374-4450-9ec2-af2f58b57e35 · outbound
Aligning Frozen LLMs by Reinforcement Learning: An Iterative Reweight-then-Optimize Approach FUDGE: Controlled Text Generation With Future Discriminators
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b5696eaf-536f-45ae-8674-e416f20989c5 · outbound
Aligning Frozen LLMs by Reinforcement Learning: An Iterative Reweight-then-Optimize Approach Llm alignment through successive policy re-weighting (spr)
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation ae25fa4c-a574-4b60-9738-0fd31626139b · outbound
Aligning Frozen LLMs by Reinforcement Learning: An Iterative Reweight-then-Optimize Approach Weak-to-Strong Search: Align Large Language Models via Searching over Small Language Models
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 70baa008-3fdb-4b7c-b6ef-c1acef246b2e · outbound
Aligning Frozen LLMs by Reinforcement Learning: An Iterative Reweight-then-Optimize Approach Starling-7b: Improving llm helpfulness & harmlessness with rlaif.2023, 2023
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 7ec3185e-4a7d-4051-88ce-a8b0f99235ff · outbound
Aligning Frozen LLMs by Reinforcement Learning: An Iterative Reweight-then-Optimize Approach However, optimizing these algorithms for peak performance requires substantial effort and resources, which are often beyond the reach of the open-source community
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation edf01e89-e595-4e80-acd8-c18e1853083e · outbound
Aligning Frozen LLMs by Reinforcement Learning: An Iterative Reweight-then-Optimize Approach Unresolved cited work
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation 8b0a2116-d1f3-4836-ab67-a5de1c6dad5d · outbound
Aligning Frozen LLMs by Reinforcement Learning: An Iterative Reweight-then-Optimize Approach Unresolved cited work
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
Observation c279e0b0-8cb8-4173-8dee-b642b7fa5d73 · outbound
Aligning Frozen LLMs by Reinforcement Learning: An Iterative Reweight-then-Optimize Approach A"or "B"to indicate your choice. Your response should use the format: Comparison: <one-sentence comparison and explanation> Preferred: <
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.
No inbound Pith citation observations are available.