Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T22:28:10.266576Z
Paper Citation Record · LEDGER
As of 22 August 2026, this Paper Citation Record lists 55 of 55 outbound references and 12 inbound Pith citation observations for arXiv:2506.21495.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T22:28:10.266576Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-04T09:50:46.549563Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-04T20:40:08.218379Z
55 of 55 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 344a4cba-33ab-4db7-a4f1-3f43136360d3 · outbound
Bridging Offline and Online Reinforcement Learning for LLMs write newline
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b4279561-afe0-46a0-89e5-01b5919863f3 · outbound
Bridging Offline and Online Reinforcement Learning for LLMs fairseq2, 2023
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation b781d302-fe36-4b2f-b531-b915a5c70a9d · outbound
Bridging Offline and Online Reinforcement Learning for LLMs Preference learning algorithms do not learn preference rankings
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation f37495a5-d296-46c0-b9df-2c80c0b198a9 · outbound
Bridging Offline and Online Reinforcement Learning for LLMs Self-Play Fine-Tuning Converts Weak Language Models to Strong Language Models
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b7ef05f0-2c6f-407e-8e1f-2280965fa0c9 · outbound
Bridging Offline and Online Reinforcement Learning for LLMs Deep reinforcement learning from human preferences
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fdc39dda-ea42-4647-896f-4e89e297eccf · outbound
Bridging Offline and Online Reinforcement Learning for LLMs The Llama 3 Herd of Models
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f133324e-40d8-4c15-997e-ea96aa6c6069 · outbound
Bridging Offline and Online Reinforcement Learning for LLMs Length-Controlled AlpacaEval: A Simple Way to Debias Automatic Evaluators
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2a3fc58e-c16e-4c0d-b117-3ac588379246 · outbound
Bridging Offline and Online Reinforcement Learning for LLMs Athene-70b: Redefining the boundaries of post-training for open models, July 2024 a
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation dafde6d9-7c3b-4dae-ac77-f8f20c1952cf · outbound
Bridging Offline and Online Reinforcement Learning for LLMs How to Evaluate Reward Models for RLHF
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cebb86d3-51cf-482d-a85e-7cc9aa5a59fb · outbound
Bridging Offline and Online Reinforcement Learning for LLMs Scaling laws for reward model overoptimization
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2cc140a8-6a72-4d09-a3b9-bf86a18b7246 · outbound
Bridging Offline and Online Reinforcement Learning for LLMs DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 84f5acc2-ba3b-43e3-946b-d31459bd35a3 · outbound
Bridging Offline and Online Reinforcement Learning for LLMs Direct Language Model Alignment from Online AI Feedback
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 587f2705-7c85-4e90-af71-3c6b3ee6d799 · outbound
Bridging Offline and Online Reinforcement Learning for LLMs Measuring Mathematical Problem Solving With the MATH Dataset
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8bc5fe93-ac70-44e6-944e-71c4a0713743 · outbound
Bridging Offline and Online Reinforcement Learning for LLMs ORPO : Monolithic preference optimization without reference model
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 648e4f9b-ba8a-4ae4-ad7c-fd7242c24883 · outbound
Bridging Offline and Online Reinforcement Learning for LLMs Lo RA : Low-rank adaptation of large language models
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bb5cea72-367c-44e7-ab17-1f8503bf2a32 · outbound
Bridging Offline and Online Reinforcement Learning for LLMs Gonzalez, Hao Zhang, and Ion Stoica
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5a1178cc-fecf-4e77-8d95-f8d3fd2ee581 · outbound
Bridging Offline and Online Reinforcement Learning for LLMs Tulu 3: Pushing Frontiers in Open Language Model Post-Training
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 53e6dc1b-db68-4490-854b-f79bda263def · outbound
Bridging Offline and Online Reinforcement Learning for LLMs Numinamath: The largest public dataset in ai4maths with 860k pairs of competition math problems and solutions
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 9745ab9e-2f20-445d-8ffa-8f388e7736f2 · outbound
Bridging Offline and Online Reinforcement Learning for LLMs From live data to high-quality benchmarks: The arena-hard pipeline, April 2024 b
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 54b4cd45-973e-41ea-a2e1-ba3a182c92cd · outbound
Bridging Offline and Online Reinforcement Learning for LLMs From Crowdsourced Data to High-Quality Benchmarks: Arena-Hard and BenchBuilder Pipeline
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 73c2a475-43b2-41d8-a0e4-231a0a7e259a · outbound
Bridging Offline and Online Reinforcement Learning for LLMs Self-Alignment with Instruction Backtranslation
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bf53e613-46e1-4248-89ae-94c0df82f4d3 · outbound
Bridging Offline and Online Reinforcement Learning for LLMs Hashimoto
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation e8167f2a-9bfc-49f2-aa9c-e62e88efe5ac · outbound
Bridging Offline and Online Reinforcement Learning for LLMs Let's Verify Step by Step
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d68b4f96-d8a1-4c78-a8f5-74ec1ec203de · outbound
Bridging Offline and Online Reinforcement Learning for LLMs Statistical Rejection Sampling Improves Preference Optimization
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 627f0544-a507-48b9-8e8d-b5dd31932f45 · outbound
Bridging Offline and Online Reinforcement Learning for LLMs Understanding r1-zero-like training: A critical perspective, 2025
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 118a8369-2f59-4425-b628-9df5572ca0e0 · outbound
Bridging Offline and Online Reinforcement Learning for LLMs Mixtral of experts: A high quality sparse mixture-of-experts
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 1af63203-9798-4f1c-948d-469e6525b206 · outbound
Bridging Offline and Online Reinforcement Learning for LLMs Ray: A distributed framework for emerging \ AI \ applications
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation ceaffc31-60bd-42df-a8c3-a4578692c358 · outbound
Bridging Offline and Online Reinforcement Learning for LLMs Training language models to follow instructions with human feedback
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d924b7eb-7f25-42fe-836d-3b7a7fe23925 · outbound
Bridging Offline and Online Reinforcement Learning for LLMs West-of-N: Synthetic Preferences for Self-Improving Reward Models
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6e887785-451e-429b-81b5-183f57a09fcb · outbound
Bridging Offline and Online Reinforcement Learning for LLMs Iterative reasoning preference optimization
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8524ab45-8871-43d7-9e42-487e27f30440 · outbound
Bridging Offline and Online Reinforcement Learning for LLMs Disentangling length from quality in direct preference optimization
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 561a7e33-73f4-4e69-9cfa-a42e91b4fffd · outbound
Bridging Offline and Online Reinforcement Learning for LLMs Disentangling Length from Quality in Direct Preference Optimization
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 24110d55-6030-444d-b2d1-75b55ad7a8e0 · outbound
Bridging Offline and Online Reinforcement Learning for LLMs Online dpo: Online direct preference optimization with fast-slow chasing, 2024
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 02239330-82da-4ba0-a3d5-42fa36a5edc1 · outbound
Bridging Offline and Online Reinforcement Learning for LLMs Direct preference optimization: Your language model is secretly a reward model
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5b8ade72-fba5-41ec-a837-618165e1337b · outbound
Bridging Offline and Online Reinforcement Learning for LLMs Direct preference optimization: Your language model is secretly a reward model
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eb2362c4-a505-453e-9bfd-7166795d23e6 · outbound
Bridging Offline and Online Reinforcement Learning for LLMs Trust region policy optimization
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3120a6ea-4cc8-4632-a8fa-ac9bc2109dc1 · outbound
Bridging Offline and Online Reinforcement Learning for LLMs Proximal Policy Optimization Algorithms
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 54604467-07e4-4c24-881d-947e81676d91 · outbound
Bridging Offline and Online Reinforcement Learning for LLMs DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a80e91d1-6705-43ff-8fc0-1e318c2c7c9a · outbound
Bridging Offline and Online Reinforcement Learning for LLMs Welcome to the era of experience
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 591e2016-2b65-4747-896a-17aba7f47326 · outbound
Bridging Offline and Online Reinforcement Learning for LLMs A long way to go: Investigating length correlations in RLHF
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 346d1015-0e56-44ca-b54d-f209525b007f · outbound
Bridging Offline and Online Reinforcement Learning for LLMs LLaMA: Open and Efficient Foundation Language Models
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ccc5ca8f-c8c3-425d-becb-b12fdd54ce2e · outbound
Bridging Offline and Online Reinforcement Learning for LLMs Llama 2: Open Foundation and Fine-Tuned Chat Models
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0d1cd526-46cd-488e-87b3-c0ba5b3fd08a · outbound
Bridging Offline and Online Reinforcement Learning for LLMs Zephyr: Direct Distillation of LM Alignment
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ce183029-a436-4bfc-a777-5425946aa887 · outbound
Bridging Offline and Online Reinforcement Learning for LLMs Thinking LLMs: General Instruction Following with Thought Generation
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation baeaa4fb-5bd9-4b9f-84da-61c5d5ffab84 · outbound
Bridging Offline and Online Reinforcement Learning for LLMs Meta-Rewarding Language Models: Self-Improving Alignment with LLM-as-a-Meta-Judge
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aadb1a06-6ba7-44f1-8ad0-d2b140535341 · outbound
Bridging Offline and Online Reinforcement Learning for LLMs Iterative Preference Learning from Human Feedback: Bridging Theory and Practice for RLHF under KL-Constraint
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a00e9344-8a24-475a-b719-bad7539cd25c · outbound
Bridging Offline and Online Reinforcement Learning for LLMs Gibbs sampling from human feedback: A provable kl-constrained framework for rlhf
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 4d8e1c05-82c5-420f-9591-4dd36d84a8f6 · outbound
Bridging Offline and Online Reinforcement Learning for LLMs WizardLM: Empowering large pre-trained language models to follow complex instructions
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ab9dd718-b78d-491d-9789-0004f8b42a36 · outbound
Bridging Offline and Online Reinforcement Learning for LLMs Some things are more CRINGE than others: Iterative Preference Optimization with the Pairwise Cringe Loss
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6a88cb36-aac5-4a67-8291-9c4fc629a398 · outbound
Bridging Offline and Online Reinforcement Learning for LLMs Is DPO Superior to PPO for LLM Alignment? A Comprehensive Study
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 757a19e2-111d-4b8c-b2d4-47fbf950d92b · outbound
Bridging Offline and Online Reinforcement Learning for LLMs BPO : Staying close to the behavior LLM creates better online LLM alignment
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 41c10192-46cd-49c6-aadf-e0e4536b6240 · outbound
Bridging Offline and Online Reinforcement Learning for LLMs Self-Rewarding Language Models
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e89d6539-c0bf-45ad-b512-df53c68879fa · outbound
Bridging Offline and Online Reinforcement Learning for LLMs WildChat: 1M ChatGPT Interaction Logs in the Wild
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 23af5f03-b49c-4648-8428-f8d344d54282 · outbound
Bridging Offline and Online Reinforcement Learning for LLMs Lima: Less is more for alignment
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c4852346-3bd9-40a7-a77f-56969f60f120 · outbound
Bridging Offline and Online Reinforcement Learning for LLMs Fine-Tuning Language Models from Human Preferences
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3ab7ef80-1dbb-46a4-821c-b519340abb04 · inbound
Task-Dependent Evaluation of LLM Output Homogenization: A Taxonomy-Guided Framework Bridging Offline and Online Reinforcement Learning for LLMs
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation d3c3b757-9212-422a-b8f0-97949b079396 · inbound
Beyond Correctness: Rewarding Faithful Reasoning in Retrieval-Augmented Generation Bridging Offline and Online Reinforcement Learning for LLMs
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 202772d5-4b27-4529-a7b7-41fb2b972c45 · inbound
Safety Alignment of LMs via Non-cooperative Games Bridging Offline and Online Reinforcement Learning for LLMs
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 089ad325-e523-45ab-a2c4-a1e5469647d6 · inbound
OGER: A Robust Offline-Guided Exploration Reward for Hybrid Reinforcement Learning Bridging Offline and Online Reinforcement Learning for LLMs
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 7465628b-7cca-4c32-8f66-7c8684b401f7 · inbound
Smaller Models are Natural Explorers for Policy-Level Diversity in GRPO Bridging Offline and Online Reinforcement Learning for LLMs
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation b7230ccc-d595-499e-98fc-4d0eb1787025 · inbound
Multi$^2$: Hierarchical Multi-Agent Decision-Making with LLM-Based Agents in Interactive Environments Bridging Offline and Online Reinforcement Learning for LLMs
Reference 72
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 507f7013-5128-4117-8a77-4d64895355cc · inbound
Multi$^2$: Hierarchical Multi-Agent Decision-Making with LLM-Based Agents in Interactive Environments Bridging Offline and Online Reinforcement Learning for LLMs
Reference 72
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cc445246-eaec-4226-b5fc-f4fefe4861de · inbound
Step-by-Step Optimization-like Reasoning in LLMs over Expanding Search Spaces Bridging Offline and Online Reinforcement Learning for LLMs
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 7ee3a065-8118-4b92-9708-1f8e7d080591 · inbound
Modularized Reinforcement Learning on LLMs: From MDP Creation to Exploration and Learning Bridging Offline and Online Reinforcement Learning for LLMs
Reference 97
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 7e0c52fc-5ece-4b58-9779-38d28c66da07 · inbound
Autodata: An agentic data scientist to create high quality synthetic data Bridging Offline and Online Reinforcement Learning for LLMs
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation f177b23e-6f8c-4e22-b30a-ef7b1335dae5 · inbound
Autodata: An agentic data scientist to create high quality synthetic data Bridging Offline and Online Reinforcement Learning for LLMs
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.
Observation 9cbe79a4-8756-444d-a1b3-b98c5f4f1484 · inbound
LeVo 2: Stable and Melodious Song Generation via Hierarchical Representation Modeling and Progressive Post-Training Bridging Offline and Online Reinforcement Learning for LLMs
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.