Pith. sign in

Paper Citation Record · LEDGER

Bridging Offline and Online Reinforcement Learning for LLMs

As of 22 August 2026, this Paper Citation Record lists 55 of 55 outbound references and 12 inbound Pith citation observations for arXiv:2506.21495.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.21495 v1

Coverage vector

measured 55 of 55 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T22:28:10.266576Z

measured 67 of 67 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 12 of 12 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T09:50:46.549563Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T20:40:08.218379Z

Reference resolution

55 of 55 outbound references displayed

  • verified exact1
  • verified fuzzy11
  • unresolved43
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 344a4cba-33ab-4db7-a4f1-3f43136360d3 · outbound

This paper cites write newline.

Bridging Offline and Online Reinforcement Learning for LLMs write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T22:28:04.879326Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:28:04.879326Z digest=sha256:615e2fbd29f4f84bdc15a747c58250422c97c9b03a49c127a8687db4f246dbe5

Observation b4279561-afe0-46a0-89e5-01b5919863f3 · outbound

This paper cites fairseq2, 2023.

Bridging Offline and Online Reinforcement Learning for LLMs fairseq2, 2023

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:28:12.612372Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-06T22:28:04.982424Z digest=sha256:25ac049784829c4cedc076709b15af19cb335dcf6e2a9cc00cba7c23cd023a62

Observation b781d302-fe36-4b2f-b531-b915a5c70a9d · outbound

This paper cites Preference learning algorithms do not learn preference rankings.

Bridging Offline and Online Reinforcement Learning for LLMs Preference learning algorithms do not learn preference rankings

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:28:12.399050Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-06T22:28:05.039479Z digest=sha256:a6b73de03f28cb49678eea56c5ab03a6459c80437293cc2c3518bf92bf0343ef

Observation f37495a5-d296-46c0-b9df-2c80c0b198a9 · outbound

This paper cites Self-Play Fine-Tuning Converts Weak Language Models to Strong Language Models.

Bridging Offline and Online Reinforcement Learning for LLMs Self-Play Fine-Tuning Converts Weak Language Models to Strong Language Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T22:28:05.128670Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:28:05.128670Z digest=sha256:10fecc8a4a3215091a55e2ae94ea551c7f770ea57e6103648c4e604e83115e7c

Observation b7ef05f0-2c6f-407e-8e1f-2280965fa0c9 · outbound

This paper cites Deep reinforcement learning from human preferences.

Bridging Offline and Online Reinforcement Learning for LLMs Deep reinforcement learning from human preferences

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T22:28:05.212164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:28:05.212164Z digest=sha256:3a8e3f37ad07ab8c3e4afbb901ff5ed79e9c022bdcd7de330a5fc88e26611dcb

Observation fdc39dda-ea42-4647-896f-4e89e297eccf · outbound

This paper cites The Llama 3 Herd of Models.

Bridging Offline and Online Reinforcement Learning for LLMs The Llama 3 Herd of Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T22:28:05.308808Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:28:05.308808Z digest=sha256:a833c503d7d8930758abe0ff3a2baa1ce39fff54fb5212f177126bb7a66302db

Observation f133324e-40d8-4c15-997e-ea96aa6c6069 · outbound

This paper cites Length-Controlled AlpacaEval: A Simple Way to Debias Automatic Evaluators.

Bridging Offline and Online Reinforcement Learning for LLMs Length-Controlled AlpacaEval: A Simple Way to Debias Automatic Evaluators

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T22:28:05.412620Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:28:05.412620Z digest=sha256:39258bf57aec1f092ed433fecb6a58d7293fba8869ae31ebf25e965651c8ee40

Observation 2a3fc58e-c16e-4c0d-b117-3ac588379246 · outbound

This paper cites Athene-70b: Redefining the boundaries of post-training for open models, July 2024 a.

Bridging Offline and Online Reinforcement Learning for LLMs Athene-70b: Redefining the boundaries of post-training for open models, July 2024 a

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:28:12.160526Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-06T22:28:05.547248Z digest=sha256:9ab9f0cd77fc32eeec891b734b5e0c57ff754bf55c97d21fdb9b9e4490451653

Observation dafde6d9-7c3b-4dae-ac77-f8f20c1952cf · outbound

This paper cites How to Evaluate Reward Models for RLHF.

Bridging Offline and Online Reinforcement Learning for LLMs How to Evaluate Reward Models for RLHF

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T22:28:05.645832Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:28:05.645832Z digest=sha256:6b182f4ba72ac21b8830ccaea72073d1750d192f487f1ab41bff1814d6a63e7f

Observation cebb86d3-51cf-482d-a85e-7cc9aa5a59fb · outbound

This paper cites Scaling laws for reward model overoptimization.

Bridging Offline and Online Reinforcement Learning for LLMs Scaling laws for reward model overoptimization

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T22:28:05.741024Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:28:05.741024Z digest=sha256:8b70c243814272c90080a6e8bbc3c476555b56c673ce78dc78e45d7790b01b19

Observation 2cc140a8-6a72-4d09-a3b9-bf86a18b7246 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Bridging Offline and Online Reinforcement Learning for LLMs DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T22:28:05.871341Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:28:05.871341Z digest=sha256:9bc4a038acf0427ba9b744b707fc2daf8aeed2091b13b00b5cc2b9c85a297e51

Observation 84f5acc2-ba3b-43e3-946b-d31459bd35a3 · outbound

This paper cites Direct Language Model Alignment from Online AI Feedback.

Bridging Offline and Online Reinforcement Learning for LLMs Direct Language Model Alignment from Online AI Feedback

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T22:28:05.981374Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:28:05.981374Z digest=sha256:630e58ec178be7b8b7604092f35737dd84c40a2df25bd7a63193108813379a2e

Observation 587f2705-7c85-4e90-af71-3c6b3ee6d799 · outbound

This paper cites Measuring Mathematical Problem Solving With the MATH Dataset.

Bridging Offline and Online Reinforcement Learning for LLMs Measuring Mathematical Problem Solving With the MATH Dataset

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T22:28:06.065494Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:28:06.065494Z digest=sha256:0cbda9f97b96c0bfe695fd532266e35c472377b1870cd1a36e5dfb7620dfc43f

Observation 8bc5fe93-ac70-44e6-944e-71c4a0713743 · outbound

This paper cites ORPO : Monolithic preference optimization without reference model.

Bridging Offline and Online Reinforcement Learning for LLMs ORPO : Monolithic preference optimization without reference model

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T22:28:06.143444Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:28:06.143444Z digest=sha256:b9e972b32a1d1db1c691bf2201b1c71038310e6adfca0fd3530b102caf556555

Observation 648e4f9b-ba8a-4ae4-ad7c-fd7242c24883 · outbound

This paper cites Lo RA : Low-rank adaptation of large language models.

Bridging Offline and Online Reinforcement Learning for LLMs Lo RA : Low-rank adaptation of large language models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T22:28:06.218419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:28:06.218419Z digest=sha256:fc0e659978392c34805a0419c33efaa6e4db78e209853e459b7fc786d7531981

Observation bb5cea72-367c-44e7-ab17-1f8503bf2a32 · outbound

This paper cites Gonzalez, Hao Zhang, and Ion Stoica.

Bridging Offline and Online Reinforcement Learning for LLMs Gonzalez, Hao Zhang, and Ion Stoica

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T22:28:06.322902Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:28:06.322902Z digest=sha256:d0a86ee28a291111dbc34f60dec303b049fca38dce8953f8fcbf0d82ccc7ef75

Observation 5a1178cc-fecf-4e77-8d95-f8d3fd2ee581 · outbound

This paper cites Tulu 3: Pushing Frontiers in Open Language Model Post-Training.

Bridging Offline and Online Reinforcement Learning for LLMs Tulu 3: Pushing Frontiers in Open Language Model Post-Training

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T22:28:06.421583Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:28:06.421583Z digest=sha256:df92774e363f4e611e4b587de208407fffe3fc55a4608693b14f06b7dab2c53e

Observation 53e6dc1b-db68-4490-854b-f79bda263def · outbound

This paper cites Numinamath: The largest public dataset in ai4maths with 860k pairs of competition math problems and solutions.

Bridging Offline and Online Reinforcement Learning for LLMs Numinamath: The largest public dataset in ai4maths with 860k pairs of competition math problems and solutions

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:28:12.015713Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-06T22:28:06.557537Z digest=sha256:fa729252d131de6b5ca05c836e063412480ee0d224aa30667a3df63276f666f7

Observation 9745ab9e-2f20-445d-8ffa-8f388e7736f2 · outbound

This paper cites From live data to high-quality benchmarks: The arena-hard pipeline, April 2024 b.

Bridging Offline and Online Reinforcement Learning for LLMs From live data to high-quality benchmarks: The arena-hard pipeline, April 2024 b

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:28:11.809112Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-06T22:28:06.681214Z digest=sha256:e9034bde3c93cbfa75e83c566b4e5f213b74384b839173df3ea8b64729e8be65

Observation 54b4cd45-973e-41ea-a2e1-ba3a182c92cd · outbound

This paper cites From Crowdsourced Data to High-Quality Benchmarks: Arena-Hard and BenchBuilder Pipeline.

Bridging Offline and Online Reinforcement Learning for LLMs From Crowdsourced Data to High-Quality Benchmarks: Arena-Hard and BenchBuilder Pipeline

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T22:28:06.760710Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:28:06.760710Z digest=sha256:947a867f7a05a38691f86a20192fb93503b7a6ba3a246465d60f6ae210628449

Observation 73c2a475-43b2-41d8-a0e4-231a0a7e259a · outbound

This paper cites Self-Alignment with Instruction Backtranslation.

Bridging Offline and Online Reinforcement Learning for LLMs Self-Alignment with Instruction Backtranslation

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T22:28:06.859418Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:28:06.859418Z digest=sha256:5e2810e55b32ba7b9e252e89b6aeb9d070440634024c7aa42755f359c102939a

Observation bf53e613-46e1-4248-89ae-94c0df82f4d3 · outbound

This paper cites Hashimoto.

Bridging Offline and Online Reinforcement Learning for LLMs Hashimoto

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:28:11.687633Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-06T22:28:06.946488Z digest=sha256:2c841aa2eeaac601ee34183f75de0de64beb4d718684f9d789ae6929e80c1c07

Observation e8167f2a-9bfc-49f2-aa9c-e62e88efe5ac · outbound

This paper cites Let's Verify Step by Step.

Bridging Offline and Online Reinforcement Learning for LLMs Let's Verify Step by Step

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T22:28:07.062465Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:28:07.062465Z digest=sha256:a698f5d531a030dcbef51884a0c26e341bcad6f66a0ab0ecf3fee122a91e18ae

Observation d68b4f96-d8a1-4c78-a8f5-74ec1ec203de · outbound

This paper cites Statistical Rejection Sampling Improves Preference Optimization.

Bridging Offline and Online Reinforcement Learning for LLMs Statistical Rejection Sampling Improves Preference Optimization

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T22:28:07.163268Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:28:07.163268Z digest=sha256:827941a75e3a2030d32444d69a9af1d4106b757c8718d99921678e56a6bdb072

Observation 627f0544-a507-48b9-8e8d-b5dd31932f45 · outbound

This paper cites Understanding r1-zero-like training: A critical perspective, 2025.

Bridging Offline and Online Reinforcement Learning for LLMs Understanding r1-zero-like training: A critical perspective, 2025

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T22:28:07.276662Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:28:07.276662Z digest=sha256:6dd7f5b27446cdea52b49d01d5d72edd78846e4bae4452d34c5567fc68f9d26e

Observation 118a8369-2f59-4425-b628-9df5572ca0e0 · outbound

This paper cites Mixtral of experts: A high quality sparse mixture-of-experts.

Bridging Offline and Online Reinforcement Learning for LLMs Mixtral of experts: A high quality sparse mixture-of-experts

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:28:11.541879Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-06T22:28:07.363506Z digest=sha256:a251b54adfd0298351be34a9815a922f65a7b28a32031a0f2b5e57efec56f184

Observation 1af63203-9798-4f1c-948d-469e6525b206 · outbound

This paper cites Ray: A distributed framework for emerging \ AI \ applications.

Bridging Offline and Online Reinforcement Learning for LLMs Ray: A distributed framework for emerging \ AI \ applications

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:28:11.357146Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-06T22:28:07.471713Z digest=sha256:1bd7eb49fedcf3b0485c91b5a2d024ce2e1d20f2b4f06eb267220d9105238916

Observation ceaffc31-60bd-42df-a8c3-a4578692c358 · outbound

This paper cites Training language models to follow instructions with human feedback.

Bridging Offline and Online Reinforcement Learning for LLMs Training language models to follow instructions with human feedback

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T22:28:07.577745Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:28:07.577745Z digest=sha256:f554f0904be64f48308e7aa3640f9c0e224f5d4056b2202b3de0bae204719497

Observation d924b7eb-7f25-42fe-836d-3b7a7fe23925 · outbound

This paper cites West-of-N: Synthetic Preferences for Self-Improving Reward Models.

Bridging Offline and Online Reinforcement Learning for LLMs West-of-N: Synthetic Preferences for Self-Improving Reward Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T22:28:07.646010Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:28:07.646010Z digest=sha256:19aed2fa04141c80d9891f1f324f5c04282cae455b305c4b2a608b6a15834059

Observation 6e887785-451e-429b-81b5-183f57a09fcb · outbound

This paper cites Iterative reasoning preference optimization.

Bridging Offline and Online Reinforcement Learning for LLMs Iterative reasoning preference optimization

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T22:28:07.735012Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:28:07.735012Z digest=sha256:816b0567a59267f3ee61ad2adfea536b707cb17e20575c69d867e2fc35689e74

Observation 8524ab45-8871-43d7-9e42-487e27f30440 · outbound

This paper cites Disentangling length from quality in direct preference optimization.

Bridging Offline and Online Reinforcement Learning for LLMs Disentangling length from quality in direct preference optimization

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T22:28:07.868238Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:28:07.868238Z digest=sha256:1a54e40c4d488ae1e318651e3cd438bcf68ccb0f64046e7f7af53dc9e2c5b72a

Observation 561a7e33-73f4-4e69-9cfa-a42e91b4fffd · outbound

This paper cites Disentangling Length from Quality in Direct Preference Optimization.

Bridging Offline and Online Reinforcement Learning for LLMs Disentangling Length from Quality in Direct Preference Optimization

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T22:28:07.964261Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:28:07.964261Z digest=sha256:41c95703668875776dc6a939fdfb698a3ceb72b1a2516fee42db7313376b3b90

Observation 24110d55-6030-444d-b2d1-75b55ad7a8e0 · outbound

This paper cites Online dpo: Online direct preference optimization with fast-slow chasing, 2024.

Bridging Offline and Online Reinforcement Learning for LLMs Online dpo: Online direct preference optimization with fast-slow chasing, 2024

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T22:28:08.036930Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:28:08.036930Z digest=sha256:3b8d0909e547d7c0e4ae9890dc238b612577b904c53007840caa9580cf54dff1

Observation 02239330-82da-4ba0-a3d5-42fa36a5edc1 · outbound

This paper cites Direct preference optimization: Your language model is secretly a reward model.

Bridging Offline and Online Reinforcement Learning for LLMs Direct preference optimization: Your language model is secretly a reward model

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T22:28:08.101830Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:28:08.101830Z digest=sha256:1c1d2a73ba9040e1e65e88203f0ea67ee06d7fe6df76a28effccc285c98ab916

Observation 5b8ade72-fba5-41ec-a837-618165e1337b · outbound

This paper cites Direct preference optimization: Your language model is secretly a reward model.

Bridging Offline and Online Reinforcement Learning for LLMs Direct preference optimization: Your language model is secretly a reward model

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T22:28:08.225651Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:28:08.225651Z digest=sha256:62ec166a09c96c8ccceb346d55b519cd46d6f2163bd98743937bba7afd31b6ea

Observation eb2362c4-a505-453e-9bfd-7166795d23e6 · outbound

This paper cites Trust region policy optimization.

Bridging Offline and Online Reinforcement Learning for LLMs Trust region policy optimization

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T22:28:08.319306Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:28:08.319306Z digest=sha256:fd4118a5d833d066065cd4746675c700351d1d3194566d5b148bc19b7c5b2c02

Observation 3120a6ea-4cc8-4632-a8fa-ac9bc2109dc1 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Bridging Offline and Online Reinforcement Learning for LLMs Proximal Policy Optimization Algorithms

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T22:28:08.531883Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:28:08.531883Z digest=sha256:8e6a3c56098e1f49f7ec2d098ca766b30a847223ce4c1bc292b8b2b7d5bfe22b

Observation 54604467-07e4-4c24-881d-947e81676d91 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Bridging Offline and Online Reinforcement Learning for LLMs DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T22:28:08.661390Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:28:08.661390Z digest=sha256:14f8dce2b49edfbb4ac0fbd859cd47308af715adef266bab522d74cbc17765f8

Observation a80e91d1-6705-43ff-8fc0-1e318c2c7c9a · outbound

This paper cites Welcome to the era of experience.

Bridging Offline and Online Reinforcement Learning for LLMs Welcome to the era of experience

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:28:11.153605Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-06T22:28:08.768169Z digest=sha256:67aeba384bae51a0e0665f7b7e51a0ea40990f35d5e8e767963effc04b8a309b

Observation 591e2016-2b65-4747-896a-17aba7f47326 · outbound

This paper cites A long way to go: Investigating length correlations in RLHF.

Bridging Offline and Online Reinforcement Learning for LLMs A long way to go: Investigating length correlations in RLHF

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:28:11.008634Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-06T22:28:08.855939Z digest=sha256:a8c4ecc4064b3fffc8bf4612004fad7516900ff39fd348405e702c5226e42a12

Observation 346d1015-0e56-44ca-b54d-f209525b007f · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Bridging Offline and Online Reinforcement Learning for LLMs LLaMA: Open and Efficient Foundation Language Models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T22:28:08.959424Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:28:08.959424Z digest=sha256:b3b10c38e9f6922b601d50ff09dda78f97d96105a027b2e29d90fb7a68574091

Observation ccc5ca8f-c8c3-425d-becb-b12fdd54ce2e · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Bridging Offline and Online Reinforcement Learning for LLMs Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T22:28:09.048795Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:28:09.048795Z digest=sha256:fd959a679077f6be1906278bec743d59497a0c56a20bcb2bcea68d445b419dab

Observation 0d1cd526-46cd-488e-87b3-c0ba5b3fd08a · outbound

This paper cites Zephyr: Direct Distillation of LM Alignment.

Bridging Offline and Online Reinforcement Learning for LLMs Zephyr: Direct Distillation of LM Alignment

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T22:28:09.134651Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:28:09.134651Z digest=sha256:f2efddd3a3b872283498e35ca2445ece2ed5678198d61a046546207c30b7e252

Observation ce183029-a436-4bfc-a777-5425946aa887 · outbound

This paper cites Thinking LLMs: General Instruction Following with Thought Generation.

Bridging Offline and Online Reinforcement Learning for LLMs Thinking LLMs: General Instruction Following with Thought Generation

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T22:28:09.217475Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:28:09.217475Z digest=sha256:e8c137ce03d3544ad6f6777dbf9f1a9a4955f506aae52e3cdc1c3ba38d81eec1

Observation baeaa4fb-5bd9-4b9f-84da-61c5d5ffab84 · outbound

This paper cites Meta-Rewarding Language Models: Self-Improving Alignment with LLM-as-a-Meta-Judge.

Bridging Offline and Online Reinforcement Learning for LLMs Meta-Rewarding Language Models: Self-Improving Alignment with LLM-as-a-Meta-Judge

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T22:28:09.310079Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:28:09.310079Z digest=sha256:f307b85aaa97166b608d498b059db224f92a3b5c4cc250c63da38102333981ca

Observation aadb1a06-6ba7-44f1-8ad0-d2b140535341 · outbound

This paper cites Iterative Preference Learning from Human Feedback: Bridging Theory and Practice for RLHF under KL-Constraint.

Bridging Offline and Online Reinforcement Learning for LLMs Iterative Preference Learning from Human Feedback: Bridging Theory and Practice for RLHF under KL-Constraint

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-06T22:28:09.397544Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:28:09.397544Z digest=sha256:754783bb906e91bfd81103e8ba12020c0c7642236c90cfa325bdddb0fe484e2e

Observation a00e9344-8a24-475a-b719-bad7539cd25c · outbound

This paper cites Gibbs sampling from human feedback: A provable kl-constrained framework for rlhf.

Bridging Offline and Online Reinforcement Learning for LLMs Gibbs sampling from human feedback: A provable kl-constrained framework for rlhf

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:28:10.850501Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-06T22:28:09.483479Z digest=sha256:49c7cc02855f6e2e59542c3265cdc6d4c07f266e9a62d4750c16c8ba0b7519fc

Observation 4d8e1c05-82c5-420f-9591-4dd36d84a8f6 · outbound

This paper cites WizardLM: Empowering large pre-trained language models to follow complex instructions.

Bridging Offline and Online Reinforcement Learning for LLMs WizardLM: Empowering large pre-trained language models to follow complex instructions

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-06T22:28:09.578482Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:28:09.578482Z digest=sha256:76ee0ec2ebd5d3308e1f3cda372adbf65b3b853941080fe5dbf8e619c336d8ca

Observation ab9dd718-b78d-491d-9789-0004f8b42a36 · outbound

This paper cites Some things are more CRINGE than others: Iterative Preference Optimization with the Pairwise Cringe Loss.

Bridging Offline and Online Reinforcement Learning for LLMs Some things are more CRINGE than others: Iterative Preference Optimization with the Pairwise Cringe Loss

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T22:28:09.691354Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:28:09.691354Z digest=sha256:2092821c7e813016d568e7d6ba29a73150a512ccd1071eda925ca77db5695798

Observation 6a88cb36-aac5-4a67-8291-9c4fc629a398 · outbound

This paper cites Is DPO Superior to PPO for LLM Alignment? A Comprehensive Study.

Bridging Offline and Online Reinforcement Learning for LLMs Is DPO Superior to PPO for LLM Alignment? A Comprehensive Study

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-06T22:28:09.790974Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:28:09.790974Z digest=sha256:6230f8d035cd19dd84d7a08f95ace75cc97124e6a775d42938735dda054bc7aa

Observation 757a19e2-111d-4b8c-b2d4-47fbf950d92b · outbound

This paper cites BPO : Staying close to the behavior LLM creates better online LLM alignment.

Bridging Offline and Online Reinforcement Learning for LLMs BPO : Staying close to the behavior LLM creates better online LLM alignment

Reference 52

Resolution
verified exact
doi, observed 2026-08-06T22:28:10.442692Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-08-06T22:28:09.869519Z digest=sha256:4935ad2ed5bd8cfac29725a5e74abc83ec1b9a85d6f65ba96129a13021900d31

Observation 41c10192-46cd-49c6-aadf-e0e4536b6240 · outbound

This paper cites Self-Rewarding Language Models.

Bridging Offline and Online Reinforcement Learning for LLMs Self-Rewarding Language Models

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-06T22:28:09.980703Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:28:09.980703Z digest=sha256:63c0ed124be68092487110e2268755d167c9cfb0393167dcfca6e5af41134de3

Observation e89d6539-c0bf-45ad-b512-df53c68879fa · outbound

This paper cites WildChat: 1M ChatGPT Interaction Logs in the Wild.

Bridging Offline and Online Reinforcement Learning for LLMs WildChat: 1M ChatGPT Interaction Logs in the Wild

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-06T22:28:10.083842Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:28:10.083842Z digest=sha256:ddbd63dc0925e3ae62466954787452ab7f8e799411368bd0f44e388f492fb0ed

Observation 23af5f03-b49c-4648-8428-f8d344d54282 · outbound

This paper cites Lima: Less is more for alignment.

Bridging Offline and Online Reinforcement Learning for LLMs Lima: Less is more for alignment

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-06T22:28:10.181387Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:28:10.181387Z digest=sha256:24a2b4fb592b695bf2aa4b8eb971fb114d7f2b40e238074d04eaaef21f73bbac

Observation c4852346-3bd9-40a7-a77f-56969f60f120 · outbound

This paper cites Fine-Tuning Language Models from Human Preferences.

Bridging Offline and Online Reinforcement Learning for LLMs Fine-Tuning Language Models from Human Preferences

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-06T22:28:10.266576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:28:10.266576Z digest=sha256:eca37eee785232bd4dbd6d46b2d16ecf24be2694a8b30a0d588a8d75212005b3

Pith citing papers

Observation 3ab7ef80-1dbb-46a4-821c-b519340abb04 · inbound

Task-Dependent Evaluation of LLM Output Homogenization: A Taxonomy-Guided Framework cites this paper.

Task-Dependent Evaluation of LLM Output Homogenization: A Taxonomy-Guided Framework Bridging Offline and Online Reinforcement Learning for LLMs

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-18T13:36:24.983666Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-05-18T13:34:22.790199Z digest=sha256:765a6b37bd55deccde342fdbf88ea4948dbc02f84f53f9c88232260969806b5b

Observation d3c3b757-9212-422a-b8f0-97949b079396 · inbound

Beyond Correctness: Rewarding Faithful Reasoning in Retrieval-Augmented Generation cites this paper.

Beyond Correctness: Rewarding Faithful Reasoning in Retrieval-Augmented Generation Bridging Offline and Online Reinforcement Learning for LLMs

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-04T09:50:46.549563Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T09:50:46.549563Z digest=sha256:ea111f92f56d7df7d82407f2bd4e758133d62ab001f58f0c6b7b080aa9f12db4

Observation 202772d5-4b27-4529-a7b7-41fb2b972c45 · inbound

Safety Alignment of LMs via Non-cooperative Games cites this paper.

Safety Alignment of LMs via Non-cooperative Games Bridging Offline and Online Reinforcement Learning for LLMs

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-03T14:24:21.679930Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T14:24:21.679930Z digest=sha256:a00c213981620c97a51125e01bca9de692881f58b85a1c7cca154bd46233e7bd

Observation 089ad325-e523-45ab-a2c4-a1e5469647d6 · inbound

OGER: A Robust Offline-Guided Exploration Reward for Hybrid Reinforcement Learning cites this paper.

OGER: A Robust Offline-Guided Exploration Reward for Hybrid Reinforcement Learning Bridging Offline and Online Reinforcement Learning for LLMs

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-11T11:56:15.425803Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-05-10T04:29:21.897215Z digest=sha256:87562b55c74e62c22ca063759cc22faa6d3026672344295bb1c94ffb885edc8b

Observation 7465628b-7cca-4c32-8f66-7c8684b401f7 · inbound

Smaller Models are Natural Explorers for Policy-Level Diversity in GRPO cites this paper.

Smaller Models are Natural Explorers for Policy-Level Diversity in GRPO Bridging Offline and Online Reinforcement Learning for LLMs

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-06-29T00:02:50.010858Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-28T23:54:32.621093Z digest=sha256:458f516e7f23d941706063ba6923a17de08d1f2d753ff295e63e30f6a9dd4139

Observation b7230ccc-d595-499e-98fc-4d0eb1787025 · inbound

Multi$^2$: Hierarchical Multi-Agent Decision-Making with LLM-Based Agents in Interactive Environments cites this paper.

Multi$^2$: Hierarchical Multi-Agent Decision-Making with LLM-Based Agents in Interactive Environments Bridging Offline and Online Reinforcement Learning for LLMs

Reference 72

Resolution
verified exact
arxiv_id, observed 2026-07-02T01:46:27.027263Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-28T11:29:46.554292Z digest=sha256:f318a5bed19cbfcbba2bf963c4ca459e5eeea42001c84bbf358ac40fc418e78c

Observation 507f7013-5128-4117-8a77-4d64895355cc · inbound

Multi$^2$: Hierarchical Multi-Agent Decision-Making with LLM-Based Agents in Interactive Environments cites this paper.

Multi$^2$: Hierarchical Multi-Agent Decision-Making with LLM-Based Agents in Interactive Environments Bridging Offline and Online Reinforcement Learning for LLMs

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-02T12:30:46.217716Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T12:30:46.217716Z digest=sha256:3296a6ef1968e39b67d9277dbcd01889b9274e039b5204b789c8f8a6bcfe9a1f

Observation cc445246-eaec-4226-b5fc-f4fefe4861de · inbound

Step-by-Step Optimization-like Reasoning in LLMs over Expanding Search Spaces cites this paper.

Step-by-Step Optimization-like Reasoning in LLMs over Expanding Search Spaces Bridging Offline and Online Reinforcement Learning for LLMs

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-07-02T08:46:48.927436Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-28T05:46:26.938277Z digest=sha256:5af1d816da6367237b1b071b8387df898ecd0281cb42c3bf15839b1ccb45b3ff

Observation 7ee3a065-8118-4b92-9708-1f8e7d080591 · inbound

Modularized Reinforcement Learning on LLMs: From MDP Creation to Exploration and Learning cites this paper.

Modularized Reinforcement Learning on LLMs: From MDP Creation to Exploration and Learning Bridging Offline and Online Reinforcement Learning for LLMs

Reference 97

Resolution
verified exact
arxiv_id, observed 2026-07-04T08:09:40.509771Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-26T12:15:08.304150Z digest=sha256:e093a365c0f96bf10ad2623f3da3c0dc1b91f9eeabaab7d3409ef1dc95c8cb02

Observation 7e0c52fc-5ece-4b58-9779-38d28c66da07 · inbound

Autodata: An agentic data scientist to create high quality synthetic data cites this paper.

Autodata: An agentic data scientist to create high quality synthetic data Bridging Offline and Online Reinforcement Learning for LLMs

Reference 40

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T20:40:08.219887Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-06-25T19:50:35.574454Z digest=sha256:e045f54c9da51faae55b7874f3cea2a85ccb1387b5e8c1c8345f8a6dcd6ff810

Observation f177b23e-6f8c-4e22-b30a-ef7b1335dae5 · inbound

Autodata: An agentic data scientist to create high quality synthetic data cites this paper.

Autodata: An agentic data scientist to create high quality synthetic data Bridging Offline and Online Reinforcement Learning for LLMs

Reference 40

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T13:19:51.287491Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=arxiv_source observed=2026-06-26T05:16:12.361470Z digest=sha256:e68a0e3fde71e3ee48ea4fa2305b4dcb8e6400067883236977ed0d3f36843b8e

Observation 9cbe79a4-8756-444d-a1b3-b98c5f4f1484 · inbound

LeVo 2: Stable and Melodious Song Generation via Hierarchical Representation Modeling and Progressive Post-Training cites this paper.

LeVo 2: Stable and Melodious Song Generation via Hierarchical Representation Modeling and Progressive Post-Training Bridging Offline and Online Reinforcement Learning for LLMs

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-07-01T15:15:47.555967Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-30T04:21:38.825926Z digest=sha256:ffd640b998adf494f3e8ac0d598bea7c6aae25522ab6c6fa0d9afcbe670b5273