Pith. sign in

Paper Citation Record · LEDGER

Bridging Offline and Online Reinforcement Learning for LLMs

As of 22 August 2026, this Paper Citation Record lists 55 of 55 outbound references and 12 inbound Pith citation observations for arXiv:2506.21495.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.21495 v1

Coverage vector

measured 55 of 55 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T22:28:10.266576Z

measured 67 of 67 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 12 of 12 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T09:50:46.549563Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T20:40:08.218379Z

Reference resolution

55 of 55 outbound references displayed

  • verified exact1
  • verified fuzzy11
  • unresolved43
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 344a4cba-33ab-4db7-a4f1-3f43136360d3 · outbound

This paper cites write newline.

Bridging Offline and Online Reinforcement Learning for LLMs write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T22:28:04.879326Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:28:04.879326Z digest=sha256:615e2fbd29f4f84bdc15a747c58250422c97c9b03a49c127a8687db4f246dbe5

Observation b4279561-afe0-46a0-89e5-01b5919863f3 · outbound

This paper cites fairseq2, 2023.

Bridging Offline and Online Reinforcement Learning for LLMs fairseq2, 2023

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:28:12.612372Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-06T22:28:04.982424Z digest=sha256:ca76330aec20a206cce5a72ae3df51f34bf8db2aba5fdd510f7af87ec98c9fa1

Observation b781d302-fe36-4b2f-b531-b915a5c70a9d · outbound

This paper cites Preference learning algorithms do not learn preference rankings.

Bridging Offline and Online Reinforcement Learning for LLMs Preference learning algorithms do not learn preference rankings

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:28:12.399050Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-06T22:28:05.039479Z digest=sha256:b900cc48e9f3af2be047f80a7668ff065b37b21cb61f775ce52c90d754b604b0

Observation f37495a5-d296-46c0-b9df-2c80c0b198a9 · outbound

This paper cites Self-Play Fine-Tuning Converts Weak Language Models to Strong Language Models.

Bridging Offline and Online Reinforcement Learning for LLMs Self-Play Fine-Tuning Converts Weak Language Models to Strong Language Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T22:28:05.128670Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:28:05.128670Z digest=sha256:10fecc8a4a3215091a55e2ae94ea551c7f770ea57e6103648c4e604e83115e7c

Observation b7ef05f0-2c6f-407e-8e1f-2280965fa0c9 · outbound

This paper cites Deep reinforcement learning from human preferences.

Bridging Offline and Online Reinforcement Learning for LLMs Deep reinforcement learning from human preferences

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T22:28:05.212164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:28:05.212164Z digest=sha256:3a8e3f37ad07ab8c3e4afbb901ff5ed79e9c022bdcd7de330a5fc88e26611dcb

Observation fdc39dda-ea42-4647-896f-4e89e297eccf · outbound

This paper cites The Llama 3 Herd of Models.

Bridging Offline and Online Reinforcement Learning for LLMs The Llama 3 Herd of Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T22:28:05.308808Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:28:05.308808Z digest=sha256:a833c503d7d8930758abe0ff3a2baa1ce39fff54fb5212f177126bb7a66302db

Observation f133324e-40d8-4c15-997e-ea96aa6c6069 · outbound

This paper cites Length-Controlled AlpacaEval: A Simple Way to Debias Automatic Evaluators.

Bridging Offline and Online Reinforcement Learning for LLMs Length-Controlled AlpacaEval: A Simple Way to Debias Automatic Evaluators

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T22:28:05.412620Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:28:05.412620Z digest=sha256:39258bf57aec1f092ed433fecb6a58d7293fba8869ae31ebf25e965651c8ee40

Observation 2a3fc58e-c16e-4c0d-b117-3ac588379246 · outbound

This paper cites Athene-70b: Redefining the boundaries of post-training for open models, July 2024 a.

Bridging Offline and Online Reinforcement Learning for LLMs Athene-70b: Redefining the boundaries of post-training for open models, July 2024 a

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:28:12.160526Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-06T22:28:05.547248Z digest=sha256:1a73b2759203b06dbbdb82ee534bff049b852c53c19aa5d32ba37422ce5d0967

Observation dafde6d9-7c3b-4dae-ac77-f8f20c1952cf · outbound

This paper cites How to Evaluate Reward Models for RLHF.

Bridging Offline and Online Reinforcement Learning for LLMs How to Evaluate Reward Models for RLHF

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T22:28:05.645832Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:28:05.645832Z digest=sha256:6b182f4ba72ac21b8830ccaea72073d1750d192f487f1ab41bff1814d6a63e7f

Observation cebb86d3-51cf-482d-a85e-7cc9aa5a59fb · outbound

This paper cites Scaling laws for reward model overoptimization.

Bridging Offline and Online Reinforcement Learning for LLMs Scaling laws for reward model overoptimization

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T22:28:05.741024Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:28:05.741024Z digest=sha256:8b70c243814272c90080a6e8bbc3c476555b56c673ce78dc78e45d7790b01b19

Observation 2cc140a8-6a72-4d09-a3b9-bf86a18b7246 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Bridging Offline and Online Reinforcement Learning for LLMs DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T22:28:05.871341Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:28:05.871341Z digest=sha256:9bc4a038acf0427ba9b744b707fc2daf8aeed2091b13b00b5cc2b9c85a297e51

Observation 84f5acc2-ba3b-43e3-946b-d31459bd35a3 · outbound

This paper cites Direct Language Model Alignment from Online AI Feedback.

Bridging Offline and Online Reinforcement Learning for LLMs Direct Language Model Alignment from Online AI Feedback

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T22:28:05.981374Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:28:05.981374Z digest=sha256:630e58ec178be7b8b7604092f35737dd84c40a2df25bd7a63193108813379a2e

Observation 587f2705-7c85-4e90-af71-3c6b3ee6d799 · outbound

This paper cites Measuring Mathematical Problem Solving With the MATH Dataset.

Bridging Offline and Online Reinforcement Learning for LLMs Measuring Mathematical Problem Solving With the MATH Dataset

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T22:28:06.065494Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:28:06.065494Z digest=sha256:0cbda9f97b96c0bfe695fd532266e35c472377b1870cd1a36e5dfb7620dfc43f

Observation 8bc5fe93-ac70-44e6-944e-71c4a0713743 · outbound

This paper cites ORPO : Monolithic preference optimization without reference model.

Bridging Offline and Online Reinforcement Learning for LLMs ORPO : Monolithic preference optimization without reference model

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T22:28:06.143444Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:28:06.143444Z digest=sha256:b9e972b32a1d1db1c691bf2201b1c71038310e6adfca0fd3530b102caf556555

Observation 648e4f9b-ba8a-4ae4-ad7c-fd7242c24883 · outbound

This paper cites Lo RA : Low-rank adaptation of large language models.

Bridging Offline and Online Reinforcement Learning for LLMs Lo RA : Low-rank adaptation of large language models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T22:28:06.218419Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:28:06.218419Z digest=sha256:fc0e659978392c34805a0419c33efaa6e4db78e209853e459b7fc786d7531981

Observation bb5cea72-367c-44e7-ab17-1f8503bf2a32 · outbound

This paper cites Gonzalez, Hao Zhang, and Ion Stoica.

Bridging Offline and Online Reinforcement Learning for LLMs Gonzalez, Hao Zhang, and Ion Stoica

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T22:28:06.322902Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:28:06.322902Z digest=sha256:d0a86ee28a291111dbc34f60dec303b049fca38dce8953f8fcbf0d82ccc7ef75

Observation 5a1178cc-fecf-4e77-8d95-f8d3fd2ee581 · outbound

This paper cites Tulu 3: Pushing Frontiers in Open Language Model Post-Training.

Bridging Offline and Online Reinforcement Learning for LLMs Tulu 3: Pushing Frontiers in Open Language Model Post-Training

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T22:28:06.421583Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:28:06.421583Z digest=sha256:df92774e363f4e611e4b587de208407fffe3fc55a4608693b14f06b7dab2c53e

Observation 53e6dc1b-db68-4490-854b-f79bda263def · outbound

This paper cites Numinamath: The largest public dataset in ai4maths with 860k pairs of competition math problems and solutions.

Bridging Offline and Online Reinforcement Learning for LLMs Numinamath: The largest public dataset in ai4maths with 860k pairs of competition math problems and solutions

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:28:12.015713Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-06T22:28:06.557537Z digest=sha256:fd759ea76b06575735f1e977863d4ea84604d120a542139585e59edeb0c7332b

Observation 9745ab9e-2f20-445d-8ffa-8f388e7736f2 · outbound

This paper cites From live data to high-quality benchmarks: The arena-hard pipeline, April 2024 b.

Bridging Offline and Online Reinforcement Learning for LLMs From live data to high-quality benchmarks: The arena-hard pipeline, April 2024 b

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:28:11.809112Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-06T22:28:06.681214Z digest=sha256:ef2927d33741e15648cc0dfd70c89a962bea6c0fdeefaae1babc6793ddc561be

Observation 54b4cd45-973e-41ea-a2e1-ba3a182c92cd · outbound

This paper cites From Crowdsourced Data to High-Quality Benchmarks: Arena-Hard and BenchBuilder Pipeline.

Bridging Offline and Online Reinforcement Learning for LLMs From Crowdsourced Data to High-Quality Benchmarks: Arena-Hard and BenchBuilder Pipeline

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T22:28:06.760710Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:28:06.760710Z digest=sha256:947a867f7a05a38691f86a20192fb93503b7a6ba3a246465d60f6ae210628449

Observation 73c2a475-43b2-41d8-a0e4-231a0a7e259a · outbound

This paper cites Self-Alignment with Instruction Backtranslation.

Bridging Offline and Online Reinforcement Learning for LLMs Self-Alignment with Instruction Backtranslation

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T22:28:06.859418Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:28:06.859418Z digest=sha256:5e2810e55b32ba7b9e252e89b6aeb9d070440634024c7aa42755f359c102939a

Observation bf53e613-46e1-4248-89ae-94c0df82f4d3 · outbound

This paper cites Hashimoto.

Bridging Offline and Online Reinforcement Learning for LLMs Hashimoto

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:28:11.687633Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-06T22:28:06.946488Z digest=sha256:d089c305ebbf308b393f8fa3417e221049bb84e80789ba2c806f2aa32fe444bd

Observation e8167f2a-9bfc-49f2-aa9c-e62e88efe5ac · outbound

This paper cites Let's Verify Step by Step.

Bridging Offline and Online Reinforcement Learning for LLMs Let's Verify Step by Step

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T22:28:07.062465Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:28:07.062465Z digest=sha256:a698f5d531a030dcbef51884a0c26e341bcad6f66a0ab0ecf3fee122a91e18ae

Observation d68b4f96-d8a1-4c78-a8f5-74ec1ec203de · outbound

This paper cites Statistical Rejection Sampling Improves Preference Optimization.

Bridging Offline and Online Reinforcement Learning for LLMs Statistical Rejection Sampling Improves Preference Optimization

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T22:28:07.163268Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:28:07.163268Z digest=sha256:827941a75e3a2030d32444d69a9af1d4106b757c8718d99921678e56a6bdb072

Observation 627f0544-a507-48b9-8e8d-b5dd31932f45 · outbound

This paper cites Understanding r1-zero-like training: A critical perspective, 2025.

Bridging Offline and Online Reinforcement Learning for LLMs Understanding r1-zero-like training: A critical perspective, 2025

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T22:28:07.276662Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:28:07.276662Z digest=sha256:6dd7f5b27446cdea52b49d01d5d72edd78846e4bae4452d34c5567fc68f9d26e

Observation 118a8369-2f59-4425-b628-9df5572ca0e0 · outbound

This paper cites Mixtral of experts: A high quality sparse mixture-of-experts.

Bridging Offline and Online Reinforcement Learning for LLMs Mixtral of experts: A high quality sparse mixture-of-experts

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:28:11.541879Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-06T22:28:07.363506Z digest=sha256:033447b634d8c0ac0e5e29d3b03d0e838c1542ecf104fb7a9009e072d7740d49

Observation 1af63203-9798-4f1c-948d-469e6525b206 · outbound

This paper cites Ray: A distributed framework for emerging \ AI \ applications.

Bridging Offline and Online Reinforcement Learning for LLMs Ray: A distributed framework for emerging \ AI \ applications

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:28:11.357146Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-06T22:28:07.471713Z digest=sha256:4895c6416213df3c9f413721a22f79e2cb242946e1454bf5119c49a431835bc4

Observation ceaffc31-60bd-42df-a8c3-a4578692c358 · outbound

This paper cites Training language models to follow instructions with human feedback.

Bridging Offline and Online Reinforcement Learning for LLMs Training language models to follow instructions with human feedback

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T22:28:07.577745Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:28:07.577745Z digest=sha256:f554f0904be64f48308e7aa3640f9c0e224f5d4056b2202b3de0bae204719497

Observation d924b7eb-7f25-42fe-836d-3b7a7fe23925 · outbound

This paper cites West-of-N: Synthetic Preferences for Self-Improving Reward Models.

Bridging Offline and Online Reinforcement Learning for LLMs West-of-N: Synthetic Preferences for Self-Improving Reward Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T22:28:07.646010Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:28:07.646010Z digest=sha256:19aed2fa04141c80d9891f1f324f5c04282cae455b305c4b2a608b6a15834059

Observation 6e887785-451e-429b-81b5-183f57a09fcb · outbound

This paper cites Iterative reasoning preference optimization.

Bridging Offline and Online Reinforcement Learning for LLMs Iterative reasoning preference optimization

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T22:28:07.735012Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:28:07.735012Z digest=sha256:816b0567a59267f3ee61ad2adfea536b707cb17e20575c69d867e2fc35689e74

Observation 8524ab45-8871-43d7-9e42-487e27f30440 · outbound

This paper cites Disentangling length from quality in direct preference optimization.

Bridging Offline and Online Reinforcement Learning for LLMs Disentangling length from quality in direct preference optimization

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T22:28:07.868238Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:28:07.868238Z digest=sha256:1a54e40c4d488ae1e318651e3cd438bcf68ccb0f64046e7f7af53dc9e2c5b72a

Observation 561a7e33-73f4-4e69-9cfa-a42e91b4fffd · outbound

This paper cites Disentangling Length from Quality in Direct Preference Optimization.

Bridging Offline and Online Reinforcement Learning for LLMs Disentangling Length from Quality in Direct Preference Optimization

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T22:28:07.964261Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:28:07.964261Z digest=sha256:41c95703668875776dc6a939fdfb698a3ceb72b1a2516fee42db7313376b3b90

Observation 24110d55-6030-444d-b2d1-75b55ad7a8e0 · outbound

This paper cites Online dpo: Online direct preference optimization with fast-slow chasing, 2024.

Bridging Offline and Online Reinforcement Learning for LLMs Online dpo: Online direct preference optimization with fast-slow chasing, 2024

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T22:28:08.036930Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:28:08.036930Z digest=sha256:3b8d0909e547d7c0e4ae9890dc238b612577b904c53007840caa9580cf54dff1

Observation 02239330-82da-4ba0-a3d5-42fa36a5edc1 · outbound

This paper cites Direct preference optimization: Your language model is secretly a reward model.

Bridging Offline and Online Reinforcement Learning for LLMs Direct preference optimization: Your language model is secretly a reward model

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T22:28:08.101830Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:28:08.101830Z digest=sha256:1c1d2a73ba9040e1e65e88203f0ea67ee06d7fe6df76a28effccc285c98ab916

Observation 5b8ade72-fba5-41ec-a837-618165e1337b · outbound

This paper cites Direct preference optimization: Your language model is secretly a reward model.

Bridging Offline and Online Reinforcement Learning for LLMs Direct preference optimization: Your language model is secretly a reward model

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T22:28:08.225651Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:28:08.225651Z digest=sha256:62ec166a09c96c8ccceb346d55b519cd46d6f2163bd98743937bba7afd31b6ea

Observation eb2362c4-a505-453e-9bfd-7166795d23e6 · outbound

This paper cites Trust region policy optimization.

Bridging Offline and Online Reinforcement Learning for LLMs Trust region policy optimization

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T22:28:08.319306Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:28:08.319306Z digest=sha256:fd4118a5d833d066065cd4746675c700351d1d3194566d5b148bc19b7c5b2c02

Observation 3120a6ea-4cc8-4632-a8fa-ac9bc2109dc1 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Bridging Offline and Online Reinforcement Learning for LLMs Proximal Policy Optimization Algorithms

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T22:28:08.531883Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:28:08.531883Z digest=sha256:8e6a3c56098e1f49f7ec2d098ca766b30a847223ce4c1bc292b8b2b7d5bfe22b

Observation 54604467-07e4-4c24-881d-947e81676d91 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Bridging Offline and Online Reinforcement Learning for LLMs DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T22:28:08.661390Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:28:08.661390Z digest=sha256:14f8dce2b49edfbb4ac0fbd859cd47308af715adef266bab522d74cbc17765f8

Observation a80e91d1-6705-43ff-8fc0-1e318c2c7c9a · outbound

This paper cites Welcome to the era of experience.

Bridging Offline and Online Reinforcement Learning for LLMs Welcome to the era of experience

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:28:11.153605Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-06T22:28:08.768169Z digest=sha256:4332fed96ca184e1d69d5a022edb9d74f2e9ed2e586ecd7d758cda7a7daae384

Observation 591e2016-2b65-4747-896a-17aba7f47326 · outbound

This paper cites A long way to go: Investigating length correlations in RLHF.

Bridging Offline and Online Reinforcement Learning for LLMs A long way to go: Investigating length correlations in RLHF

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:28:11.008634Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-06T22:28:08.855939Z digest=sha256:5c451cf9c6110ac18a3a6c362dbc519c788242dfd37b2c600318aa27a23ed2a7

Observation 346d1015-0e56-44ca-b54d-f209525b007f · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Bridging Offline and Online Reinforcement Learning for LLMs LLaMA: Open and Efficient Foundation Language Models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T22:28:08.959424Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:28:08.959424Z digest=sha256:b3b10c38e9f6922b601d50ff09dda78f97d96105a027b2e29d90fb7a68574091

Observation ccc5ca8f-c8c3-425d-becb-b12fdd54ce2e · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Bridging Offline and Online Reinforcement Learning for LLMs Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T22:28:09.048795Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:28:09.048795Z digest=sha256:fd959a679077f6be1906278bec743d59497a0c56a20bcb2bcea68d445b419dab

Observation 0d1cd526-46cd-488e-87b3-c0ba5b3fd08a · outbound

This paper cites Zephyr: Direct Distillation of LM Alignment.

Bridging Offline and Online Reinforcement Learning for LLMs Zephyr: Direct Distillation of LM Alignment

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T22:28:09.134651Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:28:09.134651Z digest=sha256:f2efddd3a3b872283498e35ca2445ece2ed5678198d61a046546207c30b7e252

Observation ce183029-a436-4bfc-a777-5425946aa887 · outbound

This paper cites Thinking LLMs: General Instruction Following with Thought Generation.

Bridging Offline and Online Reinforcement Learning for LLMs Thinking LLMs: General Instruction Following with Thought Generation

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T22:28:09.217475Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:28:09.217475Z digest=sha256:e8c137ce03d3544ad6f6777dbf9f1a9a4955f506aae52e3cdc1c3ba38d81eec1

Observation baeaa4fb-5bd9-4b9f-84da-61c5d5ffab84 · outbound

This paper cites Meta-Rewarding Language Models: Self-Improving Alignment with LLM-as-a-Meta-Judge.

Bridging Offline and Online Reinforcement Learning for LLMs Meta-Rewarding Language Models: Self-Improving Alignment with LLM-as-a-Meta-Judge

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T22:28:09.310079Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:28:09.310079Z digest=sha256:f307b85aaa97166b608d498b059db224f92a3b5c4cc250c63da38102333981ca

Observation aadb1a06-6ba7-44f1-8ad0-d2b140535341 · outbound

This paper cites Iterative Preference Learning from Human Feedback: Bridging Theory and Practice for RLHF under KL-Constraint.

Bridging Offline and Online Reinforcement Learning for LLMs Iterative Preference Learning from Human Feedback: Bridging Theory and Practice for RLHF under KL-Constraint

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-06T22:28:09.397544Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:28:09.397544Z digest=sha256:1288a0653b813a5298fcab2ec33b25cc137a7af1d2e6d62fccb924e66abf8b73

Observation a00e9344-8a24-475a-b719-bad7539cd25c · outbound

This paper cites Gibbs sampling from human feedback: A provable kl-constrained framework for rlhf.

Bridging Offline and Online Reinforcement Learning for LLMs Gibbs sampling from human feedback: A provable kl-constrained framework for rlhf

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T22:28:10.850501Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-06T22:28:09.483479Z digest=sha256:5f3b9d05092a7b2b2c7af04c7bea3f95494741bd6661cfc0ad2cf99592731ed9

Observation 4d8e1c05-82c5-420f-9591-4dd36d84a8f6 · outbound

This paper cites WizardLM: Empowering large pre-trained language models to follow complex instructions.

Bridging Offline and Online Reinforcement Learning for LLMs WizardLM: Empowering large pre-trained language models to follow complex instructions

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-06T22:28:09.578482Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:28:09.578482Z digest=sha256:76ee0ec2ebd5d3308e1f3cda372adbf65b3b853941080fe5dbf8e619c336d8ca

Observation ab9dd718-b78d-491d-9789-0004f8b42a36 · outbound

This paper cites Some things are more CRINGE than others: Iterative Preference Optimization with the Pairwise Cringe Loss.

Bridging Offline and Online Reinforcement Learning for LLMs Some things are more CRINGE than others: Iterative Preference Optimization with the Pairwise Cringe Loss

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T22:28:09.691354Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:28:09.691354Z digest=sha256:2092821c7e813016d568e7d6ba29a73150a512ccd1071eda925ca77db5695798

Observation 6a88cb36-aac5-4a67-8291-9c4fc629a398 · outbound

This paper cites Is DPO Superior to PPO for LLM Alignment? A Comprehensive Study.

Bridging Offline and Online Reinforcement Learning for LLMs Is DPO Superior to PPO for LLM Alignment? A Comprehensive Study

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-06T22:28:09.790974Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:28:09.790974Z digest=sha256:6230f8d035cd19dd84d7a08f95ace75cc97124e6a775d42938735dda054bc7aa

Observation 757a19e2-111d-4b8c-b2d4-47fbf950d92b · outbound

This paper cites BPO : Staying close to the behavior LLM creates better online LLM alignment.

Bridging Offline and Online Reinforcement Learning for LLMs BPO : Staying close to the behavior LLM creates better online LLM alignment

Reference 52

Resolution
verified exact
doi, observed 2026-08-06T22:28:10.442692Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-06T22:28:09.869519Z digest=sha256:5a902cd67992eb00583383baf65750baec159ae36ca87a77d981732975704930

Observation 41c10192-46cd-49c6-aadf-e0e4536b6240 · outbound

This paper cites Self-Rewarding Language Models.

Bridging Offline and Online Reinforcement Learning for LLMs Self-Rewarding Language Models

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-06T22:28:09.980703Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:28:09.980703Z digest=sha256:63c0ed124be68092487110e2268755d167c9cfb0393167dcfca6e5af41134de3

Observation e89d6539-c0bf-45ad-b512-df53c68879fa · outbound

This paper cites WildChat: 1M ChatGPT Interaction Logs in the Wild.

Bridging Offline and Online Reinforcement Learning for LLMs WildChat: 1M ChatGPT Interaction Logs in the Wild

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-06T22:28:10.083842Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:28:10.083842Z digest=sha256:ddbd63dc0925e3ae62466954787452ab7f8e799411368bd0f44e388f492fb0ed

Observation 23af5f03-b49c-4648-8428-f8d344d54282 · outbound

This paper cites Lima: Less is more for alignment.

Bridging Offline and Online Reinforcement Learning for LLMs Lima: Less is more for alignment

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-06T22:28:10.181387Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:28:10.181387Z digest=sha256:24a2b4fb592b695bf2aa4b8eb971fb114d7f2b40e238074d04eaaef21f73bbac

Observation c4852346-3bd9-40a7-a77f-56969f60f120 · outbound

This paper cites Fine-Tuning Language Models from Human Preferences.

Bridging Offline and Online Reinforcement Learning for LLMs Fine-Tuning Language Models from Human Preferences

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-06T22:28:10.266576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T22:28:10.266576Z digest=sha256:eca37eee785232bd4dbd6d46b2d16ecf24be2694a8b30a0d588a8d75212005b3

Pith citing papers

Observation 3ab7ef80-1dbb-46a4-821c-b519340abb04 · inbound

Task-Dependent Evaluation of LLM Output Homogenization: A Taxonomy-Guided Framework cites this paper.

Task-Dependent Evaluation of LLM Output Homogenization: A Taxonomy-Guided Framework Bridging Offline and Online Reinforcement Learning for LLMs

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-18T13:36:24.983666Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-18T13:34:22.790199Z digest=sha256:068fe12623a6638e267cb7b191f39fc3e991b0d954e476ab8f5512b7fc0c0b8b

Observation d3c3b757-9212-422a-b8f0-97949b079396 · inbound

Beyond Correctness: Rewarding Faithful Reasoning in Retrieval-Augmented Generation cites this paper.

Beyond Correctness: Rewarding Faithful Reasoning in Retrieval-Augmented Generation Bridging Offline and Online Reinforcement Learning for LLMs

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-04T09:50:46.549563Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T09:50:46.549563Z digest=sha256:ea111f92f56d7df7d82407f2bd4e758133d62ab001f58f0c6b7b080aa9f12db4

Observation 202772d5-4b27-4529-a7b7-41fb2b972c45 · inbound

Safety Alignment of LMs via Non-cooperative Games cites this paper.

Safety Alignment of LMs via Non-cooperative Games Bridging Offline and Online Reinforcement Learning for LLMs

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-03T14:24:21.679930Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T14:24:21.679930Z digest=sha256:a00c213981620c97a51125e01bca9de692881f58b85a1c7cca154bd46233e7bd

Observation 089ad325-e523-45ab-a2c4-a1e5469647d6 · inbound

OGER: A Robust Offline-Guided Exploration Reward for Hybrid Reinforcement Learning cites this paper.

OGER: A Robust Offline-Guided Exploration Reward for Hybrid Reinforcement Learning Bridging Offline and Online Reinforcement Learning for LLMs

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-11T11:56:15.425803Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-10T04:29:21.897215Z digest=sha256:3cfe8ac4d1f1e87529a0229f31f2a3d70ff62c1ee47be28c3428ee61698bc9ff

Observation 7465628b-7cca-4c32-8f66-7c8684b401f7 · inbound

Smaller Models are Natural Explorers for Policy-Level Diversity in GRPO cites this paper.

Smaller Models are Natural Explorers for Policy-Level Diversity in GRPO Bridging Offline and Online Reinforcement Learning for LLMs

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-06-29T00:02:50.010858Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-28T23:54:32.621093Z digest=sha256:c64caa3fef33c3e1aa457f287340f0c6a711693d0611cff462e06eb36fea531d

Observation b7230ccc-d595-499e-98fc-4d0eb1787025 · inbound

Multi$^2$: Hierarchical Multi-Agent Decision-Making with LLM-Based Agents in Interactive Environments cites this paper.

Multi$^2$: Hierarchical Multi-Agent Decision-Making with LLM-Based Agents in Interactive Environments Bridging Offline and Online Reinforcement Learning for LLMs

Reference 72

Resolution
verified exact
arxiv_id, observed 2026-07-02T01:46:27.027263Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-28T11:29:46.554292Z digest=sha256:82fd8dddd9d2e141007f952e484d10d7003da8559bb0f193e0cf737566a2b9fa

Observation 507f7013-5128-4117-8a77-4d64895355cc · inbound

Multi$^2$: Hierarchical Multi-Agent Decision-Making with LLM-Based Agents in Interactive Environments cites this paper.

Multi$^2$: Hierarchical Multi-Agent Decision-Making with LLM-Based Agents in Interactive Environments Bridging Offline and Online Reinforcement Learning for LLMs

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-02T12:30:46.217716Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T12:30:46.217716Z digest=sha256:3296a6ef1968e39b67d9277dbcd01889b9274e039b5204b789c8f8a6bcfe9a1f

Observation cc445246-eaec-4226-b5fc-f4fefe4861de · inbound

Step-by-Step Optimization-like Reasoning in LLMs over Expanding Search Spaces cites this paper.

Step-by-Step Optimization-like Reasoning in LLMs over Expanding Search Spaces Bridging Offline and Online Reinforcement Learning for LLMs

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-07-02T08:46:48.927436Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-28T05:46:26.938277Z digest=sha256:849671c822f3eb8db1df9e3a799afb1dceaedbc4ca891a858beb95eb8cc05d07

Observation 7ee3a065-8118-4b92-9708-1f8e7d080591 · inbound

Modularized Reinforcement Learning on LLMs: From MDP Creation to Exploration and Learning cites this paper.

Modularized Reinforcement Learning on LLMs: From MDP Creation to Exploration and Learning Bridging Offline and Online Reinforcement Learning for LLMs

Reference 97

Resolution
verified exact
arxiv_id, observed 2026-07-04T08:09:40.509771Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-26T12:15:08.304150Z digest=sha256:d036f963197d3c5e23c71a6ac896097ff50a31940e34312d78f29fb1ce237c63

Observation 7e0c52fc-5ece-4b58-9779-38d28c66da07 · inbound

Autodata: An agentic data scientist to create high quality synthetic data cites this paper.

Autodata: An agentic data scientist to create high quality synthetic data Bridging Offline and Online Reinforcement Learning for LLMs

Reference 40

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T20:40:08.219887Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-06-25T19:50:35.574454Z digest=sha256:8c21e9e729a1a854a567e7fae2f82e5bc14a3143e610731c8259cc6b92d3dc10

Observation f177b23e-6f8c-4e22-b30a-ef7b1335dae5 · inbound

Autodata: An agentic data scientist to create high quality synthetic data cites this paper.

Autodata: An agentic data scientist to create high quality synthetic data Bridging Offline and Online Reinforcement Learning for LLMs

Reference 40

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T13:19:51.287491Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-06-26T05:16:12.361470Z digest=sha256:600f0f3cb97d287e43a62966f6cad8813e41f2f72efff3c7795f674a75b31028

Observation 9cbe79a4-8756-444d-a1b3-b98c5f4f1484 · inbound

LeVo 2: Stable and Melodious Song Generation via Hierarchical Representation Modeling and Progressive Post-Training cites this paper.

LeVo 2: Stable and Melodious Song Generation via Hierarchical Representation Modeling and Progressive Post-Training Bridging Offline and Online Reinforcement Learning for LLMs

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-07-01T15:15:47.555967Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-30T04:21:38.825926Z digest=sha256:8f631223c7754c50b66768188827bcea51a19ac3f5c3112ce2da4845a88ede02